DETAILED ACTION
1. This action is responsive to Application no.19/005,475 filed 12/30/2024. All claims have been examined and are currently pending.
The claim limitations (of claim 1, and other corresponding independent claims) recite a series of steps and is a process. The claim does not recite any of the judicial exceptions (mathematical concepts, mental processes, certain methods of organizing human activity) as it recites a sequence generation model, which incorporates a trained dataset to process a received document to generate and output recommended key phrases, And therefore recites patent eligible subject matter.
Notice of Pre-AIA or AIA Status
2. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
3. The information disclosure statement (IDS) submitted is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 103
4. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
5. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
6. Claims 1-4, 6-12, 14-20 are rejected under 35 U.S.C. 103 as being unpatentable over Cheng et al (2022/0374600) “KEYPHRASE GENERATION FOR TEXT SEARCH WITH OPTIMAL INDEXING REGULARIZATION VIA REINFORCEMENT LEARNING” in view of Samuelson et al (2018/0211552).
Regarding claim 1 Cheng et al (2022/0374600) teaches A method implemented by at least one computing device (abstract: a computer-implemented method; fig 1 computing device; 0004: computer-implemented method), the method comprising:
receiving an input document (figure 8; 0047: the input to the KG algorithm is a document; 0050 Seq2Seq model receives a document d; 0053);
generating, using a sequence generation model, a sequence of key phrases based on the input document, {the sequence generation model omitting use of a self-generated sequence termination token during generation of the sequence} (figure 8; paragraph 4: predicting, by the processor device, new keyphrases using the trained policy neural network; 0047; 0049-0050: The Seq2Seq model f.sub.θ 810 receives a document d, and outputs a sequence y responsive to the document d.; 0053 transformers; 0058 keyphrase generation; transformer); and
outputting, as recommended key phrases for the input document, the sequence of key phrases (abstract: predicting …new keyphrases using the trained policy neural network; 0004; 0047: the output of the KG algorithm is a set of keyphrases; 0050; 0070 output keyphrases);
but does not specifically teach the sequence generation model omitting use of a self-generated sequence termination token during generation of the sequence.
In a similar field of endeavor, Samuelson teaches identification and adding of phrases, and where additional phrases and/or sentences can be included … until the threshold number of characters is reached (0082). It would have been obvious to one of ordinary skill in the art before the effective filing date to incorporate Samuelson and a limit for establishing a termination point in key phrase generation.
Cheng already teaches EOS is the end generation (53); [EOS] to stop the generation (55), but does not specifically teach when EOS is to be implemented. It would have been obvious to thus look to Samuelson to further incorporate a specific limit, allowing for the generating to conclude once a specific condition is met (a certain number of characters/phrases).
Regarding claim 2 Cheng teaches The method of claim 1, further comprising:
receiving a training dataset that includes a plurality of training samples, wherein each training sample includes a training document paired with one or more positive key phrase samples (abstract: the method includes pretraining, by a processor device, a policy neural network on training documents using sequence-to-sequence model. The training documents are each associated with a list of keyphrases; 0004: The method includes pretraining, by a processor device, a policy neural network on training documents using a sequence-to-sequence model. The training documents are each associated with a list of keyphrases included therein.; 0061); and
training the sequence generation model using the training dataset (abstract: the method includes pretraining, by a processor device, a policy neural network on training documents using sequence-to-sequence model.; 0004; 0061).
Regarding claim 3 Cheng teaches The method of claim 2, further comprising pairing the training document with a positive key phrase sample in the training dataset based on historical engagement with the training document in response to the positive key phrase sample being searched via a search platform (figure 2; 0004: The method includes pretraining, by a processor device, a policy neural network on training documents using a sequence-to-sequence model. The training documents are each associated with a list of keyphrases included therein; 0034-35: The system 210 receives a set of documents 220, and a query 230 directed to the set of documents 220, and returns a document(s) and/or term(s) 240 in the set of documents 220 found responsive to the query 230; fig 3; 0037-0041; 0061).
Regarding claim 4 Cheng teaches The method of claim 2, wherein training the sequence generation model further comprises:
generating, using the sequence generation model, an additional sequence of training key phrases based on the training document of a training sample, {the sequence generation model omitting use of the self-generated sequence termination token during generation of the additional sequence} (abstract: the method includes pretraining, by a processor device, a policy neural network on training documents using sequence-to-sequence model. The training documents are each associated with a list of keyphrases; 0004: The method includes pretraining, by a processor device, a policy neural network on training documents using a sequence-to-sequence model. The training documents are each associated with a list of keyphrases included therein.; 0061); and
training the sequence generation model based on a comparison of the training key phrases to the one or more positive key phrase samples of the training sample (abstract: the method includes pretraining, by a processor device, a policy neural network on training documents using sequence-to-sequence model. The training documents are each associated with a list of keyphrases; 0004: The method includes pretraining, by a processor device, a policy neural network on training documents using a sequence-to-sequence model. The training documents are each associated with a list of keyphrases included therein. The method further includes training, by the processor device, the policy neural network using reinforcement learning with a summarization reward on present annotated keyphrases in an input training document and absent annotated keyphrase from the input training document that semantically describe a concept of the input training document.; 0061);
but does not specifically teach the sequence generation model omitting use of the self-generated sequence termination token during generation of the additional sequence.
Rejected for similar rationale and reasoning as claim 1
Regarding claim 6 Cheng does not specifically teach where Samuelson teaches The method of claim 1, wherein generating the sequence of key phrases further comprises terminating the generation of the sequence of key phrases after a predefined number of key phrases have been generated by the sequence generation model (0082: In these implementations, for example, if a number of characters included between the start point 226a and the end point 226b does not exceed the threshold number of characters, additional phrases and/or sentences can be included in the note until the threshold number of characters is reached (e.g., the third phrase 254).).
Rejected for similar rationale and reasoning as claim 1
Regarding claim 7 Cheng teaches The method of claim 1, wherein generating a key phrase of the sequence of key phrases further comprises:
generating a start token that marks a start of the key phrase (figure 10; 55: present keyphrases); and
generating an end token that marks an end of the key phrase, wherein the start token and the end token delineate the key phrase from other key phrases of the sequence (fig 10; 0053: a special token that indicates the end of present keyphrases; 0055; 0059 end of one phrase – incorporates boundaries to differentiate between individual keyphrases and then also between present keyphrase and absent keyphrase groups).
Regarding claim 8 Cheng teaches The method of claim 7, wherein generating the key phrase further comprises inserting the end token of the key phrase (53; 55; 59);
But does not specifically teach where Samuelson teaches in response to generating a threshold number of content tokens for the key phrase (82).
Rejected for similar rationale and reasoning as claim 1
Regarding claim 9 Cheng teaches The method of claim 1, wherein the sequence generation model is a transformer-based natural language processing model (53; 58 transformer).
Regarding claim 10 Cheng teaches A system (0006: computer processing system…includes a memory, processor) comprising:
one or more processors (0006); and
memory storing instructions that, when executed by the one or more processors, cause the system (0006) to:
receive an input document;
generate, using a sequence generation model, a sequence of key phrases based on the input document, {the sequence generation model omitting use of a self-generated sequence termination token during generation of the sequence}; and
output, as recommended key phrases for the input document, the sequence of key phrases;
but does not specifically teach where Samuelson teaches the sequence generation model omitting use of a self-generated sequence termination token during generation of the sequence.
Recites limitations similar to claim 1 and is rejected for similar rationale and reasoning
Claim 11 recites limitations similar to claim 2 and is rejected for similar rationale and reasoning
Claim 12 recites limitations similar to claim 4 and is rejected for similar rationale and reasoning
Claim 14 recites limitations similar to claim 6 and is rejected for similar rationale and reasoning
Claim 15 recites limitations similar to claim 7 and is rejected for similar rationale and reasoning
Claim 16 recites limitations similar to claim 8 and is rejected for similar rationale and reasoning
Regarding claim 17 Cheng teaches A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations (0005: non-transitory computer readable storage medium) comprising:
receiving an input document;
generating a sequence of key phrases based on the input document using a sequence generation model {having been trained to generate the key phrases indefinitely;
terminating generation of the sequence in response to a threshold number of key phrases having been generated;} and
outputting the sequence of key phrases;
but does not specifically teach where Samuelson teaches
terminating generation of the sequence in response to a threshold number of key phrases having been generated.
Rejected for similar rationale and reasoning as claim 1/6 above.
Regarding claim 18 Cheng does not specifically teach where Samuelson teaches The non-transitory computer-readable storage medium of claim 17, wherein generating the sequence of key phrases further comprises omitting use of a self- generated sequence termination token during generation of the sequence.
Rejected for similar rationale and reasoning as claim 1 above.
Claim 19 Recites limitations similar to claim 7 and is rejected for similar rationale and reasoning
Claim 20 Recites limitations similar to claim 8 and is rejected for similar rationale and reasoning
7. Claims 5, 13 are rejected under 35 U.S.C. 103 as being unpatentable over Cheng et al (2022/0374600) in view of Samuelson in further view of Xiong et al (2021/0004416).
Regarding claim 5 Cheng (and Samuelson) does not specifically teach where Xiong teaches The method of claim 2, further comprising:
selecting, from the plurality of training samples, first training samples that include at least a threshold number of the positive key phrase samples (0006: extract one or more key phrase candidates);
generating, using the trained sequence generation model, additional key phrases for a subset of training samples of the plurality of training samples (0006 one or more key phrase candidates; incorporates machine learning);
selecting, from the subset of training samples, second training samples that include at least a threshold number of unique key phrases from the positive key phrase samples and the additional key phrases (0006: The key phrase model incorporates machine learning to filter the feature vectors to remove duplicates using a model or classifier, such as a binary classifier, that was trained on a set of key phrase pairs with manual labels indicating whether two key phrases are duplicates of each other, to produce remaining key phrase candidates. The system can use the remaining key phrase candidates in a computer-implemented application.; 32-35); and
re-training the sequence generation model using an augmented dataset that includes the first training samples and the second training samples ([0035] In more detail, the binary classifier analyzes the feature vectors, which as explained above are extracted key-phrase candidates that have been converted into feature vectors (and that preferably have undergone deduplication). The binary classifier determines which feature vectors pass and which are filtered out based on the features and based on machine learning to learn how to filter out feature vectors that are duplicates of each other. The manual labels are used to train the binary classifier. Key-phrase pairs are provided to the binary classifier with labels showing examples of key-phrase pairs that are duplicates of each other. One way to do this is by using the LambdaMart algorithm. LambdaMart is a machine learning algorithm for this prediction task. After the binary classifier using the LambdaMart algorithm is trained on these labels, the binary classifier can predict whether a pair of feature vectors are duplicates.).
It would have been obvious to one of ordinary skill in the art before the effective filing date to incorporate Xiong for an improved system, to remove duplicates ensuring key phrases are unique and adequately represent the content.
Cheng already teaches training a sequence generation model to generate key phrases. It would have been obvious to look to Xiong to further analyze and filter the key phrases to remove duplicates and to encourage generating both distinctive and informative keyphrases (Cheng 68).
Claim 13 recites limitations similar to claim 5 and is rejected for similar rationale and reasoning
Conclusion
8. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: See PTO-892.
Zhang – “Keyphrase Generation Based om Deep Seq2Seq Model”
Abstract: Keyphrase can provide highly summative information which can help us improve information utilization efficiency in the era of information overload. Though previous researches about keyphrase generation have provided some workable solutions, they generate keyphrase by ranking and selecting meaningful words from the source text. These approaches belong to an extractive method, by which they cannot effectively use semantic meaning of the source text, and are unable to generate keyphrases which do not appear in the source text. So we propose a sequence-to-sequence framework with attention mechanism, copy mechanism,and coverage mechanism,which can effectively deal with the above-mentioned drawbacks.
-R. Devika, S. Vairavasundaram, C. S. J. Mahenthar, V. Varadarajan and K. Kotecha, "A Deep Learning Model Based on BERT and Sentence Transformer for Semantic Keyphrase Extraction on Big Social Data," in IEEE Access, vol. 9, pp. 165252-165261, 2021, doi: 10.1109/ACCESS.2021.3133651.
Abstract: This work aims to extract the keyphrase from Big social data using a sentence transformer with Bidirectional Encoder Representation Transformers (BERT) deep learning model. This BERT representation retains semantic and syntactic connectivity between tweets, enhancing performance in every NLP task on large data sets. It can automatically extract the most typical phrases in the Tweets. The proposed Semkey-BERT model shows that BERT with sentence transformer accuracy of 86% is higher than the other existing models.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHAUN A ROBERTS whose telephone number is (571)270-7541. The examiner can normally be reached Monday-Friday 9-5 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached on 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov.
For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHAUN ROBERTS/Primary Examiner, Art Unit 2655