DETAILED ACTION
This is in response to the reply filed on 05/06/2026. Claims 1, 3-10, 12-16, and 21 are pending in this Action. Claims 2 and 17-20 had been previously canceled.
Remark
In response filed 05/06/2026, claims 1 and 21 have been amended, claim 11 has been cancelled, and no new claim has been added.
Note that the limitation of “storing the deduplicated dataset” as originally filed (see claimed filed 01/22/2025) is missing in claim 1. The record does not show that the applicant had deleted said limitation. In reply to this Office action, Applicant must indicate whether said limitation was inadvertently omitted from the claim, or delete it by striking a line through the limitation. The correction is required.
Response to Arguments
Applicant's arguments filed 05/06/2026 have been fully considered but they are not persuasive.
With respect to 35 USC 101 rejections:
The applicant in page 7 of the Remark alleges that:
It is respectfully submitted that claim 1 recites a number of features that cannot
practically be performed in the human mind, including for example: (a) normalizing data in metadata fields of a same metadata type; (b) harmonizing strings by removing special characters (claim 5); (c) using natural language processing (claim 7); (d) using a string metric algorithm (claim 9); and (e) using a Levenshtein distance algorithm. For example, a human does not think in terms of metadata types, harmonizing strings, using NLP, or using specific algorithms string metric or Levenshtein. Rather, these elements are utilized in a technical manner. It is respectfully submitted that the Examiner fails to provide any support for the conclusion, for example, that these features encompass a mental process that is practically performed in the human mind.
The Examiner respectfully disagrees.
The Examiner holds that above-mentioned features are directed to one of the enumerated groupings of abstract idea. The enumerated groupings of abstract ideas are defined as: 1) Mathematical concepts – mathematical relationships, mathematical formulas or equations, mathematical calculations (see MPEP § 2106.04(a)(2), subsection I); 2) Certain methods of organizing human activity – fundamental economic principles or practices (including hedging, insurance, mitigating risk); commercial or legal interactions (including agreements in the form of contracts; legal obligations; advertising, marketing or sales activities or behaviors; business relations); managing personal behavior or relationships or interactions between people (including social activities, teaching, and following rules or instructions) (see MPEP § 2106.04(a)(2), subsection II); and 3) Mental processes – concepts performed in the human mind (including an observation, evaluation, judgment, opinion) (see MPEP § 2106.04(a)(2), subsection III).
Here, the features of “(a) normalizing data in metadata fields of a same metadata type; (b) harmonizing strings by removing special characters (claim 5); (c) using natural language processing (claim 7); (d) using a string metric algorithm (claim 9); and (e) using a Levenshtein distance algorithm” are recited at high-level of generality. Giving the broadest and reasonable interpretation (BRI) to the claimed feature, they constitute mathematical concept and/or concepts performed in the human mind (including an observation, evaluation, judgment, opinion). Under BRI, the functions of “normalizing data” (which organizing and structuring data) removing special characters using multiple algorithms (which is a set of instruction or steps to solve a problem or perform computations) are operations that involves mathematical concept and/or observation, evaluation, and judgment that could be practically performed in the human mind. Although, these features could be executed by a computer, nothing preclude these features from being performed in the human mind. As such, contrary to the applicant’s allegation, the above-mentioned features recite abstract idea.
Furthermore, merely adding the feature of “executed by one or more processors of a computing system” does not make a step or feature non-abstract. It is simply an addition of a general purpose computer to an abstract idea. Merely adding a generic computer, generic computer components, or a programmed computer to perform generic computer functions does not automatically overcome an eligibility rejection. Alice Corp. Pty. Ltd. v. CLS Bank Int’l, 573 U.S. 208, 223-24, 110 USPQ2d 1976, 1983-84 (2014). See In re Alappat, 33 F.3d 1526, 1545, 31 USPQ2d 1545, 1558 (Fed. Cir. 1994); In re Bilski, 545 F.3d 943, 88 USPQ2d 1385 (Fed. Cir. 2008). It is important to note that a general purpose computer that applies a judicial exception, such as an abstract idea, by use of conventional computer functions does not qualify as a particular machine. Ultramercial, Inc. v. Hulu, LLC, 772 F.3d 709, 716-17, 112 USPQ2d 1750, 1755-56 (Fed. Cir. 2014).
The applicant further argues under Step 2A, Prong 1 analysis in pages 9-12 of the Remark that:
First, the computer-implemented method is now expressly stated to be "executed by one or more processors of a computing system," with each individual method step tied to execution "by the one or more processors" and the outputting step producing results 'from the one or more processors." This processor recitation is supported by the specification as filed in, for example, paragraphs [0019] and [0028]. For instance, the specification states: "The steps of the method can be executed or processed by a processor or an equivalent means in the order set out before" (emphasis added), and "A processor of the digital device can be configured to execute the steps as set out herein before and can be provided to process the input dataset accordingly" (emphasis added).
Second, the 'faster than..." limitation previously added in the last response filed on 08/07/2025 has been replaced with a clearer functional architecture language that describes the structural purpose of the two-stage claimed method: the similarity scoring and clustering steps constitute a computationally efficient first stage (fast) executed by the one or more processors that limits the application, within a second stage, of the filtering rules (slow) to citation pairs within the clusters, rather than to all possible pairwise combinations of citations in the input dataset. This is supported by the specification in [0044], [0047].
Third, the filtering rules step now expressly specifies the comparison mechanism:
"wherein each filtering rule verifies whether normalized data fields of the same metadata type of two citations in one cluster are equal or concurring, and each filtering rule is applied separately to more than one metadata type.", as previously defined in dependent claim 11.
This is a defined computational architecture directed to a specific technical result (time and resource efficient processing of entries from multiple bibliographic databases), not a result or abstract idea expressed in technical language.
The first and most direct basis on which Prong 1 is satisfied is the processor recitation distributed throughout amended claim 1.
The Examiner respectfully disagrees.
As explained above and in Step 2A, Prong 1 analysis (see below), the claimed features of “normalizing…data…”, “calculating…a similarity score…”, and “applying…filtering rules…” as featured in claim 1 are recited at high-level of generality. Giving the broadest and reasonable interpretation (BRI) to the claimed feature, they constitute mathematical concept and/or concepts of observation, evaluation, and/or judgment which could practically be performed in the human mind. Under BRI, the functions of “normalizing data” (which organizing and structuring data involves concepts of observation, evaluation, and/or judgment), calculating a similarity score (involves mathematical calculations), and applying filters (which also involves concepts of observation, evaluation, and/or judgment) are functions that recite a mental process and/or a mathematical concept which is one of the enumerated groups of abstract idea.
Furthermore, merely adding the feature of “executed by one or more processors
of a computing system” does not make a step or feature non-abstract. It is simply the addition of a general purpose computer to an abstract idea. Merely adding a generic computer, generic computer components, or a programmed computer to perform generic computer functions does not automatically overcome an eligibility rejection. Alice Corp. Pty. Ltd. v. CLS Bank Int’l, 573 U.S. 208, 223-24, 110 USPQ2d 1976, 1983-84 (2014). See In re Alappat, 33 F.3d 1526, 1545, 31 USPQ2d 1545, 1558 (Fed. Cir. 1994); In re Bilski, 545 F.3d 943, 88 USPQ2d 1385 (Fed. Cir. 2008). It is important to note that a general purpose computer that applies a judicial exception, such as an abstract idea, by use of conventional computer functions does not qualify as a particular machine. Ultramercial, Inc. v. Hulu, LLC, 772 F.3d 709, 716-17, 112 USPQ2d 1750, 1755-56 (Fed. Cir. 2014).
Therefore, based on above reasoning and explanation, the above mentioned functions recite abstract idea and the amended claim 1 is directed to a judicial exception.
Moreover, the applicant argues under Step 2A, Prong 2 analysis in pages 12-13 of the Remark that:
The applicant submits that amended claim 1 is integrated into a practical application since it is directed to a specific improvement in how computing systems process bibliographic citation data, which is a recognized basis for patent eligibility firmly established by Enfish, LLC v. Microsoft Corp., which held that "software can make non-abstract improvements to computer technology just as hardware improvements can," and identified the relevant inquiry as whether "the claims are directed to an improvement to computer functionality versus being directed to an abstract idea.".
The specific two-stage pair-reduction architecture combined with sequential priority- ordered field-type comparison rules is an improvement in the computational methodology for bibliographic data quality management, not a computer used as a tool to perform an otherwise abstract task. The specification confirms this characterization: "the process and the related computer-implemented method can significantly reduce the time spent on deduplication data objects containing bibliographic data obtained from different sources. The process is simple and efficient and can be implemented on a digital device resource-efficiently without fear of compromising output quality or purity" ([0118]).
The Examiner respectfully disagrees that the claims are not directed to an abstract idea based on Enfish. Applicant' s reliance on Enfish for the eligibility of the claims of the instant application have been fully considered but are likewise not deemed to be persuasive. In order to rely on Enfish, improvement in the functionality of the computer itself must be proved. The specification needs to distinguish between the invention and conventional solutions. The specification must identify improvement in the functionality of the computer itself, tying the particular features of the claims to that improvement and how the claims are directed to improvement. The Examiner is unable to match the fact of the instant application to guidance of Enfish to establish improvement in the functionality of the computer itself to render the claims eligible. Moreover, a claim directed to an improvement to computer-related technology is likely not similar to claims that have previously been identified as abstract by the courts. In in the instant case, the claims recite a concept similar to the concepts of CyberSource and Int. Ventures v. Cap One Financial.
The current claimed invention does not improve a manner in which a computer functions or technology. The claimed invention at best improves “bibliographic data quality management.” Improvement in quality of information is not equivalent to improvement in storage/database or computer functionality itself. With respect to the applicant’s assertion that the claimed invention “can significantly reduce the time spent on deduplication data objects containing bibliographic data obtained from different sources”, the Examiner notes that alleged “reducing time spent” or speed comes solely from the capability of general purpose computer. Adding a general purpose computer or computer components after the fact to an abstract idea does not integrate a judicial exception into a practical application or provide significantly more. See Affinity Labs v. DirecTV, 838 F.3d 1253, 1262, 120 USPQ2d 1201, 1207 (Fed. Cir. 2016) (cellular telephone); TLI Communications LLC v. AV Auto, LLC, 823 F.3d 607, 613, 118 USPQ2d 1744, 1748 (Fed. Cir. 2016) (computer server and telephone unit). Similarly, "claiming the improved speed or efficiency inherent with applying the abstract idea on a computer" does not integrate a judicial exception into a practical application or provide an inventive concept. Intellectual Ventures I LLC v. Capital One Bank (USA), 792 F.3d 1363, 1367, 115 USPQ2d 1636, 1639 (Fed. Cir. 2015).
It is important to note that in order for a method claim to improve computer functionality, the broadest reasonable interpretation of the claim must be limited to computer implementation. The Examiner contends that the claimed invention could be purely performed mentally.
PNG
media_image1.png
18
19
media_image1.png
Greyscale
That is, a claim whose entire scope can be performed mentally, cannot be said to improve computer technology. Synopsys, Inc. v. Mentor Graphics Corp., 839 F.3d 1138, 120 USPQ2d 1473 (Fed. Cir. 2016). Similarly, a claimed process covering embodiments that can be performed on a computer, as well as embodiments that can be practiced verbally or with a telephone, cannot improve computer technology. See RecogniCorp, LLC v. Nintendo Co., 855 F.3d 1322, 1328, 122 USPQ2d 1377, 1381 (Fed. Cir. 2017) (process for encoding/decoding facial data using image codes assigned to particular facial features held ineligible because the process did not require a computer). See MPEP 2106.05(a).
PNG
media_image1.png
18
19
media_image1.png
Greyscale
Moreover, the current specification in paragraphs 5-6 hints that the process of
deduplication of data objects could be “carried out manually”, but it might be time-consuming and susceptible to errors. As such, the claimed steps of deduplication of citations does not require a computer and could be carried out manually and mentally.
Furthermore, merely adding a general purposes computer components (e.g., one or more processors) are not be sufficient to improve the manner in which a computer functions, because it would evoke a computer merely as a tool to perform existing process.
Additionally, the Examiner respectfully disagrees that the current claimed invention is similar to McRO v. Bandai Namco Games Am. Inc. The basis for the McRO court's decision was that the claims were directed to an improvement in computer-related technology (allowing computers to produce "accurate and realistic lip synchronization and facial expressions in animated characters" that previously could only be produced by human animators), and thus did not recite a concept similar to previously identified abstract ideas. The claimed invention would improve computer-related technology by allowing computer performance of a function not previously performable by a computer. The criteria for determining whether a concept of invention is directed to a particular solution to a problem rather than a general idea of regarding an outcome of the problem itself is that the concept must improve computer-related technology by allowing a computer performance of a function not previously performed by a computer.”
Based on above explanation and reasoning, the Examiner contends that the claimed invention fails to integrate the recite judicial exception into a practical application because the claimed features are not sufficient to improve the functionality
of a computer or technology.
Furthermore, the applicant argues under Step 2B analysis in page 14 of the Remark that:
The combination of additional elements in amended claim 1 is not well-understood, conventional, or routine. The pair-reduction architecture, representing the first stage that limits the application of filtering rules to cluster-member pairs, is a specific algorithmic design choice with a defined functional purpose in the bibliographic deduplication context. The field-type specific, separately applied comparison rules, each independently evaluating equality or concurrence across defined metadata field types, implement a decision architecture that is functionally necessary for the conclusive/inconclusive escalation logic: without per-field-type separate testing, no formal conclusive or inconclusive determination can be made for each rule, and the escalation logic has no basis. The sequential priority-ordered escalation logic itself representing a deterministic control flow with a formal termination criterion at each priority level is a computational protocol, not a conventional ordering of operations.
The Examiner respectfully disagrees.
In order to determine whether a claimed invention is eligible under Step 2B, it must be determined whether the recited additional limitations to transform an abstract idea in a claim to amount to significantly more than the abstract idea.
Here, as identified in step 2B, the additional limitations of “inputting an input dataset comprising the set of citations;” and “outputting a deduplicated dataset comprising only citations that were not identified as duplicates” (and “storing the deduplicated dataset” which is no missing from claim 1) are recited at a high level of generality. Under BRI, these additional limitations are considered as well-understood, conventional, and routine activities of inputting, outputting, and storing data. See MPEP 2106.04(d) and 2106.05(g). Additionally, the claim recites generic computer components of “one or more processors of a computing system” to implement the steps of the invention. Said generic computer components are recited at a high-level of generality such that it amounts no more than mere instructions to apply the exception using a generic computer component and considered to be insignificant extra solution activities.
Thus, the additional elements individually or as whole are not sufficient to amount to significantly more than the judicial exception. Viewed as a whole, these additional claim element(s) do not provide meaningful limitation(s) to transform the abstract idea into a patent eligible application of the abstract idea such that the claim(s) amounts to significantly more than the abstract idea itself.
Therefore, rejections of claims 1, 3-10, 12-16, and 21 under 35 USC 101 as being directed to non-statutory subject matter of abstract idea are maintained.
With respect to 35 USC 103 rejections:
Applicant argues that “amended claim 1…requires that the similarity scoring and clustering steps ‘constitute a computationally efficient first stage executed by the one or more processors that limits application of the filtering rules to citation pairs within the clusters’…However, Huang does not describe or teach the purpose of clustering as being to limit the input domain of a downstream filtering operation to cluster members.”
When construing claim terminology during prosecution before the Office, claims are to be given the broadest and reasonable interpretation consistent with the Specification, reading language of the claims in light of the Specification as it would be interpreted by one of the ordinary skill in the art. In re Am. Acad. of Sci. Tech Ctr., 367 F.3d 1359, 1364 (Fed. Cir. 2004). The Examiner is mindful, however, that limitations are not to be read into claims from the Specification. In re Van Geuns, 988 F.2d 1181 (Fed.
Cir. 1993).
In response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986).
The Examiner holds that the combination of Huang, Heller, Jefferies discloses or at least suggests the feature of clustering, by the one or more processors, the citations based on the calculated similarity score” required by claim 1.
Huang in at least the highlighted sections in pages 2, 4, 8, and 11 discloses that during the first clustering, similar documents (e.g., with similar titles) are clustered to obtain a first set or collection. After clustering similar documents into a first the set or collection of documents are screened/filtered out the documents that meet the requirements. “The first screening unit 103 is configured to calculate the similarity of the documents in each of the first sets, according to the similarity of the calculated documents filters out a plurality of first sets that meet the criteria.”
“The first clustering unit 102 is configured to cluster he documents of similar
titles according to the similarity of the titles of the standardized documents to obtain a plurality of first sets, and in parallel according to the first author of the standardized
documents, publish the source and publish The similarity of the years, clustering similar
documents to obtain multiple second sets.
The first screening unit 103 is configured to calculate the similarity of the
documents in each first set, select a plurality of first sets that meet the conditions
according to the similarity of the calculated documents, and calculate the documents in each second set. Similarity, a plurality of eligible second sets are screened according to the similarity of the calculated documents.”
As such, Huang discloses first clustering the similar documents based on a first attribute (e.g., similar titles) into set/collection of documents, and then screening/filter out the set/collection of documents based on a requirement (e.g., hamming distance between documents).
Furthermore, Heller discloses Heller discloses deduplication of citation, outputting reduplicated citations (See Heller: at least Fig. 14, and 35:34 through 36:15).
Therefore, the combination of Huang, Heller, Jefferies discloses or at least suggests the feature of clustering, by the one or more processors, the citations based on the calculated similarity score” required by claim 1.
With respect to the amended limitation of "each filtering rule verifies whether normalized data fields of the same metadata type of two citations in one cluster are equal or concurring, and each filtering rule is applied separately to more than one metadata type."
The Examiner holds that Huang in at least the highlighted sections in pages 2, 4-5, 7, and 9-10 discloses screening and filtering the qualified document collection (e.g., using the Hamming distance compared to a preset threshold) to identify similar documents and identifying duplicate documents. “S123: Filter out a first cluster whose Hamming distance is less than or equal to a preset threshold, and select a plurality of eligible first clusters to be a plurality of first sets obtained by clustering documents of
similar titles.”
“The first clustering unit 102 is configured to cluster he documents of similar
titles according to the similarity of the titles of the standardized documents to obtain a plurality of first sets, and in parallel according to the first author of the standardized
documents, publish the source and publish The similarity of the years, clustering similar
documents to obtain multiple second sets.
The first screening unit 103 is configured to calculate the similarity of the documents in each first set, select a plurality of first sets that meet the conditions according to the similarity of the calculated documents, and calculate the documents in each second set. Similarity, a plurality of eligible second sets are screened according to the similarity of the calculated documents.”
Therefore, the above teachings of Huang suggests that in some embodiments for the filtering of documents the similarity of the data fields of same metadata (e.g., title, publication source, or year) could be verified separately.
Thus, the combination of Huang, Heller, Jefferies discloses or at least suggests the feature of “applying, by the one or more processors, a filtering rule on the normalized data of citations in clusters for identifying equivalent data objects and flag identified duplicates, wherein each filtering rule verifies whether normalized data fields of the same metadata type of two citations in one cluster are equal or concurring, and each filtering rule is applied separately to more than one metadata type” required by claim 1.
With respect to the applicants’ argument that neither Huang, Heller, or Jefferies discloses the limitation of “"high-priority filtering rule(s) is(are) applied first, and low-priority filtering rule(s) is(are) applied if previously applied rule(s) is(are) not conclusive,"
present in the prior version of claim 1.
In response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986)
The Examiner holds that the combination of Huang, Heller, and Jefferies discloses or at least suggests said feature.
Jefferies discloses applying filter rules in order and applying a filter rule with higher priority and then applying a lower priority rule (See Jefferies: at least 2:9-12, 9:27-33, and 4:57 to 5:7). Thus, modifying the combination of Huang and Heller with teachings of Jefferies would disclose or at least suggests the feature of “wherein the filtering rules have priorities and are sequentially applied, and wherein high-priority filtering rules(s) is (are) applied first, and low-priority filtering rules(s) is (are) applied if previously applied rule(s) is no conclusive,” as required by claim 1.
In conclusion, the Examiner contends that combined teachings of Huang, Heller, and Jefferies (not references individually) would disclose or at least suggests all the limitations and feature claim 1. Thus, the 35 USC 103 rejection of claims 1, 3-10, 12-16, and 21 are maintained.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1, 3-10, 12-16, and 21 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter of abstract ideas.
Step 1:
Claims 1, 3-10, 12-16, and 21 are directed to a method which is one of the statutory categories of invention.
Step 2A:
Prong 1:
Claim 1 is directed to an abstract idea without significantly more.
Claim 1 recites the steps of:
normalizing…data comprised in metadata fields of a same metadata type; [recited at a high level of generality and based on broadest and reasonable interpretation (BRI), it constitutes concept which could be practically performed in the human mind. A person can manually normalize (e.g., converting data types) data]
calculating…similarity score between each pair of citations using the normalized data; [recited at a high level of generality and based on BRI, it constitutes a mathematical concept]
clustering…the citations based on the calculated similarity score; [recited at a high level of generality and based on BRI, it constitutes concept which could be practically performed in the human mind. A person can manually and mentally group
data objects based on similarity scores]
applying…a filtering rule on the normalized data of citations in clusters for identifying equivalent data objects and flag identified duplicates, wherein each filtering rule verifies whether normalized data fields of the same metadata type of two citations in one cluster are equal or concurring, and each filtering rule is applied separately to more than one metadata type; [recited at a high level of generality and based on BRI, it constitutes an evaluation concept which could be practically performed in the human mind. A person can mentally apply filtering rule to identify similar data objects based on evaluation and judgment]
Wherein the steps of calculating a similarity score and clustering the citations together constitute a computationally efficient first stage that…limits application of the filtering rules to citation pairs within the clusters, and wherein said first stage requires lower computational resources per citation than applying the filtering rules to all possible pairwise combinations of citations in the input dataset, [the amount of calculation for similarity score and operations for clustering data within a cluster data would be less than all the possible pairwise combinations of data within the input dataset because a cluster(s) is a subset of a set of “input dataset”. It is a fact that the amounts of [mathematical] calculations and [mental] operations to calculate similarity scores and filtering for a subset of data are always less (and therefore more efficient) than entire set of data]
Wherein the filtering rules have priorities and are sequentially applied, and wherein high-priority filtering rules(s) is (are) applied first, and low-priority filtering rules(s) is (are) applied if previously applied rule(s) is no conclusive. [recited at a high level of generality and based on BRI, it involves observation, evaluation, and/or judgement concepts that could be practically performed in the human mind. A person can mentally determine priorities for the filter rules and apply the filtering rules accordingly]
The above-mentioned steps are processes that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. That is, nothing in the claim element precludes the step from practically being performed in a human mind or with pen and paper. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas (concepts performed in the human mind including an observation, evaluation, judgment, and opinion).
Prong 2:
This judicial exception in claim 1 is not integrated into a practical application. Claim 1 recites the additional steps of “inputting, by the one or processors, an input dataset comprising the set of citations;”, “outputting, by the one or processors, a deduplicated dataset comprising only citations that were not identified as duplicates”, and “storing the deduplicated dataset” which could be considered as insignificant extra solution activities of inputting, outputting, and
storing data. See MPEP 2106.04(d) and 2106.05(g).
Furthermore, the claim recites generic computer components of one or more processors of a computing system to implement the steps of the invention. Said generic computer components are recited at a high-level of generality such that it amounts no more than mere instructions to apply the exception using a generic
computer component and considered to be insignificant extra solution activities.
Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. See MPEP 2106.04(d) and 2106.05(g).
Step 2B:
Claim 1 does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Claim 1 recites the additional steps of “inputting an input dataset comprising the set of citations;”, “outputting a deduplicated dataset comprising only citations that were not identified as duplicates”, and “storing the deduplicated dataset” which could be considered as well-understood, conventional, and routine activities of inputting, outputting, and storing data. See MPEP 2106.04(d) and 2106.05(g).
Furthermore, the claim recites generic computer components of one or more processors of a computing system to implement the steps of the invention. Said generic computer components are recited at a high-level of generality such that it amounts no more than mere instructions to apply the exception using a generic computer component and considered to be insignificant extra solution activities.
Therefore, the claim is not patent eligible.
Regarding dependent claims 3-10, 14,15, and 21,
The dependent claims further include data description and the additional steps for normalizing, filtering, calculating scores (mathematical concept), updating, clustering, and manual verification that could be performed mentally failing to integrate the judicial exception into a practical application or to amount significantly to more than abstract idea.
Regarding dependent claims 12,
the dependent claims additional generic computer functions for notifying a user by displaying data that are considered to be insignificant extra solution and/or well-understood routine computer routine failing to integrate the judicial exception into a practical application or to amount significantly to more than abstract idea.
Regarding dependent claim 13,
The dependent claims additional steps for generic computer function of inputting data that is considered to be insignificant extra solution and/or well-understood routine computer routine failing to integrate the judicial exception into a practical application or to amount significantly to more than abstract idea.
The dependent claims further recite the additional step for merging, normalizing,
calculating similarity score, clustering, filtering, deleting, and updating data that could be performed mentally failing to integrate the judicial exception into a practical application or to amount significantly to more than abstract idea.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3-5, 7-9, and 13-20 are rejected under 35 U.S.C. 103 as being unpatentable over Huang et al., WO 2017/096777 (Huang, hereafter) in view of Heller et
al., US 12,067,366 (Heller, hereafter) and further in view of Jeffries et al., US 6,807,576 (Jeffries, hereafter).
Regarding claim 1,
Huang discloses a computer-implemented method for deduplication of bibliographic citations in a set of citations originating from different databases, wherein each data object is provided with metadata fields; the method being executed by one or more processors of a computing system and comprising the steps:
inputting, by the one or more processors, an input dataset comprising the set of [data] (See Huang: at least the highlighted sections in pages 2 and 7, acquiring documents from at least one website source);
normalizing, by the one or more processors, data comprised in metadata fields of a same metadata type (See Huang: at least the highlighted sections in pages 2 and 7, normalizing or standardizing attributes (i.e., metadata) of document including title, author, abstract, publication source, publication time, and the like);
calculating, by the one or more processors, a similarity score between each pair of [data] (See Huang: at least the highlighted sections in pages 2-4 and 8, calculating a similarity score between documents using normalized attributes or metadata (e.g., title, author, publication data));
clustering, by the one or more processors, the [data]
calculated similarity score (See Huang: at least the highlighted sections in pages 2, 4, 8, and 11, during a first clustering step, cluster similar documents to obtain second sets. For example, the documents of similar titles are clustered (steps S20, S120-S121);
applying, by the one or more processors, a filtering rule on the normalized data of [data] (See Huang: at least the highlighted sections in pages 2, 4-5, and 7, screening and filtering the qualified document collection (e.g., using the Hamming distance compared to a preset threshold) to identify similar documents and identifying duplicate documents. “S123: Filter out a first cluster whose Hamming distance is less than or equal to a preset threshold, and select a plurality of eligible first clusters to be a plurality of first sets obtained by clustering documents of similar titles.”
“The first clustering unit 102 is configured to cluster he documents of similar
titles according to the similarity of the titles of the standardized documents to obtain a plurality of first sets, and in parallel according to the first author of the standardized
documents, publish the source and publish The similarity of the years, clustering similar
documents to obtain multiple second sets.
The first screening unit 103 is configured to calculate the similarity of the documents in each first set, select a plurality of first sets that meet the conditions according to the similarity of the calculated documents, and calculate the documents in each second set. Similarity, a plurality of eligible second sets are screened according to
the similarity of the calculated documents.”
The above teachings of Huang suggests that in some embodiments for the filtering of documents the similarity of the data fields of same metadata (e.g., title, publication source, or year) could be verified separately);
outputting, from the one or more processors, a (See Huang: at least the highlighted sections in pages 4, 6, and 7, identifying and displaying the same or duplicate documents);
wherein the steps of calculating a similarity score and clustering the citations together constitute a computationally efficient first stage executed by the one or more processors that limits application of the filtering rules to citation pairs within the clusters, and wherein said first stage requires lower computational resources per citation than applying the filtering rules to all possible pairwise combinations of citations in the input dataset (Huang inherently discloses these features. For example, Huang in one the embodiments teaches that:
“The first clustering unit 102 is configured to cluster the documents of similar titles according to the similarity of the titles of the standardized documents to obtain a plurality of first sets. The first set includes at least two documents.
The first screening unit 103 is configured to calculate the similarity of the documents in each of the first sets, according to The similarity of the calculated documents filters out a plurality of first sets that meet the criteria.
Specifically, the first screening unit 103 is configured to: preset a weight
corresponding to the document attribute, where the document attribute may be an author, a summary, a publication source, a publication time, and the like. In each of the first sets, the similarity of each document in each first set is calculated according to the weight corresponding to the preset document attribute, and the first set whose similarity of each document is greater than the preset total score is determined to be in accordance with the first set of conditions.”
Therefore, the above teachings of Huang suggests that step of screening/filtering is implemented within the first clustered set implying the all the computations and operation are limited with the first clustered set which inherently results in requiring “lower computational resources per citation than applying the filtering rules to all possible pairwise combinations of citations in the input dataset”).
Although, Huang discloses identifying duplicate documents, Huang does not explicitly teach citations, deduplication of bibliographic citations, outputting a deduplicated dataset comprising only citations that are not identified as duplicates, and storing the deduplicated dataset.
On the other hand, Heller discloses deduplication of citation, outputting reduplicated citations (See Heller: at least Fig. 14, and 35:34 through 36:15). Furthermore, it is obvious to a person of ordinary skill in the art that a deduplicated citation could be stored in a database or a storage. Therefore, it would have been obvious to one of ordinary skill in the art before the time the invention was effectively filed to modify the teachings of Huang with Heller’s teaching in order to implement above function with reasonable expectation of success. The motivation for doing so would have been to help in determining significance of a citation and improve data
storage by eliminating citation duplicate.
The combination of Huang and Heller discloses the limitations as stated above including applying filters to documents. However, it does not explicitly teach Wherein the filtering rules have priorities and are sequentially applied, and wherein high-priority filtering rules(s) is (are) applied first, and low-priority filtering rules(s) is (are) applied if previously applied rule(s) is no conclusive.
On the other hand, Jefferies discloses applying filter rules in order and applying a filter rule with higher priority and then applying a lower priority rule (See Jefferies: at least 2:9-12, 9:27-33, and 4:57 to 5:7). Therefore, it would have been obvious to one of ordinary skill in the art before the time the invention was effectively filed to modify the teachings of the combination of Huang and Heller with Jefferies’s teaching in order to implement above function with reasonable expectation of success. The motivation for doing so would have been to improve efficiency of the method by removing irrelevant results based on filter rule hierarchical priorities.
Regarding claim 3,
the combination of Huang, Heller, Jefferies discloses wherein the metadata fields include data of different data types (See Huang: at least the highlighted sections in pages 2 and 7).
Regarding claim 4,
the combination of Huang, Heller, Jefferies discloses wherein the step of normalizing metadata fields comprises the step of converting the data of different data types into a common data type (See Huang: at least the highlighted sections in pages 2, 7-8, 11 and 17 and Fig. 3, converting data form various format (e.g.,
publication or author data)).
Regarding claim 5,
the combination of Huang, Heller, Jefferies discloses wherein the step of normalizing metadata fields comprises the step of harmonizing strings comprised in metadata fields of the same metadata type by removing special characters (See Huang: at least the highlighted sections in pages 2, 7-8, 11 and 17 and Fig. 3, e.g., removing special characters such as punctuations).
Regarding claim 7,
the combination of Huang, Heller, Jefferies discloses wherein the step of normalizing data included in the metadata fields comprises the step of translating strings representing a word comprised in metadata fields of the same metadata type into a common language using natural language processing (See Huang: at
least the highlighted sections in pages 2, 7-8, 11 and 17 and Fig. 3).
Regarding claim 8,
the combination of Huang, Heller, Jefferies discloses wherein the step of calculating the similarity score includes the step of calculating the similarity score based on data included in the metadata fields of the same metadata type (See Huang: at least the highlighted sections in pages 2, 4, 8, and 11).
Regarding claim 9,
the combination of Huang, Heller, Jefferies discloses wherein the similarity score being calculated using a string metric algorithm (See Huang: at least the highlighted sections in pages 2, 4, 8, and 11, e.g., Hamming distance).
Regarding claim 13,
the combination of Huang, Heller, Jefferies discloses - inputting a second input dataset comprising a different second set of citations; - inputting the deduplicated dataset comprising the deduplicated set of citations; - merging the second input dataset and the deduplicated dataset for providing a merged input dataset;- repeating the steps of normalizing data, calculating the similarity score, clustering the citations, applying the filtering rule, and deleting identified equivalent citations in the merged input dataset; - updating the deduplicated dataset using the deduplicated merged input dataset (See Huang: at least the highlighted sections in pages 2, 4-5, and 7 and Heller: at least Fig. 14, and 35:34 through 36:15).
Regarding claim 14,
the combination of Huang, Heller, Jefferies discloses wherein updating includes the step of replacing citations comprised in the deduplicated dataset with citations comprised in the deduplicated merged input dataset (See Huang: at least the highlighted sections in pages 2, 4-5, and 7 and Heller: at least Fig. 14, and 35:34 through 36:15).
Regarding claim 15,
the combination of Huang, Heller, Jefferies discloses wherein updating includes the step of adding citations not yet comprised in the deduplicated dataset from the deduplicated merged input dataset into the deduplicated dataset (See Huang: at least the highlighted sections in pages 2, 4-5, and 7 and Heller: at least Fig. 14, and 35:34 through 36:15).
Regarding claim 16,
the combination of Huang, Heller, Jefferies discloses comprising a step of outputting a second duplicate dataset containing only identified equivalent citations from the merged input dataset (See Huang: at least the highlighted sections in pages 2, 4-5, and 7 and Heller: at least Fig. 14, and 35:34 through 36:15).
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Huang et al., WO 2017/096777 (Huang, hereafter) in view of Heller et al., US 12,067,366 further in view of Jeffries et al., US 6,807,576 and further in view of Benke et al., US 2023/0090601 (Benke, hereafter).
The combination of Huang, Heller, Jefferies discloses normalizing attributes or metadata of the documents by removing characters, however, it does not explicitly teach wherein the removal of special characters comprises the step of removing an URL prefix.
On the other hand, Benke discloses normalizing a document by removing a URL (See Benke: at least para 87). Therefore, it would have been obvious to one of ordinary skill in the art before the time the invention was effectively filed to modify the teachings of the combination of Huang, Heller, Jefferies with Benke’s teaching in order to implement above function with reasonable expectation of success. The motivation for doing so would have been to improve identifying the duplicate documents by reducing the variations.
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Huang et al., WO 2017/096777 (Huang, hereafter) in view of Heller et al., US 12,067,366 further in view of Jeffries et al., US 6,807,576 and further in view of Ferreira et al., US 2024/0176795 (Ferreira, hereafter).
The combination of Huang, Heller, Jefferies discloses wherein the string metric algorithm is provided as an edit distance algorithm; however, it does not explicitly teach Levenshtein distance algorithm.
On the other hand, Ferreira discloses using Levenshtein distance algorithm to calculate similarity (See Ferreira: at least para 49). Therefore, it would have been obvious to one of ordinary skill in the art before the time the invention was effectively filed to modify the teachings of the combination of Huang, Heller, Jefferies with Ferreira’s teaching in order to implement above function with reasonable expectation of success. The motivation for doing so would have been to improve similarity calculation
taking advantage of Levenshtein distance algorithm benefits.
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Huang et al., WO 2017/096777 (Huang, hereafter) in view of Heller et al., US 12,067,366 and further in view of Jeffries et al., US 6,807,576 further in view of Cialdea, JR. et al., US 2015/0032730 (Cialdea, hereafter).
The combination of Huang, Heller, Jefferies discloses identifying equivalent data objects based on the entered filtering rule or on a pre-defined filtering rule, however, it does not explicitly teach comprising the step of notifying a user by displaying a dialog for entering the filtering rule.
On the other hand, Cialdea discloses asking a user to specify filtering rules to be applied to data on a GUI (See Cialdea: at least para 46 and Fig. 5). Therefore, it would have been obvious to one of ordinary skill in the art before the time the invention was effectively filed to modify the teachings of the combination of Huang, Heller, Jefferies with Cialdea’s teaching in order to implement above function with reasonable expectation of success. The motivation for doing so would have been to improve identifying duplicate documents by allowing a user to specify filtering conditions.
Claim 21 is rejected under 35 U.S.C. 103 as being unpatentable over Huang et al., WO 2017/096777 (Huang, hereafter) in view of Heller et al., US 12,067,366 further in view of Jeffries et al., US 6,807,576 and further in view of Sim, US 2022/0247706.
The combination of Huang, Heller, Jefferies discloses applying a filtering rule, however, it does not explicitly teach the application of the filtering rule includes a conditional step of manual verification.
On the other hand, Sim discloses a user verifies setting of filter conditions (See Sim: at least para 89 and Fig. 7). Therefore, it would have been obvious to one of ordinary skill in the art before the time the invention was effectively filed to modify the teachings of the combination of Huang, Heller, Jefferies with Sim’s teaching in order to implement above function with reasonable expectation of success. The motivation for doing so would have been to improve functionality of the method by allowing a user to verify the filtering rules.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Points of Contact
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HARES JAMI whose telephone number is (571)270-1291. The examiner can normally be reached M-F 9:00a-5:00p.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amy Ng can be reached at (571) 270-1698. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Hares Jami/ Primary Examiner, Art Unit 2164
07/14/2026