DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
Claims 21-40 are pending of which claims 22, 28 and 35 are in independent form.
Claims 21-40 are rejected on the ground of nonstatutory double patenting.
Claims 21-40 are rejected under 35 U.S.C. 101.
Claims 21-40 are rejected under 35 U.S.C. 103.
Response to Arguments
Applicant’s arguments with respect to claim(s) 21-40 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Regarding the 35 USC 101 (Abstract Idea), remarks made by the applicant.
Applicant’s arguments have been fully considered but are not persuasive. The independent claims recite the abstract mental process of identifying missing information, extracting/selecting proposed information from available information. Recitation of an analytics service, ML models, and programmatic interfaces simply implements these functions using computer technology and does not itself improve the operation of the computer, ML model, or other technology. Likewise, providing an explanation of the extraction constitutes additional information concerning the result rather than a technological improvement. Therefore, the claims do not integrate the abstract idea into a practical application; hence the 35 USC 101 rejection is maintained.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 21-40 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-20 of U.S. Patent No. US 12130863 B1. Although the claims at issue are not identical, they are not patentably distinct from each other.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 21-40 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more.
The claim(s) recite(s) using AI or ML for efficient attribute extraction.
With respect to step 1 of the patent subject matter eligibility analysis, the claims are directed to a process, machine, manufacture, or composition of matter. Independent claim 35 directed to a non-transitory computer-accessible storage media, which is directed to one of the four statutory subject matters. Independent Claim 28 is directed to a system, including one or more processors and a memory, which is a machine. Independent claim 21 is directed to a method, which is a process. All other claims depend on claims 1, 6 and 16. As such, claims 1-20 are directed to a statutory category.
With respect to step 2A, Prong One, prong one, the claims recite an abstract idea, law of nature, or natural phenomenon. Specifically, the following limitations recite mathematical concepts and/or mental processes and/or certain methods of organizing human activity.
The claims recite the following core steps:
Determining that a value of an attribute of an incomplete record is missing,
Extracting, using ML models, a proposed value for that attribute from a corpus of documents,
Providing an explanation for the extraction using the ML models.
These steps amount to, identifying missing info (mental process); analyzing a document to infer or predict a value (mathematical concept/algorithm); generating an explanation for the results (organizing/presenting information), which are all considered abstract.
The claims fall within:
Mental Process (identifying missing fields, review documents, infer values, and explain reasoning),
Information Analysis Evaluation (extract and proposing values based on review),
Data Collection and Presentation (Gathering data from corpus and providing explanation).
With respect to step 2A, Prong Two, prong two, the claims do not recite additional elements that integrate the judicial exception into a practical application. The following limitations are considered “additional elements” and explanation will be given as to why these “additional elements” do not integrate the judicial exception into a practical application.
The claims are generic computer components preforming their routine functions. The claims recite:
Analytic services,
ML models,
Programmatic interface,
A corpus of documents.
The claims FAIL to integrate the abstract idea into a practical application.
The claims do not:
Improve training techniques or inference efficiency,
Improve data storage or retrieval mechanism,
Improve ML model architecture,
Provide a new technical mechanism for handling incomplete records,
Improve computer functionality.
No step ties the claimed operation to a specific technological constraint or transforms the abstract idea into a technical solution to a technical problem.
The claims simply use ML models as tools to perform the abstract idea, which does not integrate the exception into a practical application. The analytics service is a generic processing environment. The programmatic interface simply output the results.
As mentioned above, none of the elements in the claims impose meaningful limits or provide a technological improvement. There are no recited improvements to database structures, embedding generation hardware, computer memory, or processor operation.
Everything is simply implemented by a computer performing generic data processing.
With respect to Step 2B. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The additional limitations are directed to a computer readable storage medium, computer, memory, and processor, at a very high level of generality and without imposing meaningful limitations on the scope of the claim. In addition, pages 2-5 of the published instant specification describe generic off‐the‐shelf computer‐based elements for implementing the claimed invention, which does not amount to significantly more than the abstract idea and is not enough to transform an abstract idea into eligible subject matter. Such generic, high‐level, and nominal involvement of a computer or computer‐based elements for carrying out the invention merely serves to tie the abstract idea to a particular technological environment, which is not enough to render the claims patent‐eligible, as noted at pg.74624 of Federal Register/Vol. 79, No. 241, citing Alice, which in turn cites Mayo. Further, See, e.g., Alice Corp. Pty. Ltd. v. CLS Bank Int'l, 134 S. Ct. 2347, 2359‐60, 110 USPQ2d 1976, 1984 (2014). See also OIP Techs. v. Amazon.com, 788 F.3d 1359, 1364, 115 USPQ2d 1090, 1093‐94 (Fed. Cir. 2015) ("Just as Diehr could not save the claims in Alice, which were directed to 'implement[ing] the abstract idea of intermediated settlement on a generic computer', it cannot save O/P's claims directed to implementing the abstract idea of price optimization on a generic computer.") (citations omitted). See also, Affinity Labs of Texas LLC v. DirecTV LLC, 838 F.3d 1253, 1257‐1258 (Fed. Cir. 2016) (mere recitation of a GUI does not make a claimpatent‐eligible); Intellectual Ventures I LLC v. Capital One Bank, 792 F.3d 1363, 1370 (Fed. Cir. 2015) ("the interactive interface limitation is a generic computer element".).
The additional elements are broadly applied to the abstract idea at a high level of generality ("similar to how the recitation of the computer in the claims in Alice amounted to mere instructions to apply the abstract idea of intermediated settlement on a generic computer,") as explained in MPEP § 2106.05(f)) and they operate in a well‐understood, routine, and conventional manner.
MPEP § 2106.0S(d)(II) sets forth the following:
The courts have recognized the following computer functions as well-understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity.
• Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec ... ; TLI Communications LLC v. AV Auto. LLC ... ; OIP Techs., Inc., v. Amazon.com, Inc ... ; buySAFE, Inc. v. Google, Inc ... ;
• Performing repetitive calculations, Flook ... ; Bancorp Services v. Sun Life ... ;
• Electronic recordkeeping, Alice Corp ... ; Ultramercial ... ;
• Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc ... ;
• Electronically scanning or extracting data from a physical document, Content Extraction and Transmission, LLC v. Wells Fargo Bank ... ; and
• A web browser's back and forward button functionality, Internet Patent
• Corp. v. Active Network, Inc. ...
. . . Courts have held computer-implemented processes not to be significantly more than an abstract idea (and thus ineligible) where the claim as a whole amounts to nothing more than generic computer functions merely used to implement an abstract idea, such as an idea that could be done by a human analog (i.e., by hand or by merely thinking).
In addition, when taken as an ordered combination, the ordered combination adds nothing that is not already present as when the elements are taken individually. There is no indication that the combination of elements integrate the abstract idea into a practical application. Their collective functions merely provide conventional computer implementation. Therefore, when viewed as a whole, these additional claim elements do not provide meaningful limitations to transform the abstract idea into a practical application of the abstract idea or that the ordered combination amounts to significantly more than the abstract idea itself.
The dependent claims have been fully considered as well, however, similar to the findings for claims above, these claims are similarly directed to the “Mental Processes” grouping of abstract ideas set forth in the 2019 PEG, without integrating it into a practical application and with, at most, a general purpose computer that serves to tie the idea to a particular technological environment, which does not add significantly more to the claims. The ordered combination of elements in the dependent claims (including the limitations inherited from the parent claim(s)) add nothing that is not already present as when the elements are taken individually. There is no indication that the combination of elements improves the functioning of a computer or improves any other technology. Their collective functions merely provide conventional computer implementation. Accordingly, the subject matter encompassed by the dependent claims fails to amount to significantly more than the abstract idea.
Examiner further indicates that the remaining claims also fall with 35 USC 101 (abstract idea) for at least the same reasons.
Regarding claims 22, 29 and 36,
The claim recites:
The corpus documents include a web page.
This is merely a specific type of data source, used by abstract analytic process. This does not change the nature of the abstract idea (data extraction and inference). It does not add a technical improvement to an abstract idea.
There is no practical application, and no inventive step, the claims are still considered abstract.
Regarding claims 23, 30 and 37,
The claim recites:
explanation comprises an indication of signals, obtained by applying rules to documents.
This is merely an information analysis and explanation generation (abstract). Applying “rules” to determine signals is algorithmic process, is recognized as abstract. Explanation is still presentation of information.
This does not change the nature of the abstract idea. It does not add a technical improvement to an abstract idea.
There is no practical application, and no inventive step, the claims are still considered abstract.
Regarding claims 24, 31 and 38,
The claim recites:
specific rule types (text pattern rule, absence indicator rule, enumeration-based rule).
This is merely reciting specific algorithms, which are considered mathematical concepts. Does not require any unconventional model, processing architecture, or technical improvements.
The is simply narrowing the abstract idea used in the already abstract inference step.
This does not change the nature of the abstract idea. It does not add a technical improvement to an abstract idea.
There is no practical application, and no inventive step, the claims are still considered abstract.
Regarding claims 25, 32 and 39,
The claim recites:
generating the rules based on analysis of example documents.
This is merely data analysis, which is a mental process/mathematical algorithm. There is no specific technical mechanism for rule-based generation provided; only the results. This does not integrate the exception into a practical application.
This does not change the nature of the abstract idea. It does not add a technical improvement to an abstract idea.
There is no practical application, and no inventive step, the claims are still considered abstract.
Regarding claims 26, 33 and 40,
The claim recites:
obtaining, a plurality of signals via rules, and explanation includes a fraction of signals that agree.
Calculating and comparing signals and computing a fraction is a mathematical evaluation (mathematical algorithm). No change to the generic computer implementation. This does not integrate the exception into a practical application.
This does not change the nature of the abstract idea. It does not add a technical improvement to an abstract idea.
There is no practical application, and no inventive step, the claims are still considered abstract.
Regarding claims 27, and 34,
The claim recites:
the explanation includes content from a particular section of a plurality of sections of a document.
Identifying and referencing a document section is merely information retrieval, which is a mental process. This does not improve document storage, indexing, data structure, or computer functionality. This is merely focused on selection and presenting content.
This does not change the nature of the abstract idea. It does not add a technical improvement to an abstract idea.
There is no practical application, and no inventive step, the claims are still considered abstract.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 21, 22, 24-45, 27-29, 31-, 32, 34-36, 38, and 39 are rejected under 35 U.S.C. 103 as being unpatentable over Chandrasekhar; Govind et al. (US 20200151201 A1) [Chandrasekhar] in view of Zhang; Zuohua (US 11182691 B1) [Zhang] in view of Galginaitis; Steven et al. (US 20210201228 A1) [Galginaitis].
Regarding claim 21, 28 and 35, Chandrasekhar discloses, a computer-implemented method, comprising: determining, by an analytics service, that a value of an attribute of a record which is to be included in a collection of records is missing (The system 100 may also employ the extraction heuristics building of training datasets in the training data 120 for machine learning models of the machine learning network 130. As the system 100 generates output dataset from applying the extraction heuristics to various input product catalogs to generate output datasets. The output datasets invariably represent product records enriched with product attribute information and may thereby be used as training datasets for training one or more machine learning models ¶ [0041]. For the extraction and inference mode of the system 100, an input product record 150 is first sent to the preprocessor module 106-1 which prepares the input product record for efficient attribute enrichment. The role of the preprocessor module 106-1 is to complete a number of tasks, including classification, attribute placeholder creation as well as formatting. The preprocessor module 106-1 preprocesses the input product record 150 with the classification tree. Using one or more machine learning classification models trained via the active learning training pipeline module 108, an output product category is predicted by the preprocessor module 106-1 for each input product record 150. The output predicted product category is represented by a specific node within the classification tree ¶ [0042]. One or more classification models may also be employed by the attribute inference module 106-3 to infer attribute-keys. The classification models may utilize extracted attributes as well as the attribute-master list and the attributes-allowed-values list. Attribute-keys that have a fixed set of possible key-value can be predicted by the classification output. In addition, one or more imputation models may also be employed by the attribute inference module 106-3 to infer attribute-keys. Imputation models can be used to complete absent attribute-keys that are dependent on the presence of other proximate attribute-keys. The imputation models may utilize extracted attributes as well as attribute distribution histograms, and attribute key-value distribution histograms. The task of the imputation models is framed as that of filling in the other missing attribute-keys given a surrounding context created by the presence of existing attribute-keys. The inferred attribute-keys of generated by the attribute inference module 106-3 may be merged with the extracted product record to generate a merged product record. The merged product record is sent as input for the normalization module 106-4 ¶ [0047]);
extracting, using one or more machine learning models, a particular proposed value of the attribute (The computer system initiates an extraction and inference mode that utilizes the classification tree, the plurality of attribute lists and the plurality of extraction heuristics to output a normalized record categorizing at least one product record's input data according to the classification tree's product-description taxonomy, the extraction and inference mode and generates: (i) a predicted product category by feeding the respective product record into one or more machine learning (ML) models that tracks the classification tree; (ii) at least one extracted attribute by feeding the respective product record and the predicted product category into the one or more machine learning (ML) models; (iii) an extracted product record based on the extracted attribute, the respective product record and the predicted product category; (iv) at least one inferred attribute by feeding the extracted product record into the one or more machine learning (ML) models ¶ [0007]. The system 100 may also employ the extraction heuristics building of training datasets in the training data 120 for machine learning models of the machine learning network 130. As the system 100 generates output dataset from applying the extraction heuristics to various input product catalogs to generate output datasets. The output datasets invariably represent product records enriched with product attribute information and may thereby be used as training datasets for training one or more machine learning models ¶ [0041]-[0047]; also see ¶ [0057]);
of the particular proposed value for the attribute (The attribute extraction module 106-2 may employ one or more machine learning extraction models of the machine learning (ML) network 130 to the preprocessed product record. For example, a named entity recognition (NER) model(s) may segment out product attributes from text in the preprocessed product record. The output of the attribute extraction module 106-2 may be one or more extracted attributes from application of both the extraction heuristics and the one or more machine learning extraction models. The extracted attributes are inserted into the preprocessed record to generate an extracted product record. The extracted product record is augmented with relevant metadata, such as an attribute-master list, an attribute-allowed-value list, and one or more attribute distribution histograms ¶ [0045]; The attribute inference module 106-3 takes the extracted product record as input. The attribute inference module 106-3 employs various types of ML models of the ML network 130. The attribute inference module 106-3 may employ regression ML models to predict numeric attribute-keys. An example numeric product attribute may be a “kilogram” attribute-key with key-values that are represented by numeric data—as opposed to textual data. The regression ML models may utilize extracted attributes and an attribute-master list and an attribute-allowed-values list. Numeric attribute-keys tend to have key-values within fixed bounds which can be predicted by the regression output ¶ [0046]; As shown in FIG. 4B, the attribute extraction module 106-2 matches a preprocessed product record against the results from the Taxonomy Generator module 102-2. Relevant Extraction Heuristics 102-8-1, 102-8-2, 102-8-3 are selected and applied on the input product record as well. By matching the record against the results from the Taxonomy Generator, the relevant Machine Learning models are selected and applied on the product. The extracted attributes may then be merged with the product record by inserting the extracted attributes into the product record Also add the source and confidence score associated with each extraction result ¶ [0073]).
However, Chandrasekhar does not explicitly facilitate from a corpus of documents; providing, via one or more programmatic interfaces of the analytics service.
Zhang discloses, from a corpus of documents (The tfidf measure may be intended to reflect the relative importance of words within a document in a collection or corpus; the tfidf value for a given word typically is proportional to the number of occurrences of the word in a document, offset by the frequency of the word in the collection as a whole [col. 29, ll. 26-53]);
providing, via one or more programmatic interfaces of the analytics service (number of MLS programmatic interfaces (such as application programming interfaces (APIs)) may be defined by the service, which guide non-expert users to start using machine learning best practices relatively quickly, without the users having to expend a lot of time and effort on tuning models, or on learning advanced statistics or artificial intelligence techniques [col. , ll. 40-64]; the MLS may implement a set of programmatic interfaces 161 (e.g., APIs, command-line tools, web pages, or standalone GUIs) that can be used by clients 164 (e.g., hardware or software entities owned by or assigned to customers of the MLS) to submit requests 111 for a variety of machine learning tasks or operations [col. 8, ll. 64-col. 9, ll. 21]. Also see [col. 16, ll. 45-62], [col. 20, ll. 27-65], [col. 23, ll. 24-51], [col. 43, ll. 13-42], [col. 44, ll. 28-44]).
It would have been obvious to one ordinary skilled in the art at the time of the filing of the present invention to combine the teachings of the cited references because Zhang’s system would have allowed Chandrasekhar to facilitate from a corpus of documents; providing, via one or more programmatic interfaces of the analytics service. The motivation to combine is apparent in the Chandrasekhar's reference, because it would be desirable to improve quality of the results obtained from machine learning algorithms and providing more effective and efficient results.
However, neither Chandrasekhar nor Zhang explicitly facilitates an explanation for the extraction, using the one or more machine learning models.
Galginaitis discloses, an explanation for the extraction, using the one or more machine learning models (In one embodiment, a justifier model may be provided which receives inputs from the predicted attributes, automated accept/reject recommendations, overall scores, and attribute contributions to the scores, and in response generates reasons and evaluations to be shown to the user. In an embodiment, the justifier model may provide automatic filling of explanations for recommendations for accepting or rejecting comparable entities for transfer pricing purposes. The justifier model may provide interpretable and consistent explanations of the automated recommendations for accepting or rejecting comparables for users and tax authorities. For example, such interpretable explanations may be in the form of text, tables, charts, figures, or any combination thereof ¶ [0066]; also see ¶ [0064]-[0065]; The CI system may include a number of artificial intelligence models and machine learning models such as classifiers, which together interpret the input data sources to identify the products, services, functions and other attributes of the potential comparables ¶ [0031], [0033], [0036], [0059]).
It would have been obvious to one ordinary skilled in the art at the time of the filing of the present invention to combine the teachings of the cited references because Galginaitis’ system would have allowed Chandrasekhar and Zhang to facilitate an explanation for the extraction, using the one or more machine learning models. The motivation to combine is apparent in the Chandrasekhar and Zhang's reference, because there is a need for an improved system and method to accurately and consistently identify a set of comparable companies for transfer pricing, valuation, and other purposes.
Regarding claim 22, 29 and 36, the combination of Chandrasekhar, Zhang in view of Galginaitis discloses, wherein the corpus of documents includes a web page (In addition, as shown in FIG. 3, company websites may also provide valuable information such as product lines, services, and other information. According to one embodiment, the CI system is programmed to perform web scraping of company websites for comparables to automatically pull the relevant sections of the company website and present them to the analyst in an efficient manner. The information obtained by web scraping also may be relied upon by the classifiers, such as the classifiers described above with reference to FIG. 2, to improve the models for automated analysis ¶ [0049]).
Regarding claims 24, 31 and 38, the combination of Chandrasekhar, Zhang in view of Galginaitis discloses, wherein the one or more rules comprise one or more of: (a) a text pattern based rule, (b) an attribute absence indicator rule, or (c) an enumeration based rule (Chandrasekhar: The system 100 defines key data based on the product data (Act 204). The key data comprises at least one product attribute-key (master attribute-key), a plurality of allowed key-values, and a representative key-value matched to a list of equivalent key-values. The system 100 defines heuristic data based on the product data and the key data (Act 206). The heuristic data comprises a mapping between a master attribute-key and a list of equivalent attribute-keys, at least one key-value exclusive to a particular master attribute-key and a frequent text phrase as a regex key-value that corresponds to a master attribute-key ¶ [0051]; Based on the input product data, the classification tree, and the plurality of the attribute lists, the system 100 generates a plurality of extraction heuristics (Act 218). An attribute-key-mapping heuristic has a respective master attribute-key matched to a list of equivalent attribute-keys. A key-value-dictionary heuristic defines exclusive key-values for particular master attribute-keys. A pattern heuristic identifies a frequent text phrase as a regex key-value that corresponds to a master attribute-key ¶ [0055]).
Regarding claims 25, 32 and 39, the combination of Chandrasekhar, Zhang in view of Galginaitis discloses, generating the one or more rules based at least in part on analysis of one or more example documents which contain respective values of the attribute (Chandrasekhar: Based on the input product data, the classification tree, and the plurality of the attribute lists, the system 100 generates a plurality of extraction heuristics (Act 218). An attribute-key-mapping heuristic has a respective master attribute-key matched to a list of equivalent attribute-keys. A key-value-dictionary heuristic defines exclusive key-values for particular master attribute-keys. A pattern heuristic identifies a frequent text phrase as a regex key-value that corresponds to a master attribute-key ¶ [0055], also see ¶ [0051]-[0054]; The heuristics generator module 102-7 compares each structured key-value pair to one or more attribute-allowed-values lists for master attribute-keys ¶ [0066]-[0067]. As the extraction heuristics are generated and updated, human testers may sample data from the list and send verifications of the extraction heuristics or overrides of the extraction heuristics to the human verification module 112 ¶ [0039]-[0040]).
Regarding claims 27 and 34, the combination of Chandrasekhar, Zhang in view of Galginaitis discloses, wherein the explanation comprises content of a particular section of a plurality of sections of a particular document of the corpus, wherein the particular proposed value is extracted, at least in part, from the particular section (Galginaitis: In addition, as shown in FIG. 3, company websites may also provide valuable information such as product lines, services, and other information. According to one embodiment, the CI system is programmed to perform web scraping of company websites for comparables to automatically pull the relevant sections of the company website and present them to the analyst in an efficient manner. The information obtained by web scraping also may be relied upon by the classifiers, such as the classifiers described above with reference to FIG. 2, to improve the models for automated analysis ¶ [0049]).
Claim(s) 23, 26, 30, 33, 37 and 40 are rejected under 35 U.S.C. 103 as being unpatentable over Chandrasekhar in view of Zhang in view of Galginaitis in view of Quader; Shaikh Shahriar et al. (US 20210209412 A1) [Quader].
Regarding claims 23, 30 and 37, the combination of Chandrasekhar, Zhang in view of Galginaitis discloses,
However, neither one of Chandrasekhar, Zhang, or Galginaitis explicitly facilitates wherein the explanation comprises an indication of one or more signals, obtained by applying one or more rules to one or more documents of the corpus, of presence of a value of the attribute in the one or more documents.
Quader discloses, wherein the explanation comprises an indication of one or more signals, obtained by applying one or more rules to one or more documents of the corpus, of presence of a value of the attribute in the one or more documents (In one example, the heuristics generator module 211 creates a set of weak classifiers by diving DL into a training dataset and an evaluation dataset, and employing one or more classifier algorithms and an iterative process of all possible combinations of input features to determine classifiers that perform well when applied to the evaluation dataset. Classifier algorithms may include, but are not limited to, decision stump algorithms and random forest algorithms. More specifically, in one exemplary implementation, the heuristics generator module 211 uses an ensemble of decision stumps as the inner classification model to mimic the threshold-based heuristics that users usually write ¶ [0044]-[0046], [0048], [0050], [0051]).
It would have been obvious to one ordinary skilled in the art at the time of the filing of the present invention to combine the teachings of the cited references because Quader’s system would have allowed Chandrasekhar, Zhang, and Galginaitis to facilitate wherein the explanation comprises an indication of one or more signals, obtained by applying one or more rules to one or more documents of the corpus, of presence of a value of the attribute in the one or more documents. The motivation to combine is apparent in Chandrasekhar, Zhang, and Galginaitis' reference, because there is a need for an improved automated supervised label training data that is used to train a machine learning model.
Regarding claims 26, 33 and 40, the combination of Chandrasekhar, Zhang in view of Galginaitis and Quader discloses, obtaining, by applying one or more rules, a plurality of signals associated with presence of a value of the attribute in one or more documents of the corpus, wherein the explanation comprises an indication of a fraction of the plurality of signals which agree with one another (Quader: In accordance with aspects of the invention, the heuristics generator module 211 is configured to automatically produce a set of heuristics using a labeled dataset and use the heuristics to automatically assign initial labels to data points in an unlabeled dataset. In one implementation, the heuristics generator module 211 receives the labeled dataset and the unlabeled dataset from a client device 220 via a network 230 ¶ [0037]-[0038]; In embodiments, the low confidence points are originated when either the heuristics abstain from labeling or disagree on specific points. Therefore, the data-driven learner module 212 enhances the quality of the labels by trying to eliminate the abstaining effect and resolve the disagreements between the heuristics to increase their accuracies ¶ [0049]; In embodiments, the heuristics generator module 211 generates the heuristics by employing a process of creating a set of probabilistic classification models that take one or more features as input and calculate probability distribution over a set of classes. Then, the heuristics generator module 211 uses this distribution to either assign labels to the unlabeled dataset (i.e., assigns either −1 or 1) or abstain (i.e., outputs (0)) ¶ [0043]).
Conclusion
The examiner requests, in response to this Office action, support be shown for language added to any original claims on amendment and any new claims. That is, indicate support for newly added claim language by specifically pointing to page(s) and line no(s) in the specification and/or drawing figure(s). This will assist the examiner in prosecuting the application.
When responding to this office action, Applicant is advised to clearly point out the patentable novelty which he or she thinks the claims present, in view of the state of the art disclosed by the references cited or the objections made. He or she must also show how the amendments avoid such references or objections See 37 CFR 1.111(c).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MOHAMMAD S ROSTAMI whose telephone number is (571)270-1980. The examiner can normally be reached Mon-Fri From 9 a.m. to 5 p.m..
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Boris Gorney can be reached at (571)270-5626. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
8/24/2026
/MOHAMMAD S ROSTAMI/Primary Examiner, Art Unit 2154