DETAILED ACTION
Response to Amendment
Claims 1-20 are pending. Claims 1-20 are amended directly or by dependency on an amended claim.
Response to Arguments
Applicant's arguments filed 21 July, 2026 with respect to the 35 USC 103 rejections of claims 1, 19, and 20 have been fully considered but are moot because the new ground of rejection does not rely on the combination of references applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. In particular, while claims 1, 19, and 20 now incorporate part of one of the claims indicated as allowable the claims do not incorporate all of the dependent claim and therefore are not allowed. New art is used to teach the newly added limitation.
Applicant’s arguments, see page 9, filed 21 July, 2026 with respect to the 35 USC 112b rejection of claim 5 along with accompanying amendments received on the same date have been fully considered and are persuasive. The 35 USC 112b rejection of claim 5 has been withdrawn.
All other arguments are by similarity or dependency and are addressed by the above.
Priority
Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged. Applicant has not complied with one or more conditions for receiving the benefit of an earlier filing date under 35 U.S.C. 35 USC 119(e) as follows:
The later-filed application must be an application for a patent for an invention which is also disclosed in the prior application (the parent or original nonprovisional application or provisional application). The disclosure of the invention in the parent application and in the later-filed application must be sufficient to comply with the requirements of 35 U.S.C. 112(a) or the first paragraph of pre-AIA 35 U.S.C. 112, except for the best mode requirement. See Transco Products, Inc. v. Performance Contracting, Inc., 38 F.3d 551, 32 USPQ2d 1077 (Fed. Cir. 1994).
The disclosure of the prior-filed application, Application No. 63/169789, fails to provide adequate support or enablement in the manner provided by 35 U.S.C. 112(a) or pre-AIA 35 U.S.C. 112, first paragraph for one or more claims of this application. The disclosure of the prior-filed application, Application No. 63/169789, fails to give support for (at a minimum): “generating a combined confidence score by combining an entity extraction confidence score of the one or more entity extraction confidence scores and an OCR confidence score of the one or more OCR confidence scores, wherein the OCR confidence score corresponds to a textual region from which the extracted entity was extracted; comparing the combined confidence score to an acceptance threshold; if the combined confidence score meets the acceptance threshold: transmitting the extracted entity to a database for validation against reference data stored in the database; verifying existence of the extracted entity in the database; comparing one or more attributes of the extracted entity against stored attributes maintained in the database; and validating one or more relationships of the extracted entity with other entities in the database; and if the combined confidence score does not meet the acceptance threshold: withholding the extracted entity from database validation”.
Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged. Applicant has not complied with one or more conditions for receiving the benefit of an earlier filing date under 35 U.S.C. 35 USC 120 as follows:
The later-filed application must be an application for a patent for an invention which is also disclosed in the prior application (the parent or original nonprovisional application or provisional application). The disclosure of the invention in the parent application and in the later-filed application must be sufficient to comply with the requirements of 35 U.S.C. 112(a) or the first paragraph of pre-AIA 35 U.S.C. 112, except for the best mode requirement. See Transco Products, Inc. v. Performance Contracting, Inc., 38 F.3d 551, 32 USPQ2d 1077 (Fed. Cir. 1994).
The disclosure of the prior-filed application, Application No. 17/569121, fails to provide adequate support or enablement in the manner provided by 35 U.S.C. 112(a) or pre-AIA 35 U.S.C. 112, first paragraph for one or more claims of this application. The disclosure of the prior-filed application, Application No. 17/569121, fails to give support for (at a minimum): “generating a combined confidence score by combining an entity extraction confidence score of the one or more entity extraction confidence scores and an OCR confidence score of the one or more OCR confidence scores, wherein the OCR confidence score corresponds to a textual region from which the extracted entity was extracted; comparing the combined confidence score to an acceptance threshold; if the combined confidence score meets the acceptance threshold: transmitting the extracted entity to a database for validation against reference data stored in the database; verifying existence of the extracted entity in the database; comparing one or more attributes of the extracted entity against stored attributes maintained in the database; and validating one or more relationships of the extracted entity with other entities in the database; and if the combined confidence score does not meet the acceptance threshold: withholding the extracted entity from database validation”.
Applicant is given priority to 18/357655 filed 24 July, 2023. This is the earliest priority which has adequate support.
Drawings
The drawings were received on 21 November, 2025. These drawings are accepted.
Examiner’s Comment
Claim 1 (and by similarity claims 19, 20) are NOT rejected on the ground of nonstatutory double patenting over claim 1 of copending Application No. 19/396753 (even when taking into account claim 3 of 19/396753 which discusses embedding similarities) or 19/441538 because while some of the language overlaps, there are sufficient differences between the copending applications so as to not necessitate a double patenting rejection. However, examiner notes there is some language overlap and advises the applicant to keep this copending application in mind when drafting any potential future amendments.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 3, 4, 9, 18 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wheaton et al. (US 20210110527 A1) in view of Zeng et al. (US 20220300834 A1).
Regarding claims 1 and 20, Wheaton et al. disclose a method for validating extracted document data, comprising, and non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to ([0042], [0117]): performing optical character recognition on a document (identifying semi-structured data generated by optical character recognition, [0033], text identified by an OCR processes may be provided along with the corresponding portion of the original image, [0275]) to generate one or more OCR confidence scores for textual content of the document (confidence score associated with overall text field accuracy, [0246], OCR text accuracy score, [0273]) generating one or more entity extraction confidence scores for the one or more extracted entities (estimate confidence scores at multiple levels (e.g., template and field level), [0245], confidence associated with each step, based on scores relating to one or more of assessed image quality, template matching distance, document structure overlap, identified metadata overlap, and overall text field accuracy, [0246]) for each extracted entity, generating a combined confidence score by combining an entity extraction confidence score of the one or more entity extraction confidence scores and an OCR confidence score of the one or more OCR confidence scores, wherein the OCR confidence score corresponds to a textual region from which the extracted entity was extracted (In many embodiments, rankings may be computed based on a composite score derived from the confidence associated with each step of the process. For example, the composite score may be based on scores relating to one or more of assessed image quality, template matching distance, document structure overlap, identified metadata overlap, and overall text field accuracy, [0246], generate an overall image review priority ranking score based on an assessed image quality score generated by data adjuster 1704, a template matching confidence score generated by template manager 1836, document structure overlap score generated by pixel manager 1838, and an identified metadata overlap score generated by metadata identifier 1840, and an overall text field accuracy score generated by optical analyzer 1734, [0272], the composite score may be based on scores relating to one or more of assessed image quality, template matching distance, document structure overlap, identified metadata overlap, and overall text field accuracy, [0292]) if the combined confidence score meets the acceptance threshold: transmitting the extracted entity to a database for validation against reference data stored in the database (the document image collection 2202 may be received via a user interface provided by data interpreter 1608, [0283]); verifying existence of the extracted entity in the database (the document image 2202 may be clustered against existing templates in the template database, [0284]); comparing one or more attributes of the extracted entity against stored attributes maintained in the database (data contextualizer 1306 may compute the hamming distance between document image 2202 and one or more templates, or hashes thereof, in the template database, [0284]); and validating one or more relationships of the extracted entity with other entities in the database (At block 2210 it may be determined whether document image 2202 matches an existing template in the template database. If not, the document image 2202 is utilized at block 2216 as a bases for a new template in the template database. At block 2218 an annotation for the new template may be generated, [0284]); and if the combined confidence score does not meet the acceptance threshold: withholding the extracted entity from database validation; and initiating an error remediation process (Oftentimes confidence scores may be utilized to prioritize images for manual review, such as by triggering an exception handler. The exception handler may interrupt or divert the normal process flow of an algorithm. For example, a confidence score below a minimum threshold may generate an exception that causes the exception handler to tag the corresponding image for manual review. In some such examples, the exception handler may cause the corresponding image to be removed from or added to one or more collections and/or analyses, such as in response to user input, [0245], Similarly, block 2212 may include exception handling. Thus, if issues occur, such as due to confidence levels (e.g., for matching) being below a threshold, exception handling 2212 may be triggered. In another example, if one or more word tokens from optical character recognition are corrupted, exception handling 2212 may be triggered. Exception handling 2212 may cause user input to be requested to resolve an issue, [0285], Continuing to block 1360 instances with low confidence may be identified for review, [0293]).
Wheaton et al. do not disclose the language “if the combined confidence score meets the acceptance threshold”, however, as the error remediation process only happens if the threshold condition is not satisfied, it would have been obvious at the time of filing to one of ordinary skill in the art the steps that do not involve human intervention, such as the automatic comparing to a template, will usually happen in the cases the combined confidence score meets the acceptance threshold. Additionally it is noted as claims 1 and 20 are the method and non-transitory computer-readable storage medium respectively, in the case of a conditional statement such as “if the combined confidence score meets the acceptance threshold” and “if the combined confidence score does not meet the acceptance threshold”, only one of these limitations must be found to satisfy the claim. As the structure requires both conditions, claim 19 is interpreted more rigidly and rejected further in a separate section.
Wheaton et al. do not disclose generating a combined confidence score by multiplying an entity extraction confidence score of the one or more entity extraction confidence scores by an OCR confidence score of the one or more OCR confidence scores.
Zeng et al. teach generating a combined confidence score by multiplying an entity extraction confidence score of the one or more entity extraction confidence scores by an OCR confidence score of the one or more OCR confidence scores, wherein the OCR confidence score corresponds to a textual region from which the extracted entity was extracted (“Each of the extracted entities may have an associated confidence value, which may be a composite of multiple confidence values (e.g., confidence value=(region detection confidence)*(OCR word confidence)*(entity extraction confidence)). If the validation succeeds, then the validation module 120 may calculate the calibrated confidence values for the validated entities by updating the existing confidence values by a factor boost_weight_0. In such a case, the resulting calibrated confidence value for an entity may be (region detection confidence)*(OCR word confidence)*(entity extraction confidence)*(boost_weight_0). In general, the resulting calibrated confidence value for an entity may be (region detection confidence)*(OCR word confidence)*(entity extraction confidence)*(knowledge validation weight), where the knowledge validation weight may be boost_weight_0 or another value as described below”, [0047]).
Wheaton et al. and Zeng et al. are in the same art of document processing (Wheaton et al., abstract; Zeng et al., abstract). The combination of Zeng et al. with Wheaton et al. will enable multiplying an entity extraction confidence score. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the multiplying an entity extraction confidence score of Zeng et al. with the invention of Wheaton et al. as this was known at the time of filing, the combination would have predictable results, and as Zeng et al. state, “Uncertainty in the predictions generated by a machine learning (ML) model may be represented by confidence values. For example, a confidence value may indicate how reliable a model prediction is. Confidence calibration of a machine learning (ML) model may play an important role for many applications, such as self-driving vehicles, medical diagnosis, and human-in-the-loop systems” ([0023]) and “Examples presented herein include architectures or frameworks of document entity extraction with knowledge calibration of model confidence; methods of confidence calculation through pre-processing, OCR, entity extraction and knowledge validation; methods of using knowledge validation logics to boost or downgrade the model confidence; and methods of using knowledge-based fuzzy search to search the noisy/missing entities. Such examples may be applied to a digital document or in general to a corpus of structured or unstructured text. By utilizing the techniques presented herein, the knowledge-based validation can provide information about the model ignorance (epistemic uncertainty) to the out-of-distribution (OOD) dataset which are not covered by the training data” ([0026]) providing a means of applying the invention to a variety of different industries and a way to ensure higher accuracy through the use of the multiplied confidence values.
Regarding claim 3, Wheaton et al. and Zeng et al. disclose the method of claim 1. Wheaton et al. further indicate the error remediation process comprises: presenting the extracted entity and the combined confidence score to a human reviewer; receiving corrections or confirmations from the human reviewer; and using the corrections or confirmations as training data to update machine learning models used in the entity extraction process (This may allow document images with one or more of low-quality ratings, poor document structure, missing metadata, and a high level of textual errors to be readily identified for manual review. One or more embodiments may utilize machine learning techniques to continually improve template matching accuracy and/or computed image review priority ranking scores, [0100], In various embodiments, multiple similarity scores (e.g., scores corresponding to one or more of image quality, document structure, document metadata, and document data) may be compared/utilized to confirm that images are the same (such as by creating overall image ranking scores), an operator can be prompted to manually annotate a single image for each cluster, [0242], Several embodiments may estimate confidence scores at multiple levels (e.g., template and field level). Oftentimes confidence scores may be utilized to prioritize images for manual review, such as by triggering an exception handler, [0245], Various embodiments may include a graphical user interface (GUI) for manual review. In some embodiments, manual reviews may be presented with a GUI that prioritizes which images have the greatest need for manual review. In some such embodiments, columns can be interactively filtered and sorted. In many embodiments, rankings may be computed based on a composite score derived from the confidence associated with each step of the process. For example, the composite score may be based on scores relating to one or more of assessed image quality, template matching distance, document structure overlap, identified metadata overlap, and overall text field accuracy., [0246], overall image review priority ranking score, may receive corrections, notes, and/or feedback from users, [0262], In several embodiments, one or more confidence scores may be computed, or recomputed, at block 2220 in response to the human assisted feedback with active learning, [0285], instances with low confidence may be identified for review, At block 2364 a machine learning model may be trained to adjust confidence scores based on the review assessment. For example, an operator may manually review text output for each field, which can then be used as a target for a machine learning model to predict the likelihood of a mistake, [0293]).
Regarding claim 4, Wheaton et al. and Zeng et al. disclose the method of claim 1. Wheaton et al. further indicate tracking validation results for extracted entities over time; computing validation success rates for one or more confidence score ranges; and dynamically adjusting the acceptance threshold based on the validation success rates (In some embodiments, disposition history (e.g., previous actions and/or input in response to the same or similar circumstances) is used to adjust baseline confidence scores. For instance, some templates and fields are expected to be more accurately captured than other. In another instance, baseline confidence scores may be adjusted with the objective of reducing the amount time spent reviewing accurate extractions and increasing the amount of time spent reviewing of inaccurate extractions, [0245], if issues occur, such as due to confidence levels (e.g., for matching) being below a threshold, exception handling 2212 may be triggered, a user interface may be generated for manual review of a template, and a template confidence score corresponding to the template may be increased due to the manual review confirming the template, [0285]).
Regarding claim 8, Wheaton et al. and Zeng et al. disclose the method of claim 1. Wheaton et al. further indicate the document is received in at least one format selected from: an image format, a PDF format, and an extractable document format (extracting contextually structured data from document images, abstract, multiple similarity scores (e.g., scores corresponding to one or more of image quality, document structure, document metadata, and document data) may be compared/utilized to confirm that images are the same, [0242], incoming image, [0244], contextually structured format of the document image 2202 may be communicated, [0285]).
Regarding claim 9, Wheaton et al. and Zeng et al. disclose the method of claim 1. Wheaton et al. further indicate prior to extracting the one or more entities, performing an image preprocessing operation comprising at least one of: image enhancement, noise reduction, binarization, and deskewing (The training data can be used in its raw form for training a machine-learning model or pre-processed into another form, which can then be used for training the machine-learning model. For example, the raw form of the training data can be smoothed, truncated, aggregated, clustered, or otherwise manipulated into another form, which can then be used for training the machine-learning model, [0210], At block 2332 low-quality and nonconforming images may be filtered out. In many embodiments, block 2332 includes subblock 2332-1 for initial review and annotation. At block 2334 images may be standardized and preprocessed. In several embodiments, block 2334 includes subblock 2334-1 for resizing images, subblock 2334-2 for binarizing images with adaptive thresholding, and subblock 2334-3 for morphological transformation, [0287]).
Regarding claim 18, Wheaton et al. and Zeng et al. disclose the method of claim 1. Wheaton et al. further indicate if discrepancies between extracted entities and stored attributes exceed a predetermined threshold, automatically initiating an error remediation process comprising presenting the discrepancies to a human reviewer for manual verification (For example, a confidence score below a minimum threshold may generate an exception that causes the exception handler to tag the corresponding image for manual review., [0245], Thus, if issues occur, such as due to confidence levels (e.g., for matching) being below a threshold, exception handling 2212 may be triggered, In several embodiments, one or more confidence scores may be computed, or recomputed, at block 2220 in response to the human assisted feedback with active learning. For example, a user interface may be generated for manual review of a template, and a template confidence score corresponding to the template may be increased due to the manual review confirming the template. In some such embodiments, a blended image may be presented for manual review of the template., [0285], each text block without any metadata may be combined with the closest metadata-containing text block that appears to the north or west, combining may be subject to a threshold distance, a text block with no metadata may trigger an exception for manual review, [0320]).
Claim(s) 2 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wheaton et al. (US 20210110527 A1) and Zeng et al. (US 20220300834 A1) as applied to claim 1 above, further in view of O'Neill (US 20180285835 A1).
Regarding claim 2, Wheaton et al. and Zeng et al. disclose the method of claim 1. Wheaton et al. and Zeng et al. do not disclose the entity extraction confidence score and the OCR confidence score are each expressed as a percentage or as a decimal value between 0 and 1.
O'Neill teaches the entity extraction confidence score and the OCR confidence score are each expressed as a percentage or as a decimal value between 0 and 1 (“In an embodiment, some OCR results for a single one of the plurality of images includes identification of a single amount having an assigned confidence value based on a predefined scale that can be percentage based or based on some predefined numeric range (for example 1-10 with 10 being the highest confidence. The OCR results produce, for at least one image, identification of two or more different amounts, each amount having an assigned confidence value based on some predefined range (percentage 1-100 or 1-10, etc. with the highest confidence value and lowest confidence value being known for the predefined range)” [0057]-[0058]).
Wheaton et al. and O'Neill are in the same art of document processing (Wheaton et al., abstract; O'Neill, abstract). The combination of O'Neill with Wheaton et al. and Zeng et al. will enable confidence scores expressed as a percentage or as a decimal value between 0 and 1. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the percentage of O'Neill with the invention of Wheaton et al. and Zeng et al. as this was known at the time of filing, the combination would have predictable results, and as O'Neill state, “The different confidence percentages and the dollar amounts for each of the three checks are provided to the decision manager 130. The decision manager 130 performs the balancing operation and uses the threshold value for making decisions as to which recognized dollar amount to accept from the OCR engine 120” ([0022]) which should reduce fraud when inventions are combined.
Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wheaton et al. (US 20210110527 A1) and Zeng et al. (US 20220300834 A1) as applied to claim 1 above, further in view of Desai et al. (US 20230154222 A1).
Regarding claim 5, Wheaton et al. and Zeng et al. disclose the method of claim 1. Wheaton et al. and Zeng et al. do not disclose when the combined confidence score meets the acceptance threshold, performing a three-pronged validation process comprising: an entity existence check to determine if the extracted entity exists in the database; an entity attribution verification to check attributes of the extracted entity against attributes defined in the database; and an entity relationship validation to check whether relationships of the extracted entity with other entities are defined in the database and are accurate.
Desai et al. teach when the combined confidence score meets the acceptance threshold, performing a three-pronged validation process comprising: an entity existence check to determine if the extracted entity exists in the database (duplicate document detector, [0036]); an entity attribution verification to check attributes of the extracted entity against attributes defined in the database (comparing the values of these features in a specific record against the historic values of the features for that vendor, [0036]); and an entity relationship validation to check whether relationships of the extracted entity with other entities are defined in the database and are accurate (invoices 152, 154, are scored by the anomaly detection model 206 for identifying anomalies/errors and fraud, If the values for the current record lie outside the expected distribution of values for the feature, the feature is flagged as a likely cause of an anomaly, and tagged with a corresponding reason code. This aids, for example, in the manual validation of the record, [0036]).
Wheaton et al. and Desai et al. are in the same art of document processing (Wheaton et al., abstract; Desai et al., abstract). The combination of Desai et al. with Wheaton et al. and Zeng et al. will enable using a three-pronged validation process. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the three-pronged validation process of Desai et al. with the invention of Wheaton et al. and Zeng et al. as this was known at the time of filing, the combination would have predictable results, and as Desai et al. state, “For example, documents such as invoices, machinery or process diagrams, labels, etc., can require validation in addition to more complex validations such as validations of software systems, etc. Validation of documents such as invoices can require that the entries therein are accurate in addition to complying with the format and other requirements. Particularly, invoice validation includes a thorough review of the bills to ensure that any discrepancies are highlighted, acted upon, and rectified” ([0001]) indicating multiple commercial applications for the combination of inventions.
Claim(s) 12-13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wheaton et al. (US 20210110527 A1) and Zeng et al. (US 20220300834 A1) as applied to claim 1 above, further in view of Tasnia et al. (“Exploiting stacked embeddings with LSTM for multilingual humor and irony detection”).
Regarding claim 12, Wheaton et al. and Zeng et al. disclose the method of claim 1. Wheaton et al. and Zeng et al. do not disclose wherein the entity extraction process uses a voting ensemble comprising: one or more transformer-based models including at least one of BERT models and RoBERTa models; and one or more recurrent neural network models comprising bidirectional LSTM networks utilizing embeddings from the transformer-based models.
Tasnia et al. teach the entity extraction process uses a voting ensemble comprising: one or more transformer-based models including at least one of BERT models and RoBERTa models; and one or more recurrent neural network models comprising bidirectional LSTM networks utilizing embeddings from the transformer-based models (We combine ELMo in stacked embeddings so that we can achieve the deep contextual representation of words which in turn helps to improve the performance of humor and irony detection tasks. It works with bidirectional language models (biLM), p7, LSTM (Hochreiter and Schmidhuber 1997) stands for long short-term memory network. It is a type of recurrent neural network, Document LSTM embeddings run an LSTM-type RNN over all words of a sentence and use the final state of the networks as embeddings for the whole document, p8, In the embedding section, we deployed various embedding combinations both for word embeddings and document embeddings. We amalgamed BERT, RoBERTa, GloVe, ELMo, XLNet, GPT, Flair news-forward, and Flair newsbackward for word embeddings, p11).
Wheaton et al. and Tasnia et al. are in the same art of document processing (Wheaton et al., abstract; Tasnia et al., p4). The combination of Tasnia et al. with Wheaton et al. and Zeng et al. will enable using BERT and similar models. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the models of Tasnia et al. with the invention of Wheaton et al. and Zeng et al. as this was known at the time of filing, the combination would have predictable results, and as Tasnia et al. state, “The key contribution of this paper is that we exploit a fine-tuned stacked embeddings approach to combine various word and transformer embeddings to capture diverse word semantics in context. Utilizing LSTM architecture on word and transformer embeddings, we construct a unified document vector (UDV) that efficiently demonstrates the syntactic and semantic properties of a document to derive global context. To acquire better insights into an intermediate representation, we feed the unified document vector in multiple feed-forward linear architectures and procure our final predictions from the last layer. The utilization of lightweight features from the last layer gives better delineation of text and makes our system more robust and memory efficient” (p4) indicating an accuracy improvement and memory usage benefit when inventions are combined.
Regarding claim 13, Wheaton et al. and Zeng et al. and Gupta et al. disclose the method of claim 12. Gupta et al. further indicate the embeddings comprise at least one of: contextual embeddings from hidden layers of the transformer-based models; ELMo embeddings; Flair character-based embeddings; and stacked embeddings combining multiple embedding types (Here, we perform the stacking of various fine-tuned word embeddings and transformer models including GloVe, ELMo, BERT, and Flair’s contextual embeddings to extract the diversified contextual features of texts, abstract, For a given input text, we extract embedding features using various models including GloVe, ELMo in word embeddings, BERT in transformer word embeddings, and Flair embeddings. To generate effective word representation, we combine the extracted embedding features through the stacked capsule and feed them to an LSTM architecture to learn the context of a particular word based on all of its surroundings and thus can predict the next word accurately. BERT-base model uses 12 layers of transformer encoders. The output of each token from each layer of these encoders works as word embeddings, p7-8, The representation of the stacked embedding-based framework is depicted in Fig. 1. We unify GloVe, Flair, ELMo, and BERT embeddings utilizing the stacked embeddings approach of the FLAIR framework to concatenate different embeddings of the same lexical word, p8).
Claim(s) 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wheaton et al. (US 20210110527 A1) and Zeng et al. (US 20220300834 A1) as applied to claim 1 above, further in view of Villa-Real et al. (US 20160012445 A1).
Regarding claim 15, Wheaton et al. and Zeng et al. disclose the method of claim 1. Wheaton et al. and Zeng et al. do not disclose the database comprises a collateralized debt obligation database, and wherein validating further comprises: matching one or more extracted entity names against one or more registered entity names in the database; verifying one or more extracted numerical values against one or more transaction records stored in the database; confirming one or more extracted dates fall within one or more valid transaction periods defined in the database; and validating that relationships between entities match one or more counterparty relationships in the database.
Villa-Real et al. teach database comprises a collateralized debt obligation database, and wherein validating further comprises: matching one or more extracted entity names against one or more registered entity names in the database (verify, match, control, record and consummate the safe and legal transactions and transfers of the correct amount(s) of funds according to the relevant accurate names, [0322]); verifying one or more extracted numerical values against one or more transaction records stored in the database (check, verify, match, control, record and consummate the safe and legal transactions and transfers of the correct amount(s) of funds, [0322]); confirming one or more extracted dates fall within one or more valid transaction periods defined in the database (check, verify, match, control, record and consummate the safe and legal transactions and transfers of the correct amount(s) of funds according to the relevant accurate names, dates, [0322]); and validating that relationships between entities match one or more counterparty relationships in the database (“It is important to note and made clear here in these general diagrammatic representations that both the Customer's Credit/Debit Card Processing Means 384, and the Merchant's Credit/Debit Card Processing Means 382, shall include, (though all not shown here in these FIGS. 14 and 17) the necessary relevant corresponding networks, systems, means and secured transmissions and safely protected databases in order to accomplish accurate verification, authentication, authorization, security and settlement and other necessary relevant means and methods, aimed to securely check, verify, match, control, record and consummate the safe and legal transactions and transfers of the correct amount(s) of funds according to the relevant accurate names, dates, and other identifying means for both the authorized customers and/or authorized users and the correct merchants and/or vendors, including the correct corresponding respective legitimate banks and/or lending institutions or companies concerned, with the aim of curtailing, preventing and/or minimizing the risks and dangers of frauds and identity thefts derogatory activities, while doing personal and/or business electronic transactions in e-banking and e-commerce”, [0322]).
Wheaton et al. and Villa-Real et al. are in the same art of document processing (Wheaton et al., abstract; Villa-Real et al., [0098]). The combination of Villa-Real et al. with Wheaton et al. and Zeng et al. will enable using matching one or more counterparty relationships. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the matching of Villa-Real et al. with the invention of Wheaton et al. and Zeng et al. as this was known at the time of filing, the combination would have predictable results, and as Villa-Real et al. state, “All-in-one wireless mobile telecommunication devices, methods and systems providing greater customer-control, instant-response anti-fraud/anti-identity theft protections with instant alarm, messaging and secured true-personal identity verifications for numerous registered customers/users, with biometrics and PIN security, operating with manual, touch-screen and/or voice-controlled commands, achieving secured rapid personal/business e-banking, e-commerce, accurate transactional monetary control and management, having interactive audio-visual alarm/reminder preventing fraudulent usage of legitimate physical and/or virtual credit/debit cards, with cheques anti-forgery means, curtailing medical/health/insurance frauds/identity thefts, having integrated cellular and/or satellite telephonic/internet and multi-media means, equipped with language translations, GPS navigation with transactions tagging, currency converters, with or without NFC components, minimizing potential airport risks/mishaps, providing instant aid against school bullying, kidnapping, car-napping and other crimes, applicable for secured military/immigration/law enforcements, providing guided warning/rescue during emergencies and disasters”, (abstract), providing a security benefit to combining inventions.
Claim(s) 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wheaton et al. (US 20210110527 A1) and Zeng et al. (US 20220300834 A1) as applied to claim 1 above, further in view of Deryagin et al. (US 20130223743 A1).
Regarding claim 16, Wheaton et al. and Zeng et al. disclose the method of claim 1. Wheaton et al. and Zeng et al. do not disclose performing a source classification operation to determine a type of document from which the one or more entities are being extracted; and selecting one or more entity extraction models based on the determined document type.
Deryagin et al. teach performing a source classification operation to determine a type of document from which the one or more entities are being extracted; and selecting one or more entity extraction models based on the determined document type (Advantageously, the correct recognition of the logical structure of a document enables the system to preserve the basic layout of the source document and to classify documents according to their types, including spreadsheets, magazine articles, contracts, and even faxes. As a result, the headers and footers, page numbering, footnotes, and fonts and styles of the original are retained. For example, footnotes linked with their corresponding text on the page, image captions, graphics, and tables may be automatically grouped with the appropriate object type. Headers and footers can be directly edited or even removed using the header and footer tools provided by any text editing software, A variety of additional formatting elements, including line numbering, signatures, and stamps found in legal and other documents, may be recognized and retained, [0039], Thus, verifying each block hypothesis includes generating at least one block hypothesis for each block on the page based on the document hypothesis and selecting a best block hypothesis for each block. The best block hypothesis is selected on the basis of estimation of correspondence parameters of the block to the block model and the model of the whole document. The decision about confirmation the document model hypothesis on a page is made on the basis of estimation of correspondence parameters of the page (blocks on the page) to the selected model of the whole document. If such estimation is satisfied, it considered as confirmation (105) of the selected hypothesis of the whole document on this page, and the system can go to verifying the document hypothesis on the next page, [0059], Selecting, at block 106, one or more best document models executed on the basis of confirmation one of more hypotheses. The document hypothesis that correlates best with the entire document is selected as the best model of the document. In one embodiment, the best model may be selected automatically by the OCR system. In another embodiment, the best model may be selected manually by the user from among several models. For a manual selection, options are shown on a user interface and a user may make a selection through the user interface, [0061]).
Wheaton et al. and Deryagin et al. are in the same art of document processing (Wheaton et al., abstract; Deryagin et al., abstract). The combination of Deryagin et al. with Wheaton et al. and Zeng et al. will enable performing a source classification operation to determine a document type for each document of the plurality of documents. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the source classification of Deryagin et al. with the invention of Wheaton et al. and Zeng et al. as this was known at the time of filing, the combination would have predictable results, and as Deryagin et al. state, “Advantageously, the correct recognition of the logical structure of a document enables the system to preserve the basic layout of the source document and to classify documents according to their types, including spreadsheets, magazine articles, contracts, and even faxes. As a result, the headers and footers, page numbering, footnotes, and fonts and styles of the original are retained” ([0039]), thereby improving the accuracy of the final extraction when the inventions are combined.
Regarding claim 17, Wheaton et al. and Zeng et al. and Deryagin et al. disclose the method of claim 16. Deryagin et al. further indicate the source classification operation comprises a multi-modal source classification operation that analyzes at least two of: text content of the document, document structure, and visual features of the document (FIG. 1 of the drawings shows a flowchart of steps describing the process to recognize the logical structure of a document and select its model, in accordance with one embodiment of the invention. Referring to FIG. 1, at block 100 a document image is acquired, e.g. from an imaging device. At block 102, by means of an OCR software or function, a preliminary analysis of the physical structure of the document is executed, and in particular, at least blocks, e.g. footers, headers, are detected. The blocks may comprise text, pictures, tables, etc. In one embodiment, text occurring in the block may be clustered based on the properties of its font, i.e., a font which is only slightly different from the main font. A different font in the document may be the result of incorrect OCR processing, and may also be considered as if the different font were of a main font or same font as other parts of the document. Next, at block 103, at least one document hypothesis about possible logical structure of the whole document is generated. The document hypotheses are generated on the basis of a collection of models 120 of possible document logical structures. In one embodiment, the collection of models 120 of possible logical structures may includes models of different documents, for example, a research paper, a patent, a patent application, business letter, a contract, an agreement, etc. Each model may describe a set of essential and possible elements of logical structure and their mutual arrangement within the model. In one embodiment, for example, one of possible models of a research paper may include a title, an authors information, an abstract, an issue name, an issue number, and an issue date within page footer or page header, tables, pictures, diagrams, endnotes and footnotes, bibliography, flowcharts and other, [0041]-[0042]).
The limitation “the source classification operation comprises a multi-modal source classification operation that analyzes at least two of: text content of the document, document structure, and visual features of the document” is interpreted in the conjunctive, in accordance with SuperguideCorp. v. DirecTV Enter., Inc., 358 F.3d 870, 875 (Fed. Cir. 2004) in which the Federal Circuit held that the plain meaning of “at least one of A, B, and C” means: at least one A, at least one of B and at least one of C.
Claim(s) 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wheaton et al. (US 20210110527 A1) in view of Desai et al. (US 20230154222 A1) in view of Zeng et al. (US 20220300834 A1).
Regarding claim 19, Wheaton et al. disclose a system for validating extracted document data, comprising: a processor; a memory coupled to the processor; and a machine learning kernel stored in the memory and configured to cause the processor to (memory, processor, [0117], machine learning models, kernel [0206]): perform optical character recognition on a document (identifying semi-structured data generated by optical character recognition, [0033], text identified by an OCR processes may be provided along with the corresponding portion of the original image, [0275]) to generate one or more OCR confidence scores for textual content of the document (confidence score associated with overall text field accuracy, [0246], OCR text accuracy score, [0273]) extract one or more entities from the document using an entity extraction process (extracting contextually structured data from document images, abstract, [0085], extract document image contents into a contextually structured format, [0282]) generate one or more entity extraction confidence scores for the one or more extracted entities (estimate confidence scores at multiple levels (e.g., template and field level), [0245], confidence associated with each step, based on scores relating to one or more of assessed image quality, template matching distance, document structure overlap, identified metadata overlap, and overall text field accuracy, [0246]) for each extracted entity generate a combined confidence score by combining an entity extraction confidence score of the one or more entity extraction confidence scores and an OCR confidence score of the one or more OCR confidence scores, wherein the OCR confidence score corresponds to a textual region from which the extracted entity was extracted (In many embodiments, rankings may be computed based on a composite score derived from the confidence associated with each step of the process. For example, the composite score may be based on scores relating to one or more of assessed image quality, template matching distance, document structure overlap, identified metadata overlap, and overall text field accuracy, [0246], generate an overall image review priority ranking score based on an assessed image quality score generated by data adjuster 1704, a template matching confidence score generated by template manager 1836, document structure overlap score generated by pixel manager 1838, and an identified metadata overlap score generated by metadata identifier 1840, and an overall text field accuracy score generated by optical analyzer 1734, [0272], the composite score may be based on scores relating to one or more of assessed image quality, template matching distance, document structure overlap, identified metadata overlap, and overall text field accuracy, [0292]) compare the combined confidence score to an acceptance threshold; if the combined confidence score meets the acceptance threshold: transmit the extracted entity to a database for validation against reference data stored in the database (the document image collection 2202 may be received via a user interface provided by data interpreter 1608, [0283]); verify existence of the extracted entity in the database (the document image 2202 may be clustered against existing templates in the template database, [0284]); compare one or more attributes of the extracted entity against stored attributes maintained in the database (data contextualizer 1306 may compute the hamming distance between document image 2202 and one or more templates, or hashes thereof, in the template database, [0284]); and validate one or more relationships of the extracted entity with other entities in the database (At block 2210 it may be determined whether document image 2202 matches an existing template in the template database. If not, the document image 2202 is utilized at block 2216 as a bases for a new template in the template database. At block 2218 an annotation for the new template may be generated, [0284]); and if the combined confidence score does not meet the acceptance threshold: withhold the extracted entity from database validation; and initiate an error remediation process (Oftentimes confidence scores may be utilized to prioritize images for manual review, such as by triggering an exception handler. The exception handler may interrupt or divert the normal process flow of an algorithm. For example, a confidence score below a minimum threshold may generate an exception that causes the exception handler to tag the corresponding image for manual review. In some such examples, the exception handler may cause the corresponding image to be removed from or added to one or more collections and/or analyses, such as in response to user input, [0245], Similarly, block 2212 may include exception handling. Thus, if issues occur, such as due to confidence levels (e.g., for matching) being below a threshold, exception handling 2212 may be triggered. In another example, if one or more word tokens from optical character recognition are corrupted, exception handling 2212 may be triggered. Exception handling 2212 may cause user input to be requested to resolve an issue, [0285], Continuing to block 1360 instances with low confidence may be identified for review, [0293]).
Wheaton et al. do not disclose generating a combined confidence score by multiplying an entity extraction confidence score of the one or more entity extraction confidence scores by an OCR confidence score of the one or more OCR confidence scores. Wheaton et al. do not explicitly disclose if the combined confidence score meets the acceptance threshold: transmit the extracted entity to a database for validation against reference data stored in the database.
Desai et al. teach if the combined confidence score meets the acceptance threshold (an aggregate fault score can be calculated, “The AI-based fault processor 104 can be configured with a threshold score below which an invoice is considered as a valid invoice and is allowed for further processing by the action optimizer 106. Invoices having scores above the thresholds scores are flagged for further review. Upon completion of the further review, a subset of the flagged invoices may be considered as valid and can be allowed for further processing by the action optimizer 106”, [0022], invoices associated with higher amounts and higher probabilities of errors, frauds, and/or duplicates are provided for further processing while the remaining valid invoices are provided to the action optimizer 106 for scheduling the automatic actions such as the automatic payments, [0028]): transmitting the extracted entity to a database for validation against reference data stored in the database (remaining valid invoices are provided to the action optimizer 106, [0028]); verifying existence of the extracted entity in the database (duplicate document detector 144, [0022], date-document selector 402 selects as input, the open/paid invoices from the document package 150 falling due within the specific time window, along with the payment term metadata and the scores or the anomaly probabilities 258 and the duplicate probabilities 358 of each of the selected invoices, [0030]); comparing one or more attributes of the extracted entity against stored attributes maintained in the database (feature extractor 304 can be trained to automatically extract features 356 for each of the sets based on set properties, invoice properties, and historical profiles, etc., [0028]); and validating one or more relationships of the extracted entity with other entities in the database (The data storage 170 can be used for storing data that is generated and used during the various validation and automatic execution processes, [0020], remaining valid invoices are provided to the action optimizer 106 for scheduling the automatic actions such as the automatic payments, [0028]).
Wheaton et al. and Desai et al. are in the same art of document processing (Wheaton et al., abstract; Desai et al., abstract). The combination of Desai et al. with Wheaton et al. will enable performing actions if the combined confidence score meets the acceptance threshold. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the confidence threshold of Desai et al. with the invention of Wheaton et al. as this was known at the time of filing, the combination would have predictable results, and as Desai et al. state, “For example, documents such as invoices, machinery or process diagrams, labels, etc., can require validation in addition to more complex validations such as validations of software systems, etc. Validation of documents such as invoices can require that the entries therein are accurate in addition to complying with the format and other requirements. Particularly, invoice validation includes a thorough review of the bills to ensure that any discrepancies are highlighted, acted upon, and rectified” ([0001]) indicating multiple commercial applications for the combination of inventions.
Wheaton et al. and Desai et al. do not disclose generating a combined confidence score by multiplying an entity extraction confidence score of the one or more entity extraction confidence scores by an OCR confidence score of the one or more OCR confidence scores.
Zeng et al. teach generating a combined confidence score by multiplying an entity extraction confidence score of the one or more entity extraction confidence scores by an OCR confidence score of the one or more OCR confidence scores, wherein the OCR confidence score corresponds to a textual region from which the extracted entity was extracted (“Each of the extracted entities may have an associated confidence value, which may be a composite of multiple confidence values (e.g., confidence value=(region detection confidence)*(OCR word confidence)*(entity extraction confidence)). If the validation succeeds, then the validation module 120 may calculate the calibrated confidence values for the validated entities by updating the existing confidence values by a factor boost_weight_0. In such a case, the resulting calibrated confidence value for an entity may be (region detection confidence)*(OCR word confidence)*(entity extraction confidence)*(boost_weight_0). In general, the resulting calibrated confidence value for an entity may be (region detection confidence)*(OCR word confidence)*(entity extraction confidence)*(knowledge validation weight), where the knowledge validation weight may be boost_weight_0 or another value as described below”, [0047]).
Wheaton et al. and Zeng et al. are in the same art of document processing (Wheaton et al., abstract; Zeng et al., abstract). The combination of Zeng et al. with Wheaton et al. and Desai et al. will enable multiplying an entity extraction confidence score. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the multiplying an entity extraction confidence score of Zeng et al. with the invention of Wheaton et al. and Desai et al. as this was known at the time of filing, the combination would have predictable results, and as Zeng et al. state, “Uncertainty in the predictions generated by a machine learning (ML) model may be represented by confidence values. For example, a confidence value may indicate how reliable a model prediction is. Confidence calibration of a machine learning (ML) model may play an important role for many applications, such as self-driving vehicles, medical diagnosis, and human-in-the-loop systems” ([0023]) and “Examples presented herein include architectures or frameworks of document entity extraction with knowledge calibration of model confidence; methods of confidence calculation through pre-processing, OCR, entity extraction and knowledge validation; methods of using knowledge validation logics to boost or downgrade the model confidence; and methods of using knowledge-based fuzzy search to search the noisy/missing entities. Such examples may be applied to a digital document or in general to a corpus of structured or unstructured text. By utilizing the techniques presented herein, the knowledge-based validation can provide information about the model ignorance (epistemic uncertainty) to the out-of-distribution (OOD) dataset which are not covered by the training data” ([0026]) providing a means of applying the invention to a variety of different industries and a way to ensure higher accuracy through the use of the multiplied confidence values.
Allowable Subject Matter
Claims 6, 7, 10, 11, and 14 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Examiner notes that while the application does not currently have a double patenting rejection in view of 19/396753 or 19/441538, applicant is advised to be mindful when amending to not create a double patenting rejection in view of those applications. Note on interpretation of claim 14: The limitation “wherein the document comprises a loan notice document, and wherein the extracted entities comprise at least one of: borrower name, lender name, lender address, borrower address, loan amount, interest rate, maturity date, effective date, origination date, facility name, deal name, reference rate, reference rate source, governing law state, spread margin, and day count convention” is interpreted in the conjunctive, in accordance with SuperguideCorp. v. DirecTV Enter., Inc., 358 F.3d 870, 875 (Fed. Cir. 2004) in which the Federal Circuit held that the plain meaning of “at least one of A, B, and C” means: at least one A, at least one of B and at least one of C.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHELLE ENTEZARI whose telephone number is (571)270-5084. The examiner can normally be reached 10-7 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent M Rudolph can be reached at (571) 272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHELLE M ENTEZARI HAUSMANN/Primary Examiner, Art Unit 2671