DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Introductory Remarks
In response to communications filed on 3 June 2026, claims 1 and 11 are amended per Applicant's request. No claims were cancelled. No claims were withdrawn. No new claims were added. Therefore, claims 1-20 are presently pending in the application, of which claims 1 and 11 are presented in independent form.
The previously raised 103 rejection of the pending claims is withdrawn in view of the amendments to the claims. A new ground(s) of rejection has been issued.
Response to Arguments
Applicant’s arguments filed 3 June 2026 with respect to the rejection of the claims under 35 U.S.C. 103 (see Remarks, p. 1-9) have been fully considered but are not persuasive. Applicant argues solely that the prior art references do not teach, suggest, or otherwise disclose/render obvious the amended claim language. The examiner respectfully disagrees and the rejections have been modified to conform to the current amended claim language.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
Independent Claims 1 and 11 recite “identifying the one or more discrepant data fields comprises mapping each first data field of the first data structure to a second data field of the second data structure using configuration files that specify which fields are compared, by:” (performing the transformation, determination and assigning weight, comparing weight steps, etc.).
The closest paragraph appears to be Specification, [0024-0025], where in Specification, [0024], where “in some embodiments”, the system maps first and second data fields using predefined mapping logic, the mapping logic implemented using configuration fields or schema definitions that explicitly specify which fields should be compared, e.g., “Diagnosis Code” in primary data 120 is defined in the mapping configuration to correspond to the “Required Diagnosis Code” in secondary data 136. It appears, however, that this explicit mapping is separate from, e.g., the processes of Specification, [0025], which, “in some embodiments”, uses an encoder by transforming the data into vector representations to identify discrepant data fields (i.e., which are claimed via performing the disclosed transformation, determination and assigning of weights, comparing weights, etc.).
In other words, Applicant is improperly combining the steps disclosed by Specification, [0024], with the steps disclosed in Specification, [0025], the latter of which pertains to a different process. Thus, the mapping using a configuration file is not performed “by” using the disclosed encoder process, but rather they describe two different embodiments. This also raises the question why there would need to be the steps of Specification, [0025] within the context of Specification, [0024], as Specification, [0024] describes predefined mapping using configuration files or schema definitions, whereas Specification, [0025] pertains to identifying what the mapping was meant to be, i.e., no configuration files or schema definitions present or needed. Thus, it does not appear that the claimed invention intended to have the two combined in such a manner, which is further evidenced via the use of the language “in some embodiments” in both Specification, [0024], and separately in Specification, [0025], not “Additionally”, or in any way linking the process of Specification, [0024] to Specification, [0025].
Lastly, it is unclear if the limitation “comparing document data in the mapped first data field with predefined documentation requirements in the second data field and identifying the one or more discrepant data fields when the correspondence score fails to satisfy a similarity criterion” is meant to refer to that both conditions must be met for discrepancies to be identified, or only one, depending on whether the process specified in Specification, [0024] or Specification, [0025], was utilized. However, given that it appears to indicate that both criteria must be satisfied, there is a lack of support for such a limitation (the first part occurring in the process(es) described in Specification, [0024], the latter occurring in the process(es) described in Specification, [0025], but not together).
The rest of the dependent claims are rejected for at least by virtue of their dependency on their respective independent claims.
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Independent Claims 1 and 11 recite “identifying the one or more discrepant data fields comprises mapping each first data field of the first data structure to a second data field of the second data structure using configuration files that specify which fields are compared, by:” (performing the transformation, determination and assigning weight, comparing weight steps, etc.).
As noted in the 112(a), lack of written description rejection above, there is no support that the process disclosed in Specification, [0024] is performed by the steps of Specification, [0025]. Therefore, it is unclear what is meant by such a limitation, and furthermore, there is no relationship between the configuration files specifying which fields are compared to using the encoder process, which appears to infer such information (thus making it confusing as to why there would be an explicit mapping and then an inferred relationship established between the two fields via the encoder process). In other words, Specification, [0025] appears to be for situations/contexts in which there is no predefined, explicit mapping, whereas Specification, [0024] has predefined mapping. For purposes of examination, the interpretation that the configuration file specifies which fields are to be extracted/compared has been taken.
Additionally, it is unclear if the limitation “comparing document data in the mapped first data field with predefined documentation requirements in the second data field and identifying the one or more discrepant data fields when the correspondence score fails to satisfy a similarity criterion” is meant to refer to that both conditions must be met for discrepancies to be identified, or only one, depending on whether the process specified in Specification, [0024] or Specification, [0025], was utilized. It is unclear the relationship between the two with respect to identifying discrepant data fields, as the two refer to two separate processes described in Specification, [0024] and Specification, [0025].
For purposes of examination, the interpretation that both criteria are satisfied as separate processes, has been taken.
The dependent claims are rejected for at least by virtue of their dependency on their respective independent claims, and for failing to cure the deficiencies of their respective independent claims.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-5, 7-8, 11-15, and 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Priestas et al. (“Priestas”) (US 2020/0364404 A1, incorporating by reference Tucker et al. (“IBR-Tucker”) (App. No. 16/531,848 published as US 2019/0354720 A1), Priestas et al. (“IBR-Priestas”) (App. No. 16/179,448 published as US 2019/0087395 A1) and Sacaleanu et al. (“IBR-Sacaleanu”) (App. No. 15/879,031 published as US 2019/0006027 A1) at [0001]), in view of Sathyanarayana et al. (“Sathyanarayana”) (US 2012/0102002 A1), in further view of Srinivasan et al. (“Srinivasan”) (US 2025/0111169 A1).
Regarding claim 1: Priestas teaches An apparatus for resolving data discrepancies between a first data structure and a second data structure, the apparatus comprising:
at least a processor; a memory communicatively connected to the at least a processor, wherein the memory contains instructions configuring the at least a processor to (Priestas, [0078], where the disclosed system may be implemented as software stored on a non-transitory computer readable medium with processor-executable instructions executed by one or more processors):
access a first data structure, wherein the first data structure comprises a plurality of first data fields representing primary data, wherein the primary data comprises document data; access a second data structure, wherein the second data structure comprises a plurality of second data fields representing secondary data (IBR-Tucker, [0046-0047], where the system compares values in the documents from the request (i.e., “access a first data structure”) with the information in the domain model 104. External knowledge base 108 (i.e., “second [source]”) from one or more other/external data sources can be accessed by the document processing system to identify the discrepancies, and one or more fields can be flagged for human validation which can be executed via one or more GUIs 140. Note that the required documents in different data formats can include text files, documents with structured data, and/or database files (i.e., “data structure”). See also IBR-Tucker, [0055], where the documents employed by the document comparator 304 for comparisons may include structured data such as values from a database.
See also Priestas, [0021], where the request 102 can include a message 104 with certain content and may optionally include one or more documents 106 associated with the information conveyed in the message 104, where the message 104 and the documents 106 can include certain textual content of a plurality of information types/structures (i.e., “first data structure”). See also Priestas, [0022], where the external data source that is accessed by the data processing system 100 can pertain to databases (i.e., “second data structure”));
analyze, using a language processing module, the document data of the primary data representing the plurality of first data fields by:
extracting textual data from the document data by extracting one or more words represented by strings of one or more characters; parsing the strings of the textual data into at least a token including sequences of characters (Priestas, [0029], where the parser 204 parses the text included in one or more of the message 104 and the documents 106 of the request 102. The tokenizer 206 can produce word tokens from the output of the parser 204); and
detecting associations between the one or more words extracted from the document data (IBR-Sacaleanu, [0038] and [0052-0055], where the system automatically identifies and extracts entities and relations between entities within the text of each of the multiple documents (i.e., “detecting associations between the one or more words extracted from the document data”), and may apply reasoning techniques over multiple knowledge sources, e.g., including medical ontologies 212 and knowledge graphs or databases 214 to infer condition-evidence linking (i.e., “using a language processing model”), and may further score and rank the extracted entities and condition-evidence links to generate a most-representative set of entities and condition-evidence links (i.e., “produce ranked associations”). Note that the extracted entities, prior to linking medical condition entities to relevant supporting evidence entities, may categorize or label extracted entities (i.e., “using at least a classification model to process the parsed textual data”) (IBR-Sacaleanu, [0053-0054]));
… identify one or more discrepant data fields in the plurality of first data fields by comparing the first data structure, including the document data analyzed by the language processing module, to the second data structure, wherein: identifying the one or more discrepant data fields comprises mapping each first data field of the first data structure to a second data field of the second data structure using configuration files that specify which fields are compared … and the one or more discrepant data fields is identified when the mapped first data field of the plurality of first data fields does not match the corresponding mapped data field of the plurality of second data fields (IBR-Tucker, [0047-0048], where a discrepancy processor 112 determines or identifies discrepancies between compared documents having different data formats, where when a discrepancy is identified, the discrepancy processor 112 can analyze the reason for the discrepancy include identifying those data fields, e.g., field names, where the comparisons failed to produce a positive result (i.e., “the one or more discrepant data field is identified when the mapped first data field of the plurality of first data fields does not match the corresponding mapped data field of the plurality of second data fields”), where various data models can be employed for comparing the data fields/types.
See also, e.g., IBR-Tucker, [0055-0056], where for a given field including a name-value pair extracted from one or more documents 154, 156 (i.e., “mapping a first data field of the first data structure…”), the similarities between the data extracted from the documents 154, 156 and the values as specified by one or more of the process rules 322 (i.e. “to a corresponding data field from the plurality of second data fields”) can be compared to similarity thresholds, e.g., extracted fields can be compared to the fields specified by the process rules 322 (i.e., “specify[ing] which fields are compared”), comparison on name-value pairs are determined to correspond to those as specified in the process rules, etc.
See IBR-Tucker, [0055], [0062] and [0065], where when comparing information from documents in a patient’s file history to a list of documents as specified by the process rules 322, fields may be extracted and compared to the fields specified by the process rules 322, where a domain model can be employed. More specifically, a domain model 104 (i.e., corresponding to the claimed “configuration file”) supplies information or fields necessary for the automatic execution of the process, and the fields necessary for automatic execution of the process are subsequently extracted from the selected input data.
Although IBR-Tucker does not appear to explicitly state that the domain model 104 is a “file” as claimed, or that it explicitly maps a field to a field, one of ordinary skill in the art would have found it obvious to have (1) utilized a “file” for IBR-Tucker’s domain model with the motivation of simplicity in storing such models and ease of transmitting/transferring due to low overhead, low complexity, and portability; and (2) utilized field-to-field mapping with the motivation of faster lookup and comparisons (e.g., mapping the “SSN” field to “Social Security Number” is faster to look up and use than performing the various computational analysis before the system makes such a determination).
See Priestas, [0029], where the parser 204 parses documents 106 of the request 102, and the tokenizer 206 can produce word tokens from the output of the parser 204. The tokens allow the process analyzer 124 to identify the automatic document processing task 112 to be executed (see also IBR-Tucker, [0052-0053]). See IBR-Tucker, [0062-0063], where in an example document processing method, a request is analyzed using textual processing techniques, and fields necessary for the automatic execution of the process are extracted from the selected input data based on the intent derived from the textual processing technique (see also, e.g., IBR-Tucker, [0052-0053]). Note that textual processing includes parsing, tokenization, and natural language processing (IBR-Tucker, [0043]). The system determines whether the fields are valid, i.e., no discrepancies exist, in which case a domain model 104 (which was selected/identified via the textual processing technique) can be used to validate the fields (IBR-Tucker, [0043]));
generate an alteration datum as a function of the one or more discrepant data fields and the secondary data in the corresponding data field (IBR-Tucker, [0049], where the mismatched/unmatched fields from the discrepancy processor 112 can be communicated to a data resolver 114 for an intelligent resolution, where the data resolver 114 can automatically identify a resolution to the discrepancy, e.g., by access alternative formats, synonyms, etc., to automatically resolve the discrepancies (IBR-Tucker, [0057]). See IBR-Tucker, [0064], where discrepancies may be resolved based on data from one or more of the intent 164, the domain model 104, and the external knowledge base 108. In addition, the resolution of the discrepancies can require human intervention, i.e., user edits for resolving discrepancies.
In addition, the system may automatically update information, e.g., IBR-Tucker, [0074], where the fields for processing forms are obtained at 1104, where the fields are compared against member information associated with the healthcare plan at 1106, e.g., in the external knowledge base 108. Based on the comparison, an add, update, or delete operation on member information can be identified at 1108. At 1110, discrepancies or errors, if any, are identified and resolved, e.g., if updates to member information indicate a change in a social security number, it may indicate an error which needs to be resolved, where automatic resolution routines may be executed to identify similar information from multiple other resources and may automatically update the SSN information of the member. Alternately, human intervention may be sought to fix the error) … .
Priestas does not appear to explicitly teach an encoder component comprising a transformer architecture configured to use: a self-attention mechanism for detecting the associations between the one or more words extracted from the document data; and a positional encoding mechanism for encoding a position of an extracted word in a sequence of the textual data, wherein the encoder component is configured to [identify one or more discrepant data fields]; [identifying the one or more discrepant data fields] by: transforming at least the textual data of the document data analyzed by the language processing module into a vector format; dynamically determining and assigning weights, using an attention mechanism of the encoder component, to different vectorized tokens of the at least a token of the parsed textual data; comparing the weighted vectorized tokens of the primary data to the secondary data using a similarity function operating on attention-weighted vector representations to compute a correspondence score between the mapped first data field and the corresponding data field, wherein the correspondence score comprises a cosine similarity between the attention-weighted vector representations, the cosine similarity being computed using a dot product of the attention-weighted vector representations normalized by lengths of the attention-weighted vector representations; comparing document data in the mapped first data field with predefined documentation requirements in the second data field and identifying the one or more discrepant data fields when the correspondence score fails to satisfy a similarity criterion; [and] updat[ing] the one or more discrepant data fields of the plurality of first data fields in the first data structure as a function of the alteration datum, wherein the alteration datum modifies the document data of the one or more discrepant data fields to conform with the secondary data of the corresponding data field.
Sathyanarayana teaches updat[ing] the one or more discrepant data fields of the plurality of first data fields in the first data structure as a function of the alteration datum, wherein the alteration datum modifies the document data of the one or more discrepant data fields to conform with the secondary data of the corresponding data field (Sathyanarayana, [0035-0036], where the data manager extracts data from a document, database, or other data source. The data manager extracts contextual information and/or business rules, and can build an index or database to organize identified contextual information. After data extraction, the data manager refines or further processes data elements among extracted data, and identifies and refines specific types of data elements for generating similar data elements (i.e., “data field”). See Sathyanarayana, [0031], where the data manager identifies one or more anomalies from a given data set using both contextual information and validation rules (i.e., “secondary data”), and then automatically corrects any identified anomalies or missing information, e.g., involving business rules compliance, data formatting information, standardization, correcting format, etc. (Sathyanarayana, [0032]) (i.e., “to conform with the secondary data of the corresponding data field”), e.g., via identifying a unique best weight of potential corrections (see, e.g., Sathyanarayana, [0038]). See Sathyanarayana, [0039], where in response to identifying a unique best weight, the data manager can automatically correct and validate the associated data element including updating the received document or source from which the data element was extracted (i.e., “update the one or more discrepant data fields of the plurality of first data fields in the first [source data] as a function of the alteration datum”), where data validation can include all of the sub steps of indicating that data is valid, correcting incorrect data, and completing missing values).
It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to have combined the teachings of Priestas and Sathyanarayana (hereinafter “Priestas as modified”), Priestas discloses that changes to data are enabled, e.g., user edits may be performed. Therefore, one of ordinary skill in the art would have found it obvious to have combined the updating steps of Sathyanarayana with Priestas with the motivation of generating consistent data across data sources, e.g., persisting changes that were inputted (thereby enabling data to be saved for later retrieval).
Priestas as modified does not appear to explicitly teach an encoder component comprising a transformer architecture configured to use: a self-attention mechanism for detecting the associations between the one or more words extracted from the document data; and a positional encoding mechanism for encoding a position of an extracted word in a sequence of the textual data, wherein the encoder component is configured to [identify one or more discrepant data fields]; [identifying the one or more discrepant data fields] by: transforming at least the textual data of the document data analyzed by the language processing module into a vector format; dynamically determining and assigning weights, using an attention mechanism of the encoder component, to different vectorized tokens of the at least a token of the parsed textual data; comparing the weighted vectorized tokens of the primary data to the secondary data using a similarity function operating on attention-weighted vector representations to compute a correspondence score between the mapped first data field and the corresponding data field, wherein the correspondence score comprises a cosine similarity between the attention-weighted vector representations, the cosine similarity being computed using a dot product of the attention-weighted vector representations normalized by lengths of the attention-weighted vector representations; [and] comparing document data in the mapped first data field with predefined documentation requirements in the second data field and identifying the one or more discrepant data fields when the correspondence score fails to satisfy a similarity criterion.
Srinivasan teaches an encoder component comprising a transformer architecture (Srinivasan, [0034], where the architecture of the disclosed LLMs commonly includes multiple layers of neural networks and may include a transformer architecture) configured to use:
a self-attention mechanism for detecting the associations between the one or more words extracted from the document data (Srinivasan, [0043] and [0056], where encoder instructions understand and extract relevant information from the interim output and outputs (as an encoder output) a continuous representation or embedding of the interim output (using self-attention heads or layers different types of relationships and dependencies in the interim output), for processing by decoder instructions. Stated differently, the encoder instructions can capture contextual relationships between different portions of the words and generate an attention vector for each portion. See Priestas above with regards to the “document” data); and
a positional encoding mechanism for encoding a position of an extracted word in a sequence of the textual data (Srinivasan, [0042], where positional encodings (comprising positional encoding vectors for each tokenized set of information, e.g., words) are added to vectorized format token sequences, to define the relative positions of the corresponding information in the input data), wherein the encoder component is configured to [identify one or more discrepant data fields] (Srinivasan, [0059], where the system compares embedding vectors generated from selected pair of responses to determine a level of agreement or disagreement between the responses and source LLMs. See Priestas above with regards to the “discrepant data fields”, as claimed);
[identifying the one or more discrepant data fields] by: transforming at least the textual data of the document data analyzed by the language processing module into a vector format (Srinivasan, [0042], where the input data is transformed into individual word tokens, and converted into input embeddings or high-dimensional vectors);
dynamically determining and assigning weights, using an attention mechanism of the encoder component, to different vectorized tokens of the at least a token of the parsed textual data (Srinivasan, [0070], where the attention mechanism calculates “soft” weights for each token, more precisely for its embedding, by using multiple attention heads, each with its own “relevance” for calculating its own soft weights);
comparing the weighted vectorized tokens of the primary data to the secondary data using a similarity function operating on attention-weighted vector representations to compute a correspondence score between the mapped first data field and the corresponding data field, wherein the correspondence score comprises a cosine similarity between the attention-weighted vector representations, the cosine similarity being computed using a dot product of the attention-weighted vector representations normalized by lengths of the attention-weighted vector representations (Srinivasan, [0059-0060], where the system compares embedding vectors (i.e., “operating on attention-weighted vector representations”1) generated from selected pair of responses to determine a level of agreement or disagreement between the responses and source LLMs. More specifically, the vector embeddings of an input are compared to a reference input, e.g., using cosine similarity scoring, where cosine similarity can be used to compare the vector embeddings of an input to that of a reference input. See Priestas above with regards to the “primary data”, “secondary data”, “first data field” and “corresponding data field” as claimed.
Note that one of ordinary skill in the art would have recognized that the claimed limitation of “the cosine similarity being computed using a dot product of the attention-weighted vector representations normalized by lengths of the attention-weighted vector representations” is, by definition, how cosine similarity is computed, i.e., the dot product of normalized vectors is the cosine similarity, where the numerator is the dot product of both vectors, and the denominator (i.e., “normalized by”) containing the product of their magnitude, i.e., their length (i.e., “lengths of the attention-weighted vector representation”)2,3); [and]
comparing document data in the mapped first data field with predefined documentation requirements in the second data field and identifying the one or more discrepant data fields when the correspondence score fails to satisfy a similarity criterion (Srinivasan, [0059-0060], where the system can compare and assign a level of pairwise similarity or dissimilarity between selected pairs of responses, e.g., assigning a score, by comparing the embedding vectors between pairs of responses to determine a level of agreement or disagreement between the responses and source LLMs (i.e., “identifying one or more discrepant [responses using] the correspondence score”). See also IBR-Tucker, [0055-0056], where for a given field including a name-value pair extracted from one or more documents 154, 156, the similarities between the data extracted from the documents 154, 156 and the values as specified by one or more of the process rules 322 can be compared to similarity thresholds (i.e., “identifying the one or more discrepant data fields when the [value] fails to satisfy a similarity criterion”). See IBR-Tucker, [0055-0057], [0062-0063], [0080-0082], where extracted fields are compared to fields specified by process rules 322. More specifically, process rules define particular discrepancies that can be raised based on various field mismatches likely to occur, such as certain formats, values specified by the process rules including a mismatch of provider name such as “John Doe” versus “J. Doe”, etc. (thus, the process rules corresponding to the claimed “predefined documentation requirements”)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Priestas as modified and Srinivasan (hereinafter “Priestas as modified”) with the motivation of (1) utilizing a transformer architecture for faster training (Srinivasan, [0034]), (2) improving the accuracy of the output and therefore classification of the set of tokens4, (3) capturing feature representations, patterns, and relationships within the input data5, (4) enabling the model to focus on the relevant parts of the input sequence when processing each token by calculating attention weights6, and (5) determining how similar two vectors are irrespective of their size by using cosine similarity (i.e., via the normalization aspect), thereby enabling the similarities between two vectors to be compared based on their content, not based on their relative sizes.
Furthermore, it would have been obvious to one of ordinary skill in the art to have substituted Priestas’ values (used to compare against similarity thresholds) with Srinivasan’s correspondence score with the motivation of quantifying the degree/level of similarity/dissimilarity between two values (thereby enabling the system some manner of being able to make quantifiable, objective comparisons/determinations with some degree of flexibility, as opposed to binary and/or subjective determinations).
Regarding claim 2: Priestas as modified teaches The apparatus of claim 1, wherein the secondary data comprises an output of a requirement machine-learning model, wherein the requirement machine-learning model is configured to generate the secondary data (IBR-Tucker, [0054], where the document selector 302 enables selection of documents to be compared in order to enable automated execution of the process pertaining to the request 152. The document selection is enabled by the external knowledge base 108 and domain model 104, where process rules 322 included in the external knowledge base 108 enable identification of the relevant documents needed for the automatic execution of the process. See IBR-Tucker, [0047], where external knowledge base 108 can include explicit knowledge such as rules, inputs from subject matter experts (SMEs), machine-generated inputs generated using machine learning (i.e., “an output of a…machine-learning model”). See also Priestas, [0018] and [0020], where guidelines for executing a specific automatic document processing task are retrieved from one or more external data sources, where the guidelines can include requirement such as data requirements. The disclosed system processes involve extracting data from data sources having a plurality of formats to meet requirements in guidelines for complex processes).
Although Priestas does not appear to explicitly state that the machine-learning model is a requirement machine learning model as claimed, Priestas suggests that requirements/guidelines are part of the process of performing automatic document processing tasks. Therefore, one of ordinary skill in the art would have been suggested to have modified IBR-Tucker’s disclosed machine learning models to Priestas’ plurality of ML models that are selected and trained to meet one or more requirements of each of the guidelines (Priestas, [0020]) with the motivation of ensuring that accurate data is extracted for that guideline (Priestas, [0020]).
Regarding claim 3: Priestas as modified teaches The apparatus of claim 1, wherein the secondary data comprises a third party input (IBR-Tucker, [0054], where the document selector 302 enables selection of documents to be compared in order to enable automated execution of the process pertaining to the request 152. The document selection is enabled by the external knowledge base 108 and domain model 104, where process rules 322 included in the external knowledge base 108 enable identification of the relevant documents needed for the automatic execution of the process. See IBR-Tucker, [0047], where external knowledge base 108 can include explicit knowledge such as rules, inputs from subject matter experts (SMEs) (i.e., “third party input”), machine-generated inputs generated using machine learning, predictive modeling algorithms, etc.).
Regarding claim 4: Priestas as modified teaches The apparatus of claim 1, wherein generating the alteration datum comprises:
determining a correction method datum as a function of the one or more discrepant data fields (IBR-Tucker, [0047-0050], where external knowledge base 108 can be accessed to identify the discrepancies, and when a discrepancy is identified, the reason for the discrepancy can be analyzed, include identifying those data fields wherein the comparisons failed to produce a positive result. Threshold probabilities can be defined for the data models wherein the compared fields that meet the thresholds are deemed as matching while those that fail to meet the thresholds are considered as mismatched/unmatched fields (i.e., “as a function of the one or more discrepant data fields”), where mismatched/unmatched fields from the discrepancy processor can be communicated to a data resolver 114 for an intelligent resolution, where if the data resolver 114 fails to automatically resolve the discrepancy, the information can be displayed for user review using one of the GUIs 104. See also IBR-Tucker, [0057], where whenever a discrepancy is recorded, an auto resolver 308 may initially process the discrepancy for automatic resolution by accessing alternative formats, synonyms, etc. to automatically resolve discrepancies. If the discrepancy cannot be automatically resolved (i.e., “as a function of the one or more discrepant data fields”), a manual resolver 310 can alert a user via one of the GUIs 140 to receive manual input for the discrepancy resolution), wherein determining the correction method datum comprises determining whether the one or more discrepant data fields needs user-provided corrections or automatic system-generated corrections (IBR-Priestas, [0040], and [0052], [0058], where any discrepancies that are identified are resolved at 908 via one or more of an automatic resolution (i.e., “system-generated corrections”) or manual resolution (i.e., “user-provided corrections”). More particularly, if the data resolver 114 fails to automatically resolve the discrepancy (e.g., based on data from intent 164, the domain model 104, and external knowledge base 108), the information can be displayed for user review using one of the GUIs 140. See also IBR-Priestas, [0073], where any discrepancies between documents 154, 156, and data from the external knowledge base 108 can be automatically and/or manually resolved using a disability/life insurance domain model and based on the process rules 322 corresponding to the insurance procedures. The output of the automated procedure can be presented to a user who can either approve or disapprove the reimbursement. The document processing system 100 can also produce a recommendation on whether or not the reimbursement can be approved based on the results of the various categorizations, comparisons, validations, etc., which the user may decide to accept or decline); and
generating the alteration datum as a function of the correction method datum (IBR-Priestas, [0040], where upon user review and confirmation, the information or the required fields augmented with the matches, discrepancies and resolutions are communicated to the document builder 116 which builds an internal master document 172. If the user does not approve the data, the user can make the changes via the GUI. See also IBR-Tucker, [0049], [0057], [0064], and [0074] in claim 1 above for further detail on the generation of the alteration datum by both the system and the user).
Regarding claim 5: Priestas as modified teaches The apparatus of claim 1, wherein generating the alteration datum comprises generating a user prompt as a function of the one or more discrepant data fields (IBR-Priestas, [0036] and [0040], where one or more fields can be flagged for human validation which can be executed via one or more GUIs 140. More specifically, information can be displayed for user review to resolve a discrepancy using one of the GUIs).
Regarding claim 7: Priestas as modified teaches The apparatus of claim 5, wherein generating the user prompt comprises:
transmitting the user prompt to a user device; receiving a user input from the user device as a function of the user prompt; and generating the alteration datum as a function of the user input (IBR-Priestas, [0040], where information can be displayed for user review to resolve a discrepancy using one of the GUIs. Upon user review and confirmation, the information or the required fields augmented with the matches, discrepancies and resolutions are communicated to the document builder 116 which builds an internal master document 172. If the user does not approve the data, the user can make the changes via the GUI).
Regarding claim 8: Priestas as modified teaches The apparatus of claim 1, wherein generating the alteration datum comprises categorizing the one or more discrepant data fields based on its severity, wherein a discrepant data field with a high severity triggers generation of a user prompt, wherein the user prompt comprises a notification (Sathyanarayana, [0032], where data validation includes automatically identifying one or more anomalies. The system selects one or more steps for correction of the anomalies, as well as one or more correction methodologies. The system can complete a data correction process which involves calculating missing values, correcting format, adjusting for inference from other data elements/sources, and correcting erroneous data generally. The data validation process can select a most appropriate and corrected data element package from a weighted list, though can also flag for apparent indecision of the automated process, failing to meet a predetermined threshold for accuracy, and can also highlight residual cases for further analysis including manual review. See, e.g., Sathyanarayana, [0039], where the data manager identifies a unique best weight and can automatically correct and validate the associated data element. Having a unique highest weight does not need to be determinative, e.g., if a next highest weight is relatively close (such as within a few percentage points) (i.e., “categorizing the one or more discrepant data fields based on its severity”), then the data manager may flag those data elements as too close for automatic correction and validation in step 169.
See IBR-Tucker, [0057], where a manual resolver 310 alerts a user via one of the GUIs 140 to receive manual input for the discrepancy resolution (i.e., “generation of a user prompt, wherein the user prompt comprises a notification”)).
It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to have combined the teachings of Priestas as modified and Sathyanarayana with the motivation of reducing the amount of manual review needed for corrections, while taking advantage of manual review in some cases.
Regarding claim 11: Claim 11 recites substantially the same claim limitations as claim 1, and is rejected for the same reasons.
Regarding claim 12: Claim 12 recites substantially the same claim limitations as claim 2, and is rejected for the same reasons.
Regarding claim 13: Claim 13 recites substantially the same claim limitations as claim 3, and is rejected for the same reasons.
Regarding claim 14: Claim 14 recites substantially the same claim limitations as claim 4, and is rejected for the same reasons.
Regarding claim 15: Claim 15 recites substantially the same claim limitations as claim 5, and is rejected for the same reasons.
Regarding claim 17: Claim 17 recites substantially the same claim limitations as claim 7, and is rejected for the same reasons.
Regarding claim 18: Claim 18 recites substantially the same claim limitations as claim 8, and is rejected for the same reasons.
Claims 6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Priestas et al. (“Priestas”) (US 2020/0364404 A1, incorporating by reference Tucker et al. (“IBR-Tucker”) (App. No. 16/531,848 published as US 2019/0354720 A1), Priestas et al. (“IBR-Priestas”) (App. No. 16/179,448 published as US 2019/0087395 A1) and Sacaleanu et al. (“IBR-Sacaleanu”) (App. No. 15/879,031 published as US 2019/0006027 A1) at [0001]), in view of Sathyanarayana et al. (“Sathyanarayana”) (US 2012/0102002 A1), in further view of Srinivasan et al. (“Srinivasan”) (US 2025/0111169 A1), in further view of Lucioni et al. (“Lucioni”) (US 2024/0256424 A1).
Regarding claim 6: Priestas as modified teaches The apparatus of claim 5, but does not appear to explicitly teach wherein generating the user prompt comprises: generating language training data, wherein the language training data comprises exemplary discrepancy data fields correlated to exemplary user prompts; training a large language model using the language training data; and generating the user prompt using the trained large language model.
Lucioni teaches wherein generating the user prompt comprises: generating language training data, wherein the language training data comprises exemplary discrepancy data fields correlated to exemplary user prompts; training a large language model using the language training data; and generating the user prompt using the trained large language model (Lucioni, [0049], where the system processes data during session(s) of a software application using a trained language model (e.g., a large language model (LLM)) to obtain a natural language description of an issue that occurred in the session(s); see also Lucioni, [0134], where the language model 712 may be a large language model (LLM). See Lucioni, [0139], where the language model 712 may be trained using a corpus of textual data, where the language model generate a set of training data using data collected during software application sessions, and be fine-tuned for processing data related to software application sessions. See also Lucioni, [0155], where the session representations are processed using the language model 712 to obtain descriptions of issues that occurred. See IBR-Priestas, [0036] and [0040] above with regards to the “user prompt”).
It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to have combined the teachings of Priestas as modified and Lucioni with the motivation of generating natural language descriptions of issues in order to make such issues more understandable to users and thus potentially facilitate greater effective resolutions by the user.
Regarding claim 16: Claim 16 recites substantially the same claim limitations as claim 6, and is rejected for the same reasons.
Claims 9 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Priestas et al. (“Priestas”) (US 2020/0364404 A1, incorporating by reference Tucker et al. (“IBR-Tucker”) (App. No. 16/531,848 published as US 2019/0354720 A1), Priestas et al. (“IBR-Priestas”) (App. No. 16/179,448 published as US 2019/0087395 A1) and Sacaleanu et al. (“IBR-Sacaleanu”) (App. No. 15/879,031 published as US 2019/0006027 A1) at [0001]), in view of Sathyanarayana et al. (“Sathyanarayana”) (US 2012/0102002 A1), in further view of Srinivasan et al. (“Srinivasan”) (US 2025/0111169 A1), in further view of Ankur et al. (“Ankur”) (US 11,875,109 B1).
Regarding claim 9: Priestas as modified teaches The apparatus of claim 1, but does not appear to explicitly teach wherein generating the alteration datum comprises: generating alteration training data, wherein the alteration training data comprises exemplary discrepant data fields correlated to exemplary alteration data; training an alteration machine-learning model using the alteration training data; and generating the alteration datum using the trained alteration machine-learning model.
Ankur teaches wherein generating the alteration datum comprises: generating alteration training data, wherein the alteration training data comprises exemplary discrepant data fields correlated to exemplary alteration data; training an alteration machine-learning model using the alteration training data; and generating the alteration datum using the trained alteration machine-learning model (Ankur, [6:9-49] and [8:34-67]-[9:1-32], where the ML-based computing system 104 scans one or more documents for obtaining one or more mis-captured data fields present inside the received one or more documents. The ML-based computing system 104 obtains a historical correction data associated with the one or more customers, and determines one or more deltas associated with the one or more mis-captured data fields based on the obtained historical correction data by using a trained data correction-based ML model (implying that the “alteration machine-learning model” was previously “trained”). The system then automatically replaces the one or more mis-captured data fields with the generated one or more correct data fields based on one or more predefined rules. See, e.g., Ankur, [9:48-67]-[10:1-25], where with respect to obtaining training data/test data, where all patterns are captured for corrections done on one or more mis-captured data fields (i.e., “using the alteration training data”)).
It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to have combined the teachings of Priestas as modified and Ankur with the motivation of facilitating automated correction of data via historical examples, which reduces the amount of manual labor required, which can be time consuming, inefficient, and tedious (Ankur, [Background]).
Regarding claim 19: Claim 19 recites substantially the same claim limitations as claim 9, and is rejected for the same reasons.
Claims 10 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Priestas et al. (“Priestas”) (US 2020/0364404 A1, incorporating by reference Tucker et al. (“IBR-Tucker”) (App. No. 16/531,848 published as US 2019/0354720 A1), Priestas et al. (“IBR-Priestas”) (App. No. 16/179,448 published as US 2019/0087395 A1) and Sacaleanu et al. (“IBR-Sacaleanu”) (App. No. 15/879,031 published as US 2019/0006027 A1) at [0001]), in view of Sathyanarayana et al. (“Sathyanarayana”) (US 2012/0102002 A1), in further view of Srinivasan et al. (“Srinivasan”) (US 2025/0111169 A1), in further view of Bell (“Bell”) (US 6,272,506 B1).
Regarding claim 10: Priestas as modified teaches The apparatus of claim 1, but does not appear to explicitly teach further comprising reverting the one or more updated discrepant data fields to a previous state if the one or more updated discrepant data fields does not comply with the corresponding data field.
Bell teaches reverting the one or more updated discrepant data fields to a previous state if the one or more updated discrepant data fields does not comply with the corresponding data field (Bell, [7:29-51], where the system determines whether a value has changed. If so, then the field is flagged for sign off. If the sign off is not accepted, then the field value is reset to its previous value. See Priestas in claim 1 above with respect to the “discrepant data field”).
It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to have combined the teachings of Priestas as modified and Bell with the motivation of ensuring data integrity and compliance.
Regarding claim 20: Claim 20 recites substantially the same claim limitations as claim 10, and is rejected for the same reasons.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. See the enclosed 892 form. vom Lehn (“Understanding Vector Similarity for Machine Learning: Cosine Similarity, Dot Product, Manhattan Distance L1, Euclidian Distance L2”) and Crettaz (“Vector similarity measures and scoring”) are cited to show that one of ordinary skill in the art would have recognized that cosine similarity involves the claimed limitation of “the cosine similarity being computed using a dot product of the [attention-weighted] vector representation normalized by lengths of the [attention-weighted] vector representation”7 (see vom Lehn at [“Cosine Similarity”] on p. 3 of the attachment; Crettaz at [“Cosine Similarity”] on p. 7-9 of the attachment).
The prior art should be considered to define the claims over the art of record.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to IRENE BAKER whose telephone number is (408)918-7601. The examiner can normally be reached M-F 8-5PM PT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Boris Gorney can be reached at (571) 270-5626. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/IRENE BAKER/Primary Examiner, Art Unit 2154
27 June 2026
1 Recall from Srinivasan, [0042] and [0070] in the preceding paragraphs with respect to the input data’s embedding being derived by using multiple attention heads for calculating its own soft weights.
2 vom Lehn (“Understanding Vector Similarity for Machine Learning: Cosine Similarity, Dot Product, Manhattan Distance L1, Euclidian Distance L2”) at [“Cosine Similarity”] on p. 3 of attachment.
3 Crettaz (“Vector similarity measures and scoring”) at [“Cosine Similarity”] on p. 7-9 of attachment.
4 Birru et al. US 2024/0289609 A1 at [0097].
5 D’Agostino. US 2025/0307222 A1 at [0157].
6 D’Agostino at [0158].
7 Note that vom Lehn and Crettaz pertain to vector representations generally; the “attention-weighted” aspect was disclosed by the cited prior art reference in the rejection.