DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
Claims 1-20 were previously pending and subject to non-final action filed on 01/13/2026. In the response filed 04/13/2026, claims 1, 8, 11, 13, 14, 17 and 18 were amended. Therefore, claims 1-20 are currently pending and subject to the non-final action below.
Response to Arguments
Applicant’s arguments, see page 13, filed 04/13/2026, with respect to Drawing have been fully considered and are persuasive. The objection of the Drawing has been withdrawn.
Applicant’s arguments, see page 13, filed 04/13/2026, with respect to Specification have been fully considered and are persuasive. The objection of the Specification has been withdrawn.
Applicant’s arguments, see page 14-16, filed 04/13/2026, with respect to claim 1-20 under 35 U.S.C. 103 have been fully considered and are persuasive. The 103 rejection of claims 1-20 has been withdrawn.
Applicant’s argument: Applicant initially notes that Rimchala is not prior art under 35 U.S.C. § 102(b)(2)(C) because the disclosure of the subject matter on which the rejection is based (Rimchala) and the claimed invention were owned by the same person (Intuit, Inc.) or subject to an obligation of assignment to the same person not later than the effective filing date of the claimed invention. See MPEP § 717.02(a).
Examiner Response: Applicant arguments have been fully considered and are persuasive.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 1-2, 4-8, 11, and 13-19 are rejected under 35 U.S.C. 103 as being unpatentable over Dancewicz (US 11455468 B2, Filed Date: Feb. 16, 2022) in view of ACHIWA (US PGPUB: 20250078549 A1, Filed Date: Aug. 31, 2023).
Regarding independent claim 1, Dancewicz teaches: A method comprising:
receiving a digital image, (Dancewicz − [Col. 3 ll. 40-45] TILT 206 receives real world data 206 including text data, layout data, and image data electronically via any type of data network 210.) wherein the digital image comprises text arranged in a layout within the digital image; (Dancewicz − [Col. 3 ll. 35-45] FIG. 2 is a system diagram of an embodiment of a real-world document processing system 200 as described herein. NLP system 202 in an embodiment is a text-image-layout transformer (TILT). TILT 206 receives real world data 206 including text data, layout data, and image data electronically via any type of data network 210.)
generating, by an optical character recognition model, (Dancewicz – Fig. 4B [Col. 5 ll. 45-52] An image, represented as a matrix of pixels, is processed by an OCR system to obtain text tokens.) a layout text vector that encodes at least one word in the text of the digital image and also encodes a position of the at least one word in the layout of the digital image; (Dancewicz – Fig. 4B [col. 5 ll. 45-57] An image, represented as a matrix of pixels, is processed by an OCR system to obtain text tokens. result of this operation is summed with corresponding attention biases combining linear 1D relations as well as spatial 2D relations; the spatial 2D relations are, in turn, determined using the distances of bounding boxes of each token, as obtained with OCR.)
generating, by a visual encoder model, a visual representation vector embedding a content of the digital image; (Dancewicz – Fig. 4B [col. 5 ll. 45-57] An image, represented as a matrix of pixels, is processed by an OCR system to obtain text tokens. he contextualized visual features obtained directly from the image, each text token is assigned distinct visual features relative to its position and surroundings. The joint embeddings are mapped into queries, keys and values, using learnable linear projections. Image embeddings are produced representing visual features of the image.)
converting both the layout text vector and the visual representation vector into a projected text vector, (Dancewicz – Fig. 4B [col. 5 ll. 45-57] Text embedding are added to contextualized visual features to form joint embeddings. The joint embeddings are mapped into queries, keys and values, using learnable linear projection. Thereby generating contextualized embeddings for subsequent Transformer processing.)
wherein the projected text vector comprises a digital format suitable for input to a large language model (Dancewicz – [Col. 4 ll. 4-10] Most NLP tasks can be unified under one framework by casting them as Language Modeling, The T5 Transformer is a prominent prior art example of largescale Transformers achieving state-of-the-art results on varied NLP benchmarks. [col. 5 ll. 45-65] [Col. 6 ll. 1-11] The contextualized embedding are input to the next Transformer layer and the contextualized image-region embeddings are input into the Transformer by adding them to semantic embeddings; In order to inject visual information to the Transformer, a matrix of contextualized image-region embeddings I is added to semantic embedding). Examiner Note: Under the BRI, the T5 Transformer is considered a large language that performs language modeling using transformer-based contextualized text embeddings as input.
and generating an output comprising a key-value pair, wherein: a key of the key-value pair represents a type of the text and a value of the key-value pair represents the value of the type, (Dancewicz – [Col. 3 ll. 45-50] TILT generates output 212 which includes key information, document classification and answers to questions 208. Fig. 4B [col. 5 ll. 45-57The joint embeddings are mapped into queries, keys and values, using learnable linear projections. Queries are matched against keys using dot product.)
Dancewicz does not explicitly teach: combining, into a prompt, the projected text vector, a system message, and a task instruction;
However, ACHIWA teaches: combining, into a prompt, (ACHIWA − [0046] The large language model 116 is a model called LLM (Large Language Model) capable of generating sentences in an interactive manner, and generates replies to input instruction messages (prompts).)
the projected text vector, a system message, (ACHIWA − [0109] In S802, the CPU 261 obtains an instruction message template from the storage 265. The instruction message template, which has been prepared in advance, may be a template prepared as a preset template by the engineer or the user or such a preset template to which a correction or an addition has been made by the system or the user. Examiner NOTE: instruction message template can be system message prepared in advance)
and a task instruction; (ACHIWA − [0165-0168] message “output only the data and total amount from the following text in the JSON format”, as shown in Fig. 14B element 1410. Is a task instruction; [0168] For example, the instruction message 1410 in FIG. 14B includes an instruction to answer the item values corresponding to the items determined to have been unextracted or erroneously extracted from among the group of character strings recognized from the document image.)
the output is generated by the large language model which takes, as input, the prompt, (ACHIWA − [0165-0168] message “output only the data and total amount from the following text in the JSON format”, as shown in Fig. 14B element 1410. Is a task instruction; [0168] For example, the instruction message 1410 in FIG. 14B includes an instruction to answer the item values corresponding to the items determined to have been unextracted or erroneously extracted from among the group of character strings recognized from the document image. A reply 1411 to the instruction message 1410 from the large language model 116 indicates that the large language model 116 has returned “June 2, 2023” as the character string of the item “date” and “¥13,000” as the character string of the item “total amount”.)
and the large language model establishes a semantic association between a key layout of the key in the digital image relative to a value layout of the value in the digital image. (ACHIWA − [0165] the CPU 261 inserts the group of character strings recognized from the document image by the OCR process into a character string group region 1402. The character strings are inserted in the order of recognition, for example. As a result, an instruction message 1410 in FIG. 14B is generated. [0166] In S1205, the CPU 261 performs a process of inputting the instruction message generated in S1204 into the large language model 116. [0167] In S1206, the CPU 261 receives a reply to the instruction message input in S1205 from the large language model 116. 0168] For example, the instruction message 1410 in FIG. 14B includes an instruction to answer the item values corresponding to the items determined to have been unextracted or erroneously extracted from among the group of character strings recognized from the document image. A reply 1411 to the instruction message 1410 from the large language model 116 indicates that the large language model 116 has returned “June 2, 2023” as the character string of the item “date” and “¥13,000” as the character string of the item “total amount”.) The returned values (“June 2, 2023” and “¥13,000”) are semantically associated with their corresponding requested items (“date” and “total amount”).
Accordingly, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the teachings of Dancewicz and ACHIWA because both references relate to document understanding using OCR recognized text. ACHIWA teaches using a large language model (LLM) with instruction messages (prompts) to interpret and correct OCR-recognized document content. One of ordinary skill in the art would have been motivated to incorporate ACHIWA prompt-based LLM processing into Dancewicz layout-aware document understanding system in order to improve extraction accuracy and correct misrecognized or unextracted document information.
Regarding dependent claim 2, depends on claim 1, Dancewicz teaches: wherein the visual representation vector comprises a hidden representation vector output by a plurality of inner layers of the visual encoder model. (Dancewicz – [Col. 2 ll. 59-60] FIG. 5 is an illustration of a U-NET network according to an embodiment. [Col. 4 ll. 10-15] The T5 Transformer is a prominent prior art example of largescale Transformers achieving state-of-the-art results on varied NLP benchmarks. [Col. 5 ll. 15-25] In an embodiment, to produce image embeddings, a convolutional network that consumes the whole page image of size 512×384 is used, and it produces a feature map of 64×48×128. An embodiment uses U-Net as a backbone encoder network since this architecture provides access to not only the information in the near neighborhood of the token, such as font and style, but also to more distant regions of the page, which is useful in cases where the text is related to other structures, e.g., where the text is the description of a picture.)
Regarding dependent claim 4, depends on claim 1, Dancewicz teaches: wherein: converting comprises inputting a combination of the layout text vector and the visual representation vector to a projection network model, the projection network model outputs the projected text vector, and the projection network model projects the layout text vector and the visual representation vector into a textual token embedding space. (Dancewicz – [Col. 2 ll. 59-60] FIG. 5 is an illustration of a U-NET network according to an embodiment. [Col. 4 ll. 10-15] The T5 Transformer is a prominent prior art example of largescale Transformers achieving state-of-the-art results on varied NLP benchmarks. [Col. 5 ll. 15-25] In an embodiment, to produce image embeddings, a convolutional network that consumes the whole page image of size 512×384 is used, and it produces a feature map of 64×48×128. An embodiment uses U-Net as a backbone encoder network since this architecture provides access to not only the information in the near neighborhood of the token, such as font and style, but also to more distant regions of the page, which is useful in cases where the text is related to other structures, e. g, where the text is the description of a picture.)
Regarding dependent claim 5, depends on claim 1, Dancewicz teaches: extract a plurality of key-value pairs from the digital image in the document, (Dancewicz – Fig. 4B [col. 5 ll. 45-57] An image, represented as a matrix of pixels, is processed by an OCR system to obtain text tokens. he contextualized visual features obtained directly from the image, each text token is assigned distinct visual features relative to its position and surroundings. The joint embeddings are mapped into queries, keys and values, using learnable linear projections. Image embeddings are produced representing visual features of the image.) but does not explicitly teach: the task instruction
However, ACHIWA teaches: wherein: the task instruction is to extract the key-value pair, or the task instruction is to extract a plurality of key-value pairs from the digital image, and the key-value pair is one of the plurality of key-value pairs representing key information entities in a document. (ACHIWA − [0046] The external information processing server 105 is an apparatus that utilizes a large language model 116. The large language model 116 is a model called LLM (Large Language Model) capable of generating sentences in an interactive manner, and generates replies to input instruction messages (prompts). For example, ChatGPT (registered trademark), [0165-0168] message “output only the data and total amount from the following text in the JSON format”, as shown in Fig. 14B element 1410. Is a task instruction; [0168] A reply 1411 to the instruction message 1410 from the large language model 116 indicates that the large language model 116 has returned “June 2, 2023” as the character string of the item “date” and “¥13,000” as the character string of the item “total amount”.) return instruction message is key information entities in document.
Accordingly, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the teachings of Dancewicz and ACHIWA because both references relate to document understanding using OCR recognized text. ACHIWA teaches using a large language model (LLM) with instruction messages (prompts) to interpret and correct OCR-recognized document content. One of ordinary skill in the art would have been motivated to incorporate ACHIWA prompt-based LLM processing into Dancewicz layout-aware document understanding system in order to improve extraction accuracy and correct misrecognized or unextracted document information.
Regarding dependent claim 6, depends on claim 1, Dancewicz does not explicitly teach: the prompt is generated for a zero shot inference without demonstration, and the prompt is additionally generated to include both the task instruction and a demonstration comprising a multimodal input followed by an expected output represented as a known key-value pair in a structured format.
However, ACHIWA teaches: wherein: the prompt is generated for a zero shot inference without demonstration, and the prompt is additionally generated to include both the task instruction and a demonstration comprising a multimodal input followed by an expected output represented as a known key-value pair in a structured format. (ACHIWA − [0046] The external information processing server 105 is an apparatus that utilizes a large language model 116. The large language model 116 is a model called LLM (Large Language Model) capable of generating sentences in an interactive manner, and generates replies to input instruction messages (prompts). For example, ChatGPT (registered trademark), [0165-0168] message “output only the data and total amount from the following text in the JSON format”, as shown in Fig. 14B element 1410. Is a task instruction; [0168] A reply 1411 to the instruction message 1410 from the large language model 116 indicates that the large language model 116 has returned “June 2, 2023” as the character string of the item “date” and “¥13,000” as the character string of the item “total amount”.) JSON is a structured format; processing server 105 operates like ChatGPT (registered trademark), at zero-shot inference, a core capability where it performs tasks (like summarizing, translating, or classifying) with a prompt,
Accordingly, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the teachings of Dancewicz and ACHIWA because both references relate to document understanding using OCR recognized text. ACHIWA teaches using a large language model (LLM) with instruction messages (prompts) to interpret and correct OCR-recognized document content. One of ordinary skill in the art would have been motivated to incorporate ACHIWA prompt-based LLM processing into Dancewicz layout-aware document understanding system in order to improve extraction accuracy and correct misrecognized or unextracted document information.
Regarding dependent claim 7, depends on claim 1, Dancewicz does not explicitly teach: wherein: the prompt is generated for a zero shot inference without demonstration, and the prompt is generated to include meta-information including a document type of the digital image.
However, ACHIWA teaches: wherein: the prompt is generated for a zero shot inference without demonstration, and the prompt is generated to include meta-information including a document type of the digital image. (ACHIWA −[0046] The external information processing server 105 is an apparatus that utilizes a large language model 116. The large language model 116 is a model called LLM (Large Language Model) capable of generating sentences in an interactive manner, and generates replies to input instruction messages (prompts). For example, ChatGPT (registered trademark), [0165-0168] message “output only the data and total amount from the following text in the JSON format”, as shown in Fig. 14B element 1410. Is a task instruction; [0168] A reply 1411 to the instruction message 1410 from the large language model 116 indicates that the large language model 116 has returned “June 2, 2023” as the character string of the item “date” and “¥13,000” as the character string of the item “total amount”.) date is a type of data; processing server 105 operates like ChatGPT (registered trademark), at zero-shot inference, a core capability where it performs tasks (like summarizing, translating, or classifying) with a prompt,
Accordingly, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the teachings of Dancewicz and ACHIWA because both references relate to document understanding using OCR recognized text. ACHIWA teaches using a large language model (LLM) with instruction messages (prompts) to interpret and correct OCR-recognized document content. One of ordinary skill in the art would have been motivated to incorporate ACHIWA prompt-based LLM processing into Dancewicz layout-aware document understanding system in order to improve extraction accuracy and correct misrecognized or unextracted document information.
Regarding independent claim 8, Dancewicz teaches: A method comprising:
receiving training data comprising a reference output for a digital image, (Dancewicz − [Col. 3 ll. 40-45] TILT 206 receives real world data 206 including text data, layout data, and image data electronically via any type of data network 210.) wherein the digital image comprises text arranged in a layout within the digital image; (Dancewicz − [Col. 3 ll. 35-45] FIG. 2 is a system diagram of an embodiment of a real-world document processing system 200 as described herein. NLP system 202 in an embodiment is a text-image-layout transformer (TILT). TILT 206 receives real world data 206 including text data, layout data, and image data electronically via any type of data network 210.)
performing a first sub-method comprising: generating, by a visual encoder model, a visual representation vector embedding a content of the digital image; (Dancewicz – Fig. 4B [col. 5 ll. 45-57] An image, represented as a matrix of pixels, is processed by an OCR system to obtain text tokens. he contextualized visual features obtained directly from the image, each text token is assigned distinct visual features relative to its position and surroundings. The joint embeddings are mapped into queries, keys and values, using learnable linear projections. Image embeddings are produced representing visual features of the image.)
converting, using a projection network model, the visual representation vector into a projected text vector, wherein the projected text vector comprises a digital format suitable for input to a large language model; (Dancewicz – [Col. 4 ll. 4-10] Most NLP tasks can be unified under one framework by casting them as Language Modeling, The T5 Transformer is a prominent prior art example of largescale Transformers achieving state-of-the-art results on varied NLP benchmarks. [col. 5 ll. 45-65] [Col. 6 ll. 1-11] The contextualized embedding are input to the next Transformer layer and the contextualized image-region embeddings are input into the Transformer by adding them to semantic embeddings; In order to inject visual information to the Transformer, a matrix of contextualized image-region embeddings I is added to semantic embedding). Examiner Note: Under the BRI, the T5 Transformer is considered a large language that performs language modeling using transformer-based contextualized text embeddings as input.
generating a loss function by comparing the output to the reference output; (Dancewicz – [Col. 3 ll. 59-62] Further improvements focus on the training and inference aspects by the inclusion of the area masking loss function or achieving independence from sequential order in decoding respectively.)
adjusting, based on the loss function, one or more parameters in the projection network model; and training a first trained projection network model by iterating, until convergence, (Dancewicz – Fig. 6 [Col. 3 ll. 59-62] Further improvements focus on the training and inference aspects by the inclusion of the area masking loss function or achieving independence from sequential order in decoding respectively [Col. 7 ll. 20-35] The training process is iterative. The result of each iteration is a working prediction model able to process a document and return a value for each data point. The model is evaluated at the end of each iteration. The user verifies the correctness of the predictions provided by the model comparing the extracted values with the information present in the document. The model quality is measured as a percentage of correct predictions. The quality of the model improves over time. The process is stopped when the user achieves satisfactory quality.)
receiving the training data, performing the first sub-method, generating the loss function, and adjusting the one or more parameters, wherein upon convergence the projection network model is transformed into the first trained projection network model. (Dancewicz – Fig. 6 [Col. 7 ll. 30-50] Referring to FIG. 6, an overview of an embodiment of a training process 600 is illustrated. Each iteration consists of the following steps performed by the user. Sample documents are documents of a type that includes the kind of document form and included information expected to be processed. At 602, the documents are processed with the model. In a first iteration a generic neural network is used. The generic neural network returns predictions, answering the questions asked in natural language. In next iterations a fine-tuned neural network is used. At 604, the extracted values returned by the prediction model are validated for the sample documents. Validated data for each data point consists of validated extracted value where correct answers are confirmed, incorrect or missing answers are completed. At 606, the quality of the model is evaluated with the use of validated data. At 608, the prediction model is trained using validated data for few sample documents.)
Dancewicz does not explicitly teach: combining, into a prompt, the projected text vector, a system message, and a task instruction;
However, ACHIWA teaches: combining, into a prompt, (ACHIWA − [0046] The large language model 116 is a model called LLM (Large Language Model) capable of generating sentences in an interactive manner, and generates replies to input instruction messages (prompts).)
the projected text vector, a system message, (ACHIWA − [0109] In S802, the CPU 261 obtains an instruction message template from the storage 265. The instruction message template, which has been prepared in advance, may be a template prepared as a preset template by the engineer or the user or such a preset template to which a correction or an addition has been made by the system or the user. Examiner NOTE: instruction message template can be system message prepared in advance)
and a task instruction; (ACHIWA − [0165-0168] message “output only the data and total amount from the following text in the JSON format”, as shown in Fig. 14B element 1410. Is a task instruction; [0168] For example, the instruction message 1410 in FIG. 14B includes an instruction to answer the item values corresponding to the items determined to have been unextracted or erroneously extracted from among the group of character strings recognized from the document image.)
and generating, using the large language model that takes the prompt as input, an output comprising a sequence of next tokens in an optical character recognition text determined for the text in the image; (ACHIWA – [0045] The display control unit 158 performs control for displaying the item values extracted by the document image analysis unit 154 and the reply to the instruction message obtained from the large language model 116 to the user. [0061] The information processing server 103 and execute an information processing program stored in the storage 265 to execute information processing such as character recognition (OCR) and information extraction. [0165-0168] message “output only the data and total amount from the following text in the JSON format”, as shown in Fig. 14B element 1410. Is a task instruction; [0168] For example, the instruction message 1410 in FIG. 14B includes an instruction to answer the item values corresponding to the items determined to have been unextracted or erroneously extracted from among the group of character strings recognized from the document image. A reply 1411 to the instruction message 1410 from the large language model 116 indicates that the large language model 116 has returned “June 2, 2023” as the character string of the item “date” and “¥13,000” as the character string of the item “total amount”.)
wherein the large language model establishes a semantic association between a key layout of a key in the digital image relative to a value layout of a value in the digital image; (ACHIWA − [0165] the CPU 261 inserts the group of character strings recognized from the document image by the OCR process into a character string group region 1402. The character strings are inserted in the order of recognition, for example. As a result, an instruction message 1410 in FIG. 14B is generated. [0166] In S1205, the CPU 261 performs a process of inputting the instruction message generated in S1204 into the large language model 116. [0167] In S1206, the CPU 261 receives a reply to the instruction message input in S1205 from the large language model 116. 0168] For example, the instruction message 1410 in FIG. 14B includes an instruction to answer the item values corresponding to the items determined to have been unextracted or erroneously extracted from among the group of character strings recognized from the document image. A reply 1411 to the instruction message 1410 from the large language model 116 indicates that the large language model 116 has returned “June 2, 2023” as the character string of the item “date” and “¥13,000” as the character string of the item “total amount”.) The returned values (“June 2, 2023” and “¥13,000”) are semantically associated with their corresponding requested items (“date” and “total amount”).
Accordingly, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the teachings of Dancewicz and ACHIWA because both references relate to document understanding using OCR recognized text. ACHIWA teaches using a large language model (LLM) with instruction messages (prompts) to interpret and correct OCR-recognized document content. One of ordinary skill in the art would have been motivated to incorporate ACHIWA prompt-based LLM processing into Dancewicz layout-aware document understanding system in order to improve extraction accuracy and correct misrecognized or unextracted document information.
Regarding dependent claim 11, depends on claim 8, Dancewicz teaches: wherein receiving, performing, generating, adjusting, and training comprise a first training operation, and wherein the method further comprises a second training operation comprising: (Dancewicz – Fig. 6 [Col. 7 ll. 30-50] Referring to FIG. 6, an overview of an embodiment of a training process 600 is illustrated. Each iteration consists of the following steps performed by the user. Sample documents are documents of a type that includes the kind of document form and included information expected to be processed. At 602, the documents are processed with the model. In a first iteration a generic neural network is used. The generic neural network returns predictions, answering the questions asked in natural language. In next iterations a fine-tuned neural network is used. At 604, the extracted values returned by the prediction model are validated for the sample documents. Validated data for each data point consists of validated extracted value where correct answers are confirmed, incorrect or missing answers are completed. At 606, the quality of the model is evaluated with the use of validated data. At 608, the prediction model is trained using validated data for few sample documents.)
re-receiving the training data; (Dancewicz − [Col. 3 ll. 40-45] TILT 206 receives real world data 206 including text data, layout data, and image data electronically via any type of data network 210.)
performing a second sub-method comprising: generating, by an optical character recognition model, (Dancewicz – Fig. 4B [Col. 5 ll. 45-52] An image, represented as a matrix of pixels, is processed by an OCR system to obtain text tokens.)
a layout text vector that encodes at least one word in the text of the digital image and also encodes a position of the at least one word in the layout of the digital image; (Dancewicz – Fig. 4B [col. 5 ll. 45-57] An image, represented as a matrix of pixels, is processed by an OCR system to obtain text tokens. result of this operation is summed with corresponding attention biases combining linear 1D relations as well as spatial 2D relations; the spatial 2D relations are, in turn, determined using the distances of bounding boxes of each token, as obtained with OCR.)
generating, by the visual encoder model, the visual representation vector embedding the content of the digital image; (Dancewicz – Fig. 4B [col. 5 ll. 45-57] An image, represented as a matrix of pixels, is processed by an OCR system to obtain text tokens. he contextualized visual features obtained directly from the image, each text token is assigned distinct visual features relative to its position and surroundings. The joint embeddings are mapped into queries, keys and values, using learnable linear projections. Image embeddings are produced representing visual features of the image.)
converting, using the projection network model, both the layout text vector and the visual representation vector into a second projected text vector, wherein the projected text vector comprises the digital format suitable for input to the large language model; (Dancewicz – [Col. 4 ll. 4-10] Most NLP tasks can be unified under one framework by casting them as Language Modeling, The T5 Transformer is a prominent prior art example of largescale Transformers achieving state-of-the-art results on varied NLP benchmarks. [col. 5 ll. 45-65] [Col. 6 ll. 1-11] The contextualized embedding are input to the next Transformer layer and the contextualized image-region embeddings are input into the Transformer by adding them to semantic embeddings; In order to inject visual information to the Transformer, a matrix of contextualized image-region embeddings I is added to semantic embedding). Examiner Note: Under the BRI, the T5 Transformer is considered a large language that performs language modeling using transformer-based contextualized text embeddings as input.
and generating a second output comprising a second key-value pair, wherein: a second key of the second key-value pair represents a second type of the text and a second value of the key-value pair represents the second value of the type, (Dancewicz – [Col. 3 ll. 45-50] TILT generates output 212 which includes key information, document classification and answers to questions 208. Fig. 4B [col. 5 ll. 45-57The joint embeddings are mapped into queries, keys and values, using learnable linear projections. Queries are matched against keys using dot product.)
generating a second loss function by comparing the second output to the reference output; (Dancewicz – [Col. 3 ll. 59-62] Further improvements focus on the training and inference aspects by the inclusion of the area masking loss function or achieving independence from sequential order in decoding respectively.)
adjusting, based on the loss function, the one or more parameters of both the first trained projection network model and the large language model; (Dancewicz – Fig. 6 [Col. 3 ll. 59-62] Further improvements focus on the training and inference aspects by the inclusion of the area masking loss function or achieving independence from sequential order in decoding respectively [Col. 7 ll. 20-35] The training process is iterative. The result of each iteration is a working prediction model able to process a document and return a value for each data point. The model is evaluated at the end of each iteration. The user verifies the correctness of the predictions provided by the model comparing the extracted values with the information present in the document. The model quality is measured as a percentage of correct predictions. The quality of the model improves over time. The process is stopped when the user achieves satisfactory quality.)
and training a second trained projection network model and a trained large language model by iterating, until convergence, re-receiving the training data, re-performing the second sub-method, generating the second loss function, and adjusting the one or more parameters, (Dancewicz – Fig. 6 [Col. 3 ll. 59-62] Further improvements focus on the training and inference aspects by the inclusion of the area masking loss function or achieving independence from sequential order in decoding respectively [Col. 7 ll. 20-35] The training process is iterative. The result of each iteration is a working prediction model able to process a document and return a value for each data point. The model is evaluated at the end of each iteration. The user verifies the correctness of the predictions provided by the model comparing the extracted values with the information present in the document. The model quality is measured as a percentage of correct predictions. The quality of the model improves over time. The process is stopped when the user achieves satisfactory quality.)
wherein: upon convergence the first trained projection network model is transformed into the second trained projection network model and the large language model is transformed into the trained large language model; (Dancewicz – Fig. 6 [Col. 7 ll. 30-50] Referring to FIG. 6, an overview of an embodiment of a training process 600 is illustrated. Each iteration consists of the following steps performed by the user. Sample documents are documents of a type that includes the kind of document form and included information expected to be processed. At 602, the documents are processed with the model. In a first iteration a generic neural network is used. The generic neural network returns predictions, answering the questions asked in natural language. In next iterations a fine-tuned neural network is used. At 604, the extracted values returned by the prediction model are validated for the sample documents. Validated data for each data point consists of validated extracted value where correct answers are confirmed, incorrect or missing answers are completed. At 606, the quality of the model is evaluated with the use of validated data. At 608, the prediction model is trained using validated data for few sample documents.)
and the trained large language model is adapted for extraction of the key-value pair. (Dancewicz – Fig. 9 [Col 8 ll. 35-40] At 910, the neural network answers the formulated questions and generates extracted values for each question and for each document.)
Dancewicz does not explicitly teach: combining, into a second prompt, the second projected text vector, a second system message, and a second task instruction;
However, ACHIWA teaches: combining, into a second prompt, the second projected text vector, a second system message, (ACHIWA − [0109] In S802, the CPU 261 obtains an instruction message template from the storage 265. The instruction message template, which has been prepared in advance, may be a template prepared as a preset template by the engineer or the user or such a preset template to which a correction or an addition has been made by the system or the user. Examiner NOTE: instruction message template can be system message prepared in advance)
and a second task instruction; (ACHIWA − [0165-0168] message “output only the data and total amount from the following text in the JSON format”, as shown in Fig. 14B element 1410. Is a task instruction; [0168] For example, the instruction message 1410 in FIG. 14B includes an instruction to answer the item values corresponding to the items determined to have been unextracted or erroneously extracted from among the group of character strings recognized from the document image.)
and the second output is generated by the large language model which takes, as input, the second prompt; (ACHIWA − [0165-0168] message “output only the data and total amount from the following text in the JSON format”, as shown in Fig. 14B element 1410. Is a task instruction; [0168] For example, the instruction message 1410 in FIG. 14B includes an instruction to answer the item values corresponding to the items determined to have been unextracted or erroneously extracted from among the group of character strings recognized from the document image. A reply 1411 to the instruction message 1410 from the large language model 116 indicates that the large language model 116 has returned “June 2, 2023” as the character string of the item “date” and “¥13,000” as the character string of the item “total amount”.)
Accordingly, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the teachings of Dancewicz and ACHIWA because both references relate to document understanding using OCR recognized text. ACHIWA teaches using a large language model (LLM) with instruction messages (prompts) to interpret and correct OCR-recognized document content. One of ordinary skill in the art would have been motivated to incorporate ACHIWA prompt-based LLM processing into Dancewicz layout-aware document understanding system in order to improve extraction accuracy and correct misrecognized or unextracted document information.
Regarding dependent claim 13, depends on claim 12, Dancewicz teaches: receiving, after the second training operation, a new digital image; performing the second sub-method a third time on the new digital image; and presenting the second key-value pair extracted from the new digital image. (Dancewicz – Fig. 4B [Col. 5 ll. 45-52] An image, represented as a matrix of pixels, is processed by an OCR system to obtain text tokens. Fig. 6 [Col. 3 ll. 59-62] Further improvements focus on the training and inference aspects by the inclusion of the area masking loss function or achieving independence from sequential order in decoding respectively [Col. 7 ll. 20-35] The training process is iterative. The result of each iteration is a working prediction model able to process a document and return a value for each data point. The model is evaluated at the end of each iteration. The user verifies the correctness of the predictions provided by the model comparing the extracted values with the information present in the document. The model quality is measured as a percentage of correct predictions. The quality of the model improves over time. The process is stopped when the user achieves satisfactory quality.)
Regarding independent claim 14, Dancewicz teaches: A system comprising: (Dancewicz –)
a computer processor; (Dancewicz – [Col. 9 ll. 50-54] one or more processors that employ any one of a variety of operating systems or platforms)
a data repository in communication with the computer processor and storing: (Dancewicz – [Col. 9 ll. 59-67] The computer readable medium or media can be transportable, such that the program or programs stored thereon can be loaded onto one or more different computers or other processors to implement various aspects of the present invention as discussed above.)
a digital image comprising text arranged in a layout within the digital image, (Dancewicz − [Col. 3 ll. 40-45] TILT 206 receives real world data 206 including text data, layout data, and image data electronically via any type of data network 210.)
a layout text vector that encodes at least one word in the text of the digital image and also encodes a position of the at least one word in the layout of the digital (Dancewicz – Fig. 4B [col. 5 ll. 45-57] An image, represented as a matrix of pixels, is processed by an OCR system to obtain text tokens. result of this operation is summed with corresponding attention biases combining linear 1D relations as well as spatial 2D relations; the spatial 2D relations are, in turn, determined using the distances of bounding boxes of each token, as obtained with OCR.)
a projected text vector, wherein the projected text vector comprises a digital format suitable for input to a large language model, (Dancewicz – Fig. 4B [col. 5 ll. 45-57] Text embedding are added to contextualized visual features to form joint embeddings. The joint embeddings are mapped into queries, keys and values, using learnable linear projection. Thereby generating contextualized embeddings for subsequent Transformer processing. Dancewicz – [Col. 4 ll. 4-10] Most NLP tasks can be unified under one framework by casting them as Language Modeling, The T5 Transformer is a prominent prior art example of largescale Transformers achieving state-of-the-art results on varied NLP benchmarks. [col. 5 ll. 45-65] [Col. 6 ll. 1-11] The contextualized embedding are input to the next Transformer layer and the contextualized image-region embeddings are input into the Transformer by adding them to semantic embeddings; In order to inject visual information to the Transformer, a matrix of contextualized image-region embeddings I is added to semantic embedding).)
and an output comprising a key-value pair, wherein a key of the key-value pair represents a type of the text and a value of the key-value pair represents a value of the type; (Dancewicz – [Col. 3 ll. 45-50] TILT generates output 212 which includes key information, document classification and answers to questions 208. Fig. 4B [col. 5 ll. 45-57The joint embeddings are mapped into queries, keys and values, using learnable linear projections. Queries are matched against keys using dot product.)
an optical character recognition model which, when executed by the computer processor, is programmed to generate the layout text vector; (Dancewicz – Fig. 4B [Col. 5 ll. 45-52] An image, represented as a matrix of pixels, is processed by an OCR system to obtain text tokens. [col. 5 ll. 45-57] Text embedding are added to contextualized visual features to form joint embeddings. The joint embeddings are mapped into queries, keys and values, using learnable linear projection. Thereby generating contextualized embeddings for subsequent Transformer processing.))
a visual encoder model which, when executed by the computer processor, is programmed to generate visual representation vector; (Dancewicz – Fig. 4B [col. 5 ll. 45-57] An image, represented as a matrix of pixels, is processed by an OCR system to obtain text tokens. he contextualized visual features obtained directly from the image, each text token is assigned distinct visual features relative to its position and surroundings. The joint embeddings are mapped into queries, keys and values, using learnable linear projections. Image embeddings are produced representing visual features of the image.)
a projection network model which, when executed by the computer processor, is programmed to generate the projected text vector; (Dancewicz – [Col. 2 ll. 59-60] FIG. 5 is an illustration of a U-NET network according to an embodiment. [Col. 4 ll. 10-15] The T5 Transformer is a prominent prior art example of largescale Transformers achieving state-of-the-art results on varied NLP benchmarks. [Col. 5 ll. 15-25] In an embodiment, to produce image embeddings, a convolutional network that consumes the whole page image of size 512×384 is used, and it produces a feature map of 64×48×128. An embodiment uses U-Net as a backbone encoder network since this architecture provides access to not only the information in the near neighborhood of the token, such as font and style, but also to more distant regions of the page, which is useful in cases where the text is related to other structures, e. g, where the text is the description of a picture.)
Dancewicz does not explicitly teach: a prompt, a system message, a task instruction;
However, ACHIWA teaches: a prompt, a system message, (ACHIWA − [0046] The large language model 116 is a model called LLM (Large Language Model) capable of generating sentences in an interactive manner, and generates replies to input instruction messages (prompts). [0109] In S802, the CPU 261 obtains an instruction message template from the storage 265. The instruction message template, which has been prepared in advance, may be a template prepared as a preset template by the engineer or the user or such a preset template to which a correction or an addition has been made by the system or the user. Examiner NOTE: instruction message template can be system message prepared in advance)
a task instruction, (ACHIWA − [0165-0168] message “output only the data and total amount from the following text in the JSON format”, as shown in Fig. 14B element 1410. Is a task instruction; [0168] For example, the instruction message 1410 in FIG. 14B includes an instruction to answer the item values corresponding to the items determined to have been unextracted or erroneously extracted from among the group of character strings recognized from the document image.)
a prompt generator which, when executed by the computer processor, is programmed to generate the prompt by combining the projected text vector, the system message, and the task instruction; (ACHIWA − [0046] The large language model 116 is a model called LLM (Large Language Model) capable of generating sentences in an interactive manner, and generates replies to input instruction messages (prompts). [0109] In S802, the CPU 261 obtains an instruction message template from the storage 265. The instruction message template, which has been prepared in advance, may be a template prepared as a preset template by the engineer or the user or such a preset template to which a correction or an addition has been made by the system or the user. [0165-0168] message “output only the data and total amount from the following text in the JSON format”, as shown in Fig. 14B element 1410. Is a task instruction; [0168] For example, the instruction message 1410 in FIG. 14B includes an instruction to answer the item values corresponding to the items determined to have been unextracted or erroneously extracted from among the group of character strings recognized from the document image.)
and to establish a semantic association between a key layout of the key in the digital image relative to a value layout of the value in the digital image. (ACHIWA − [0046] The external information processing server 105 is an apparatus that utilizes a large language model 116. The large language model 116 is a model called LLM (Large Language Model) capable of generating sentences in an interactive manner, and generates replies to input instruction messages (prompts). For example, ChatGPT (registered trademark), [0165-0168] message “output only the data and total amount from the following text in the JSON format”, as shown in Fig. 14B element 1410. Is a task instruction; [0168] A reply 1411 to the instruction message 1410 from the large language model 116 indicates that the large language model 116 has returned “June 2, 2023” as the character string of the item “date” and “¥13,000” as the character string of the item “total amount”.) return instruction message is key information entities in document.
Accordingly, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the teachings of Dancewicz and ACHIWA because both references relate to document understanding using OCR recognized text. ACHIWA teaches using a large language model (LLM) with instruction messages (prompts) to interpret and correct OCR-recognized document content. One of ordinary skill in the art would have been motivated to incorporate ACHIWA prompt-based LLM processing into Dancewicz layout-aware document understanding system in order to improve extraction accuracy and correct misrecognized or unextracted document information.
Regarding dependent claim 15, depends on claim 14, Dancewicz teaches: further comprising: a training controller which, when executed by the computer processor, is programmed to train only the projection network model in a first training stage to generate a trained projection network model. (Dancewicz – Fig. 6 [Col. 7 ll. 30-50] Referring to FIG. 6, an overview of an embodiment of a training process 600 is illustrated. Each iteration consists of the following steps performed by the user. Sample documents are documents of a type that includes the kind of document form and included information expected to be processed. At 602, the documents are processed with the model. In a first iteration a generic neural network is used. The generic neural network returns predictions, answering the questions asked in natural language. In next iterations a fine-tuned neural network is used. At 604, the extracted values returned by the prediction model are validated for the sample documents. Validated data for each data point consists of validated extracted value where correct answers are confirmed, incorrect or missing answers are completed. At 606, the quality of the model is evaluated with the use of validated data. At 608, the prediction model is trained using validated data for few sample documents.)
Regarding dependent claim 16, depends on claim 15, Dancewicz teaches: wherein the training controller is further programmed to train both the trained projection network model and the large language model in a second training stage. (Dancewicz – Fig. 6 [Col. 7 ll. 30-50] Referring to FIG. 6, an overview of an embodiment of a training process 600 is illustrated. Each iteration consists of the following steps performed by the user. Sample documents are documents of a type that includes the kind of document form and included information expected to be processed. At 602, the documents are processed with the model. In a first iteration a generic neural network is used. The generic neural network returns predictions, answering the questions asked in natural language. In next iterations a fine-tuned neural network is used. At 604, the extracted values returned by the prediction model are validated for the sample documents. Validated data for each data point consists of validated extracted value where correct answers are confirmed, incorrect or missing answers are completed. At 606, the quality of the model is evaluated with the use of validated data. At 608, the prediction model is trained using validated data for few sample documents.)
Regarding dependent claim 17, depends on claim 14, Dancewicz teaches: further comprising: a training controller which, when executed by the computer processor, is programmed to: receive training data comprising a reference output comprising a reference digital image; (Dancewicz − [Col. 3 ll. 40-45] TILT 206 receives real world data 206 including text data, layout data, and image data electronically via any type of data network 210.)
perform a first sub-method comprising: generating, by the visual encoder model, the visual representation vector; (Dancewicz – Fig. 4B [col. 5 ll. 45-57] An image, represented as a matrix of pixels, is processed by an OCR system to obtain text tokens. he contextualized visual features obtained directly from the image, each text token is assigned distinct visual features relative to its position and surroundings. The joint embeddings are mapped into queries, keys and values, using learnable linear projections. Image embeddings are produced representing visual features of the image.)
converting the visual representation vector into the projected text vector; (Dancewicz – Fig. 4B [col. 5 ll. 45-57] Text embedding are added to contextualized visual features to form joint embeddings. The joint embeddings are mapped into queries, keys and values, using learnable linear projection. Thereby generating contextualized embeddings for subsequent Transformer processing.)
generate a loss function by comparing the output to the reference output; (Dancewicz – [Col. 3 ll. 59-62] Further improvements focus on the training and inference aspects by the inclusion of the area masking loss function or achieving independence from sequential order in decoding respectively.)
adjust, based on the loss function, at least one parameter of the projection network model; and Dancewicz – Fig. 6 [Col. 3 ll. 59-62] Further improvements focus on the training and inference aspects by the inclusion of the area masking loss function or achieving independence from sequential order in decoding respectively [Col. 7 ll. 20-35] The training process is iterative. The result of each iteration is a working prediction model able to process a document and return a value for each data point. The model is evaluated at the end of each iteration. The user verifies the correctness of the predictions provided by the model comparing the extracted values with the information present in the document. The model quality is measured as a percentage of correct predictions. The quality of the model improves over time. The process is stopped when the user achieves satisfactory quality.)
training a first trained projection network model by iterating, until convergence, receiving the training data, performing the sub- method, generating the loss function, adjusting the at least one parameter, (Dancewicz – Fig. 6 [Col. 3 ll. 59-62] Further improvements focus on the training and inference aspects by the inclusion of the area masking loss function or achieving independence from sequential order in decoding respectively [Col. 7 ll. 20-35] The training process is iterative. The result of each iteration is a working prediction model able to process a document and return a value for each data point. The model is evaluated at the end of each iteration. The user verifies the correctness of the predictions provided by the model comparing the extracted values with the information present in the document. The model quality is measured as a percentage of correct predictions. The quality of the model improves over time. The process is stopped when the user achieves satisfactory quality.)
wherein upon convergence the projection network model is transformed into the first trained projection network model. (Dancewicz – Fig. 6 [Col. 7 ll. 30-50] Referring to FIG. 6, an overview of an embodiment of a training process 600 is illustrated. Each iteration consists of the following steps performed by the user. Sample documents are documents of a type that includes the kind of document form and included information expected to be processed. At 602, the documents are processed with the model. In a first iteration a generic neural network is used. The generic neural network returns predictions, answering the questions asked in natural language. In next iterations a fine-tuned neural network is used. At 604, the extracted values returned by the prediction model are validated for the sample documents. Validated data for each data point consists of validated extracted value where correct answers are confirmed, incorrect or missing answers are completed. At 606, the quality of the model is evaluated with the use of validated data. At 608, the prediction model is trained using validated data for few sample documents.)
Dancewicz does not explicitly teach: combining, into the prompt, the projected text vector, the system message, and the task instruction;
However, ACHIWA teaches: combining, into the prompt, the projected text vector, the system message, (ACHIWA − [0109] In S802, the CPU 261 obtains an instruction message template from the storage 265. The instruction message template, which has been prepared in advance, may be a template prepared as a preset template by the engineer or the user or such a preset template to which a correction or an addition has been made by the system or the user. Examiner NOTE: instruction message template can be system message prepared in advance)
and the task instruction; (ACHIWA − [0165-0168] message “output only the data and total amount from the following text in the JSON format”, as shown in Fig. 14B element 1410. Is a task instruction; [0168] For example, the instruction message 1410 in FIG. 14B includes an instruction to answer the item values corresponding to the items determined to have been unextracted or erroneously extracted from among the group of character strings recognized from the document image.)
and generating, using the large language model that takes the prompt as input, a model output comprising a sequence of next tokens in an optical character recognition text determined for the text in the digital image; (ACHIWA – [0045] The display control unit 158 performs control for displaying the item values extracted by the document image analysis unit 154 and the reply to the instruction message obtained from the large language model 116 to the user. [0061] The information processing server 103 and execute an information processing program stored in the storage 265 to execute information processing such as character recognition (OCR) and information extraction. [0165-0168] message “output only the data and total amount from the following text in the JSON format”, as shown in Fig. 14B element 1410. Is a task instruction; [0168] For example, the instruction message 1410 in FIG. 14B includes an instruction to answer the item values corresponding to the items determined to have been unextracted or erroneously extracted from among the group of character strings recognized from the document image. A reply 1411 to the instruction message 1410 from the large language model 116 indicates that the large language model 116 has returned “June 2, 2023” as the character string of the item “date” and “¥13,000” as the character string of the item “total amount”.)
Accordingly, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the teachings of Dancewicz and ACHIWA because both references relate to document understanding using OCR recognized text. ACHIWA teaches using a large language model (LLM) with instruction messages (prompts) to interpret and correct OCR-recognized document content. One of ordinary skill in the art would have been motivated to incorporate ACHIWA prompt-based LLM processing into Dancewicz layout-aware document understanding system in order to improve extraction accuracy and correct misrecognized or unextracted document information.
Regarding dependent claim 18, depends on claim 17, Dancewicz teaches: wherein receiving, performing the sub-method, generating the loss function,
adjusting, and training comprise a first training operation, (Dancewicz – Fig. 6 [Col. 7 ll. 30-50] Referring to FIG. 6, an overview of an embodiment of a training process 600 is illustrated. Each iteration consists of the following steps performed by the user. Sample documents are documents of a type that includes the kind of document form and included information expected to be processed. At 602, the documents are processed with the model. In a first iteration a generic neural network is used. The generic neural network returns predictions, answering the questions asked in natural language. In next iterations a fine-tuned neural network is used. At 604, the extracted values returned by the prediction model are validated for the sample documents. Validated data for each data point consists of validated extracted value where correct answers are confirmed, incorrect or missing answers are completed. At 606, the quality of the model is evaluated with the use of validated data. At 608, the prediction model is trained using validated data for few sample documents.)
and wherein the training controller is further programmed to perform a second training operation comprising: (Dancewicz – Fig. 4B [Col. 5 ll. 45-52] An image, represented as a matrix of pixels, is processed by an OCR system to obtain text tokens.)
re-receiving the training data comprising: the reference digital image comprising reference text; wherein a reference key of the reference key-value pair represents a reference type of the text and a reference value of the reference key-value pair represents a reference type value; (Dancewicz – Fig. 6 [Col. 3 ll. 59-62] Further improvements focus on the training and inference aspects by the inclusion of the area masking loss function or achieving independence from sequential order in decoding respectively [Col. 7 ll. 20-35] The training process is iterative. The result of each iteration is a working prediction model able to process a document and return a value for each data point. The model is evaluated at the end of each iteration. The user verifies the correctness of the predictions provided by the model comparing the extracted values with the information present in the document. The model quality is measured as a percentage of correct predictions. The quality of the model improves over time. The process is stopped when the user achieves satisfactory quality.)
perform a second sub-method comprising: generating, by the optical character recognition model, (Dancewicz – Fig. 4B [Col. 5 ll. 45-52] An image, represented as a matrix of pixels, is processed by an OCR system to obtain text tokens.)
the layout text vector that encodes the at least one word in the reference text of the reference digital image and also encodes a second position of the at least one word in the layout of the reference digital image; (Dancewicz – Fig. 4B [col. 5 ll. 45-57] An image, represented as a matrix of pixels, is processed by an OCR system to obtain text tokens. result of this operation is summed with corresponding attention biases combining linear 1D relations as well as spatial 2D relations; the spatial 2D relations are, in turn, determined using the distances of bounding boxes of each token, as obtained with OCR.)
generating, by the visual encoder model, the visual representation vector embedding the content of the reference digital image; (Dancewicz – Fig. 4B [col. 5 ll. 45-57] An image, represented as a matrix of pixels, is processed by an OCR system to obtain text tokens. he contextualized visual features obtained directly from the image, each text token is assigned distinct visual features relative to its position and surroundings. The joint embeddings are mapped into queries, keys and values, using learnable linear projections. Image embeddings are produced representing visual features of the image.)
converting, using the projection network model, both the layout text vector and the visual representation vector into a second projected text vector, wherein the second projected text vector comprises the digital format suitable for input to the large language model; (Dancewicz – [Col. 4 ll. 4-10] Most NLP tasks can be unified under one framework by casting them as Language Modeling, The T5 Transformer is a prominent prior art example of largescale Transformers achieving state-of-the-art results on varied NLP benchmarks. [col. 5 ll. 45-65] [Col. 6 ll. 1-11] The contextualized embedding are input to the next Transformer layer and the contextualized image-region embeddings are input into the Transformer by adding them to semantic embeddings; In order to inject visual information to the Transformer, a matrix of contextualized image-region embeddings I is added to semantic embedding). Examiner Note: Under the BRI, the T5 Transformer is considered a large language that performs language modeling using transformer-based contextualized text embeddings as input.
generating a second output comprising a second key-value pair, wherein: a second key of the second key-value pair represents a second type of the text and the second value of the key-value pair represents a value of the type, (Dancewicz – [Col. 3 ll. 45-50] TILT generates output 212 which includes key information, document classification and answers to questions 208. Fig. 4B [col. 5 ll. 45-57The joint embeddings are mapped into queries, keys and values, using learnable linear projections. Queries are matched against keys using dot product.)
and generating a second loss function by comparing the second output to the reference key-value pair of the reference output; (Dancewicz – [Col. 3 ll. 59-62] Further improvements focus on the training and inference aspects by the inclusion of the area masking loss function or achieving independence from sequential order in decoding respectively.)
adjusting, based on the second loss function, both the first trained projection network model and the large language model; (Dancewicz – Fig. 6 [Col. 3 ll. 59-62] Further improvements focus on the training and inference aspects by the inclusion of the area masking loss function or achieving independence from sequential order in decoding respectively [Col. 7 ll. 20-35] The training process is iterative. The result of each iteration is a working prediction model able to process a document and return a value for each data point. The model is evaluated at the end of each iteration. The user verifies the correctness of the predictions provided by the model comparing the extracted values with the information present in the document. The model quality is measured as a percentage of correct predictions. The quality of the model improves over time. The process is stopped when the user achieves satisfactory quality.)
and training a second trained projection network model and a trained large language model by iterating, until convergence, (Dancewicz – Fig. 6 [Col. 3 ll. 59-62] Further improvements focus on the training and inference aspects by the inclusion of the area masking loss function or achieving independence from sequential order in decoding respectively [Col. 7 ll. 20-35] The training process is iterative. The result of each iteration is a working prediction model able to process a document and return a value for each data point. The model is evaluated at the end of each iteration. The user verifies the correctness of the predictions provided by the model comparing the extracted values with the information present in the document. The model quality is measured as a percentage of correct predictions. The quality of the model improves over time. The process is stopped when the user achieves satisfactory quality.)
wherein upon convergence the first trained projection network model is transformed into the second trained projection network model and the large language model is transformed into the trained large language model. (Dancewicz – Fig. 6 [Col. 7 ll. 30-50] Referring to FIG. 6, an overview of an embodiment of a training process 600 is illustrated. Each iteration consists of the following steps performed by the user. Sample documents are documents of a type that includes the kind of document form and included information expected to be processed. At 602, the documents are processed with the model. In a first iteration a generic neural network is used. The generic neural network returns predictions, answering the questions asked in natural language. In next iterations a fine-tuned neural network is used. At 604, the extracted values returned by the prediction model are validated for the sample documents. Validated data for each data point consists of validated extracted value where correct answers are confirmed, incorrect or missing answers are completed. At 606, the quality of the model is evaluated with the use of validated data. At 608, the prediction model is trained using validated data for few sample documents.)
Dancewicz does not explicitly teach: a reference prompt comprising the reference digital image, a reference system message, and a reference task instruction, and a reference output comprising a reference key-value pair, combining, into a prompt, the projected text vector, a system message, and a task instruction;
However, ACHIWA teaches: a reference prompt comprising the reference digital image, a reference system message, and a reference task instruction, and a reference output comprising a reference key-value pair, (ACHIWA − [0109] In S802, the CPU 261 obtains an instruction message template from the storage 265. The instruction message template, which has been prepared in advance, may be a template prepared as a preset template by the engineer or the user or such a preset template to which a correction or an addition has been made by the system or the user. [0165-0168] message “output only the data and total amount from the following text in the JSON format”, as shown in Fig. 14B element 1410. Is a task instruction; [0168] For example, the instruction message 1410 in FIG. 14B includes an instruction to answer the item values corresponding to the items determined to have been unextracted or erroneously extracted from among the group of character strings recognized from the document image. Examiner NOTE: instruction message template can be system message prepared in advance)
combining, into a second prompt, the second projected text vector, a system message, (ACHIWA − [0109] In S802, the CPU 261 obtains an instruction message template from the storage 265. The instruction message template, which has been prepared in advance, may be a template prepared as a preset template by the engineer or the user or such a preset template to which a correction or an addition has been made by the system or the user. Examiner NOTE: instruction message template can be system message prepared in advance)
and a second task instruction; (ACHIWA − [0165-0168] message “output only the data and total amount from the following text in the JSON format”, as shown in Fig. 14B element 1410. Is a task instruction; [0168] For example, the instruction message 1410 in FIG. 14B includes an instruction to answer the item values corresponding to the items determined to have been unextracted or erroneously extracted from among the group of character strings recognized from the document image.)
and the output is generated by the large language model which takes, as input, the prompt; (ACHIWA − [0165-0168] message “output only the data and total amount from the following text in the JSON format”, as shown in Fig. 14B element 1410. Is a task instruction; [0168] For example, the instruction message 1410 in FIG. 14B includes an instruction to answer the item values corresponding to the items determined to have been unextracted or erroneously extracted from among the group of character strings recognized from the document image. A reply 1411 to the instruction message 1410 from the large language model 116 indicates that the large language model 116 has returned “June 2, 2023” as the character string of the item “date” and “¥13,000” as the character string of the item “total amount”.)
Accordingly, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the teachings of Dancewicz and ACHIWA because both references relate to document understanding using OCR recognized text. ACHIWA teaches using a large language model (LLM) with instruction messages (prompts) to interpret and correct OCR-recognized document content. One of ordinary skill in the art would have been motivated to incorporate ACHIWA prompt-based LLM processing into Dancewicz layout-aware document understanding system in order to improve extraction accuracy and correct misrecognized or unextracted document information.
Regarding dependent claim 19, depends on claim 14, Dancewicz teaches: wherein the visual representation vector comprises a hidden representation vector output by a plurality of inner layers of the visual encoder model. (Dancewicz – [Col. 2 ll. 59-60] FIG. 5 is an illustration of a U-NET network according to an embodiment. [Col. 4 ll. 10-15] The T5 Transformer is a prominent prior art example of largescale Transformers achieving state-of-the-art results on varied NLP benchmarks. [Col. 5 ll. 15-25] In an embodiment, to produce image embeddings, a convolutional network that consumes the whole page image of size 512×384 is used, and it produces a feature map of 64×48×128. An embodiment uses U-Net as a backbone encoder network since this architecture provides access to not only the information in the near neighborhood of the token, such as font and style, but also to more distant regions of the page, which is useful in cases where the text is related to other structures, e.g., where the text is the description of a picture.)
Claim(s) 3 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Dancewicz and ACHIWA as applied to claims 2 and 19 above, and further in view of Baker (US PGPUB: 20120163707 A1, Filed Date: 28, 2010).
Regarding dependent claim 3, depends on claim 2, Dancewicz does not explicitly teach: wherein the visual representation vector excludes a caption text for the digital image.
However, BAKER teaches: wherein the visual representation vector excludes a caption text for the digital image. (BAKER − [0070] Some embodiments may use various filters or heuristics to eliminate positive examples that may be noise. For example, a caption or replacement text for an image that is not descriptive may be removed as a positive example.)
Accordingly, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the teaching of Dancewicz, ACHIWA and BAKER as each inventions relates to character recognition of documents. One of ordinary skill in the art would have been motivated for correcting and updating misrecognized characters. Therefore, improving character recognition of document images.
Regarding dependent claim 20, depends on claim 19, Dancewicz does not explicitly teach: wherein the visual representation vector excludes a caption text for the digital image.
However, BAKER teaches: wherein the visual representation vector excludes a caption text for the digital image. (BAKER − [0070] Some embodiments may use various filters or heuristics to eliminate positive examples that may be noise. For example, a caption or replacement text for an image that is not descriptive may be removed as a positive example.)
Accordingly, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the teaching of Dancewicz, ACHIWA and BAKER as each inventions relates to character recognition of documents. One of ordinary skill in the art would have been motivated for correcting and updating misrecognized characters. Therefore, improving character recognition of document images.
Claim(s) 9-10 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Dancewicz and ACHIWA as applied to claims 8 and 11 above, and further in view of Lester (US PGPUB: 20230325725 A1, Filed Date: Apr. 12, 2022).
Regarding dependent claim 9, depends on claim 8, Dancewicz teaches: a the visual encoder model but does not explicitly teach: wherein the visual encoder model and the large language model are frozen.
However, Lester teaches: wherein the visual encoder model and the large language model are frozen such that only the one or more parameters of the projection network model are trained. (Lester − [0030] The plurality of pre-trained parameters for the pre-trained machine-learned model can be fixed during prompt tuning (e.g., the pre-trained machine-learned model can be frozen such that the parameters are not adjusted during training of the prompt parameters) [0034] prompt tuning can involve inputting parameters with the input data into the frozen model such that only those parameters are updated.)
Accordingly, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the teaching of Dancewicz, ACHIWA and Lester as each inventions relates to character recognition of documents. One of ordinary skill in the art would have been motivated for correcting training data based on specific task and reduce the computation resources used during training.
Regarding dependent claim 10, depends on claim 8, Dancewicz teaches: a the visual encoder model but does not explicitly teach: wherein the visual encoder model is frozen, and wherein adjusting further comprises adjusting both the projection network model and the large language model.
However, Lester teaches: wherein the visual encoder model is frozen, and wherein adjusting further comprises adjusting both the projection network model and the large language model. (Lester − [0030] The plurality of pre-trained parameters for the pre-trained machine-learned model can be fixed during prompt tuning (e.g., the pre-trained machine-learned model can be frozen such that the parameters are not adjusted during training of the prompt parameters) [0034] prompt tuning can involve inputting parameters with the input data into the frozen model such that only those parameters are updated. [0113] prompt parameter training can involve a plurality of iterations of output generation and comparison. During such training, the parameters of the pre-trained machine-learned model 906 can remain unadjusted, or “frozen.” Adjust parameters of one model while other model 906 is frozen.)
Accordingly, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the teaching of Dancewicz, ACHIWA and Lester as each inventions relates to character recognition of documents. One of ordinary skill in the art would have been motivated for correcting training data based on specific task and reduce the computation resources used during training.
Regarding dependent claim 12, depends on claim 11, Dancewicz teaches: a the visual encoder model but does not explicitly teach: wherein: the visual encoder model and the large language model are frozen during the first training operation such that the projection network model is trained during the first training operation, and the visual encoder model is frozen during the second training operation such that both the first trained projection network model and the large language model are trained during the second training operation.
However, Lester teaches: wherein: the visual encoder model and the large language model are frozen during the first training operation such that the projection network model is trained during the first training operation, and the visual encoder model is frozen during the second training operation such that both the first trained projection network model and the large language model are trained during the second training operation. (Lester − [0030] The plurality of pre-trained parameters for the pre-trained machine-learned model can be fixed during prompt tuning (e.g., the pre-trained machine-learned model can be frozen such that the parameters are not adjusted during training of the prompt parameters) [0034] prompt tuning can involve inputting parameters with the input data into the frozen model such that only those parameters are updated. [0101] As depicted, model tuning 202 can include retraining a machine-learned model for each task.)
Accordingly, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the teaching of Dancewicz, ACHIWA and Lester as each inventions relates to character recognition of documents. One of ordinary skill in the art would have been motivated for correcting training data based on specific task and reduce the computation resources used during training.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Paula US 20230409624 A1. Generating semantic embeddings for both textual content and visual content and organizing those multimodal embeddings into a searchable hierarchical graph for sematic retrieval.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CARL E BARNES JR whose telephone number is (571)270-3395. The examiner can normally be reached Monday-Friday 9am-6pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Stephen Hong can be reached at (571) 272-4124. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CARL E BARNES JR/Examiner, Art Unit 2178
/STEPHEN S HONG/Supervisory Patent Examiner, Art Unit 2178