Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-20 are presented for examination.
Drawings
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(4) because reference character “110” has been used to designate both OCR and Network. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Specification
Applicant is reminded of the proper language and format for an abstract of the disclosure.
The abstract should be in narrative form and generally limited to a single paragraph on a separate sheet within the range of 50 to 150 words in length. The abstract should describe the disclosure sufficiently to assist readers in deciding whether there is a need for consulting the full patent text for details.
The language should be clear and concise and should not repeat information given in the title. It should avoid using phrases which can be implied, such as, “The disclosure concerns,” “The disclosure defined by this invention,” “The disclosure describes,” etc. In addition, the form and legal phraseology often used in patent claims, such as “means” and “said,” should be avoided.
The abstract of the disclosure is objected to because "Aspects of the present disclosure provide" is used and should be avoided. A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b).
The disclosure is objected to because of the following informalities:
Paragraph [0044] line 12 recites “item extraction model 150” in which “150” appears to be a typographical error for “250,” the number referring to the item extraction model in the drawings (see, for example, the same paragraph [0044] line 2 reciting “Item extraction model 250” and paragraph [0045] line 3 and paragraph [0046] lines 1 and 2)
Paragraph [0059] line 6 recites “structured text 220” in which “220” appears to be a typographical error for “330,” the number referring to the structured text in the drawings (see, for example, the same paragraph [0059] line 2 reciting “Structured text 330” and paragraph [0077] line 4)
Paragraph [0073] lines 2 and 3 recite “I/O devices 514 (e.g., keyboards, displays, mouse devices, pen input, etc.) in which “514” appears to be a typographical error for “504,” the number referring to I/O devices in the drawings (see, for example, the same paragraph [0073] lines 1 and 2 reciting “I/O device interfaces 504” and paragraph [0074] line 4)
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 5, 6, and 19 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding claims 5 and 19, it is not clearly understood what is meant by “wherein the evaluating of the objective function is based on comparing an output produced by the language processing machine learning model to a schema”. Based on claim 1 and applicant specification, paragraph [0024], the language processing machine learning module generated label which have a schema. The item extraction machine learning model is the element that produced the output and the comparison involves evaluating an objective function that considers correspondence between the output and the label. For the purpose of examination, the examiner will interpret the claim in light of the specification the limitation as “wherein the evaluating of the objective function is based on comparing an output produced by the item extraction machine learning model to a schema”.
Regarding claim 6, it is not clearly understood what is meant by “wherein the evaluating of the objective function is based on determining whether the structured text indicates that an output produced by the language processing machine learning model is contained within a table in the structured document”. Based on claim 1 and application specification, paragraph [0024], the item extraction machine learning model is the element that produced the output and the determination involves evaluating an objective function that considers whether individual values in the output are contained in one or more tables in the structured document. For the purpose of examination, the examiner will interpret the claim in light of the specification the limitation as “wherein the evaluating of the objective function is based on determining whether the structured text indicates that an output produced by the item extraction machine learning model is contained within a table in the structured document”.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 2, 8, 10, 11, 12, 13, 15, 16, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Cohen et al. (US 11514244 B2, filed 12/22/2015), hereinafter Cohen, and in view of Muralidharan et al. (US 20230024040 A1, filed 07/26/2021), hereinafter Muralidharan.
Regarding claim 1, Cohen teaches: A method of training an item extraction machine learning model, comprising: (Cohen; [Abstract], Briefly described, training a machine learning model for extraction of knowledge from images and associated texts):
extracting text and bounding box coordinates from a structured document: (Cohen;
[Column 10 lines 63 – Column 11 line 15], Briefly described, selecting the bounding box coordinate to ground a caption statement and extract the coordinate and text from it; [Column 7 lines 33-55], Briefly described, structured knowledge as text being taken or extracted from existing image captions and surrounding text in documents):
creating structured text by adjusting formatting of the extracted text based on the
extracted bounding box coordinates: (Cohen; [Column 10 lines 63 – Column 11 line 15], Briefly described, a tuple formatting containing structured text extracted based on the extracted bounding box coordinates):
providing the structured text to a language processing machine learning model along
with a prompt instructing the language processing machine learning model to generate a label indicating variables present in the structured text and values for the variables: (Cohen; [Column 3 lines 40-57], Briefly described, structured text being provided to a model using natural language processing; [Column 9 lines 65 – Column 10 line 5], Briefly described, classifier labels configured to or prompted to classify objects; [Column 10 lines 6-37], Briefly described, a classification label indicating the subject, attribute tuples and subject, predicate, object tuples as variables present in the structured text, and a bounding box to consist of values for the tuples as a key):
receiving the label from the language processing machine learning model in response to
the structured text and the prompt: (Cohen; [Column 3 lines 40-57], Briefly described, structured text being provided to a model using natural language processing; [Column 9 lines 65 – Column 10 line 5], Briefly described, classifier labels configured to or prompted to classify objects as a result of the natural language processing model as a CNN localizing objects in an image; [Column 10 lines 25-37], Briefly described, the classifier labels being obtained or received from the natural language processing model;
training the item extraction machine learning model through a supervised training
process based on training data comprising the structured text and the label; (Cohen; [Column 11 lines 51-60], Briefly described, the item extraction machine learning model being trained on text features of structured text based on the classifier labels with image features of the image in training data; [Column 14 lines 42-52], Briefly described, a loss function where incorrect matching of labels calculates loss as a supervised training process where a training tuple does not match a training image).
However, Cohen fails to expressly teach – adding table delimiter tags to the extracted text based on detecting one or more tables in the structured document.
In the same field of endeavor, Muralidharan:
adding table delimiter tags to the extracted text based on detecting one or more tables
in the structured document: (Muralidharan; [0063], Briefly described, content segmentation units generated based on detecting tables in the text; [0003], Briefly described, tags applied to content segmentation units; [0043] Briefly described, structured document consisting of the text and tables to be extracted).
It would have been obvious to one of ordinary skill in the art before the effective filing data of the invention to have incorporated – adding table delimiter tags to the extracted text based on detecting one or more tables in the structured document as suggested by Cohen and Muralidharan. Doing so would be desirable because the extracted text and bounding box coordinates from Cohen would be able to have applied tags from the combination of Cohen and Muralidharan. The combined method would allow for the structured document to be analyzed for a detection of tables which Cohen is unable to do without Muralidharan. The combination of Cohen and Muralidharan provides the item extraction machine learning model with this extra capability to enact on the structured document.
Regarding claim 2, the combination of Cohen and Muralidharan teaches the invention as claimed in claim 1 above including wherein the training is based on one or more confidence sources associated with the extracting of the text, the detecting of the one or more tables, or the receiving of the label from the language processing machine learning model: (Cohen; [Column 3 lines 40-57], Briefly described, structured text being provided to a model using natural language processing; [Column 4 lines 31-55], Briefly described, a confidence source that is calculated to describe the correspondence between the extracted text and the image; [column 9 lines 6-21], Briefly described, confidence measures assigned to the tuples with a classification label; [Column 9 lines 65 – Column 10 line 5], Briefly described, classifier labels configured to classify objects as a result of the natural language processing model as a CNN localizing objects in an image; [Column 10 lines 25-37], Briefly described, Briefly described, the classifier labels being obtained or received from the natural language processing model) (Muralidharan; [0063], Briefly described, content segmentation units generated based on detecting tables in the text).
Regarding claim 8, the combination of Cohen and Muralidharan teaches the invention as claimed in claim 1 above including wherein the detecting of the one or more tables comprises providing the structured document to a table detection machine learning model and receiving bounding coordinates of the one or more tables and corresponding confidence scores from the table detection machine learning model in response to the structured document: (Cohen; [Column 10 lines 63 – Column 11 line 15], Briefly described, a tuple formatting containing structured text extracted based on the extracted bounding box coordinates; [Column 4 lines 31-55], Briefly described, tuples assigned a confidence value) (Muralidharan; [0063], Briefly described, content segmentation units generated based on detecting tables in the text; [0003], Briefly described, tags applied to content segmentation units; [0043] Briefly described, structured document consisting of the text and tables to be extracted; [0034], Briefly described, a detection machine learning model for detecting the content segmentation units generated based on detecting tables).
Regarding claim 10, the combination of Cohen and Muralidharan teaches the invention as claimed in claim 1 above including wherein the prompt specified that the label is to conform to a schema that specifies a structure for indicating variables and the values for the variables: (Cohen; [Column 10 lines 6-37], Briefly described, the output of class labels indicating subject and object tuples from the language processing machine learning model being mapped or compared with a pre-defined subset of database objects or a schema, and conforms to a localization module schema through a prompt or rules and heuristics where tuples or variables are associated with bounding box values as a key).
Regarding claim 11, the combination of Cohen and Muralidharan teaches the invention as claimed in claim 1 above including wherein the training of the item extraction machine learning model comprises generating an embedding based on the structured document and providing the embedding along with the structured document as training inputs to the item extraction machine learning model, wherein the embedding comprises a multimodal representation vector: (Cohen; [Column 13 lines 32-47], Briefly described, an embedding created or generated on the structured document where text features “t” and image features “x” are embedded and thus indicate a multimodal representation vector based on multiple forms of media in the item extraction machine learning model used as training inputs; [Column 16 lines 41-57], Briefly described, embeddings for images which are structured to be used as a training input).
Regarding claim 12, the combination of Cohen and Muralidharan teaches the invention as claimed in claim 1 above including wherein the item extraction machine learning model is a compact multimodal large language model (MLLM) having a smaller number of tunable parameters than the language processing machine learning model used to generate the label: (Cohen; [Column 13 lines 20-31], Briefly described, tunable parameters through fine tuning the language processing machine learning model, and the item extraction machine learning model fine tuning the language processing model for correlating text and image features indicating it is multimodal and that the tunable parameters of the language processing machine learning model are greater; [Column 19 lines 31 – Column 20 line 36]. Briefly described, two different models, one item extraction model and one language processing model where the item extraction model can have a smaller number of tunable parameters than the language processing model).
Regarding claim 13, the combination of Cohen and Muralidharan teaches the invention as claimed in claim 1 above including wherein the training of the item extraction machine learning model comprises instruction fine-tuning of a compact multimodal large language model (MLLM): (Cohen; [Column 13 lines 20-31], Briefly described, the item extraction machine learning model being fine-tuned through features learning for classification in an image during training, and these features consist of text and image features, indicating it is multimodal).
Regarding claims 15 and 16, they are apparatus claims that correspond to method claims 1 and 2. Therefore, they are rejected for the same reason as claims 1 and 2 above.
Regarding claim 20, it is a computer-readable medium claim that corresponds to claim 1. Therefore, it is rejected for the same reason as claim 1 above.
Claims 3, 5, 6, 14, 17, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Cohen in view of Muralidharan, as applied in the rejection of claims 1, 2, and 16 above, further in view of Kim et al. (US 20220101123 A1, filed 06/30/2021), hereinafter Kim.
Regarding claim 3, the combination of Cohen and Muralidharan teaches the invention of claim 2.
However, the combination of Cohen and Muralidharan fail to expressly teach – wherein the training of the item extraction machine learning model comprises performing a noise aware training process that involves adjusting one or more parameters of the item extraction machine learning model based on evaluating an objective function.
In the same field of endeavor, Kim teaches:
wherein the training of the item extraction machine learning model comprises
performing a noise aware training process that involves adjusting one or more parameters of the item extraction machine learning model based on evaluating an objective function: (Kim; [0099], Briefly described, a machine learning model with the ability to extract features from input data; [0027], Briefly described, the machine learning model being trained to obtain an objective quality assessment score; [0030], Briefly described, a noise aware training process based off the objective quality assessment score with adjustable parameters like noise reduction).
It would have been obvious to one of ordinary skill in the art before the publishing date of the invention to have incorporated – wherein the training of the item extraction machine learning model comprises performing a noise aware training process that involves adjusting one or more parameters of the item extraction machine learning model based on evaluating an objective function as suggested by Cohen, Muralidharan, and Kim. Doing so would be desirable because the item extraction machine that extracts and creates structured text from a structured document can be filtered using the noise aware training process to improve the accuracy of label generation for the tuple variations of Cohen. The combined method would provide for the ability of the item extraction machine learning model to receive input from a structured document that Kim otherwise could not do on their own.
Regarding claim 5, the combination of Cohen, Muralidharan, and Kim teaches the invention as claimed in claim 3 above including wherein the evaluating of the objective function is based on: (Kim; [0027], Briefly described, an objective function in the form of an objective quality assessment score):
comparing an output produced by the language processing machine learning model to a
schema: (Cohen; [Column 10 lines 6-37], Briefly described, the output of class labels indicating subjects and objects from the item extraction machine learning model being mapped or compared with a pre-defined subset of database objects or a schema).
Regarding claim 6, the combination of Cohen, Muralidharan, and Kim teaches the invention as claimed in claim 3 above including wherein the evaluating of the objective function is based on determining whether the structured text indicates that an output produced by the language processing machine learning model is contained within a table in the structured document: (Kim; [0027], Briefly described, the machine learning model being trained to obtain an objective quality assessment score) (Muralidharan; [0038], Briefly described, the language processing machine learning model processing the content segmentation unit to determine if there are action items that are an output of the machine learning model; [0063], Briefly described, content segmentation units generated based on tables, where the machine learning model can process these content segmentation units to determine if they contain tables in the content data; [0030], Briefly described, a structured document data object being organized based on tables).
Regarding claim 14, the combination of Cohen and Muralidharan teaches the invention of claim 1.
However, the combination of Cohen and Muralidharan fail to expressly teach – wherein the training data further comprises an instruction prompt.
In the same field of endeavor, Kim teaches:
wherein the training data further comprises an instruction prompt: (Kim; [0027], Briefly
described, instructions to be executed by a processor and to obtain a subjective quality assessment score and distorted image to be used as a training data set).
It would have been obvious to one of ordinary skill in the art before the filing date of the invention to have incorporated – wherein the training data further comprises an instruction prompt as suggested by Cohen, Muralidharan, and Kim. Doing so would be desirable because having instructions in the training data allows for formatting while the structured text is being extracted. Without instructions, the item extraction machine learning model would not have a directive as to how to extraction the text from a structured document. The combination of Cohen, Muralidharan, and Kim allows for a functioning item extraction machine learning model where prompts are instructions for performing the steps of the method.
Regarding claims 17 and 19, they are apparatus claims that correspond to method claims 3 and 5. Therefore, they are rejected for the same reason as claims 3 and 5 above.
Claims 4 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Cohen in view of Muralidharan and of Kim as applied in the rejection of claims 3 and 17 above, further in view of CSER et al. (WO 2020086773 A1, filed 10/23/2019), hereinafter Cser.
Regarding claim 4, the combination of Cohen, Muralidharan, and Kim teaches the invention as claimed in claim 3 above including wherein the evaluating of the objective function comprises: (Kim; [0027], Briefly described, an objective function in the form of an objective quality assessment score):
loss based on computing an aggregation of a text extraction: (Cohen; [Column 13 lines
32-36], Briefly described, a machine learning module mapping text features “t” and image features “x” as an aggregation of a text extraction; [Column 14 lines 34-52], Briefly described, a loss function based on the aggregation of the extracted text using text features “t” and image features “x”).
However, the combination of Cohen, Muralidharan, and Kim fail to teach – confidence score of the one or more confidence scores and a language processing machine learning model confidence score of the one or more confidence scores, wherein the computer aggregation is used to determine a weight associated with the label during the training.
In the same field of endeavor, Cser teaches:
confidence score of the one or more confidence scores and a language processing
machine learning model confidence score of the one or more confidence scores, wherein the computed aggregation is used to determine a weight associated with the label during training: (Cser; [Page 14 Paragraph 2], Briefly described, a scoring machine learning model where confidence is measured through an element match score, and the model has a confidence score to identify the correct element of text that gets extracted; [Page 18 Paragraph 2], Briefly described, a computed aggregation for element language scoring of extracted text being labeled to produce labeled extracted text; [Page 19 Paragraph 3], Briefly described, weighted attributes associated with the element scoring in text extraction).
It would have been obvious to one of ordinary skill in the art before the filing date of the invention to have incorporated – confidence score of the one or more confidence scores and a language processing machine learning model confidence score of the one or more confidence scores, wherein the computed aggregation is used to determine a weight associated with the label during training as suggested by Cohen, Muralidharan, Kim, and Cser. Doing so would be desirable because an implemented loss function to the machine learning model would be helpful to ensure the extracted text is being extracted and matched correctly from a document to a tuple as disclosed in Cohen. The added confidence scores describe the accuracy of how the extracted text is being labeled, so a user can ensure the higher confidence scores of the models get weighted more as they are closer in accuracy than those with lower confidence scores. The combination of Cohen, Muralidharan, Kim, and Cser ensures a model’s accuracy has a score to understand what a model is doing right or wrong.
Regarding claim 18, it is an apparatus claim that corresponds to method claim 4. Therefore, it is rejected for the same reason as claim 4 above.
Claims 7 and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Cohen in view of Muralidharan as applied in the rejection of claims 2 and 8 above, further in view of CSER et al. (WO 2020086773 A1, filed 10/23/2019), hereinafter Cser.
Regarding claim 7, the combination of Cohen and Muralidharan teaches the invention of claim 2.
However, the combination of Cohen and Muralidharan fail to teach – further comprising determining to use the training data for the training of the item extraction machine learning model based on the one or more confidence scores and a confidence score threshold:
In the same field of endeavor Cser teaches:
further comprising determining to use the training data for the training of the item
extraction machine learning model based on the one or more confidence scores and a confidence score threshold: (Cser; [Page 14 Paragraph 2], Briefly described, training data consisting of the element render data, element text data, element code data, or element context data, being determined to use in the extraction model based on the one or more confidence scores of the model output indicating a training data is meeting a sufficient confidence threshold; [Page 28 Paragraph 4], Briefly described, the training data consisting of element attributes being used on the pattern recognition machine learning system for item extraction; [Page 2 Paragraph 5], Briefly described, element attributes consisting of element render data, element text data, element code data, and element context data).
It would have been obvious to one of ordinary skill in the art before the filing date of the invention to have incorporated – further comprising determining to use the training data for the training of the item extraction machine learning model based on the one or more confidence scores and a confidence score threshold as suggested by Cohen, Muralidharan and Cser. Doing so would be desirable because being able to enact a confidence score based on the training data would allow for a confidence score of the tables detected in the structured document to be scored on accuracy. For each label that is assigned to structured text, a confidence score will also indicate if what is labeled is accurate in a supervised learning process. The combination of Cohen, Muralidharan, and Cser have training data that consist of the structured text and the label that goes through a confidence scoring process to determine if the training data is to be used in the model.
Regarding claim 9, the combination of Cohen and Muralidharan teaches the invention as claimed in claim 8 above including wherein the table detection machine learning model is: (Muralidharan; [0038], Briefly described, the machine learning model processing the content segmentation unit to detect if there are action items that are an output of the machine learning model; [0063], Briefly described, content segmentation units generated based on tables, where the machine learning model can process these content segmentation units to determine if they contain tables in the content data);
accepts an image of the structured document as an input: (Muralidharan; [0029], Briefly
described, a document data object being referred to as an image; [0030], Briefly described, a structured document data object described a document data object that structured data to be used as input).
However, the combination of Cohen and Muralidharan fail to teach – a computer vision neural network that; and – and that is trained for object detection through a supervised learning process.
In the same field of endeavor Cser teaches:
a computer vision neural network that: (Cser; [Page 24 Paragraph 4] Briefly described, a
computer vision application that is used in neural networks):
and that is trained for object detection through a supervised learning process: (Cser;
[Page 27 Paragraph 2], Briefly described, an object detection function to be used in the neural network for computer vision application; [Page 28 Paragraph 4], Briefly described, the neural network using the computer vision application using training data for a supervised learning process).
It would have been obvious to one of ordinary skill in the art before the filing date of the invention to have incorporated – a computer vision neural network that; and – and that is trained for object detection through a supervised learning process as suggested by Cohen, Muralidharan and Cser. Doing so would be desirable because the computer vision application that a neural network uses can be the table detection machine learning model as described by Muralidharan. The combination of Cohen, Muralidharan, and Cser can also be used to combine this table detection machine learning model with an object detection function that the computer vision application can use through a supervised learning process as disclosed by Cser. This allows for a structured document to be analyzed under computer vision application to detect objects in the structured document after accepting it as an image.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Subramani et al. (arXiv:2011.13534v2 [cs.CL], 4 Feb 2021) teaches in Page 8, Paragraph 1, extraction of information from a structured document, and Page 8 Paragraph 4, an image of the structured document to be extracted using computer vision.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ROLANDO PATRICK VIRREIRA whose telephone number is (571)270-1570. The examiner can normally be reached Monday – Friday, 8:30AM-5PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached on (571)272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of the published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions. Contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ROLANDO PATRICK VIRREIRA/Examiner, Art Unit 2143
/JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143