DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
Applicant’s Amendments filed on 05/21/2026 has been entered and made of record.
Status of Claims
Currently pending Claim(s):
Amended claim(s):
Canceled claim(s):
1, 4, and 6-14
1, 4, 6-8, and 14
2-3, 5, and 15
Response to Arguments
This office action is responsive to Applicant’s Arguments/Remarks made in an Amendment received on 05/21/2026.
In view of amendments filed on 05/21/2026 to the title, the objection to the specification is withdrawn.
In view of amendments filed on 05/21/2026 to the claim, the objection to claims 1, 14, and 15 are withdrawn.
In view of the new claim amendments and applicant arguments, Remarks filed on 05/21/2026, with respect to the 35 U.S.C. 101 claim rejections have been carefully considered and the claims rejections to claims 1, 2 and 9-15 under 35 U.S.C. 101 are withdrawn.
In view of the Remarks filed on 05/21/2026 the double patenting rejection stands.
In view of applicant's argument, Remarks filed on 05/21/2026, with respect to independent claims 1, 7 and 14 under 35 U.S.C. 103, the arguments have been fully considered but they are not persuasive. The Applicant argues on page 10:
PNG
media_image1.png
276
684
media_image1.png
Greyscale
Applicant's Reply includes substantive amendments to the claims. Claim 1 is amended with, in part, “wherein the cell information includes a row order in the document image”. This changes the claim’s interpretation and scope of the claim of originally filed claim where “acquire cell information…based on coordinates of each of the plurality of cells” was recited in claim 1. This Office action has been updated with new grounds of rejection addressing those amendments. Please note that, as stated by the Applicant on page 9 of the Remarks, previous claim 5 is now amended into claims 1 and 14 – changing the scope of the claim, as a whole. Further Applicant's Arguments/Remarks with respect the pending claims have been considered but are moot because the arguments do not apply to any of the references being used in the current rejection and the arguments are now rejected by newly cited art 'Pinho et al. (US 2023/0237272 A1)' as explained in the body of the rejection below.
Applicant argues on page 11:
PNG
media_image2.png
458
710
media_image2.png
Greyscale
The Examiner respectfully disagrees. The pending claim limitation “…wherein the learning model is a Vision Transformer-based model.”, refers to the machine learning model being a Vision Transformer-based model and is not limited to inputting the sorted cell information into the model. As such, Li teaches ““…wherein the learning model is a Vision Transformer-based model.” at paragraph [0064] “"implementation method may include: inputting multiple face image samples into the visual transformation model, and obtaining the attention matrix corresponding to each face image sample output by each layer of network; merging all the obtained attention matrices to obtain each image.".
Finally, in view of applicant's argument, Remarks filed on 05/21/2026, with respect to claims 6 and 8 under 35 U.S.C. 103, the arguments have been fully considered and are persuasive. Claim 6 and 8 are now objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1 and 4 provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claim 1-4 of copending Application No. 18/696,991 in view of Chatzistamatiou et al. (US 2023/0410543 A1) (hereinafter, “Chatzistamatiou”). Although the claims at issue are not identical, they are not patentably distinct from each other because of the following.
Regarding claims 1 and 4 of the instant application, claims 1-4 of the copending application has the similarity limitation as underlined below:
18/696,981 (Instant Application)
18/696,991 (Copending Application)
Claim 1:
A layout analysis system, comprising at least one processor configured to:
detect a plurality of cells from a document image showing a document including a plurality of components;
acquire cell information relating to at least one of a row or a column of each of the plurality of cells based on coordinates of each of the plurality of cells;
wherein the cell information includes a row order in the document image; and analyze a layout relating to the document by arranging the cell information on each of the plurality of cells in the row order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order, and acquiring a result of analysis of the layout by the learning model.
Claim 1:
A layout analysis system, comprising at least one processor configured to:
detect a cell of each of a plurality of scales from in a document image showing a document including a plurality of components;
acquire cell information relating to the cell of each of the plurality of scales; and
analyze a layout relating to the document based on the cell information on each of the plurality of scales.
Claim 2:
The layout analysis system according to claim 1, wherein the at least one processor is configured to analyze the layout based on a learning model which has learned a for-training layout relating to a for-training document
Claim 3:
The layout analysis system according to claim 2, wherein the at least one processor is configured to analyze the layout by arranging the cell information on each of the plurality of scales under a predetermined condition, inputting the arranged cell information to the learning model, and acquiring a result of analysis of the layout by the learning model.
Claim 4:
The layout analysis system according to claim 1, wherein the learning model is a Vision Transformer-based model.
Claim 4:
The layout analysis system according to claim 3, wherein the learning model is a Vision Transformer-based model.
Table 1
The table (Table 1) above shows that independent claim 1 of the Instant Application is not identical to the claims of 18/696,991 (Copending Application), however, the claims are not patentably distinct but are obvious in view of Chatzistamatiou et al. (US 2023/0410543 A1) (hereinafter, “Chatzistamatiou”). Chatzistamatiou teaches wherein the cell information includes a row order in the document image (Paragraph [0049] “In the first stage 610, signal analysis labeling is used to demarcate the scanned document 702 as shown, with a plurality of horizontal lines 780 and a plurality of vertical lines 770. The plurality of horizontal lines 780 are used to automatically identify each row (e.g., shown as “row 0”, “row 1”, “row 2”. . . “row 14”) by the system. In addition, row classification is performed, labeling a first section 710 (“other”), a second section 720 (“header”), and a third section 730 (“table”).”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify 18/696,991 (Copending Application) reference to include wherein the cell information includes a row order in the document image taught by Chatzistamatiou’s reference. The motivation for doing so would have been to perform row and column segmentation and export data to a desired format and arrangement as suggested by Chatzistamatiou (see Chatzistamatiou, Abstract).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Chatzistamatiou with 18/696,991 (Copending Application) to obtain the invention specified in claim 1.
This is a provisional nonstatutory double patenting rejection.
Claim Objections
Claims 1, 7, and 14 are objected to because of the following informalities:
Claim 1 is objected to because, at the end of the claim, the phrase “by the learning model” seems to belong to “acquiring” and not to “analysis”, therefore the Examiner suggest amending the claim to be “from the learning model”. Claims 7 and 14 are similarly objected to. Appropriate correction is suggested.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 7, 9-10, and 12-14 are rejected under 35 U.S.C. 103 as being unpatentable over Chatzistamatiou et al. (US 2023/0410543 A1) (hereinafter, “Chatzistamatiou”) in view of Pinho et al. (US 2023/0237272 A1) (hereinafter; “Pinho”).
Regarding claim 1, Chatzistamatiou discloses a layout analysis system, comprising at least one processor configured to (Paragraph [0008] “instructions cause the processor to map each item of data extracted from a cell in the first table to a field using semantic data understanding, and to generate a first digital table representing data extracted from the first table for presentation in a user interface.”):
detect a plurality of cells from a document image showing a document including a plurality of components (Paragraph [0026] “the proposed techniques offer an end-to-end solution toward the organization of a set of documents based on similar characteristics. In particular, documents processed by the disclosed extraction system may be generated by photography or scanning of physical documents.”; Paragraph [0031] “data from the image is extracted, even where there are no boundaries for the tables or lists (“boundaryless”). In one embodiment, column segmentation is performed based on signal analysis on column wise mean pixel values, line detection based on Computer Vision (CV) techniques, and clustering models. Furthermore, row segmentation is performed using OCR bounding boxes.”; See Figure 4C);
acquire cell information relating to at least one of a row or a column of each of the plurality of cells [based on coordinates of each of the plurality of cells] (Paragraph [0049] “In the first stage 610, signal analysis labeling is used to demarcate the scanned document 702 as shown, with a plurality of horizontal lines 780 and a plurality of vertical lines 770. The plurality of horizontal lines 780 are used to automatically identify each row (e.g., shown as “row 0”, “row 1”, “row 2”. . . “row 14”) by the system. In addition, row classification is performed, labeling a first section 710 (“other”), a second section 720 (“header”), and a third section 730 (“table”).”);
wherein the cell information includes a row order in the document image (Paragraph [0049] “In the first stage 610, signal analysis labeling is used to demarcate the scanned document 702 as shown, with a plurality of horizontal lines 780 and a plurality of vertical lines 770. The plurality of horizontal lines 780 are used to automatically identify each row (e.g., shown as “row 0”, “row 1”, “row 2”. . . “row 14”) by the system. In addition, row classification is performed, labeling a first section 710 (“other”), a second section 720 (“header”), and a third section 730 (“table”).”); and
analyze a layout relating to the document (Paragraph [0006] “where this has been applied is in museum archiving where documents may not be categorically organized by document type. In addition, the system offers highly accurate table segmentation, where columns are differentiated based on the signal analysis of column-wise mean pixel values, and rows are differentiated based on the textboxes from OCR results.”; Paragraph [0047] “CRF models are used to perform row/column classification (rather than word classification). For rows the task is to cluster into three classes: “header row”, “table row”, and an “other row”. Based on this, the position of the actual table on the page can be determined accurately. The CRF model was selected as it considers the context of information, rather than just a single aspect of data at a time. In other words, the model will attempt to predict a certain goal based not only on the individual row content being focused on, but also on the previous (above) and next (below) row. This larger view allows for improved labeling of each row.”) [by arranging the cell information on each of the plurality of cells in the row order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order], and acquiring a result of analysis of the layout by the learning model (Paragraph [0047] “A set of training data was generated to train the CRF model to classify the rows and columns with such precision before model deployment.”).
However, Chatzistamatiou fails to teach based on coordinates of each of the plurality of cells and by arranging the cell information on each of the plurality of cells in the row order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order.
Pinho teach based on coordinates of each of the plurality of cells (bounding box associated with each of the words in Paragraph [0112] equate to the plurality of cells) (Paragraph [0052] “the word positions can include (x,y) coordinates for the words inside the tables, which confer information about the tables' layout.”; Paragraph [0112] “the word position is based on a coordinate corresponding to a bounding box associated with the each of the one or more words, and the coordinate is in a coordinate space associated with a document containing the table.”) and by arranging the cell information on each of the plurality of cells in the row order (Paragraph [0081] “words are sorted by the given direction. For instance, if the present system is in the process of considering the horizontal direction, words are sorted in {y, x} order.”; See also the Algorithm on Paragraphs [0082-0083] lines 14-16
PNG
media_image3.png
474
677
media_image3.png
Greyscale
), inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order (Paragraph [0050] “Training documents may include documents that have been processed and are annotated with, for example, a reference to one or more field values of interest associated with the document. Training documents may be used to train and verify machine learning models, such as the semantic association model 112 and the table column model 114 (FIG. 1 ).”; Paragraph [0077] “Thus, the example algorithm discussed in further detail below addresses this misalignment in both directions (e.g., the vertical as well as the horizontal direction). Once obtained, target and context words, respectively, can be provided as input and output for the semantic association model (e.g., the embedding model) in a similar way as in the conventional Word2vec algorithm.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Chatzistamatiou’s reference to include based on coordinates of each of the plurality of cells and by arranging the cell information on each of the plurality of cells in the row order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order taught by Pinho’s reference. The motivation for doing so would have been to define the context window so as to capture the neighborhood of a table word in the horizontal and vertical directions as suggested by Pinho (see Pinho, Paragraph [0077]).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Pinho with Chatzistamatiou to obtain the invention specified in claim 1.
Regarding claim 7, Chatzistamatiou discloses a layout analysis system, comprising at least one processor configured to (Paragraph [0008] “instructions cause the processor to map each item of data extracted from a cell in the first table to a field using semantic data understanding, and to generate a first digital table representing data extracted from the first table for presentation in a user interface.”):
detect a plurality of cells from a document image showing a document including a plurality of components (Paragraph [0026] “the proposed techniques offer an end-to-end solution toward the organization of a set of documents based on similar characteristics. In particular, documents processed by the disclosed extraction system may be generated by photography or scanning of physical documents.”; Paragraph [0031] “data from the image is extracted, even where there are no boundaries for the tables or lists (“boundaryless”). In one embodiment, column segmentation is performed based on signal analysis on column wise mean pixel values, line detection based on Computer Vision (CV) techniques, and clustering models. Furthermore, row segmentation is performed using OCR bounding boxes.”; See Figure 4C);
acquire cell information relating to at least one of a row or a column of each of the plurality of cells [based on coordinates of each of the plurality of cells] (Paragraph [0049] “In the first stage 610, signal analysis labeling is used to demarcate the scanned document 702 as shown, with a plurality of horizontal lines 780 and a plurality of vertical lines 770. The plurality of horizontal lines 780 are used to automatically identify each row (e.g., shown as “row 0”, “row 1”, “row 2”. . . “row 14”) by the system. In addition, row classification is performed, labeling a first section 710 (“other”), a second section 720 (“header”), and a third section 730 (“table”).”);
[wherein the cell information includes a column order in the document image];
analyze a layout relating to the document (Paragraph [0006] “where this has been applied is in museum archiving where documents may not be categorically organized by document type. In addition, the system offers highly accurate table segmentation, where columns are differentiated based on the signal analysis of column-wise mean pixel values, and rows are differentiated based on the textboxes from OCR results.”; Paragraph [0047] “CRF models are used to perform row/column classification (rather than word classification). For rows the task is to cluster into three classes: “header row”, “table row”, and an “other row”. Based on this, the position of the actual table on the page can be determined accurately. The CRF model was selected as it considers the context of information, rather than just a single aspect of data at a time. In other words, the model will attempt to predict a certain goal based not only on the individual row content being focused on, but also on the previous (above) and next (below) row. This larger view allows for improved labeling of each row.”) [by arranging the cell information on each of the plurality of cells in the column order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the column order], and acquiring a result of analysis of the layout by the learning model (Paragraph [0047] “A set of training data was generated to train the CRF model to classify the rows and columns with such precision before model deployment.”).
However, Chatzistamatiou fails to teach based on coordinates of each of the plurality of cells; wherein the cell information includes a column order in the document image; and by arranging the cell information on each of the plurality of cells in the column order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the column order.
Pinho teach based on coordinates of each of the plurality of cells (bounding box associated with each of the words in Paragraph [0112] equate to the plurality of cells) (Paragraph [0052] “the word positions can include (x,y) coordinates for the words inside the tables, which confer information about the tables' layout.”; Paragraph [0112] “the word position is based on a coordinate corresponding to a bounding box associated with the each of the one or more words, and the coordinate is in a coordinate space associated with a document containing the table.”); wherein the cell information includes a column order in the document image (See Elements 19-24 of Algorithm on Paragraphs [0082-0083]); and by arranging the cell information on each of the plurality of cells in the column order (Paragraph [0081] “words are sorted by the given direction. For instance, if the present system is in the process of considering the horizontal direction, words are sorted in {y, x} order.”; See also the Algorithm on Paragraphs [0082-0083] lines 18-20
PNG
media_image4.png
474
679
media_image4.png
Greyscale
), inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the column order (Paragraph [0050] “Training documents may include documents that have been processed and are annotated with, for example, a reference to one or more field values of interest associated with the document. Training documents may be used to train and verify machine learning models, such as the semantic association model 112 and the table column model 114 (FIG. 1 ).”; Paragraph [0077] “Thus, the example algorithm discussed in further detail below addresses this misalignment in both directions (e.g., the vertical as well as the horizontal direction). Once obtained, target and context words, respectively, can be provided as input and output for the semantic association model (e.g., the embedding model) in a similar way as in the conventional Word2vec algorithm.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Chatzistamatiou’s reference to include based on coordinates of each of the plurality of cells; wherein the cell information includes a column order in the document image; and by arranging the cell information on each of the plurality of cells in the column order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the column order taught by Pinho’s reference. The motivation for doing so would have been to define the context window so as to capture the neighborhood of a table word in the horizontal and vertical directions as suggested by Pinho (see Pinho, Paragraph [0077]).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Pinho with Chatzistamatiou to obtain the invention specified in claim 7.
Regarding claim 9, which claim 1 is incorporated, Chatzistamatiou fails to teach wherein the at least one processor is configured to acquire the cell information relating to the row of each of the plurality of cells based on a y-coordinate of the each of the plurality of cells so that cells having a distance from each other in a y-axis direction of less than a threshold value are arranged in the same row.
Pinho teaches wherein the at least one processor is configured to acquire the cell information relating to the row of each of the plurality of cells based on a y-coordinate of the each of the plurality of cells so that cells having a distance from each other in a y-axis direction of less than a threshold value are arranged in the same row (See the Algorithm on Paragraphs [0082-0083] lines 3-5.
PNG
media_image5.png
474
759
media_image5.png
Greyscale
)
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Chatzistamatiou’s reference to include wherein the at least one processor is configured to acquire the cell information relating to the row of each of the plurality of cells based on a y-coordinate of the each of the plurality of cells so that cells having a distance from each other in a y-axis direction of less than a threshold value are arranged in the same row taught by Pinho’s reference. The motivation for doing so would have been to define the context window so as to capture the neighborhood of a table word in the horizontal and vertical directions as suggested by Pinho (see Pinho, Paragraph [0077]).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Pinho with Chatzistamatiou to obtain the invention specified in claim 9.
Regarding claim 10, which claim 1 is incorporated, Chatzistamatiou fails to teach wherein the at least one processor is configured to acquire the cell information relating to the column of each of the plurality of cells based on an x-coordinate of the each of the plurality of cells so that cells having a distance from each other in an x-axis direction of less than a threshold value are arranged in the same column.
Pinho teaches wherein the at least one processor is configured to acquire the cell information relating to the column of each of the plurality of cells based on an x-coordinate of the each of the plurality of cells so that cells having a distance from each other in an x-axis direction of less than a threshold value are arranged in the same column (See the Algorithm on Paragraphs [0082-0083] lines 3-5.
PNG
media_image6.png
474
759
media_image6.png
Greyscale
)
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Chatzistamatiou’s reference to include wherein the at least one processor is configured to acquire the cell information relating to the column of each of the plurality of cells based on an x-coordinate of the each of the plurality of cells so that cells having a distance from each other in an x-axis direction of less than a threshold value are arranged in the same column taught by Pinho’s reference. The motivation for doing so would have been to define the context window so as to capture the neighborhood of a table word in the horizontal and vertical directions as suggested by Pinho (see Pinho, Paragraph [0077]).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Pinho with Chatzistamatiou to obtain the invention specified in claim 10.
Regarding claim 12, which claim 9 is incorporated, Chatzistamatiou fails to teach wherein the at least one processor is configured to determine the threshold value based on a size of each of the plurality of cells.
Pinho teaches wherein the at least one processor is configured to determine the threshold value based on a size of each of the plurality of cells (Paragraph [0077] “begin by obtaining a set of words inside a document table along with their associated word positions, e.g., so that such words are associated with their (x, y) coordinates relative to the origin of the table within the coordinate space of the document. For each target word, the present system is configured to select context words within a vertical and a horizontal window of predetermined sizes.”; Paragraph [0081] “Conceptually, the function receives as input all table words 552 in a dataset, grouped by document, a horizontal tolerance 554 and a vertical tolerance 556, and a window size. Then, for each document, the example algorithm is configured to retrieve all word records and calls the previous function to process them in horizontal and vertical orders, from lines 15 to 20.”; See the Algorithm on Paragraphs [0082-0083] lines 3-5).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Chatzistamatiou’s reference to include wherein the at least one processor is configured to determine the threshold value based on a size of each of the plurality of cells taught by Pinho’s reference. The motivation for doing so would have been to define the context window so as to capture the neighborhood of a table word in the horizontal and vertical directions as suggested by Pinho (see Pinho, Paragraph [0077]).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Pinho with Chatzistamatiou to obtain the invention specified in claim 12.
Regarding claim 13, which claim 1 is incorporated, Chatzistamatiou discloses wherein the at least one processor is configured to detect the plurality of cells by executing optical character recognition on the document image (Paragraph [0036] “Document image binarization is often performed in the preprocessing stage of different document image processing related applications such as optical character recognition (OCR) and document image retrieval.”).
Regarding claim 14, Chatzistamatiou discloses a layout analysis method, comprising: detecting a plurality of cells from a document image showing a document including a plurality of components (Paragraph [0026] “the proposed techniques offer an end-to-end solution toward the organization of a set of documents based on similar characteristics. In particular, documents processed by the disclosed extraction system may be generated by photography or scanning of physical documents.”; Paragraph [0031] “data from the image is extracted, even where there are no boundaries for the tables or lists (“boundaryless”). In one embodiment, column segmentation is performed based on signal analysis on column wise mean pixel values, line detection based on Computer Vision (CV) techniques, and clustering models. Furthermore, row segmentation is performed using OCR bounding boxes.”; See Figure 4C);
acquiring cell information relating to at least one of a row or a column of each of the plurality of cells [based on coordinates of each of the plurality of cells] (Paragraph [0049] “In the first stage 610, signal analysis labeling is used to demarcate the scanned document 702 as shown, with a plurality of horizontal lines 780 and a plurality of vertical lines 770. The plurality of horizontal lines 780 are used to automatically identify each row (e.g., shown as “row 0”, “row 1”, “row 2”. . . “row 14”) by the system. In addition, row classification is performed, labeling a first section 710 (“other”), a second section 720 (“header”), and a third section 730 (“table”).”);
wherein the cell information includes a row order in the document image (Paragraph [0049] “In the first stage 610, signal analysis labeling is used to demarcate the scanned document 702 as shown, with a plurality of horizontal lines 780 and a plurality of vertical lines 770. The plurality of horizontal lines 780 are used to automatically identify each row (e.g., shown as “row 0”, “row 1”, “row 2”. . . “row 14”) by the system. In addition, row classification is performed, labeling a first section 710 (“other”), a second section 720 (“header”), and a third section 730 (“table”).”), and
analyzing a layout relating to the document (Paragraph [0006] “where this has been applied is in museum archiving where documents may not be categorically organized by document type. In addition, the system offers highly accurate table segmentation, where columns are differentiated based on the signal analysis of column-wise mean pixel values, and rows are differentiated based on the textboxes from OCR results.”; Paragraph [0047] “CRF models are used to perform row/column classification (rather than word classification). For rows the task is to cluster into three classes: “header row”, “table row”, and an “other row”. Based on this, the position of the actual table on the page can be determined accurately. The CRF model was selected as it considers the context of information, rather than just a single aspect of data at a time. In other words, the model will attempt to predict a certain goal based not only on the individual row content being focused on, but also on the previous (above) and next (below) row. This larger view allows for improved labeling of each row.”) [by arranging the cell information on each of the plurality of cells in the row order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order], and acquiring a result of analysis of the layout by the learning model (Paragraph [0047] “A set of training data was generated to train the CRF model to classify the rows and columns with such precision before model deployment.”).
However, Chatzistamatiou fails to teach based on coordinates of each of the plurality of cells and by arranging the cell information on each of the plurality of cells in the row order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order.
Pinho teach based on coordinates of each of the plurality of cells (bounding box associated with each of the words in Paragraph [0112] equate to the plurality of cells) (Paragraph [0052] “the word positions can include (x,y) coordinates for the words inside the tables, which confer information about the tables' layout.”; Paragraph [0112] “the word position is based on a coordinate corresponding to a bounding box associated with the each of the one or more words, and the coordinate is in a coordinate space associated with a document containing the table.”) and by arranging the cell information on each of the plurality of cells in the row order (Paragraph [0081] “words are sorted by the given direction. For instance, if the present system is in the process of considering the horizontal direction, words are sorted in {y, x} order.”; See also the Algorithm on Paragraphs [0082-0083]
PNG
media_image3.png
474
677
media_image3.png
Greyscale
), inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order (Paragraph [0050] “Training documents may include documents that have been processed and are annotated with, for example, a reference to one or more field values of interest associated with the document. Training documents may be used to train and verify machine learning models, such as the semantic association model 112 and the table column model 114 (FIG. 1 ).”; Paragraph [0077] “Thus, the example algorithm discussed in further detail below addresses this misalignment in both directions (e.g., the vertical as well as the horizontal direction). Once obtained, target and context words, respectively, can be provided as input and output for the semantic association model (e.g., the embedding model) in a similar way as in the conventional Word2vec algorithm.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Chatzistamatiou’s reference to include based on coordinates of each of the plurality of cells and by arranging the cell information on each of the plurality of cells in the row order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order taught by Pinho’s reference. The motivation for doing so would have been to define the context window so as to capture the neighborhood of a table word in the horizontal and vertical directions as suggested by Pinho (see Pinho, Paragraph [0077]).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Pinho with Chatzistamatiou to obtain the invention specified in claim 14.
Claims 4 is rejected under 35 U.S.C. 103 as being unpatentable over Chatzistamatiou et al. (US 2023/0410543 A1) (hereinafter, “Chatzistamatiou”) in view of Pinho et al. (US 2023/0237272 A1) (hereinafter; “Pinho”) as applied to claim 1 above, and further in view of Li (CN 113,901,904 A).
Regarding claim 4, which claim 1 is incorporated, Chatzistamatiou and Pinho fail to teach wherein the learning model is a Vision Transformer-based model.
Li teaches wherein the learning model is a Vision Transformer-based model (Paragraph [0064] “implementation method may include: inputting multiple face image samples into the visual transformation model, and obtaining the attention matrix corresponding to each face image sample output by each layer of network; merging all the obtained attention matrices to obtain each image.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective
filing date to modify Chatzistamatiou in view of Pinho to include wherein the learning model is a Vision Transformer-based model taught by Li’s reference. The motivation for doing so would have been to divide the document into multiple patches as suggested by Li (see Li, Paragraph [0052]).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Li with Chatzistamatiou and Pinho to obtain the invention specified in claim 4.
Claims 11 is rejected under 35 U.S.C. 103 as being unpatentable over Chatzistamatiou et al. (US 2023/0410543 A1) (hereinafter, “Chatzistamatiou”) in view of Pinho et al. (US 2023/0237272 A1) (hereinafter; “Pinho”) as applied to claim 9 above, and further in view of Yebes Torres et al. (US 2023/0005286 A1) (hereinafter “Torres”).
Regarding claim 11, which claim 9 is incorporated, Chatzistamatiou fails to teach wherein the at least one processor is configured to determine the threshold value based on a size of the whole document.
Pinho teaches wherein the at least one processor is configured to determine the threshold value (Paragraph [0081] “Conceptually, the function receives as input all table words 552 in a dataset, grouped by document, a horizontal tolerance 554 and a vertical tolerance 556, and a window size. Then, for each document, the example algorithm is configured to retrieve all word records and calls the previous function to process them in horizontal and vertical orders, from lines 15 to 20. It will be noted that, advantageously, the processing order is defined solely by the field_order parameter. With the horizontal and vertical target 552 and context word sets 558, the only remaining action is to concatenate and accumulate them into lists containing target and context words for all documents.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective
filing date to modify Chatzistamatiou reference to include wherein the at least one processor is configured to determine the threshold value taught by Pinho’s reference. The motivation for doing so would have been to define the context window so as to capture the neighborhood of a table word in the horizontal and vertical directions as suggested by Pinho (see Pinho, Paragraph [0077]).
However, Chatzistamatiou and Pinho both fail to teach based on a size of the whole document.
Torres teaches based on a size of the whole document (Element 108 in Figure 9 equates to the whole document) (Paragraph [0133] “the bounding box generating circuitry 316 outputs a plurality of bounding boxes including respective coordinates corresponding to lines of the receipt region.”; Paragraph [0164] “example of FIG. 9…receives an example receipt image 108 (e.g., from the basket datastore 112)…The extraction circuitry 118 applies an example regions detection model 306 to the receipt image 108 to detect an example receipt region 902 and an example products region 904. In some examples, the extraction circuitry 118 applies a cropping operation (e.g., via the image cropping circuitry 308) to the receipt image 108 based on the detected regions).
Figure 9:
PNG
media_image7.png
714
540
media_image7.png
Greyscale
Therefore, it would have been obvious to one of ordinary skill of the art before the effective
filing date to modify Chatzistamatiou in view of Pinho to include based on a size of the whole document taught by Torres’s reference. The motivation for doing so would have been to align the respective coordinates to the document region as suggested by Torres (see Torres, Paragraph [0133]).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Torres with Chatzistamatiou and Pinho to obtain the invention specified in claim 11.
Allowable Subject Matter
Claims 6 and 8 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claims 6 and 8 contain subject matter that is not disclosed or made obvious in the cited art:
In regards to claim 6, when considering claim 6 as a whole, the features highlighted below are considered an improvement over the prior art and have not been found to be anticipated or rendered obvious by a combination of the prior art:
“[…]
sort the cell information on each of the plurality of cells based on the row order of the each of the plurality of cells,
insert predetermined row change information into a portion of the cell information which has a row change in the row order, and
input the cell information having the inserted predetermined row change information to the learning model in the row order.”
In regards to claim 8, when considering claim 8 as a whole, the features highlighted below are considered an improvement over the prior art and have not been found to be anticipated or rendered obvious by a combination of the prior art:
“[…]
sort the cell information on each of the plurality of cells based on the column order of the each of the plurality of cells,
insert predetermined column change information into a portion of the cell information which has a column change in the column order, and
input the cell information having the inserted predetermined column change information to the learning model in the column order.”
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Zakharov (US 2022/0067320 A1) discloses a document classification technique that converts text content to images and apply image classification to the images. Zakharov further discloses a system comprising an image generator configured to generate graphical code based on topological features. The topological features comprise of tables where text is arranged in columns and rows, graphics such as charts, logos, etc., and blank spaces separating text and graphics.
Inaki et al. (US 5,835,916 A) discloses an apparatus comprising a processing unit which allows an operator to designate a single cell in the table and to relocate the designated cell in an area of the table together with text data registered in the designated cell.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to UROOJ FATIMA whose telephone number is (571)272-2096. The examiner can normally be reached M-F 8:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Henok Shiferaw can be reached at (571) 272-4637. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/UROOJ FATIMA/Examiner, Art Unit 2676
/Henok Shiferaw/Supervisory Patent Examiner, Art Unit 2676