Prosecution Insights
Last updated: October 02, 2026
Application No. 18/696,981

LAYOUT ANALYSIS SYSTEM, LAYOUT ANALYSIS METHOD, AND PROGRAM

Final Rejection §103§DOUBLEPATENT
Filed
Mar 28, 2024
Priority
Aug 30, 2022 — nonprovisional of PCTJP2022032643
Examiner
FATIMA, UROOJ
Art Unit
2676
Tech Center
2600 — Communications
Assignee
Rakuten Group Inc.
OA Round
2 (Final)
75%
Grant Probability
Favorable
3-4
OA Rounds
2m
Est. Remaining
75%
With Interview

Examiner Intelligence

Grants 75% — above average
75%
Career Allowance Rate
6 granted / 8 resolved
+13.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
23 currently pending
Career history
29
Total Applications
across all art units

Statute-Specific Performance

§101
13.3%
-26.7% vs TC avg
§103
60.8%
+20.8% vs TC avg
§102
7.7%
-32.3% vs TC avg
§112
14.7%
-25.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 8 resolved cases

Office Action

§103 §DOUBLEPATENT
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment Applicant’s Amendments filed on 05/21/2026 has been entered and made of record. Status of Claims Currently pending Claim(s): Amended claim(s): Canceled claim(s): 1, 4, and 6-14 1, 4, 6-8, and 14 2-3, 5, and 15 Response to Arguments This office action is responsive to Applicant’s Arguments/Remarks made in an Amendment received on 05/21/2026. In view of amendments filed on 05/21/2026 to the title, the objection to the specification is withdrawn. In view of amendments filed on 05/21/2026 to the claim, the objection to claims 1, 14, and 15 are withdrawn. In view of the new claim amendments and applicant arguments, Remarks filed on 05/21/2026, with respect to the 35 U.S.C. 101 claim rejections have been carefully considered and the claims rejections to claims 1, 2 and 9-15 under 35 U.S.C. 101 are withdrawn. In view of the Remarks filed on 05/21/2026 the double patenting rejection stands. In view of applicant's argument, Remarks filed on 05/21/2026, with respect to independent claims 1, 7 and 14 under 35 U.S.C. 103, the arguments have been fully considered but they are not persuasive. The Applicant argues on page 10: PNG media_image1.png 276 684 media_image1.png Greyscale Applicant's Reply includes substantive amendments to the claims. Claim 1 is amended with, in part, “wherein the cell information includes a row order in the document image”. This changes the claim’s interpretation and scope of the claim of originally filed claim where “acquire cell information…based on coordinates of each of the plurality of cells” was recited in claim 1. This Office action has been updated with new grounds of rejection addressing those amendments. Please note that, as stated by the Applicant on page 9 of the Remarks, previous claim 5 is now amended into claims 1 and 14 – changing the scope of the claim, as a whole. Further Applicant's Arguments/Remarks with respect the pending claims have been considered but are moot because the arguments do not apply to any of the references being used in the current rejection and the arguments are now rejected by newly cited art 'Pinho et al. (US 2023/0237272 A1)' as explained in the body of the rejection below. Applicant argues on page 11: PNG media_image2.png 458 710 media_image2.png Greyscale The Examiner respectfully disagrees. The pending claim limitation “…wherein the learning model is a Vision Transformer-based model.”, refers to the machine learning model being a Vision Transformer-based model and is not limited to inputting the sorted cell information into the model. As such, Li teaches ““…wherein the learning model is a Vision Transformer-based model.” at paragraph [0064] “"implementation method may include: inputting multiple face image samples into the visual transformation model, and obtaining the attention matrix corresponding to each face image sample output by each layer of network; merging all the obtained attention matrices to obtain each image.". Finally, in view of applicant's argument, Remarks filed on 05/21/2026, with respect to claims 6 and 8 under 35 U.S.C. 103, the arguments have been fully considered and are persuasive. Claim 6 and 8 are now objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 1 and 4 provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claim 1-4 of copending Application No. 18/696,991 in view of Chatzistamatiou et al. (US 2023/0410543 A1) (hereinafter, “Chatzistamatiou”). Although the claims at issue are not identical, they are not patentably distinct from each other because of the following. Regarding claims 1 and 4 of the instant application, claims 1-4 of the copending application has the similarity limitation as underlined below: 18/696,981 (Instant Application) 18/696,991 (Copending Application) Claim 1: A layout analysis system, comprising at least one processor configured to: detect a plurality of cells from a document image showing a document including a plurality of components; acquire cell information relating to at least one of a row or a column of each of the plurality of cells based on coordinates of each of the plurality of cells; wherein the cell information includes a row order in the document image; and analyze a layout relating to the document by arranging the cell information on each of the plurality of cells in the row order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order, and acquiring a result of analysis of the layout by the learning model. Claim 1: A layout analysis system, comprising at least one processor configured to: detect a cell of each of a plurality of scales from in a document image showing a document including a plurality of components; acquire cell information relating to the cell of each of the plurality of scales; and analyze a layout relating to the document based on the cell information on each of the plurality of scales. Claim 2: The layout analysis system according to claim 1, wherein the at least one processor is configured to analyze the layout based on a learning model which has learned a for-training layout relating to a for-training document Claim 3: The layout analysis system according to claim 2, wherein the at least one processor is configured to analyze the layout by arranging the cell information on each of the plurality of scales under a predetermined condition, inputting the arranged cell information to the learning model, and acquiring a result of analysis of the layout by the learning model. Claim 4: The layout analysis system according to claim 1, wherein the learning model is a Vision Transformer-based model. Claim 4: The layout analysis system according to claim 3, wherein the learning model is a Vision Transformer-based model. Table 1 The table (Table 1) above shows that independent claim 1 of the Instant Application is not identical to the claims of 18/696,991 (Copending Application), however, the claims are not patentably distinct but are obvious in view of Chatzistamatiou et al. (US 2023/0410543 A1) (hereinafter, “Chatzistamatiou”). Chatzistamatiou teaches wherein the cell information includes a row order in the document image (Paragraph [0049] “In the first stage 610, signal analysis labeling is used to demarcate the scanned document 702 as shown, with a plurality of horizontal lines 780 and a plurality of vertical lines 770. The plurality of horizontal lines 780 are used to automatically identify each row (e.g., shown as “row 0”, “row 1”, “row 2”. . . “row 14”) by the system. In addition, row classification is performed, labeling a first section 710 (“other”), a second section 720 (“header”), and a third section 730 (“table”).”). Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify 18/696,991 (Copending Application) reference to include wherein the cell information includes a row order in the document image taught by Chatzistamatiou’s reference. The motivation for doing so would have been to perform row and column segmentation and export data to a desired format and arrangement as suggested by Chatzistamatiou (see Chatzistamatiou, Abstract). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Chatzistamatiou with 18/696,991 (Copending Application) to obtain the invention specified in claim 1. This is a provisional nonstatutory double patenting rejection. Claim Objections Claims 1, 7, and 14 are objected to because of the following informalities: Claim 1 is objected to because, at the end of the claim, the phrase “by the learning model” seems to belong to “acquiring” and not to “analysis”, therefore the Examiner suggest amending the claim to be “from the learning model”. Claims 7 and 14 are similarly objected to. Appropriate correction is suggested. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 7, 9-10, and 12-14 are rejected under 35 U.S.C. 103 as being unpatentable over Chatzistamatiou et al. (US 2023/0410543 A1) (hereinafter, “Chatzistamatiou”) in view of Pinho et al. (US 2023/0237272 A1) (hereinafter; “Pinho”). Regarding claim 1, Chatzistamatiou discloses a layout analysis system, comprising at least one processor configured to (Paragraph [0008] “instructions cause the processor to map each item of data extracted from a cell in the first table to a field using semantic data understanding, and to generate a first digital table representing data extracted from the first table for presentation in a user interface.”): detect a plurality of cells from a document image showing a document including a plurality of components (Paragraph [0026] “the proposed techniques offer an end-to-end solution toward the organization of a set of documents based on similar characteristics. In particular, documents processed by the disclosed extraction system may be generated by photography or scanning of physical documents.”; Paragraph [0031] “data from the image is extracted, even where there are no boundaries for the tables or lists (“boundaryless”). In one embodiment, column segmentation is performed based on signal analysis on column wise mean pixel values, line detection based on Computer Vision (CV) techniques, and clustering models. Furthermore, row segmentation is performed using OCR bounding boxes.”; See Figure 4C); acquire cell information relating to at least one of a row or a column of each of the plurality of cells [based on coordinates of each of the plurality of cells] (Paragraph [0049] “In the first stage 610, signal analysis labeling is used to demarcate the scanned document 702 as shown, with a plurality of horizontal lines 780 and a plurality of vertical lines 770. The plurality of horizontal lines 780 are used to automatically identify each row (e.g., shown as “row 0”, “row 1”, “row 2”. . . “row 14”) by the system. In addition, row classification is performed, labeling a first section 710 (“other”), a second section 720 (“header”), and a third section 730 (“table”).”); wherein the cell information includes a row order in the document image (Paragraph [0049] “In the first stage 610, signal analysis labeling is used to demarcate the scanned document 702 as shown, with a plurality of horizontal lines 780 and a plurality of vertical lines 770. The plurality of horizontal lines 780 are used to automatically identify each row (e.g., shown as “row 0”, “row 1”, “row 2”. . . “row 14”) by the system. In addition, row classification is performed, labeling a first section 710 (“other”), a second section 720 (“header”), and a third section 730 (“table”).”); and analyze a layout relating to the document (Paragraph [0006] “where this has been applied is in museum archiving where documents may not be categorically organized by document type. In addition, the system offers highly accurate table segmentation, where columns are differentiated based on the signal analysis of column-wise mean pixel values, and rows are differentiated based on the textboxes from OCR results.”; Paragraph [0047] “CRF models are used to perform row/column classification (rather than word classification). For rows the task is to cluster into three classes: “header row”, “table row”, and an “other row”. Based on this, the position of the actual table on the page can be determined accurately. The CRF model was selected as it considers the context of information, rather than just a single aspect of data at a time. In other words, the model will attempt to predict a certain goal based not only on the individual row content being focused on, but also on the previous (above) and next (below) row. This larger view allows for improved labeling of each row.”) [by arranging the cell information on each of the plurality of cells in the row order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order], and acquiring a result of analysis of the layout by the learning model (Paragraph [0047] “A set of training data was generated to train the CRF model to classify the rows and columns with such precision before model deployment.”). However, Chatzistamatiou fails to teach based on coordinates of each of the plurality of cells and by arranging the cell information on each of the plurality of cells in the row order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order. Pinho teach based on coordinates of each of the plurality of cells (bounding box associated with each of the words in Paragraph [0112] equate to the plurality of cells) (Paragraph [0052] “the word positions can include (x,y) coordinates for the words inside the tables, which confer information about the tables' layout.”; Paragraph [0112] “the word position is based on a coordinate corresponding to a bounding box associated with the each of the one or more words, and the coordinate is in a coordinate space associated with a document containing the table.”) and by arranging the cell information on each of the plurality of cells in the row order (Paragraph [0081] “words are sorted by the given direction. For instance, if the present system is in the process of considering the horizontal direction, words are sorted in {y, x} order.”; See also the Algorithm on Paragraphs [0082-0083] lines 14-16 PNG media_image3.png 474 677 media_image3.png Greyscale ), inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order (Paragraph [0050] “Training documents may include documents that have been processed and are annotated with, for example, a reference to one or more field values of interest associated with the document. Training documents may be used to train and verify machine learning models, such as the semantic association model 112 and the table column model 114 (FIG. 1 ).”; Paragraph [0077] “Thus, the example algorithm discussed in further detail below addresses this misalignment in both directions (e.g., the vertical as well as the horizontal direction). Once obtained, target and context words, respectively, can be provided as input and output for the semantic association model (e.g., the embedding model) in a similar way as in the conventional Word2vec algorithm.”). Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Chatzistamatiou’s reference to include based on coordinates of each of the plurality of cells and by arranging the cell information on each of the plurality of cells in the row order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order taught by Pinho’s reference. The motivation for doing so would have been to define the context window so as to capture the neighborhood of a table word in the horizontal and vertical directions as suggested by Pinho (see Pinho, Paragraph [0077]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Pinho with Chatzistamatiou to obtain the invention specified in claim 1. Regarding claim 7, Chatzistamatiou discloses a layout analysis system, comprising at least one processor configured to (Paragraph [0008] “instructions cause the processor to map each item of data extracted from a cell in the first table to a field using semantic data understanding, and to generate a first digital table representing data extracted from the first table for presentation in a user interface.”): detect a plurality of cells from a document image showing a document including a plurality of components (Paragraph [0026] “the proposed techniques offer an end-to-end solution toward the organization of a set of documents based on similar characteristics. In particular, documents processed by the disclosed extraction system may be generated by photography or scanning of physical documents.”; Paragraph [0031] “data from the image is extracted, even where there are no boundaries for the tables or lists (“boundaryless”). In one embodiment, column segmentation is performed based on signal analysis on column wise mean pixel values, line detection based on Computer Vision (CV) techniques, and clustering models. Furthermore, row segmentation is performed using OCR bounding boxes.”; See Figure 4C); acquire cell information relating to at least one of a row or a column of each of the plurality of cells [based on coordinates of each of the plurality of cells] (Paragraph [0049] “In the first stage 610, signal analysis labeling is used to demarcate the scanned document 702 as shown, with a plurality of horizontal lines 780 and a plurality of vertical lines 770. The plurality of horizontal lines 780 are used to automatically identify each row (e.g., shown as “row 0”, “row 1”, “row 2”. . . “row 14”) by the system. In addition, row classification is performed, labeling a first section 710 (“other”), a second section 720 (“header”), and a third section 730 (“table”).”); [wherein the cell information includes a column order in the document image]; analyze a layout relating to the document (Paragraph [0006] “where this has been applied is in museum archiving where documents may not be categorically organized by document type. In addition, the system offers highly accurate table segmentation, where columns are differentiated based on the signal analysis of column-wise mean pixel values, and rows are differentiated based on the textboxes from OCR results.”; Paragraph [0047] “CRF models are used to perform row/column classification (rather than word classification). For rows the task is to cluster into three classes: “header row”, “table row”, and an “other row”. Based on this, the position of the actual table on the page can be determined accurately. The CRF model was selected as it considers the context of information, rather than just a single aspect of data at a time. In other words, the model will attempt to predict a certain goal based not only on the individual row content being focused on, but also on the previous (above) and next (below) row. This larger view allows for improved labeling of each row.”) [by arranging the cell information on each of the plurality of cells in the column order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the column order], and acquiring a result of analysis of the layout by the learning model (Paragraph [0047] “A set of training data was generated to train the CRF model to classify the rows and columns with such precision before model deployment.”). However, Chatzistamatiou fails to teach based on coordinates of each of the plurality of cells; wherein the cell information includes a column order in the document image; and by arranging the cell information on each of the plurality of cells in the column order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the column order. Pinho teach based on coordinates of each of the plurality of cells (bounding box associated with each of the words in Paragraph [0112] equate to the plurality of cells) (Paragraph [0052] “the word positions can include (x,y) coordinates for the words inside the tables, which confer information about the tables' layout.”; Paragraph [0112] “the word position is based on a coordinate corresponding to a bounding box associated with the each of the one or more words, and the coordinate is in a coordinate space associated with a document containing the table.”); wherein the cell information includes a column order in the document image (See Elements 19-24 of Algorithm on Paragraphs [0082-0083]); and by arranging the cell information on each of the plurality of cells in the column order (Paragraph [0081] “words are sorted by the given direction. For instance, if the present system is in the process of considering the horizontal direction, words are sorted in {y, x} order.”; See also the Algorithm on Paragraphs [0082-0083] lines 18-20 PNG media_image4.png 474 679 media_image4.png Greyscale ), inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the column order (Paragraph [0050] “Training documents may include documents that have been processed and are annotated with, for example, a reference to one or more field values of interest associated with the document. Training documents may be used to train and verify machine learning models, such as the semantic association model 112 and the table column model 114 (FIG. 1 ).”; Paragraph [0077] “Thus, the example algorithm discussed in further detail below addresses this misalignment in both directions (e.g., the vertical as well as the horizontal direction). Once obtained, target and context words, respectively, can be provided as input and output for the semantic association model (e.g., the embedding model) in a similar way as in the conventional Word2vec algorithm.”). Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Chatzistamatiou’s reference to include based on coordinates of each of the plurality of cells; wherein the cell information includes a column order in the document image; and by arranging the cell information on each of the plurality of cells in the column order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the column order taught by Pinho’s reference. The motivation for doing so would have been to define the context window so as to capture the neighborhood of a table word in the horizontal and vertical directions as suggested by Pinho (see Pinho, Paragraph [0077]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Pinho with Chatzistamatiou to obtain the invention specified in claim 7. Regarding claim 9, which claim 1 is incorporated, Chatzistamatiou fails to teach wherein the at least one processor is configured to acquire the cell information relating to the row of each of the plurality of cells based on a y-coordinate of the each of the plurality of cells so that cells having a distance from each other in a y-axis direction of less than a threshold value are arranged in the same row. Pinho teaches wherein the at least one processor is configured to acquire the cell information relating to the row of each of the plurality of cells based on a y-coordinate of the each of the plurality of cells so that cells having a distance from each other in a y-axis direction of less than a threshold value are arranged in the same row (See the Algorithm on Paragraphs [0082-0083] lines 3-5. PNG media_image5.png 474 759 media_image5.png Greyscale ) Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Chatzistamatiou’s reference to include wherein the at least one processor is configured to acquire the cell information relating to the row of each of the plurality of cells based on a y-coordinate of the each of the plurality of cells so that cells having a distance from each other in a y-axis direction of less than a threshold value are arranged in the same row taught by Pinho’s reference. The motivation for doing so would have been to define the context window so as to capture the neighborhood of a table word in the horizontal and vertical directions as suggested by Pinho (see Pinho, Paragraph [0077]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Pinho with Chatzistamatiou to obtain the invention specified in claim 9. Regarding claim 10, which claim 1 is incorporated, Chatzistamatiou fails to teach wherein the at least one processor is configured to acquire the cell information relating to the column of each of the plurality of cells based on an x-coordinate of the each of the plurality of cells so that cells having a distance from each other in an x-axis direction of less than a threshold value are arranged in the same column. Pinho teaches wherein the at least one processor is configured to acquire the cell information relating to the column of each of the plurality of cells based on an x-coordinate of the each of the plurality of cells so that cells having a distance from each other in an x-axis direction of less than a threshold value are arranged in the same column (See the Algorithm on Paragraphs [0082-0083] lines 3-5. PNG media_image6.png 474 759 media_image6.png Greyscale ) Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Chatzistamatiou’s reference to include wherein the at least one processor is configured to acquire the cell information relating to the column of each of the plurality of cells based on an x-coordinate of the each of the plurality of cells so that cells having a distance from each other in an x-axis direction of less than a threshold value are arranged in the same column taught by Pinho’s reference. The motivation for doing so would have been to define the context window so as to capture the neighborhood of a table word in the horizontal and vertical directions as suggested by Pinho (see Pinho, Paragraph [0077]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Pinho with Chatzistamatiou to obtain the invention specified in claim 10. Regarding claim 12, which claim 9 is incorporated, Chatzistamatiou fails to teach wherein the at least one processor is configured to determine the threshold value based on a size of each of the plurality of cells. Pinho teaches wherein the at least one processor is configured to determine the threshold value based on a size of each of the plurality of cells (Paragraph [0077] “begin by obtaining a set of words inside a document table along with their associated word positions, e.g., so that such words are associated with their (x, y) coordinates relative to the origin of the table within the coordinate space of the document. For each target word, the present system is configured to select context words within a vertical and a horizontal window of predetermined sizes.”; Paragraph [0081] “Conceptually, the function receives as input all table words 552 in a dataset, grouped by document, a horizontal tolerance 554 and a vertical tolerance 556, and a window size. Then, for each document, the example algorithm is configured to retrieve all word records and calls the previous function to process them in horizontal and vertical orders, from lines 15 to 20.”; See the Algorithm on Paragraphs [0082-0083] lines 3-5). Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Chatzistamatiou’s reference to include wherein the at least one processor is configured to determine the threshold value based on a size of each of the plurality of cells taught by Pinho’s reference. The motivation for doing so would have been to define the context window so as to capture the neighborhood of a table word in the horizontal and vertical directions as suggested by Pinho (see Pinho, Paragraph [0077]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Pinho with Chatzistamatiou to obtain the invention specified in claim 12. Regarding claim 13, which claim 1 is incorporated, Chatzistamatiou discloses wherein the at least one processor is configured to detect the plurality of cells by executing optical character recognition on the document image (Paragraph [0036] “Document image binarization is often performed in the preprocessing stage of different document image processing related applications such as optical character recognition (OCR) and document image retrieval.”). Regarding claim 14, Chatzistamatiou discloses a layout analysis method, comprising: detecting a plurality of cells from a document image showing a document including a plurality of components (Paragraph [0026] “the proposed techniques offer an end-to-end solution toward the organization of a set of documents based on similar characteristics. In particular, documents processed by the disclosed extraction system may be generated by photography or scanning of physical documents.”; Paragraph [0031] “data from the image is extracted, even where there are no boundaries for the tables or lists (“boundaryless”). In one embodiment, column segmentation is performed based on signal analysis on column wise mean pixel values, line detection based on Computer Vision (CV) techniques, and clustering models. Furthermore, row segmentation is performed using OCR bounding boxes.”; See Figure 4C); acquiring cell information relating to at least one of a row or a column of each of the plurality of cells [based on coordinates of each of the plurality of cells] (Paragraph [0049] “In the first stage 610, signal analysis labeling is used to demarcate the scanned document 702 as shown, with a plurality of horizontal lines 780 and a plurality of vertical lines 770. The plurality of horizontal lines 780 are used to automatically identify each row (e.g., shown as “row 0”, “row 1”, “row 2”. . . “row 14”) by the system. In addition, row classification is performed, labeling a first section 710 (“other”), a second section 720 (“header”), and a third section 730 (“table”).”); wherein the cell information includes a row order in the document image (Paragraph [0049] “In the first stage 610, signal analysis labeling is used to demarcate the scanned document 702 as shown, with a plurality of horizontal lines 780 and a plurality of vertical lines 770. The plurality of horizontal lines 780 are used to automatically identify each row (e.g., shown as “row 0”, “row 1”, “row 2”. . . “row 14”) by the system. In addition, row classification is performed, labeling a first section 710 (“other”), a second section 720 (“header”), and a third section 730 (“table”).”), and analyzing a layout relating to the document (Paragraph [0006] “where this has been applied is in museum archiving where documents may not be categorically organized by document type. In addition, the system offers highly accurate table segmentation, where columns are differentiated based on the signal analysis of column-wise mean pixel values, and rows are differentiated based on the textboxes from OCR results.”; Paragraph [0047] “CRF models are used to perform row/column classification (rather than word classification). For rows the task is to cluster into three classes: “header row”, “table row”, and an “other row”. Based on this, the position of the actual table on the page can be determined accurately. The CRF model was selected as it considers the context of information, rather than just a single aspect of data at a time. In other words, the model will attempt to predict a certain goal based not only on the individual row content being focused on, but also on the previous (above) and next (below) row. This larger view allows for improved labeling of each row.”) [by arranging the cell information on each of the plurality of cells in the row order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order], and acquiring a result of analysis of the layout by the learning model (Paragraph [0047] “A set of training data was generated to train the CRF model to classify the rows and columns with such precision before model deployment.”). However, Chatzistamatiou fails to teach based on coordinates of each of the plurality of cells and by arranging the cell information on each of the plurality of cells in the row order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order. Pinho teach based on coordinates of each of the plurality of cells (bounding box associated with each of the words in Paragraph [0112] equate to the plurality of cells) (Paragraph [0052] “the word positions can include (x,y) coordinates for the words inside the tables, which confer information about the tables' layout.”; Paragraph [0112] “the word position is based on a coordinate corresponding to a bounding box associated with the each of the one or more words, and the coordinate is in a coordinate space associated with a document containing the table.”) and by arranging the cell information on each of the plurality of cells in the row order (Paragraph [0081] “words are sorted by the given direction. For instance, if the present system is in the process of considering the horizontal direction, words are sorted in {y, x} order.”; See also the Algorithm on Paragraphs [0082-0083] PNG media_image3.png 474 677 media_image3.png Greyscale ), inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order (Paragraph [0050] “Training documents may include documents that have been processed and are annotated with, for example, a reference to one or more field values of interest associated with the document. Training documents may be used to train and verify machine learning models, such as the semantic association model 112 and the table column model 114 (FIG. 1 ).”; Paragraph [0077] “Thus, the example algorithm discussed in further detail below addresses this misalignment in both directions (e.g., the vertical as well as the horizontal direction). Once obtained, target and context words, respectively, can be provided as input and output for the semantic association model (e.g., the embedding model) in a similar way as in the conventional Word2vec algorithm.”). Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Chatzistamatiou’s reference to include based on coordinates of each of the plurality of cells and by arranging the cell information on each of the plurality of cells in the row order, inputting the arranged cell information into a learning model which has learned a for-training layout relating to a for-training document in the row order taught by Pinho’s reference. The motivation for doing so would have been to define the context window so as to capture the neighborhood of a table word in the horizontal and vertical directions as suggested by Pinho (see Pinho, Paragraph [0077]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Pinho with Chatzistamatiou to obtain the invention specified in claim 14. Claims 4 is rejected under 35 U.S.C. 103 as being unpatentable over Chatzistamatiou et al. (US 2023/0410543 A1) (hereinafter, “Chatzistamatiou”) in view of Pinho et al. (US 2023/0237272 A1) (hereinafter; “Pinho”) as applied to claim 1 above, and further in view of Li (CN 113,901,904 A). Regarding claim 4, which claim 1 is incorporated, Chatzistamatiou and Pinho fail to teach wherein the learning model is a Vision Transformer-based model. Li teaches wherein the learning model is a Vision Transformer-based model (Paragraph [0064] “implementation method may include: inputting multiple face image samples into the visual transformation model, and obtaining the attention matrix corresponding to each face image sample output by each layer of network; merging all the obtained attention matrices to obtain each image.”). Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Chatzistamatiou in view of Pinho to include wherein the learning model is a Vision Transformer-based model taught by Li’s reference. The motivation for doing so would have been to divide the document into multiple patches as suggested by Li (see Li, Paragraph [0052]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Li with Chatzistamatiou and Pinho to obtain the invention specified in claim 4. Claims 11 is rejected under 35 U.S.C. 103 as being unpatentable over Chatzistamatiou et al. (US 2023/0410543 A1) (hereinafter, “Chatzistamatiou”) in view of Pinho et al. (US 2023/0237272 A1) (hereinafter; “Pinho”) as applied to claim 9 above, and further in view of Yebes Torres et al. (US 2023/0005286 A1) (hereinafter “Torres”). Regarding claim 11, which claim 9 is incorporated, Chatzistamatiou fails to teach wherein the at least one processor is configured to determine the threshold value based on a size of the whole document. Pinho teaches wherein the at least one processor is configured to determine the threshold value (Paragraph [0081] “Conceptually, the function receives as input all table words 552 in a dataset, grouped by document, a horizontal tolerance 554 and a vertical tolerance 556, and a window size. Then, for each document, the example algorithm is configured to retrieve all word records and calls the previous function to process them in horizontal and vertical orders, from lines 15 to 20. It will be noted that, advantageously, the processing order is defined solely by the field_order parameter. With the horizontal and vertical target 552 and context word sets 558, the only remaining action is to concatenate and accumulate them into lists containing target and context words for all documents.”). Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Chatzistamatiou reference to include wherein the at least one processor is configured to determine the threshold value taught by Pinho’s reference. The motivation for doing so would have been to define the context window so as to capture the neighborhood of a table word in the horizontal and vertical directions as suggested by Pinho (see Pinho, Paragraph [0077]). However, Chatzistamatiou and Pinho both fail to teach based on a size of the whole document. Torres teaches based on a size of the whole document (Element 108 in Figure 9 equates to the whole document) (Paragraph [0133] “the bounding box generating circuitry 316 outputs a plurality of bounding boxes including respective coordinates corresponding to lines of the receipt region.”; Paragraph [0164] “example of FIG. 9…receives an example receipt image 108 (e.g., from the basket datastore 112)…The extraction circuitry 118 applies an example regions detection model 306 to the receipt image 108 to detect an example receipt region 902 and an example products region 904. In some examples, the extraction circuitry 118 applies a cropping operation (e.g., via the image cropping circuitry 308) to the receipt image 108 based on the detected regions). Figure 9: PNG media_image7.png 714 540 media_image7.png Greyscale Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Chatzistamatiou in view of Pinho to include based on a size of the whole document taught by Torres’s reference. The motivation for doing so would have been to align the respective coordinates to the document region as suggested by Torres (see Torres, Paragraph [0133]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Torres with Chatzistamatiou and Pinho to obtain the invention specified in claim 11. Allowable Subject Matter Claims 6 and 8 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claims 6 and 8 contain subject matter that is not disclosed or made obvious in the cited art: In regards to claim 6, when considering claim 6 as a whole, the features highlighted below are considered an improvement over the prior art and have not been found to be anticipated or rendered obvious by a combination of the prior art: “[…] sort the cell information on each of the plurality of cells based on the row order of the each of the plurality of cells, insert predetermined row change information into a portion of the cell information which has a row change in the row order, and input the cell information having the inserted predetermined row change information to the learning model in the row order.” In regards to claim 8, when considering claim 8 as a whole, the features highlighted below are considered an improvement over the prior art and have not been found to be anticipated or rendered obvious by a combination of the prior art: “[…] sort the cell information on each of the plurality of cells based on the column order of the each of the plurality of cells, insert predetermined column change information into a portion of the cell information which has a column change in the column order, and input the cell information having the inserted predetermined column change information to the learning model in the column order.” Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Zakharov (US 2022/0067320 A1) discloses a document classification technique that converts text content to images and apply image classification to the images. Zakharov further discloses a system comprising an image generator configured to generate graphical code based on topological features. The topological features comprise of tables where text is arranged in columns and rows, graphics such as charts, logos, etc., and blank spaces separating text and graphics. Inaki et al. (US 5,835,916 A) discloses an apparatus comprising a processing unit which allows an operator to designate a single cell in the table and to relocate the designated cell in an area of the table together with text data registered in the designated cell. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to UROOJ FATIMA whose telephone number is (571)272-2096. The examiner can normally be reached M-F 8:00-5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Henok Shiferaw can be reached at (571) 272-4637. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /UROOJ FATIMA/Examiner, Art Unit 2676 /Henok Shiferaw/Supervisory Patent Examiner, Art Unit 2676
Read full office action

Prosecution Timeline

Mar 28, 2024
Application Filed
Feb 23, 2026
Non-Final Rejection mailed — §103, §DOUBLEPATENT
Apr 30, 2026
Interview Requested
May 06, 2026
Applicant Interview (Telephonic)
May 07, 2026
Examiner Interview Summary
May 21, 2026
Response Filed
Sep 18, 2026
Final Rejection mailed — §103, §DOUBLEPATENT (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705860
COMPUTER-IMPLEMENTED OBJECT DETECTION METHOD, OBJECT DETECTION APPARATUS, AND COMPUTER-READABLE MEDIUM
2y 8m to grant Granted Aug 11, 2026
Patent 12693409
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND NON-TRANSITORY STORAGE MEDIUM
2y 8m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 2 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
75%
Grant Probability
75%
With Interview (+0.0%)
2y 8m (~2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 8 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month