Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed on 2 January, 2025.
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 25 April, 2025 and 22 July, 2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner.
Claim Rejections - 35 USC § 101
Claim 35 is rejected under 35 U.S.C. 101 because the claimed invention is directed to a “computer program product, comprising machine-executable instructions” that is non-statutory subject matter. The broadest reasonable interpretation of a claim drawn to a computer-readable medium (also called machine readable medium and other such variations) typically covers forms of non-transitory tangible media and transitory propagating signals per se in view of the ordinary and customary meaning of computer readable media, particularly when the specification is silent. See MPEP 2111.01. A review of the specification reveals that on page 22, lines 15 – 23 state “Such media may be any available media accessible to electronic device 700, including but not limited to volatile and non-volatile media…”
A claim drawn to such a computer-readable medium that covers both transitory and non-transitory embodiments may be amended to narrow the claim to cover only statutory embodiments to avoid a rejection under 35 U.S.C. § 101 by adding the limitation "non-transitory computer-readable medium comprising machine-executable instructions" to the claim.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of pre-AIA 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a) the invention was known or used by others in this country, or patented or described in a printed publication in this or a foreign country, before the invention thereof by the applicant for a patent.
(b) the invention was patented or described in a printed publication in this or a foreign country or in public use or on sale in this country, more than one year prior to the date of application for patent in the United States.
Claims 16 – 21, 24, 27 – 31, 34, and 35 are rejected under pre-AIA 35 U.S.C. 102(a)(2) as being anticipated by Price et al (U.S. Patent Publication No. 2020/00151444 A1, hereinafter “Price”).
Regarding claim 16, Price teaches a computer-implementation method comprising:
determining, based on a first feature map generated from an image including a table (¶ 0065: The convolutional neural network 222 receives the input features (the image 210 as well as any additional table features 212)…), a first set of reference points in the image, the first set of reference points being candidate points on separation lines of a first type of the table (¶ 0067: Specifically, the convolution and ReLU activation stage 404 applies six 7x7 kernels for each of the dilation factors of 2, 3, and 4, which together produces 18 feature maps. The various dilation factors are used so that the features can be examined at multiple scales. ReLU activation is also applied element-wise. The output of the convolution and ReLU activation stage 404 is referred to as X1, which is calculated as X1= ReLU(conv2 (X0)||conv3(X0)||conv4(X0)), where || refers to channel-wise concatenation and convN refers to a convolution layer with dilation factor N.; Examiner’s note: Examiner is interpreting the 18 feature maps which produce row features at different scales to be the “candidate points” on separation lines of a first type.);
determining, based on at least a part of the first feature map and features of the first set of reference points, a set of predicted separation lines of the first type for the table from the image (Figure 3; ¶ 0068: The row prediction machine learning system 204 generates a 1-dimensional (1D) projection across each row of pixels of the table that is a probability that the row of pixels is a row separator.); and
determine a structure of the table based at least on the set of predicted separation lines of the first type (Figure 3, 5, and 6; ¶ 0145: A layout of the table is identified using the probabilities of each row being a row separator and each column being a column separator (block 1008). These probabilities used in block 1008 are the probabilities determined in blocks 1004 and 1006. The identified layout is, for example, a set of row separator identifiers and column separator identifiers, although other formats can be used (e.g., coordinates of cells in the table).).
Regarding claim 17, Price teaches the method of claim 16.
Additionally, Price teaches wherein the first set of reference points are distributed in a direction perpendicular to a predetermined direction of the separation lines of the first type (Figure 10, Ref. No. 1006 and 1008; Such processing includes identifying a layout of a table on the digital medium 106, and outputting an indication of the table layout, such as a set of column separator identifiers 110 and a set of row separator identifiers 112.; Examiner’s note: As the system of Price determines both row and column separators, it is understood that when looking for rows (i.e. a predetermined horizontal direction) the distribution will be horizontal lines across a vertical axis. Similarly, when looking for columns (i.e., a predetermined vertical direction) the distribution will be across a horizontal axis.).
Regarding claim 18, Price teaches the method of claim 16.
Additionally, Price teaches wherein determining the set of predicted separation lines of the first type comprises:
extracting sampled features of the image from the first feature map (¶ 0067: Specifically, the convolution and ReLU activation stage 404 applies six 7x7 kernels for each of the dilation factors of 2, 3, and 4, which together produces 18 feature maps. The various dilation factors are used so that the features can be examined at multiple scales. ReLU activation is also applied element-wise. The output of the convolution and ReLU activation stage 404 is referred to as X1, which is calculated as X1= ReLU(conv2 (X0)||conv3(X0)||conv4(X0)), where || refers to channel-wise concatenation and convN refers to a convolution layer with dilation factor N.; ¶ 0068: The value X1 is provided to a max pooling stage 406.)
determining, based on the sampled features of the image and the features of the first set of reference points, predicted pixels in the image located on the separation lines of the first type (¶ 0068: The row prediction machine learning system 204 generates a 1-dimensional (1D) projection across each row of pixels of the table that is a probability that the row of pixels is a row separator.; ¶ 0071: This predicted signal y is output by the last of the five convolutional blocks of the row prediction machine learning system 204 as the probabilities of the rows being row separators.); and
determining the set of predicted separation lines of the first type based on positions of the predicted pixel in the image (Figure 3 – 5; ¶ 0071: The predicted signal y is a sequence of row predictions [y1, y2,. . . , yn], where yi ϵ [0,1]H.).
Regarding claim 19, Price teaches the method of claim 18.
Additionally, Price teaches wherein determining the predicted pixels comprises:
updating the features of the first set of reference points based on the sampled features of the image and the features of the first set of reference points (¶ 0068: The value X1 is provided to a max pooling stage 406… The output of the max pooling stage 406 is referred to as X2; ¶ 0069: The output of the max pooling stage 406 is input to the projection pooling and prediction stage 408. The projection pooling and prediction stage 408 computes row features, illustrated as the top branch 410 of the projection pooling and prediction stage 408.), wherein the updated features of a reference point reflect a correlation between the reference point and individual pixels in a sampling portion of the image (¶ 0068: The row prediction machine learning system 204 generates a 1-dimensional (1D) projection across each row of pixels of the table that is a probability that the row of pixels is a row separator.);
selecting reference points from the first set of reference points based on the updated features of the first set of reference points (¶ 0068: The row prediction machine learning system 204 generates a 1-dimensional (1D) projection across each row of pixels of the table that is a probability that the row of pixels is a row separator.); and
determining the predicted pixels based on the updated features of the selected reference points (Figure 3 – 5; ¶ 0071: The predicted 1D signal is y=f(horzproj(conv(X2))), where cony refers to a convolution layer, and horzproj refers to horizontally projecting ( e.g., averaging) over the row. This predicted signal y is output by the last of the five convolutional blocks of the row prediction machine learning system204 as the probabilities of the rows being row separators. The predicted signal y is a sequence of row predictions [y1, y2,. . . , yn], where yi ϵ [0,1]H.).
Regarding claim 20, Price teaches the method of claim 18.
Additionally, Price teaches wherein extracting the sampled features of the image comprises:
extracting features of a plurality of pixel blocks of the image from the first feature map (¶ 0067: The block input X0 is provided to a convolution and ReLU activation stage 404 where convolution operations with various dilation factors are applied to X0.), the plurality of pixel blocks spaced along a predetermined direction of the separation lines of the first type (¶ 0067: Specifically, the convolution and ReLU activation stage 404 applies six 7x7 kernels for each of the dilation factors of 2, 3, and 4, which together produces 18 feature maps.), each pixel block being sampled in a direction perpendicular to the predetermined direction (Examiner’s note: It is understood that a kernel samples an image by moving in both the X and Y directions across the image. For the instance of detecting row separation (i.e., predetermined horizontal direction), the kernel would move across the image vertically as each pixel of the image is sampled. This is understood to mean that the pixel blocks (kernel) would be sampled in a direction perpendicular to the horizontal direction.).
Regarding claim 21, Price teaches the method of claim 16.
Additionally, Price teaches wherein determining the first set of reference points comprises:
extracting features of a reference pixel block of the image from the first feature map (¶ 0067: The block input X0 is provided to a convolution and ReLU activation stage 404 where convolution operations with various dilation factors are applied to X0.), the reference pixel block being sampled in a direction perpendicular to a predetermined direction of the separation lines of the first type (¶ 0067: Specifically, the convolution and ReLU activation stage 404 applies six 7x7 kernels for each of the dilation factors of 2, 3, and 4, which together produces 18 feature maps.; Examiner’s note: It is understood that a kernel samples an image by moving in both the X and Y directions across the image. For the instance of detecting row separation (i.e., predetermined horizontal direction), the kernel would move across the image vertically as each pixel of the image is sampled. This is understood to mean that the pixel blocks (kernel) would be sampled in a direction perpendicular to the horizontal direction.); and
selecting a set of pixels from the reference pixel block based on the features of the reference pixel block as the first set of reference points (¶ 0067: The output of the convolution and ReLU activation stage 404 is referred to as X1, which is calculated as X1= ReLU(conv2 (X0)||conv3(X0)||conv4(X0)), where || refers to channel-wise concatenation and convN refers to a convolution layer with dilation factor N.).
Regarding claim 24, Price teaches the method of claim 16.
Additionally, Price teaches wherein determining the structure of the table comprises:
determining, based on a second feature map generated from the image, a second set of reference points in the image, the second set of reference points being candidate points on separation lines of a second type of the table, the separation lines of the second type being different from the separation lines of the first type (¶ 0075: The column prediction machine learning system 206 includes a convolutional neural network 226 that is implemented similarly to convolutional neural network 224 of the row prediction machine learning system 204, having similar layers and performing similar operations as the convolutional neural network 224.; ¶ 0076: The block input X0 is provided to the convolution and ReLU activation stage 404 where convolution operations with various dilation factors are applied to X0 as discussed above.);
determining a set of predicted separation lines of the second type of the table in the image based on at least a part of the second feature map features of the second set of reference points (¶ 0078: In the bottom branch 412 of the projection pooling and prediction stage 408, the 2D output map is projected to 1D by averaging over columns, and the predicted 1D signal is z=f (vertproj(conv(X2))) where cony refers to a convolution layer, and vetproj refers to vertically projecting (e.g., averaging) over the column.); and
determining the structure of the table based on the set of predicted separation lines of the first type and the set of predicted separation lines of the second type (Figure 3 – 5; ¶ 0078: The predicted signal z is a sequence of column predictions [z1, z2,. . . , zn], where zi ϵ [0,1]W.).
The rejection of method claim 16 above applies mutatis mutandis to the corresponding limitations of device claim 27 while noting that the rejection above cites to both method and device disclosures.
For the device limitations of claim 27 see Price’s teaching on:
An electronic device, comprising:
A processing unit (0151: Accordingly, the processing system 1104 is illustrated as including hardware element 1110 that may be configured as processors, functional blocks, and so forth.); and
A memory coupled to the processing unit and comprising instructions stored thereon, the instructions, when executed by the processing unit, causing the device to perform acts comprising (¶ 0156: The computer-readable storage media includes hardware such as volatile and non-volatile, removable and non-removable media and/or storage devices implemented in a method or technology suitable for storage of information such as computer readable instructions, data structures, program modules, logic elements/circuits, or other data.):
The rejection of method claim 17 above applies mutatis mutandis to the corresponding limitations of device claim 28 while noting that the rejection above cites to both method and device disclosures.
The rejection of method claim 18 above applies mutatis mutandis to the corresponding limitations of device claim 29 while noting that the rejection above cites to both method and device disclosures.
The rejection of method claim 19 above applies mutatis mutandis to the corresponding limitations of device claim 30 while noting that the rejection above cites to both method and device disclosures.
The rejection of method claim 21 above applies mutatis mutandis to the corresponding limitations of device claim 31 while noting that the rejection above cites to both method and device disclosures.
The rejection of method claim 24 above applies mutatis mutandis to the corresponding limitations of device claim 34 while noting that the rejection above cites to both method and device disclosures.
The rejection of method claim 16 above applies mutatis mutandis to the corresponding limitations of manufacture claim 35 while noting that the rejection above cites to both method and manufacture disclosures.
For the manufacture limitations of claim 35 see Price’s teaching on:
A computer program product, comprising machine-executable instructions which, when executed by a device, cause the device to perform acts comprising (¶ 0156: The computer-readable storage media includes hardware such as volatile and non-volatile, removable and non-removable media and/or storage devices implemented in a method or technology suitable for storage of information such as computer readable instructions, data structures, program modules, logic elements/circuits, or other data.):
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 22, 23, 32, and 33 are rejected under 35 U.S.C. 103 as being unpatentable over Price et al (U.S. Patent Publication No. 2020/00151444 A1, hereinafter “Price”) in view of Dong et al (U.S. Patent Publication No. 2021/0209297 A1, hereinafter “Dong”) .
Regarding claim 22, Price teaches the method of claim 16.
Price does not explicitly teach wherein determining the structure of the table comprises:
dividing at least a part of the image into a plurality of cells based at least on the set of predicted separation lines of the first type; generating a cell feature map for the plurality of cells, a feature from the cell feature map corresponding to one of the plurality of cells; and determining a layout of the cells in the table based on the cell feature map.
However, Dong does teach wherein determining the structure of the table comprises:
dividing at least a part of the image into a plurality of cells based at least on the set of predicted separation lines of the first type (Figure 2, 3; ¶ 0036: As shown in FIG. 2, multiple attributes may be extracted from a given cell among multiple cells comprised in the spreadsheet 170.; ¶ 0039: At block 310, respective multiple attributes of multiple cells comprised in the spreadsheet 170 may be extracted.);
generating a cell feature map for the plurality of cells, a feature from the cell feature map corresponding to one of the plurality of cells (¶ 0041: At block 320, respective features of the multiple cells may be determined based on the extracted multiple attributes.); and
determining a layout of the cells in the table based on the cell feature map (¶ 0042: At block 330, the multiple cells may be divided into at least one candidate region 182 and 184 based on the features.; ¶ 0043: At block 340, at least one candidate table in the spreadsheet 170 may be determined based on the at least one candidate region 182 and 184. Where it is determined that the one or more candidate regions 182 and 184 are comprised in the spreadsheet 170, one candidate table may be determined from each candidate region.).
Dong is considered to be analogous art as it pertains to table image feature extraction. Therefore, it would have been obvious to one of ordinary skill in the art to combine the table layout determination using machine learning system (as taught by Price) and the table detection in spreadsheet system (as taught by Dong) before the effective filing date of the claimed invention. The motivation for this combination of references would be the system of Dong extracts multiple features from cells based on characters of data, format of data, and style of the corresponding cell which improves the accuracy of table detection in the spreadsheet (See ¶ 0044).
This motivation for the combination of Price and Dong is supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Regarding claim 23, the Price and Dong combination teaches the method of claim 22.
Additionally, Dong teaches further comprising:
determining, based on the cell feature map, a type of content filled in cells in the plurality of cells (¶ 0042: Where the extracted attributes comprise a background color of the cell and an indication whether a character string within the cell is empty, different candidate regions comprising cells with different background colors in the spreadsheet 170 may be obtained through clustering various cells by "background color."; ¶ 0044: … multiple attributes of a corresponding cell may be extracted based on at least any of: characters of data in the corresponding cell, format of data in the corresponding cell, and style of the corresponding cell.).
The rejection of method claim 22 above applies mutatis mutandis to the corresponding limitations of device claim 32 while noting that the rejection above cites to both method and device disclosures.
The rejection of method claim 23 above applies mutatis mutandis to the corresponding limitations of device claim 33 while noting that the rejection above cites to both method and device disclosures.
Allowable Subject Matter
Claims 25 and 26 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding claim 25, the closest prior art, Price et al, in combination with secondary art Dong et al and Zhang et al )(Zhang, Zhenrong, et al. "Split, embed and merge: An accurate table structure recognizer." Pattern Recognition 126 (2022): 108565.), and additionally in combination with any other arts does not appear to teach updating the series of feature sub-maps by applying a feature transformation for extracting context information on the series of feature sub-maps in accordance with the predetermined direction and an opposite direction of the predetermined direction.
Regarding claim 26, claim 26 would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims as it depends from claim 25 and therefor contains all the limitations of claim 25 which are not found in the prior art.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Paliwal et al (U.S. Patent Publication No. 2022/0319217) teaches a system that uses a deep learning network for table detection and structure recognition by segmenting areas containing rows and columns of data using multiple masks. This prior art is understood to read entirely on claims 16 – 18, 20, and 27 – 29.
Zhang et al (Zhang, Zhenrong, et al. "Split, embed and merge: An accurate table structure recognizer." Pattern Recognition 126 (2022): 108565.) teaches a table detection method comprising splitting a feature map of a table into respective columns and rows, extracting feature representations of each grid, and creates a sequence of merged maps which provides the spanning of each table cell along the rows and columns.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANDREW JONES whose telephone number is (703)756-4573. The examiner can normally be reached Monday - Friday 8:00-5:00 EST, off Every Other Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matthew Bella can be reached at (571) 272-7778. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ANDREW B. JONES/Examiner, Art Unit 2667