DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
The United States Patent & Trademark Office appreciates the response filed for the current application that is submitted on 04/22/2026. The United States Patent & Trademark Office reviewed the following documents submitted and has made the following comments below.
Amendment
Applicant submitted amendments on 04/22/2026. The Examiner acknowledges the amendment and has reviewed the claims accordingly.
Overview
Claims 1-5, 7-15, and 17-20 are pending in this application and have been considered below.
Claims 6 and 16 are cancelled.
Claims 1-5, 7-15, and 17-20 are rejected.
Applicant Arguments
In regards to Argument 1, Applicant/s state/s “The Examiner has rejected claims 1, 2, 6, 11, 12, and 16 under 35 USC § 112(b) as being allegedly indefinite. Applicant has amended these claims in a manner believed fully responsive to the points the Examiner raised” therefore the 35 USC § 112(b) rejection should be withdrawn.
In regards to Argument 2, Applicant/s state/s “The claims clearly recite elements that recite in this improved generation of training data, including for example the limiting of the random translation of each of the first bounding boxes to be within its region, instead of anywhere on the form” therefore the 35 USC § 101 rejection should be withdrawn.
In regards to Argument 3, Applicant/s state/s “Applicant submits that no reasonable combination of the five applied references discloses, describes, or otherwise suggests the subject matter of independent claims 1 and 11” and therefore the 35 USC § 103 rejections should be withdrawn.
Examiner’s Responses
In response to Argument 1, see remarks, filed 04/22/2026, with respect to the rejection(s) of claim(s) 1, 2, 6, 11, 12, and 16 under 35 USC § 112(b) have been fully considered and are persuasive. Therefore, the Examiner has withdrawn the rejections for 35 USC § 112 in pending claims 1, 2, 11 and 12 in response to Applicants amendments. Claims 6 and 16 are cancelled.
In response to Argument 2, see remarks, filed 04/22/2026, with respect to the rejection(s) of claim(s) 1-20 under 35 USC § 101 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn from pending claims 1-5, 7-15, and 17-20. Claims 6 and 16 are cancelled.
In response to Argument 3, see remarks, filed 04/22/2026, the Examiner respectfully disagrees. The Applicant has incorporated claim limitation of cancelled claim 6, which was rejected in Office Action dated 01/27/2026, with prior art reference Ast in view of Streltsov. The Applicant has amended Claim 1 and changed the scope of Claim 1. Similarly, Applicant has incorporated claim limitation of cancelled claim 16, which was rejected in Office Action dated 01/27/2026, with prior art reference Ast in view of Streltsov. The Applicant has amended Claim 11 and changed the scope of Claim 11.
However, upon further review of Ast in view of Streltsov, the Examiner finds that the prior art references read on the new amended claims. Therefore, the Examiner has maintained the rejection.
Specifically, Streltsov teaches a system, method and computer-readable media for generating a synthetic training data set from an original unstructured electronic document. The original electronic documents contain bounding boxes, and may be separated into various sections, paragraphs [0036-38]. Various data augmentations may be applied to the original electronic documents to create structural and syntactic variance from the original electronic documents, paragraph [0036]. Both geometric (shift, clone, swap, delete, crop) and semantic augmentations (changing/swapping text strings) are applied to the original electronic document to generate training data, paragraphs [0041-48]. Augmentations to the documents may be performed by the user or performed automatically by the system, paragraph [0052]. This process is repeated until a sufficiently large set of synthetic electronic documents has been created, and the training dataset is provided to the deep learning model for training, paragraph [0036], Fig. 4. The Examiner finds that Streltsov does not teach using a deep learning system to perform the augmentations to the form to generate the training data set. However, Ast teaches a document processing system and method that leverages the power of deep learning to improve the performance of machine learning engines for data extraction and classification from document images, reducing the programming required to implement traditional algorithmic rules used in training ML engines, Col. 2, lines 18-32. Ast teaches a deep learning network used for extracting regions from the form and extracting data from the regions, Col. 2, lines 33-51. The deep learning network also updates a field type knowledge base by extracting regions from semantic images, Col. 3, lines 59-67. The Examiner finds it would be obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine the deep learning system in Ast to extract regions on the form and automate the application of geometric and semantic augmentations to the bounding boxes to generate training documents for the deep learning system in Streltsov having sufficient variance from the original unstructured electronic documents.
Claims 1-2, 4-5, 7-8, 11-12, 14-15, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Streltsov et al. (U.S. Patent Pub. No. 2023/0334309, hereafter referred to as Streltsov) in view of Ast (U.S. Patent No. 11,776,244, hereafter referred to as Ast).
Claims 3, 13 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Streltsov in view of Ast in further view of Gohari (U.S. Patent Pub. No. 2022/0318492, hereafter referred to as Gohari).
Claims 9-10 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Streltsov in view of Ast in further view of Buban et al. (U.S. Patent No. 12,094,231, hereafter referred to as Buban).
Claim Interpretation
Under MPEP 2143.03, “All words in a claim must be considered in judging the patentability of that claim against the prior art.” In re Wilson, 424 F.2d 1382 1385, 165 USPQ494, 496 (CCPA 1970). As a general matter, the grammar and ordinary meaning of terms as understood by one having ordinary skill in the art used in a claim will dictate whether, and to what extent, the language limits the claim scope. Language that suggests or makes a feature or step optional but does not require the feature or step does not limit the scope of the claim under the broadest reasonable claim interpretation. In addition, when a claim requires selection of an element from a list of alternatives, the prior art teaches the element if one of the alternatives is taught by the prior art. See, e.g., Fresenius USA, Inc. v. Baxter Int’l, Inc., 582 F .3d 1288, 1298, 92 USPQ2d 1163, 1171 (Fed. Cir. 2009).
Claims 1 and 11 recite “performing one of” then listing “randomly scaling or randomly translating.” Since “one of” is disjunctive, any one of the elements found in the prior art is sufficient to reject the claim. While citations have been provided for completeness and rapid prosecution, only one element is required. Because, on balance, it appears the disjunctive interpretation enjoys the most specification support and for that reason the disjunctive interpretation (one of A OR B) is being adopted for the purposes of this Office Action.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-2, 4-5, 7-8, 11-12, 14-15, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Streltsov et al. (U.S. Patent Pub. No. 2023/0334309, hereafter referred to as Streltsov) in view of Ast (U.S. Patent No. 11,776,244, hereafter referred to as Ast).
Regarding Claim 1, Streltsov teaches a method (Abstract, Streltsov discloses a method for generating a synthetic training data set from an original unstructured electronic document.) comprising: a) responsive to receipt of an input image of a form (Paragraphs [0038], [0056], Fig. 1A, Streltsov teaches obtaining an original electronic document (100) which may be an electronic document such as an invoice, payment advice, a paycheck, a purchase order, a receipt, or any other electronic document.) placing first bounding boxes around text in the form (Paragraphs [0039-40], Fig. 1A, reference character 110, Streltsov teaches the original electronic document (100) which may comprise a plurality of annotated data fields (108). Each section may comprise annotated data fields (108). Each annotated data field may comprise a bounding box (110) and an associated label. Each bounding box (110) may have coordinate data stored therefore (e.g., pixel positions of each corner of the bounding box (110). Bounding boxes may be automatically determined for original electronic document (100).);
PNG
media_image1.png
806
582
media_image1.png
Greyscale
one of randomly scaling or randomly translating each of the one or more of the first bounding boxes (Paragraphs [0061], [0063], Streltsov teaches performing “micro-operations” such as applying a shift operation to all labels in the original electronic document (100). The shift operations comprise a substantially small percent shift. For example, a label may be shifted by 2% downwards by modifying coordinates of the bounding box. Micro operations may comprise applying a shift operation to all annotated data fields, each of which comprise a label and a bounding box. Micro operations, such as shift operations may be randomly generated to reduce the likelihood that identical synthetic electronic documents are created.) within its own region (Paragraphs [0043], [0061], Streltsov teaches defining a distance limit for the shift operations such that the shifted annotated data fields may not be shifted outside of a specified region of the synthetic electronic document. The distance limit may vary for each annotated data field. The distance limit may be a pixel limit or a percentage limit based on the coordinates of the bounding box. The Examiner interprets a defined distance limit applied for shifts as a specified region in the document. The Examiner interprets “region” broadly as an area or location since the claim is silent to the specifications of the “region.”); e) identifying first entities in a region containing semantic information (Paragraph [0039], Fig. 1A, reference character 108, Streltsov teaches an original electronic document comprising a plurality of annotated data fields. Each section may comprise annotated data fields. Each annotated data field comprises a bounding box and an associated label. Each bounding box may have coordinate data in the form of pixel positions of the corners stored therefor. The Examiner interprets “semantic information” to include a text string including, for example, a date, a name of an industrial part, a telephone number, a street address, or the like. The Examiner interprets “region” broadly as an area or location since the claim is silent to the specifications of the “region.”);
PNG
media_image2.png
787
627
media_image2.png
Greyscale
f) replacing the identified first entities with second entities (Paragraph [0004], Streltsov teaches “semantic augmentations,” which may comprise changing a text string in the data field. For example, an address field in the original electronic document may be changed to a random address in the synthetic electronic document. The Examiner interprets “changing” and “replacing” to be synonymous in this context. Additionally, the second entities show new/different text that has been replaced, as shown below in Fig. 1B.),
PNG
media_image3.png
793
580
media_image3.png
Greyscale
wherein the replacing comprises: 1) randomly selecting a second entity (Paragraph [0052], Streltsov teaches retrieving a random address from the dictionary. The Examiner interprets a random address to be a “second entity” since it is different from the address (first entity) to be replaced in the original electronic document (Fig. 1A).); 2) replacing one of the first entities with the second entity (Paragraph [0052], Streltsov teaches within an address field, replacing the address (first entity) with a random address (second entity) from the dictionary.); 3) adding the second entity to a dictionary (Paragraph [0048], Streltsov teaches a dictionary for storing randomly generated addresses. The Examiner interprets that the second entities are automatically added to the dictionary when they are stored in it.); 4) and repeating said randomly selecting, said replacing one of the first entities with the second entity, and said adding for all of the first entities (Paragraphs [0062], [0004], Figs. 1A and 1B, Streltsov teaches semantically augmenting each annotated data field in an original synthetic document (100). The semantic augmentations may comprise changing a string in the data field, requiring each of the steps mentioned above, including selecting the first entity, replacing the first entity with the second entity, and adding the second entity to a dictionary. The Figures (1A and 1B) shown side-by-side below depict all of the first entities (original text) have been replaced with second entities (new/different text). See dashed arrows below for examples.);
PNG
media_image4.png
785
1123
media_image4.png
Greyscale
g) placing second bounding boxes around the text in the second entities (Fig. 1B, reference character 108, Paragraph [0062], Streltsov teaches modifying bounding boxes to account for the new text. For example, if semantic augmentation adds two new lines of text, the size of the bounding box may be increased accordingly. The Examiner interprets “placing second bounding boxes around the text in second entities” to include updating the position and/or size of the bounding box to account for the newly inserted text in the second entity, because the claim is silent as to how the second bounding boxes differ. See Fig. 1B below, the sizes of the bounding boxes have been adjusted to account for the newly replaced text.);
PNG
media_image5.png
785
1123
media_image5.png
Greyscale
forming text images to generate new training data for the deep learning system
(Abstract, Paragraphs [0001], [0028], [0036], Fig. 1B, Streltsov teaches generating synthetic electronic documents forming a training data set to train a deep learning model. Fig. 1B (below) depicts an example synthetic electronic document generated from the example original electronic document. The Examiner interprets the synthetic electronic document to be a “text image” since the document contains text and the claim is silent to the specifications of “text image.” Additionally, the Examiner interprets the synthetic electronic documents to be “new” training data since they are generated by performing geometric and/or semantic data augmentations to the original electronic document to add structural and/or syntactic variance from the original electronic documents.);
PNG
media_image6.png
811
582
media_image6.png
Greyscale
and i) training the deep learning system using the new training data (Abstract, Paragraphs [0069], [0005], Fig. 4, Fig. 1B, reference character 412, Streltsov teaches providing the training data set comprising a plurality of synthetic electronic documents, such as the one shown in Fig. 1B, to the learning model for training thereon.).
PNG
media_image7.png
591
554
media_image7.png
Greyscale
Streltsov does not explicitly disclose b) inputting semantic information for text in the first bounding boxes and c) using a deep learning system, identifying regions on the form, the regions to contain one or more of the first bounding boxes.
Ast is in the same field of art of using deep learning to improve the performance of machine learning engines for performing classification and data extraction from document images. Further, Ast teaches b) inputting semantic information for text in the first bounding boxes (Col. 9, lines 5-15, Fig. 8C, Ast teaches a semantic image with semantic information inputted for textual information, positioned in accordance with the bounding boxes associated with the text strings.); and c) using a deep learning system, identifying regions on the form, the regions to contain one or more of the first bounding boxes (Col. 3, lines 45-51, Col. 7, lines 17-37, Fig. 8D, Fig. 1, Ast teaches a region-based convolutional neural network (R-CNN) that extracts regions from the semantic image, each of which contains one or more bounding boxes. The Examiner interprets the extracted/identified regions contain one or more of the first bounding boxes since the bounding boxes depicted the semantic image (Fig. 1, 125) are the same size and in the same position/location as the bounding boxes in the output image (OCR/bounding boxes defined) (Fig. 1, 115). Additionally, the semantic information in 125 was coded using the text information in 115 and the position information (i.e., geometric coordinates of the bounding boxes) remains the same.).
PNG
media_image8.png
552
815
media_image8.png
Greyscale
Therefore it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Streltsov by inputting semantic information for text in the bounding boxes in the original document and using a R-CNN to extract regions from the semantic image containing first bounding boxes that is taught by Ast to make the invention that uses a deep learning system to automate the generation of training data with diverse layouts and text; thus, one of ordinary skilled in the art would be motivated to combine the references to automate the rule-based or manual process of altering forms by employing a deep learning system to carry out the semantic and geometric augmentations to the annotated data fields and bounding boxes, thereby reducing or eliminating the programming previously required to implement algorithmic rules used in training ML engines (Ast, Col. 2, lines 20-32).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
In regards to Claim 2, Streltsov in view of Ast teaches The method of claim 1, wherein the input image of the form comprises at least one table (Paragraphs [0037-38], Fig. 1A, reference character 104, Streltsov teaches the electronic document may comprise a tabular layout, such as often seen on an electronic invoice with line-item information.), the method further comprising randomly moving one or more columns and/or one or more rows within the table (Paragraph [0052], Streltsov teaches cloning a random number of line items in a table section (104) to create a plurality of synthetic electronic documents (150) with various sized tables for training the model. Semantic and geometric augmentations such as shifts to the cloned line items may also be performed. The Examiner interprets that by cloning a random number of line items (rows) in a table, as well as applying shifts to the line items, this moves the rows within the table. The claim is silent as to how the rows are moved.).
In regards to Claim 4, Streltsov in view of Ast teaches the method of claim 1, further comprising performing text spotting in the input image of the form (Paragraph [0056], Streltsov teaches automatically generating annotated data fields (108).), and optical character recognition (OCR) on the input image of the form (Paragraphs [0038], [0056], Fig. 3, reference character 306, Streltsov teaches an optical character recognition (OCR) file (306) generated by an OCR model for the original electronic document.).
In regards to Claim 5, Streltsov in view of Ast teaches the method of claim 1, wherein there are first bounding boxes for all of the first entities (Paragraph [0039], Streltsov teaches each annotated data field (108) in the original electronic document comprises a bounding box.).
In regards to Claim 7, Streltsov in view of Ast teaches the method of claim 1, wherein the randomly translating the one or more of the first bounding boxes within one or more of the regions (Paragraphs [0061], [0052], Streltsov teaches applying a “micro-operation,” to the original electronic document (100) such as applying a shift operation to all annotated data fields (108) in the original electronic document. The amount of shift applied to the field may be randomly applied.) comprises, for a bounding box in a region, translating the bounding box only within said region (Paragraph [0061], Streltsov teaches performing a substantially small percent shift, such as downwards by 2% by modifying the coordinates of the bounding box (110). The randomness of the geometric augmentations may also be controlled by the user by setting upper and lower limits on the shift distance. “Region” is being interpreted broadly as any different area or spot on the form since the claim does not specify the constraints of the “region” of the form.).
In regards to Claim 8, Streltsov in view of Ast teaches the method of claim 1, wherein the randomly translating the one or more of the first bounding boxes within one or more of the regions comprises, for a bounding box in a region, translating the bounding box from said region to another of the one or more of the regions (Paragraphs [0010], [0041], Streltsov teaches identifying the header (102) and footer section (106) in the original electronic document (100), and responsive to identifying, deleting the header section and the footer section, and shifting the table section in an arbitrary direction in the sub-template. The position of sections (102, 104, 106) may change due to segmenting annotated data fields (108). The Examiner interprets “shifting the table section in an arbitrary direction” to be a random translation. Additionally, the Examiner interprets shifting the table section, which contains bounding boxes to an arbitrary region since the claim is silent to how the regions differ. “Region” is being interpreted broadly as any different area or spot on the form since the claim is silent to the definition of “region.”).
In regards to Claim 11, Streltsov teaches an apparatus (Paragraph [0004], Streltsov teaches a system for generating synthetic electronic documents from an original electronic document for training learning models.) comprising: a deep learning system (Paragraph [0005], Streltsov teaches a deep learning model.) comprising at least one processor (Paragraph [0019], Streltsov teaches a system including at least one processor.) and a non-transitory memory that contains instructions that, when executed, enable the deep learning system to perform a method (Paragraph [0019], Streltsov teaches one or more non-transitory computer-readable media storing computer-executable instructions that, when executed by the at least one processor perform a method for generating the synthetic training data set for training the deep learning model.) comprising a) responsive to receipt of an input image of a form (Paragraphs [0038], [0056], Fig. 1A, Streltsov teaches obtaining an original electronic document (100) which may be an electronic document such as an invoice, payment advice, a paycheck, a purchase order, a receipt, or any other electronic document.) placing first bounding boxes around text in the form (Paragraphs [0039-40], Fig. 1A, reference character 110, Streltsov teaches the original electronic document (100) which may comprise a plurality of annotated data fields (108). Each section may comprise annotated data fields (108). Each annotated data field may comprise a bounding box (110) and an associated label. Each bounding box (110) may have coordinate data stored therefore (e.g., pixel positions of each corner of the bounding box (110). Bounding boxes may be automatically determined for original electronic document (100).); b) one of randomly scaling or randomly translating each of the one or more of the first bounding boxes (Paragraphs [0061], [0063], Streltsov teaches performing micro-operations such as applying a shift operation to all labels in the original electronic document (100). The shift operations comprise a substantially small percent shift. For example, a label may be shifted by 2% downwards by modifying coordinates of the bounding box. Micro operations may comprise applying a shift operation to all annotated data fields. Micro operations may be randomly generated to reduce the likelihood that identical synthetic electronic documents are created.) within its own region (Paragraphs [0043], [0061], Streltsov teaches defining a distance limit for the shift operations such that the shifted annotated data fields may not be shifted outside of a specified region of the synthetic electronic document. The distance limit may vary for each annotated data field. The distance limit may be a pixel limit or a percentage limit based on the coordinates of the bounding box.); e) identifying first entities in a region containing semantic information (Paragraph [0039], Fig. 1A, reference character 108, Streltsov teaches an original electronic document comprising a plurality of annotated data fields. Each section may comprise annotated data fields. Each annotated data field comprises a bounding box and an associated label. Each bounding box may have coordinate data in the form of pixel positions of the corners stored therefor. The Examiner interprets “semantic information” to include a text string including, for example, a date, a name of an industrial part, a telephone number, a street address, or the like. The Examiner interprets “region” broadly as an area or location since the claim is silent to the specifications of the “region.”); f) replacing the identified first entities with second entities (Paragraph [0004], Streltsov teaches “semantic augmentations,” which may comprise changing a text string in the data field. For example, an address field in the original electronic document may be changes to a random address in the synthetic electronic document. The Examiner interprets “changing” and “replacing” to be synonymous in this context.), wherein the replacing comprises: 1) randomly selecting a second entity (Paragraph [0052], Streltsov teaches retrieving a random address from the dictionary. The Examiner interprets a random address to be a “second entity” since it is different from the address to be replaced.); 2) replacing one of the first entities with the second entity (Paragraph [0052], Streltsov teaches within an address field, replacing the address with a random address from the dictionary.); 3) adding the second entity to a dictionary (Paragraph [0048], Streltsov teaches a dictionary for storing randomly generated addresses. The Examiner interprets that the second entities are automatically added to the dictionary when they are stored in it.); and 4) repeating said randomly selecting, said replacing one of the first entities with the second entity, and said adding for all of the first entities (Paragraphs [0062], [0004], Streltsov teaches semantically augmenting each annotated data field in a document. The semantic augmentations may comprise changing a string in the data field, requiring each of the steps mentioned above, including selecting, replacing, and adding.); g) placing second bounding boxes around the text in the second entities (Fig. 1B, reference character 108, Paragraph [0062], Streltsov teaches modifying bounding boxes to account for the new text. For example, if semantic augmentation adds two new lines of text, the size of the bounding box may be increased accordingly. The Examiner interprets “placing second bounding boxes around the text in second entities” to include updating the position and/or size of the bounding box to account for the newly inserted text in the second entity, because the claim is silent as to how the second bounding boxes differ.); forming text images to generate new training data for the deep learning system (Abstract, Paragraphs [0001], [0028], [0036], Fig. 1B, Streltsov teaches generating synthetic electronic documents forming a training data set to train a deep learning model. Fig. 1B depicts an example synthetic electronic document generated from the example original electronic document. The Examiner interprets the synthetic electronic document to be a “text image” since the document contains text and the claim is silent to the specifications of “text image.” Additionally, the Examiner interprets the synthetic electronic documents to be “new” training data since they are generated by performing geometric and semantic data augmentations to the original electronic documents to add structural and/or syntactic variance from the original electronic documents.); and i) training the deep learning system using the new training data (Abstract, Paragraphs [0069], [0005], Fig. 4, reference character 412, Streltsov teaches providing the training data set comprising a plurality of synthetic electronic documents to the learning model for training thereon.).
Streltsov does not explicitly disclose b) inputting semantic information for text in the first bounding boxes; and c) using a deep learning system, identifying regions on the form, the regions to contain one or more of the first bounding boxes.
Ast is in the same field of art of generating training data from input images of documents and utilizing the training documents to improve the training and performance of a machine learning engine for classification and data extraction. Further, Ast teaches b) inputting semantic information for text in the first bounding boxes (Col. 9, lines 5-15, Fig. 8C, Ast teaches a semantic image with semantic information inputted for textual information, positioned in accordance with the bounding boxes associated with the text strings.); and c) using a deep learning system, identifying regions on the form, the regions to contain one or more of the first bounding boxes (Col. 3, lines 45-51, Col. 7, lines 17-37, Fig. 8D, Fig. 1, Ast teaches a region-based convolutional neural network (R-CNN) that extracts regions from the semantic image, each of which contains one or more bounding boxes. The Examiner interprets the extracted regions contain one or more of the first bounding boxes since the bounding boxes depicted the semantic image (Fig. 1, 125) are the same size and in the same position/location as the bounding boxes in the output image (OCR/bounding boxes defined) (Fig. 1, 115). The difference between 115 and 125 is simply that the text has been coded with semantic information, however the position information (i.e., geometric coordinates) remain the same.).
Therefore it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Streltsov by inputting semantic information for text in the bounding boxes in the original document and using a R-CNN to extract regions from the semantic image containing first bounding boxes that is taught by Ast to make the invention that uses a deep learning system to automate the generation of training data with diverse layouts and text; thus, one of ordinary skilled in the art would be motivated to combine the references to automate the rule-based or manual process of altering forms by employing a deep learning system to carry out the semantic and geometric augmentations to the annotated data fields and bounding boxes, thereby reducing or eliminating the programming previously required to implement algorithmic rules used in training ML engines (Ast, Col. 2, lines 20-32).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
In regards to Claim 12, Streltsov in view of Ast teaches the apparatus of claim 11, wherein the input image of the form comprises at least one table (Paragraphs [0037-38], Fig. 1A, reference character 104, Streltsov teaches the electronic document may comprise a tabular layout, such as often seen on an electronic invoice with line-item information.), the method further comprising randomly moving one or more columns and/or one or more rows within the table (Paragraph [0052], Streltsov teaches cloning a random number of line items in a table section (104) to create a plurality of synthetic electronic documents (150) with various sized tables for training the model. Semantic and geometric augmentations such as shifts to the cloned line items may also be performed. The Examiner interprets that by cloning a random number of line items (rows) in a table, as well as applying shifts to the line items, this moves the rows within the table.).
In regards to Claim 14, Streltsov in view of Ast teaches the apparatus of claim 11, further comprising performing text spotting in the input image of the form (Paragraph [0056], Abstract, Streltsov teaches automatically generating annotated data fields (108) for the original electronic document (100).) Each annotated data field may comprise a label.), and optical character recognition (OCR) on the input image of the form (Paragraphs [0038], [0056], Fig. 3, reference character 306, Streltsov teaches an optical character recognition (OCR) file (306) generated by an OCR model.).
In regards to Claim 15, Streltsov in view of Ast teaches the apparatus of claim 11, wherein there are first bounding boxes for all of the first entities (Paragraph [0039], Streltsov teaches each annotated data field (108) in the original electronic document comprises a bounding box.).
In regards to Claim 17, Streltsov in view of Ast teaches the apparatus of claim 15, wherein the randomly translating the one or more of the first bounding boxes within one or more of the regions (Paragraphs [0061], [0052], Streltsov teaches applying a “micro-operation,” to the original electronic document (100) such as applying a shift operation to all annotated data fields (108) in the original electronic document. The amount of shift applied to the field may be randomly applied.) comprises, for a bounding box in a region, translating the bounding box only within said region (Paragraph [0061], Streltsov teaches performing a substantially small percent shift, such as downwards by 2% by modifying the coordinates of the bounding box (110). The randomness of the geometric augmentations may also be controlled by the user by setting upper and lower limits on the shift distance. “Region” is being interpreted broadly as any different area or spot on the form since the claim does not specify the constraints of the region on the form.).
Claims 3, 13 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Streltsov et al. (U.S. Patent Pub. No. 2023/0334309, hereafter referred to as Streltsov) in view of Ast (U.S. Patent No. 11,776,244, hereafter referred to as Ast) in further view of Gohari (U.S. Patent Pub. No. 2022/0318492, hereafter referred to as Gohari).
Regarding Claim 3, Streltsov in view of Ast disclose the method of Claim 1.
Streltsov in view of Ast does not explicitly disclose updating weights of nodes in the deep learning system being trained, responsive to one or more of said randomly scaling, said randomly translating, said replacing, or said forming.
Gohari is in the same field of art of form generation and data extraction using a deep learning system. Further, Gohari teaches updating weights of nodes in the deep learning system being trained (Paragraphs [0016-20], [0033], Fig. 7, Gohari teaches updating weights of nodes in the deep learning model responsive to updating one or more of the identifying text differences and the identifying graphics differences.),
PNG
media_image9.png
543
568
media_image9.png
Greyscale
responsive to one or more of said randomly scaling, said randomly translating, said replacing, or said forming (Paragraphs [0033], [0036], [0059], Figs. 1A and 1B, Fig. 4, Gohari teaches inputting sets of synthetically generated blank forms to the neural network, and comparing the forms to identify differences between them. The synthetically generated documents may be similar, but may have minor changes from form to form, to enable weights of nodes to be altered properly. For example, the header in the form containing the word “Part may be replaced with the word “Widget.” If there are differences in text between two forms, those differences are identified and categorized. Weights for nodes in the neural network are updated and flow is returned to the next pair of forms for training.).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Streltsov in view of Ast by updating weights of nodes based on the differences between sets of synthetically generated forms/documents that is taught by Gohari, to make the invention that effectively trains the deep learning system to recognize relevant differences between otherwise similar forms, including the types and locations of keywords and potential locations of values corresponding to keywords; thus, one of ordinary skilled in the art would be motivated to combine the references since automating the manual template creation process would save both time and labor associated with manual form creation for training the deep learning model to extract data (Gohari, Abstract, Paragraph [0004]).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
In regards to Claim 13, Streltsov in view of Ast disclose the apparatus of claim 11.
Streltsov in view of Ast does not explicitly disclose updating weights of nodes in the deep learning system being trained, responsive to one or more of said randomly scaling, said randomly translating, said replacing, or said forming.
Gohari is in the same field of art of form generation and data extraction using a deep learning system. Further, Gohari teaches updating weights of nodes in the deep learning system being trained (Paragraphs [0016-20], [0033], Fig. 7, Gohari teaches updating weights of nodes in the deep learning model responsive to updating one or more of the identifying text differences and the identifying graphics differences.), responsive to one or more of said randomly scaling, said randomly translating, said replacing, or said forming (Paragraphs [0033], [0036], [0059], Figs. 1A and 1B, Fig. 4, Gohari teaches inputting sets of synthetically generated blank forms to the neural network, and comparing the forms to identify differences between them. The synthetically generated documents may be similar, but may have minor changes from form to form, to enable weights of nodes to be altered properly. For example, the header in the form containing the word “Part may be replaced with the word “Widget.” If there are differences in text between two forms, those differences are identified and categorized. Weights for nodes in the neural network are updated and flow is returned to the next pair of forms for training.).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Streltsov in view of Ast by updating weights of nodes based on the differences between sets of synthetically generated forms/documents that is taught by Gohari, to make the invention that effectively trains the deep learning system to recognize relevant differences between otherwise similar forms, including the types and locations of keywords and potential locations of values corresponding to keywords; thus, one of ordinary skilled in the art would be motivated to combine the references since automating the manual template creation process would save both time and labor associated with manual form creation for training the deep learning model to extract data (Gohari, Abstract, Paragraph [0004]).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
In regards to Claim 18, Streltsov in view of Ast in further view of Gohari disclose the apparatus of claim 13, wherein the randomly translating the one or more of the bounding boxes within one or more of the regions comprises, for a bounding box in a region, translating the bounding box from said region to another of the one or more of the regions (Paragraphs [0010], [0041], Streltsov teaches identifying the header (102) and footer section (106) in the original electronic document (100), responsive to identifying, deleting the header section and the footer section, and shifting the table section in an arbitrary direction in the sub-template. The position of sections (102, 104, 106) may change due to segmenting annotated data fields (108). The Examiner interprets “shifting the table section in an arbitrary direction” to be a random translation. Additionally, the Examiner interprets shifting the table section, which contains bounding boxes to an arbitrary region since the claim is silent to how the regions differ. “Region” is being interpreted broadly as any different area or spot on the form.).
Claims 9, 10, 19, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Streltsov et al. (U.S. Patent Pub. No. 2023/0334309, hereafter referred to as Streltsov) in view of Ast (U.S. Patent No. 11,776,244, hereafter referred to as Ast) in further view of Buban et al. (U.S. Patent No. 12,094,231, hereafter referred to as Buban).
Regarding Claim 9, Ast in view of Streltsov discloses the method of claim 1.
Ast in view of Streltsov does not explicitly disclose wherein the randomly scaling comprises, for a bounding box in a region, enlarging or shrinking the bounding box.
Buban is in the same field of art of form generation for training a machine learning model. Further, Buban discloses wherein the randomly scaling comprises, for a bounding box in a region, enlarging or shrinking the bounding box (Col. 5, lines 28-43, Buban teaches augmenting documents by scaling them. The bounding boxes on the document are transformed according to the augmentation applied to the document. The Examiner interprets the term “scaling” by its well-known definition in the art, in which scaling is defined as a linear transformation that either enlarges or shrinks an object while maintaining its shape.).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Streltsov in view of Ast by, randomly scaling the bounding boxes in the augmented document by enlarging or shrinking the bounding box that is taught by Buban, to make the invention that augments a smaller set of human-labeled documents to be used to generate a larger set of documents that introduces variance to the training data set (Buban, Col. 5, lines 28-43); thus, one of ordinary skill in the art would be motivated to combine the references to improve the accuracy of identifying and classifying form fields by training the deep learning system on a set of forms with sufficient variance (Buban, Col. 5, lines 65-67 and Col. 12, lines 4-7).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
In regards to Claim 10, Ast in view of Streltsov in further view of Buban discloses the method of claim 9, further comprising enlarging or shrinking all of the bounding boxes in the region (Col. 5, lines 37-40, Buban teaches transforming the labeled bounding boxes according to the augmentation applied to the entire document. For example, if a document is scaled, the bounding box coordinates of the augmented version of the document are all adjusted accordingly.).
In regards to Claim 19, Ast in view of Streltsov discloses the apparatus of claim 11.
Ast in view of Streltsov does not explicitly disclose wherein the randomly scaling comprises, for a bounding box in a region, enlarging or shrinking the bounding box.
Buban is in the same field of art of form generation for training a machine learning model. Further, Buban discloses wherein the randomly scaling comprises, for a bounding box in a region, enlarging or shrinking the bounding box (Col. 5, lines 28-43, Buban teaches augmenting documents by scaling them. The bounding boxes on the document are transformed according to the augmentation applied to the document. The Examiner interprets the term “scaling” by its well-known definition in the art, in which scaling is defined as a linear transformation that either enlarges or shrinks an object while maintaining its shape.).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Streltsov in view of Ast by, randomly scaling the bounding boxes in the augmented document by enlarging or shrinking the bounding box that is taught by Buban, to make the invention that augments a smaller set of human-labeled documents to be used to generate a larger set of documents that introduces variance to the training data set (Buban, Col. 5, lines 28-43); thus, one of ordinary skill in the art would be motivated to combine the references to improve the accuracy of identifying and classifying form fields by training the deep learning system on a set of forms with sufficient variance (Buban, Col. 5, lines 65-67 and Col. 12, lines 4-7).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
In regards to Claim 20, Ast in view of Streltsov in further view of Buban discloses the apparatus of claim 19, wherein the method further comprises enlarging or shrinking all of the bounding boxes in the region (Col. 5, lines 37-40, Buban teaches transforming the labeled bounding boxes according to the augmentation applied to the entire document. For example, if a document is scaled, the bounding box coordinates of the augmented version of the document are all adjusted accordingly.).
Pertinent Prior Art
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Zeng et al. (U.S. Patent Pub. No. 2023/0351115 A1)
Xylouris (U.S. Patent No. 12,183,106 B1)
Ramezani (U.S. Patent No. 10,956,673 B1)
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SYDNEY L BLACKSTEN whose telephone number is 571-272-7651. The examiner can normally be reached 8:30am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Oneal Mistry can be reached at 313-446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SYDNEY L BLACKSTEN/Examiner, Art Unit 2674
/ONEAL R MISTRY/Supervisory Patent Examiner, Art Unit 2674