Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. JP2024-023409, filed on 02/20/2024.
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 02/12/2025 is being considered by the examiner.
Claim Objections
Claim 5 objected to because of the following informalities: wherein for the leaner the machine learning should be wherein for the learner for the machine learning. Appropriate correction is required.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: a target image acquiring unit, object detecting unit, process executing unit, in claim 1 and area extracting unit in claims 1-4.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Serry (US 20240177514 A1) and further in view of Agarwal (US 20250391189 A1).
Regarding claim 1, Serry discloses a target image acquiring unit configured to acquire as a target image a document image of a document ([0027] The techniques described herein implement a system that learns the structure of a form. The structure of the form can be learned from a single image (e.g., a scanned document that captures the form, a photograph that includes the form) without user annotation. The form includes typewritten text entries and handwritten text entries.);
an object detecting unit configured to detect area specifying objects additionally written to the document by handwriting in the document image using object detection with a learner for which ([0030] FIG. 1 illustrates an example environment 100 in which a system 102 learns the structure of a form 104 that includes typewritten text entries 106 and handwritten text entries 108. In various examples, the form 104 is included in an image 110 (e.g., a scanned document, a photograph).),
and derive respective confidences of the detected area specifying objects ("[0050] Described below is a specific example of a pairing algorithm 136 that identifies the optimal pairing solution 138 from the possible pairing solutions based on defined metric functions that utilize the distance properties and the angle properties of the edges 132. ");
an area extracting unit configured to determine a pair of area specifying objects that specify a rectangular area among the detected area specifying objects (fig. 5: two boundary pairs with distance, vector and angle measurements between each other to determine if it is a pair)
PNG
media_image1.png
382
544
media_image1.png
Greyscale
and extract the rectangular area on the basis of positions of the area specifying objects in the determined pair ([0005] As described herein, a location can be represented by a bounding box that contains the typewritten text or the handwritten text. The line detection module detects lines that separate the text entries in the form. The detected lines serve as constraints when determining the structure of the form for the purposes of accurately extracting pairings between the typewritten text entries and the handwritten text entries.); and
a process executing unit configured to execute a predetermined process for the rectangular area extracted in the target image ("[0033] FIG. 2B illustrates the example form 200 of FIG. 2A, in which the locations of the typewritten text entries 202, 204, 206, 208, 210, 212 and the locations of the handwritten text entries 214, 216, 218, 220, 222, 224 are illustrated as bounding boxes, represented by the dashed lines that surround the text (e.g., an example bounding box 226 is called out for the “PRODUCTION” typewritten text entry 202).
[0034] Turning back to FIG. 1, the line detection module 116 detects lines 122, in the form 104, that separate text entries. The detected lines serve as constraints when determining the structure of the form 104 for the purposes of accurately extracting pairings between the typewritten text entries 106 and the handwritten text entries 108. Moreover, the outer edges of the image 110 and/or the form 104 also serve as constraints. Switching the attention back to FIG. 2B, the line detection module 116 detects a first line 228 that separates text entries 202, 214 from text entries 204, 206, 216, 218, 208, 220. Moreover, the line detection module 116 detects a second line 230 that separates text entries 204, 206, 216, 218, 208, 220 from text entries 210, 212, 222, 224.").
Serry implicitly teaches machine learning ([0030] FIG. 1 illustrates an example environment 100 in which a system 102 learns the structure of a form 104 that includes typewritten text entries 106 and handwritten text entries 108. In various examples, the form 104 is included in an image 110 (e.g., a scanned document, a photograph). The system 102 is tasked with learning the structure of the form 104 in scenarios where the form 104 lacks a sufficient amount of lines 112 that clearly define pairings for the typewritten text entries 106 and the handwritten text entries 108.).
In a similar field of endeavor of object detection in a pdf document, Agarwal teaches machine learning ([0020] A neural network may include a machine-learning model that can be tuned (e.g., trained) based on training input to approximate unknown functions.
[0044] In one or more embodiments, the repeating objects detection module 112 first identifies the closest pair of objects of the listing of objects predicted by the page segmentation model 108. For example, repeating objects detection module 112 may identify object 412 and object 414 as being the closest pair of objects based on the distance between the two objects and a distance threshold value.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention, to combine the known system of pairing solution in handwritten form recognition, as taught by Serry, with the known teaching of machine learning, as taught by Agarwal, in order to yield the predictable results of automating data extraction by matching individual pen strokes or text blocks to predefined form fields with high precision.
Regarding claim 2, Serry discloses wherein the area specifying objects includes a first area specifying object that specifies a first angle and a second area specifying object that specifies a second angle among four angles of the rectangular area, the second angle being an opposite angle to the first angle ([0045] An element 404 (e.g., a handwritten text entry) that is directly to the right of the base element 402 has an angle of zero degrees (or alternatively three hundred and sixty degrees). An element 406 (e.g., a handwritten text entry) that is directly above the base element 402 has an angle of ninety degrees. An element 408 (e.g., a handwritten text entry) that is directly to the left of the base element 404 has an angle of one hundred and eighty degrees. And an element 410 (e.g., a handwritten text entry) that is directly below the base element 404 has an angle of two hundred and seventy degrees. Turning back to FIG. 2B, an example of an angle measurement 234 between the “DIRECTOR” typewritten text entry 210 and the “Jane D.” handwritten text entry 222 is two hundred and seventy degrees.);
the learner distinctively detects the first area specifying object and the second area specifying object ([0045] Turning back to FIG. 2B, an example of an angle measurement 234 between the “DIRECTOR” typewritten text entry 210 and the “Jane D.” handwritten text entry 222 is two hundred and seventy degrees.); and
the area extracting unit determines a pair of the first area specifying object and the second area specifying object, and extracts the rectangular area on the basis of positions of the first area specifying object and the second area specifying object in the determined pair ("[0046] Considering the distance property and the angle property, each edge 132 in the bipartite graph 126 can be treated as a vector 502 from a location 504 of a typewritten text entry (e.g., “DATE”) to a 506 location of a handwritten text entry (e.g., “Oct. 20, 2022”), as shown in FIG. 5. The vector 502 includes the measured distance 508 and the measured angle 510.
[0047] Turning back to FIG. 1, the graph generation module 124 passes the bipartite graph 126 for a group of text entries to a pairing optimization module 134. The pairing optimization module 134 applies a pairing algorithm 136 to the bipartite graph 126. The pairing algorithm 136 uses the distance properties and the angle properties associated with the edges 132 between the first vertices 128 and the second vertices 130 to identify an optimal pairing solution 138, from the possible pairing solutions.").
Regarding claim 3, Serry discloses wherein the area extracting unit considers one of four angles of a bounding box of the first area specifying object as the first angle, and considers one of four angles of a bounding box of the second area specifying object as the second angle ([0045] An element 404 (e.g., a handwritten text entry) that is directly to the right of the base element 402 has an angle of zero degrees (or alternatively three hundred and sixty degrees). An element 406 (e.g., a handwritten text entry) that is directly above the base element 402 has an angle of ninety degrees. An element 408 (e.g., a handwritten text entry) that is directly to the left of the base element 404 has an angle of one hundred and eighty degrees. And an element 410 (e.g., a handwritten text entry) that is directly below the base element 404 has an angle of two hundred and seventy degrees. Turning back to FIG. 2B, an example of an angle measurement 234 between the “DIRECTOR” typewritten text entry 210 and the “Jane D.” handwritten text entry 222 is two hundred and seventy degrees.).
Regarding claim 4, Serry discloses wherein if the area extracting unit determines plural pairs of area specifying objects that specifies plural rectangular areas among the detected area specifying objects (fig. 2b, plurality of rectangles around text and handwritten words between lines 228 and 230.),
the area extracting unit (a) determines plural combination patterns of the plural pairs on the basis of positions of the first area specifying objects and the second area specifying objects ("[0038] Looking back to FIG. 1, the bipartite graph 126 represents all the possible pairing solutions. Each pairing solution pairs a typewritten text entry 106 with a handwritten text entry 108. Accordingly, the bipartite graph 126 includes first vertices 128 on one side (e.g., the left side) and second vertices 130 on the other side (e.g., the right side). The first vertices 128 correspond to the typewritten text entries in the set of typewritten text entries of the group and the second vertices 130 correspond to the handwritten text entries in the set of handwritten text entries of the group. Furthermore, the bipartite graph 126 includes edges 132 that connect the first vertices 128 and the second vertices 130.
[0039] FIG. 3A illustrates an example bipartite graph 300 for the group of text entries that includes text entries 204, 206, 216, 218, 208, 220 from FIGS. 2A and 2B. The bipartite graph 302 includes possible pairing solutions 304, 306, 308, 310, 312, 314. Each pairing solution includes first vertices 128 (e.g., represented by the ovals) on the left for the typewritten text entries “SCENE” 204, “TAKE” 206, and “LOCATION” 208."),
(b) among the plural combination patterns, selects a combination pattern that has a largest sum of the confidences of the first and second area specifying objects ([0048] Consequently, the pairing algorithm 136 considers the possible pairing solutions and, based on the assumptions, identifies the optimal pairing solution 138 by minimizing the standard deviation of the measured distances between paired text entries, by minimizing the circular standard deviation of the measured angles for the paired text entries, by minimizing the sum of the measured distances between the paired text entries, and/or by minimizing a sum of unlikelihood scores for the paired text entries, which are calculated based on the measured angles.), and
(c) determines the plural pairs with the selected combination pattern ([0039] Moreover, each pairing solution includes second vertices 130 on the right for the handwritten text entries “Romantic Walk” 216, “15” 218, and “City Park” 220. While the vertices are consistent across the possible pairing solutions 304, 306, 308, 310, 312, 314, the edges 132 that connect the first vertices 128 representing the typewritten text entries to the second vertices 130 representing the handwritten text entries vary from one pairing solution to the next.).
Regarding claim 5, Serry discloses the area specifying objects included by the plural document images in the training data have ([0005] As described herein, a location can be represented by a bounding box that contains the typewritten text or the handwritten text. The line detection module detects lines that separate the text entries in the form
[0080] Example Clause D, the method of any one of Example Clauses A through C, further comprising normalizing the distance determined for each typewritten text entry in the set of typewritten text entries based on a height and a width of an image that contains the form.).
Serry does not explicitly disclose but Agarwal teaches wherein for the leaner the machine learning has been performed using as training data plural document images of which each document image includes a pair of area specifying objects that specify a rectangular area ("[0041] In one or more embodiments, the page segmentation model 108 uses machine learning to segment the document 400 based on predicted component or object types. In some embodiments, the component or object types predicted by the page segmentation model 108 can include heading, text, figure, footnote, table, and list-items.
[0047] In one or more embodiments, the repeating objects detection module 112 can further compare font types, font size, styles, etc. of the objects in each merged object cluster unit as an additional check to determine whether the merged object cluster units are properly identified as part of a repeating structure group of objects. The result of this process is identifying merged object cluster units 502-508 as being part of a repeating structure group of objects (e.g., repeating structure group of objects 510). While the example of FIG. 5 includes a single repeating structure group of objects 510, other documents can include multiple repeating structure groups of objects.").
the area specifying objects included by the plural document images in the training data have plural colors ([0058] The repeating objects detection module 708 then compares attributes of the objects within each merged object cluster unit, such as font size, font type, text color, etc., to identify merged object cluster units that include a similar distribution of predicted objects.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention, to combine the known system of pairing solution in handwritten form recognition, as taught by Serry, with the known teaching of machine learning, as taught by Agarwal, in order to yield the predictable results of automating data extraction by matching individual pen strokes or text blocks to predefined form fields with high precision.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 20110270851 A1 with regards to claim 5: [0028] FIG. 3 illustrates the properties of an object in relation to a node and an edge. Features that are possessed by a node when document data is converted to a labeled directed graph mainly include text, a bitmap image, and graphical properties. The content of text includes a character string. A bitmap image includes the user ID of the author and the area. Graphical properties include a foreground color, a background color, a line style, a width, a height, a shape, and an area.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to AHMED A NASHER whose telephone number is (571)272-1885. The examiner can normally be reached Mon - Fri 0800 - 1700.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emily Terrell can be reached at (571) 270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AHMED A NASHER/Examiner, Art Unit 2675
/EMILY C TERRELL/Supervisory Patent Examiner, Art Unit 2666