Detailed Action
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-4 are rejected under 35 U.S.C. 101.
Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea of detecting, recognizing, and comparing image information of character regions and objects, without significantly more.
The claim recites: “A character recognition apparatus for processing an input image to recognize characters included in the input image, the character recognition apparatus comprising: a computing circuit; and a memory that stores instructions being executable by the computing circuit, wherein, when executing the instructions, the computing circuit: detects a plurality of character regions in the input image, each of the plurality of character regions including a character or a character string made of a plurality of characters; determines a direction of the character or the character string in each of the character regions; recognizes the character or the character string in each of the character regions; generates a connected region by connecting at least two character regions including the character(s) or the character string(s) having a same direction, the at least two character regions being close to each other at a distance less than a threshold; and connects the character(s) or the character string(s) included in the connected region to each other, and wherein, when executing the instructions, the computing circuit: detects a first object region and a second object region in the input image, the first object region including a first object, and the second object region including a second object;
detects a first character region or a first connected region included in the first object region;
detects a second character region or a second connected region included in the second object region;
recognizes a first character or character string included in the first character region or the first connected region; recognizes a second character or character string included in the second character region or the second connected region; compares the first character or character string with a third character or character string stored in advance in association with the first object; and compares the second character or character string with a fourth character or character string stored in advance in association with the second object.”
The limitations, as drafted, are processes that, under their broadest reasonable interpretation,
cover performance of the limitation in the human mind. A person can visually detect character regions from an image, determine the direction of the characters, and recognize the characters. The person can further generate connected regions by observing distances between character regions and mentally grouping those that are in the same orientation and within a distance threshold. The person can also visually detect objects and character regions within them, and then mentally compare the recognized characters to characters stored in memory associated with those objects. These steps correspond to common pattern recognition and verification processes that can be performed entirely in the human mind.
The judicial exception is not integrated into a practical application. For example, the claim
recites the additional elements, “a computing circuit; and a memory that stores instructions being executable by the computing circuit”. These additional elements are recited at a high level of generality such that they amount to generic computer components performing generic computer functions. Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial expectation. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements are recited at a high-level of generality. It is therefore a judicial exception that is not integrated into a practical application, and does not include additional elements that are sufficient to amount to significantly more than the judicial exception. This claim is not patent eligible.
Claim 2 is rejected under 35 U.S.C. 101 because the claimed invention is directed to a further limitation
of the same abstract idea identified in the analysis of claim 1. For example, the person can mentally draw a box around each word or line of text and extend these boxes in a direction perpendicular to the text. The person can then determine if the extended regions overlap for region grouping. This claim is not patent eligible.
Claim 3 contains elements found analogous to claim 1. Therefore, claim 3 is similarly rejected under 35 U.S.C. 101.
Claim 4 contains elements found analogous to claim 1, with the addition of “A program including instructions executed by a computing circuit implemented in a character recognition apparatus…”. The additional element is recited at a high level of generality such that it amounts to merely using a computer as a tool to implement the abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. Therefore, claim 8 is similarly rejected under 35 U.S.C. 101.
Claim 4 is rejected under 35 U.S.C. 101 because the claimed invention is directed to nonstatutory subject matter (i.e., program per se. See MPEP § 2106.03).
Claim 4 recites “A program including instructions executed by a computing circuit implemented in a character recognition apparatus…”. While the claim recites a program executed by a computing circuit, the claim as a whole is directed to a set of instructions, or code, which is non-statutory subject matter. The broadest reasonable interpretation of “a program including instructions” encompasses products that do not have a physical or tangible form. The claim does not contain any structural limitations that would impart a physical or tangible form to the program, such as “non-transitory medium”.
The specification fails to provide any basis for interpreting the claim term “program” as anything other than the code itself. For example, paragraph 14 describes “the programs stored in the memory 12 and the storage device 13 are examples of instructions executable by the CPU 11”, which defines the program in terms of its function as executable instructions without structural recitations. Thus, the claim is not considered patent-eligible subject matter because it does not fall within any of the four statutory categories of appropriate subject matter for a patent: process, machines, manufactures and composition of matter. Therefore, claim 4 is rejected under 35 U.S.C. 101.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 3, and 4 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Ishitani (“Document image layout analysis through the interaction of data-driven and concept-driven processing”, Journal of Information Processing Society of Japan 42.11, 2001).
Regarding claim 1, Ishitani teaches a character recognition apparatus (Ishitani, see Fig. 3, Layout Analysis System) for processing an input image to recognize characters included in the input image, the character recognition apparatus comprising:
a computing circuit; and a memory that stores instructions being executable by the computing circuit (Ishitani, pg. 2720, 2nd column, 1st full paragraph, lines 7-9, “The experimental system was a PC. The system is written in C language on a CPU with a clock speed of 366 MHz.”), wherein, when executing the instructions, the computing circuit:
detects a plurality of character regions in the input image, each of the plurality of character regions including a character or a character string made of a plurality of characters (Ishitani, pg. 2713, 2nd column, 4th full paragraph, lines 1-8, “Prior to layout analysis processing, binarization and tilt correction processing19) are sequentially applied to the input document image . The resulting image has the upper left corner as the origin. Labeling is applied to this image to extract connected components of black pixels 20). Each connected component is represented by its bounding rectangle. Connected components are classified into one of the following categories: "character component", "field separator", or "other" according to the following conditional expression.”, see Fig. 1, (c) Layout objects (Text lines), A document layout analysis system performs detection of text lines and their bounding boxes, as shown in Fig. 1 (c), for input images.);
determines a direction of the character or the character string in each of the character regions; recognizes the character or the character string in each of the character regions (Ishitani, pg. 2714, 1st column, lines 18-33 and 2nd column, lines 1-8, “The layout element Oi is composed of a set of character lines Si, a set of character components Ci, and text parameters Pi, as defined below… The bounding rectangle data of the character line pattern, described by SY 1iu), the position coordinate of the bottom right corner SP2iu = (SX2iu, SY 2iu), width SWiu, and height SHiu , and the data calculated during character recognition processing. It consists of the confidence score SCFiu of a character line. The character component Cig is represented by the bounding rectangle data of the character pattern, described by the top-left corner coordinates CP1ig = (CX1ig, CY 1ig), the bottom-right corner coordinates CP2ig = (CX2ig, CY 2ig), the width CWig, and the height CHig, as well as the confidence score CCFig and character code Codeig obtained by the character recognition process. The text parameter Pi is the character line direction TDi ( horizontal : 0, vertical: 1, unknown: -1 ), and the character size OCSi (OCWi =(OCWi, OCHi)).”, For each of the detected character region, the system determines a direction of the characters and performs character recognition.);
generates a connected region by connecting at least two character regions including the character(s) or the character string(s) having a same direction, the at least two character regions being close to each other at a distance less than a threshold; and connects the character(s) or the character string(s) included in the connected region to each other (Ishitani, pg. 2714, 2nd column, 2nd full paragraph, lines 1-20, “When an initial set of layout elements O is given to the layout analysis system, the domain integration module first performs grouping of layout elements based on the laws of proximity, similarity, and continuity, forming partial clusters as shown in Figures 4 (b) and (c) . The grouping integrates two layout elements Oi and Oj in the following procedure to form a new layout element. It is based on the principle of forming [something]. (1) Setting the integration range of layout elements The horizontal integration range HMPi and the vertical integration range VMPi are set for the layout element Oi… (2) Grouping based on Gestalt laws Layout elements. Figure 5. Oi and Oj satisfy all of the following conditions: proximity condition 1 , similarity conditions 1-2 , and continuity condition 1.”, pg. 2715, 1st column, lines 1-21, “If the conditions are met, they are combined to generate a new layout… DISTH(i, j) is the minimum value d1 of the horizontal distance between character components that make up the layout element . Proximity condition 1: Oi and Oj are considered to be in close proximity if they satisfy either of the following horizontal proximity conditions or vertical proximity conditions. Horizontal proximity condition: DISTH(i, j) < min(HMPi, HMPj ) and OV LPH(i, j) > 0. Vertical proximity condition: DISTV (i, j) < min(VMPi, VMPj ) and OV LPV (i, j) > 0. Similarity condition 1: Oi and Oj are assumed to have the same character row direction.”, see Figs. 5 and 6, Connected regions are generated by grouping character regions with the same direction that are positioned within a threshold distance from each other.), and
wherein, when executing the instructions, the computing circuit:
detects a first object region and a second object region in the input image, the first object region including a first object, and the second object region including a second object (Ishitani, pg. 2719, 2nd column, lines 5-8, “First, clusters ( including those consisting of only one character) that satisfy the following condition are defined as "small layout objects." The identification of small layout objects is performed by the region recognition module.”, see Fig. 1, (b) Layout objects (Text blocks), First and second object region are detected from the input image, as shown in Fig. 1 (b), where each object region contains an object (text block).);
detects a first character region or a first connected region included in the first object region; detects a second character region or a second connected region included in the second object region; recognizes a first character or character string included in the first character region or the first connected region; recognizes a second character or character string included in the second character region or the second connected region (Ishitani, pg. 2713, 1st column, lines 3-10, “Assuming clusters are sub-regions of text, text parameters such as line direction, font size, and inter-character distance are estimated, and lines of text are extracted from the clusters based on these parameters. Region recognition: After extracting character patterns from the character lines that make up the cluster, character recognition processing is performed to calculate the confidence level of each character pattern and the confidence level of the character lines.”, see Fig. 1, (c) Layout objects (Text lines), Within each detected object region, the system detects character regions (text lines) and performs character recognition.);
compares the first character or character string with a third character or character string stored in advance in association with the first object; and compares the second character or character string with a fourth character or character string stored in advance in association with the second object (Ishitani, pg. 2713, 1st column, lines 6-10, “Region recognition: After extracting character patterns from the character lines that make up the cluster, character recognition processing is performed to calculate the confidence level of each character pattern and the confidence level of the character lines.”, pg. 2717, 1st column, lines 1-5, “The region recognition module first extracts a character row image from the layout element Oi by collecting character patterns for each character row using character row Siu and character component Cig . Next, character extraction/recognition processing based on reference 22) is applied to the character row image”, The character recognition process uses pattern matching and confidence scoring to compare recognized characters with character patterns stored in advance in a recognition database. These stored patterns can be considered associated with the object region in that the recognition process is applied to characters detected within that region.).
Claim 3 corresponds to claim 1, reciting a character recognition method to perform the steps according to claim 1. Ishitani teaches a character recognition method to perform the steps according to claim 1 (Ishitani, see Fig. 3, layout Analysis system and pgs. 2712-2713, section 2. Configuration of the layout analysis system). As indicated in the analysis of claim 1, Ishitani teaches all the limitations according to claim 1. Therefore, claim 3 is rejected for the same reason as claim 1.
Claim 4 corresponds to claim 1, with the addition of a program including instructions executed by a computing circuit implemented in a character recognition apparatus to execute the functions according to claim 1. Ishitani teaches the addition of a program including instructions executed by a computing circuit implemented in a character recognition apparatus to execute the functions according to claim 1 (Ishitani, pg. 2720, 2nd column, 1st full paragraph, lines 7-9, “The experimental system was a PC. The system is written in C language on a CPU with a clock speed of 366 MHz.”). As indicated in the analysis of claim 1, Ishitani teaches all the limitations according to claim 1. Therefore, claim 4 is rejected for the same reason as claim 1.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Ishitani (“Document image layout analysis through the interaction of data-driven and concept-driven processing”, Journal of Information Processing Society of Japan 42.11, 2001) in view of Kageyama et al. (JP 2020144719 A), (hereinafter Kageyama).
Regarding claim 2, Ishitani teaches the character recognition apparatus as claimed in claim 1, wherein, when executing the instructions, the computing circuit: generates extended regions; and generates the connected region by connecting at least two character regions including the character(s) or the character string(s) having the same direction, the at least two character regions being close to each other at the distance less than the threshold (Ishitani, pg. 2714, 2nd column, 2nd full paragraph, lines 1-20 and pg. 2715, 1st column, lines 1-19, “When an initial set of layout elements O is given to the layout analysis system, the domain integration module first performs grouping of layout elements based on the laws of proximity, similarity, and continuity, forming partial clusters as shown in Figures 4 (b) and (c)… (1) Setting the integration range of layout elements The horizontal integration range HMPi and the vertical integration range VMPi are set for the layout element Oi according to the following formula (see eq.) The updated integration range is constant less than 1. For example, the horizontal integration range is obtained by adding a margin ÿ to the horizontal character spacing HDi of the layout element Oi , so that the elements are close together as shown in Figure 5. (2) Grouping based on Gestalt laws Layout elements Oi and Oj satisfy all of the following conditions: proximity condition 1 , similarity conditions 1-2 , and continuity condition 1. If the conditions are met, they are combined to generate a new layout (see eq.) DISTH(i, j) is the minimum value d1 of the horizontal distance between character components that make up the layout element . Proximity condition 1: Oi and Oj are considered to be in close proximity if they satisfy either of the following horizontal proximity conditions or vertical proximity conditions. Horizontal proximity condition: DISTH(i, j) < min(HMPi, HMPj ) and OV LPH(i, j) > 0 Vertical proximity condition: DISTV (i, j) < min(VMPi, VMPj ) and OV LPV (i, j) > 0 Similarity condition 1: Oi and Oj are assumed to have the same character row direction.”).
Ishitani does not teach generates extended regions by extending each of the character regions in a direction perpendicular to the direction of the character or the character string included in the character region; and the at least two character regions having the extended regions overlapping each other, respectively.
However, Kageyama teaches generates extended regions by extending each of the character regions in a direction perpendicular to the direction of the character or the character string included in the character region; and the at least two character regions having the extended regions overlapping each other, respectively (Kageyama, “The text area classification unit 18 creates a text classification mask image that distinguishes the reference line area 3, the heading area 4, the photo area 5, and the advertising area 6 from the other areas from the paper image 2, and creates a text classification mask. The image is labeled, the circumscribing rectangle of the area with the same label is used as the text candidate area, and the writing orthogonal expansion processing that expands in the direction orthogonal to the writing direction with respect to the text candidate area is performed. As a result of the writing orthogonal direction expansion processing, the circumscribing rectangle when the group of the text candidate regions in which the pixels overlap is made into a single figure is defined as the text region 7.”, pg. 15, lines 15-22).
Ishitani teaches integrating character regions based on conditions of proximity and direction (Ishitani, pg. 2714, 2nd column, 2nd full paragraph, lines 1-6). Ishitani further teaches expanding character regions in horizontal and vertical directions to compute distance measurements between regions (Ishitani, pg. 2715, 1st column, lines 1-19), but does not teach integrating character regions by extending regions in a direction perpendicular to the character direction and connecting regions whose extended regions overlap. Kageyama teaches an orthogonal direction expansion process which groups text regions based on overlap of extended regions (see above). Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified the character region integration of Ishitani to include the region expansion and overlap analysis as taught by Kageyama (Kageyama, pg. 15, lines 15-22). The motivation for doing so would have been to prevent text areas from being classified in an excessively fragmented state, thereby improving character layout analysis accuracy (as suggested by Kageyama, pg. 15, lines 22-25, “As a result, even old newspapers in which the text area 7 and other areas are close to each other can be classified with high accuracy. In addition, it is possible to prevent the text area 7 from being classified in an excessively fragmented state.”). Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine the teachings of Ishitani with Kageyama to obtain the invention as specified in claim 2.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CONNOR LEVI HANSEN whose telephone number is (703)756-5533. The examiner can normally be reached Monday-Friday 9:00-5:00 (ET).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sumati Lefkowitz can be reached at (571) 272-3638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CONNOR L HANSEN/Examiner, Art Unit 2672
/SUMATI LEFKOWITZ/Supervisory Patent Examiner, Art Unit 2672