DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Oath/Declaration
Oath/Declaration as filed on February 19, 2025 is noted by the Examiner.
Claim Objections
Claim 1 is objected to because of the following informalities:
In particular, the limitations “reinforce processing” and “blending the character area” in tenth and twelfth lines of the claim renders the claim indefinite, because the meaning of the coined terms “reinforce processing” and “blending the character area” recited in the tenth and twelfth lines of the claim are not apparent in light of the specification. See MPEP § 2173.05(a). Examiner recommends applicant amend the claim, without adding new matter, to positively recite in definite terms more clearly what “reinforce processing” and “blending the character area” actually are. In addition, any claim(s) dependent on claim 1 is objected to based on same above reasoning.
Claim 12 is objected to because of the following informalities:
In particular, the limitations “reinforce processing” and “blending the character area” in sixth and eighth lines of the claim renders the claim indefinite, because the meaning of the coined terms “reinforce processing” and “blending the character area” recited in the sixth and eighth lines of the claim are not apparent in light of the specification. See MPEP § 2173.05(a). Examiner recommends applicant amend the claim, without adding new matter, to positively recite in definite terms more clearly what “reinforce processing” and “blending the character area” actually are. In addition, any claim(s) dependent on claim 12 is objected to based on same above reasoning.
Claim 20 is objected to because of the following informalities:
In particular, the limitations “reinforce processing” and “blending the character area” in eighth and tenth lines of the claim renders the claim indefinite, because the meaning of the coined terms “reinforce processing” and “blending the character area” recited in the eighth and tenth lines of the claim are not apparent in light of the specification. See MPEP § 2173.05(a). Examiner recommends applicant amend the claim, without adding new matter, to positively recite in definite terms more clearly what “reinforce processing” and “blending the character area” actually are.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 12, and 20 are rejected under 35 U.S.C. 102(a)(1) and 102(a)(2) as being anticipated by Vincent et al., U.S. Patent Application Publication 2008/0002893 A1 (hereinafter Vincent).
Regarding claim 1, Vincent teaches an electronic device comprising: at least one processor, comprising processing circuitry; and at least one camera; and memory storing instructions that, when executed by the at least one processor individually and/or collectively, cause the electronic device to (200, 202 FIGS. 2-3, paragraphs[0034]-[0035] of Vincent teaches FIG. 2 is a block diagram of an example text recognition system 200; the text recognition system 200 includes an image component 202, an image preprocessing module 204, a text detection component 206, a text box enhancement component 208, and character recognition component 210; image component 202 collects, stores, or otherwise manages one or more images for text recognition; image component 202 can include one or more image databases or can retrieve images from a data store such as one or more remote image databases; alternatively, the image component 202 can receive images for text recognition in realtime from a remote location, for example, as part of an image or video feed; the process of collecting and storing images can be automated or user driven; and the images can be retrieved, for example, as a result of a user input selecting one or more images for use in the text recognition process, and See also at least ABSTRACT, and paragraphs[0043]-[0044], and [0099]-[0102] of Vincent (i.e., Vincent teaches a text recognition system, which includes an image component that receives one or more images from at least cameras, that is implemented as one or more computer program instructions encoded on a computer-readable medium for execution by a processor and computer)):
obtain a plurality of images through the at least one camera; generate a first image using the plurality of images (FIGS. 2-3, paragraphs[0043]-[0044] of Vincent teaches the first step in the text recognition process 300 is to receive one or more images (e.g., from the image component 202) (step 302); the images can be received from numerous sources including local storage on a single computer or multiple computing devices distributed across a network; for example, the images can be retrieved from one or more, local or remote, image databases or can be collected in realtime for processing; the received images may have been captured, for example, using conventional digital cameras or video recording devices; the resulting captured images can include panoramic images, still images, or frames of digital video; and the captured images can also be associated with three-dimensional ranging data as well as location information, which can be used in processing the images, and See also at least ABSTRACT, and paragraphs[0034]-[0035], and [0099]-[0102] of Vincent (i.e., Vincent teaches the text recognition system, which includes the image component that receives one or more images from at least cameras, that is implemented as the one or more computer program instructions encoded on the computer-readable medium for execution by the processor and computer));
based on identifying that the plurality of images are related to a text, identify a character area within the first image; generate a second image on which reinforce processing is performed on the character area within the first image; and generate an output image by blending the character area within the first image and a character area within the second image based on a text property of the character area within the first image (FIGS. 2-3, and 6-7D, paragraphs[0070]-[0072] of Vincent teaches the detected candidate text regions from a number of images that include the same text can be combined using the superresolution process to provide an enhanced candidate text region; FIG. 6 is an example process 600 for generating a superresolution image that provides an enhanced candidate text region; A number of frames or consecutive images are extracted (step 602); the number of extracted images depends on the capture rate of the camera as well as the number of images; typically, a greater number of images leads to a higher quality superresolution result; the candidate text regions from each extracted image are optionally enlarged to compensate for text detection errors (step 604) (i.e., to include text which may extend beyond the candidate text region detected by the classifier); FIG. 7A shows a set of similar images extracted for superresolution; and specifically, FIG. 7A shows a collection 700 of slightly different images 702, 704, 706, 708, and 710, each image including the same street sign for the street "LYTTON", and See also at least ABSTRACT, and paragraphs[0034]-[0035], [0043]-[0044], and [0099]-[0102] of Vincent (i.e., Vincent teaches the text recognition system, which includes the image component that receives one or more images from at least cameras, that is implemented as the one or more computer program instructions encoded on the computer-readable medium for execution by the processor and computer, that detects candidate text regions from a number of images that include the same text and extracts the candidate text regions enlarges them and scales them up to a high resolution image)).
Regarding claim 12, Vincent teaches a method performed by an electronic device, the method comprising (200 FIGS. 2-3, paragraphs[0034]-[0035] of Vincent teaches FIG. 2 is a block diagram of an example text recognition system 200; the text recognition system 200 includes an image component 202, an image preprocessing module 204, a text detection component 206, a text box enhancement component 208, and character recognition component 210; image component 202 collects, stores, or otherwise manages one or more images for text recognition; image component 202 can include one or more image databases or can retrieve images from a data store such as one or more remote image databases; alternatively, the image component 202 can receive images for text recognition in realtime from a remote location, for example, as part of an image or video feed; the process of collecting and storing images can be automated or user driven; and the images can be retrieved, for example, as a result of a user input selecting one or more images for use in the text recognition process, and See also at least ABSTRACT, and paragraphs[0043]-[0044], and [0099]-[0102] of Vincent (i.e., Vincent teaches a text recognition system, which includes an image component that receives one or more images from at least cameras, that is implemented as one or more computer program instructions encoded on a computer-readable medium for execution by a processor and computer)):
obtaining a plurality of images through at least one camera; generating a first image using the plurality of images (FIGS. 2-3, paragraphs[0043]-[0044] of Vincent teaches the first step in the text recognition process 300 is to receive one or more images (e.g., from the image component 202) (step 302); the images can be received from numerous sources including local storage on a single computer or multiple computing devices distributed across a network; for example, the images can be retrieved from one or more, local or remote, image databases or can be collected in realtime for processing; the received images may have been captured, for example, using conventional digital cameras or video recording devices; the resulting captured images can include panoramic images, still images, or frames of digital video; and the captured images can also be associated with three-dimensional ranging data as well as location information, which can be used in processing the images, and See also at least ABSTRACT, and paragraphs[0034]-[0035], and [0099]-[0102] of Vincent (i.e., Vincent teaches the text recognition system, which includes the image component that receives one or more images from at least cameras, that is implemented as the one or more computer program instructions encoded on the computer-readable medium for execution by the processor and computer));
based on identifying that the plurality of images are related to a text, identifying a character area within the first image; generating a second image on which reinforce processing is performed on the character area; and generating an output image by blending the character area within the first image and a character area within the second image based on a text property of the character area within the first image (FIGS. 2-3, and 6-7D, paragraphs[0070]-[0072] of Vincent teaches the detected candidate text regions from a number of images that include the same text can be combined using the superresolution process to provide an enhanced candidate text region; FIG. 6 is an example process 600 for generating a superresolution image that provides an enhanced candidate text region; A number of frames or consecutive images are extracted (step 602); the number of extracted images depends on the capture rate of the camera as well as the number of images; typically, a greater number of images leads to a higher quality superresolution result; the candidate text regions from each extracted image are optionally enlarged to compensate for text detection errors (step 604) (i.e., to include text which may extend beyond the candidate text region detected by the classifier); FIG. 7A shows a set of similar images extracted for superresolution; and specifically, FIG. 7A shows a collection 700 of slightly different images 702, 704, 706, 708, and 710, each image including the same street sign for the street "LYTTON", and See also at least ABSTRACT, and paragraphs[0034]-[0035], [0043]-[0044], and [0099]-[0102] of Vincent (i.e., Vincent teaches the text recognition system, which includes the image component that receives one or more images from at least cameras, that is implemented as the one or more computer program instructions encoded on the computer-readable medium for execution by the processor and computer, that detects candidate text regions from a number of images that include the same text and extracts the candidate text regions enlarges them and scales them up to a high resolution image)).
Regarding claim 20, Vincent teaches a non-transitory computer-readable storage medium configured to store instructions that, when executed by at least one processor individually or collectively, cause an electronic device to perform operations including (200, 202 FIGS. 2-3, paragraphs[0034]-[0035] of Vincent teaches FIG. 2 is a block diagram of an example text recognition system 200; the text recognition system 200 includes an image component 202, an image preprocessing module 204, a text detection component 206, a text box enhancement component 208, and character recognition component 210; image component 202 collects, stores, or otherwise manages one or more images for text recognition; image component 202 can include one or more image databases or can retrieve images from a data store such as one or more remote image databases; alternatively, the image component 202 can receive images for text recognition in realtime from a remote location, for example, as part of an image or video feed; the process of collecting and storing images can be automated or user driven; and the images can be retrieved, for example, as a result of a user input selecting one or more images for use in the text recognition process, and See also at least ABSTRACT, and paragraphs[0043]-[0044], and [0099]-[0102] of Vincent (i.e., Vincent teaches a text recognition system, which includes an image component that receives one or more images from at least cameras, that is implemented as one or more computer program instructions encoded on a computer-readable medium for execution by a processor and computer)):
obtaining a plurality of images through at least one camera; generating a first image using the plurality of images (FIGS. 2-3, paragraphs[0043]-[0044] of Vincent teaches the first step in the text recognition process 300 is to receive one or more images (e.g., from the image component 202) (step 302); the images can be received from numerous sources including local storage on a single computer or multiple computing devices distributed across a network; for example, the images can be retrieved from one or more, local or remote, image databases or can be collected in realtime for processing; the received images may have been captured, for example, using conventional digital cameras or video recording devices; the resulting captured images can include panoramic images, still images, or frames of digital video; and the captured images can also be associated with three-dimensional ranging data as well as location information, which can be used in processing the images, and See also at least ABSTRACT, and paragraphs[0034]-[0035], and [0099]-[0102] of Vincent (i.e., Vincent teaches the text recognition system, which includes the image component that receives one or more images from at least cameras, that is implemented as the one or more computer program instructions encoded on the computer-readable medium for execution by the processor and computer));
based on identifying that the plurality of images are related to a text, identifying a character area within the first image; generating a second image on which reinforce processing is performed on the character area; and generating an output image by blending the character area within the first image and a character area within the second image based on a text property of the character area within the first image (FIGS. 2-3, and 6-7D, paragraphs[0070]-[0072] of Vincent teaches the detected candidate text regions from a number of images that include the same text can be combined using the superresolution process to provide an enhanced candidate text region; FIG. 6 is an example process 600 for generating a superresolution image that provides an enhanced candidate text region; A number of frames or consecutive images are extracted (step 602); the number of extracted images depends on the capture rate of the camera as well as the number of images; typically, a greater number of images leads to a higher quality superresolution result; the candidate text regions from each extracted image are optionally enlarged to compensate for text detection errors (step 604) (i.e., to include text which may extend beyond the candidate text region detected by the classifier); FIG. 7A shows a set of similar images extracted for superresolution; and specifically, FIG. 7A shows a collection 700 of slightly different images 702, 704, 706, 708, and 710, each image including the same street sign for the street "LYTTON", and See also at least ABSTRACT, and paragraphs[0034]-[0035], [0043]-[0044], and [0099]-[0102] of Vincent (i.e., Vincent teaches the text recognition system, which includes the image component that receives one or more images from at least cameras, that is implemented as the one or more computer program instructions encoded on the computer-readable medium for execution by the processor and computer, that detects candidate text regions from a number of images that include the same text and extracts the candidate text regions enlarges them and scales them up to a high resolution image)).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 9 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Vincent, in view of Jones et al., U.S. Patent Application Publication 2017/0200296 A1 (hereinafter Jones).
Regarding claim 9, Vincent teaches the electronic device of the electronic device of wherein the instructions, when executed by the at least one processor individually and/or collectively, cause the electronic device to: (FIGS. 2-3, paragraphs[0034]-[0035] of Vincent teaches FIG. 2 is a block diagram of an example text recognition system 200; the text recognition system 200 includes an image component 202, an image preprocessing module 204, a text detection component 206, a text box enhancement component 208, and character recognition component 210; image component 202 collects, stores, or otherwise manages one or more images for text recognition; image component 202 can include one or more image databases or can retrieve images from a data store such as one or more remote image databases; alternatively, the image component 202 can receive images for text recognition in realtime from a remote location, for example, as part of an image or video feed; the process of collecting and storing images can be automated or user driven; and the images can be retrieved, for example, as a result of a user input selecting one or more images for use in the text recognition process, and See also at least ABSTRACT, and paragraphs[0043]-[0044], and [0099]-[0102] of Vincent (i.e., Vincent teaches a text recognition system, which includes an image component that receives one or more images from at least cameras, that is implemented as one or more computer program instructions encoded on a computer-readable medium for execution by a processor and computer)); but does not expressly teach identify a text area that has a probability of containing text greater than or equal to a reference value within the first image; and identify the character area within the identified text area.
However, Jones teaches identify a text area that has a probability of containing text greater than or equal to a reference value within the first image; and identify the character area within the identified text area (Claim 1 of Jones teaches a near-to-eye (NR2I) system providing improved legibility of text within an image to a user based upon a process comprising the steps of: acquiring an original image; processing the original image to establish a region of a plurality of regions, each region having a probability of character based content exceeding a threshold probability; determining whether the region of the plurality of regions is relevant to the user; and upon a positive determination: processing the region of the plurality of regions to extract character-based content; processing the extracted character based content in dependence upon an aspect of the user of the NR2I system to generate a modified region; and displaying the modified region in combination with the original image, and See also at least ABSTRACT, Claim 11, paragraphs[0016], [0022], and [0243]-[0249] of Jones (i.e., Jones teaches processing an images to establish each region that has a probability of character based content exceeding a threshold probability and relevant to a user, and processing each region to extract character-based content)).
Furthermore, Vincent and Jones are considered to be analogous art because they are from the same field of endeavor with respect to a text recognition system, and involve the same problem of forming the system to suitably identify a region of characters. Therefore, before the effective filing date of the claimed invention it would have been obvious to one of ordinary skill in the art to modify the method of Vincent based on Jones to identify a text area that has a probability of containing text greater than or equal to a reference value within the first image; and identify the character area within the identified text area. One reason for the modification as taught by Jones is to enhance textual based content displayed on a near-to-eye system (paragraph[0002] of Jones). The same motivation and rationale to combine for claim 9 mentioned above, in light of corresponding statement of grounds of rejection, applies to each dependent claim mentioned in the corresponding statement of grounds of rejection.
Regarding claim 11, Vincent teaches the electronic device of The electronic device of wherein the instructions, when executed by the at least one processor individually and/or collectively, cause the electronic device to: (FIGS. 2-3, paragraphs[0034]-[0035] of Vincent teaches FIG. 2 is a block diagram of an example text recognition system 200; the text recognition system 200 includes an image component 202, an image preprocessing module 204, a text detection component 206, a text box enhancement component 208, and character recognition component 210; image component 202 collects, stores, or otherwise manages one or more images for text recognition; image component 202 can include one or more image databases or can retrieve images from a data store such as one or more remote image databases; alternatively, the image component 202 can receive images for text recognition in realtime from a remote location, for example, as part of an image or video feed; the process of collecting and storing images can be automated or user driven; and the images can be retrieved, for example, as a result of a user input selecting one or more images for use in the text recognition process, and See also at least ABSTRACT, and paragraphs[0043]-[0044], and [0099]-[0102] of Vincent (i.e., Vincent teaches a text recognition system, which includes an image component that receives one or more images from at least cameras, that is implemented as one or more computer program instructions encoded on a computer-readable medium for execution by a processor and computer)); but does not expressly teach identify a plurality of characters within the text area within the first image, and wherein the character area within the first image includes individual characters among the plurality of characters.
However, Jones teaches identify a plurality of characters within the text area within the first image, and wherein the character area within the first image includes individual characters among the plurality of characters (Claim 1 of Jones teaches a near-to-eye (NR2I) system providing improved legibility of text within an image to a user based upon a process comprising the steps of: acquiring an original image; processing the original image to establish a region of a plurality of regions, each region having a probability of character based content exceeding a threshold probability; determining whether the region of the plurality of regions is relevant to the user; and upon a positive determination: processing the region of the plurality of regions to extract character-based content; processing the extracted character based content in dependence upon an aspect of the user of the NR2I system to generate a modified region; and displaying the modified region in combination with the original image, and See also at least ABSTRACT, Claim 11, paragraphs[0016], [0022], and [0243]-[0249] of Jones (i.e., Jones teaches processing an images to establish each region that has a probability of character based content exceeding a threshold probability and relevant to a user, and processing each region to extract character-based content)).
Potentially Allowable Subject Matter
Claims 2-8, 10, and 13-19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims, because for each of claims 2-8, 10, and 13-19 the prior art references of record do not teach the combination of all element limitations as presently claimed.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ABDUL-SAMAD A ADEDIRAN whose telephone number is (571)272-3128. The examiner can normally be reached on Monday through Thursday, 8:00 am to 5:00 pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amr Awad can be reached on 571-272-7764. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ABDUL-SAMAD A ADEDIRAN/Primary Examiner, Art Unit 2621