DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1–4, 6–11, 19–23, and 25–27 are pending. Claims 5, 12–18, and 24 are canceled. Claims 1, 3, 6, 9, 10, 19, 20, 22, 25, 26, and 27 are currently amended. Claims 11, 21, and 23 are previously presented. This action is **Non-Final**.
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-27 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2,9, 19-21, 26 is/are rejected under 35 U.S.C. 103 as being unpatentable over “Real-time Scene Text Detection with Differentiable Binarization”, hereafter referred to as Liao, in view of “Image compression based on octave convolution and semantic segmentation” hereafter referred to as Liu in further view of Woo “CBAM: Convolutional Block Attention Module”.
Liao discloses 1.A text recognition method, comprising:
acquiring
performing an M-level convolution process on the
merging the M pairs of target (Liao, Fig. 3 ,
PNG
media_image1.png
410
640
media_image1.png
Greyscale
Concat reads on merging; pairs of data; Liao takes an image and perform)
determining a probability map and a threshold map of the target image based on the target feature map, and calculating a binarization map of the target image based on the probability map and the threshold map; and determining a text area in the target image based on the binarization map, and recognizing text information in the text area. (Liao, Fig. 3,
PNG
media_image2.png
300
594
media_image2.png
Greyscale
; see probablity map, threshold map, binary map and the determined text)
Liao discloses normal convolution, but does not expressly disclose octave convolution, in particular
“acquiring a first high-frequency feature map and a first low-frequency feature map of a target image;
performing an M-level convolution process on the first high-frequency feature map and the first low-frequency feature map by M cascaded convolution modules to obtain M pairs of target high-frequency feature map and target low-frequency feature map of the target image, where M is a positive integer;
merging the M pairs of target high-frequency feature map and target low-frequency feature map to obtain a target feature map of the target image;”
Liu discloses “acquiring a first high-frequency feature map and a first low-frequency feature map of a target image;
performing an M-level convolution process on the first high-frequency feature map and the first low-frequency feature map by M cascaded convolution modules to obtain M pairs of target high-frequency feature map and target low-frequency feature map of the target image, where M is a positive integer;
merging the M pairs of target high-frequency feature map and target low-frequency feature map to obtain a target feature map of the target image;”(Liu, pg. 1-2,
PNG
media_image3.png
232
342
media_image3.png
Greyscale
PNG
media_image4.png
90
352
media_image4.png
Greyscale
PNG
media_image5.png
240
163
media_image5.png
Greyscale
Fig. 1. On the left is octave convolution, where f(·) is vanilla convolution, W is convolution kernel, pool(·) is average pooling, and upsample(·) is an up-sampling using nearest neighbor interpolation.)
It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to replace the standard convolution of Liao with the octave convolution of Liu.
The suggestion/motivation for doing so would have been to decrease the number of calculations and storage space.
Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Liao in view of Liu does not expressly disclose “wherein each convolution module comprises an attention unit, and wherein the method further comprises: adjusting a feature weight output by the convolution module through the attention unit.”
Woo discloses “wherein each convolution module comprises an attention unit, and wherein the method further comprises: adjusting a feature weight output by the convolution module through the attention unit.” (Woo, Fig. 1
PNG
media_image6.png
202
478
media_image6.png
Greyscale
)
It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to add the attention module as shown by Woo into the convolutional blocks.
The suggestion/motivation for doing so would have been to improve the representation of interests (tells where to focus).
Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Liao with Liu and to obtain the invention as specified in claim 1.
Liao in view of Liu in view of Woo discloses 2. The text recognition method according to claim 1, wherein the convolution module performs a convolution process on the first high-frequency feature map and the first low-frequency feature map, and the convolution process comprises: performing a first convolution process on the input first high-frequency feature map to obtain a second high-frequency feature map, and performing an up-sampling convolution process on the input first low-frequency feature map to obtain a second low-frequency feature map; acquiring the target high-frequency feature map based on the second high-frequency feature map and the second low-frequency feature map; performing a second convolution process on the input first low-frequency feature map to obtain a third low-frequency feature map, and performing a down-sampling convolution process on the input first high-frequency feature map to obtain a third high-frequency feature map; and acquiring the target low-frequency feature map based on the third low-frequency feature map and the third high-frequency feature map. (Liu, Fig. 1
PNG
media_image5.png
240
163
media_image5.png
Greyscale
Fig. 1. On the left is octave convolution, where f(·) is vanilla convolution, W is convolution kernel, pool(·) is average pooling, and upsample(·) is an up-sampling using nearest neighbor interpolation )
Liao in view of Liu in view of Woo discloses 9. (Currently Amended) The text recognition method according to claim [[5]] 1, wherein the determining of the probability map and the threshold map of the target image based on the target feature map and the calculating of the binarization map of the target image based on the probability map and the threshold map,(see claim 1), comprise: predicting a probability that each pixel in the target image is text based on the target feature map to obtain the probability map of the target image; predicting a binary result that each pixel in the target image is text based on the target feature map to obtain the threshold map of the target image; and performing an adaptive learning process by using a differentiable binarization function in combination with the probability map and the threshold map to obtain a best adaptive threshold, and acquiring the binarization map of the target image based on the best adaptive threshold and the probability map. (Liao, Fig. 3)
Claim 19 is rejected under similar grounds as claim 1.
Claim 20 is rejected under similar grounds as claim 1.
Claim 21 is rejected under similar grounds as claim 2.
Claim 26 is rejected under similar grounds as claim 9.
Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liao in view of Lui in further view of “AMultiplexed Network for End-to-End, Multilingual OCR”, hereafter referred to as Huang.
Liao in view of Liu in view of Woo discloses 11. The text recognition method according to claim 1,
But doesn’t expressly disclose “wherein the method further comprises: predicting a language in which the target image contains text based on the target feature map; and the recognizing of the text information in the text area comprises: determining a corresponding text recognition model according to the language in which the target image contains the text to recognize the text information in the text area.”
Huang discloses “wherein the method further comprises: predicting a language in which the target image contains text based on the target feature map; and the recognizing of the text information in the text area comprises: determining a corresponding text recognition model according to the language in which the target image contains the text to recognize the text information in the text area.” (Huang, Fig. 1,
PNG
media_image7.png
398
680
media_image7.png
Greyscale
)
It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to add the language detection and selection of language model of Huang to the system of Liao in view of Liu.
The suggestion/motivation for doing so would have been to provide translations for the users.
Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Liao in view of Lui and Huang to obtain the invention as specified in claim 11.
Allowable Subject Matter
Claims 3-4, 6-8, 10, 22-23,25, 27 would be allowable if rewritten to include all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GANDHI THIRUGNANAM whose telephone number is (571)270-3261. The examiner can normally be reached M-F 8:30-5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sumati Lefkowitz can be reached at 571-272-3638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/GANDHI THIRUGNANAM/Primary Examiner, Art Unit 2672