DETAILED ACTION
Contents
Notice of Pre-AIA or AIA Status 2
Claim Rejections - 35 USC § 102 2
Claim Rejections - 35 USC § 103 9
Allowable Subject Matter 12
Conclusion 12
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is responsive to applicant’s claim set received on 9/18/24. Claims 1-10 are currently pending.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless - (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 6 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Rao et al (US 10,262,229 B1). Regarding claim 1, Rao discloses an object detection device, comprising: an input device for receiving an input image (see col. 9, lines 40-54, col. 7, lines 1-30; (38) (3.1) Overall System Architecture
(39) FIG. 3 illustrates the overall system architecture, consisting of a stationary camera 300 that captures input image frames 302 and maps them to a set of detected object bounding boxes 304 that are output on a display 306. The pipeline includes one or more of the following operations/modules. (1) Construct Channels (element 308) From each consecutive pair of RGB (red, green, blue) image frames, the system constructs one or more of four channels including the usual three “color opponents” used in bio-inspired computer vision (red-green, blue-yellow, and intensity) were constructed (as described in Literature Reference No. 3), as well as a motion channel consisting of the difference in intensity between the two consecutive frames. (2) Resize Images (element 310) The one or more image channels are resized into one or multiple scales that specify the relative size of the objects of interest. In the image processing pipeline, it was found that using two scales with downsampling factors of 4× and 16×, respectively, provides adequate performance while keeping the number of required circuit components sufficiently small. (3) Apply Patch-based Spectral Saliency (element 312) To each resized image channel, a spectral saliency operation is applied to a grid of overlapping patches that tile the image that highlights the saliency pixels in the patch. As shown in FIG. 4, the spectral saliency operation includes the following steps: 1. A 2-D spectral transformation of the patch is performed (or equivalently a 1-D spectral transformation performed on all columns of the patch followed by a 1-D spectral transformation performed on all rows of the patch) (element 328). 2. A “sign” operation is performed that maps all positive values to +1, all negative values to −1, and zero values to 0 (element 330). 3. An inverse 2-D spectral transformation is performed (element 332). 4. A square operation is performed (element 334). In one embodiment, the system described herein uses, as the spectral transformation, the Walsh-Hadamard transform, which is described in detail below. (4) Combine Patches (element 314) The saliency responses from overlapping patches are combined in a weighted sum, where each contribution is weighted by its distance from its respective patch center. (5) Combine Channels (element 316) The saliency maps from each image channel are combined into one aggregate saliency map via weighted averaging. (6) Apply Smoothing Filter (element 318) The raw saliency maps detect aperiodic salient edges (or boundaries). To turn these into salient object regions, the raw combined saliency map is smoothed with a simple 5×5 Gaussian filter. (7) Apply Adaptive Threshold (element 320) To determine which pixels in the saliency map correspond to salient objects, a simple adaptive threshold is applied that marks a pixel as part of a detected object if its saliency value is greater than three times the average value in the saliency map. (8) Apply Size and Morphological Filtering (elements 322, 324, and 326) To prune out regions that almost surely do not correspond to objects of interest, detected object regions are filtered based on simple size-based and morphological filtering. Specifically, regions with a number of pixels that is too small or too large, and thin regions where the number of pixels divided by the bounding box area is below a threshold are filtered out. Simple morphological operations, such as erosion and dilation, are also used to further prune regions. (9) Generate Bounding Boxes (element 304) A simple serial algorithm is used to determine the coordinates of the bounding box of each detected salient object region. These bounding box coordinates are the input to the back end.);
an object detection model for generating a plurality of input images of different resolutions using the input image (see col. 9, lines 40-col. Line 10, col. 8, lines 35-67; (33) Described is an object detection system that can be used to detect multiple salient objects of interest, such as cars, trucks, buses, persons and cyclists, in a sequence of wide-area video image frames taken from a stationary camera and output the bounding boxes of detect objects to a display (e.g., computer monitor, touchscreen). The system according to embodiments of the present disclosure finds salient objects by applying an efficient spectral transform to find salient object features that differ from smooth and periodically textured regions. The algorithms and modules have been designed to be implementable in low-power hardware, in particular, recently developed emerging device hardware, such as spiking circuits and coupled oscillators, that are capable of rapidly performing parallel operations that approximate “degree-of-match”, inner products, convolution, and filtering using very little power.
(34) Some embodiments of the system include one or more of several innovations that improve upon conventional software object detection pipeline that enable efficient (i.e., low-power and low circuit complexity) implementation in hardware. The spectral transformation according to various embodiments of the present disclosure is applied on an overlapping grid of small (64×64 pixels) local patches in the image, instead of on the whole image at one time. This reduces the number of inputs needed in the spectral saliency circuit by a factor of 30 (i.e., 30×). To mitigate the reduction in detection performance induced by operating on local patches, object detection is performed at two different scales (e.g., 4× and 16× downsampling, respectively), and the saliency responses from overlapping patches are combined using two-dimensional (2-D) triangular windows that reduce boundary effects at the boundary (or edge) of patches.), and obtaining a plurality of pyramid images of different resolutions using the input image (see col. 10,l ines 1-35; (3) Apply Patch-based Spectral Saliency (element 312) To each resized image channel, a spectral saliency operation is applied to a grid of overlapping patches that tile the image that highlights the saliency pixels in the patch. As shown in FIG. 4, the spectral saliency operation includes the following steps: 1. A 2-D spectral transformation of the patch is performed (or equivalently a 1-D spectral transformation performed on all columns of the patch followed by a 1-D spectral transformation performed on all rows of the patch) (element 328). 2. A “sign” operation is performed that maps all positive values to +1, all negative values to −1, and zero values to 0 (element 330). 3. An inverse 2-D spectral transformation is performed (element 332). 4. A square operation is performed (element 334). In one embodiment, the system described herein uses, as the spectral transformation, the Walsh-Hadamard transform, which is described in detail below. (4) Combine Patches (element 314) The saliency responses from overlapping patches are combined in a weighted sum, where each contribution is weighted by its distance from its respective patch center. (5) Combine Channels (element 316) The saliency maps from each image channel are combined into one aggregate saliency map via weighted averaging. (6) Apply Smoothing Filter (element 318) The raw saliency maps detect aperiodic salient edges (or boundaries). To turn these into salient object regions, the raw combined saliency map is smoothed with a simple 5×5 Gaussian filter. (7) Apply Adaptive Threshold (element 320) To determine which pixels in the saliency map correspond to salient objects, a simple adaptive threshold is applied that marks a pixel as part of a detected object if its saliency value is greater than three times the average value in the saliency map. (8) Apply Size and Morphological Filtering (elements 322, 324, and 326) To prune out regions that almost surely do not correspond to objects of interest, detected object regions are filtered based on simple size-based and morphological filtering. Specifically, regions with a number of pixels that is too small or too large, and thin regions where the number of pixels divided by the bounding box area is below a threshold are filtered out. Simple morphological operations, such as erosion and dilation, are also used to further prune regions. (9) Generate Bounding Boxes (element 304) A simple serial algorithm is used to determine the coordinates of the bounding box of each detected salient object region. These bounding box coordinates are the input to the back end.);
a processor for obtaining an important object image representative of a predetermined important object in the input image based on the plurality of pyramid images (see col. 2, lines 40-67. Col. 7, lines 1-30, col. 10, lines 1-35, col. 8, lines 1-50; FIG. 3 illustrates the overall system architecture, consisting of a stationary camera 300 that captures input image frames 302 and maps them to a set of detected object bounding boxes 304 that are output on a display 306. The pipeline includes one or more of the following operations/modules. (1) Construct Channels (element 308) From each consecutive pair of RGB (red, green, blue) image frames, the system constructs one or more of four channels including the usual three “color opponents” used in bio-inspired computer vision (red-green, blue-yellow, and intensity) were constructed (as described in Literature Reference No. 3), as well as a motion channel consisting of the difference in intensity between the two consecutive frames. (2) Resize Images (element 310) The one or more image channels are resized into one or multiple scales that specify the relative size of the objects of interest. In the image processing pipeline, it was found that using two scales with downsampling factors of 4× and 16×, respectively, provides adequate performance while keeping the number of required circuit components sufficiently small. (3) Apply Patch-based Spectral Saliency (element 312) To each resized image channel, a spectral saliency operation is applied to a grid of overlapping patches that tile the image that highlights the saliency pixels in the patch. As shown in FIG. 4, the spectral saliency operation includes the following steps: 1. A 2-D spectral transformation of the patch is performed (or equivalently a 1-D spectral transformation performed on all columns of the patch followed by a 1-D spectral transformation performed on all rows of the patch) (element 328). 2. A “sign” operation is performed that maps all positive values to +1, all negative values to −1, and zero values to 0 (element 330). 3. An inverse 2-D spectral transformation is performed (element 332). 4. A square operation is performed (element 334). In one embodiment, the system described herein uses, as the spectral transformation, the Walsh-Hadamard transform, which is described in detail below. (4) Combine Patches (element 314) The saliency responses from overlapping patches are combined in a weighted sum, where each contribution is weighted by its distance from its respective patch center. (5) Combine Channels (element 316) The saliency maps from each image channel are combined into one aggregate saliency map via weighted averaging. (6) Apply Smoothing Filter (element 318) The raw saliency maps detect aperiodic salient edges (or boundaries). To turn these into salient object regions, the raw combined saliency map is smoothed with a simple 5×5 Gaussian filter. (7) Apply Adaptive Threshold (element 320) To determine which pixels in the saliency map correspond to salient objects, a simple adaptive threshold is applied that marks a pixel as part of a detected object if its saliency value is greater than three times the average value in the saliency map. (8) Apply Size and Morphological Filtering (elements 322, 324, and 326) To prune out regions that almost surely do not correspond to objects of interest, detected object regions are filtered based on simple size-based and morphological filtering. Specifically, regions with a number of pixels that is too small or too large, and thin regions where the number of pixels divided by the bounding box area is below a threshold are filtered out. Simple morphological operations, such as erosion and dilation, are also used to further prune regions. (9) Generate Bounding Boxes (element 304) A simple serial algorithm is used to determine the coordinates of the bounding box of each detected salient object region. These bounding box coordinates are the input to the back end.); and
an output device for outputting the important object image (see col. 9, lines 40-67, col. 7, lines 40-67; The computer system 100 presented herein is an example computing environment in accordance with an aspect. However, the non-limiting example of the computer system 100 is not strictly limited to being a computer system. For example, an aspect provides that the computer system 100 represents a type of data processing analysis that may be used in accordance with various aspects described herein. Moreover, other computing systems may also be implemented. Indeed, the spirit and scope of the present technology is not limited to any single data processing environment. Thus, in an aspect, one or more operations of various aspects of the present technology are controlled or implemented using computer-executable instructions, such as program modules, being executed by a computer. In one implementation, such program modules include routines, programs, objects, components and/or data structures that are configured to perform particular tasks or implement particular abstract data types. In addition, an aspect provides that one or more aspects of the present technology are implemented by utilizing one or more distributed computing environments, such as where tasks are performed by remote processing devices that are linked through a communications network, or such as where various program modules are located in both local and remote computer-storage media including memory-storage devices).
Regarding claim 6, the claim is analyzed as a method that implements the limitations of claim 1 (see rejection of claim 1).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimedinvention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2-3, 7-8 are rejected under 35 U.S.C. 103 as being unpatentable over Rao et al (US 10,262,229 B1) in view of Xie et al (IEEE: “Pyramid Grafting Network for One-Stage High Resolution Saliency Detection”).
Regarding claims 2-3, Rao teaches all elements as mentioned above in claim 1. Rao does not teach expressly includes a backbone network including a plurality of layers for outputting a feature map from which semantic information is extracted from the input image; is configured to generate a highest resolution pyramid image using a plurality of first feature maps having a smaller size among the plurality of feature maps output by the backbone network; and is configured to generate a plurality of rest of pyramid images using a plurality of first feature maps having a smaller size among the plurality of feature maps output by the backbone network;
generate a low resolution input image and a high resolution input image; and is configured to obtain a plurality of low resolution pyramid images and a plurality of high resolution pyramid images using the low resolution input image and the high resolution input image.
Xie, in the same field of endeavor, teaches a backbone network including a plurality of layers for outputting a feature map from which semantic information is extracted from the input image (see section 4.1, 4.2); is configured to generate a highest resolution pyramid image using a plurality of first feature maps having a smaller size among the plurality of feature maps output by the backbone network (see section 4.1); and is configured to generate a plurality of rest of pyramid images using a plurality of first feature maps having a smaller size among the plurality of feature maps output by the backbone network (see section 4.1);
generate a low resolution input image and a high resolution input image (see 4.1, abstract); and is configured to obtain a plurality of low resolution pyramid images and a plurality of high resolution pyramid images using the low resolution input image and the high resolution input image (see section 4.1-4.2).
It would have been obvious (before the effective filing date of the claimed invention) or (at the time the invention was made) to one of ordinary skill in the art to modify Rao to utilize the cited limitations as suggested by Xie. The suggestion/motivation for doing so would have been to provide superior performance (see abstract). Furthermore, the prior art collectively includes each element claimed (though not all in the same reference), and one of ordinary skill in the art could have combined the elements in the manner explained above using known engineering design, interface and/or programming techniques, without changing a “fundamental” operating principle of Rao, while the teaching of Xie continues to perform the same function as originally taught prior to being combined, in order to produce the repeatable and predictable result. It is for at least the aforementioned reasons that the examiner has reached a conclusion of obviousness with respect to the claim in question.
Regarding claims 7-8, the claim is analyzed as a method that implements the limitations of claims 2-3 (see rejection of claims 2-3).
Allowable Subject Matter
Claims 4-5, 9-10 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding claims 4-5, 9-10, none of the references of record alone or in combination suggest or fairly teach wherein the processor: is configured to generate a plurality of low resolution important object images using the plurality of low resolution pyramid images; is configured to generate a plurality of high resolution important object images using the plurality of high resolution pyramid images; and is configured to obtain a final important object image by synthesizing the plurality of low resolution important object images and the plurality of high resolution important object images.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to EDWARD PARK. The examiner’s contact information is as follows:
Telephone: (571)270-1576 | Fax: 571.270.2576 | Edward.Park@uspto.gov
For email communications, please notate MPEP 502.03, which outlines procedures pertaining to communications via the internet and authorization. A sample authorization form is cited within MPEP 502.03, section II.
The examiner can normally be reached on M-F 9-6 CST.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, John M. Villecco, can be reached on (571) 272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/EDWARD PARK/Primary Examiner, Art Unit 2661