Prosecution Insights
Last updated: October 02, 2026
Application No. 18/965,318

Quad-Photodiode (QPD) Image Deblurring Using Convolutional Neural Network

Non-Final OA §103
Filed
Dec 02, 2024
Priority
Jan 02, 2024 — provisional 63/616,874
Examiner
PHAM, NHUT HUY
Art Unit
Tech Center
Assignee
OmniVision Technologies Inc.
OA Round
1 (Non-Final)
82%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
64 granted / 78 resolved
+22.1% vs TC avg
Strong +24% interview lift
Without
With
+23.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
15 currently pending
Career history
91
Total Applications
across all art units

Statute-Specific Performance

§101
8.9%
-31.1% vs TC avg
§103
60.3%
+20.3% vs TC avg
§102
14.2%
-25.8% vs TC avg
§112
14.5%
-25.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 78 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION The United States Patent & Trademark Office appreciates the application that is submitted by the inventor/assignee. The United States Patent & Trademark Office reviewed the following application and has made the following comments below. Information Disclosure Statement The information disclosure statement (IDS) submitted on 12/02/2024 is considered and attached. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 3, 9-10, 12 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (US20240127407A1, filed 2022, hereinafter Li) in view of Yang et al. (US-20240303963-A1, filed 2023, hereinafter Yang), and further in view of Dudhane et al. (Dudhane, Akshay, et al. "Burst Image Restoration and Enhancement" IEEE, published 2022, hereinafter Dudhane). CLAIM 1 In regards to Claim 1, Li teaches a system for deblurring (Li, Abstract: “This application describes apparatuses and systems for rendering the Bokeh effect using multi-pixel microlenses”, ¶ [0002]: “The Bokeh effect is used in photography to produce images where the closer objects look sharp and everything else stays out-of-focus.”) quad-photodiode (QPD) image (Li, ¶ [0055]: “the image sensing module 4020 may include a plurality of pixels (also called a pixel array) and a plurality of quad-pixel microlenses 4010 to capture a target image. The resolution of the target image may be proportional to the number of pixels in the image sensor. The plurality of pixels may be segmented into a plurality of pixel groups, each pixel group including multiple pixels (e.g., two-by-two pixels). Each of the plurality of quad-pixel microlens covers one pixel group. The microlens includes a spherical surface and a plane surface”), comprising: a blurred input QPD image comprising a plurality of pixel units, each pixel unit comprising four pixels, an Up-left (Ul) pixel, an Up-right (Ur) pixel, a Down-left (Dl) pixel, and a Down-right (Dr) pixel, under a microlens (Li, ¶ [0041]: “ For instance, if the microlens 2010 is a quad-pixel microlens (covering two-by-two pixels), the underlying four pixels may capture four different views from four different angles caused by the refractions. The views may be referred to as top-left view, top-right view, bottom-left view, and bottom-right view, respectively”; ¶ [0055]: “the image sensing module 4020 may include a plurality of pixels (also called a pixel array) and a plurality of quad-pixel microlenses 4010 to capture a target image. The resolution of the target image may be proportional to the number of pixels in the image sensor. The plurality of pixels may be segmented into a plurality of pixel groups, each pixel group including multiple pixels (e.g., two-by-two pixels). Each of the plurality of quad-pixel microlens covers one pixel group. The microlens includes a spherical surface and a plane surface”); an input unit collecting the Ul pixels in a Ul view, the Ur pixels in a Ur view, the Dl pixels in a Dl view, and the Dr pixels in a Dr view (Li, ¶ [0056-0057]: “The image signal processor 4030 may be configured to process the multi-angle views captured by the image sensing module 4020 … the angle information computation module 4031 may extract angle information based on the multiple angle views from the image sensing module 4020. For instance, the multiple angle views collected by each pixel group may include a top-left view, a top-right view, a bottom-left view, and a bottom-right view.” Li teaches an image processor that collects and processes pixels from 4 views of a quad-pixel image); the input unit (Li, ¶ [0056-0058]: “The image signal processor 4030 … ”) defines a U view (Li, ¶ [0056-0058]: “The image signal processor 4030 … The computation may be based on one or more combinations of two views selected from the top-left view, the top-right view, the bottom-left view, and the bottom-right view” Li teaches combined view of two views selected from 4 views of a quad-pixel image) of the Ul view and the Ur view , a D view of the Dl view and the Dr view (Li, ¶ [0058]: “a top view and a bottom view from the top-left view, the top-right view, the bottom-left view, and the bottom-right view” Li teaches obtaining a top view by combining top-left and top-right views; and obtaining a bottom view by combining bottom-left and bottom-right views), a L view of the Ul view and the Dl view, and a R view of the Ur view and the Dr view (Li, ¶ [0058]: “a left view and a right view from the top-left view, the top-right view, the bottom-left view, and the bottom-right view” Li teaches obtaining a left view by combining top-left and bottom-left views; and obtaining a right view by combining top-right and bottom-right views); Li does not explicitly discloses a U view as a mean of the Ul view and the Ur view, a D view as a mean of the Dl view and the Dr view, a L view as a mean of the Ul view and the Dl view, and a R view as a mean of the Ur view and the Dr view; (underlined emphasizes the limitation not taught by Li) Yang is in the same field of art of processing quad-pixel images. Further, Yang teaches a U view as a mean of the Ul view and the Ur view, a D view as a mean of the Dl view and the Dr view, a L view as a mean of the Ul view and the Dl view, and a R view as a mean of the Ur view and the Dr view. (Yang, ¶ [0002]: “Quad-pixel image sensor is a type of digital image sensor used in cameras or other imaging devices. It may include quad-pixel units, also called quaternions (or quad), each of which includes four pixels arranged in a square or rectangular pattern”; ¶ [0036]: “ the quaternion binning calculation module 200 may be configured to compute the signal intensity (e.g., brightness level) and a corresponding standard deviation for each quaternion. FIG. 3 illustrates an example process in a quaternion binning calculation module 200, according to some embodiments of this specification. In particular, the signal intensity may be represented as an average pixel value of the four pixels in the quaternion, denoted as a bin”. Yang teaches Pixel Averaging/Binning - combining data from neighboring pixels (0_a, PNG media_image1.png 404 364 media_image1.png Greyscale 0_b, 0_c and 0_d) into a single larger pixel/bin, see modified FIG. 3 below) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method of combining pixels of Li with the pixel averaging/binning technique of Yang to create a system that combine adjacent pixels into a larger pixel (Li) by using pixel averaging/binning technique (Yang) because Yang teaches that pixel averaging/binning provides benefits of reducing image noise and improving image quality (Yang, ¶ [0033]: “The input 210 may include a quad-image 211 captured by a quad-pixel image sensor and data from a programmable memory, such as an one-time programmable memory 212. The output 230 may include a compensated quad image in which some of the pixels are compensated. Here, the compensation refers to adjusting the gain or offset (e.g., pixel value) of the pixel so that the signal levels of the pixels are smoothened, thereby reducing visual artifacts and improving the overall quality of the image”), and one of ordinary skill would have recognized that incorporating this feature into the pixel combining method of Li would allow the pixel combining to smooth the differences in gain and offset between the individual image pixels within a quad, and between different quads (Yang, ¶ [0003]: “ The variations and imbalance would result in visible artifacts in the final image captured by a quad image sensor, such as uneven brightness and color between the different parts of the image. The goal of the correction process is to smooth the differences in gain and offset between the individual image pixels within a quad, and between different quads.”) The combination Li and Yang does not explicitly disclose a convolutional neural network (CNN), wherein the U view, the D view, the L view, and the R view are input to the CNN; wherein the CNN outputs an output Bayer image, which is a deblurred image of the blurred input QPD image. Dudhane is in the same field of art of image enhancement and restoration. Further, Dudhane teaches a convolutional neural network (CNN) (Dudhane, pages 5749-5750, Introduction: “The goal of burst imaging is to composite a high-quality image by merging desired information from a collection of (degraded) frames of the same scene captured in a rapid succession. However, burst image acquisition presents its own challenges. For example, during image burst capturing, any movement in camera and/or scene objects will cause misalignment issues, thereby leading to ghosting and blurring artifacts in the output image [53]. Therefore, there is a pressing need to develop a multi-frame processing algorithm PNG media_image2.png 687 761 media_image2.png Greyscale that is robust to alignment problems and requires no special burst acquisition conditions... we present a burst image processing approach, named BIPNet …”; see annotated FIG. 1 below. Dudhane teaches a CNN based network that take multiple frames of a scene as input, and output an enhanced image without blurring artifacts. Page 5753, section 3.3 and Fig. 3: “Pseudo bursts are processed with (shared) U-Net to extract multi-scale features” The Examiner notes U-Net is a convolutional neural network), wherein the U view, the D view, the L view, and the R view are input to the CNN (Dudhane, pages 5755, right col, section 4.2: “We prepare 28k patches of spatial size 128×128 with burst size 8 from the trainings et of Sony subset of SID to train the network for 50 epochs.” The Examiner notes burst size is the number of images in a burst, a burst size of 8 means the network takes 2 to 8 images as input; page 5754, left col, section 4.1: “Each burst contains 14 LR RAW images (each of size 48×48 pixels) that are synthetically generated from a single sRGB image” The Examiner notes the network supports up to 14 images input); wherein the CNN outputs an output Bayer image (Dudhane, page 5753, section 3.3, see reconstructed text below; See “output image” in annotated FIG. 1 above), PNG media_image3.png 183 748 media_image3.png Greyscale which is a deblurred image of the blurred input QPD image. (Dudhane, pages 5749-5750, Introduction: “The goal of burst imaging is to composite a high-quality image by merging desired information from a collection of (degraded) frames of the same scene captured in a rapid succession. However, burst image acquisition presents its own challenges. For example, during image burst capturing, any movement in camera and/or scene objects will cause misalignment issues, thereby leading to ghosting and blurring artifacts in the output image [53]. Therefore, there is a pressing need to develop a multi-frame processing algorithm that is robust to alignment problems and requires no special burst acquisition conditions... we present a burst image processing approach, named BIPNet …” The output image is free of ghosting and blurring artifacts) Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Li and Yang by incorporating the image enhancement network BIPNet that is taught by Dudhane, to make a system to output an enhanced image from multiple frames input; thus, one of ordinary skilled in the art would be motivated to combine the references since among its several aspects, the present invention recognizes there is a need to remove blurring artifacts due to misalignment/shifting between multiple images, and generate an enhanced, high-quality and artifact-free image (Dudhane, pages 5749-5750, Introduction: “The goal of burst imaging is to composite a high-quality image by merging desired information from a collection of (degraded) frames of the same scene captured in a rapid succession. However, burst image acquisition presents its own challenges. For example, during image burst capturing, any movement in camera and/or scene objects will cause misalignment issues, thereby leading to ghosting and blurring artifacts in the output image [53]. Therefore, there is a pressing need to develop a multi-frame processing algorithm that is robust to alignment problems and requires no special burst acquisition conditions... we present a burst image processing approach, named BIPNet …”). Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. CLAIM 3 Regarding claim 3, the combination of Li, Zhang and Dudhane teaches the system of Claim 1. In addition, the combination of Li, Zhang and Dudhane teaches wherein in the CNN: the U view is input to a first feature extraction unit; the D view is input to a second feature extraction unit; the L view is input to a third feature extraction unit; and the R view is input to a fourth feature extraction unit. (Dudhane, pages 5752-5753, section 3.2: “Existing burst image processing techniques [4,5] separately extract and align features of burst images and usually employ late feature fusion mechanisms, which can hinder flexible information exchange between frames. We instead propose a pseudo-burst feature fusion (PBFF) mechanism …”, see modified FIG. 3 below. Dudhane teaches a PBFF module to extract features from multiple images e1 to eB (B goes up to 14), e1 is input into its corresponding first feature extraction PNG media_image4.png 738 1089 media_image4.png Greyscale unit, e2 is input into its corresponding second feature extraction unit, etc.) CLAIM 9 Regarding claim 9, the combination of Li, Zhang and Dudhane teaches the system of Claim 1. In addition, the combination of Li, Zhang and Dudhane teaches the CNN is trained using input-output pairs of blurred QPD images and ground truth deblurred Bayer images. (Dudhane, page 5754, section 4.1: “(2) BurstSR dataset consists of 200 RAW bursts, each containing 14 images. To gather these burst sequences, the LR images and the corresponding (ground-truth) HR images are captured with a smartphone camera and a DSLR camera, respectively. From 200 bursts, 5,405 patches are cropped for training and 882 for validation. Each input crop is of size 80×80 pixels.”) CLAIM 10 Regarding Claim 10, Li teaches A method for deblurring (Li, Abstract: “This application describes apparatuses and systems for rendering the Bokeh effect using multi-pixel microlenses”, ¶ [0002-0003]: “The Bokeh effect is used in photography to produce images where the closer objects look sharp and everything else stays out-of-focus…Various embodiments of this specification may include hardware circuits, systems, and methods”) quad-photodiode (QPD) image (Li, ¶ [0055]: “the image sensing module 4020 may include a plurality of pixels (also called a pixel array) and a plurality of quad-pixel microlenses 4010 to capture a target image. The resolution of the target image may be proportional to the number of pixels in the image sensor. The plurality of pixels may be segmented into a plurality of pixel groups, each pixel group including multiple pixels (e.g., two-by-two pixels). Each of the plurality of quad-pixel microlens covers one pixel group. The microlens includes a spherical surface and a plane surface”) comprising: providing a blurred input QPD image comprising a plurality of pixel units, each pixel unit comprising four pixels, an Up-left (Ul) pixel, an Up-right (Ur) pixel, a Down-left (Dl) pixel, and a Down-right (Dr) pixel, under a microlens (Li, ¶ [0041]: “ For instance, if the microlens 2010 is a quad-pixel microlens (covering two-by-two pixels), the underlying four pixels may capture four different views from four different angles caused by the refractions. The views may be referred to as top-left view, top-right view, bottom-left view, and bottom-right view, respectively”; ¶ [0055]: “the image sensing module 4020 may include a plurality of pixels (also called a pixel array) and a plurality of quad-pixel microlenses 4010 to capture a target image. The resolution of the target image may be proportional to the number of pixels in the image sensor. The plurality of pixels may be segmented into a plurality of pixel groups, each pixel group including multiple pixels (e.g., two-by-two pixels). Each of the plurality of quad-pixel microlens covers one pixel group. The microlens includes a spherical surface and a plane surface”); collecting the Ul pixels in a Ul view, the Ur pixels in a Ur view, the Dl pixels in a Dl view, and the Dr pixels in a Dr view (Li, ¶ [0056-0057]: “The image signal processor 4030 may be configured to process the multi-angle views captured by the image sensing module 4020 … the angle information computation module 4031 may extract angle information based on the multiple angle views from the image sensing module 4020. For instance, the multiple angle views collected by each pixel group may include a top-left view, a top-right view, a bottom-left view, and a bottom-right view.” Li teaches an image processor that collects and processes pixels from 4 views of a quad-pixel image); the input unit (Li, ¶ [0056-0058]: “The image signal processor 4030 … ”) defines a U view (Li, ¶ [0056-0058]: “The image signal processor 4030 … The computation may be based on one or more combinations of two views selected from the top-left view, the top-right view, the bottom-left view, and the bottom-right view” Li teaches combined view of two views selected from 4 views of a quad-pixel image) of the Ul view and the Ur view , a D view of the Dl view and the Dr view (Li, ¶ [0058]: “a top view and a bottom view from the top-left view, the top-right view, the bottom-left view, and the bottom-right view” Li teaches obtaining a top view by combining top-left and top-right views; and obtaining a bottom view by combining bottom-left and bottom-right views), a L view of the Ul view and the Dl view, and a R view of the Ur view and the Dr view (Li, ¶ [0058]: “a left view and a right view from the top-left view, the top-right view, the bottom-left view, and the bottom-right view” Li teaches obtaining a left view by combining top-left and bottom-left views; and obtaining a right view by combining top-right and bottom-right views); Li does not explicitly discloses a U view as a mean of the Ul view and the Ur view, a D view as a mean of the Dl view and the Dr view, a L view as a mean of the Ul view and the Dl view, and a R view as a mean of the Ur view and the Dr view; (underlined emphasizes the limitation not taught by Li) Yang is in the same field of art of processing quad-pixel images. Further, Yang teaches a U view as a mean of the Ul view and the Ur view, a D view as a mean of the Dl view and the Dr view, a L view as a mean of the Ul view and the Dl view, and a R view as a mean of the Ur view and the Dr view. (Yang, ¶ [0002]: “Quad-pixel image sensor is a type of digital image sensor used in cameras or other imaging devices. It may include quad-pixel units, also called quaternions (or quad), each of which includes four pixels arranged in a square or rectangular pattern”; ¶ [0036]: “ the quaternion binning calculation module 200 may be configured to compute the signal intensity (e.g., brightness level) and a corresponding standard deviation for each quaternion. FIG. 3 illustrates an example process in a quaternion binning calculation module 200, according to some embodiments of this specification. In particular, the signal intensity may be represented as an average pixel value of the four pixels in the quaternion, denoted as a bin”. Yang teaches Pixel Averaging/Binning - combining data from neighboring PNG media_image1.png 404 364 media_image1.png Greyscale pixels (0_a, 0_b, 0_c and 0_d) into a single larger pixel/bin, see modified FIG. 3 below) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method of combining pixels of Li with the pixel averaging/binning technique of Yang to create a system that combine adjacent pixels into a larger pixel (Li) by using pixel averaging/binning technique (Yang) because Yang teaches that pixel averaging/binning provides benefits of reducing image noise and improving image quality (Yang, ¶ [0033]: “The input 210 may include a quad-image 211 captured by a quad-pixel image sensor and data from a programmable memory, such as an one-time programmable memory 212. The output 230 may include a compensated quad image in which some of the pixels are compensated. Here, the compensation refers to adjusting the gain or offset (e.g., pixel value) of the pixel so that the signal levels of the pixels are smoothened, thereby reducing visual artifacts and improving the overall quality of the image”), and one of ordinary skill would have recognized that incorporating this feature into the combing pixels method of Li would smooth the differences in gain and offset between the individual image pixels within a quad, and between different quads (Yang, ¶ [0003]: “ The variations and imbalance would result in visible artifacts in the final image captured by a quad image sensor, such as uneven brightness and color between the different parts of the image. The goal of the correction process is to smooth the differences in gain and offset between the individual image pixels within a quad, and between different quads.”) The combination Li and Yang does not explicitly disclose inputting the U view, the D view, the L view, and the R view into a convolutional neural network (CNN); outputting an output Bayer image from the CNN, wherein the output Bayer image is a deblurred image of the QPD image. Dudhane is in the same field of art of image enhancement and restoration. Further, Dudhane teaches inputting the U view, the D view, the L view, and the R view (Dudhane, pages 5755, right col, section 4.2: “We prepare 28k patches of spatial size 128×128 with burst size 8 from the trainings et of Sony subset of SID to train the network for 50 epochs.” The Examiner notes burst size is the number of images in a burst, a burst size of 8 means the network takes 2 to 8 images as input; page 5754, left col, section 4.1: “Each burst contains 14 LR RAW images (each of size 48×48 pixels) that are synthetically generated from a single sRGB image” The Examiner notes the network supports up to 14 images input) into a convolutional neural network (CNN) (Dudhane, pages 5749-5750, Introduction: “The goal of burst imaging is to composite a high-quality image by merging desired information from a collection of (degraded) frames of the same scene captured in a rapid succession. However, burst image acquisition presents its own challenges. For example, during image burst capturing, any movement in camera and/or scene objects will cause misalignment issues, thereby leading to ghosting and blurring artifacts in the output image [53]. Therefore, there is a pressing need to develop a multi-frame processing algorithm that is robust to alignment problems and requires no special burst acquisition conditions... we present a burst image processing approach, named BIPNet …”; see annotated FIG. 1 below. Dudhane teaches a CNN based network that take multiple frames of a scene as input, and output an enhanced image without blurring artifacts. Page 5753, section 3.3 and Fig. 3: “Pseudo bursts are processed with (shared) U-Net to extract multi-scale features” The Examiner notes U-Net is a PNG media_image2.png 687 761 media_image2.png Greyscale convolutional neural network); PNG media_image3.png 183 748 media_image3.png Greyscale outputting an output Bayer image from the CNN (Dudhane, page 5753, section 3.3, see reconstructed text below; See “output image” in annotated FIG. 1 above), , wherein the output Bayer image is a deblurred image of the QPD image. (Dudhane, pages 5749-5750, Introduction: “The goal of burst imaging is to composite a high-quality image by merging desired information from a collection of (degraded) frames of the same scene captured in a rapid succession. However, burst image acquisition presents its own challenges. For example, during image burst capturing, any movement in camera and/or scene objects will cause misalignment issues, thereby leading to ghosting and blurring artifacts in the output image [53]. Therefore, there is a pressing need to develop a multi-frame processing algorithm that is robust to alignment problems and requires no special burst acquisition conditions... we present a burst image processing approach, named BIPNet …” The output image is free of ghosting and blurring artifacts) CLAIM 12 Regarding claim 12, the combination of Li, Zhang and Dudhane teaches the method of Claim 10. In addition, the combination of Li, Zhang and Dudhane teaches steps in the CNN: PNG media_image4.png 738 1089 media_image4.png Greyscale inputting the U view to a first feature extraction unit; inputting the D view to a second feature extraction unit; inputting the L view to a third feature extraction unit; and inputting R view to a fourth feature extraction unit. (Dudhane, pages 5752-5753, section 3.2: “Existing burst image processing techniques [4,5] separately extract and align features of burst images and usually employ late feature fusion mechanisms, which can hinder flexible information exchange between frames. We instead propose a pseudo-burst feature fusion (PBFF) mechanism …”, see modified FIG. 3 below. Dudhane teaches a PBFF module to extract features from multiple images e1 to eB (B goes up to 14), e1 is input into its corresponding first feature extraction unit, e2 is input into its corresponding second feature extraction unit, etc.) CLAIM 18 Regarding claim 18, the combination of Li, Zhang and Dudhane teaches the method of Claim 10. In addition, the combination of Li, Zhang and Dudhane teaches the CNN is trained using input-output pairs of blurred QPD images and ground truth deblurred Bayer images. (Dudhane, page 5754, section 4.1: “(2) BurstSR dataset consists of 200 RAW bursts, each containing 14 images. To gather these burst sequences, the LR images and the corresponding (ground-truth) HR images are captured with a smartphone camera and a DSLR camera, respectively. From 200 bursts, 5,405 patches are cropped for training and 882 for validation. Each input crop is of size 80×80 pixels.”) Claim(s) 2 and 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li, in view of Zhang in view of Dudhane, and further in view of PixelCraft (PixelCraft “Smartphone cameras, pixel-binning, and the art of megapixel hype” https://pixelcraft.photo.blog/2023/11/14/smartphone-cameras-pixel-binning-and-the-art-of-megapixel-hype/, published 2023, hereinafter PixelCraft). CLAIM 2 In regards to Claim 2, the combination of Li, Zhang and Dudhane teaches the system of claim 1. In addition, the combination of Li, Zhang and Dudhane teaches using quad-Bayer filter for capturing colored images (Li, ¶ [0036-0038]: “To capture colored images, the pixels are also covered by a CFA following the Bayer pattern … The color filters (e.g., CFA) may be arranged in the Bayer pattern, in which every two by two pixels are covered by one red filter, two green filters, and one blue filter.”) and pixel binning. (Zhang, ¶ [0035-0036], see modified FIG. 3 in the rejection of claim 1.) The combination of Li, Zhang and Dudhane does not explicitly disclose a size of the input QPD image is mxm and a size of the output Bayer image is ¼ (mxm), m is an integer. PixelCraft is in the same field of art of quad-Bayer filter and pixel binning. Further, PixelCraft teaches a size of the input QPD image is mxm and a size of the output Bayer image is ¼ (mxm), m is an integer. (PixelCraft, page 2: “Photosite binning artificially groups smaller pixels into larger ones, potentially boosting the amount of light that can be gathered. The example in Figure 1 shows part of the quad-Bayer sensor of the iPhone 14 Pro. It illustrates how a 12MP is generated from a 48MP sensor. Here four photosites (2×2) are binned from the sensor, producing a 12 megapixel image, i.e. 48÷4=12.”, see modified FIG. 1 below. PixelCraft teaches combining 4 neighboring pixels into 1 large pixel, which reduces the resolution 4 times, from 48mp to 12mp) PNG media_image5.png 453 643 media_image5.png Greyscale Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Li, Zhang and Dudhane by substituting Li’s quad-Bayer color filter with PixelCraft’s quad-Bayer color filter; thus, one of ordinary skilled in the art would be motivated to combine the references since it’s a simple substitution, Li teaches using quad-Bayer color filter and PixelCraft teaches a specific example of quad-Bayer filter used in iPhone 14 Pro. (PixelCraft, page 2: “Photosite binning artificially groups smaller pixels into larger ones, potentially boosting the amount of light that can be gathered. The example in Figure 1 shows part of the quad-Bayer sensor of the iPhone 14 Pro. It illustrates how a 12MP is generated from a 48MP sensor. Here four photosites (2×2) are binned from the sensor, producing a 12 megapixel image, i.e. 48÷4=12.”, see modified FIG. 1 below. PixelCraft teaches combining 4 neighboring pixels into 1 large pixel, which reduces the resolution 4 times, from 48mp to 12mp) Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. CLAIM 11 In regards to Claim 11, the combination of Li, Zhang and Dudhane teaches the method of claim 10. In addition, the combination of Li, Zhang and Dudhane teaches using quad-Bayer filter for capturing colored images (Li, ¶ [0036-0038]: “To capture colored images, the pixels are also covered by a CFA following the Bayer pattern … The color filters (e.g., CFA) may be arranged in the Bayer pattern, in which every two by two pixels are covered by one red filter, two green filters, and one blue filter.”) and pixel binning. (Zhang, ¶ [0035-0036], see modified FIG. 3 in the rejection of claim 1.) The combination of Li, Zhang and Dudhane does not explicitly disclose a size of the input QPD image is mxm and a size of the output Bayer image is ¼ (mxm), m is an integer. PixelCraft is in the same field of art of quad-Bayer filter and pixel binning. Further, PixelCraft teaches a size of the input QPD image is mxm and a size of the output Bayer image is ¼ (mxm), m is an integer. (PixelCraft, page 2: “Photosite binning artificially groups smaller pixels into larger ones, potentially boosting the amount of light that can be gathered. The example in Figure 1 shows part of the quad-Bayer sensor of the iPhone 14 Pro. It illustrates how a 12MP is generated from a 48MP sensor. Here four photosites (2×2) are binned from the sensor, producing a 12 megapixel image, i.e. 48÷4=12.”, see modified FIG. 1 below. PixelCraft teaches combining 4 neighboring pixels into 1 large pixel, which reduces the resolution 4 times, from 48mp to 12mp) PNG media_image5.png 453 643 media_image5.png Greyscale Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Li, Zhang and Dudhane by substituting Li’s quad-Bayer color filter with PixelCraft’s quad-Bayer color filter; thus, one of ordinary skilled in the art would be motivated to combine the references since it’s a simple substitution, Li teaches using quad-Bayer color filter and PixelCraft teaches a specific example of quad-Bayer filter used in iPhone 14 Pro. (PixelCraft, page 2: “Photosite binning artificially groups smaller pixels into larger ones, potentially boosting the amount of light that can be gathered. The example in Figure 1 shows part of the quad-Bayer sensor of the iPhone 14 Pro. It illustrates how a 12MP is generated from a 48MP sensor. Here four photosites (2×2) are binned from the sensor, producing a 12 megapixel image, i.e. 48÷4=12.”, see modified FIG. 1 below. PixelCraft teaches combining 4 neighboring pixels into 1 large pixel, which reduces the resolution 4 times, from 48mp to 12mp) Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Allowable Subject Matter Claims 4-8 and 13-17 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Pertinent Arts The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Zamir_1 et al. (Zamir, Syed Waqas, et al. "Learning enriched features for fast image restoration and enhancement." IEEE, published 2022) teaches a CNN-based network named MIRNet-v2 that learns enriched feature representations for image restoration and enhancement. Specifically, the system includes Selective Kernel Feature Fusion (SKFF) module that dynamically combines multi-resolution features using a self-attention mechanism. Its purpose is to adaptively adjust receptive fields, allowing the network to select informative features from different scales while keeping complementary details for tasks like image denoising and super-resolution without heavy computational overhead Zamir_2 et al. (Zamir, Syed Waqas, et al. "Restormer: Efficient transformer for high-resolution image restoration." IEEE, published 2022) teaches a Transformer model named Restoration Transformer (Restormer), which achieves state-of-the-art results on several image restoration tasks, including image deraining, single-image motion deblurring, defocus deblurring (single-image and dual-pixel data), and image denoising (Gaussian grayscale/color denoising, and real image denoising). Specifically, the system includes multiple Transformer blocks consists of (a) multi-Dconv head transposed attention (MDTA) and (b) gated-Dconv feed-forward network (GDFN). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to NHUT HUY (JEREMY) PHAM whose telephone number is (703)756-5797. The examiner can normally be reached Mo - Fr. 8:30am - 6pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, O'Neal Mistry can be reached on (313)446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /NHUT HUY PHAM/Examiner, Art Unit 2674 /ONEAL R MISTRY/Supervisory Patent Examiner, Art Unit 2674
Read full office action

Prosecution Timeline

Dec 02, 2024
Application Filed
Aug 24, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743877
TRAINING DATASET AUGMENTATION METHOD AND SYSTEM FOR TRAINING DEEP LEARNING NETWORK
2y 8m to grant Granted Sep 22, 2026
Patent 12744876
CROSS-VIEW ATTENTION FOR VISUAL PERCEPTION TASKS USING MULTIPLE CAMERA INPUTS
3y 0m to grant Granted Sep 22, 2026
Patent 12737844
METHOD, APPARATUS AND SOFTWARE PROGRAM FOR INCREASING RESOLUTION IN MICROSCOPY
3y 10m to grant Granted Sep 15, 2026
Patent 12737883
PORTABLE DEVICE FOR ENUMERATION AND SPECIATION OF FOOD ANIMAL PARASITES
2y 7m to grant Granted Sep 15, 2026
Patent 12731236
METHOD OF MEASURING STRUCTURE DISPLACEMENT BASED ON FUSION OF ASYNCHRONOUS VISION MEASUREMENT DATA OF NATURAL TARGET AND ACCELERATION DATA OF STRUCTURE, AND SYSTEM FOR THE SAME
2y 11m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
82%
Grant Probability
99%
With Interview (+23.5%)
2y 10m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 78 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month