DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
1. This action is in response to the amendment filed on March 30th, 2026. Claims 1-6 and 12-18 have been amended. Claims 1-10 and 12-21 are pending. Claims 1-10 and 12-21 remain rejected in the application.
Response to Arguments
2. Applicant’s arguments with respect to claim 1, and similarly claims 12 and 13, filed on 3/30/2026, with respect to the rejection under 35 U.S.C. 102 have been fully considered, but are moot because of new grounds for rejection. Claim 1, and similarly claim 12, are now disclosed by Barkatullah. Claim 13 is now disclosed by Barkatullah, Niu, and Bi.
3. Regarding arguments to claims 2-10 and 14-21, they are dependent on independent claims 1 and 12 respectively. Applicant does not argue anything other than independent claim 1, and similarly claims 12 and 13. The limitations in those claims, in conjunction with combination, has previously been established and explained.
Claim Rejections - 35 USC § 102
4. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
5. Claims 1, 9, 12, and 21 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Barkatullah (US-2018/0139432-A1).
6. As per claim 1, Barkatullah discloses: A binocular image generation method, comprising:
(Barkatullah, Abstract, “A method for adjusting and generating enhanced 3D-effects for 2D to 3D image and video conversion applications …”)
obtaining an original image, wherein the original image is a monocular image; (Barkatullah, [0015], “Embodiments here relate to a method, apparatus, system, and computer program for modifying, enhancing or exaggerating 3D-image rendered given a mono-ocular (2D) image source and its depth map.”)
determining, based on depth values of a plurality of pixels in a salient region of the original image, (Barkatullah, [0006], “A greyscale image represents the depth map of an image in which each pixel is assigned a value between and including 0 and 255.” and [0015], “Embodiments here relate to a method, apparatus, system, and computer program for modifying, enhancing or exaggerating 3D-image rendered given a mono-ocular (2D) image source and its depth map.” and Abstract, “A method for adjusting and generating enhanced 3D-effects for 2D to 3D image and video conversion applications includes controlling a depth location of a zero parallax plane within a depth field of an image scene to adjust parallax of objects in the image scene, controlling a depth volume of objects in the image scene to one of either exaggerate or reduce 3D-effect of the image scene …”) a first depth value corresponding to a zero-disparity plane in a binocular image to be generated, (Barkatullah, Abstract, “A method for adjusting and generating enhanced 3D-effects for 2D to 3D image and video conversion applications includes controlling a depth location of a zero parallax plane within a depth field of an image scene to adjust parallax of objects in the image scene, controlling a depth volume of objects in the image scene to one of either exaggerate or reduce 3D-effect of the image scene …” and [0023], “FIG. 3 illustrates one embodiment of graphical user interface (GUI) 202 to enable the user to adjust the location of the zero plane …” and [0006], “A greyscale image represents the depth map of an image in which each pixel is assigned a value between and including 0 and 255. … The depth value of a pixel is used to calculate the horizontal (x-axis) offset of the pixel for left and right eye view images. In particular, if the calculated offset is w for pixel at position (x,y) in the original image, then this pixel is placed at position (x+w, y) in the left image and (x−w, y) in the right image. ... If the offset w is zero, the pixel appears on the screen plane.”) wherein the salient region is a region corresponding to an object of interest to a user, or a region in the original image that comprises main information; and (Barkatullah, Abstract, “A method for adjusting and generating enhanced 3D-effects for 2D to 3D image and video conversion applications includes controlling a depth location of a zero parallax plane within a depth field of an image scene to adjust parallax of objects in the image scene, controlling a depth volume of objects in the image scene to one of either exaggerate or reduce 3D-effect of the image scene …”)
generating the binocular image based on the first depth value and depth values of a plurality of pixels in the original image. (Barkatullah, Claim 1, “1. A method for adjusting and generating enhanced 3D-effects for real time and offline 2D to 3D image and video conversion applications consisting of: (a) controlling a depth location of a zero parallax plane within a depth field of an image scene to adjust parallax of objects in the image scene; (b) controlling a depth volume of objects in the image scene to one of either exaggerate or reduce 3D-effect of the image scene … and (f) generating an updated depth map file for a 2D-image based upon the controlling the depth location, the controlling the depth volume, the controlling the depth location, the increasing and decreasing depth volume, and the increasing and decreasing depth separation.” and Claim 2, “2. The method of claim 1, further comprising rendering an enhanced 3D-image using the updated depth map.” and [0006], “A greyscale image represents the depth map of an image in which each pixel is assigned a value between and including 0 and 255.”)
7. As per claim 9, Barkatullah discloses: The method according to claim 1, wherein the original image is a video frame; and (Barkatullah, Claim 7, “7. The method of claim 1, further comprising generating a depth map for each 2D-still image or a sequence of depth maps for each frame in a 2D-video.”)
after the generating the binocular image, the method further comprises: (See Barkatullah, [0015], [0016], and Claim 11 below.)
generating a stereoscopic video based on a plurality of binocular images. (Barkatullah, [0015], “The embodiments can take advantage of the computing power of general purpose CPU, GPU or dedicated FPGA or ASIC chip to process sequence of images from video frames of a streaming 2D-video to generate 3D video frames. Depending on the available processing capabilities of the processing unit and complexity of desired transformations, the conversion of 2D video frames to 3D can be done in real.” and [0016], “A user receives a streaming 2D-video from the internet or from a file stored on a local storage device. The user then uses the application GUI to adjust the quality and attributes of 3D-video in an automatic 2D video to 3D conversion and display it on the attached 3D display in real time.” and Claim 11, “11. The method of claim 2, further comprising one of displaying generated 3D image or video on and attached 3D display in real time ...”)
8. Claim 12 is similar in scope to claim 1 except for additional limitations that Barkatullah discloses: An electronic device, comprising: (Barkatullah, [0017], “In one embodiment, the 2D to 3D conversion process is implemented as a software application running on a computing device such as a personal computer, tablet computer or smart-phone.”)
at least one processor; and (Barkatullah, [0015], “The embodiments can take advantage of the computing power of general purpose CPU, GPU or dedicated FPGA or ASIC chip to process sequence of images from video frames of a streaming 2D-video to generate 3D video frames.”)
a storage apparatus configured to store at least one program, wherein the at least one program, when executed by the at least one processor, causes the at least one processor to implement a binocular image generation method, wherein the binocular image generation method comprises: (Barkatullah, [0015], “Embodiments here relate to a method, apparatus, system, and computer program for modifying, enhancing or exaggerating 3D-image rendered given a mono-ocular (2D) image source and its depth map. … Optionally, such control settings can be presented to the 3D-render engine as commands stored in a file and read by 3D-rendering application or routine. … The embodiments can take advantage of the computing power of general purpose CPU, GPU or dedicated FPGA or ASIC chip to process sequence of images from video frames of a streaming 2D-video to generate 3D video frames.”)
9. Claim 21, which is similar in scope to dependent claim 9 and independent claim 12, is thus rejected under the same rationale as described above.
Claim Rejections - 35 USC § 103
10. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
11. Claims 2-3 and 13-15 are rejected under 35 U.S.C. 103 as being unpatentable over Barkatullah (US-2018/0139432-A1) in view of Niu et al. (US-2012/0008852-A1, hereinafter "Niu"), and further in view of Bi et al. (US-2012/0140038-A1, hereinafter "Bi").
12. As per claim 2, Barkatullah discloses: The method according to claim 1, wherein the determining, based on depth values of a plurality of pixels in a salient region of the original image, (Barkatullah, [0006], “A greyscale image represents the depth map of an image in which each pixel is assigned a value between and including 0 and 255.” and [0015], “Embodiments here relate to a method, apparatus, system, and computer program for modifying, enhancing or exaggerating 3D-image rendered given a mono-ocular (2D) image source and its depth map.” and Abstract, “A method for adjusting and generating enhanced 3D-effects for 2D to 3D image and video conversion applications includes controlling a depth location of a zero parallax plane within a depth field of an image scene to adjust parallax of objects in the image scene, controlling a depth volume of objects in the image scene to one of either exaggerate or reduce 3D-effect of the image scene …”) a first depth value corresponding to a zero-disparity plane in a binocular image to be generated comprises: (Barkatullah, Abstract, “A method for adjusting and generating enhanced 3D-effects for 2D to 3D image and video conversion applications includes controlling a depth location of a zero parallax plane within a depth field of an image scene to adjust parallax of objects in the image scene, controlling a depth volume of objects in the image scene to one of either exaggerate or reduce 3D-effect of the image scene …” and [0023], “FIG. 3 illustrates one embodiment of graphical user interface (GUI) 202 to enable the user to adjust the location of the zero plane …” and [0006], “A greyscale image represents the depth map of an image in which each pixel is assigned a value between and including 0 and 255. … The depth value of a pixel is used to calculate the horizontal (x-axis) offset of the pixel for left and right eye view images. In particular, if the calculated offset is w for pixel at position (x,y) in the original image, then this pixel is placed at position (x+w, y) in the left image and (x−w, y) in the right image. ... If the offset w is zero, the pixel appears on the screen plane.”)
[[generating a first histogram based on the depth values of the plurality of pixels in the salient region of the original image; and]]
[[determining, based on a distribution of the depth values in the first histogram, the first depth value corresponding to the zero-disparity plane in the binocular image to be generated.]]
13. Barkatullah doesn't explicitly disclose but Niu discloses: generating a first histogram based on the depth values of the plurality of pixels in the salient region of the original image; and (Niu, Fig. 4A; [0024], “The depth histogram is a distribution (usually depicted as a graph) of the depth levels of pixels, in which each depth level has its counted number of pixels. If the histogram is depicted as a graph, the horizontal axis represents the depth levels and the vertical axis represents the corresponding number of pixels. FIG. 4A shows an exemplary depth histogram according to an original depth map ...”)
14. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 1 of Barkatullah to include the disclosure of generating a first histogram based on the depth values of the plurality of pixels in the salient region of the original image, of Niu. The motivation for this modification could have been to help determine a distribution of depth values of the original image based on the histogram. This not only can help inform which depth plane has the most “information” or objects represented, but also may allow adjustment of the depth planes by expanding or contracting how the data in the histogram is distributed.
15. Barkatullah in view of Niu doesn't explicitly disclose but Bi discloses: determining, based on a distribution of the depth values in the first histogram, the first depth value corresponding to the zero-disparity plane in the binocular image to be generated. (Bi, Fig. 7; [0055], “According to one example, ZDP determination module 226 may determine the typical disparity of the ROI by creating a histogram representing relative disparity of pixels of the ROI, and selecting a bin of the histogram with a largest number of pixels.” and [0054], “ZDP determination module 226 may, based on the identified ROI, determine a zero disparity plane (ZDP) for the display of 3D images.” and [0074], “According to another example, image processing module 120 may determine a typical disparity for the ROI by assigning pixels of the ROI to a histogram that comprises a number of bins that correspond to various disparity ranges as described below with respect to FIG. 7. According to these examples, image processing module 120 may identify a disparity range of a bin for which a largest number of pixels was assigned as a typical disparity of an ROI.”)
16. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 1 of Barkatullah in view of Niu to include the disclosure of determining, based on a distribution of the depth values in the first histogram, the first depth value corresponding to the zero-disparity plane in the binocular image to be generated, of Bi. The motivation for this modification could have been to utilize the histogram to determine a zero-disparity plane for the binocular image. Since the zero-disparity plane is a plane where no depth is perceived for a user, it is often chosen to provide visual accuracy of content and optical comfort for the user.
17. As per claim 3, Barkatullah in view of Niu, and further in view of Bi discloses: The method according to claim 2, wherein the determining, based on a distribution of the depth values in the first histogram, the first depth value corresponding to the zero-disparity plane in the binocular image to be generated comprises: (See rejection for claim 2.)
determining the first depth value corresponding to the zero-disparity plane in the binocular image to be generated, based on a depth value range in the first histogram within which a largest number of pixels are distributed. (Bi, Fig. 7; [0055], “According to one example, ZDP determination module 226 may determine the typical disparity of the ROI by creating a histogram representing relative disparity of pixels of the ROI, and selecting a bin of the histogram with a largest number of pixels.” and [0054], “ZDP determination module 226 may, based on the identified ROI, determine a zero disparity plane (ZDP) for the display of 3D images.” and [0074], “According to another example, image processing module 120 may determine a typical disparity for the ROI by assigning pixels of the ROI to a histogram that comprises a number of bins that correspond to various disparity ranges as described below with respect to FIG. 7. According to these examples, image processing module 120 may identify a disparity range of a bin for which a largest number of pixels was assigned as a typical disparity of an ROI.”)
18. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 2 of Barkatullah in view of Niu to include the disclosure of determining the first depth value corresponding to the zero-disparity plane in the binocular image to be generated, based on a depth value range in the first histogram within which a largest number of pixels are distributed, of Bi. The motivation for this modification could have been to utilize the histogram to determine a zero-disparity plane by determining which depth plane has the largest distribution of pixels for the binocular image. The depth plane with the largest distribution of pixels would contain the highest amount of image content among the depth planes. Since the zero-disparity plane is a plane where no depth is perceived for a user, it would likely be chosen to provide visual accuracy of the largest amount of content and provide optical comfort for the user.
19. Claim 13 is similar in scope to claims 1 and 12 except for additional limitations that Barkatullah in view of Niu, and further in view of Bi discloses: A non-transitory computer-readable storage medium … (Bi, [0107], “In one or more of these examples, the functions described herein may be implemented at least partially in hardware, such as specific hardware components or a processor. More generally, the techniques may be implemented in hardware, processors, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted over as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, computer-readable media generally may correspond to (1) tangible computer-readable storage media which is non-transitory or (2) a communication medium such as a signal or carrier wave.” and [0054], “As also shown in FIG. 2, image processing module 220 further includes a ZDP determination module 226. ZDP determination module 226 may be configured to receive, from ROI identification module 224, an indication of an ROI of captured images. ZDP determination module 226 may, based on the identified ROI, determine a zero disparity plane (ZDP) for the display of 3D images.”)
20. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of Barkatullah in view of Niu to include the disclosure of utilizing a non-transitory computer-readable storage medium, of Bi. The motivation for this modification could have been to provide an additional computing medium to perform a binocular image generation method.
21. Claim 14, which is similar in scope to dependent claim 2 and independent claim 12, is thus rejected under the same rationale as described above. The motivation for this modification is the same as claim 2.
22. Claim 15, which is similar in scope to dependent claims 2 and 3 and independent claim 12, is thus rejected under the same rationale as described above. The motivation for this modification is the same as claim 3.
23. Claims 4 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Barkatullah (US-2018/0139432-A1) in view of Ha et al. (US-2007/0081716-A1, hereinafter "Ha").
24. As per claim 4, Barkatullah discloses: The method according to claim 1, wherein the generating the binocular image based on the first depth value and depth values of a plurality of pixels in the original image comprises:
determining, based on the first depth value and the depth values of the plurality of pixels in the original image, (See rejection for claim 1.)
25. Barkatullah doesn't explicitly disclose but Ha discloses: [[determining, based on the first depth value and the depth values of the plurality of pixels in the original image,]] a plurality of displacement vectors of the plurality of pixels in the original image in each of a left-eye image and a right-eye image that are to be generated; and (Ha, [0045]-[0046], “FIG. 7 illustrates block-based disparity estimation (DE) according to an embodiment of the present invention. Referring to FIG. 7, a left-eye image is divided into N×N blocks of equal size. Blocks of a right-eye image which are most similar to corresponding blocks in the left-eye image are estimated using a sum of absolute difference (SAD) or a mean of absolute difference (MAD). In this case, a distance between a reference block and an estimated block is defined as a disparity vector (DV). Generally, a DV is assigned to each pixel in the reference image. However, to reduce the amount of computation required, it is assumed that the DVs of all pixels in a block are approximately the same in the block-based DE. The performing of DE on each pixel to obtain the DV for each pixel is called pixel-based DE.”)
processing the plurality of pixels in the original image based on the plurality of displacement vectors, to generate the left-eye image and the right-eye image. (Ha, [0068], “The left-eye image and the right-eye image are horizontally moved based on the determined horizontal movement value and the disparities between the left-eye image and the right-eye image are adjusted (S126). The disparity-adjusted left- and right-eye images are output and displayed.” and [0046]-[0047], “The performing of DE on each pixel to obtain the DV for each pixel is called pixel-based DE. The block-based DE or the pixel-based DE is used to estimate a disparity.” and [0062]-[0063], “The histogram generation unit 13 estimates the disparities between the right-eye image and the left-eye image, measures the frequency with which the estimated disparities occur, and generates a histogram for the disparities and the frequency. In this case, the block-based DE or the pixel-based DE described above or other methods may be used. The horizontal movement value determination unit 15 receives the generated histogram from the histogram generation unit 13 and determines a horizontal movement value for the left- and right-eye images.”)
26. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 1 of Barkatullah to include the disclosure of determining displacement vectors between pixels of the left and right eye images and processing the pixels based on the vectors to generate a stereo image pair, of Ha. The motivation for this modification could have been to allow a user to fully adjust the image disparity and the zero-disparity plane. Some users of a stereo viewing system may have higher or lower tolerance for image disparity. This would allow a user to customize how they view stereo content so they are comfortable while viewing.
27. Claim 16, which is similar in scope to dependent claim 4 and independent claim 12, is thus rejected under the same rationale as described above. The motivation for this modification is the same as claim 4.
28. Claims 5-6 and 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Barkatullah (US-2018/0139432-A1) in view of Rowell et al. (US-2019/0158813-A1, hereinafter "Rowell").
29. As per claim 5, Barkatullah discloses: The method according to claim 1, wherein before the generating the binocular image based on the first depth value and depth values of a plurality of pixels in the original image, the method further comprises: (Barkatullah, [0006], “A greyscale image represents the depth map of an image in which each pixel is assigned a value between and including 0 and 255.” and [0015], “Embodiments here relate to a method, apparatus, system, and computer program for modifying, enhancing or exaggerating 3D-image rendered given a mono-ocular (2D) image source and its depth map.” and Abstract, “A method for adjusting and generating enhanced 3D-effects for 2D to 3D image and video conversion applications includes controlling a depth location of a zero parallax plane within a depth field of an image scene to adjust parallax of objects in the image scene, controlling a depth volume of objects in the image scene to one of either exaggerate or reduce 3D-effect of the image scene …”)
30. Barkatullah doesn't explicitly disclose but Rowell discloses: filtering out a pixel whose depth value is less than a first threshold and a pixel whose depth value is greater than a second threshold from the original image, wherein the first threshold is less than the second threshold. (Rowell, [0187], “A depth filtering function is a third filtering function that may be implemented in the filtering module 1705. The depth filtering function excludes image data from image sections incorporating objects positioned a certain distance away from the stereo camera device. In one example, the depth filtering function rejects image data included in image sections containing close objects because the extreme horizontal and/or vertical shifts applied to image sections including close objects elsewhere in projection process interfere with re-calibration.” and [0189], “One non-limiting example depth filtering function determines a depth metric for each image section then filters image data in the image sections by comparing the depth metrics to a depth filtering threshold. The depth metric describes the distance between the stereo camera device and the objects included in an image section. Depth maps, point clouds, 3D scans, and distance measurements may all be used to generate depth metrics. One non-limiting example distance measurement is equivalent to 1/the horizontal shift (in pixels) applied to images captured by the stereo camera module. Depth filtering thresholds for evaluating depth metrics include 20 cm for camera modules having small zoom ranges and short focal lengths and 1 m for camera modules having moderate to large zoom ranges and average to long focal lengths.” and [0190], “In filtering routines including three filtering layers, image data may be required to pass three filtering thresholds to be incorporated into a disparity analysis. Alternatively, image data may only need to pass a majority or at least one of the three filtering thresholds (e.g., 2 of 3 or 1 of 3). In embodiments where image data is required to meet or exceed a subset of the filtering thresholds, the required filtering thresholds may be the same or different (e.g., image sections must pass both the depth filtering function and the standard deviation filtering function; image sections must pass the correlation filtering function and at least one other filtering function; or image sections must pass and any two filtering functions).”; Examiner’s note: The filtering routines disclosed by Rowell would have the capability to filter out pixels for either being less than and/or greater certain defined thresholds.)
31. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 1 of Barkatullah to include the disclosure of filtering out pixels whose depth value is either less than or greater than defined thresholds from the image, of Rowell. The motivation for this modification could have been to assist in the visual removal of objects that are either too close or too far away in the stereo images. In doing so, this would help prevent awkward distractions and place more emphasis on the objects closer to the zero-disparity plane where a user is more likely to be focused.
32. As per claim 6, Barkatullah in view of Rowell discloses: The method according to claim 1, wherein before the generating the binocular image based on the first depth value and depth values of a plurality of pixels in the original image, the method further comprises: (Barkatullah, [0006], “A greyscale image represents the depth map of an image in which each pixel is assigned a value between and including 0 and 255.” and [0015], “Embodiments here relate to a method, apparatus, system, and computer program for modifying, enhancing or exaggerating 3D-image rendered given a mono-ocular (2D) image source and its depth map.” and Abstract, “A method for adjusting and generating enhanced 3D-effects for 2D to 3D image and video conversion applications includes controlling a depth location of a zero parallax plane within a depth field of an image scene to adjust parallax of objects in the image scene, controlling a depth volume of objects in the image scene to one of either exaggerate or reduce 3D-effect of the image scene …”)
performing filtering on the depth values of the plurality of pixels in the original image. (Rowell, [0187], “A depth filtering function is a third filtering function that may be implemented in the filtering module 1705. The depth filtering function excludes image data from image sections incorporating objects positioned a certain distance away from the stereo camera device. In one example, the depth filtering function rejects image data included in image sections containing close objects because the extreme horizontal and/or vertical shifts applied to image sections including close objects elsewhere in projection process interfere with re-calibration.” and [0189], “One non-limiting example depth filtering function determines a depth metric for each image section then filters image data in the image sections by comparing the depth metrics to a depth filtering threshold.”)
33. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 1 of Barkatullah to include the disclosure of performing filtering on the depth values of the plurality of pixels in the original image, of Rowell. The motivation for this modification could have been to help smooth out any potential image artifacts in the stereo imagery. In addition, it could also help to make sure any drastic changes in image disparity are dampened so that viewing the stereo imagery is more comfortable to the user.
34. Claim 17, which is similar in scope to dependent claim 5 and independent claim 12, is thus rejected under the same rationale as described above. The motivation for this modification is the same as claim 5.
35. Claim 18, which is similar in scope to dependent claim 6 and independent claim 12, is thus rejected under the same rationale as described above. The motivation for this modification is the same as claim 6.
36. Claims 7-8 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Barkatullah (US-2018/0139432-A1) in view of Liao et al. (US-2012/0113093-A1, hereinafter "Liao").
37. As per claim 7, Barkatullah discloses: The method according to claim 1, wherein after the generating the binocular image, the method further comprises: (See rejection for claim 1.)
38. Barkatullah doesn't explicitly disclose but Liao discloses: determining a void region in the binocular image; and (Liao, [0034], “To reliably fill the holes and produce a convincing, realistic result, the system should specifically target the image inpainting for a stereo setting, taking into account the available depth (disparity) information, the local geometric saliency, as well as intensity information.” and [0036], “Referring to FIG. 5, with different techniques more suitable for different types of content (low frequency and high frequency), where natural images typically exhibit both types of content, a hybrid framework facilitates the inpainting for both cases. The first pass fills holes in low-frequency image regions using a first filter, such as an adaptive linear interpolation, and afterwards the second pass processes high-frequency image regions using a second filter, such as a non-linear exemplar-based inpainting.” and [0037], “During the first pass for low frequency inpainting, the holes in a low-frequency image region in IS are typically surrounded by pixels that are similar in intensity values as well as in disparity values. On a single scanline, a hole is represented as a succession of empty pixels of length γ.” and [0039], “The second pass of inpainting addresses high-frequency content in the image, filling pixels that are left unprocessed (left empty) by the first pass in IS′, DS′, and ES′. In order to preserve high-frequency details, the technique may employ a non-linear exemplar-based approach to synthesize the remaining holes. Since an exemplar-based inpainting relies both on intensity similarity and disparity similarity in evaluating a matching cost, the technique may make an explicit assumption regarding occlusion. For example, the technique may assume all occlusions in the stereo image pair to be two-layer occlusions.”)
performing image gradient diffusion from an edge with a large depth value in the void region to an edge with a small depth value, to fill the void region. (Liao, [0023], “Since the technique performs disparity scaling, rather than constant offsetting, holes 320 are created in the synthesized view IS ... The view optimization 400 may employ a one or a two-pass, depth-assisted inpainting approach 410 to optimize the synthesis by filling in the holes, producing a complete 420 (i.e. no-holes) new view IP and a complete reference disparity map DP.” and [0034], “To reliably fill the holes and produce a convincing, realistic result, the system should specifically target the image inpainting for a stereo setting, taking into account the available depth (disparity) information, the local geometric saliency, as well as intensity information.” and [0021], “To summarize, based on the original stereo image pair, the technique preferably generates a “continuous range” (e.g., three or more) of interpolation as well as extrapolation of new view points with adjusted depth, according to a viewer's preference. By employing scaling of disparity values, rather than constant offsetting, this preserves the spatial variation of disparity (inverse depth) across the image, which in turn preserves the underlying scene geometry (depth). Also, by synthesizing a new view using local image intensity values, disparity values, and geometric saliency information, as opposed to relying only on intensity further refines the synthesis process by separately processing the low and high frequency content in the stereo pair. Further, the reduction of grid quantization artifacts in a local manner also preserves details and reduces blurring the synthesized virtual view.” and [0035], “The inpainting technique may be linear or non-linear. Such linear techniques may include, for example, scanline interpolation and extrapolation. Such linear techniques tend to be computationally efficient, and tend to be robust when intensity variations are low (low frequency) in an image region requiring inpainting. Such non-linear techniques tend to be more computationally complex, and may include for example, partial differential equations or utilize exemplars in filling holes.”)
39. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 1 of Barkatullah to include the disclosure of determining a void region in the binocular image and performing image gradient diffusion to fill the void region, of Liao. The motivation for this modification could have been to prevent annoying or distracting “gaps” (or void) within the stereo images. In a stereo application, these “gaps” may make a user uncomfortable or even nauseous, especially at the edges of objects of different disparities. By filling these holes, it helps preserve the stereo experience, removing the distraction, and making the watching experience more pleasant.
40. As per claim 8, Barkatullah in view of Liao discloses: The method according to claim 7, wherein after the void region is filled, the method further comprises:
performing filtering on the filled void region. (Liao, [0019]-[0020], “To provide in-painting of the holes in the synthesized image created as a result of scaled offsets should, in addition to using local intensity values, model the local disparity values and local geometric saliency. ... The technique may further eliminate grid quantization artifacts that are a result of the inherently discrete, or quantized, nature of disparity values, where a given pixel is only allowed to assume an integer lateral offset number. This may be done by adaptively filtering the synthesized virtual view in a small neighborhood where such artifacts are detected. By restricting the filtering to a small neighborhood, the technique reduces the artifacts while preserving the details in the synthesized view, without introducing excessive undesirable blurring.” and [0052], “With knowledge of the locations of artifacts, the technique proceeds to filter out the artifacts. ... The operation effectively low-pass filters a very small neighborhood, typically 4×1 or 3×1, centered on and overwriting the artifact pixel. By restricting the filtering to a small neighborhood, the system is able to eliminate the artifacts while preserving the details in the synthesized view, without introducing undesired blurring.” and Claim 14. “The method of claim 7 further comprising a filter to reduce grid quantization artifacts.” and Claim 15, “The method of claim 14 wherein adaptive filtering is used to reduce grid quantization artifacts.”)
41. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 7 of Barkatullah to include the disclosure of performing filtering on the filled void region, of Liao. The motivation for this modification could have been to help smooth out any strange image artifacts the image filling process created when removing the “gaps” (or void). This process would help making the filled regions more seamless with the stereo images and much less likely to be noticed or distracting.
42. Claim 19, which is similar in scope to dependent claim 7 and independent claim 12, is thus rejected under the same rationale as described above. The motivation for this modification is the same as claim 7.
43. Claim 20, which is similar in scope to dependent claims 7 and 8 and independent claim 12, is thus rejected under the same rationale as described above. The motivation for this modification is the same as claim 8.
44. Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Barkatullah (US-2018/0139432-A1) in view of Freeman et al. (US-2018/0160098-A1, hereinafter "Freeman").
45. As per claim 10, Barkatullah discloses: The method according to claim 9, (See rejection for claim 9.)
46. Barkatullah doesn't explicitly disclose but Freeman discloses: wherein the video frame comprises a video frame in a live stream, and the salient region comprises a facial region of a live streamer. (Freeman, Abstract, “A system, method and software for producing 3D effects in a video of a physical scene. The 3D effects can be observed when the video is viewed, either during a live stream or later when viewing the recorded video. A reference plane is defined. The reference plane has peripheral boundaries. A live event is viewed with stereoscopic video cameras.” and [0037], “Depending upon the 3D effect being created, the stereoscopic cameras 20L, 20R are oriented so that their focal points are on the reference plane 46 and/or their lines of sight intersect at the reference plane 46. That is, the two stereoscopic cameras 20L, 20R achieve zero parallax at the reference plane 46.” and [0053], “The video producer and/or 3D effects technician also selects a reference plane 46 within the physical scene. See Block 82. The video producer and/or 3D effects technician then selects objects in the view of the stereoscopic cameras 20L, 20R that will be identified in production as primary subjects 42, secondary subjects 60 and background subjects 62. See Block 84. Using the reference plane 46 and the selected subjects, the video producer and/or 3D effects technician can determine the boundaries for the video scene 40 being produced. See Block 86.” and [0057], “Referring to FIG. 14, a studio setting 102 is shown. The studio setting 102 has one of more sets of stereoscopic cameras 104 positioned to image a person or other real object within a known defined area. In the shown example, the person is a teacher 106 standing in front of a blackboard 108. In this scenario, the blackboard 108 can be selected as the reference plane.”; Examiner’s note: As disclosed in ¶ [0053], a video producer can select a reference plane within a physical scene. For the example of the teacher and the blackboard described in ¶ [0057], the producer has the ability to choose the reference plane. Instead of the blackboard, the producer can rather choose the teacher’s face as a reference plane.)
47. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 9 of Barkatullah to include the disclosure of a live stream of stereo video frames wherein the salient region comprises of a facial region of a live streamer, of Freeman. The motivation for this modification could have been to adapt the zero-disparity plane for a live stream where the plane is focused on a live streamer. In a stereo live stream application, a user would likely be focused on a live streamer. By adjusting the zero-disparity plane to the facial region of the streamer, the user may find watching the stream more comfortable and easier to focus on.
Conclusion
48. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. These are as follows: Curti et al. (EP-1353518-A1) and Guo et al. (CN-103384340-A) disclose methods to generate stereo/binocular imaging from a single, monocular image and a depth map. Zheng et al. (CN-109147027-A) discloses a method that comprises a segmentation of a two-dimensional monocular image to obtain a plurality of reference planes and performs three-dimensional reconstruction corresponding to the two-dimensional image and plurality of reference planes.
49. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
50. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW CLOTHIER whose telephone number is (571)272-4667. The examiner can normally be reached Mon-Fri 8:00am-4:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached at (571)272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MATTHEW CLOTHIER/Examiner, Art Unit 2614
/KENT W CHANG/Supervisory Patent Examiner, Art Unit 2614