DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Election/Restrictions
Applicant’s election without traverse of Invention I, Species I.A (claims 1–5, and 14–15) in the reply filed on July 13th 2026 is acknowledged. The application has pending claims 1-15 (withdrawn claims 6-13 are withdrawn from further consideration), Claims 16-20 have been cancelled.
Specification
The abstract of the disclosure is objected to because: there is one minor editorial issue in the first sentence: “the seamless transition of subject into an image or scene”; for correct grammar, this could read: “the seamless transition of a subject into an image or scene”. A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b).
The disclosure is objected to because of the following informalities: in [0006], the phrase “non-transitory program storage device” is not a usual term. The phrase should be revised to read “non-transitory computer readable storage medium”.
Appropriate correction is required.
The disclosure is objected to because of the following informalities: In paragraph [0054], the phrase “one more non-transitory storage mediums” is grammatically incorrect. The phrase should be revised to read “one or more non-transitory computer readable storage mediums”.
Appropriate correction is required.
Claim Rejections - 35 USC § 112(b)
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 4 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding claim 4, the phrase “the image segmentation operation identifies persons or other objects of interest in an image” renders the claim indefinite because: Claim 1 explicitly recites "a first image", it is unclear whether claim 4 refers to the same first image processed throughout claim 1 or a different, unclaimed image. Consequently, the relationship between the image recited in claim 4 and the first image recited in claim 1 is unclear. Applicant may clarify the claim by amending “an image” to “the first image” if that is the intended meaning.
Claim Rejections - 35 USC § 112(d)
The following is a quotation of 35 U.S.C. 112(d):
(d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
The following is a quotation of pre-AIA 35 U.S.C. 112, fourth paragraph:
Subject to the following paragraph [i.e., the fifth paragraph of pre-AIA 35 U.S.C. 112], a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
Claim 5 is rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends.
Claim 1 already recites “applying the second alpha mask modifies an opacity of at least some portions of the first image corresponding to the location of the first subject”. Claim 5 merely recites that “the second alpha mask modifies a transparency level of portions of the first image”. The specification expressly describes transparency and opacity as alternative expressions of the same alpha-mask property. See paragraph [0042], “transparency (or opacity)”. Accordingly, claim 5 restates the opacity-modification limitation of claim 1 using equivalent terminology and does not specify a further limitation of the subject matter of claim 1.
Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1–2, 4–5 and 14 are rejected under 35 U.S.C. §102(a)(1) as being anticipated by Zhou (Zhou et al, US 2021/0019892 A1, 2019).
Regarding claim 1, Zhou teaches a non-transitory computer readable storage medium storing program instructions that, when executed by a processing device ( [0031], [0044], [0099], [Fig. 9]: Zhou expressly teaches a non-transitory computer-readable medium having instructions/ software applications stored thereon that, when executed by one or more hardware processors, cause the hardware processors to perform the disclosed segmentation operations. ), cause the processing device to:
obtain a first image of a scene, the first image comprising at least a first subject;
( [0005], [0047], [0058], [0094], [Fig. 7]: Zhou teaches receiving a plurality of video frames, each frame comprising color data and depth data for a plurality of pixels. A selected video frame constitutes the claimed first image. Zhou further teaches that the video frame may depict one or more persons and that a head bounding box identifies pixels corresponding to the head of a person, thereby teaching an image comprising a first subject. )
generate a first alpha mask for the first image, wherein the first alpha mask is generated based on an image segmentation operation, and wherein the image segmentation operation identifies a location of the first subject within the first image;
( [0005], [0058], [Fig. 2, Step 214], [Claim 9]: Zhou teaches generating an initial segmentation mask for the selected frame that categorizes each pixel as foreground or background. The initial segmentation mask includes a respective mask value for each pixel, such as 255 for a foreground pixel and 0 for a background pixel, and therefore generating an alpha mask. This initial segmentation mask is a first alpha-type mask generated by segmentation, and it identifies the foreground subject region (e.g., the head, hair, neck, shoulders) within the image. Zhou further teaches detecting a head region or multiple person regions using the initial segmentation mask, thereby identifying the location of the subject within the frame. )
generate a depth map for the first image;
( [0005], [0047], [0052], [0118]: Zhou teaches that each received frame includes depth data for a plurality of pixels and that the processor directly downsamples the depth data while preserving the bit depth; the initial segmentation mask is generated by comparing each pixel’s depth value to a depth range, thereby using per-pixel depth information. The spatially corresponding per-pixel depth data for the selected frame constitutes a depth map for the image. )
determine a foreground depth for the scene;
( [0006 & Claim 2], [0054]: Zhou teaches setting a depth range (e.g., 0.5–1.5 m) and categorizing pixels whose depth values fall within that range as foreground pixels and those outside as background pixels. This effectively determines a foreground depth region for the scene; i.e., a depth interval that defines the foreground subject. )
generate a second alpha mask for the first image, wherein the second alpha mask is generated by modifying the first alpha mask based, at least in part, on comparisons between values in corresponding portions of the depth map and the determined foreground depth;
( [0015-0016, Claim 18], [0082], [0088]: Zhou first uses comparisons of each pixel’s depth value to the selected depth range to generate the initial segmentation mask (first alpha mask) as said above. Zhou then generates a trimap based on the initial segmentation mask, performs fine segmentation using the trimap to obtain a further binary mask, and applies a Gaussian filter to the binary mask to provide alpha matting. Accordingly, the Gaussian-filtered alpha-matted mask is a second alpha mask produced by refining or modifying the first mask and remains based at least in part on the original per-pixel depth comparisons because the depth-generated initial mask is used to derive the trimap and subsequent binary mask. )
apply the second alpha mask to the first image to create a final image,
( [0088], [0090–0092]: Zhou teaches using the Gaussian-filtered alpha-matted binary mask to obtain a foreground mask for the frame, identify the image pixels included in the foreground video, and render a resulting frame containing the segmented foreground, optionally together with a blank or replacement background. The rendered foreground or composite frame constitutes the claimed final image. )
wherein applying the second alpha mask modifies an opacity of at least some portions of the first image corresponding to the location of the first subject; and
( [0088], [0094]: Zhou teaches applying a Gaussian filter to the binary mask to provide alpha matting and smooth the segmentation boundary for hairy or fuzzy portions of the foreground subject. The alpha-matted mask is then used to include the foreground-subject pixels while excluding or replacing background pixels, effectively describing how the mask cuts out or blends an image [modifying opacities]. Accordingly, alpha matting modifies the opacity or transparency corresponding to the segmented subject, including portions corresponding to the subject’s hair, head, and neck. )
display the final image.
( [0091–0092], [0098], [Fig. 1]: Zhou states that the foreground video is rendered and displayed, including in real-time, and that the foreground can be shown with a blank background or with a substituted background scene. )
Regarding claim 2, Zhou teaches the non-transitory computer readable storage medium of claim 1, wherein the first alpha mask comprises a plurality of segmentation values, wherein each segmentation value corresponds to a pixel in the first image.
( [0005], [0009, Claim 9], [0054], [Fig. 5]: Zhou teaches generating an initial segmentation mask that categorizes each pixel of the selected frame as a foreground pixel or a background pixel. The initial segmentation mask includes a respective mask value for each pixel, including a value of 255 for a foreground pixel and a value of 0 for a background pixel, thereby teaching a plurality of segmentation values respectively corresponding to pixels of the first image. )
Regarding claim 4, Zhou teaches the non-transitory computer readable storage medium of claim 1, wherein the image segmentation operation identifies persons or other objects of interest in an image.
( [0058], [0073], [0075], [0088], [0094]: Zhou teaches detecting a head bounding box that specifies pixels corresponding to the head of a person, including respective image regions corresponding to multiple persons. Zhou further teaches identifying skin regions corresponding to a hand, an arm, or other portions of the body, and separating hairy or fuzzy foreground objects from the background. Accordingly, Zhou’s segmentation operation identifies persons or other foreground objects of interest in the image. )
Regarding claim 5, Zhou teaches the non-transitory computer readable storage medium of claim 1, wherein the second alpha mask modifies a transparency level of portions of the first image.
( [0088], [0090–0092], [0131–0132]: Zhou teaches applying a Gaussian filter to the binary mask to provide alpha matting and smooth the segmentation boundary, including boundaries of hairy or fuzzy foreground objects. Zhou further teaches applying the resulting mask to obtain and render a foreground image with the original background removed or replaced. Thereby, the alpha-matted boundary values modify transparency or opacity levels of portions of the image corresponding to the foreground subject; by using these non-binary alpha values to blend foreground pixels with background pixels, the system effectively modifies the transparency level of portions of the first image (especially near object boundaries). )
Regarding claim 14, Zhou teaches the non-transitory computer readable storage medium of claim 1, wherein the foreground depth is determined based on one of the following: an estimated depth of the first subject; an estimated depth of a second subject identified by the image segmentation operation; a predetermined value; a focus setting of a camera; or a user- controllable value.
( [0054]: Zhou teaches determining a foreground-depth criterion using predetermined depth values defining a foreground depth range, such as 0.5 meters to 1.5 meters. Zhou compares each pixel depth value with the predetermined depth range and categorizes the pixel as foreground when its depth is within the range. This depth range corresponds to a foreground depth region and is selected based on expected subject distance from the camera in video conferencing scenarios, effectively an estimated depth of the subject or a predetermined range reflecting camera usage and setup. )
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 3 and 15 are rejected under 35 U.S.C. §103 as being unpatentable over Zhou in view of Lindskog (Lindskog et al, US 2020/0082535 A1, 2020), as provided by Applicant’s disclosure filed on Nov 14th, 2024.
Regarding claim 3, Zhou teaches the non-transitory computer readable storage medium of claim 1,
Zhou teaches generating and using a first alpha mask in processing an image, but fails to expressly disclose obtaining the first alpha mask as an output from a neural network, where Lindskog teaches:
wherein the first alpha mask is obtained as an output from a neural network.
( [0008], [0016], [0032], [0044]: Lindskog teaches using an alpha matte having alpha values that control an amount of blending between a base or background layer and a segmented overlay layer. Lindskog teaches that segmentation masks and corresponding confidence masks “may be obtained as an output of a neural network, e.g., a Convolutional Neural Network (CNN)”. )
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to obtain Zhou’s first alpha mask using Lindskog’s neural-network-based segmentation technique because neural networks were known to generate pixel-level masks for image segmentation, thereby providing accurate mask values for blending image layers with predictable results.
Regarding claim 15, Zhou teaches the non-transitory computer readable storage medium of claim 1,
Zhou teaches generating and using a depth map in processing an image, but fails to expressly disclose generating the depth map using a monocular depth neural network, stereo-camera depth information, a time-of-flight camera, structured-light sensors, or phase-detection pixels, where Lindskog teaches:
wherein the depth map is generated using one of the following: a monocular depth neural network, stereo camera depth information, a time of flight camera, structured light sensors, or phase detection pixels.
( [0011], [0042], [0065]: Lindskog teaches obtaining initial depth or disparity information using a secondary stereo camera, focus pixels, and other depth or disparity sensors. Lindskog further teaches generating an initial depth/ disparity map using a secondary stereo camera, structured-light sensors, or focus pixels, and expressly explains that the focus pixels are pixels used for phase-detection autofocus. )
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to generate Zhou’s depth map using Lindskog’s stereo-camera, structured-light, or phase-detection techniques because these were known alternative depth-sensing modalities suitable for providing the depth information required by Zhou’s image-processing system, with the selection of a particular modality yielding predictable results according to implementation requirements.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEN KUDO whose telephone number is (571)272-4498. The examiner can normally be reached M-F 8am - 5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent Rudolph can be reached at 571-272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
KEN KUDO
Examiner
Art Unit 2671
/KEN KUDO/Examiner, Art Unit 2671
/VINCENT RUDOLPH/Supervisory Patent Examiner, Art Unit 2671