DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 01/30/2025 and 12/09/2025 have been considered by the examiner.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Shen et al. (“High Quality Segmentation for Ultra High-resolution Images”) in view of Kirillov et al. (“PointRend: Image Segmentation as Rendering”).
Regarding claim 1, Shen discloses a computer-implemented method comprising: by at least one processor, (See Shen p. 5 left column 5th para, “In addition, high-resolution training is directly limited by the constraint of GPU memory and batch size.”)
generating, simulated masks for objects in a plurality of digital images by modifying masked regions in a plurality of ground-truth masks for the objects utilizing one or more mask modification operations; (See Shen p. 5 left column 6th para, “Mcoarse is generated by morphological perturbation of the provided ground truth mask Mgt.”)
generating, by the at least one processor utilizing a mask refinement neural network, a plurality of estimated refined masks for the objects in the plurality of digital images based on the plurality of digital images and the simulated masks; (See Shen p. 3 Fig. 3, where the Icoarse input mask and the original digital image are input to the CRM (Continuous Refinement Model) which is a neural network. The output of the CRM is a refined image Mrefined.
This is done for the plurality of training images as disclosed by Shen p. 5 left column 5th para, “It has 2K images as ground truth and generates any low-resolution images as input.”)
and adjusting, by the at least one processor, parameters of the mask refinement neural network by utilizing a matting loss (See Shen p. 5 left column 6th para, “We design the training loss in a simple way on the final prediction Mrefined without different loss functions on different resolution stages [9]. Our loss term L (Ө, ϕ) is calculated on the refinement target as L (Ө, ϕ) = Ʃ i=1…4 wi, Li (Mrefined, Mgt) (7) where Li; i ϵ 2 [1; 2; 3; 4] denote cross-entropy loss, L1 loss, L2 loss, and gradient loss.”)
Shen discloses the above limitations, but he fails to disclose, based on a plurality of separate point-sampling operations to reduce differences between the plurality of estimated refined masks and the plurality of ground-truth masks.
However, Kirillov discloses, based on a plurality of separate point-sampling operations to reduce difference between the plurality of estimated refined masks and the plurality of ground-truth masks. (See Kirillov p. 4 right column 3rd para, “The sampling strategy selects N points on a feature map to train on.1 It is designed to bias selection towards uncertain regions, while also retaining some degree of uniform coverage, using three principles. (i) Over generation: we over-generate candidate points by randomly sampling kN points (k>1) from a uniform distribution. (ii) Importance sampling: we focus on points with uncertain coarse predictions by interpolating the coarse prediction values at all kN points and computing a task-specific uncertainty estimate (defined in x4 and x5). The most uncertain βN points (β ϵ [0; 1]) are selected from the kN candidates. (iii) Coverage: the remaining (1-β) N points are sampled from a uniform distribution. At training time, predictions and loss functions are only computed on the N sampled points.”)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include the computing a loss by sampling the refined mask and the ground truth mask as suggested by Kirilov to Shen’s computing a loss for a refinement neural network. This can be done using known engineering techniques, with a reasonable expectation of success. The motivation for doing so is because processing high resolution images requires immense GPU memory. Sampling a subset of points keeps memory consumption manageable.
Regarding claim 16, Shen discloses, a non-transitory computer readable medium storing instructions thereon that, when executed by at least one processor, cause the at least one processor to perform operations comprising: (See Shen p. 5 left column 5th para, “In addition, high-resolution training is directly limited by the constraint of GPU memory and batch size.” Where the memory is a computer readable medium that will inherently store instructions, and GPU processor will inherently perform operations for the image processing.)
generating simulated masks for objects in a plurality of digital images by modifying masked regions in a plurality of ground-truth masks for the objects utilizing one or more mask modification operations; (See Shen p. 5 left column 6th para, “Mcoarse is generated by morphological perturbation of the provided ground truth mask Mgt.”)
generating, utilizing a mask refinement neural network, a plurality of estimated refined masks for the objects in the plurality of digital images based on the plurality of digital images and the simulated masks; (See Shen p. 3 Fig. 3, where the Icoarse input mask and the original digital image are input to the CRM (Continuous Refinement Model) which is a neural network. The output of the CRM is a refined image Mrefined.
This is done for the plurality of training images as disclosed by Shen p. 5 left column 5th para, “It has 2K images as ground truth and generates any low-resolution images as input.”)
determining a matting loss indicating differences between the plurality of estimated refined masks and the plurality of ground-truth masks and adjusting parameters of the mask refinement neural network by utilizing the matting loss to reduce the differences between the plurality of estimated refined masks and the plurality of ground-truth masks. (See Shen p. 5 left column 6th para, “We design the training loss in a simple way on the final prediction Mrefined without different loss functions on different resolution stages [9]. Our loss term L (Ө, ϕ) is calculated on the refinement target as L (Ө, ϕ) = Ʃ i=1…4 wi, Li (Mrefined, Mgt) (7) where Li; i ϵ 2 [1; 2; 3; 4] denote cross-entropy loss, L1 loss, L2 loss, and gradient loss.”)
Shen discloses the above limitations, but he fails to disclose determining a matting loss indicating differences between the plurality of estimated refined masks and the plurality of ground-truth masks based on comparison pixels sampled via a plurality of separate point-sampling operations.
However, Kirillov discloses, determining a matting loss indicating differences between the plurality of estimated refined masks and the plurality of ground-truth masks based on comparison pixels sampled via a plurality of separate point-sampling operations; (See Krillov p. 4 right column 3rd para, “The sampling strategy selects N points on a feature map to train on.1 It is designed to bias selection towards uncertain regions, while also retaining some degree of uniform coverage, using three principles. (i) Over generation: we over-generate candidate points by randomly sampling kN points (k>1) from a uniform distribution. (ii) Importance sampling: we focus on points with uncertain coarse predictions by interpolating the coarse prediction values at all kN points and computing a task-specific uncertainty estimate (defined in x4 and x5). The most uncertain βN points (β ϵ [0; 1]) are selected from the kN candidates. (iii) Coverage: the remaining (1-β) N points are sampled from a uniform distribution. At training time, predictions and loss functions are only computed on the N sampled points.”)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include the computing a loss by sampling the refined mask and the ground truth mask as suggested by Kirilov to Shen’s computing a loss for a refinement neural network. This can be done using known engineering techniques, with a reasonable expectation of success. The motivation for doing so is because processing high resolution images requires immense GPU memory. Sampling a subset of points keeps memory consumption manageable.
Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Shen et al. (“High Quality Segmentation for Ultra High-resolution Images”) in view of Kirillov et al. (“PointRend: Image Segmentation as Rendering”) and in further view of Gazit (US Pub. No. 2016/0300351 A1).
Regarding claim 2, Shen and Krillov disclose the computer-implemented method of claim 1, but they fail to disclose, wherein generating the simulated masks comprises: detecting one or more holes in a masked region of a ground-truth mask of the plurality of ground-truth masks; and generating a simulated mask by synthetically filling the one or more holes in the masked region
However, Gazit discloses, wherein generating the simulated masks comprises: detecting one or more holes in a masked region of a ground-truth mask of the plurality of ground-truth masks; and generating a simulated mask by synthetically filling the one or more holes in the masked region. (See Gazit ¶177, “At 820, the mask generated in 818 is optionally morphologically improved, for example filling in holes in its interior that are completely surrounded by mask voxels or filling in holes that are completely surrounded by mask voxels … 4) Filling in any cavities that are completely surrounded by mask voxels in their axial slice. … The resulting mask, after the morphological improvements, may be considered a rough segmentation of the target organ, in this image.”)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include the generating a rough segmentation mask by filling in holes as suggested by Gazit to Shen and Kirillov’s generation of a coarse mask. This can be done using known engineering techniques, with a reasonable expectation of success. The motivation for doing so is because filling in holes in mask removes high frequency interval variations so the refinement network can focus strictly on correcting boundaries and edges.
Allowable Subject Matter
Claims 3-9 and 17-20 are objected to as being dependent upon a rejected base claim but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding claim 3, the computer-implemented method of claim 2, wherein detecting the one or more holes comprises: determining a size ratio indicating a size of a hole in the masked region relative to a size of the masked region; and selecting the hole for synthetically filling in response to determining that the size ratio is lower than a size ratio threshold. (The disclosed prior art of record fails to disclose the limitations of this claim.)
Regarding claim 4, the computer-implemented method of claim 1, but they fail to disclose, wherein generating the simulated masks comprises: generating downscaled masks by downscaling a subset of ground-truth masks from one or more initial sizes to a plurality of randomly selected sizes; and generating a subset of simulated masks by upscaling the downscaled masks from the plurality of randomly selected sizes to the one or more initial sizes. (The disclosed prior art of record fails to disclose the limitations of this claim.)
Regarding claim 5, the computer-implemented method of claim 1, wherein generating the plurality of estimated refined masks comprises: generating, utilizing a coarse mask generation neural network, coarse masks for objects in a plurality of additional digital images; determining a training dataset comprising the simulated masks and the coarse masks; and generating the plurality of estimated refined masks based on the training dataset comprising the simulated masks and the coarse masks. (The disclosed prior art of record fails to disclose the limitations of this claim.)
Regarding claim 6, the computer-implemented method of claim 1, wherein adjusting the parameters of the mask refinement neural network comprises determining the matting loss by: sampling a first comparison pixel in a first estimated refined mask utilizing a first point-sampling operation of the plurality of separate point-sampling operations; and sampling a second comparison pixel in a second estimated refined mask utilizing a second point-sampling operation of the plurality of separate point-sampling operations. (The disclosed prior art of record fails to disclose the limitations of this claim.)
Regarding claims 7-8, these claims are objected to, since they depend on objected to claim 6.
Regarding claim 9, the computer-implemented method of claim 1, wherein adjusting the parameters of the mask refinement neural network comprises: determining, for a digital image, that a masked region of a simulated mask is within a threshold distance of a boundary of the simulated mask; generating a padded mask by inserting a boundary padding at the boundary of the simulated mask in response to determining that the masked region is within the threshold distance; and adjusting the parameters of the mask refinement neural network based on the padded mask. (The disclosed prior art of record fails to disclose the limitations of this claim.)
Regarding claim 17, the non-transitory computer readable medium of claim 16, wherein generating the simulated masks comprises: detecting, in masked portions of the plurality of ground-truth masks, holes that meet a size ratio threshold; and generating simulated masks by synthetically filling the holes in the masked portions. (The disclosed prior art of record fails to disclose the limitations of this claim.)
Regarding claim 18, the non-transitory computer readable medium of claim 16, wherein the operations further comprise: generating, utilizing a coarse mask generation neural network, coarse masks for a plurality of additional digital images; and generating a training dataset comprising the coarse masks and the simulated masks, wherein determining the matting loss comprises determining the matting loss based on the training dataset comprising the coarse masks and the simulated masks. (The disclosed prior art of record fails to disclose the limitations of this claim.)
Regarding claim 19, the non-transitory computer readable medium of claim 16, wherein determining the matting loss comprises: sampling comparison pixels from the plurality of estimated refined masks utilizing the plurality of separate point-sampling operations; determining, for the comparison pixels, corresponding pixels of the plurality of ground-truth masks; and determining the matting loss by determining differences between the comparison pixels and the corresponding pixels. (The disclosed prior art of record fails to disclose the limitations of this claim.)
Regarding claim 20, this claim is objected to since it depends on objected to claim 19.
Claims 10-15 are allowed.
The following is an examiner’s statement of reasons for allowance:
The closest found prior art references are Shen et al. (“High Quality Segmentation for Ultra High-resolution Images”) and Kirillov et al. (“PointRend: Image Segmentation as Rendering”). These references disclosed some, but not all of the limitations of claim 10 as follows:
Regarding claim 10, a system comprising: one or more memory devices comprising a plurality of digital images; and one or more processors coupled to the one or more memory devices that cause the system to perform operations comprising: (See Shen p. 5 left column 5th para, “In addition, high-resolution training is directly limited by the constraint of GPU memory and batch size.”)
generating simulated masks for objects in a plurality of digital images by modifying masked regions in a plurality of ground-truth masks for the objects utilizing one or more mask modification operations; (See Shen p. 5 left column 6th para, “Mcoarse is generated by morphological perturbation of the provided ground truth mask Mgt.”)
generating, utilizing a coarse mask generation neural network, coarse masks for objects in the plurality of digital images; (The disclosed prior art of record fails to disclose this limitation of the claim in combination with the rest of the claim.)
generating, utilizing a mask refinement neural network, a plurality of estimated refined masks for the objects in the plurality of digital images based on the plurality of digital images and a set of masks comprising the simulated masks and the coarse masks; (See Shen p. 3 Fig. 3, where the Icoarse input mask and the original digital image are input to the CRM (Continuous Refinement Model) which is a neural network. The output of the CRM is a refined image Mrefined.
This is done for the plurality of training images as disclosed by Shen p. 5 left column 5th para, “It has 2K images as ground truth and generates any low-resolution images as input.”)
and adjusting parameters of the mask refinement neural network by utilizing a matting loss (See Shen p. 5 left column 6th para, “We design the training loss in a simple way on the final prediction Mrefined without different loss functions on different resolution stages [9]. Our loss term L (Ө, ϕ) is calculated on the refinement target as L (Ө, ϕ) = Ʃ i=1…4 wi, Li (Mrefined, Mgt) (7) where Li; i ϵ 2 [1; 2; 3; 4] denote cross-entropy loss, L1 loss, L2 loss, and gradient loss.”)
based on randomly selected point-sampling operations to reduce differences between the plurality of estimated refined masks and the plurality of ground-truth masks. (See Kirillov p. 4 right column 3rd para, “The sampling strategy selects N points on a feature map to train on.1 It is designed to bias selection towards uncertain regions, while also retaining some degree of uniform coverage, using three principles. (i) Over generation: we over-generate candidate points by randomly sampling kN points (k>1) from a uniform distribution. (ii) Importance sampling: we focus on points with uncertain coarse predictions by interpolating the coarse prediction values at all kN points and computing a task-specific uncertainty estimate (defined in x4 and x5). The most uncertain βN points (β ϵ [0; 1]) are selected from the kN candidates. (iii) Coverage: the remaining (1-β) N points are sampled from a uniform distribution. At training time, predictions and loss functions are only computed on the N sampled points.”
However, Kirillov fails to disclose “randomly selecting point-sampling operations.”)
Regarding claims 11-15, these claims are allowed since they depend on allowed independent claim 10.
Any comments considered necessary by applicant must be submitted no later than the payment of the issue fee and, to avoid processing delays, should preferably accompany the issue fee. Such submissions should be clearly labeled “Comments on Statement of Reasons for Allowance.”
Conclusion
Listed below are the prior arts made of record and not relied upon but are considered pertinent to applicant’s disclosure.
Brada et al. (US Pub. No. 2020/0294239 A1) The example embodiments are directed to refinement process for generating an accurate image segmentation map. A refinement network may enhance an initially generated segmentation map using a model that is trained using synthetic images. In one example, the method may include storing an image of content which includes a plurality of categories of data, receiving an initial segmentation map of the image, the initial segmentation map comprising pixel probability values with respect to the plurality of categories, executing a refinement predictive model on the initial segmentation map and the image to generate a refined segmentation map, wherein the predictive model is trained using synthetic images of the plurality of categories of data, and generating a segmented image based on the refined segmentation map.
Kohli et al. (US Pub. No. 2012/0288186 A1) An enhanced training sample set containing new synthesized training images that are artificially generated from an original training sample set is provided to satisfactorily increase the accuracy of an object recognition system. The original sample set is artificially augmented by introducing one or more variations to the original images with little to no human input. There are a large number of possible variations that can be introduced to the original images, such as varying the image's position, orientation, and/or appearance and varying an object's context, scale, and/or rotation.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAVID PERLMAN whose telephone number is (571) 270-1417.
The examiner can normally be reached on Monday - Friday; 10:00am -6:30pm.
Examiner interviews are available via telephone and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is
(571) 273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at (866) 217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call (800) 786-9199 (IN USA OR CANADA) or (571) 272-1000.
/DAVID PERLMAN/Primary Examiner, Art Unit 2673