Prosecution Insights
Last updated: October 02, 2026
Application No. 18/526,875

ONE-STAGE PROGRESSIVE DICHOTOMOUS SEGMENTATION

Non-Final OA §103§112
Filed
Dec 01, 2023
Priority
May 18, 2023 — provisional 63/467,570
Examiner
DRYDEN, EMMA ELIZABETH
Art Unit
2677
Tech Center
2600 — Communications
Assignee
Samsung Electronics Co., Ltd.
OA Round
3 (Non-Final)
68%
Grant Probability
Favorable
3-4
OA Rounds
2m
Est. Remaining
80%
With Interview

Examiner Intelligence

Grants 68% — above average
68%
Career Allowance Rate
19 granted / 28 resolved
+5.9% vs TC avg
Moderate +12% lift
Without
With
+12.5%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
18 currently pending
Career history
51
Total Applications
across all art units

Statute-Specific Performance

§101
9.7%
-30.3% vs TC avg
§103
59.3%
+19.3% vs TC avg
§102
12.9%
-27.1% vs TC avg
§112
13.3%
-26.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 28 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Applicant claims the benefit of US Provisional Application No. 63/467,570, filed May 18, 2023. Claims 1-20 have been afforded the benefit of this filing date. Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's RCE submission filed on 08/17/2026 has been entered. Response to Amendment The amendment filed 07/15/2026 has been entered. Claims 1-20 remain pending in the application. Response to Arguments Applicant’s arguments, see pg. 9-11 of the remarks filed 07/15/2026, have been considered but are moot because the new ground of rejection does not rely on any combination of references applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Interpretation Regarding claim 15 and dependent claims 16-17, the two multilayer perceptrons of the “one hamburger head” are interpreted to be MLPs that utilize no activation or only linear activation to perform a linear transform on the input data (consistent with the hamburger head embodiment of FIG 12 and para 71-74). The MLPs cannot utilize nonlinear activation. Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 1-20 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. The amendment to independent claims 1, 19, and 20 (“wherein the prediction image segments the foreground object from the scene…such that the foreground object remains in the prediction image and the scene is removed in the prediction image.”) now requires that the foreground object (which is ultimately segmented from the input image) is in the prediction image. In the specification, the output image of the foreground object is generated by combining the input image with an image representing a prediction segmentation map (Para 44: “The output image of the foreground object is obtained by performing a logical AND operation between the input image (for example, FIG. 3) and the output segmentation map (for example, FIG. 7).”). While the specification describes wherein the prediction image may be a prediction image of the foreground object (i.e., Figure 2), as described in the amended independent claims, the specification does not disclose wherein an output image of the foreground object is generated by combining the input image with an image such that the foreground object remains in the prediction image and the scene is removed in the prediction image. Furthermore, it is unclear what the difference is between 1) the output image of the foreground object wherein the foreground object is segmented from a scene and 2) the prediction image such that the foreground object remains and the scene is removed. The specification fails to disclose combining the input image with an image such that the foreground object is remaining (prediction image in the amended claims) to generate a segmentation of the same foreground object (output image). In view of the foregoing, the subject matter identified above was not described in the specification at the time the application was filed. Accordingly, claims 1, 19, and 20 fail to comply with the written description requirement set forth in 35 U.S.C. 112(a). Dependent claims 2-18 are similarly rejected due to their dependence on a rejected base claim. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Guo et al. (Guo, M. H., Lu, C. Z., Hou, Q., Liu, Z., Cheng, M. M., & Hu, S. M. (2022). Segnext: Rethinking convolutional attention design for semantic segmentation. Advances in neural information processing systems, 35, 1140-1156.), hereinafter Guo, in view of Lin et al. (CN Patent No. 111340186 A), hereinafter Lin, in further view of Masud (Masud, Umar. Background Removal using Semantic Segmentation, 10 Sept 2021, Github [online], [retrieved on 2026-08-21]. Retrieved from the Internet <URL: https://github.com/umar07/Background_Removal_Semantic_Segmentation>) and Rosebrock (Rosebrock, Adrian. Image Segmentation with Mask R-CNN, GrabCut, and OpenCV. PyImageSearch, 28 Sept 2020, Internet Archive [online], [retrieved on 2026-05-13]. Retrieved from the Internet <URL: https://pyimagesearch.com/2020/09/28/image-segmentation-with-mask-r-cnn-grabcut-and-opencv/>), hereinafter Rosebrock. Regarding claim 1, Guo teaches a method of segmenting a foreground object from a scene for image editing or augmented reality, wherein the scene is represented in an input image (Guo, pg. 1, abstract: “SegNeXt, a simple convolutional network architecture for semantic segmentation”; trained and validated on image datasets with foreground/background classes, see 1st para on pg. 6), the method comprising: obtaining a plurality of feature vectors using a feature extractor (Guo, convolutional encoder, pg. 4, FIG 2 caption: “We extract multi-scale features”; pg. 4, section 3.1: “multi-scale convolutional attention (MSCA) module”), wherein the feature extractor comprises a plurality of multi-scale convolutional attention blocks (Guo, hierarchical structure to extract features at each resolution stage, pg. 4, section 3.1: “For MSCAN, we adopt a common hierarchical structure, which contains four stages with decreasing spatial resolutions”); obtaining, based on the plurality of feature vectors and performing an operation using one or more hamburger heads (Guo, see the decoder in FIG 3c, pg. 5 below, emphasis added to underlined text), a prediction image (Guo, performs image segmentation, 2nd to last para on pg. 7: “decoder designs for segmentation”; see image outputs in FIG 4, pg. 8), wherein the prediction image segments the foreground object from the scene (Guo, see pg. 8, FIG 4 and foreground/background classes in 1st para on pg. 6). PNG media_image1.png 318 679 media_image1.png Greyscale Guo fails to explicitly teach wherein a first hamburger head of the one or more hamburger heads comprises, in sequence, a first multilayer perceptron (MLP), a matrix decomposition, and a second MLP (the hamburger cited in Guo teaches a matrix decomposition in between two linear transform operations, but MLPs are not explicitly disclosed). Guo also fails to explicitly teach generating an output image of the foreground object from the scene by combining the input image with the prediction image and such that the foreground object remains in the prediction image and the scene is removed in the prediction image. However, Lin teaches a hamburger head comprising, in sequence, a first multilayer perceptron (MLP), a matrix decomposition, and a second MLP (Lin, para 79: “The “hamburger” module used for illustrative purposes in this invention consists of two linear transformations and a matrix factorization model. As shown in Figure 1, two linear transformations, acting as the “lower bun” and “upper bun,” are placed before and after the “patty” of the tensor decomposition model or matrix decomposition model, respectively, forming a “hamburger module.””; regarding MLPs - see para 66 wherein the lower and upper breads are composed of micro neural network that are linear transforms). Thus, Lin discloses a hamburger head comprising two MLPs, or linear transforms (Lin, para 66), while Guo discloses a hamburger head with two linear transform operations. A person of ordinary skill in the art, before the effective filing date of the claimed invention, would have recognized that the hamburger head of Lin could have been substituted for the hamburger head of Guo because both perform a linear transform operation before and after a matrix decomposition. Furthermore, a person of ordinary skill in the art would have been able to carry out the substitution. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to substitute the hamburger head of Lin for that of Guo according to known methods to yield the predictable result of implementing a light attention mechanism for a computer vision task (Lin, pg. 12, para 56: “our proposed method, as a highly interpretable representation learning approach, outperforms commonly used attention mechanisms in computer vision with a lighter computational and parameter count”). Additionally, Guo teaches an output including semantically segmented objects (Guo, see pg. 8, FIG 4), but doesn’t explicitly teach a prediction image such that the foreground object remains in the prediction image and the scene is removed in the prediction image. However, Masud teaches a segmentation method including a prediction image such that the foreground object remains in the prediction image and the scene is removed in the prediction image (Masud, Final output generated from a semantically segmented image in the figure from pg. 2, attached below). Removing the background object(s) isolates a foreground object. Guo discloses a base method for segmenting objects in an image, but does not specify specific methods for a prediction image such that the foreground object remains in the prediction image and the scene is removed in the prediction image. Masud teaches a known technique of removing the background from a semantically segmented image to isolate a foreground object. A person having ordinary skill in the art, before the effective filing date of the claimed invention, could have applied the known technique, as taught by Masud, in the same way to the method of Guo and achieved predictable results of isolating a foreground object in an input image. PNG media_image2.png 248 850 media_image2.png Greyscale Lastly, Rosebrock teaches a method for generating an output image of a foreground object from a scene (Rosebrock, segmented output image of the horse, see FIG. 1 from pg. 2 attached below) by combining the input image (Rosebrock, top left input image) with a prediction image (Rosebrock, top right segmentation mask; pg. 2: “Using Mask R-CNN, we can automatically compute pixel-wise masks for objects in the image, allowing us to segment the foreground from the background. An example mask computed via Mask R-CNN can be seen in Figure 1 at the top of this section. On the top-left, we have an input image of a barn scene. Mask R-CNN has detected a horse and then automatically computed its corresponding segmentation mask (top-right). And on the bottom, we can see the results of applying the computed mask to the input image — notice how the horse has been automatically segmented.”). Guo in view of Masud teaches an image input to the segmentation model (Guo, trained and validated on image datasets, 1st para on pg. 6) and a prediction image (see combination above). Therefore, Guo discloses a base method for segmenting objects in an input image, but does not specify specific methods for generating an output image of a foreground object from a scene by combining the input image with a prediction image. Rosebrock teaches a known technique of combining images, one of which has the background removed (pixels set to a black value and/or value of 0), to segment a foreground object. A person having ordinary skill in the art, before the effective filing date of the claimed invention, could have applied the known technique, as taught by Rosebrock, in the same way to the method of Guo in view of Lin and Masud and achieved predictable results of generating an output image containing a segmented foreground object with the background removed. Setting the removed background values in the prediction image to black (or a binary value of 0), as taught by Rosebrock, and combining that image with the input image results in a segmented foreground object. PNG media_image3.png 714 787 media_image3.png Greyscale Regarding claim 19, Guo teaches an apparatus comprising: one or more memories; and one or more processors, wherein the one or more processors are configured to execute instructions stored in the one or more memories to perform the claimed operation steps (Guo, method is a computer task, pg. 9, Table 11 caption: “We test our method with a single RTX-3090 GPU and AMD EPYC 7543 32-core processor CPU”). All further claim limitations are met and rendered obvious by Guo in view of Lin, Masud, and Rosebrock because the method steps of claim 1 are the same as the apparatus steps of claim 19. Regarding claim 20, Guo teaches a non-transitory computer readable medium storing instructions to be executed by one or more processors, wherein the instructions are configured to cause the one or more processors to perform the claimed operation steps (Guo, method is a computer task, pg. 9, Table 11 caption: “We test our method with a single RTX-3090 GPU and AMD EPYC 7543 32-core processor CPU”). All further claim limitations are met and rendered obvious by Guo in view of Lin, Masud, and Rosebrock because the method steps of claim 1 are the same as the executed steps of claim 20. Claims 2-5 and 14-17 are rejected under 35 U.S.C. 103 as being unpatentable over Guo in view of Lin, Masud, and Rosebrock, in further view of Woo et al. (Woo, S., Park, J., Lee, J. Y., & Kweon, I. S. (2018). Cbam: Convolutional block attention module. In Proceedings of the European conference on computer vision (ECCV) (pp. 3-19).), hereinafter Woo. Regarding claim 2 (dependent on claim 1), Guo in view of Lin, Masud, and Rosebrock teaches wherein the feature extractor comprises a convolutional layer (Guo, pg. 4, section 3.1: “depth-wise convolution to aggregate local information”), but fails to teach a plurality of multi-scale convolutional attention modules. However, Woo teaches a convolutional model comprising a plurality multi-scale convolutional attention modules (Woo, pg. 1, abstract: “our module sequentially infers attention maps along two separate dimensions, channel and spatial”; see FIG 3 on pg. 7, attached below). It would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the use of a plurality of attention modules, taught by Woo, in each multi-scale convolutional block in the method of Guo in view of Lin, Masud, and Rosebrock in order to improve the performance of the convolutional model (Woo, pg. 14, section 5: “We apply attention based feature refinement with two distinctive modules, channel and spatial, and achieve considerable performance improvement while keeping the overhead small”). PNG media_image4.png 112 490 media_image4.png Greyscale Regarding claim 3 (dependent on claim 2), Guo in view of Lin, Masud, Rosebrock, and Woo teaches wherein the obtaining the plurality of feature vectors comprises obtaining the plurality of feature vectors corresponding to the plurality of multi-scale convolutional attention blocks (Guo, features are output at each resolution stage – 4, 8, 26, 32, pg. 4, section 3.1: “For MSCAN, we adopt a common hierarchical structure, which contains four stages with decreasing spatial resolutions”). Regarding claim 4 (dependent on claim 3), Guo in view of Lin, Masud, Rosebrock, and Woo teaches wherein the prediction image is a final prediction image (Guo, see image outputs in FIG 4, pg. 8 with segmentation results). Regarding claim 5 (dependent on claim 4), Guo in view of Lin, Masud, Rosebrock, and Woo teaches wherein the plurality of feature vectors comprises a first feature vector, a second feature vector, a third feature vector and a fourth feature vector (Guo, encoder output, pg. 4, last para: “four stages with decreasing spatial resolutions”; see also 4 stages in FIG 3c), the method further comprising: applying the first hamburger head to the third feature vector and the fourth feature vector to obtain a first intermediate vector (Guo, output of head in FIG 3c which is then input to MLP; pg. 5, 1st para in section 3.2: “We aggregate features from the last three stages and use a lightweight Hamburger [22] to further model the global context”, this includes a third and fourth feature vector of the four stages); and applying a third MLP to the first intermediate vector to obtain a first image (Guo, see MLP sequentially after head in FIG 3c; see image outputs in FIG 4, pg. 8 - first image is the prediction image, in the same way as claim 14 wherein there is one hamburger head). Regarding claim 14 (dependent on claim 5), Guo in view of Lin, Masud, Rosebrock, and Woo teaches wherein the one or more hamburger heads is only one hamburger head (Guo, see one hamburger head in pg. 5, 1st para of section 3.2 and FIG. 3c) and the first image is the prediction image (Guo, see image outputs in FIG 4, pg. 8). Regarding claim 15 (dependent on claim 14), Guo in view of Lin, Masud, Rosebrock, and Woo teaches wherein the performing the operation comprises in sequence a first linear transforming, a matrix decomposing, and a second linear transforming (Lin, pg. 5, para 11: “lower-level linear transformation Wl, a matrix decomposition model M, and an upper-level linear transformation Wu connected sequentially”). Regarding claim 16 (dependent on claim 15), Guo in view of Lin, Masud, Rosebrock, and Woo teaches wherein the matrix decomposing comprises decomposition into a product and a summand (Lin, pg. 21, para 79-80: “transformed representation X is decomposed into the sum of the product of a dictionary matrix D and a reconstruction coefficient matrix C and the residual matrix E. X = DC + E (Equation 3)”). Regarding claim 17 (dependent on claim 16), Guo in view of Lin, Masud, Rosebrock, and Woo teaches wherein the matrix decomposing further comprises discarding the summand, so that a noise is reduced in the final prediction image (Lin, pg. 21, para 81: “In Equation 3, the residual E is discarded as invalid information”; pg. 16, para 67: “tensor decomposition and matrix decomposition models can be used to recover observations without noise and missing values”). Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Guo in view of Lin, Masud, Rosebrock, Woo, and Fu et al. (CN Patent No. 114419449 A), hereinafter Fu. Regarding claim 6 (dependent on claim 5), Guo in view of Lin, Masud, Rosebrock, and Woo fails to teach wherein the first image is an initial prediction image; however, Fu teaches a similar segmentation method (Fu, pg. 2, para 1: “semantic segmentation of remote sensing images using self-attention multi-scale feature fusion”) wherein a first image is an initial prediction image (Fu, the prediction results from one scale, pg. 6, para 8: “The prediction results from the four scales are then fused to obtain the final remote sensing image semantic segmentation result”). It would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the initial and final prediction images, taught by Fu, with the method of Guo in view of Lin, Masud, Rosebrock, and Woo in order to improve segmentation results by combining the segmentation predictions from different feature scales (Fu, pg. 6, para 8: “This method effectively fuses remote sensing semantic features of different scales, improving segmentation performance”). Allowable Subject Matter Claims 7-13 and 18 are currently rejected under 35 U.S.C. 112(a) due to the written description rejections applied to independent claim 1, but would be allowable if rewritten 1) to overcome the outstanding rejections and 2) in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: regarding claim 7, the closest prior art, cited herein, fails to teach “applying a second hamburger head to the second feature vector and the first intermediate vector to obtain a second intermediate vector”. Analogous prior art teaching similar multi-scale decoder models, such as Fu cited above, also fails to teach the limitations of claim 7. These reasons similarly apply to dependent claims 8-13 and 18. In view of the foregoing, the prior art references alone or in reasonable combination are insufficient to teach the invention as a whole, as claimed in claim 7-13 and 18. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to EMMA E DRYDEN whose telephone number is (571)272-1179. The examiner can normally be reached M-F 9-5 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ANDREW BEE can be reached at (571) 270-5183. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /EMMA E DRYDEN/Examiner, Art Unit 2677 /ANDREW W BEE/Supervisory Patent Examiner, Art Unit 2677
Read full office action

Prosecution Timeline

Dec 01, 2023
Application Filed
Dec 04, 2025
Non-Final Rejection mailed — §103, §112
Apr 06, 2026
Response Filed
May 15, 2026
Final Rejection mailed — §103, §112
Jul 15, 2026
Response after Non-Final Action
Aug 17, 2026
Request for Continued Examination
Aug 18, 2026
Response after Non-Final Action
Aug 27, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743830
GENERATING DIGITAL MATERIALS FROM DIGITAL IMAGES USING A CONTROLLED DIFFUSION NEURAL NETWORK
3y 1m to grant Granted Sep 22, 2026
Patent 12731271
ACTIVE LEARNING SYSTEM AND METHOD
2y 7m to grant Granted Sep 08, 2026
Patent 12705722
Real Time Inconsistency Detection During Composite Material Manufacturing
2y 6m to grant Granted Aug 11, 2026
Patent 12664680
LOCALIZATION AND MAPPING BY A GROUP OF MOBILE COMMUNICATIONS DEVICES
3y 8m to grant Granted Jun 23, 2026
Patent 12632966
METHOD, ELECTRONIC DEVICE, AND COMPUTER PROGRAM PRODUCT FOR RECOGNIZING OBJECT REGIONS IN IMAGE
2y 11m to grant Granted May 19, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
68%
Grant Probability
80%
With Interview (+12.5%)
3y 1m (~2m remaining)
Median Time to Grant
High
PTA Risk
Based on 28 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month