Prosecution Insights
Last updated: August 17, 2026
Application No. 18/941,579

GENERATING MULTIPLE SEGMENTATION MASKS IN A SINGLE MODEL WITH MULTI-TASK QUERY DECODERS

Non-Final OA §103§112
Filed
Nov 08, 2024
Examiner
LEMIEUX, IAN L
Art Unit
2669
Tech Center
2600 — Communications
Assignee
Adobe Inc.
OA Round
1 (Non-Final)
87%
Grant Probability
Favorable
1-2
OA Rounds
4m
Est. Remaining
96%
With Interview

Examiner Intelligence

Grants 87% — above average
87%
Career Allowance Rate
510 granted / 587 resolved
+24.9% vs TC avg
Moderate +9% lift
Without
With
+8.9%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 2m
Avg Prosecution
19 currently pending
Career history
610
Total Applications
across all art units

Statute-Specific Performance

§101
11.2%
-28.8% vs TC avg
§103
42.7%
+2.7% vs TC avg
§102
17.8%
-22.2% vs TC avg
§112
21.9%
-18.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 587 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are currently pending in U.S. Patent Application No. 18/941,579 and an Office action on the merits follows. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim(s) 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim(s) 1/10/17, at line 6 in that second ‘generating’ step, recite(s) the limitation in part “utilizing a plurality of query decoder neural networks in connection with a plurality of segmentation tasks”. Given the association between a ‘query’ and ‘a plurality of segmentation tasks’, (Applicant’s Specification at e.g. [0080] “person wearing blue shirt”), it is not clear if and how (to what extent), each of the ‘plurality’ of query decoders necessarily differ structurally, and accordingly if they may in fact be realized by a shared/common/multi-layered architecture (i.e. not decoupled and distinct primarily in terms of the query received and/or impacts of a non-recited training). In other words, under at least one apparently permissible interpretation, the query decoders in question may be realized by a ‘single’/ shared/common/multi-layered decoder architecture, distinct primarily in terms of the query received (which in turn is associated with a different/query-specific task e.g. instance segmentation, “leftmost elephant” vs “small elephant in front”; Applicant’s specification/claims does not appear to limit tasks to any that are e.g. instance vs. semantic vs. panoptic). An alternative interpretation (further giving rise to interpretation ambiguity – see MPEP 2173.02), although less of an issue, may also be one where the query decoders differ primarily in terms of an associated scale/resolution. With reference to MPEP 2173.02 sub-section II, the threshold requirements for clarity and precision (analyzed not in a vacuum but in light of A-C) may depend on a consideration of (B) the teachings of the prior art. At least the following potential prior art references serve to illustrate query decoder architecture that gives rise to the question(s) above regarding the decoder architecture as recited/claimed. Jain et al. “OneFormer: One Transformer to Rule Universal Image Segmentation”, Fig. 2 reproduced below (note x L at Transformer Decoder): PNG media_image1.png 546 1144 media_image1.png Greyscale Shah et al. “LQMFormer: Language-aware Query Mask Transformer for Referring Image Segmentation”, Fig. 2 (note L x) reproduced below: PNG media_image2.png 408 768 media_image2.png Greyscale Li et al. “OMG-Seg: Is One Model Good Enough for All Segmentation?”, Fig 2 reproduced below: PNG media_image3.png 574 1100 media_image3.png Greyscale With reference to Li for example (OMG-Seg), Li discloses two types of mask queries, semantic (Li Fig. 2, see also Applicant’s Specification at [0043]), and location (Li Fig. 2, see also Applicant’s Specification at [0079]). Jain’s OneFormer Fig. 2b illustrates semantic, instance and panoptic task-oriented object queries. The references identified above each/all appear to disclose a “set of queries” comparable to Applicant’s 506, and transformer decoder architecture adapted to receive such queries as input, but would Applicant assert that all of the references identified herein each fail to disclose a ‘plurality’ of query decoders? Applicant’s Fig. 5, illustrates query decoder 500, and Examiner’s doubt/suspicion is that e.g. 408a-n (606a-n, 806a-n, etc.,) even if disclosed as ‘separate’ decoders, are/may be realized by 500 (a single/shared instance of 500 – with N masks produced based on N queries). The literature identified herein appears to illustrate instances of ‘single’ transformer decoders, understood to produce a plurality of (per-query) masks. Even Mask2Former – Cheng et al. “Masked-attention Mask Transformer for Universal Image Segmentation” (Applicant’s 1/30/2025 IDS, NPL Citation No. 2) illustrates a ‘single’ transformer decoder (Lx) – with information suggesting it also produces two distinct e.g. ‘leaf’ and ‘plant’ masks – from https://mask2former.com/ PNG media_image4.png 650 882 media_image4.png Greyscale PNG media_image5.png 856 660 media_image5.png Greyscale Also, Cavagnero et al. “PEM: Prototype-based Efficient MaskFormer for Image Segmentation” (06 May 2024) – Figure 2 reproduced below: see “x2” in Transformer Decoder portion: PNG media_image6.png 350 976 media_image6.png Greyscale Cavagnero (PEM) discloses in e.g. Section 3.1 N learnable queries and N generated masks. Cheng’s Mask2Former (referenced above - 1/30/2025 IDS citation No. 2) explicitly discloses at Section 3.2.2 “We repeat this 3-layer Transformer decoder L times”. Applicant’s assistance is requested in distinguishing the invention, particularly as claimed and otherwise, over the state of the prior art, and clarifying the record at large. There is a considerable body of potential prior art, limited time for Examination, and Examiner requests Applicant’s assistance in ensuring the independent claims clearly and precisely capture at least those elements considered by Applicant to be distinguishing/essential. See also MPEP 2173.02 sub-section I, identifying how claims under examination are construed differently and with a potentially lower threshold for ambiguity. During prosecution the Office construes claims by giving them their broadest reasonable interpretation consistent with the specification in an effort to establish a clear record of what the applicant intends to claim. Such claim construction during prosecution may effectively result in a lower threshold for ambiguity than a court's determination. Packard, 751 F.3d at 1323-24, 110 USPQ2d at 1796-97 (Plager, J., concurring). However, applicant has the ability to amend the claims during prosecution to ensure that the meaning of the language is clear and definite prior to issuance or provide a persuasive explanation (with evidence as necessary) that a person of ordinary skill in the art would not consider the claim language unclear. In re Buszard, 504 F.3d 1364, 1366, 84 USPQ2d 1749, 1750 (Fed. Cir. 2007) Dependent claims 2-9, 11-16 and 18-20 inherit and fail to cure that/those deficiencies identified above for the case of independent claim(s) 1/10/17 and are rejected accordingly. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 1. Claims 1-3, 8-13 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Cavagnero et al. “PEM: Prototype-based Efficient MaskFormer for Image Segmentation” (06 May 2024) in view of Cheng et al. “Masked-attention Mask Transformer for Universal Image Segmentation” (Applicant’s 1/30/2025 IDS, NPL Citation No. 2), and “RAP-SAM: Towards Real-Time All-Purpose Segment Anything” (18 Jan 2024). As to claim 1, Cavagnero discloses a computer-implemented method comprising: extracting, by at least one processor utilizing an image encoder neural network, encoded feature maps from a digital image (Fig. 2 backbone in view of input image, page 3 Section 3.1 “To achieve this goal, we follow the framework provided by the MaskFormer architecture [4, 5, 37] that consists of three main components, depicted in Fig. 2: (i) a backbone extracting feature maps Fi ∈ RHi×Wi×Bi from the image I”); generating, by the at least one processor utilizing a pixel decoder neural network, a set of mask features from the encoded feature maps generated by the image encoder neural network (Fig. 2 Pixel Decoder stage following Backbone, page 3 Section 3.1 “(ii) a pixel decoder that processes Fi to produce high-resolution multi-scale features F̂i ∈ RHi×Wi×C, with i ∈ {1, 2, 3, 4} and Hi, Wi equal to the image resolution divided by, respectively, 4, 8, 16, and 32, and”); and generating, by the at least one processor utilizing a plurality of query decoder neural networks in connection with a plurality of segmentation tasks for the digital image, a plurality of object segmentation masks from the set of mask features generated by the pixel decoder neural network according to a plurality of separate sets of learned queries (page 3 Section 3.1 “(iii) a transformer decoder which accepts three multi-scale features F̂i, i ∈ {2, 3, 4} as input together with N learnable queries Q ∈ RN×C and it generates N refined queries Q̂ ∈ RN×C. To generate the N binary masks”, page 2 Section 2 “We benchmark the proposed PEM architecture on two distinct tasks, semantic and panoptic segmentation, on two datasets: Cityscapes [6] and ADE20K [38] (see Fig. 1)”, etc.,). Under any assertion that Cavagnero fails to fairly disclose at least two query decoder equivalents despite that “x2” illustrated, and/or in further view of those considerations identified above with respect to 112(b), numerous references as identified in the corresponding 112(b) rejection above appear to evidence the obvious nature of relying on such a plurality of query decoder equivalents (see at least Jain and Shah above “Lx” and “x L” respectively) (decoupled for the purposes of e.g. minimizing task interference). Cheng (Mask2Former) further evidences the obvious nature of that final generating as recited (Fig. 2 Transformer Decoder portion following Pixel Decoder and that “Lx”, page 1293 Section 3.2.2 “We repeat this 3-layer Transformer decoder L times”). Xu’s RAP-SAM further evidences the obvious nature of varying decoder architecture between shared and decoupled embodiments – see Fig. 3(b) and (d) in particular disclosing decoupled decoder designs as obvious variants of shared embodiments. PNG media_image7.png 324 1340 media_image7.png Greyscale Xu further discloses at page 7 Section 5.2, that a decoupled decoder would have more capacity, and may boost final performance by being better able to exploit benefits associated with larger/known feature extraction/backbone portions (see also reference therein to [42] Panoptic-Partformer). It would have been obvious to a person of ordinary skill in the art, before the effective filing date, to modify the system and method of Cavagnero, so as to implement a plurality of query decoder equivalents as taught/suggested in Cheng, and/or so as to implement a shared/single decoder alternatively in a decoupled manner as taught/ suggested by Xu, the motivation as similarly taught/suggested therein that such a plurality of decoupled decoders may make better use of extracted features and/or minimize task interference. As to claim 2, Cavagnero in view of Cheng and Xu teaches/suggests the method of claim 1. Cavagnero in view of Cheng and Xu further teaches/suggests the method wherein generating the set of mask features from the encoded feature maps comprise generating the set of mask features as a single set of mask features based on the encoded feature maps utilizing Cavagnero discloses high resolution multi-scale features F̂i derived from multi-scale convolutional pyramid network relying on deformable convolutions and self-modulation CSM – at page 2 Section 1 “While previous works enhanced a feature pyramid network (FPN) with transformer-based attention modules [4, 5, 37], we employ a more efficient fully convolutional FPN. We supplement it with a context-based self-modulation module to recover the global context and deformable convolutions [7] to allow each kernel to dynamically modify its receptive field to focus on relevant regions”; Cavagnero’s attention disclosure is limited to the transformer architecture following the pixel decoder, with the exception of that reference to [4, 5, 37] suggesting the obvious nature of pixel decoder architecture implementing attention mechanisms (FPN with transformer based attention) – Examiner notes [5] in Cavagnero is Cheng, Mask2Former); and generating the plurality of object segmentation masks comprises generating the plurality of object segmentation masks from the single set of mask features utilizing the plurality of query decoder neural networks (Cavagnero as proposed above, in view of F̂i as passed between the pixel decoder and transformer decoder(s) of Cavagnero Fig. 2, and at least those N masks Section 3.1 “we aim to predict a set of N binary masks M ∈ {0, 1}N×H×W each associated with a probability distribution pi ∈ ΔK+1, where (H,W) is the height and width of the image, and K + 1 is the number of classes plus an additional “no object” class”). Cheng further suggests the obvious nature of a pixel decoder utilizing a transformer neural network (page 1294, Section 4.1 Implementation details “We adopt settings from [14] with the following differences: Pixel decoder. Mask2Former is compatible with any existing pixel decoder module. In MaskFormer [14], FPN [33] is chosen as the default for its simplicity. Since our goal is to demonstrate strong performance across different segmentation tasks, we use the more advanced multi-scale deformable attention Transformer (MSDeformAttn) [66] as our default pixel decoder. Specifically, we use 6 MSDeformAttn layers applied to feature maps with resolution 1/8, 1/16 and 1/32, and use a simple upsampling layer with lateral connection on the final 1/8 feature map to generate the feature map of resolution 1/4 as the per-pixel embedding”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date, to further modify the system and method of Cavagnero in view of Cheng and Xu, so as to implement a known/established pixel decoder alternative, implementing a transformer neural network e.g. MSDeformAttn as taught/suggested by Cheng (see MPEP 2143 Rationale B, in further view of predictable results as evidenced by Cheng, and/or references referenced therein [66]), the motivation as similarly taught/suggested therein that such an attention mechanism (as an alternative to e.g. CNN with localized kernels) may allow for resultant feature maps to dynamically attend to information outside of said localized kernel (more context aware at the image level). As to claim 3, Cavagnero in view of Cheng and Xu teaches/suggests the method of claim 1. Cavagnero in view of Cheng and Xu further teaches/suggests the method wherein generating the plurality of object segmentation masks (Cavagnero generating those N masks, Section 3.1 “we aim to predict a set of N binary masks M”) comprises: generating, utilizing a first query decoder neural network for a first segmentation task (see claim 1 above), a first object segmentation mask from the set of mask features generated by the pixel decoder neural network (any one of the disclosed masks in view of Section 3.1, see also e.g. page 2 Section 1 “we incorporate a prototype selection mechanism which leverages a single visual token for each object descriptor”, page 4 Prototype Selection “The goal of the cross-attention is to refine each input query based on the visual features of the object it represents. However, we argue that using all the pixels belonging to an object for the refinement is redundant, since pixels associated with a specific object query will naturally become close to each other as training progresses. We can leverage this inherent redundancy to focus solely on the most relevant feature for each object, i.e. the prototype”, etc.,); and generating, utilizing a second query decoder neural network for a second segmentation task, a second object segmentation mask from the set of mask features generated by the pixel decoder neural network (see above, for any second instance/object – see also final panoptic prediction and multiple object classes and object instances). As to claim 10, this claim is the system claim corresponding to the method of claim 1 and is rejected accordingly. In response to any argument/assertion that e.g. Cavagnero fails to explicitly disclose generic computer processor and/or memory, Examiner would assert a reasonable reading of Cavagnero at least implies if not inherently requires such structural components. Cheng further explicitly discloses associated GPU memory constraints (Cheng at §§ 4.2 and 1 with reference to MaskFormer) further evidencing an implied and/or inherent nature of such structural components. As to claim 17, this claim is the non-transitory CRM claim sufficiently incorporating and corresponding to the method of claim 1, and is rejected accordingly. Additional References Prior art made of record and not relied upon that is considered pertinent to applicant's disclosure: Additionally cited references (see attached PTO-892) otherwise not relied upon above have been made of record in view of the manner in which they evidence the general state of the art. As suggested above, alternative combination(s) of references may serve in rejections to at least the independent claim(s) as recited, as multiple pieces of literature appear to disclose at least those three components that are a backbone, pixel decoder, and Lx Transformer decoder (query decoder equivalents). Allowable Subject Matter Claims 4-9, 11-16 and 18-20 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims. References of record fail to serve in any obvious combination teaching each and every limitation as required therein. With respect to claim(s) 4, 14, 16 and 19-20 in particular, Examiner has considered e.g. Chen at al. “Vision Transformer Adapter for Dense Predictions”, which may serve to evidence the obvious nature of an adapter neural network broadly, while arguably failing to serve in any obvious combination teaching the claim(s) as a whole, particularly in view of the previously proposed combination of references as required for proper rejection of independent and/or intervening claim(s), and additional/subsequent modification(s) thereto as would be necessary/required. Inquiry Any inquiry concerning this communication or earlier communications from the examiner should be directed to IAN L LEMIEUX whose telephone number is (571)270-5796. The examiner can normally be reached Mon - Fri 9:00 - 6:00 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chan Park can be reached on 571-272-7409. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /IAN L LEMIEUX/Primary Examiner, Art Unit 2669
Read full office action

Prosecution Timeline

Nov 08, 2024
Application Filed
Jul 24, 2026
Non-Final Rejection mailed — §103, §112
Jul 28, 2026
Examiner Interview (Telephonic)
Jul 28, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705792
APPARATUS AND METHODS FOR IMPROVING DRIVER MONITORING SYSTEMS
3y 10m to grant Granted Aug 11, 2026
Patent 12705879
SYSTEMS AND METHODS FOR UNDERWATER FLOW ANALYSIS
3y 4m to grant Granted Aug 11, 2026
Patent 12705484
IMAGE CLASSIFICATION USING BATCH NORMALIZATION LAYERS
2y 7m to grant Granted Aug 11, 2026
Patent 12700135
System And Method For Image Based Registration And Calibration
2y 10m to grant Granted Aug 04, 2026
Patent 12700136
GARMENT PALLETIZATION
2y 8m to grant Granted Aug 04, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
87%
Grant Probability
96%
With Interview (+8.9%)
2y 2m (~4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 587 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month