Prosecution Insights
Last updated: October 01, 2026
Application No. 18/680,579

ROBUST AND CONSISTENT VIDEO INSTANCE SEGMENTATION

Final Rejection §103
Filed
May 31, 2024
Examiner
SOFRONIOU, MICHAEL MARIO
Art Unit
2661
Tech Center
2600 — Communications
Assignee
Adobe Inc.
OA Round
2 (Final)
100%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 100% — above average
100%
Career Allowance Rate
3 granted / 3 resolved
+38.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 3m
Avg Prosecution
21 currently pending
Career history
22
Total Applications
across all art units

Statute-Specific Performance

§101
5.6%
-34.4% vs TC avg
§103
42.3%
+2.3% vs TC avg
§102
12.7%
-27.3% vs TC avg
§112
33.8%
-6.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 3 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Disposition of the Claims In response to applicant’s amendment received on 07/08/2026, all requested changes to the claims and specification have been entered. Claim(s) 1-20 were previously pending. No claim(s) have been added. No Claim(s) have been cancelled. Claim(s) 1-20 are currently pending. Claims(s) 1, 8 & 15 have been amended. Response to Amendment Specification Objections Applicant has amended ¶0018 of the specification to correct the previously indicated minor informality. The objection to the specification is withdrawn. Response to Arguments Claim Rejections – 35 U.S.C. § 103 Claims 1 & 8 Applicant’s arguments, see Applicant Remarks, pg. 7-9, filed 07/08/2026, with respect to the rejection(s) of claim(s) 1 & 8 under 35 U.S.C. § 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. Claim 15 Applicant's arguments, see Applicant Remarks, pg. 9-10, filed 07/08/2026 have been fully considered but they are not persuasive. Applicant argues that the cited references do not teach or fail to suggest, “generating a spatial identity of a previous frame by encoding an embedding of the object depicted in the previous frame into respective spatial regions defined by a mask of the previous frame”. Applicant states that neither of the previously cited references, Heo, or Oh, disclose generating a spatial identity by encoding an object embedding into respective spatial regions defined by a mask. Applicant argues that Oh merely generates an embedding using a frame and a mask, and that neither Oh nor Heo disclose encoding an object embedding into mask-defined spatial regions to create a spatial identity representation. The examiner respectfully disagrees. When looking at Fig. 3 of Oh, previous frames with a masked object are used to encode memory embeddings (i.e., a spatial identity) in the forms of keys and values via the memory encoder EncM. These memory embeddings inherently encode the masked object of the previous frame to convey the spatial regions defined by the mask. Prior art mappings have been updated to reflect the amended claim scope in 35 U.S.C. § 103 rejections below. Therefore, applicant’s argument is not persuasive. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 15-17 & 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Heo et al (“VITA: Video Instance Segmentation via Object Token Association”, 2022, NeurIPS) in view of Oh et al (“Space-Time Memory Networks for Video Object Segmentation with User Guidance”, 2022, IEEE). Regarding claim 15, Heo teach A system (Heo: the video instance segmentation (VIS) system “VITA” of Fig. 1(c), further outlined in Fig. 2) comprising: a memory component (Heo: the system is implemented via 4 NVIDIA A100 GPUs with 40 GB of memory [Sec A.1 – Implementation - ¶01]); and a processing device coupled to the memory component (Heo: 4 NVIDIA A100 GPUs with 40 GB of memory [Sec A.1 – Implementation - ¶01]), the processing device to perform operations comprising: obtaining a frame of a video sequence (Heo: frames are input into the frame-level detector (Mask2Former) [Sec. 3.1 – Frame-level Detector - ¶01], further depicted as input frames of Fig. 2), wherein the frame depicts an object (Heo: the frame-level detector parses object queries, later referred to as frame queries [Sec. 3.1 – Frame-level Detector - ¶01]); determining frame features using the frame (Heo: the backbone extracts frame features from input frames [Sec. 3.1 – Frame-level Detector - ¶01; Fig. 2]); and generating a masked frame using a pixel embedding and an embedding of the object depicted in the frame (masks are generated in VITA using pixel embeddings (pink parallelograms) and object tokens [Fig. 2; Sec. 3.2 – Object Decoder and Output Heads - ¶03]), wherein the pixel embedding is based on the … frame features (the frame-level detector generates pixel embeddings from spatially-encoded features of the backbone [Sec 3.1 – Frame-level Detector - ¶01; Fig 2.]), but do not teach that the pixel embedding is based on an augmented spatial identity, nor do they teach generating a spatial identity using the object embedding or encoding the background of a previous frame. Oh, on the other hand, teach generating a spatial identity of a previous frame by encoding an embedding of the object depicted in the previous frame into respective spatial regions defined by a mask of the previous frame (Oh: the memory encoder encodes memory embeddings (i.e., a spatial identity) in the form of keys and values (feature tensors with an associated key [Sec 3.1.2 – Memory Encoder - ¶02]) which provide spatial feature representations for respective regions of masks of previous frames [Sec 3 – Space-Time Memory Networks - ¶01-02; Fig. 3]); and generating an augmented spatial identity using the spatial identity and encoding a background of the previous frame; (Oh: the space-time memory read operation uses keys and values (that provide spatial feature representations) from previous frames, that also encode visual semantics regarding whether a feature pixel belongs to a foreground and background [Sec 3 – Space-Time Memory Networks - ¶03], and concatenates them with the current frame keys and values to form an aligned feature vector y / y-I that conveys a spatial identity of an object [Sec. 3.2 – Space-Time Memory Read - ¶01-02; Fig. 4; Eq. 1]). Oh further disclose that their method performs memory-driven object detection through offline training, which frees it from propagation driven problems [Sec. 1 – Introduction - ¶07-08]. Therefore, it would have been obvious to one of ordinary skill before the effective filing date of the present application to incorporate the spatial identity generation provided by Oh that represent an object spatial context to further define the pixel embeddings generated in Heo to arrive at the invention of the instant application. The motivation for doing so would be to provide both temporally and spatially enriched context of objects for improved segmentation and temporal consistency. Regarding claim 16, Heo in view of Oh teach The system of claim 15 (as described above), wherein the processing device performs further operations comprising: determining the embedding of the object depicted in the frame using the object depicted in the previous frame and the pixel embedding (Heo: the frame-level detector, outputs pixel embeddings via the pixel decoder, which are then utilized by the transformer decoder to inform the object token generation performed by VITA [Sec. 3.1 – Frame-level Detector; Fig. 2, with particular attention to the parallelograms of the pixel encoder, representing pixel embeddings], the object encoder of VITA, employs self-attention along the temporal axis, where object tokens (embeddings of objects) from different frames (i.e., a past frame), can exchange object-wise information [Sec. 3.2 – VITA – Object Encoder - ¶01]). Regarding claim 17, Heo in view of Oh teach The system of claim 15 (as described previously), wherein the masked frame includes a masked object corresponding to the object (Heo: the object detector predicts masks of objects in videos [Sec. 3.2 – VITA - ¶01]). Regarding claim 19, Heo in view of Oh teach The system of claim 15 (as described previously), wherein generating the masked frame using the pixel embedding and the embedding of the object depicted in the frame (Heo: masks are generated in VITA using pixel embeddings (pink parallelograms) and object tokens [Fig. 2; Sec. 3.2 – Object Decoder and Output Heads - ¶03]) includes the processing device performing further operations comprising: convolving the embedding of the object depicted in the frame with the pixel embedding to generate a probability distribution, wherein the probability distribution indicates a likelihood of each pixel of the frame belonging to a masked object of the masked frame (Heo: the frame-level detector generates a 1 × 1 convolutional weight from frame / object queries (token representations) and per-pixel embeddings, then applying a dot product between the two embeddings to segment objects [Sec 3.1 – Frame-level Detector], this segmentation process further defines a probability for a frame containing a segmented object [Sec. 4.4 – Ablation Studies: Pruning Tokens - ¶01]). Regarding claim 20, Heo in view of Oh teach The system of claim 15 (as described previously), wherein the masked frame comprises (due to the use of the disjunctive “or”, only of either of these limitations is required to be mapped to) or more masked objects (Heo: the object detector predicts masks of objects in videos [Sec. 3.2 – VITA - ¶01]). Claim 18 are rejected under 35 U.S.C. 103 as being unpatentable over Heo et al (“VITA: Video Instance Segmentation via Object Token Association”, 2022, NeurIPS) in view of Oh et al (“Space-Time Memory Networks for Video Object Segmentation with User Guidance”, 2022, IEEE), further in view of Kirillov et al (“Panoptic Feature Pyramid Networks”, IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019). Regarding claim 18, Heo in view of Oh teach The system of claim 15 (as described previously), however, they do not teach that the background of a past frame is learned through end-to-end supervised learning. Kirillov, however, teach wherein encoding the background of the previous frame is learned during end-to-end supervised learning (Kirillov: the semantic segmentation branch of the panoptic feature pyramid network generates per-pixel class labels “stuff” for non-object pixels [Sec 3.1 – Model Architecture: Semantic Segmentation Branch - ¶01], note that Kirillov et al splits classification into object “thing” and non-object “stuff” classes [ Sec 1 – Introduction - ¶01], further illustrated in Fig. 2, wherein “thing” classes are highlighted with blue or red overlays, while “stuff” classes, relating to the background of the scene, are highlighted in green). Kirillov further states that their approach allows for facile incorporation of semantic segmentation with known instance segmentation methods to a yield lightweight and effective method for frame classification [Abstract]. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to utilize the semantic segmentation of Kirillov to isolate background features and incorporate that in parallel with the instance segmentation outlined by Heo in view of Oh to arrive at the invention of the instant application. Allowable Subject Matter Claims 1 – 14 are allowed. Regarding claim 1, the primary reason for allowance is that the prior art does not teach or fails to suggest “determining an object token for the frame by inputting a past object token associated with the past frame and the pixel embedding to a decoder”. The closest cited prior art, Heo disclose object tokens informed by past object tokens via a self-attention layer, and that these object tokens are generated using pixel embeddings input into a transformer decoder, but this decoder does not take previous object tokens in as inputs for determining current object tokens . The closest prior art of record, Liu (Liu et al, “InstMove: Instance Motion for Object-centric Video Segmentation”, 2023, IEEE/CVF) describe an object-centric video segmentation method levering pixel-wise motion for accurate object tracking. Liu teach using feature maps generated by upscaling low-level image features (which the examiner interprets to be a form of pixel embedding in accordance with the disclosure of the current application [¶0079, 97]) and input them into a decoder along with the previous motion feature or motion pattern for a respective pattern. These motion features / patterns could be construed as a form of object token, however, Liu fails to describe inputting these feature maps and past motion features to determine current motion features, and thus fails to read on the language of the claim. Girdhar (Girdhar et al, US 2023/0386203 A1), prior art also made of record, describe a video transformer model leveraging a spatial-attention encoder and temporal-attention decoder to predict future features and anticipate future actions. Girdhar disclose inputting past frame features (which the examiner construes as a past object token) along with temporal position embeddings into a casual transformer decoder to predict future frame features, however, these temporal position embeddings fail to describe embedding any particular pixel or image patches and simply provide temporal context for associated past frame features. Claims 2-7, which depend on claim 1, are also allowed given their dependence on a claim already indicated as allowable. Regarding claim 8, a similar analysis to that performed for claim 1 can also be applied as they recite substantially identical subject matter. Accordingly, claims 9-14, which depend on claim 8, are also all allowed given their dependence on a claim already indicated as allowable. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael M. Sofroniou whose telephone number is (571)272-0287. The examiner can normally be reached M-F: 8:30 AM - 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, John M. Villecco can be reached at (571) 272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MICHAEL M SOFRONIOU/Examiner, Art Unit 2661 /JOHN VILLECCO/Supervisory Patent Examiner, Art Unit 2661
Read full office action

Prosecution Timeline

May 31, 2024
Application Filed
Apr 09, 2026
Non-Final Rejection mailed — §103
Jul 02, 2026
Interview Requested
Jul 06, 2026
Examiner Interview Summary
Jul 06, 2026
Applicant Interview (Telephonic)
Jul 08, 2026
Response Filed
Aug 31, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12711652
IMAGE PROCESSING APPARATUS, IMAGE PICKUP APPARATUS, CONTROL METHOD FOR IMAGE PROCESSING APPARATUS, AND STORAGE MEDIUM CAPABLE OF NOTIFYING USER OF BLUR INFORMATION, OR OF DISPLAYING INDICATOR
2y 8m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 1 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
100%
Grant Probability
99%
With Interview (+0.0%)
2y 3m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 3 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month