DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Preliminary Amendment
In response to applicant’s preliminary amendment received on 10/31/2024, all requested changes to the claims have been entered.
Claim(s) 1-26 were previously pending.
No Claim(s) have been added.
Claim(s) 5, 10, 20, 24 & 26 have been cancelled.
Claim(s) 1-4,6-9,11-14,16-19,21-23 and 25 are currently pending.
Priority
Acknowledgement is made of applicant’s claim for priority under 35 U.S.C. 119(e) to US provisional applications, 63/610,544, filed 12/15/2023 and 63/656,777, filed 06/06/2024.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 10/31/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are:
Claim 11 recites:
“an embedding module configured to refine query embeddings…”
“a fusion module configured to fuse the forward embeddings and the backward embeddings…”
“a classification module configured to generate a classification…”
Claim 12 recites:
“a forward offline refiner configured to refine the merged online embeddings…”
“a backward offline refiner configured to refine the reversed online embeddings…”
Claim 19 recites:
“a visual transformer adapter configured to generate multiple-scale features…”
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
The following limitations are computer-implemented means-plus-function limitations, wherein the corresponding structure and associated algorithm as disclosed in the specification for each limitation are being interpreted (see MPEP 2181(B)):
“an embedding module” (element 130 of Fig. 1; step 204 of Fig. 2; element 330 of Fig. 3; process 400 of Fig. 4 performed by 330; element 530 of Fig. 5; process 600 of Fig. 6, performed by 530 | element 1030 of Fig. 10; element 1201 of Fig. 12A & B; ¶ 0037, 41, 48-55, 64, 69, 72)
The associated structure of the embedding module 130 is processor 1420 of Fig. 14.
The associated algorithm for performing the claimed function is as follows: Learned query embeddings are provided to the embedding module 130 for refinement to generate forward and backward embeddings to perform operation 204. The embedding module 330 may comprise an online tracker 331, a forward offline refiner 332 and a backward offline refiner 333 to perform offline refining. The embedding module 330 may perform the entirety of process 400. The embedding module 330 may reverse a time order of the merged online embeddings to obtain reversed online embeddings at operation 404. It may then reverse a time order of the refined reversed online embeddings from the backward offline refiner 333 to generate the backward embeddings. The embedding module 530 may comprise a forward online tracker 531 and a backward online tracker 532 to perform online refining. The embedding module 530 may perform the entirety of process 600.
“a fusion module” (element 140 of Fig. 1; step 205 of Fig. 2, element 740 of Fig. 7, element 741 of Fig. 8, element 1040 of Fig. 10; ¶0037-39, 42, 49, 56-58, 64-65, 69)
The associated structure of the fusion module is processor 1420 of Fig. 14.
The associated algorithm for performing the claimed function is as follows: The fusion module 140, otherwise referred to as forward and backward embedding fusion (FBEF) module, generates fused embeddings from the forward and backward refined embeddings according to step 205. The FBEF module may comprise fusion learning blocks 741, 742… outlined in Figs. 7 and 8 to output fusion weights. The fusion learning blocks may comprise a long-term temporal self-attention module 8411, a short-term temporal convolution block 8412, instance self-attention module 8413, cross-attention module 8414, and feed-forward network module 8415.
“a classification module” (element 150 of Fig. 1; operation 206 of Fig. 2, element 1050 of Fig. 10; element 1150 of Fig. 11; ¶0037, 43, 50, 64-69)
The associated structure of the classification module is processor 1420 of Fig. 14.
The associated algorithm for performing the claimed function is as follows: The classification module 150 obtains predicted classification corresponding to one or more objects in the video based on the fused embeddings according to step 206 of Fig. 2. It may use the fused embeddings to generated predicted classification masks which may be applied to one or more frames of the video in order to indicate the class of one or more objects included in the video. The classification module 1050 may also further generate forward and backward classification masks.
“an online tracking module” (element 331 of Fig. 3, operation 401 of Fig. 4; 531 and 532 of Fig. 5; operations 601 & 603 of Fig. 6, element 1131 of Fig. 11; ¶0051-55, 66-69)
The associated structure of the online tracking module is processor 1420 of Fig. 14.
The associated algorithm for performing the claimed function is as follows: The online tracking module, referred to as "online tracker" 331, is part of the embedding module 350 to perform online refining. The online tracker generates online embeddings by refining query embeddings according to operation 401 of Fig. 4. The online tracker may refine groups of initial query embeddings to generate online embeddings for each clip. In Fig. 5, the online tracker is divided into a forward (531) and backward (532) online tracker. The forward online tracker may refine each group of query embeddings to generate forward online embeddings for each clip according to operation 601 of Fig. 6. The backward online tracker may perform online refining on the reversed online embeddings to generate refined reversed query embeddings according to operation 603.
“a forward offline refiner” (element 332 of Fig. 3; ¶0051, 53, 66)
The associated structure of the forward offline refiner is processor 1420 of Fig. 14.
The associated algorithm for performing the claimed function is as follows: The forward offline refiner 332 is a component of the of embedding module 330. It may perform offline refining on merged online embeddings to generate forward embeddings at step 403 of Fig. 4.
“a backward offline refiner” (element 333 of Fig. 3; ¶0051, 53, 66, 69)
The associated structure of the backward offline refiner is processor 1420 of Fig. 14.
The associated algorithm for performing the claimed function is as follows: The backward offline refiner 333 is a component of the of embedding module 330. It may process reversed online embeddings to generate refined reversed online embeddings at step 405 of Fig. 4.
“a visual transformer adapter” (element 1112 of Fig. 11, 1112A & B of Figs 12A & B; ¶0067, 70-76)
The associated structure of the visual transformer adapter is processor 1420 of Fig. 14.
The associated algorithm for performing the claimed function is as follows: The visual transformer (ViT) adapter 1112 may process clips in tandem with the ViT which are provided to pixel decoder 1160. The image encoder 1110 may include the ViT adapter which may be used together with the ViT 1111 to generate multiple-scale features. The ViT adapter 1112A/B may include a spatial prior module 1211/1221 and one or more extractors 1212/1222. Outputs of the ViT blocks may be split into stages and interact with a corresponding extractor 1212 and ViT adapter 1112A. The ViT adapter 1112B may include one or more multi-receptive field feature pyramid modules 1223 which may be inserted before each extractor module 1222. These may include a feature pyramid and multi-receptive field convolution layers. The ViT adapter 1112B may be based on a vision transformer with convolutional multiple-scale feature interaction (ViT-CoMer).
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112(b)
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 25 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 25 recites the limitation "The method of claim 21". Claim 21 is directed to a non-transitory computer-readable medium storing instructions, with no explicit mention of an associated method. There is insufficient antecedent basis for this limitation in the claim.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 2, 4, 7, 21 & 22 is/are rejected under 35 U.S.C. 103 as being unpatentable over He et al (US 2025/0113087 A1), hereinafter referred to as “He” in view of Mehl et al (“M-FUSE: Multi-frame Fusion for Scene Flow Estimation”, IEEE/CVF Winter Conference on Applications of Computer Vision, 2023), hereinafter referred to as “M-FUSE”.
Regarding claim 1, He teach A method of performing video segmentation (process 500 of Fig. 5 [¶0040]), the method comprising:
obtaining a plurality of frames from an input video (at step 502, a video is obtained to be divided into a plurality of clips -- a video by nature is comprised of a plurality of frames [¶0045; Fig. 5]);
extracting a plurality of features from the plurality of frames (at step 506, clip features are generated by the first sub-model 102 for each [¶0045; Fig. 5]);
obtaining query embeddings corresponding to the plurality of features (at step 508, a set of object queries based on the clips features are generated by the transformer decoder [¶0046; Fig. 5]);
and
generating a classification prediction corresponding to the input video based on the fused embeddings (the second sub-model 106 performs trajectory-attention operations to encode clip-level predictions for an object class [¶0059], performing panoptic segmentation to segment "thing" and "stuff" classes in a video [¶0042]).
While He disclose generating classification predictions, they fail to disclose distinct forward and backward refining of query embeddings and fusing them in order to make the classification prediction. M-FUSE, on the other hand, is analogous art pertinent to the field of endeavor of the present application and disclose a neural network system for flow estimation that creates and fuses forward and backward embeddings. M-FUSE teach refining the query embeddings in a forward time order to generate forward embeddings, and in a backward time order to generate backward embeddings (M-FUSE: at step 2, embVec (embeddings) are generated for both the forward and backward directions of scene flow [Sec 3.2: Multi-frame Fusion Network, steps 1-2; Fig. 2]);
fusing the forward embeddings and the backward embeddings to obtain fused embeddings; (M-FUSE: at step 7, forward and backward embeddings are fused together [Sec. 3.2: Multi-frame Fusion Network, steps 6-7; Fig. 2]).
M-FUSE further explain that by estimating forward and backward flow, the model is able to learn flow of objects in occluded regions without an explicit motion model [ Sec 2: Related work, subsec: multi-frame optical flow - ¶01-03], with the fusion of forward and backward embeddings allowing the integration of temporal information on demand. [ Sec: Abstract]. Therefore, it would have been obvious to one of ordinary skill to utilize the forward and backward embedding fusion of M-FUSE to inform the classification prediction of He to better track and classify objects to better account for occluded objects during video segmentation.
Considering claim 2, He in view of M-FUSE teach The method of claim 1 (as described above), wherein the plurality of frames are divided into a plurality of clips (He: at step 502, a video is divided into a plurality of clips [¶0045; Fig. 2]), and
wherein the refining comprises:
performing online refining on a group of query embeddings corresponding to each clip of the plurality of clips to generate online embeddings for each clip of the plurality of clips (He: in some embodiments, the first sub-model 102 performs segmentation in a near-online fashion [¶0036], with object queries corresponding to each clip being generated in step 508 [0046; Fig. 2]);
merging the online embeddings for the plurality of clips to obtain merged online embeddings; (He: at step 510, trajectory attention is applied by the second sub-model 106, to refine object queries and to track objects across the plurality of clips (i.e., merged across the plurality of clips) [¶0057; Fig. 2]).
He, however, fails to recite explicit forward and backward embedding fusion. M-FUSE, on the other hand, teach performing offline refining on the merged online embeddings to generate the forward embeddings (M-FUSE: at step 2, forward embeddings are generated for a forward time order [Sec 3.2: Multi-frame Fusion Network, steps 1-2; Fig. 2]);
reversing a time order of the merged online embeddings to obtain reversed online embeddings (M-FUSE: at step 1, the time order is reversed (t -> t - 1) M-FUSE: at step 2, backward embeddings are generated for a reverse time order [Sec 3.2: Multi-frame Fusion Network, step 1; Fig. 2]);
performing the offline refining on the reversed online embeddings to generate refined reversed online embeddings (M-FUSE: at step 2, backward embeddings are generated for a reverse time order [Sec 3.2: Multi-frame Fusion Network, steps 1-2; Fig. 2]); and
reversing a time order of the refined reversed online embeddings to generate the backward embeddings (M-FUSE: at step 5, the time order for the reversed for embedding vectors to enable backward-to-forward prediction [Sec 3.2: Multi-frame Fusion Network, step 5; Fig. 2]).
M-FUSE further explain that by estimating forward and backward flow, the model is able to learn flow of objects in occluded regions without an explicit motion model [ Sec 2: Related work, subsec: multi-frame optical flow - ¶01-03], with the fusion of forward and backward embeddings allowing the integration of temporal information on demand. [ Sec: Abstract]. Therefore, it would have been obvious to one of ordinary skill to utilize the forward and backward embedding fusion of M-FUSE to inform the classification prediction of He to better track and classify objects to better account for occluded objects during video segmentation.
With respect to claim 4, He in view of M-FUSE teach The method of claim 1 (as described previously), wherein the fusing comprises generating a plurality of fusion weights using a fusion module comprising a plurality of fusion learning blocks (M-FUSE: the fusion model's weights are learned via the original training of the RAFT-3D approach [ Sec 3.3: Supervision - ¶01; Sec 4: Experiments - ¶01-02; eq. 3]).
M-FUSE further explain that by estimating forward and backward flow, the model is able to learn flow of objects in occluded regions without an explicit motion model [ Sec 2: Related work, subsec: multi-frame optical flow - ¶01-03], with the fusion of forward and backward embeddings allowing the integration of temporal information on demand. [ Sec: Abstract]. Therefore, it would have been obvious to one of ordinary skill to utilize the fusion weighting of M-FUSE to inform the classification prediction of He to better track and classify objects to better account for occluded objects during video segmentation.
As for claim 7, He in view of M-FUSE teach The method of claim 1 (as described previously), further comprising generating a plurality of classification masks corresponding to the input video based on the classification prediction and the plurality of features (He: The segmentation sytem 100 deploys segmentation masks for "thing" and "stuff" classes based on clip-level predictions and object features [¶0042-43, 59]).
Concerning 21, He teach A non-transitory computer-readable medium storing instructions (computer-readable storage medium (ROM) 1420 [¶0076; Fig. 14]) which, when executed by at least one processor of a device for performing video segmentation (CPU(s) 1404 [¶0073; Fig. 14]), causes the at least one processor to:
obtain a plurality of frames from an input video (at step 502, a video is obtained to be divided into a plurality of clips -- a video by nature is comprised of a plurality of frames [¶0045; Fig. 5]);
extract a plurality of features from the plurality of frames (at step 506, clip features are generated by the first sub-model 102 for each [¶0045; Fig. 5]);
obtain query embeddings corresponding to the plurality of features (at step 508, a set of object queries based on the clips features are generated by the transformer decoder [¶0046; Fig. 5]);
and
generate a classification prediction corresponding to the input video based on the fused embeddings (the second sub-model 106 performs trajectory-attention operations to encode clip-level predictions for an object class [¶0059], performing panoptic segmentation to segment "thing" and "stuff" classes in a video [¶0042]).
While He disclose generating classification predictions, they fail to disclose distinct forward and backward refining of query embeddings and fusing them in order to make the classification prediction. M-FUSE, on the other hand, teach refine the query embeddings in a forward time order to generate forward embeddings, and in a backward time order to generate backward embeddings (M-FUSE: at step 2, embVec (embeddings) are generated for both the forward and backward directions of scene flow [Sec 3.2: Multi-frame Fusion Network, steps 1-2; Fig. 2]);
fuse the forward embeddings and the backward embeddings to obtain fused embeddings; (M-FUSE: at step 7, forward and backward embeddings are fused together [Sec. 3.2: Multi-frame Fusion Network, steps 6-7; Fig. 2]).
M-FUSE further explain that by estimating forward and backward flow, the model is able to learn flow of objects in occluded regions without an explicit motion model [ Sec 2: Related work, subsec: multi-frame optical flow - ¶01-03], with the fusion of forward and backward embeddings allowing the integration of temporal information on demand. [ Sec: Abstract]. Therefore, it would have been obvious to one of ordinary skill to utilize the forward and backward embedding fusion of M-FUSE to inform the classification prediction of He to better track and classify objects to better account for occluded objects during video segmentation.
Considering claim 22, He in view of M-FUSE teach The non-transitory computer-readable medium of claim 21 (as described above), wherein the plurality of frames are divided into a plurality of clips (He: at step 502, a video is divided into a plurality of clips [¶0045; Fig. 2]), and
wherein to perform the refining, the instructions further cause the at least one processor (CPU(s) 1404 [¶0073; Fig. 14]) to:
perform online refining on a group of query embeddings corresponding to each clip of the plurality of clips to generate online embeddings for each clip of the plurality of clips (He: in some embodiments, the first sub-model 102 performs segmentation in a near-online fashion [¶0036], with object queries corresponding to each clip being generated in step 508 [0046; Fig. 2]);
merge the online embeddings for the plurality of clips to obtain merged online embeddings; (He: at step 510, trajectory attention is applied by the second sub-model 106, to refine object queries and to track objects across the plurality of clips (i.e., merged across the plurality of clips) [¶0057; Fig. 2]).
He, however, fails to recite explicit forward and backward embedding fusion. M-FUSE, on the other hand, teach perform offline refining on the merged online embeddings to generate the forward embeddings (M-FUSE: at step 2, forward embeddings are generated for a forward time order [Sec 3.2: Multi-frame Fusion Network, steps 1-2; Fig. 2]);
reverse a time order of the merged online embeddings to obtain reversed online embeddings (M-FUSE: at step 1, the time order is reversed (t -> t - 1) M-FUSE: at step 2, backward embeddings are generated for a reverse time order [Sec 3.2: Multi-frame Fusion Network, step 1; Fig. 2]);
perform the offline refining on the reversed online embeddings to generate refined reversed online embeddings (M-FUSE: at step 2, backward embeddings are generated for a reverse time order [Sec 3.2: Multi-frame Fusion Network, steps 1-2; Fig. 2]); and
reverse a time order of the refined reversed online embeddings to generate the backward embeddings (M-FUSE: at step 5, the time order for the reversed for embedding vectors to enable backward-to-forward prediction [Sec 3.2: Multi-frame Fusion Network, step 5; Fig. 2]).
M-FUSE further explain that by estimating forward and backward flow, the model is able to learn flow of objects in occluded regions without an explicit motion model [ Sec 2: Related work, subsec: multi-frame optical flow - ¶01-03], with the fusion of forward and backward embeddings allowing the integration of temporal information on demand. [ Sec: Abstract]. Therefore, it would have been obvious to one of ordinary skill to utilize the forward and backward embedding fusion of M-FUSE to inform the classification prediction of He to better track and classify objects to better account for occluded objects during video segmentation.
Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over He et al (US 2025/0113087 A1), hereinafter referred to as “He” in view of Mehl et al (“M-FUSE: Multi-frame Fusion for Scene Flow Estimation”, IEEE/CVF Winter Conference on Applications of Computer Vision, 2023), hereinafter referred to as “M-FUSE”, further in view of Zhang et al (“DVIS: Decoupled Video Instance Segmentation Framework”, arXiv, 2023), hereinafter referred to as “DVIS”.
Regarding claim 6, He in view of M-FUSE teach The method of claim 4 (as described previously), and while M-FUSE generally recite learned fusion blocks for obtaining a plurality of fusion weights, they do not recite applying long-term temporal self-attention, short-term temporal convolution, instance self-attention, cross-attention and a feed-forward network for obtaining these weights.
DVIS, however, is analogous art pertinent to the field of endeavor. DVIS disclose a decoupled referring tracker and temporal refiner for video instance segmentation. More specifically, DVIS teach wherein for each fusion learning block of the plurality of fusion learning blocks, the fusing comprises:
applying long-term temporal self-attention in a forward time dimension corresponding to the forward time order;
performing short-term temporal convolution in the forward time dimension;
applying instance self-attention along a channel dimension, wherein each channel of the channel dimension corresponds to an instance of an object included in the input video;
applying cross-attention on the forward embeddings and the backward embeddings; and
applying a feed forward network to obtain the plurality of fusion weights (DVIS: the framework of the temporal refiner takes input queries applies long-term temporal self-attention, short-term temporal convolution, instance self-attention per object instance, cross-attention, and a feed forward network (FFN) for queries in a forward time order for the temporally weighted sum QRf [Sec 3.2: Temporal Refiner - ¶01-3; Fig. 3]).
DVIS explain that their temporal refiner address issues with online video instance segmentation that lack a proper refinement step. The provided temporal refiner utilizes information across the entire video the refine outputs to correct instance representation [Sec 1: Introduction - ¶05-07; Sec 3.2: Temporal Refiner - ¶01] Therefore, it would have been obvious to implement the temporal refiner architecture of DVIS and apply it to both the forward and reverse embeddings to inform the fusion weights of He in view of M-FUSE to account for correct instance representations from both a forward and backward refinement direction.
Claim(s) 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over He et al (US 2025/0113087 A1), hereinafter referred to as “He” in view of Mehl et al (“M-FUSE: Multi-frame Fusion for Scene Flow Estimation”, IEEE/CVF Winter Conference on Applications of Computer Vision, 2023), hereinafter referred to as “M-FUSE”, further in view of Nishimura et al (“Weakly-Supervised Cell Tracking via Backward-and-Forward Propagation”, arXiv, 2020), hereinafter referred to as “Nishimura”.
Considering claim 8, He in view of M-FUSE teach The method of claim 7 (as described previously), wherein the plurality of features comprises a plurality of forward features corresponding to the forward time order and a plurality of backward features corresponding to the backward time order (M-FUSE: the forward and backward embeddings generated in step 2 encode features from frames in a forward and backward time order [Sec 3.2: Multi-frame Fusion Network, steps 1-2; Fig. 2 and while M-FUSE teach generating forward and backward masks, they fail to teach an explicit fusion of these masks.
In contrast, Nishimura is analogous art pertinent to the field of endeavor of the instant application and disclose a backward-and-forward mask propagation for tracking cells during live-cell microscopy. Nishimura teach
and
wherein the method further comprises:
generating a plurality of forward masks based on the classification prediction and the plurality of forward features (Nishimura: a co-detection CNN that detects cells using weak labels (i.e. a classification prediction for what is or is not considered a cell) and cell appearance features [Sec 3: Weakly-supervised cell tracking – Subsec 3.1 – 3.2], which inform the cell position likelihood map determined during forward propagation [Sec 3.3: Backward-and-Forward propagation - ¶01, 05-07; Fig. 2]);
generating a plurality of backward masks based on the classification prediction and the plurality of backward features (Nishimura: a co-detection CNN that detects cells using weak labels (i.e. a classification prediction for what is or is not considered a cell) and cell appearance features [Sec 3: Weakly-supervised cell tracking – Subsec 3.1 – 3.2], which inform the cell position likelihood map determined during backward propagation [Sec 3.3: Backward-and-Forward propagation - ¶01-04, 07; Fig. 2]); and
fusing the plurality of forward masks and the plurality of backward masks to obtain the plurality of classification masks (Nishimura: forward and backward masks are used to generate pseudo labels [Sec 1: Introduction - ¶05-06; Fig. 2]).
Nishimura further disclose that their method utilizes backward-and-forward propagation methods for analyzing the correspondence of cell positions in detection maps through weak-supervision, minimizing the need for extensive training data [Sec: Abstract; Sec 1: Introduction ¶06]. One of ordinary skill in the art would recognize the utility using the forward and backward masks generated from their respective directional features and classification predictions and use that to inform the overall video segmentation system of He in view M-FUSE without extensive training procedures.
Claim(s) 9 & 25 is/are rejected under 35 U.S.C. 103 as being unpatentable over He et al (US 2025/0113087 A1), hereinafter referred to as “He” in view of Mehl et al (“M-FUSE: Multi-frame Fusion for Scene Flow Estimation”, IEEE/CVF Winter Conference on Applications of Computer Vision, 2023), hereinafter referred to as “M-FUSE”, further in view of Chen et al (“Vision Transformer Adapter for Dense Predictions, ICLR, 2023), hereinafter referred to as “Chen”.
Regarding claim 9, He in view of M-FUSE teach The method of claim 1 (as described previously), but fail to disclose an image encoder with a visual transformer model and adapter architecture for generating multi-scale features.
Chen, per contra, is analogous art pertinent to the field of endeavor of the present application. Chen teach a ViT Adapter framework to reconstruct multi-scale features needed for dense prediction tasks. More particularly, Chen teach wherein the features are extracted using an image encoder comprising a visual transformer model, and a visual transformer adapter† (†limitations previously interpreted under 35 U.S.C. § 112(f) are mapped to their corresponding structure and function as disclosed in the specification of the current application) configured to generate multiple-scale features based on an output of the visual transformer model (Chen: a vision transformer adapter (ViT-Adapter) is proposed to work in tandem with a plain ViT model that includes a multi-scale feature extractor to reconstruct multi-scale features [Sec: Introduction - 03; Figs. 1 & 3]).
Chen further describe that their ViT adapter architecture improve dense prediction performance without pre-training the adapter on large-scale image datasets [Sec: Abstract; Fig. 1]. It would have been obvious to one of ordinary skill before the effective filing date of the present application to implement the ViT-Adapter disclosed by Chen as part of the overall image encoder taught by He in view of M-FUSE to extract multi-scale features without the need for large training datasets.
With respect to claim 25, He in view of M-FUSE teach The method of claim 21 (as described previously), but fail to disclose an image encoder with a visual transformer model and adapter architecture for generating multi-scale features.
Chen, on the other hand, teach wherein the features are extracted using an image encoder comprising a visual transformer model, and a visual transformer adapter† configured to generate multiple-scale features based on an output of the visual transformer model (Chen: a vision transformer adapter (ViT-Adapter) is proposed to work in tandem with a plain ViT model that includes a multi-scale feature extractor to reconstruct multi-scale features [Sec: Introduction - 03; Figs. 1 & 3]).
Chen further describe that their ViT adapter architecture improve dense prediction performance without pre-training the adapter on large-scale image datasets [Sec: Abstract; Fig. 1]. It would have been obvious to one of ordinary skill before the effective filing date of the present application to implement the ViT-Adapter disclosed by Chen as part of the overall image encoder taught by He in view of M-FUSE to extract multi-scale features without the need for large training datasets.
Allowable Subject Matter
Claims 3 & 23 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The primary reason for indication of allowable subject matter is that no prior art completely teach the sequence of steps outlined for performing online forward and backward refinement of query embeddings. The closest prior art, M-FUSE, generally disclose forward and backward embedding refinement, but this is conducted in an offline manner, with the full range of the video available prior to the start of refinement for global temporal context. In contrast, the language of claim 3 necessitate that this refinement is conducted completely online, i.e., instantaneously for the clips generated. Similarly, Nishimura’s forward and backward propagation is conducted completely offline. No prior art cited or made of record disclose an online forward and backward refinement of embeddings as described in claim 3.
Claims 11-14 & 16-19 are allowed. The following is a statement of reasons for the indication of allowable subject matter:
Claim 11 recites a plurality of limitations interpreted under 35 U.S.C. § 112(f), which necessitate mapping to the structure and associated algorithm for performing the claimed function as indicated in the claim interpretation section of this office action. None of the prior art cited nor made of record teach the claimed embedding module in accordance with its associated algorithms described in the specification [¶ 0037, 41, 48-55, 64, 69, 72]. More specifically, the embedding module is described as performing the online refining of embeddings as described in process 600 of Fig. 6, which conducts the online refining as described in claim 3 that was previously indicated as containing allowable subject matter.
Therefore claim 11 is allowed. Claims 12-14 & 16-19 are indicated as allowed by virtue of their dependence on claim 11.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Liu et al (US 2026/0134704 A1) disclose a panoptic segmentation system for refining query embeddings that identify class and instance associations across frames.
Yang et al (“Stacking-Based Attention Temporal Convolutional Network for Action Segmentation”, 2023, IEEE) teach an action segmentation algorithm using attention-based temporal convolution block to capture temporal dependencies across frames.
Wu et al (“SeqFormer: Sequential Transformer for Video Instance Segmentation”, arXiv, 2022) describe a foundational video instance segmentation method that aggregates temporal information for each instance in a frame to make dynamic mask predictions.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael M. Sofroniou whose telephone number is (571)272-0287. The examiner can normally be reached M-F: 8:30 AM - 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, John M. Villecco can be reached at (571) 272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL M SOFRONIOU/Examiner, Art Unit 2661
/AARON W CARTER/Primary Examiner, Art Unit 2661