Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier.
Such claim limitation(s) is/are: “a video encoding module”, “an eigen clustering module”, “an object-centric contrastive learning module” in claim 1 and dependent claims 2-10.
See para. 133 “Referring to FIG. 3, the device for object-centric representation learning through unsupervised semantic segmentation 100 may include a processor 210, a memory 230, a user input and output unit 250, a network input and output unit 270, and a communication port unit 290.”
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 4-8, 11 are rejected under 35 U.S.C. 103 as being unpatentable over Salehi et al. (WO 2024102510 A1), hereinafter Salehi, in view of Melas-Kyriazi et al.: "Deep Spectral Methods: A Surprisingly Strong Baseline for Unsupervised Semantic Segmentation and Localization", CVF Open Access, Submitted June 2022 [retrieved on 7-21-2026]. Retrieved from the internet <https://openaccess.thecvf.com/content/CVPR2022/html/Melas-Kyriazi_Deep_Spectral_Methods_A_Surprisingly_Strong_Baseline_for_Unsupervised_Semantic_CVPR_2022_paper.html>, hereinafter Melas-Kyriazi, and Gansbeke et al.: "Unsupervised Semantic Segmentation by Contrasting Object Mask Proposals", CVF Open Access, Submitted October 2021 [retrieved on 7-21-2026]. Retrieved from the internet <https://openaccess.thecvf.com/content/ICCV2021/html/Van_Gansbeke_Unsupervised_Semantic_Segmentation_by_Contrasting_Object_Mask_Proposals_ICCV_2021_paper.html>, hereinafter Gansbeke.
Regarding claim 1, Salehi teaches A device for object-centric representation learning through unsupervised semantic segmentation, the object-centric representation learning device comprising: (Para. 21 see "Segmentation results can include one or more segmentation masks generated to indicate one or more locations, areas, and/or pixels within a frame of image data that belong to a given semantic segment (e.g., a particular object, class of objects, etc.). For example, each pixel of a segmentation mask can include a value indicating a particular semantic segment (e.g., a particular object, class of objects, etc.) to which each pixel belongs." Para. 29 see "the systems and techniques can be used to perform unsupervised semantic segmentation based on using temporally-propagated cluster maps." Para. 49 see "the systems and techniques can be used to perform unsupervised semantic segmentation based on using temporally-propagated cluster maps of similar patch representations."). a video encoding module configured to receive an input video and generate a feature map; (Para. 17 see "FIG. 5 is a flow diagram illustrating an example of a process for processing image and/or video data." Para. 79 see "the systems and techniques can use a self-supervised clustering approach on the patch representations of frames (e.g., instead of the image-level representations) to generate a cluster map for each image." Para. 131 see "process, using a machine learning model, a source image of the image data to generate a first set of features for the source image; process, using the machine learning model, a target image to generate a second set of features for the target image." Examiner note: This does not explicitly use the term 'feature map' but teaches the same concept as a set of features.).
While Salehi teaches unsupervised semantic segmentation to separate objects in an image or video by generating a cluster map using image patches and image features, Salehi does not teach an eigen clustering module configured to calculate an eigenvector representing a semantic structure of patches in the input video based on color affinity and semantic similarity of the input video, and generate a patch cluster for the patches in the input video through the eigenvector; and an object-centric contrastive learning module configured to generate an object prototype based on the patch cluster and distinguish objects in the input video through semantic coherence based on the contrastive learning for the object prototype.
However, Melas-Kyriazi teaches an eigen clustering module configured to calculate an eigenvector representing a semantic structure of patches in the input image (Abstract see "we examine the eigenvectors of the Laplacian of a feature affinity matrix from self-supervised networks. We find that these eigenvectors already decompose an image into meaningful segments." Pg. 2, Col. 1, Para. 3 see "the eigenvectors of the Laplacian of this graph directly correspond to semantically meaningful image regions."). based on color affinity and semantic similarity of the input image, (Pg. 1, Fig. 1 see "We present a simple approach based on spectral methods that decomposes an image using the eigenvectors of a Laplacian matrix constructed from a combination of color information and unsupervised deep features." Pg. 2, Col. 1, Para. 3 see "We then construct a weighted graph over patches, where edge weights give the semantic affinity of pairs of patches."). and generate a patch cluster for the patches in the input image through the eigenvector; (Abstract see "by clustering the features associated with these segments across a dataset, we can obtain well-delineated, nameable regions, i.e. semantic segmentations." Pg. 2, Col. 1, Para. 3 see "the eigenvectors of the Laplacian of this graph directly correspond to semantically meaningful image regions." Pg. 2, Col. 1, Para. 4 see "We first convert the eigensegments into discrete image regions by thresholding and associate each region with a semantic feature vector from the network." Pg. 6, Col. 1, Para. 1 see "we discretize the first m eigenvectors {y1,··· ,ym} of L by clustering them across the eigenvector dimension using K-means clustering (for every image separately).").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Salehi to incorporate the teachings of Melas-Kyriazi to calculate an eigenvector representing a semantic structure of patches in the input video based on color affinity and semantic similarity of the input video and generate a patch cluster for the patches in the input video through the eigenvector. Doing so would predictably make the patch clusters more accurate and better at including pixels of actual objects by using color affinity and semantic similarity.
Furthermore, Gansbeke teaches and an object-centric contrastive learning module configured to generate an object prototype based on the patch cluster (Pg. 3, Col. 1, Para. 3 see "we pull pixels belonging to the same object together, and contrast them against pixels from other objects." Pg. 5, Col. 1, Para. 4 see "We modify the contrastive loss from Equation 1 to include the proposed pull and push-forces. Positive pairs of object-centric... The pixel embedding function Φθ maximizes the agreement between pixels and an augmented view of the object they belong to, while minimizing the agreement with other objects."). and distinguish objects in the input image through semantic coherence based on the contrastive learning for the object prototype. (Abstract see "the learned pixel embeddings can be directly clustered in semantic groups using K-Means." Pg. 3, Col. 1, Para. 3 see "we pull pixels belonging to the same object together, and contrast them against pixels from other objects, as shown in Figure 2. This forces the model to map pixels from visually similar objects closer together, while pushing pixels from dissimilar objects further apart.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Salehi and Melas-Kyriazi to incorporate the teachings of Gansbeke to use object-centric contrastive learning to generate an object prototype based on the patch cluster and distinguish objects in the input video through semantic coherence based on the contrastive learning for the object prototype. Doing so would predictably separate objects more accurately by clearly separating pixels that belong to an object from other pixels that belong to other objects.
Regarding claim 4, Salehi in view of Melas-Kyriazi and Gansbeke teaches The device for object-centric representation learning through unsupervised semantic segmentation of claim 1.
While Salehi teaches unsupervised semantic segmentation to separate objects in an image or video by generating a cluster map using image patches and image features, Salehi does not teach wherein the eigen clustering module segments the input video into patch units and calculates color affinity based on color information of each of the patches to generate a color affinity matrix.
However, Melas-Kyriazi teaches wherein the eigen clustering module segments the input image into patch units and calculates color affinity based on color information of each of the patches to generate a color affinity matrix. (Pg. 1, Fig. 1 see "We present a simple approach based on spectral methods that decomposes an image using the eigenvectors of a Laplacian matrix constructed from a combination of color information and unsupervised deep features." Pg. 2, Col. 1, Para. 3 see "Our method first utilizes a self-supervised network to extract dense features corresponding to image patches. We then construct a weighted graph over patches, where edge weights give the semantic affinity of pairs of patches.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Salehi and Melas-Kyriazi and Gansbeke to incorporate the teachings of Melas-Kyriazi to segment the video into patches and calculate color affinity to generate a color affinity matrix. Doing so would predictably make the patch clusters more accurate and better at including pixels of actual objects by using color affinity.
Regarding claim 5, Salehi in view of Gansbeke and Melas-Kyriazi teaches The device for object-centric representation learning through unsupervised semantic segmentation of claim 4.
While Salehi teaches unsupervised semantic segmentation to separate objects in an image or video by generating a cluster map using image patches and image features, Salehi does not teach wherein the Eigen clustering module performs an inner product between the patches on the feature map to generate a semantic similarity matrix indicating how semantically similar the respective patches are.
However, Melas-Kyriazi teaches wherein the Eigen clustering module performs an inner product between the patches on the feature map to generate a semantic similarity matrix indicating how semantically similar the respective patches are. (Abstract see "Specifically, we examine the eigenvectors of the Laplacian of a feature affinity matrix from self-supervised networks." Pg. 1, Fig. 1 see "We present a simple approach based on spectral methods that decomposes an image using the eigenvectors of a Laplacian matrix constructed from a combination of color information and unsupervised deep features." Pg. 2, Col. 1, Para. 3 see "Our method first utilizes a self-supervised network to extract dense features corresponding to image patches. We then construct a weighted graph over patches, where edge weights give the semantic affinity of pairs of patches." Pg. 2, Col. 2, Para. 3 see "Since self-attention involves self-comparisons of image patch features, it is natural to construct a semantic affinity matrix over patches, as we do in our approach." Examiner note: Edge weights are used to determine the semantic affinity of pairs of patches and therefore calculates an inner product.).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Salehi and Gansbeke and Melas-Kyriazi to incorporate the teachings of Melas-Kyriazi to perform inner product between the patches on the feature map to generate a semantic similarity matrix indicating how semantically similar the respective patches are. Doing so would predictably make the patch clusters more accurate and better at including pixels of actual objects by using semantic similarity to characterize objects.
Regarding claim 6, Salehi in view of Gansbeke and Melas-Kyriazi teaches The device for object-centric representation learning through unsupervised semantic segmentation of claim 5.
While Salehi teaches unsupervised semantic segmentation to separate objects in an image or video by generating a cluster map using image patches and image features, Salehi does not teach wherein the Eigen clustering module merges the color affinity matrix and the semantic similarity matrix to generate a Laplacian matrix, and eigendecomposes the Laplacian matrix to calculate the eigenvector.
However, Melas-Kyriazi teaches wherein the Eigen clustering module merges the color affinity matrix and the semantic similarity matrix to generate a Laplacian matrix, and eigendecomposes the Laplacian matrix to calculate the eigenvector. (Abstract see "we examine the eigenvectors of the Laplacian of a feature affinity matrix from self-supervised networks. We find that these eigenvectors already decompose an image into meaningful segments." Pg. 1, Fig. 1 see "We present a simple approach based on spectral methods that decomposes an image using the eigenvectors of a Laplacian matrix constructed from a combination of color information and unsupervised deep features." Pg. 2, Col. 1, Para. 3 see "We then construct a weighted graph over patches, where edge weights give the semantic affinity of pairs of patches... the eigenvectors of the Laplacian of this graph directly correspond to semantically meaningful image regions.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Salehi and Gansbeke and Melas-Kyriazi to incorporate the teachings of Melas-Kyriazi to merge the color affinity matrix and the semantic similarity matrix to generate a Laplacian matrix, and eigendecomposes the Laplacian matrix to calculate the eigenvector. Doing so would predictably make the patch clusters more accurate and better at including pixels of actual objects by merging the color affinity and semantic similarity to characterize objects.
Regarding claim 7, Salehi in view of Gansbeke and Melas-Kyriazi teaches The device for object-centric representation learning through unsupervised semantic segmentation of claim 6.
While Salehi teaches applying K-Means on the representation of the given pre-trained model to produce a cluster map for each input data (Para. 88), Salehi does not teach wherein the Eigen clustering module performs K-means clustering for the patches in the input video through the eigenvector and classifies similar patches into the same object to generate the patch cluster (EiCue).
However, Melas-Kyriazi teaches wherein the Eigen clustering module performs K-means clustering for the patches in the input image through the eigenvector and classifies similar patches into the same object to generate the patch cluster (EiCue). (Abstract see "by clustering the features associated with these segments across a dataset, we can obtain well-delineated, nameable regions, i.e. semantic segmentations." Pg. 2, Col. 1, Para. 3 see "the eigenvectors of the Laplacian of this graph directly correspond to semantically meaningful image regions." Pg. 6, Col. 1, Para. 1 see "we discretize the first m eigenvectors {y1,··· ,ym} of L by clustering them across the eigenvector dimension using K-means clustering (for every image separately).").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Salehi and Gansbeke and Melas-Kyriazi to incorporate the teachings of Melas-Kyriazi to perform K-means clustering for the patches in the input video through the eigenvector and classify similar patches into the same object to generate the patch cluster. Doing so would predictably make the patch clusters more accurate and better at including pixels of distinct objects by using the eigenvectors that already breaks the image into meaningful regions as opposed to clustering raw features.
Regarding claim 8, Salehi in view of Melas-Kyriazi and Gansbeke teaches The device for object-centric representation learning through unsupervised semantic segmentation of claim 1.
While Salehi teaches determining an average between the Intersection over union of the segmented objects over all the video frames (Para. 88), Salehi does not teach wherein the object-centric contrastive learning module selects a center vector from the patch cluster or calculates a mean vector to determine the object prototype.
However, Gansbeke teaches wherein the object-centric contrastive learning module selects a center vector from the patch cluster or calculates a mean vector to determine the object prototype. (Pg. 3, Col. 1, Para. 3 see "we pull pixels belonging to the same object together, and contrast them against pixels from other objects." Pg. 5, Col. 1, Para. 1 see "let the mean pixel embedding zMn of an object mask Mn be defined aszMn = 1|Mn| i∈Mnzi." Pg. 5, Col. 1, Para. 4 "Positive pairs of object-centric crops (Ψη(X),Ψη(X+)) are replaced with positive pairs of pixel embeddings: (zi,zMX+) for i ∈ MX. In a similar way, the negative pairs (Ψη(X),Ψη(X−k)) are replaced with (zi, zMX−k). We obtain the following optimization criterion for a pixel i ∈ MX.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Salehi and Melas-Kyriazi and Gansbeke to incorporate the teachings of Gansbeke to calculate a mean vector to determine the object prototype. Doing so would predictably separate objects more accurately by clearly separating pixels that belong to an object from other pixels that belong to other objects.
Claim 11 is rejected under the same analysis as claim 1 above.
Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Salehi et al. (WO 2024102510 A1), hereinafter Salehi, in view of Melas-Kyriazi et al.: "Deep Spectral Methods: A Surprisingly Strong Baseline for Unsupervised Semantic Segmentation and Localization", CVF Open Access, Submitted June 2022 [retrieved on 7-21-2026]. Retrieved from the internet <https://openaccess.thecvf.com/content/CVPR2022/html/Melas-Kyriazi_Deep_Spectral_Methods_A_Surprisingly_Strong_Baseline_for_Unsupervised_Semantic_CVPR_2022_paper.html>, hereinafter Melas-Kyriazi, and Gansbeke et al.: "Unsupervised Semantic Segmentation by Contrasting Object Mask Proposals", CVF Open Access, Submitted October 2021 [retrieved on 7-21-2026]. Retrieved from the internet <https://openaccess.thecvf.com/content/ICCV2021/html/Van_Gansbeke_Unsupervised_Semantic_Segmentation_by_Contrasting_Object_Mask_Proposals_ICCV_2021_paper.html>, hereinafter Gansbeke, and Ke et al.: "Unsupervised Hierarchical Semantic Segmentation with Multiview Cosegmentation and Clustering Transformers", arxiv.org, Submitted 25 April 2022, [retrieved on 7-22-2026]. Retrieved from the internet <https://arxiv.org/abs/2204.11432>, hereinafter Ke.
Regarding claim 2, Salehi in view of Melas-Kyriazi and Gansbeke teaches The device for object-centric representation learning through unsupervised semantic segmentation of claim 1.
In addition, Salehi teaches wherein the video encoding module receives an original video (Para. 17 see "FIG. 5 is a flow diagram illustrating an example of a process for processing image and/or video data." Para. 79 see "the systems and techniques can use a self-supervised clustering approach on the patch representations of frames (e.g., instead of the image-level representations) to generate a cluster map for each image." Para. 131 see "process, using a machine learning model, a source image of the image data to generate a first set of features for the source image; process, using the machine learning model, a target image to generate a second set of features for the target image."). obtained by transforming the original video through a vision transformer (ViT) as input videos. (Para. 30 see "the systems and techniques can address a dense image segmentation task. In some aspects, one or more pre-trained vision transformers (ViTs) can be utilized. ViTs can be used to maintain the spatial relationship of input patches in the final patch representations." Para. 33 see "the one or more patch representations can be obtained from and/or generated by a pre-trained machine learning model, such as a ViT and/or ViT-based machine learning model.").
While Salehi teaches transforming a video using ViT, Salehi does not teach receiving the original video and a transformed video.
However, Ke teaches receiving an original image and a transformed image (Abstract see "We enforce spatial consistency of grouping and bootstrap feature learning with co-segmentation among multiple views of the same image." Pg. 2, Col. 2, Para. 1 see "given the pixel-wise feature, we perform hierarchical groupings within and across images and their transformed versions (i.e.,views)." Pg. 6, Fig. 5 see "Starting with input X0 of an image and its augmented views, we conduct feature clustering to merge G0 into G1,and then, G1 into G2.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Salehi and Melas-Kyriazi and Gansbeke to incorporate the teachings of Ke to feed both the original video and a transformed version of it into the video encoding module that uses a vision transformer. Doing so would predictably make the features more robust and accurate by giving the model different views of the same scene with different appearances.
Claims 9-10 are rejected under 35 U.S.C. 103 as being unpatentable over Salehi et al. (WO 2024102510 A1), hereinafter Salehi, in view of Melas-Kyriazi et al.: "Deep Spectral Methods: A Surprisingly Strong Baseline for Unsupervised Semantic Segmentation and Localization", CVF Open Access, Submitted June 2022 [retrieved on 7-21-2026]. Retrieved from the internet <https://openaccess.thecvf.com/content/CVPR2022/html/Melas-Kyriazi_Deep_Spectral_Methods_A_Surprisingly_Strong_Baseline_for_Unsupervised_Semantic_CVPR_2022_paper.html>, hereinafter Melas-Kyriazi, and Gansbeke et al.: "Unsupervised Semantic Segmentation by Contrasting Object Mask Proposals", CVF Open Access, Submitted October 2021 [retrieved on 7-21-2026]. Retrieved from the internet <https://openaccess.thecvf.com/content/ICCV2021/html/Van_Gansbeke_Unsupervised_Semantic_Segmentation_by_Contrasting_Object_Mask_Proposals_ICCV_2021_paper.html>, hereinafter Gansbeke, and Araslanov et al.: "Dense Unsupervised Learning for Video Segmentation", arxiv.org, Submitted 11 November 2021 [retrieved on 7-23-2026]. Retrieved from the internet <https://arxiv.org/abs/2111.06265>, hereinafter Araslanov.
Regarding claim 9, Salehi in view of Melas-Kyriazi and Gansbeke teaches The device for object-centric representation learning through unsupervised semantic segmentation of claim 8.
While Salehi teaches performing intra-video object-centric learning (Para. 33 see “the one or more patch representations can be obtained from and/or generated by a pre-trained machine learning model, such as a ViT and/or ViT-based machine learning model. Based on tracking patch locations in the local windows of temporally close frames (e.g., adjacent frames in time, etc.), different object views can be detected across time.”), Salehi does not teach wherein the object-centric contrastive learning module performs intra-video contrastive learning and inter-video contrastive learning for the object prototype to learn semantic coherence of the object.
However, Araslanov teaches wherein the object-centric contrastive learning module performs intra-video contrastive learning and inter-video contrastive learning for the object prototype to learn semantic coherence of the object. (Abstract see "We present a novel approach to unsupervised learning for video object segmentation (VOS)... We rely on uniform grid sampling to extract a set of anchors and train our model to disambiguate between them on both inter- and intra-video levels." Pg. 1, Para. 3 see "Using a contrastive formulation [12], our approach learns to represent the temporally proximate frames to the reference in terms of these anchors." Pg. 3, Para. 5 see "We further assume that the semantic content of video clips remains unchanged, at least for a short time span. Specifically, if we represent a given reference frame with a set of distinct features (following Assumption 1), the semantic content of temporally close frames can be faithfully represented with the same feature set." Pg. 6, Para. 2 see "In Eq. (1), we compute the affinity of the features to the anchors extracted from multiple videos in the training batch. Note that we select the dominant anchors in Eq. (2) for self-training only from the same video sequence as the feature itself. This implies that (i) the features will be attracted only to the anchors originating from the same video, and (ii) the distance between the anchors and the features from different video sequences will increase by virtue of our contrastive formulation of the affinity.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Salehi and Melas-Kyriazi and Gansbeke to incorporate the teachings of Araslanov to perform intra-video and inter-video contrastive learning for the object prototype to learn semantic coherence of the object. Doing so would predictably improve the consistency of object representations across frames and videos by attracting temporally related features to the same object anchors while discriminating features from different video sequences.
Regarding claim 10, Salehi in view of Gansbeke and Melas-Kyriazi and Araslanov teaches The device for object-centric representation learning through unsupervised semantic segmentation of claim 9.
While Salehi teaches learning semantic distinction of objects, Salehi does not teach wherein the object-centric contrastive learning module learns semantic distinction of the objects through contrastive learning between patch clusters.
However, Gansbeke teaches wherein the object-centric contrastive learning module learns semantic distinction of the objects through contrastive learning between patch clusters. (Abstract see "the learned pixel embeddings can be directly clustered in semantic groups using K-Means." Pg. 3, Col. 1, Para. 3 see "we pull pixels belonging to the same object together, and contrast them against pixels from other objects, as shown in Figure 2. This forces the model to map pixels from visually similar objects closer together, while pushing pixels from dissimilar objects further apart." Pg. 5, Col. 1, Para. 4 see "We modify the contrastive loss from Equation 1 to include the proposed pull and push-forces. Positive pairs of object-centric... The pixel embedding function Φθ maximizes the agreement between pixels and an augmented view of the object they belong to, while minimizing the agreement with other objects.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Salehi and Gansbeke and Melas-Kyriazi and Araslanov to incorporate the teachings of Gansbeke to learn semantic distinction of the objects through contrastive learning between patch clusters. Doing so would predictably separate different objects more accurately by pushing the representations of dissimilar patch clusters farther apart in the embedding space while maintaining attraction within each cluster.
Allowable Subject Matter
Claim(s) 3 is/are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding claim 3, none of Salehi, Melas-Kyriazi et al., Gansbeke et al., Ke et al., and Araslanov et al. teach integrating key features of different layers from the original video and the transformed video.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Chen et al. (US 20230154139 A1) discloses an intelligent method to select instances, by utilizing unsupervised tracking for videos.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEXANDER VAUGHN whose telephone number is (571) 272-5253. The examiner can normally be reached M-F 11am-7pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JENNIFER MEHMOOD can be reached on (571) 272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALEXANDER VAUGHN/Examiner, Art Unit 2675
/JENNIFER MEHMOOD/Supervisory Patent Examiner, Art Unit 2664