Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
The United States Patent & Trademark Office appreciates the application that is submitted by the inventor/assignee. The United States Patent & Trademark Office reviewed the following application and has made the following comments below.
Priority
This application claims benefit of foreign priority under 35 U.S.C. 119(a)-(d) of KR10-2023-0008665, filed in Korea on 01/20/2023.
Preliminary Amendment
Applicant submitted amendments on 05/11/2026. The Examiner acknowledges the amendment and has reviewed the claims accordingly.
Overview
Claims 1-8 are pending in this application and have been considered below.
Claims 1-4 and 7-8 are rejected
Claims 5-6 are objected to.
Applicants Arguments:
In regards to Argument 1, Applicant/s state/s that “the Office Action objects to the specification, for the informality issues in pages 4, 5 and 16. In response, the specification is amended to address the noted informality issues. Applicant submits that there is no informality issue in the amended specification. Thus, Applicant submits that the specification objection is overcome. Accordingly, Applicant respectfully requests reconsideration and withdrawal of the specification objection.”
In regards to Argument 2, Applicant/s state/s that “the Office Action objects to claim 8, for the informality issue in claim 8. In response, claim 8 is amended to address the noted informality issue. Applicant submits that there is no informality issue in amended claim 8. Thus, Applicant submits that the claim objection is overcome. Accordingly, Applicant respectfully requests reconsideration and withdrawal of the claim objection.”
In regards to Argument 3, Applicant/s state/s “the cited references, i.e., Bear, Ter Haar Romenij, Tachella, Melas-Kyriazi, and Kanakis, considered either alone or in combination, fail to disclose, suggest, or otherwise render obvious the claimed combination of features presently set forth in independent claim 1. … including the combination of features of "generating, by a processor, a plurality of eigenvectors for an image according to an affinity matrix of the image, wherein the plurality of eigenvectors are generated by an encoder-decoder neural network model and correspond to structural features of respective regions of the image clustered according to the affinity matrix; and generating, by the processor, the scene structure by convolutioning the plurality of eigenvectors, and outputting the scene structure to the baseline network.”
In regards to Argument 4, Applicant/s state/s “Bear and Ter Haar Romenij merely disclose conventional architectures in which scene- related features or eigenvector-derived filters are used within general image analysis or CNN frameworks. Neither reference discloses or suggests the more specific framework, "the plurality of eigenvectors are generated by an encoder-decoder neural network model and correspond to structural features ofrespective regions of the image clustered according to the affinity matrix," as explicitly recited in amended independent claim 1… Thus, amended independent claim 1 distinguishes over Ter Haar Romenij's PCA-based eigenvector extraction and Bear's PSG-based scene representation.” (Remarks, page 6.)
In regards to Argument 5, Applicant/s state/s “the other cited references, i.e., Tachella, Melas-Kyriazi, and Kanakis, considered either alone or in combination, also fail to disclose the above features of amended independent claim 1, and thus, fail to remedy the deficiencies of Bear and Ter Haar Romenij.”
In regards to Argument 6, Applicant/s state/s “in view of the above, Applicant submits that amended independent claim 1 recites allowable subject matter. Dependent claims 2-4, 7 and 8 are also allowable at least for their dependency on amended independent claim 1.”
In regards to Argument 7, Applicant/s state/s “thus, Applicant submits that the rejections of claims 1-4, 7 and 8 are overcome. Accordingly, Applicant respectfully requests reconsideration and withdrawal of the rejections under 35 U.S.C. § 103.”
Examiner’s Responses
In response to Argument 1, see remarks, filed 05/11/2026, with respect to the objections to Applicant’s specification for the informality issues on pages 4, 5 and 16 has been fully considered and is persuasive. Therefore, the objections to the specification have been withdrawn.
In response to Argument 2, see remarks, filed 05/11/2026, with respect to the claim objection for claim 8 has been fully considered and is persuasive. Therefore, the claim 8 objection has been withdrawn.
In response to Argument 3, see remarks, filed 05/11/2026, regarding applicant’s argument that the cited references, i.e., Bear, ter Haar Romenij, Tachella, Melas-Kyriazi, and Kanakis, considered alone or in combination, fail to disclose, suggest, or otherwise render obvious the claimed combination of features presently set forth in independent claim 1. Specifically, "generating, by a processor, a plurality of eigenvectors for an image according to an affinity matrix of the image, wherein the plurality of eigenvectors are generated by an encoder-decoder neural network model and correspond to structural features of respective regions of the image clustered according to the affinity matrix; and generating, by the processor, the scene structure by convolutioning the plurality of eigenvectors, and outputting the scene structure to the baseline network” has been fully considered and are persuasive. Therefore, the rejection has been withdrawn due to the amendment. However, upon further consideration, a new ground(s) of rejection is made for Claim 1 under 35 U.S.C. 103 in view of Bear et al. (NPL “Learning Physical Graph Representations from Visual Scenes,” 2020, hereafter referred to as Bear) in view of ter Haar Romenij et al. (U.S. Patent No. 10713563 B2, hereafter referred to as ter Haar Romenij) in further view of Yang et al. (NPL, “Deep Spectral Clustering using Dual Autoencoder Network,” 2019, hereafter referred to as Yang).
The Examiner finds that Yang teaches on the amended claim language “wherein the plurality of eigenvectors are generated by an encoder-decoder neural network model” and ter Haar Romenij teaches the amended claim language “and correspond to structural features of respective regions of the image clustered according to the affinity matrix.”
Specifically, regarding the claim limitation “wherein the plurality of eigenvectors are generated by an encoder-decoder neural network model,” Yang teaches performing deep spectral clustering using a dual autoencoder network (1. Introduction). The unified framework consists of two main components: a dual autoencoder and a deep spectral clustering network. The dual autoencoder reconstructs the inputs using latent representations and their noise versions, to make the latent representations more robust. The mutual information estimation between inputs and latent representations is applied to preserve the input information. Next, the deep spectral clustering network is utilized to embed the latent representations into the eigenspace and subsequently, clustering is performed. The two networks are merged together into a unified framework and jointly optimized (3. Methodology). The claim does not specify that the eigenvectors specifically be generated by the encoder or decoder in the neural network model. Instead, the claim states that the eigenvectors are generated by “an encoder-decoder neural network model.” Therefore, under Broadest Reasonable Interpretation, the Examiner has interpreted the unified framework autoencoder architecture to be a “encoder-decoder neural network model” since the model includes an encoder-decoder for performing steps for embedding the latent representations into the eigenspace (generating eigenvectors).
In addition, regarding the claim limitation, “correspond to structural features of respective regions of the image clustered according to the affinity matrix,” ter Haar Romenij teaches performing linear principal component analysis of patches of the affinity matrix to produce eigenvectors (Col. 2, lines 65-67). The eigenvectors are partitioned into square patches, which form the output convolutional filters. The kernels/filters show the next conceptual perceptual groups, the parts (for faces: elements of mouths, noses, eyes) (Col. 5, lines 23-36). The output filters of this PCA on affinity (spectral clustering) stage are used to convolve the next layer, giving rise to a new set of more complex descriptive features (Col. 2, lines 8-29). The Examiner interprets the square patches (filters/kernels), which are partitioned eigenvectors, represent “structural features” of “respective regions” of the image, i.e., the “parts” of an image such as: (elements of mouths, noses, eyes). Under BRI, parts of a face in the image such as mouth, nose, eyes are “structural features” since they describe the structure of the image contents. Additionally, the Examiner interprets respective regions are the square patches of the image since the patches correspond to different respective regions of the image. Details of the rejection are below.
In response to Argument 4, see remarks, filed 05/11/2026, regarding applicant’s argument that “Bear and Ter Haar Romenij merely disclose conventional architectures in which scene- related features or eigenvector-derived filters are used within general image analysis or CNN frameworks. Neither reference discloses or suggests the more specific framework, "the plurality of eigenvectors are generated by an encoder-decoder neural network model and correspond to structural features of respective regions of the image clustered according to the affinity matrix," as explicitly recited in amended independent claim 1… Thus, amended independent claim 1 distinguishes over Ter Haar Romenij's PCA-based eigenvector extraction and Bear's PSG-based scene representation” has been fully considered and is persuasive. Therefore, the rejection has been withdrawn due to the amendment. However, upon further consideration, a new ground(s) of rejection is made for Claim 1 under 35 U.S.C. 103 in view of Bear in view of ter Haar Romenij in further view of Yang.
Examiner states see response to Argument 3 above regarding the teachings of Yang and ter Haar Romenij on the amended claim 1 language. Details of the rejection are below.
In response to Argument 5, see remarks, filed 05/11/2026, regarding applicant’s argument that “the other cited references, i.e., Tachella, Melas-Kyriazi, and Kanakis, considered either alone or in combination, also fail to disclose the above features of amended independent claim 1, and thus, fail to remedy the deficiencies of Bear and Ter Haar Romenij” has been fully considered and is persuasive. Therefore, the rejection has been withdrawn due to the amendment. However, upon further consideration, a new ground(s) of rejection is made for Claim 1 under 35 U.S.C. 103 in view of Bear in view of ter Haar Romenij in further view of Yang.
Examiner states see response to Argument 3 above regarding the teachings of Yang and ter Haar Romenij on the amended claim 1 language. Details of the rejection are below.
In response to Argument 6, see remarks, filed 05/11/2026, regarding applicant’s argument that “in view of the above, Applicant submits that amended independent claim 1 recites allowable subject matter. Dependent claims 2-4, 7 and 8 are also allowable at least for their dependency on amended independent claim 1” has been fully considered but is not persuasive. Specifically, upon further consideration, a new ground(s) of rejection due to the amendment is made in view of Bear in view of ter Haar Romenij in further view of Yang.
Specifically, regarding Claim 1, Bear teaches a self-supervised neural network architecture, called PSGNet, that learns to estimate Physical Scene Graphs (PSGs) from visual inputs, such as images (1 Introduction). PSGNet may be used for performing scene segmentation tasks (Abstract, Fig. 2) and generates depth maps (Fig. 1, “Graph Rendering”). The Examiner interprets both scene segmentation and/or depth map outputs to be “task-specific scene structures.” Further, Bear teaches providing depth maps and normal maps as inputs along with RGB images to baseline CNN models (MONet and IODINE) for performing the scene segmentation task (Page 23, Assessing the dependence of baselines on geometric feature maps). The Examiner asserts this aspect reads on the “plug-and-play” limitation since the scene structures (depth maps) are provided to the baseline CNN models to perform an image processing task (semantic segmentation). Bear does not explicitly disclose generating by a processor, a plurality of eigenvectors for an image according to an affinity matrix of the image wherein the plurality of eigenvectors are generated by an encoder-decoder neural network model and correspond to structural features of respective regions of the image clustered according to the affinity matrix and generating by a processor, the scene structure by convolutioning the plurality of eigenvectors. ter Haar Romenij is in the same field of art of performing an image processing task using a CNN. Further, ter Haar Romenij teaches generating, by a GPU (Col. 5, lines 62-67), a set of eigenvectors for an image by performing linear principal component analysis on the affinity matrix (Col. 5, lines 28-33). Further, ter Haar Romenij teaches partitioning the eigenvectors into square patches, which form the convolutional filters/kernels. The kernels show the next conceptual groups, the “parts” (regions) of the image (for example, for faces: elements of mouths, noses, eyes). The Examiner interprets elements/parts of faces to be structural features corresponding to regions of the image. Additionally, ter Haar Romenij teaches generating, by the GPU, detections for edges, corners, lines, and other primitives, which, under BRI, is interpreted as a scene structure since these detections provides structural information regarding the scene. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Bear by incorporating in the pre-processing steps, the generation of eigenvectors for an image based on pixel affinities and partitioning the eigenvectors into kernels which perform convolutions on the image and detect features such as edges, corners, and lines that is taught by ter Haar Romenij to make the invention that uses the eigenvectors according to the affinity matrix to detect complex and abstract features of the image to produce a scene structure for an image; thus one of ordinary skill in the art would have been motivated to combine the references because eigenvectors reveal the most prominent features in the images, such as edges, corners, shapes, and even objects (Col. 2, lines 20-25). Additionally, the reduction technique using PCA on eigenvectors and PCA allow the image to be projected into a lower-dimensional space while preserving the most important information using only about 5-8 eigenvectors (Col. 2, lines 1-6). This reduction technique improves efficiency and accuracy of the method (Col. 2, lines 6-7). Neither Bear nor ter Haar Romenij disclose generating the eigenvectors by an encoder-decoder neural network model. Yang is in the same field of art of performing image analysis tasks using a neural network model. Further, Yang teaches deep spectral clustering using a dual autoencoder network. Specifically, Yang teaches a unified neural network architecture which includes a pre-trained dual autoencoder to embed inputs into a latent space, and reconstruction results are obtained by the latent representations and their noise versions (3. Methodology). Latent representations are assigned to the ideal clusters by a deep spectral clustering model. A spectral clustering method is used to embed latent representations into the eigenspace of their associated graph Laplacian matrix. The dual autoencoder and spectral clustering network are jointly optimized and merged (3.2. Deep Spectral Clustering). Therefore, the Examiner interprets the unified framework to be an “encoder-decoder neural network model” since the claim does not specify that the eigenvectors specifically by generated by the encoder/decoder in the neural network model. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Bear in view of ter Haar Romenij by generating the eigenvectors using a autoencoder model that is taught by Yang to make the invention that embeds the latent representations into the eigenspace and clusters them to exploit the relationship between inputs to achieve optimal clustering results (Abstract); thus one of ordinary skill in the art would have been motivated to combine the references because autoencoders (encoder-decoder) have a powerful ability to capture high dimensional probability distributions of the inputs without needing supervised information (1. Introduction).
In response to Argument 7, see remarks, filed 05/11/2026, regarding applicant’s argument that “the rejections of claims 1-4, 7 and 8 are overcome. Accordingly, Applicant respectfully requests reconsideration and withdrawal of the rejections under 35 U.S.C. § 103” has been fully considered but is not persuasive. Specifically, upon further consideration, a new ground(s) of rejection due to the amendment is made in view of Bear in view of ter Haar Romenij in further view of Yang.
Examiner states see response to Argument 6 above regarding the teachings of Yang in view of ter Haar Romenij in further view of Yang for Claim 1 35 U.S.C. 103 rejection. Details of the rejection are found below.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action:
(a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103(a) are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim 1 is rejected under 35 U.S.C. 103(a) as being unpatentable over Bear et al. (NPL “Learning Physical Graph Representations from Visual Scenes,” 2020, hereafter referred to as Bear) in view of ter Haar Romenij et al. (U.S. Patent No. 10713563 B2, hereafter referred to as ter Haar Romenij) in further view of Yang et al. (NPL, “Deep Spectral Clustering using Dual Autoencoder Network,” 2019, hereafter referred to as Yang).
Regarding Claim 1, Bear teaches a method for generating a task-specific scene structure using a neural network model (1 Introduction, Fig. 2, Bear teaches a self-supervised neural network architecture, PSGNet, that learns to estimate PSGs (Physical Scene Graphs) from visual inputs. PSGNet may perform scene segmentation on an image dataset (See Fig. 2).) applied to a baseline network (Assessing the dependence on baselines on geometric feature maps, S3 Comparing PSGNet to Baseline Models, Bear teaches giving depth and normal maps as inputs in addition to the RGB image to MONet and IONINE, which are CNN baseline models. The Examiner interprets depth/normal (geometric feature) maps to be task-specific scene structures.) performing an image processing task by a plug-and-play scheme by using the scene structure (3 Experiments and Analysis, Fig. 2, Bear teaches giving ground truth RGB, depth, and normals maps as input to baseline networks (MONet, IODINE, and OP3 and one-non-learned baseline Quickshift++ (Q++). The baseline networks perform a scene segmentation task (see Fig. 2). The Examiner interprets a depth map to be a scene structure since depth maps contain distance information for objects in the scene (i.e., information about how the scene is structured).), and outputting the scene structure to the baseline network (Assessing the dependence of baselines on geometric feature maps, Fig. S4, Bear teaches providing depth and normal maps as inputs in addition to the RGB image to each baseline network, MONet and IODINE. PSGNet generates depth and normal maps, such as those shown in Fig. S4, which provide detailed geometric information. Under BRI, the Examiner interprets depth and normal maps to be “scene structures” because they include detailed information representing boundaries, shapes, structures, and textures of objects in images.).
Bear does not explicitly disclose the method comprising: generating, by a processor, a plurality of eigenvectors for an image according to an affinity matrix of the image, wherein the plurality of eigenvectors are generated by an encoder-decoder neural network model and correspond to structural features of respective regions of the image clustered according to the affinity matrix; and generating, by the processor, the scene structure by convolutioning the plurality of eigenvectors.
Ter Haar Romenij is in the same field of art of performing an image processing task using a neural network model. Further, ter Haar Romenij teaches the method comprising: generating, by a processor (Col. 5, lines 62-67 and Col. 6, lines 1-2, ter Haar Romenij teaches the method is realized using appropriate hardware conventionally used for implementing a CNN. For example, it may be implemented using a GPU.), a plurality of eigenvectors for an image according to an affinity matrix of the image (Col. 5, lines 28-33, ter Haar Romenij teaches performing PCA on the affinity matrix to produce a set of eigenvectors.), (Col. 5, lines 24-36, ter Haar Romenij teaches the eigenvectors are partitioned into square patches, which form the convolutional filters. These kernels are visualized as square filters, and show typically the next contextual perceptual groups, the parts (for faces: elements of mouths, noses, eyes). PCA on this affinity matrix automatically learns features or learns affinities among image pixels directly. The Examiner interprets the kernels correspond to parts of the image (elements of mouths, noses, eyes), which are being interpreted as “structural features” of the image.); and generating, by the processor (Col. 5, lines 62-67 and Col. 6, lines 1-2, ter Haar Romenij teaches the method is realized using appropriate hardware conventionally used for implementing a CNN. For example, it may be implemented using a GPU.), the scene structure by convolutioning the plurality of eigenvectors (Col. 4, lines 38-48, Col. 5, lines 30-33, ter Haar Romenij teaches partitioning eigenvectors into kernels (filters), which produce, after convolution of the image with these kernels, a rich set of features per pixel. These kernels detect edges, corners, lines, and other primitives. Under Broadest Reasonable Interpretation (BRI) in light of Applicant’s specification, the Examiner interprets the edges, corners, lines, and other primitives identified by the kernels/filters to be scene structures since Applicant’s specification states, “the scene structure…may include arbitrary features representing the texture of the image, the boundary of the object in the image, a structure, a shape, etc…. and may include all concepts expressed as … a structural feature (Applicant’s specification, Page 11, lines 2-9).” ),
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Bear by incorporating in the preprocessing steps, the generation of eigenvectors for an image based on pixel affinities (similarities) and partitioning the eigenvectors into kernels/filters which perform convolutions on the image and detect features such as edges, corners, and lines, that is taught by ter Haar Romenij, to make the invention that uses eigenvectors according to an affinity matrix to learn and detect complex and abstract features of the image including low-level features such as edges and lines, and high-level features such as shapes and arrangements of shapes to produce a scene structure for an image (ter Haar Romenij, Col. 1, lines 31-38); thus, one of ordinary skilled in the art would be motivated to combine the references to reduce the dimensionality of the data using the proposed reduction technique, which is far more accurate and efficient than existing pooling techniques (ter Haar Romenij, Col. 2, lines 6-7).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
Bear in view of ter Haar Romenij does not explicitly disclose wherein the plurality of eigenvectors are generated by an encoder-decoder neural network model.
Yang is in the same field of art of performing image analysis tasks using a neural network model. Further, Yang teaches wherein the plurality of eigenvectors are generated by an encoder-decoder neural network model (Fig. 2 caption, 3. Methodology, 3.2. Deep Spectral Clustering, Yang teaches deep spectral clustering using a dual autoencoder network. Specifically, the unified framework consists of two main components: a dual autoencoder and a deep spectral clustering network. The dual autoencoder reconstructs inputs using the latent representations and their noise versions. The mutual information estimation between the inputs and latent representations is applied to preserve the input information. Then, a deep spectral clustering network is used to embed the latent representations into the eigenspace of their associated graph Laplacian and clustering is subsequently performed. The two networks are merged into a unified framework and jointly optimized with KL divergence. The Examiner interprets the merged architecture to be an “encoder-decoder neural network model” since the claim does not specify that the eigenvectors specifically be produced by an encoder/decoder in the neural network model it only states that the “model” generates the plurality of eigenvectors, which the Examiner has interpreted to be the unified framework containing an autoencoder (encoder-decoder) which performs steps required to embed the latent representations into the eigenspace of their associated graph Laplacian.)
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Bear in view of ter Haar Romenij by generating the eigenvectors by using an architecture containing a dual autoencoder (encoder-decoder) that is taught by Yang, to make the invention that exploits the relationship between inputs to achieve optimal clustering results; thus, one of ordinary skilled in the art would be motivated to combine the references since the autoencoder captures high dimensional probability distributions of the inputs without supervised information (Yang, 1. Introduction). In addition, the features of the latent space obtained by the autoencoder are robust to noise and are more discriminative (Yang, 5. Conclusion).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
Claim 2 is rejected under 35 U.S.C. 103(a) as being unpatentable over Bear et al. (NPL “Learning Physical Graph Representations from Visual Scenes,” 2020, hereafter referred to as Bear) in view of ter Haar Romenij et al. (U.S. Patent No. 10713563 B2, hereafter referred to as ter Haar Romenij) in view of Yang et al. (NPL, “Deep Spectral Clustering using Dual Autoencoder Network,” 2019, hereafter referred to as Yang) in further view of Tachella et al. (NPL “The Neural Tangent Link Between CNN Denoisers and Non-Local Filters,” 2021, hereafter referred to as Tachella).
Regarding Claim 2, Bear in view of ter Haar Romenij in further view of Yang discloses the method of claim 1.
Bear in view of ter Haar Romenij in further view of Yang does not explicitly disclose wherein the baseline network performs at least one task of denoising, image deblurring, image super-resolution, image inpainting, or depth upsampling, and depth completion.
Tachella is in the same field of art of image processing using a neural network model. Further, Tachella teaches wherein the baseline network performs at least one task of denoising, image deblurring, image super-resolution, image inpainting, or depth upsampling, and depth completion (Fig. 1, Tachella teaches a convolutional neural network trained on a single corrupted image can perform powerful denoising.).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Bear in view of ter Haar Romenij in further view of Yang by providing the scene structure to a baseline network such as a CNN to perform denoising of an image that is taught by Tachella, to make the invention that performs image restoration tasks on an image such as denoising to improve image quality; thus, one of ordinary skilled in the art would be motivated to combine the references since CNNs are frequently used to perform denoising steps, especially in the context of plug-and-play methods (Tachella, 1. Introduction).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
Claims 3 and 4 are rejected under 35 U.S.C. 103(a) as being unpatentable over Bear et al. (NPL “Learning Physical Graph Representations from Visual Scenes,” 2020, hereafter referred to as Bear) in view of ter Haar Romenij et al. (U.S. Patent No. 10713563 B2, hereafter referred to as ter Haar Romenij) in view of Yang et al. (NPL, “Deep Spectral Clustering using Dual Autoencoder Network,” 2019, hereafter referred to as Yang) in further view of Melas-Kyriazi et al (NPL “Deep Spectral Methods: A Surprisingly Strong Baseline for Unsupervised Semantic Segmentation and Localization,” 2022, hereafter referred to as Melas-Kyriazi).
Regarding Claim 3, Bear in view of ter Haar Romenij in view of Yang discloses the method of claim 1, wherein the generating of the eigenvector includes generating an eigenvector corresponding to a structure (Col. 5, lines 23-36, ter Haar Romenij discloses producing a set of eigenvectors by performing PCA on the affinity matrix. The eigenvectors are partitioned into square patches, which form the output convolutional filters. These kernels are visualized as square filters, and show typically the next perceptual groups, the parts (for faces: elements of mouths, noses, eyes). PCA on the affinity matrix automatically learns features.) (1. Introduction, Yang teaches a dual autoencoder network for deep spectral clustering. Deep spectral clustering is used to embed the latent representations into the eigenspace, which is followed by clustering.).
Bear in view of ter Haar Romenij in view of Yang does not explicitly disclose for each region of the image clustered according to the affinity matrix.
Melas-Kyriazi is in the same field of art of decomposing an image into segments using eigenvectors. Further, Melas-Kyriazi teaches generating an eigenvector corresponding to a structure for each region of the image clustered according to the affinity matrix (Fig. 6, 3.3. Object Localization, Melas-Kyriazi discloses eigensegments are derived from the affinity matrix and correspond to semantically consistent regions. The eigenvectors of the Laplacian of a feature affinity matrix are examined. The eigenvectors already decompose an image into meaningful segments, and can be readily used to localize objects in a scene. The eigenvectors are discretized by clustering them across the eigenvector dimension using K-means clustering.)
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Bear in view of ter Haar Romenij in view of Yang by generating eigenvectors corresponding to a structure for each region of the image based on the feature affinity matrix that is taught by Melas-Kyriazi, to make the invention that decomposes the image into meaningful segments to be used for an image processing task such as semantic segmentation; thus, one of ordinary skilled in the art would be motivated to combine the references since the method is significantly simpler and outperforms state of the art methods in unsupervised segmentation and object localization (Melas-Kyriazi, Fig. 1 caption).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
Regarding Claim 4, Bear in view of ter Haar Romenij in further view of Yang discloses the method of claim 1.
Bear in view of ter Haar Romenij in further view of Yang does not explicitly disclose wherein the generating of the eigenvector includes transforming the affinity matrix, and deriving a Laplacian matrix, and generating an eigenvector which makes a value of a quadratic form of the Laplacian matrix for the eigenvector become the minimum.
Melas-Kyriazi is in the same field of art of decomposing an image into segments using eigenvectors. Further, Melas-Kyriazi teaches wherein the generating of the eigenvector includes transforming the affinity matrix (Abstract, Melas-Kyriazi discloses examining the eigenvectors of the Laplacian of a feature affinity matrix.), and deriving a Laplacian matrix (3.1. Background, Melas-Kyriazi discloses computing the Laplacian matrix L of the graph given by L=D-W or in the normalized case, where D is the diagonal matrix whose entries contain the row-wise sum of W.), and generating an eigenvector which makes a value of a quadratic form of the Laplacian matrix for the eigenvector become the minimum (3.1. Background, Melas-Kyriazi discloses the eigenvectors and eigenvalues of L are the central objects of study in spectral graph theory. The eigenvectors yi span an orthogonal basis for functions on G that is the smoothest possible orthogonal basis defined by:
PNG
media_image1.png
43
321
media_image1.png
Greyscale
xTLx represents the quadratic form (See 3.1. Background). In addition, “argmin” indicates finding the minimum value of the quadratic form to generate the eigenvector yi.).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Bear in view of ter Haar Romenij in further view of Yang by transforming the affinity matrix and deriving the Laplacian matrix that makes the value of the quadratic form of the Laplacian matrix become a minimum that is taught by Melas-Kyriazi, to make the invention that decomposes/partitions the image into meaningful segments using the eigenvectors thus, one of ordinary skilled in the art would be motivated to combine the references since existing unsupervised approaches struggle with complex scenes containing multiple objects and the proposed method (graoh partitioning problem) can easily obtain meaningful image segments and thus well-delineated, namable regions, i.e., semantic segmentations can be obtained (Abstract, Melas-Kyriazi).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
Claims 7 and 8 are rejected under 35 U.S.C. 103(a) as being unpatentable over Bear et al. (NPL “Learning Physical Graph Representations from Visual Scenes,” 2020, hereafter referred to as Bear) in view of ter Haar Romenij et al. (U.S. Patent No. 10713563 B2, hereafter referred to as ter Haar Romenij) in view of Yang et al. (NPL, “Deep Spectral Clustering using Dual Autoencoder Network,” 2019, hereafter referred to as Yang) in further view of Lin et al. (NPL “Lightweight Convolutional Neural Networks with Model-Switching Architecture for Multi-Scenario Road Semantic Segmentation,” 2021, hereafter referred to as Lin).
Regarding Claim 7, Bear in view of ter Haar Romenij in further view of Yang disclose the method of claim 1.
Bear in view of ter Haar Romenij in further view of Yang does not explicitly disclose wherein the generating of the scene structure includes generating the scene structure by inputting the plurality of eigenvectors into a single convolution layer.
Lin is in the same field of art of performing image analysis tasks, such as semantic segmentation using a neural network. Further, Lin teaches wherein the generating of the scene structure includes generating the scene structure by inputting the plurality of eigenvectors into a single convolution layer (3.2.2. Reduction in Convolutional Layers, Lin teaches reducing the number of convolutional layers in the CNN as shown in Fig. 4. As shown in Fig. 4, the 5x5 layer is reduced to a 1x1 feature layer.).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Bear in view of ter Haar Romenij in further view of Yang by reducing the number of layers in the CNN to a single convolution layer that is taught by Lin, to make the invention that more efficiently convolutions the plurality of eigenvectors (patches/kernels) to generate the scene structure; thus, one of ordinary skilled in the art would be motivated to combine the references since reducing the number of layers in the CNN lowers the weight and the amount of calculation required, and increases the execution speed of the CNN (Lin, 3.2.2. Reduction in Convolutional Layers).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
Regarding Claim 8, Bear in view of ter Haar Romenij in view of Yang in further view of Lin discloses the method claim 7, wherein the single convolution layer (3.2.2. Reduction in Convolution, Lin teaches reducing the layers in the CNN to a single 1x1 convolution layer.) convolutions the plurality of eigenvectors (Col. 4, lines 44-48, ter Haar Romenij teaches eigenvectors are partitioned into δxδ kernels, or filters, giving, after convolution of the image with these kernels, a rich set of features per pixel. These kernels detect edges, corners, lines, and other priitives. They are δxδ patches.) based on a weight learned according to a task of the baseline network (3.1 Model-Switching Architecture, 5. Conclusions, Lin teaches a model-switching architecture to eliminate the mutual suppression of weights. The architecture uses multiple CNN models to store the weights of various states individually and uses one or more CNN classifiers to identify the current state and switch to a suitable model for semantic segmentation. The Examiner interprets “semantic segmentation” to be a specific task of the CNN/baseline network.).
Allowable Subject Matter
Claims 5 and 6 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claims 5 and 6 were previously indicated as containing allowable subject matter in the previous office action, dated 02/11/2026.
Found below are reasons for indication of allowable subject matter:
Regarding Claim 5, no prior art teaches wherein the generating of the eigenvector includes generating an eigenvector which makes a loss function expressed by [Equation 1] below become the minimum.
PNG
media_image2.png
65
271
media_image2.png
Greyscale
[Equation 1]
(Where Y represents the eigenvector, k represents a channel of the eigenvector, and L represents the Laplacian matrix).
Regarding Claim 6, no prior art teaches wherein the generating of the eigenvector includes generating an eigenvector which makes linear combination of two loss functions expressed by [Equation 1] above and [Equation 2] below become the minimum.
PNG
media_image3.png
68
512
media_image3.png
Greyscale
[Equation 2]
(Where y is a hyperparameter.)
Pertinent Prior Art
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Shaham et al. (NPL,“SpectralNet: Spectral Clustering using Deep Neural Networks,” 2018) teaches a deep learning approach to spectral clustering that overcomes scalability and generalization limitations. The network, called SpectralNet, learns a map that embeds input data points into an eigenspace of their associated Laplacian matrix and subsequently clusters them. Once trained, SpectralNet provides a parametric function whose image for the training points is approximately the eigenvectors of the graph Laplacian. In addition, Shaham suggests improvement can be achieved by applying the network to code representations produced, e.g., by standard autoencoders (encoder-decoder network).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SYDNEY L BLACKSTEN whose telephone number is (571)272-7120. The examiner can normally be reached 8:30am-4:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Oneal Mistry can be reached at 313-446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SYDNEY L BLACKSTEN/Examiner, Art Unit 2674
/ONEAL R MISTRY/Supervisory Patent Examiner, Art Unit 2674