DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 12/05/2025 is/are compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claim(s) 1, 9, 12, and 20 is/are objected to because of the following informalities:
In claim 1, line(s) 5-6, “object in the image using to obtain a segmentation mask […]” should read “object in the image .
In claim 9, line(s) 3, “[…] ML model is performed by a using a trained ML […]” should read “[…] ML model is performed .
In claim 12, line(s) 4, “wherein the second trained ML was […]” should read “wherein the second trained ML model was […]”.
In claim 20, line(s) 6, “[…] ML model is neural network model […]” should read “[…] ML model is a neural network model […]”.
Appropriate correction is required.
Office Action Summary
Claim(s) 8 and 18-20 is/are rejected under 35 U.S.C. 112(b).
Claim(s) 1, 4-6, 10, 12-13, and 17 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by He et al (US 10,713,794 B1).
Claim(s) 2-3 and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over He et al (US 10,713,794 B1) in view of Wang et al (CN 103488988 A; See translation provided by Examiner).
Claim(s) 7-8, 11, and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over He et al (US 10,713,794 B1) in view of Bajpai et al (US 2023/0072400 A1).
Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over He et al (US 10,713,794 B1) in view of Sun et al (US 2021/0397966 A1).
Claim(s) 14 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lee et al (US 2019/0311202 A1) in view of Bajpai et al (US 2023/0072400 A1).
Claim(s) 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lee et al (US 2019/0311202 A1) in view of Bajpai et al (US 2023/0072400 A1), further in view of Chidlovskii (US 2021/0174513 A1).
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim(s) 8 and 18-20 is/are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding claim 8, the claim recites “wherein the at least some parameter values of the first trained ML model were estimated by training on image data including multiple types of objects of a different type than the object's type.” However, claim 8 depends directly from claim 1, and claim 1 does not introduce or otherwise provide antecedent basis for “the at least some parameter values of the first trained ML model.” Although claim 7 recites “at least some parameter values of the first trained ML model,” claim 8 does not depend from claim 7 and therefore does not incorporate that limitation. Accordingly, it is unclear what previously recited subject matter is intended to be referenced by “the at least some parameter values of the first trained ML model,” thereby rendering the scope of claim 8 indefinite.
Regarding claim 18, the claim recites “wherein the at least some parameter values of the first encoder were estimated by training on image data including multiple types of objects of a different type than the object's type.” However, claim 18 depends directly from claim 14, and claim 14 does not introduce or otherwise provide antecedent basis for “the at least some parameter values of the first encoder.” Claim 14 recites “a first encoder whose weights are obtained by transfer learning,” but does not recite “at least some parameter values of the first encoder.” Although claim 16 recites “parameter values of the first encoder,” claim 18 does not depend from claim 16 and therefore does not incorporate that limitation. Accordingly, it is unclear what previously recited subject matter is intended to be referenced by “the at least some parameter values of the first encoder,” thereby rendering the scope of claim 18 indefinite.
Regarding claim 19, the claim recites “wherein dimensionality and/or shape of the transform space of the trained AE model or the trained VAE model matches dimensionality and/or shape of the second trained ML model.” It is unclear what is meant by the “dimensionality and/or shape of the second trained ML model.” A trained ML model may have different dimensionalities and/or shapes associated with, for example, its input, one or more intermediate layers or feature spaces, and its output. Claim 19 does not specify which dimensionality and/or shape associated with the second trained ML model is intended to be matched by the transform space of the trained AE model or trained VAE model. Accordingly, the metes and bounds of claim 19 cannot be determined with reasonable certainty.
It is noted that claim 9 recites a corresponding limitation in which the dimensionality and/or shape of the transform space “matches dimensionality and/or shape of input to second trained ML model.” However, claim 19 does not recite that the dimensionality and/or shape being matched is that of the input to the second trained ML model. Therefore, it is unclear from the language of claim 19 which dimensionality and/or shape associated with the second trained ML model is intended to be matched by the transform space.
Regarding Claim 20 is rejected under 35 U.S.C. 112(b) for at least the same reasons set forth above with respect to claim 19, from which claim 20 depends. Claim 20 does not cure the indefiniteness of claim 19 concerning the recited “dimensionality and/or shape of the second trained ML model.”
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1, 4-6, 10, 12-13, and 17 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by He et al (US 10,713,794 B1).
Regarding claim(s) 1 and 13, He teaches a system, comprising:
at least one computer hardware processor (Figure 17; and Col. 40, lines 56-59); and
at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform a method for detecting objects in images using machine learning (ML) (Figure 17; and Col. 40, lines 56-59), the method comprising:
obtaining an image of an object (Figure 14B: Access image; and Col. 37, lines 55-56: “At step 1482, a system using the trained model may access an image of interest”); and
segmenting the object in the image from background of the object in the image to obtain a segmentation mask for the image (Figure 3B; Figure 4; Col. 21, lines 47-49: “the system identifies objects in the forefront, which are differentiated from less important objects in the background”; Col. 22, lines 7-8: “an object proposal 430 (e.g., a binary map that identifies the pixels in the image that correspond to the object)”; and Col. 29, lines 12-21: “a machine-learning model may be trained to detect object instances depicted in an image, classify the detected object instances, and/or segment the object instances from the image […] identify particular pixels that correspond to the object instance”), the segmenting comprising:
processing the image with a first trained ML model to obtain first image features (Figure 5; Figure 14B; Col. 37, lines 56-58: “At step 1484, the system may generate a feature map for the input image (e.g., using ResNet, FPN, or any other suitable network)”; and Col. 23, lines 15-24: “the system may have a feature-extraction convolutional neural network (e.g., the first convolutional neural network 510) that may take as inputs patches of images 410 and output features 520 of the patch/image. The features 520 may be represented by a feature map that encodes various features of the image 410 […] the feature-extraction network 510 may be pre-trained to perform classification on the image. The feature-extraction model 510 may be fine-tuned for object proposals during training”);
transforming the first image features to second image features for processing by a second trained ML model (Figure 14B; Col. 36, lines 9-10: “At step 1440, the system may generate a regional feature map using RoIAlign […] the regional feature map may be generated based on sampling locations (e.g., 1241-1244 shown in FIG. 12B) defined by a sampling region (e.g., 1220 shown in FIG. 12B)”; Col. 36, lines 37-39: “The feature values associated with the bins may be used to generate the regional feature map for the RoI”; and Col. 34, lines 20-23: “for each sampling point, RolAlign may compute the value for that sampling point by bilinear interpolation from the nearby grid points on the feature map”); and
processing the second image features using the second trained ML model to obtain the segmentation mask for the image (Col. 36, lines 52-56: “the system may generate instance segmentation masks associated with the RoI by processing the regional feature map using a second neural network, such as a fully convolutional network”; and Col. 37, lines 47-50: “once trained, the neural network of the mask branch (e.g., the fully convolutional network) would be configured to generate instance segmentation masks for object instances depicted in images”),
wherein the segmentation mask identifies pixels associated with the object in the image and pixels associated with background of the object in the image (Figure 14; and Col. 36, lines 56-58: “an instance segmentation mask may be a pixel-wise binary classification of whether each pixel belongs to the instance or not”).
Regarding claim(s) 4, He teaches the method of the claim 1, wherein processing the image with the first trained ML model to obtain first image features comprises encoding the image using the first trained ML model to obtain the first image features, wherein the first image features comprise feature maps determined by processing the image using the first trained ML model (Figure 5; and Col. 23, lines 15-31: “the system may have a feature-extraction convolutional neural network (e.g., the first convolutional neural network 510) that may take as inputs patches of images 410 and output features 520 of the patch/image. The features 520 may be represented by a feature map that encodes various features of the image 410 […] the feature-extraction network 510 may be pre-trained to perform classification on the image. The feature-extraction model 510 may be fine-tuned for object proposals during training […] the feature-extraction layers may take an input image of dimension 3 x h x w, and the output 520 may be a feature map of dimensions 512 x h/16 x w/16”).
Regarding claim(s) 5, He teaches the method of claim 1, wherein the first trained ML model is a neural network model (Figure 5; and Col. 23, lines 15-31: “the system may have a feature-extraction convolutional neural network (e.g., the first convolutional neural network 510) that may take as inputs patches of images 410 and output features 520 of the patch/image. The features 520 may be represented by a feature map that encodes various features of the image 410 […] the feature-extraction network 510 may be pre-trained to perform classification on the image. The feature-extraction model 510 may be fine-tuned for object proposals during training”).
Regarding claim(s) 6 and 17, He teaches the method of claim 5, wherein the first trained ML model is a portion of a trained neural network model having a ResNet architecture, a visual geometry group (VGG) architecture, or an Inception architecture (Figure 9; Col. 29, lines 59-62: “A deep network, such as ResNet, Feature Pyramid Network (FPN), or any other suitable convolutional backbone, may process the input image 910 and generate a feature map 920”; and Col. 37, lines 56-58: “At step 1484, the system may generate a feature map for the input image (e.g., using ResNet, FPN, or any other suitable network)”).
Regarding claim(s) 10, He teaches the method of claim 1,
wherein the second trained ML model is trained to segment, from background, objects of a same type as the type of the object (Col. 32, lines 15-20: “the mask branch may have a Km2-dimensional output for each RoI, which encodes K binary masks (one for each of the K classes) of resolution m x m each […] for an RoI associated with ground-truth class k, Lmask may only be defined for the mask that corresponds to class k”; and Col. 38, lines 20-22: “the system may select, for each of the selected RoIs, the associated instance segmentation mask that corresponds to the predicted class […] for a particular RoI, the system may generate K masks that correspond to K predetermined classifications (e.g., a mask for an instance of "people," another mask for an instance of "car," etc.)”),
wherein the second trained ML model is a neural network model (Col. 36, lines 52-56: “the system may generate instance segmentation masks associated with the RoI by processing the regional feature map using a second neural network, such as a fully convolutional network”; and Col. 37, lines 47-50: “once trained, the neural network of the mask branch (e.g., the fully convolutional network) would be configured to generate instance segmentation masks for object instances depicted in images”), and
wherein the second trained ML model comprises a convolutional neural network model comprising a series of convolutional layers (Col. 35, lines 25-29: “the ‘x4’ denotes a stack of four consecutive convolutions and the generation of the 28 x 28 x 256 map is by deconvolution. In particular embodiments, all convolutions are 3x3, except for the output convolution, which is 1 x 1”).
Regarding claim(s) 12, He teaches the method of claim 1, wherein the first trained ML model was trained on image data including images of numerous types of objects each of a different type than the object's type (Figure 10; Col. 30, lines 20-25: “[…] result of performing image processing using Mask R-CNN on the Common Objects in Context (COCO) test set (which contains various images of particular objects). FIG. 10 may alternatively represent a training image with ground-truth labels for classification (e.g., labels of "car" or "person") […]”; and Col. 31, lines 61-64: “during training, each sample image in the training dataset may be processed by a neural network (e.g., ResNet 50 or any other suitable network) to generate a feature map”); and
wherein the second trained ML was trained on image data including images of objects of a same type as the object's type (Col. 32, lines 15-20: “the mask branch may have a Km2-dimensional output for each RoI, which encodes K binary masks (one for each of the K classes) of resolution m x m each […] for an RoI associated with ground-truth class k, Lmask may only be defined for the mask that corresponds to class k”; and Col. 38, lines 20-22: “the system may select, for each of the selected RoIs, the associated instance segmentation mask that corresponds to the predicted class […] for a particular RoI, the system may generate K masks that correspond to K predetermined classifications (e.g., a mask for an instance of "people," another mask for an instance of "car," etc.)”).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 2-3 and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over He et al (US 10,713,794 B1) in view of Wang et al (CN 103488988 A; See translation provided by Examiner).
Regarding claim(s) 2 and 15, He teaches the method of claim 1, but do not specifically teach wherein obtaining the image of the object comprises capturing the image of the object using an unmanned aerial vehicle. However, Wang teaches wherein obtaining the image of the object comprises capturing the image of the object using an unmanned aerial vehicle (Paragraph [0043]: “The visible image obtained from the unmanned plane inspection system is the RGB image […] the present invention first carries out filtering by Gaussian filter to image […] The main task of this paper is exactly automatically to extract porcelain insulator from different transformer stations”; and Paragraph [0016]: “The visible visible image to take power equipment, obtain the power equipment image, include the insulator image that at least one can be identified in this power equipment image”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify He to capture the image of the object using an unmanned aerial vehicle, as taught by Wang, in order to facilitate automated inspection and detection of power-equipment objects using aerially captured images. The motivation for this combination of references would have been to automate the detection of objects from images obtained during unmanned aerial inspection, because Wang expressly recognizes that the robotization that realizes the electric power circuit insulator detection based on the unmanned plane line walking is more and more Important (Paragraph [0013]). This motivation for the combination of He and Wang is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Regarding claim(s) 3, He teaches the method of claim 1, but do not specifically teach wherein the image of the object is an image of power electronics equipment. However, Wang teaches wherein the image of the object is an image of power electronics equipment (Paragraph [0015]: “[…] the purpose of this invention is to provide insulator in the power equipment based on unmanned plane line walking visible image, this extracting method efficient and automatically detecting the insulator in electrical equipment”; and Paragraph [0016]: “The visible visible image to take power equipment, obtain the power equipment image, include the insulator image that at least one can be identified in this power equipment image”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify He to process an image of power electronics equipment, as taught by Wang, in order to facilitate automated detection and extraction of electrical power-equipment objects from images. The motivation for this combination of references would have been to provide efficient and automated detection of insulators in electrical equipment, because Wang expressly teaches that the extracting method that the purpose of this invention is to provide insulator in the power equipment based on unmanned plane line walking visible image, this extracting method efficient and automatically detecting the insulator in electrical equipment (Paragraph [0015]). This motivation for the combination of He and Wang is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Claim(s) 7-8, 11, and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over He et al (US 10,713,794 B1) in view of Bajpai et al (US 2023/0072400 A1).
Regarding claim(s) 7, He teaches the method of claim 1, but do not specifically teach wherein at least some parameter values of the first trained ML model are obtained by transfer learning. However, Bajpai teaches wherein at least some parameter values of the first trained ML model are obtained by transfer learning (Paragraph [0042]: “Transfer Learning is used to initialize the starting point of a neural network to be trained on a specific task, with the parameters of a neural network already trained on a similar task”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the first trained ML model of He such that at least some parameter values of the first trained ML model are obtained by transfer learning, as taught by Bajpai, by initializing the neural network using parameters of a previously trained neural network for training on a specific task. The motivation for this combination of references would have been to reduce the amount of target-task training data while improving model performance and accuracy, because Bajpai teaches that “Firstly, even the reduced amount of data for the target task may provide good performance. Secondly, results in a more effective and accurate model for the task at hand” (Paragraph [0046]). This motivation for the combination of He and Bajpai is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Regarding claim(s) 8 and 18, He teaches the method of claim 1, but do not specifically teach wherein the at least some parameter values of the first trained ML model were estimated by training on image data including multiple types of objects of a different type than the object's type. However, Bajpai teaches wherein the at least some parameter values of the first trained ML model were estimated by training on image data including multiple types of objects of a different type than the object's type (Paragraph [0042]: “Transfer Learning is used to initialize the starting point of a neural network to be trained on a specific task, with the parameters of a neural network already trained on a similar task”; Paragraph [0050]: “Pre-trained ImageNet models have provided a significant boost in both natural and medical imaging applications [...] pre-trained ImageNet models are supervised and based on 2D natural images […]”; Paragraph [0052]: “By employing transfer learning to nnU-Net: by transferring generic image representation learned from the massive images to a specific target task. To train the target task, the starting point was utilized from Models Genesis based on nnU-Net architecture”; and Paragraph [0070]: “the ImageNet dataset is large scale and contains 1.3 million labeled natural images annotated by humans”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify He's pre-trained feature-extraction network by employing Bajpai's transfer-learning technique, such that parameter values of the feature-extraction network are obtained from prior training on image data encompassing object types different from the object type of a particular target task. The motivation for this combination of references would have been to transfer “generic image representation learned from the massive images to a specific target task” (Bajpai, Paragraph [0052]), because transfer learning allows “even the reduced amount of data for the target task” to “provide good performance” and results “in a more effective and accurate model for the task at hand” (Bajpai, Paragraph [0046]). This motivation for the combination of He and Bajpai is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Regarding claim(s) 11, He as modified by Bajpai teaches the method of claim 10, but do not specifically teach wherein the second trained ML model has a U-Net architecture. However, Bajpai teaches wherein the second trained ML model has a U-Net architecture (Paragraph [0065]: “extended fully convolutional networks may be utilized by introducing upsampling layers in order to increase the resolution of the output and proposed an encoder-decoder architecture named U-Net. To get the segmentation map, the encoder extracts the features, then a decoder projects these features to a higher resolution”; and Paragraph [0101]: “The nnU-Net framework trains three variations of architecture for each task i.e. 2D U-Net, 3D U-Net, and 3D U-Net Cascade […] the 3D U-Net architecture was utilized”).
Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over He et al (US 10,713,794 B1) in view of Sun et al (US 2021/0397966 A1).
Regarding claim(s) 9, He teaches the method of claim 1, but do not specifically teach wherein transforming the first image features to second image features for processing by the second trained ML model is performed by a using a trained ML transformation model, wherein the trained ML transformation model comprises an encoder of a trained autoencoder (AE) model or a trained variational autoencoder (VAE) model; and wherein dimensionality and/or shape of the transform space of the trained AE model or the trained VAE model matches dimensionality and/or shape of input to second trained ML model.
However, Sun teaches wherein transforming the first image features to second image features for processing by the second trained ML model is performed by a using a trained ML transformation model, wherein the trained ML transformation model comprises an encoder of a trained autoencoder (AE) model or a trained variational autoencoder (VAE) model (Paragraph [0016]: “the neural network system 100 may be implemented using an autoencoder structure, a variational autoencoder structure, and/or other types of structures suitable for the functionality described herein. When a variational autoencoder is used, the representation (e.g., latent variable z) may be regularized using a prior normal distribution”; Paragraph [0013]: “The encoder network 102 may be configured to receive an input image 106 such as a medical image (e.g., a CT or MRI scan) and produce a representation 108 (e.g., a low-resolution or low-dimension representation) of the input image that indicates one or more features of the input image […] Each of these neural networks may comprises a plurality of layers, and the encoder network 102 may be configured to produce the representation 108 by performing a series of down-sampling and/or convolution operations on the input image 106 through the layers of the neural networks […] The output of the encoder network 102 may include the representation 108, which may be in various forms depending on the specific task to be accomplished. For instance, the representation 108 may include a latent space variable Z”); and
wherein dimensionality and/or shape of the transform space of the trained AE model or the trained VAE model matches dimensionality and/or shape of input to second trained ML model (Paragraph [0014]: “The decoder network 104 may be configured to receive the representation 108 produced by the encoder 102 and reconstruct (e.g., recover the details of) the input image 106 based on the representation 108. The decoder network 104 may generate a mask 110 (e.g., a pixel- or voxel-wise segmentation mask) for segmenting an object (e.g., a body part such as an organ) from the image 106 […] Using the un-pooling layers, the decoder network 104 may up-sample the representation 108 produced by the encoder network 102, e.g., based on pooled indices stored by the encoder. The up-sampled representation may then be processed through the convolutional layers […]”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the feature transformation of He to employ a trained ML transformation model comprising an encoder of a trained autoencoder or variational autoencoder, as taught by Sun, such that the first image features are transformed into a lower-dimensional or latent representation having dimensionality and/or shape suitable for input to the downstream trained segmentation model. The motivation for this combination of references would have been to provide a compact feature representation suitable for downstream segmentation processing while improving segmentation performance, because Sun teaches that the encoder reduces “the redundancy and/or dimension of the features” (Paragraph [0013]) and further teaches that the encoder and/or decoder may be trained to use a shape prior during image segmentation “to prevent or reduce problems associated with over-segmentation […] and/or under-segmentation […] and to improve the success rate of the segmentation operation” (Paragraph [0017]). This motivation for the combination of He and Sun is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Claim(s) 14 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lee et al (US 2019/0311202 A1) in view of Bajpai et al (US 2023/0072400 A1).
Regarding claim(s) 14, Lee teaches a method for detecting objects in images using machine learning (ML), the method comprising:
using at least one computer hardware processor (Figure 19; and Paragraph [0124]) to perform:
obtaining an image of an object (Figure 1A; and Paragraph [0042]: “Example video stream 100 includes a set of n video frames 110-1, 110-2, 110-3, . . ., and 110-n (collectively video frames 110) that are sequential in time”); and
segmenting the object in the image from background of the object in the image to obtain a segmentation mask for the image by using multiple trained encoders (Figure 6; Paragraph [0041]: “video object segmentation can be used to segment an object from a background and output a mask of the object in each frame of a video stream”; Paragraph [0042]: “Each video frame 110 includes a foreground object 120 (e.g., a car) to be segmented from the background in each video frame 110”; Paragraph [0061]: “Neural network 600 includes a first encoder 620 and a second encoder 630. First encoder 620 takes an input 610, which includes a reference video frame and the ground-truth mask, and extracts a feature map 625 from input 610. Second encoder 630 takes an input 615, which includes a target video frame in a video stream and an estimated mask of the previous video frame in the video stream, and extracts a feature map 635 from input 615”) including:
a second encoder whose weights are trained using images of objects of a same type as the object (Paragraph [0073]: “can be used to synthesize training samples, which include both the reference images and the corresponding target images, where each pair of reference image and target image include a same object”; Paragraph [0087]: “One encoder takes a reference image and a ground-truth object mask that identifies an object in the reference image as inputs, and the other encode takes a target image that includes the same object and a guidance object mask as inputs. Thus, two images including the same object may be needed”; and Paragraph [0088]: “the neural network including two encoders is trained using the pair of training images and the corresponding object masks, where one training image is fed to a first encoder of the two encoders as a reference image and the other training image is fed to a second encoder as a target image”), wherein the segmentation mask identifies pixels associated with the object in the image and pixels associated with background of the object in the image (Paragraph [0035]: “segmentation data includes a set of labels, such as pairwise labels […] indicating whether a given pixel in the image is part of an image region depicting a human figure. In some cases, labels have multiple available values, such as a set of labels indicating whether a given pixel depicts, for example, a human figure, an animal figure, or a background region”; and Paragraph [0036]: “A mask, objet mask, or segmentation mask may refer to an image where the intensity values for pixels in a region of interest are non-zero, while the intensity values for pixels in other regions of the image are set to the background value (e.g., zero)”).
Lee fails to teach a first encoder whose weights are obtained by transfer learning. However, Bajpai teaches a first encoder whose weights are obtained by transfer learning (Paragraph [0042]: “Transfer Learning is used to initialize the starting point of a neural network to be trained on a specific task, with the parameters of a neural network already trained on a similar task”; and Paragraph [0138]: “All the layers were fine-tuned from encoder and decoder blocks in the above segmentation tasks. The weights were transferred for all except the last layer from the pre-trained model”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the multiple-encoder segmentation network of Lee such that the weights of a first encoder are obtained by transfer learning, as taught by Bajpai, while retaining Lee's trained second encoder receiving training images including the same object, in order to initialize the first encoder using parameters learned by a previously trained neural network and thereby improve training efficiency and segmentation performance. The motivation for this combination of references would have been to reduce the amount of target-task training data required while improving model effectiveness and accuracy, because Bajpai teaches that transfer learning allows a neural network to be initialized using parameters of a neural network already trained on a similar task (Paragraph [0042]) and further teaches that “even the reduced amount of data for the target task may provide good performance” and results in “a more effective and accurate model for the task at hand” (Paragraph [0046]). This motivation for the combination of Lee and Bajpai is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Regarding claim(s) 16, Lee as modified by Bajpai teaches the method of claim 14, where Lee teaches wherein segmenting the objects comprises:
encoding the image using the first encoder to obtain first image features (Figure 6; and Paragraph [0061]: “Neural network 600 includes a first encoder 620 and a second encoder 630. First encoder 620 takes an input 610, which includes a reference video frame and the ground-truth mask, and extracts a feature map 625 from input 610”),
wherein parameter values of the first encoder are obtained from a portion of another trained neural network model (Paragraph [0065]: “Each of the reference frame encoder subnetwork and the target frame encoder subnetwork may include a fully convolutional neural network […] a known convolutional neural network for static image classification, such as ResNet 50 or VGG-16, is adopted and modified (e.g., adding a fourth channel for the mask in addition to the R, G, and B channels, and removing the fully connected layers) for use as the reference frame encoder subnetwork and the target frame encoder subnetwork […] the network parameters are initialized from an ImageNet pre-trained neural network, such as ResNet 50 or VGG-16”), and
wherein the first image features comprise feature maps determined by processing the image using the first encoder (Figure 7; Paragraph [0061]: “First encoder 620 takes an input 610, which includes a reference video frame and the ground-truth mask, and extracts a feature map 625 from input 610”; and Paragraph [0066]: “Each convolution layer in block 722 or 724 performs convolutions with a filter bank to produce a set of feature maps.”).
Claim(s) 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lee et al (US 2019/0311202 A1) in view of Bajpai et al (US 2023/0072400 A1), further in view of Chidlovskii (US 2021/0174513 A1).
Regarding claim(s) 19, Lee as modified by Bajpai teaches the method of claim 16, but do not specifically teach wherein segmenting the object further comprises: transforming, using the second encoder, the first image features to second image features for processing by a second trained ML model, wherein the second encoder is an encoder of a trained autoencoder (AE) model or a trained variational autoencoder (VAE) model, and wherein dimensionality and/or shape of the transform space of the trained AE model or the trained VAE model matches dimensionality and/or shape of the second trained ML model.
However, Chidlovskii teaches wherein segmenting the object further comprises:
transforming, using the second encoder, the first image features to second image features for processing by a second trained ML model (Figure 2B; Paragraph [0049]: “the feature maps generated by the encoder branch 210 are fed to the third subnetwork S3 206 (i.e., the common representation network 207) […] the common representation network 207 utilizes a multi-view autoencoder that enables the extraction of a common representation from either one or two views”; and Paragraph [0050]: “the third subnetwork S3 206 reconstructs the inputted first feature map […] and the inputted second feature map […] from the outputted first feature map […] and the outputted second feature map […] More specifically, the third subnetwork S3 206 generates the common representation network 207 from the outputted first feature map […] and the outputted second feature map […] and then generates the inputted first feature map […] and the inputted second feature map […] using the common representation network 207”),
wherein the second encoder is an encoder of a trained autoencoder (AE) model or a trained variational autoencoder (VAE) model (Paragraph [0052]: “the third subnetwork S3 206 is a multi-view autoencoder network, with at least one hidden layer. The inputs to the hidden layer of the third subnetwork S3 206 are the outputted feature map”; Paragraph [0053] – Paragraph [0054]: “the hidden layer typically computes an encoded representation […]”; Paragraph [0067]: “The CNN 200 as a whole is trained at 308 by minimizing a weighted sum of the first loss function […] the second loss function […] and a third loss function […] This objective function to minimize may then be defined […] In the above formulation, the semantic, depth and reconstruction losses are optimized jointly”; and Paragraph [0070]: “the first term is the usual autoencoder objective function which helps in learning meaningful hidden representations”), and
wherein dimensionality and/or shape of the transform space of the trained AE model or the trained VAE model matches dimensionality and/or shape of the second trained ML model (Figure 2A; Paragraph [0052]: “Similar to conventional autoencoders, the input and output layer of the auto encoder network have the same shape as the input, d × H′ × W′, whereas the hidden layer is shaped as k × H′ × W′, where k is frequently smaller than d”; and Paragraph [0080]: “From the common representation computed at 406b, outputted first feature map […] is processed at 406c by the RGB Decoder 218 of the first subnetwork S1 202 to generate a semantically segmented image (RGB-SS branch)”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the second encoder of the multiple-encoder segmentation system of Lee, as modified by Bajpai, to employ an encoding portion of a trained autoencoder as taught by Chidlovskii, such that first image features are transformed through an encoded common representation into second image features having a shape suitable for processing by a downstream trained segmentation model. The motivation for this combination of references would have been to obtain a common feature representation that permits feature information to be reconstructed and processed by a downstream segmentation network, because Chidlovskii teaches that the multi-view autoencoder “enables the extraction of a common representation from either one or two views” (Paragraph [0049]) and that the autoencoder objective “helps in learning meaningful hidden representations” (Paragraph [0070]). This motivation for the combination of Lee, Bajpai, and Chidlovskii is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Regarding claim(s) 20, Lee as modified by Bajpai and Chidlovskii teaches the method of claim 19,
where Lee teaches wherein the second trained ML model is trained to segment, from background, objects of a same type as the type of the object (Paragraph [0073]: “[…] used to synthesize training samples, which include both the reference images and the corresponding target images, where each pair of reference image and target image include a same object; and Paragraph [0105]: “When there are multiple objects to be segmented from a video stream, the same network can be used, and the training can be based on a single object”),
where Chidlovskii teaches wherein segmenting the object further comprises processing the second image features using the second trained ML model to obtain the segmentation mask for the image (Paragraph [0045]: “the decoder branch 212 up samples the feature map back to the original input resolution by semantically projecting low-resolution features learned by the encoders 214 and 216 onto the pixel high resolution space to output a dense, pixel-level classification […] the RGB Decoder 218 generates a semantically segmented image (xₛ) from an inputted first feature map”; and Paragraph [0080]: “From the common representation computed at 406b, outputted first feature map […] is processed at 406c by the RGB Decoder 218 of the first subnetwork S1 202 to generate a semantically segmented image (RGB-SS branch)”), and
wherein the second trained ML model is neural network model comprising a series of convolutional layers (Figure 8A-8B; Paragraph [0068] – Paragraph [0069]: “an example k × k global convolution block is achieved using a 1 × k convolution layer 810 and a k × 1 convolution layer 820 on one path, and a k × 1 convolution layer 830 and a 1 × k convolution layer 840 on another path […] The residual mapping path includes two or more convolution and ReLU layers 852 and 854”; and Paragraph [0102]: “a decoder (e.g., decoder 750) of the neural network extracts a target segmentation mask for the target frame from the combined feature map”).
Relevant Prior Art Directed to State of Art
Davies et al (US 2022/0309633 A1) are relevant prior art not applied in the rejection(s) above. Davies discloses a computer system configured to automatically interpolate or extrapolate visual modifications from a set of keyframes extracted from a set of target video frames, the system comprising: a computer processor, operating in conjunction with computer memory maintaining a machine learning model architecture, the computer processor configured to: receive the set of target video frames; identify, from the set of target video frames, the set of keyframes; provide the set of keyframes for visual modification by a human; receive a set of modified keyframes; train the machine learning model architecture using the set of modified keyframes and the set of keyframes, the machine learning model architecture including a first autoencoder configured for unity reconstruction of the set of modified keyframes from the set of keyframes to obtain a trained machine learning model architecture; and process one or more frames of the set of target video frames to generate a corresponding set of modified target video frames having the automatically interpolated or extrapolated visual modifications.
Myronenko (US 12,536,664 B2) are relevant prior art not applied in the rejection(s) above. Myronenko discloses a method comprising: generating an encoding representation based, at least in part, on a first image; generating a second image from the encoding representation; generating a segmentation based, at least in part, on the encoding representation; calculating an error based, at least in part, on the encoding representation, the segmentation, and the second image; and updating one or more neural networks based, at least in part, on the error.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JONGBONG NAH whose telephone number is (571) 272-1361. The examiner can normally be reached M - F: 9:00 AM - 5:30 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ONEAL MISTRY can be reached on 313-446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JONGBONG NAH/Examiner, Art Unit 2674