DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 12/05/2025 is/are compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Office Action Summary
Claim(s) 1-2, 4-5, 9-10, 12-13, and 16-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Terra et al (US 2025/0005450 A1) in view of Faulhaber et al (US 2019/0156247 A1), further in view of Lutchoomun et al (WO 2024211555 A1).
Claim(s) 3, 6, 11, and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Terra et al (US 2025/0005450 A1) in view of Faulhaber et al (US 2019/0156247 A1) and Lutchoomun et al (WO 2024211555 A1), further in view of Chidlovskii et al (US 2021/0174513 A1).
Claim(s) 7. 14, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Terra et al (US 2025/0005450 A1) in view of Faulhaber et al (US 2019/0156247 A1) and Lutchoomun et al (WO 2024211555 A1), further in view of Bajpai et al (US 2023/0072400 A1).
Claim(s) 8 and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Terra et al (US 2025/0005450 A1) in view of Faulhaber et al (US 2019/0156247 A1), Lutchoomun et al (WO 2024211555 A1) and Bajpai et al (US 2023/0072400 A1), further in view of Lee et al (US 2019/0311202 A1) and Chidlovskii (US 2021/0174513 A1).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-2, 4-5, 9-10, 12-13, and 16-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Terra et al (US 2025/0005450 A1) in view of Faulhaber et al (US 2019/0156247 A1), further in view of Lutchoomun et al (WO 2024211555 A1).
Regarding claim(s) 1, 9, and 16, Terra teaches at least one non-transitory computer-readable storage medium storing processor- executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform a method for online detection of objects in images using a set of trained machine learning (ML) models (Figure 2A; and Paragraph [0032]), the method comprising:
(a) responsive to receiving images of objects of a first type, processing the received images, using a selected trained ML model from among the set of trained ML models, to segment objects of the first type in the images to obtain respective segmentation masks (Figure 6B; Paragraph [0026]: “the risk value determined for the state of the environment, based on the safety information, is then used to dynamically select one of a plurality of ML models available to the ML agent […] The ML models available to the ML agent have all been trained (using any suitable training mechanism, such as supervised learning, reinforcement learning, unsupervised learning and so on) to execute a task. In some embodiments, all of the plurality of ML models are trained to perform the same task”; Paragraph [0042]: “In image segmentation, obstacles are identified in captured images (the object may be said to occupy a segment of the image) using a suitable ML model such as a trained deep neural network model (DNN) […]”; and Paragraph [0043]: “Two examples of models that may be used for image segmentation are Multi-level Scene Description Networks (MSDN) and Mask-Recursive Convolutional Neural Networks (Mask-RCNN) […] generate more precise segmentation masks identifying the position of the obstacle in the image with pixel accuracy”).
Terra fails to teach (b) during performance of (a) and either in accordance with a prespecified schedule and/or responsive to one or more triggering events, evaluating performance of the selected trained ML model and one or more other trained ML models in the set of trained ML models using one or more measures of performance; identifying, from among the selected trained ML model and the one or more other trained ML models, a particular trained ML based on its performance on the one or more measures of performance and in accordance with one or more ML model selection rules; when the selected trained model is identified as the particular trained ML model, continuing performance of (a) using the selected trained ML model; and when the selected trained ML model is not identified as the particular trained ML model, continuing performance of (a) using the particular trained ML model instead of the selected trained ML model and treating the particular trained ML model as the selected trained ML model next time (b) is performed.
However, Faulhaber teaches (b) during performance of (a) (Paragraph [0017]: “one or more primary ML models can be used to actively service inference requests while one or more secondary ML models can similarly—but without direct visibility for users or influence over the results provided to users—perform inference using the same input data, allowing the secondary ML model(s) to be evaluated for actual performance under the same conditions and environment as the “live” primary ML model(s)”; and Paragraph [0018]: “the quality of multiple ML models can be measured and further traffic (e.g., inference requests) can be redirected in a controlled manner to cause more traffic to be processed by those ML models that are performing better […]”);
identifying, from among the selected trained ML model and the one or more other trained ML models, a particular trained ML based on its performance on the one or more measures of performance and in accordance with one or more ML model selection rules (Paragraph [0038]: “each of these twenty different models can operate upon the request using the same data and environment, allowing for a true apples-to-apples comparison of the performance and results of these models. The analytics engine 122 thus can, for example, watch the outputs of each model and/or measure the performance (e.g., required time to execute, resource utilization such as processing, memory, etc.) of each model. The results of these parallel “shadow” executions, on a per-request basis and/or in an aggregate form (e.g., across multiple requests grouped according to time, type, etc.) can be provided to the user to provide the useful information needed to select the best model or models for future jobs, or can be used (e.g., with a set of user-defined model transition rules) to update the model selector 110 to use different models”);
when the selected trained model is identified as the particular trained ML model, continuing performance of (a) using the selected trained ML model; and when the selected trained ML model is not identified as the particular trained ML model, continuing performance of (a) using the particular trained ML model instead of the selected trained ML model and treating the particular trained ML model as the selected trained ML model next time (b) is performed (Paragraph [0036]: “the dynamic router 108 can apply both an old model (or models) and the new model for incoming requests that are actually serviced by an old model […] the analytics engine 122 can measure how the “new” model would have performed if it had actually been set as the “live” model”; Paragraph [0037]: “sending an update message 138 to the dynamic router 108 to cause the model selector 110 to switch over some or all traffic to a “new” model (e.g., if its performance meets or exceeds some threshold, such as having an accuracy value that is greater than the “old” model's corresponding accuracy value), sending analytic results 202 to a logging system or client, etc.”; and Paragraph [0047]: “the machine learning service 140 can have multiple models performing the same tasks (or the same “type” of inference), discover which model is performing better, start shifting over traffic to the more performant one(s), continue to monitor the performance of the models, and continue to adjust the shifting of traffic accordingly. Thus, a shift may occur in one direction (e.g., only from a first model A to a second model B) and/or in two directions (e.g., from model A towards model B, and then later from model B back towards model A)”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Terra's dynamic selection of trained ML models for image segmentation with Faulhaber's dynamic performance evaluation and model-transition techniques so that the selected segmentation model and other available trained segmentation models are evaluated during operation and a better-performing model is selected and used for subsequent processing. The motivation for this combination of references would have been to “select the best model or models for future jobs” and to “start shifting over traffic to the more performant one(s), continue to monitor the performance of the models, and continue to adjust the shifting of traffic accordingly.” This motivation for the combination of Terra and Faulhaber is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Terra and Faulhaber fail to teach either in accordance with a prespecified schedule and/or responsive to one or more triggering events. However, Lutchoomun teaches either in accordance with a prespecified schedule and/or responsive to one or more triggering events (Paragraph [0110]: “the WTRU may be configured to perform performance monitoring with a minimum frequency (e.g., an assessment every X s/min) etc. […] In the case the WTRU is performing performance monitoring for multiple models, the WTRU may send one performance report for each model (possible associated with a model ID and/or metadata about the model being assessed) or one performance report aggregating the assessment of multiple models”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the dynamic model-evaluation system of Terra as modified by Faulhaber according to the performance-monitoring technique of Lutchoomun by performing the evaluation of the selected trained ML model and the other trained ML models in accordance with a prespecified schedule. The motivation for this combination of references would have been to provide periodic performance assessment of the available ML models at a configured minimum frequency, thereby enabling the performance-monitoring and model-selection process to be performed at predetermined intervals. This motivation for the combination of Terra, Faulhaber, and Lutchoomun is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Regarding claim(s) 2, 10, and 17, Terra as modified by Faulhaber and Lutchoomun teaches the method of claim 1, wherein the one or more measures of performance comprises a segmentation performance metric (where Terra teaches Paragraph [0043]: “Mask-RCNN models generate more precise segmentation masks identifying the position of the obstacle in the image with pixel accuracy”; and where Faulhaber further teaches Paragraph [0036]: “the analytics engine 122 can interact with a ground truth collector 124 to obtain ground truth for a set of requests, and compare this obtained ground truth with the inference results generated by the ML model(s) 118A-118C under scrutiny to identify the true accuracy of these models”; and Paragraph [0071]: “[…] comparing the plurality of inference results generated by the plurality of ML models with the obtained ground truth values, and assigning an accuracy score for a model based on how frequent and/or how similar that model's inference results match the corresponding ground truth values”).
Regarding claim(s) 4, 12, and 18, Terra as modified by Faulhaber and Lutchoomun teaches the method of claim 1, wherein the one or more measures of performance comprises a computational performance metric providing a measure of an amount of time used by an ML model to segment objects of the first type in a set of one or more images (where Terra teaches Paragraph [0042]: “In image segmentation, obstacles are identified in captured images (the object may be said to occupy a segment of the image) using a suitable ML model such as a trained deep neural network model (DNN) […]”; and Paragraph [0043]: “Two examples of models that may be used for image segmentation are Multi-level Scene Description Networks (MSDN) and Mask-Recursive Convolutional Neural Networks (Mask-RCNN)”; and where Faulhaber teaches Paragraph [0038]: “each of these twenty different models can operate upon the request using the same data and environment, allowing for a true apples-to-apples comparison of the performance and results of these models. The analytics engine 122 thus can, for example, watch the outputs of each model and/or measure the performance (e.g., required time to execute, resource utilization such as processing, memory, etc.) of each model”).
Regarding claim(s) 5, 13, and 19, Terra as modified by Faulhaber and Lutchoomun teaches the method of claim 1, where Faulhaber teaches wherein the one or more ML model selection rules comprises a rule for selecting an ML model, from among a set of ML models, based on values of a computational performance metric and/or a segmentation performance metric for ML models in the set of ML models (Paragraph [0038]: “The analytics engine 122 thus can, for example, watch the outputs of each model and/or measure the performance (e.g., required time to execute, resource utilization such as processing, memory, etc.) of each model. The results of these parallel “shadow” executions, on a per-request basis and/or in an aggregate form (e.g., across multiple requests grouped according to time, type, etc.) can be provided to the user to provide the useful information needed to select the best model or models for future jobs, or can be used (e.g., with a set of user-defined model transition rules) to update the model selector 110 to use different models”; and Paragraph [0047]: “the machine learning service 140 can have multiple models performing the same tasks (or the same “type” of inference), discover which model is performing better, start shifting over traffic to the more performant one(s), continue to monitor the performance of the models, and continue to adjust the shifting of traffic accordingly”).
Claim(s) 3, 6, 11, and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Terra et al (US 2025/0005450 A1) in view of Faulhaber et al (US 2019/0156247 A1) and Lutchoomun et al (WO 2024211555 A1), further in view of Chidlovskii et al (US 2021/0174513 A1).
Regarding claim(s) 3, 11, and 17, Terra as modified by Faulhaber and Lutchoomun teaches the method of claim 2, but do not specifically teach wherein the segmentation performance metric is an intersection of over union (IOU) metric for measuring accuracy of segmenting objects of the first type in a set of one or more images.
However, Chidlovskii teaches wherein the segmentation performance metric is an intersection of over union (IOU) metric for measuring accuracy of segmenting objects of the first type in a set of one or more images (Equation 6; Equation 7; and Paragraph [0087]: “Two common criteria that are used to evaluate a neural network's performance of a segmentation task are the pixel accuracy and the intersection-over-union (IoU) score […] intersection-over-union (IoU) calculates the average value of the intersection between the ground truth and the predictions”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the segmentation performance evaluation of Terra as modified by Faulhaber and Lutchoomun according to Chidlovskii by using an intersection-over-union (IoU) score as the segmentation performance metric for measuring the accuracy of segmentation. The motivation for this combination of references would have been because Chidlovskii expressly teaches that “Two common criteria that are used to evaluate a neural network's performance of a segmentation task are the pixel accuracy and the intersection-over-union (IoU) score,” thereby providing a known criterion for evaluating the performance of a segmentation task. This motivation for the combination of Terra, Faulhaber, Lutchoomun, and Chidlovskii by is supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Regarding claim(s) 6, Terra as modified by Faulhaber and Lutchoomun teaches the method of claim 5, where Faulhaber teaches wherein the one or more ML model selection rules comprises a rule for selecting an ML model, from among a set of ML models, based on values of a computational performance metric (Paragraph [0038]: “The analytics engine 122 thus can, for example, watch the outputs of each model and/or measure the performance (e.g., required time to execute, resource utilization such as processing, memory, etc.) of each model. The results of these parallel “shadow” executions, on a per-request basis and/or in an aggregate form (e.g., across multiple requests grouped according to time, type, etc.) can be provided to the user to provide the useful information needed to select the best model or models for future jobs, or can be used (e.g., with a set of user-defined model transition rules) to update the model selector 110 to use different models”; Paragraph [0036]: “the analytics engine 122 can interact with a ground truth collector 124 to obtain ground truth for a set of requests, and compare this obtained ground truth with the inference results generated by the ML model(s) 118A-118C under scrutiny to identify the true accuracy of these models”; and Paragraph [0037]: “the model selector 110 to switch over some or all traffic to a “new” model (e.g., if its performance meets or exceeds some threshold, such as having an accuracy value that is greater than the “old” model's corresponding accuracy value)”).
Terra, Faulhaber, and Lutchoomun fail to teach a segmentation performance metric. However, Chidlovskii teaches a segmentation performance metric (Paragraph [0087]: “Two common criteria that are used to evaluate a neural network's performance of a segmentation task are the pixel accuracy and the intersection-over-union (IoU) score”).
Therefore, It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify Terra as modified by Faulhaber and Lutchoomun with Chidlovskii to select an ML model based on values of both a computational performance metric and a segmentation performance metric. The motivation for this combination of references would have been to evaluate the ML models using both computational performance and segmentation performance, where Faulhaber teaches measuring the “required time to execute” and using the resulting performance information to “select the best model or models for future jobs,” and Chidlovskii teaches that “pixel accuracy and the intersection-over-union (IoU) score” are criteria used to evaluate “a neural network's performance of a segmentation task.” This motivation for the combination of Terra, Faulhaber, Lutchoomun, and Chidlovskii to is supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Claim(s) 7. 14, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Terra et al (US 2025/0005450 A1) in view of Faulhaber et al (US 2019/0156247 A1) and Lutchoomun et al (WO 2024211555 A1), further in view of Bajpai et al (US 2023/0072400 A1).
Regarding claim(s) 7, 14, and 20, Terra as modified by Faulhaber and Lutchoomun teaches the method of claim 1, but do not specifically teach wherein the particular trained ML model is one of a U-Net neural network model, a multi-encoder neural network segmentation model, and a K-means clustering algorithm.
However, Bajpai teaches wherein the particular trained ML model is one of a U-Net neural network model, a multi-encoder neural network segmentation model, and a K-means clustering algorithm (Figure 2; Paragraph [0065]: “extended fully convolutional networks may be utilized by introducing upsampling layers in order to increase the resolution of the output and proposed an encoder-decoder architecture named U-Net. To get the segmentation map, the encoder extracts the features, then a decoder projects these features to a higher resolution”; and Paragraph [0101]: “The nnU-Net framework trains three variations of architecture for each task i.e. 2D U-Net, 3D U-Net, and 3D U-Net Cascade as shown in FIG. 2. Based on experimental results, 3D U-Net is the most favorable option for 3D imaging tasks. Due to this reason, the 3D U-Net architecture was utilized, extracted from the nnU-Net framework, and demonstrated that the performance can be enhanced by fine-tuning Models Genesis for lung tumor, liver organ, and liver tumor segmentation tasks”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the image-segmentation ML model of Terra as modified by Faulhaber and Lutchoomun according to Bajpai by using a U-Net neural network model as the particular trained ML model. The motivation for this combination of references would have been to provide a trained segmentation model demonstrated to be favorable based on experimental performance for use among the trained ML models available for selection, thereby enabling selection of a suitable segmentation model based on its performance for the particular imaging task. This motivation for the combination of Terra, Faulhaber, Lutchoomun, and Bajpai is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Claim(s) 8 and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Terra et al (US 2025/0005450 A1) in view of Faulhaber et al (US 2019/0156247 A1), Lutchoomun et al (WO 2024211555 A1) and Bajpai et al (US 2023/0072400 A1), further in view of Lee et al (US 2019/0311202 A1) and Chidlovskii (US 2021/0174513 A1).
Regarding claim(s) 8 and 15, Terra as modified by Faulhaber, Lutchoomun, and Bajpai teaches the method of claim 7, where Bajpai teaches wherein the multi-encoder neural network model comprises:
a first encoder whose weights are obtained by transfer learning (Paragraph [0042]: “Transfer Learning is used to initialize the starting point of a neural network to be trained on a specific task, with the parameters of a neural network already trained on a similar task”; and Paragraph [0138]: “All the layers were fine-tuned from encoder and decoder blocks in the above segmentation tasks. The weights were transferred for all except the last layer from the pre-trained model”).
Terra, Faulhaber, Lutchoomun, and Bajpai fail to teach a second encoder whose weights are trained using images of objects of a same type as the object; and a neural network model, wherein the first encoder is trained to process an image to obtain first image features, the second encoder is trained to transform first image features to second image features for subsequent processing by the neural network model, and the neural network model is trained to process the second image features to obtain a segmentation mask for the image, wherein the segmentation mask identifies pixels associated with the object in the image and pixels associated with background of the object in the image.
However, Lee teaches a second encoder whose weights are trained using images of objects of a same type as the object Paragraph [0073]: “can be used to synthesize training samples, which include both the reference images and the corresponding target images, where each pair of reference image and target image include a same object”; Paragraph [0087]: “One encoder takes a reference image and a ground-truth object mask that identifies an object in the reference image as inputs, and the other encode takes a target image that includes the same object and a guidance object mask as inputs. Thus, two images including the same object may be needed”; and Paragraph [0088]: “the neural network including two encoders is trained using the pair of training images and the corresponding object masks, where one training image is fed to a first encoder of the two encoders as a reference image and the other training image is fed to a second encoder as a target image”); and
a neural network model (Paragraph [0061]: “Neural network 600 includes a first encoder 620 and a second encoder 630),
wherein the first encoder is trained to process an image to obtain first image features and the neural network model is trained to process the second image features to obtain a segmentation mask for the image (Figure 1A; Figure 6; Paragraph [0042]: “Example video stream 100 includes a set of n video frames 110-1, 110-2, 110-3, . . ., and 110-n (collectively video frames 110) that are sequential in time”; Paragraph [0041]: “video object segmentation can be used to segment an object from a background and output a mask of the object in each frame of a video stream”; Paragraph [0042]: “Each video frame 110 includes a foreground object 120 (e.g., a car) to be segmented from the background in each video frame 110”; Paragraph [0061]: “[…] First encoder 620 takes an input 610, which includes a reference video frame and the ground-truth mask, and extracts a feature map 625 from input 610. Second encoder 630 takes an input 615, which includes a target video frame in a video stream and an estimated mask of the previous video frame in the video stream, and extracts a feature map 635 from input 615”),
wherein the segmentation mask identifies pixels associated with the object in the image and pixels associated with background of the object in the image (Paragraph [0035]: “segmentation data includes a set of labels, such as pairwise labels […] indicating whether a given pixel in the image is part of an image region depicting a human figure. In some cases, labels have multiple available values, such as a set of labels indicating whether a given pixel depicts, for example, a human figure, an animal figure, or a background region”; and Paragraph [0036]: “A mask, objet mask, or segmentation mask may refer to an image where the intensity values for pixels in a region of interest are non-zero, while the intensity values for pixels in other regions of the image are set to the background value (e.g., zero)”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify Terra as modified by Faulhaber, Lutchoomun, and Bajpai with Lee to provide a multi-encoder neural network model having multiple encoders trained using images including the same object and to process extracted image features to obtain a segmentation mask. The motivation for this combination of references would have been to improve segmentation of a target object by using features extracted from both a reference image identifying the target object and a target image containing the same object, because Lee teaches that the extracted features are combined and used to extract the segmentation mask for the target frame, thereby allowing the network to both detect the target object and track the segmentation mask. This motivation is supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Terra, Faulhaber, Lutchoomun, Bajpai, and Lee fail the second encoder is trained to transform first image features to second image features for subsequent processing by the neural network model.
However, Chidlovskii teaches the second encoder is trained to transform first image features to second image features for subsequent processing by the neural network model (Figure 2B; Paragraph [0049]: “the feature maps generated by the encoder branch 210 are fed to the third subnetwork S3 206 (i.e., the common representation network 207) […] the common representation network 207 utilizes a multi-view autoencoder that enables the extraction of a common representation from either one or two views”; and Paragraph [0050]: “the third subnetwork S3 206 reconstructs the inputted first feature map […] and the inputted second feature map […] from the outputted first feature map […] and the outputted second feature map […] More specifically, the third subnetwork S3 206 generates the common representation network 207 from the outputted first feature map […] and the outputted second feature map […] and then generates the inputted first feature map […] and the inputted second feature map […] using the common representation network 207”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify Terra as modified by Faulhaber, Lutchoomun, Bajpai, and Lee with Chidlovskii such that first image features generated by an encoder are further transformed by an additional encoder network into second image features for subsequent processing by the neural network model. The motivation for this combination of references would have been to provide an additional encoded representation of the image features for subsequent segmentation processing, because Chidlovskii teaches feeding feature maps generated by an encoder branch to a multi-view autoencoder that “enables the extraction of a common representation from either one or two views” (Paragraph [0049]), thereby providing a common feature representation for subsequent processing by the segmentation network. This motivation is supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Relevant Prior Art Directed to State of Art
He et al (US 10,713,794 B1) are relevant prior art not applied in the rejection(s) above. He discloses a method comprising, by a computing system: accessing a training image; generating a feature map for the training image using a first neural network; identifying a region of interest in the feature map; generating a regional feature map for the region of interest based on sampling locations defined by a sampling region, wherein the sampling region and the region of interest align to the same region in the feature map; generating an instance segmentation mask associated with the region of interest by processing the regional feature map using a second neural network; and training the second neural network using the instance segmentation mask; wherein the trained second neural network is configured to generate instance segmentation masks for object instances depicted in images.
Davies et al (US 2022/0309633 A1) are relevant prior art not applied in the rejection(s) above. Davies discloses a computer system configured to automatically interpolate or extrapolate visual modifications from a set of keyframes extracted from a set of target video frames, the system comprising: a computer processor, operating in conjunction with computer memory maintaining a machine learning model architecture, the computer processor configured to: receive the set of target video frames; identify, from the set of target video frames, the set of keyframes; provide the set of keyframes for visual modification by a human; receive a set of modified keyframes; train the machine learning model architecture using the set of modified keyframes and the set of keyframes, the machine learning model architecture including a first autoencoder configured for unity reconstruction of the set of modified keyframes from the set of keyframes to obtain a trained machine learning model architecture; and process one or more frames of the set of target video frames to generate a corresponding set of modified target video frames having the automatically interpolated or extrapolated visual modifications.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JONGBONG NAH whose telephone number is (571) 272-1361. The examiner can normally be reached M - F: 9:00 AM - 5:30 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ONEAL MISTRY can be reached on 313-446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JONGBONG NAH/Examiner, Art Unit 2674