Prosecution Insights
Last updated: October 02, 2026
Application No. 18/857,412

IMAGE ASSESSMENT METHOD AND APPARATUS, AND DEVICE, STORAGE MEDIUM AND PROGRAM PRODUCT

Non-Final OA §103
Filed
Oct 16, 2024
Priority
May 13, 2022 — CN 202210524525.X +1 more
Examiner
HAUSMANN, MICHELLE M
Art Unit
Tech Center
Assignee
Beijing Zitiao Network Technology Co., Ltd.
OA Round
1 (Non-Final)
77%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
677 granted / 883 resolved
+16.7% vs TC avg
Strong +21% interview lift
Without
With
+21.1%
Interview Lift
resolved cases with interview
Typical timeline
2y 12m
Avg Prosecution
25 currently pending
Career history
907
Total Applications
across all art units

Statute-Specific Performance

§101
14.0%
-26.0% vs TC avg
§103
67.3%
+27.3% vs TC avg
§102
6.4%
-33.6% vs TC avg
§112
7.3%
-32.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 883 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claims 8 and 23 objected to because of the following informalities: Claim 8 is directed to the “method according to claim 7, wherein the set color space comprises one or more of: an RGB color space, an HSV color space, an LAB color space, and a Grayscale color space”. However claim 7 is written in the alternative, therefore in one case such as choosing rotating the sample image with the preset resolution by a preset angle instead of converting there would be no antecedent basis for this limitation. Similar language used in claims 23 and 22 respectively. Appropriate correction is required. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 2, 3, 13, 14, 17, 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jing et al. (IDS: CN114299358A [Machine Translation]) in view of Kim et al. (US 20200167943 A1). Regarding claims 1, 13, 14, Jing et al. disclose an image assessment method, comprising, an electronic device, comprising: one or more processors; and a storage device configured to store one or more programs, the one or more programs, when executed by the one or more processors, causing the one or more processors to implement a method comprising, and non-transitory computer-readable storage medium having thereon stored a computer program which, when executed by a processor, implements a method comprising ([0011]): acquiring an image to be assessed (The image to be evaluated is input into the image quality assessment model to obtain the evaluation scores of N different distortion types of the image to be evaluated, [0006]); and inputting the image to be assessed into an image assessment model to obtain a quality assessment result corresponding to the image to be assessed, wherein the image assessment model comprises: a multilevel transformation network, a fusion network and a fully connected layer, the multilevel transformation network being configured for processing the image to be assessed to obtain image features output by each level of transformation network (output features after the first processing of each layer are upsampled to obtain the output features after the second processing of each layer, [0043], Accordingly, before performing feature fusion on the output features of each layer of the feature extraction layer, the output features of each layer can be processed to have the same number of channels and the same feature map size, [0047], output features of each layer of the feature extraction layer can be processed, [0048], Image features are processed based on an attention mechanism to determine the feature dependencies of image features in the spatial and/or channel dimensions, [0054]), the fusion network being configured for fusing the image features output by the each level of transformation network to obtain a fused image feature (feature enhancement processing may include, but is not limited to, feature fusion processing and/or attention mechanism processing, [0039], The output features after the second processing of each layer are subjected to feature fusion processing to obtain fused features, [0044]), and the fully connected layer being configured for processing the fused image feature to obtain the quality assessment result (fractional regression model can include N fully connected layers, each corresponding to a type of distortion, [0032], a fractional regression model could include a fully connected layer with N outputs, each corresponding to a type of distortion, [0033], For example, the fractional regression module may include N (N = number of distortion types) fully connected layers, each corresponding to one distortion type; or, the fractional regression module may have one fully connected layer, which includes N outputs, each corresponding to one distortion type, [0103]). Jing et al. do not explicitly disclose the multilevel transformation network being configured for processing the image to be assessed to obtain image features output by each level of transformation network. Kim et al. teach acquiring an image to be assessed (input image, [0040], [0041]); and inputting the image to be assessed into an image assessment model to obtain a quality assessment result corresponding to the image to be assessed, wherein the image assessment model comprises: a multilevel transformation network, a fusion network and a fully connected layer, the multilevel transformation network being configured for processing the image to be assessed to obtain image features output by each level of transformation network (At step 102, a pyramid of feature maps is generated based on an input image. The input image can be a single image frame with one or more channels (e.g., RGB, monochromatic, etc.). In an embodiment, a feature pyramid network is implemented to process the input image and generate a pyramid of feature maps. As used herein, the pyramid of feature maps refers to a plurality of feature maps at different scales relative to a scale of the input image. The pyramid of feature maps can comprise a number of levels, with each level including one or more feature maps at a particular scale, with the scale (e.g., resolution in a pixel space) increasing when moving from the top to the bottom of the pyramid. In an embodiment, the feature pyramid network is based on a residual network that extracts features of the image. The feature map is then up-sampled and combined with intermediate feature maps from the residual network in order to generate the pyramid of feature maps, [0040], Importantly, the FPN 212 extracts features from the image 202 at different resolution levels, which enables different RoIs 602 to be defined and analyzed using a sliding window 604 associated with a fixed set of RoIs 602, [0124]), the fusion network being configured for fusing the image features output by the each level of transformation network to obtain a fused image feature (the feature maps produced by the CNN for the segmentation mask are concatenated with mean feature maps for all other segmentation masks for the other plane objects before being passed to a second CNN, which produces a refined mask for the plane object, [0036], Once all of the feature maps in the feature map pyramid have been concatenated with the outputs of the stages of the DN 616, [0144], Returning now to FIG. 9A, once the inputs 902 are processed by one or more subsequent layers of the neural network, the output of the last set of ConvAccu modules 222 is a refined mask 920 for each instance of the plane objects 810 detected by the PDN 210. The refined masks 920 are concatenated to generate a refined segmentation mask 930 for the image 202., [0150], mean feature map 1040-1 that is an average of all of the corresponding feature maps 1020, [0153]), and the fully connected layer being configured for processing the fused image feature to obtain the quality assessment result (For example, in an embodiment, the FPN 212 is implemented based on a ResNet-50 residual network architecture as defined in He, et al., “Deep Residual Learning for Image Recognition,” Computer Vision and Pattern Recognition, Dec. 10, 2015, which is incorporated herein by reference in its entirety. It will be appreciated that, in some embodiments, the FPN 212 can be implemented using other deep residual network architectures, such as ResNet-101, [0117], As used herein, a predictor head refers to a neural network that includes one or more layers such as convolution layers or fully connected layers that process the RoI 602 input to generate an output, such as various plane object parameters., [0125], The output of the convolution layer is then processed by one or more fully connected layers that calculate a multi-dimensional feature vector (e.g., a 100-d feature vector). The feature vector can then be provided to the input of one or more additional fully connected layers that are implemented by each branch of the RN 612. The fully connected layer(s) of each branch estimate the various parameters associated with the RoI 602., [0126]). Kim et al. discusses fully connected layers but also indicates use of ResNet architectures which include fully connected layers. Kim et al. does not use the term fusing but the pooling and concatenating are interpreted as a type of fusion. The feature pyramid network that generates a pyramid of feature maps comprising a number of levels is interpreted as a multilevel transformation network. Jing et al. and Kim et al. are in the same art of fusion (Jing et al., [0039]; Kim et al., [0144]). The combination of Kim et al. with Jing et al. will enable obtain image features output by each level. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the transformation of Kim et al. with the invention of Jing et al. as this was known at the time of filing, the combination would have predictable results, and as Kim et al. indicate “A deep neural network (DNN) model includes multiple layers of many connected nodes (e.g., perceptrons, Boltzmann machines, radial basis functions, convolutional layers, etc.) that can be trained with enormous amounts of input data to quickly solve complex problems with high accuracy” ([0112]) and “Discarding certain objects based on the classification of the objects as planar or non-planar enables the PDN 210 to detect an arbitrary number of plane objects in the image 202 and reduces the computational load of the PDN 210 by avoiding the unnecessary execution of the CNN 614 in cases where the detected object has a low confidence score of being planar.” (0135]) providing an efficiency benefit to combining inventions. Regarding claims 2 and 17, Jing et al. and Kim et al. disclose the method and electronic device according to claims 1 and 13. Kim et al. further teach the image assessment model further comprises: a sliding window; the sliding window being configured for segmenting the input image to be assessed to obtain a plurality of image blocks and inputting the image blocks to the multilevel transformation network (At 104, regions of interest sampled from the pyramid of feature maps are processed to identify a number of plane object in the input image. In an embodiment, a sliding window is applied to each of the feature maps in the pyramid of feature maps to sample regions of interest. A region of interest can refer to a region of a feature map that corresponds to a particular subset of the input image. While the sliding window can have a fixed size, as applied to a particular feature map at a given scale in the pyramid of feature maps, the region of interest is associated with a variable sized region of the input image. For example, the sliding window can be defined as a 7×7 pixel region relative to a down-sampled size of a particular feature map in the pyramid of feature maps. The 7×7 pixel region in the feature map can corresponds to, e.g., a 14×14 pixel region, a 28×28 pixel region, or a 56×56 pixel region (or larger) in the input image based on the relative difference in scale, in pixel space, of the feature map(s) and the input image, [0041], As used herein, a RoI 602 refers to an anchor bounding box associated with a sliding window 604 applied to one of the feature maps generated by the FPN 212, and each anchor bounding box represents a different scale and aspect ratio for a bounding box centered on the sliding window, [0116]). Regarding claims 3 and 18, Jing et al. and Kim et al. disclose the method and electronic device according to claims 1 and 13. Jing et al. and Kim et al. further indicate wherein the method further comprises: dividing the image to be assessed to obtain a plurality of sub-images to be assessed; for one or more of the sub-images to be assessed, inputting the sub-image to be assessed into the image assessment model to obtain a quality assessment result corresponding to the sub-image to be assessed; and determining the quality assessment result of the image to be assessed based on the quality assessment results corresponding to the one or more of the sub-images to be assessed (Jing et al., quality assessment, [0001]; The sliding window is then moved over the feature maps for the image at each scale, computing a fixed-size feature vector for each of the region of interest associated with the current sliding window location, [0033], While the sliding window can have a fixed size, as applied to a particular feature map at a given scale in the pyramid of feature maps, the region of interest is associated with a variable sized region of the input image. For example, the sliding window can be defined as a 7×7 pixel region relative to a down-sampled size of a particular feature map in the pyramid of feature maps. The 7×7 pixel region in the feature map can corresponds to, e.g., a 14×14 pixel region, a 28×28 pixel region, or a 56×56 pixel region (or larger) in the input image based on the relative difference in scale, in pixel space, of the feature map(s) and the input image, [0041]) [finding feature vectors for blocks of Kim + quality assessment of Jing together teach the claim limitations]. Claim(s) 4 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jing et al. (IDS: CN114299358A [Machine Translation]) and Kim et al. (US 20200167943 A1) as applied to claims 3 and 18 above, further in view of Woodard et al. (US 20080273743 A1). Regarding claims 4 and 19, Jing et al. and Kim et al. disclose the method and electronic device according to claims 3 and 18. Kim et al. partly indicate the determining the quality assessment result of the image to be assessed based on the quality assessment results corresponding to the plurality of sub-images to be assessed, comprises: calculating a harmonic mean corresponding to the quality assessment results corresponding to the plurality of sub-images to be assessed; and determining the harmonic mean as the quality assessment result of the image to be assessed (feature map 1020-1 is then combined with a mean feature map 1040-1 that is an average of all of the corresponding feature maps 1020, [0153]) however another reference is added to make this explicit. Woodard et al. teach calculating a harmonic mean corresponding to the quality assessment results corresponding to the plurality of sub-images to be assessed; and determining the harmonic mean as the quality assessment result of the image to be assessed (The SFM can be computed by taking the ratio of the harmonic to the arithmetic mean of the magnitude-squared, discrete spectrum. The two means involved in the ratio are special cases of the generalized mean (GM). Below, the application of GMs in image processing is briefly reviewed. Next, a mathematical formulation is developed where ratios of linear combinations of GMs define a class of full-reference image quality measures, [0052], One of the first applications of GMs was as a type of nonlinear filter to remove impulse noise. GMs have also been used to compare the illumination conditions of two images. Ratios of GM in both spatial and frequency domains have been shown to be useful as no-reference measures of image quality. However, other reported image processing applications of GMs are rare. GMs also appear infrequently in other engineering applications, although they have been sparingly used in pattern recognition studies, which is beyond the scope of this specification, [0053]). Jing et al. and Woodard et al. are in the same art of quality assessment (Jing et al., [0001]; Woodard et al., [0052]). The combination of Woodard et al. with Jing et al. and Kim et al. will enable calculating a harmonic mean. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the calculating a harmonic mean of Woodard et al. with the invention of Jing et al. and Kim et al. as this was known at the time of filing, the combination would have predictable results, and as Woodard et al. indicate this is a good way to measure image quality ([0053]) which will improve the quality analysis of Jing et al. Claim(s) 5 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jing et al. (IDS: CN114299358A [Machine Translation]) and Kim et al. (US 20200167943 A1) as applied to claims 1 and 13 above, further in view of Wu et al. (IDS: CN109727246A [Machine Translation]). Regarding claims 5 and 20, Jing et al. and Kim et al. disclose the method and electronic device according to claims 1 and 13 above. Jing et al. and Kim et al. do not explicitly disclose the image assessment model is trained by: inputting a first sample image in a sample image set to a first branch network comprised in a model to be trained to obtain a first quality assessment result; inputting a second sample image in the sample image set to a second branch network comprised in the model to be trained to obtain a second quality assessment result, wherein the first branch network and the second branch network are twin networks with a same structure, and the first branch network and the second branch network each comprises a multilevel transformation network, a fusion network and a fully connected layer; and training the model to be trained based on the first quality assessment result, the second quality assessment result, labeling information corresponding to the first sample image and labeling information corresponding to the second sample image to obtain a trained image assessment model. Wu et al. teach inputting a first sample image in a sample image set to a first branch network comprised in a model to be trained to obtain a first quality assessment result; inputting a second sample image in the sample image set to a second branch network comprised in the model to be trained to obtain a second quality assessment result, wherein the first branch network and the second branch network are twin networks with a same structure, and the first branch network and the second branch network each comprises a multilevel transformation network, a fusion network and a fully connected layer; and training the model to be trained based on the first quality assessment result, the second quality assessment result, labeling information corresponding to the first sample image and labeling information corresponding to the second sample image to obtain a trained image assessment model (Step S21: Design a twin convolutional neural network structure. The network consists of two sub-networks: sub-network I and sub-network II. Sub-network I consists of two identical branch structures, and the two branch structures share weights. Each branch structure consists of N stacked convolutional structures. The task of sub-network I is to extract features from two input image patches. Sub-network II consists of M fully connected layers. The features extracted by sub-network I are fused, and the fused features are used as the input of sub-network II. Sub-network II distinguishes the quality of the two input images based on the fused features. Step S22: The Siamese convolutional neural network uses N stacked convolutional layers to abstract and learn image information, and then extracts image features through two fully connected layers. At the same time, it is input into a classification network for quality evaluation score optimization learning. The task of the classification network is to distinguish the quality of two input image patches. That is, the final output of the classification network is the probability of the quality of the two input image patches. The image patch with the higher probability is of better quality than the image patch with the lower probability. Step S23: During the training phase, cross-entropy is used as the loss function, [0021]-[0023]). Jing et al. and Wu et al. are in the same art of quality assessment (Jing et al., [0001]; Wu et al., [0022]). The combination of Wu et al. with Jing et al. and Kim et al. will enable labeling information corresponding to the first sample image and labeling information corresponding to the second sample image to obtain a trained image assessment model. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the configuration of Wu et al. with the invention of Jing et al. and Kim et al. as this was known at the time of filing, the combination would have predictable results, and as Wu et al. indicate “The purpose of this invention is to provide a contrastive learning image quality assessment method based on Siamese networks, which is beneficial for improving the performance of image quality assessment without reference” ([0007]) which will improve the performance of the combination of inventions. Claim(s) 6-7 and 21-22 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jing et al. (IDS: CN114299358A [Machine Translation]) and Kim et al. (US 20200167943 A1) and Wu et al. (IDS: CN109727246A [Machine Translation]) as applied to claims 5 and 20 above, further in view of Brandt et al. (US 20240394874 A1). Regarding claims 6 and 21, Jing et al. and Kim et al. and Wu et al. disclose the method and electronic device according to claims 5 and 20 above. Jing et al. and Kim et al. and Wu et al. do not explicitly disclose the method further comprises: for at least one sample image in the sample image set, zooming the sample image to obtain a sample image with a preset resolution; and preprocessing the sample image with the preset resolution by using an enhancement strategy, the enhancement strategy being configured for improving richness of the sample image set. Brandt et al. teach for at least one sample image in the sample image set, zooming the sample image to obtain a sample image with a preset resolution; and preprocessing the sample image with the preset resolution by using an enhancement strategy, the enhancement strategy being configured for improving richness of the sample image set (trained modular neural networks on OCT data with manual quality grading to analyze image quality features, abstract, Common data augmentation as shearing, shifts, rotation, zoom and flipping was applied to all model datasets after splitting them into train, validate and test sets. We utilized augmentation to compensate for device and class imbalance in the data as well as to further increase data variability up to 15-fold it original class size, [0094]). Jing et al. and Brandt et al. are in the same art of quality assessment (Jing et al., [0001]; Brandt et al., abstract). The combination of Brandt et al. with Jing et al. and Kim et al. and Wu et al. will enable zooming samples. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the zooming of Brandt et al. with the invention of Jing et al. and Kim et al. and Wu et al. as this was known at the time of filing, the combination would have predictable results, and as Brandt et al. indicate this will “increase data variability up to 15-fold it original class size” ([0094]) which will improve the training in the combination of inventions. Regarding claims 7 and 22, Jing et al. and Kim et al. and Wu et al. and Brandt et al. disclose the method according to claims 6 and 21. Brandt et al. further indicate the preprocessing the sample image with the preset resolution by using the enhancement strategy, comprises at least one of: rotating the sample image with the preset resolution by a preset angle; or converting the sample image with the preset resolution into a set color space (trained modular neural networks on OCT data with manual quality grading to analyze image quality features, abstract, Common data augmentation as shearing, shifts, rotation, zoom and flipping was applied to all model datasets after splitting them into train, validate and test sets. We utilized augmentation to compensate for device and class imbalance in the data as well as to further increase data variability up to 15-fold it original class size, [0094]). Claim(s) 8 and 23 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jing et al. (IDS: CN114299358A [Machine Translation]) and Kim et al. (US 20200167943 A1) and Wu et al. (IDS: CN109727246A [Machine Translation]) and Brandt et al. (US 20240394874 A1) as applied to claims 7 and 22 above, further in view of Cheng et al. (US 20210312629 A1). Regarding claims 8 and 23, Jing et al. and Kim et al. and Wu et al. and Brandt et al. disclose the method and electronic device according to claims 7 and 22 above. Jing et al. and Kim et al. and Wu et al. and Brandt et al. do not disclose the set color space comprises one or more of: an RGB color space, an HSV color space, an LAB color space, and a Grayscale color space. Cheng et al. teach the set color space comprises one or more of: an RGB color space, an HSV color space, an LAB color space, and a Grayscale color space (determining a quality of the chest image, [0007], A plurality of CXR sample images may be prepared to serve as a training dataset for the rib segmentation model. The preparation may include, for example, discarding images that are of poor quality, reformatting the images into a suitable format, converting color images to grayscale, resizing the images into unified dimensions, and/or the like. The training images may then be provided (e.g., as an input) to the rib segmentation model, [0078]). Jing et al. and Cheng et al. are in the same art of quality assessment (Jing et al., [0001]; Cheng et al., [0007]). The combination of Cheng et al. with Jing et al. and Kim et al. and Wu et al. and Brandt et al. will enable converting color images to grayscale. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the converting color images to grayscale of Cheng et al. with the invention of Jing et al. and Kim et al. and Wu et al. and Brandt et al. as this was known at the time of filing, the combination would have predictable results, and as improving the quality of the training data will improve the training in the combination of inventions. Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jing et al. (IDS: CN114299358A [Machine Translation]) and Kim et al. (US 20200167943 A1) and Wu et al. (IDS: CN109727246A [Machine Translation]) as applied to claim 5 above, further in view of Kolouri et al. (US 20180307936 A1). Regarding claim 9, Jing et al. and Kim et al. and Wu et al. disclose the method and electronic device according to claim 5. Jing et al. and Kim et al. and Wu et al. do not explicitly disclose if the labeling information corresponding to each sample image in the sample image set is unevenly distributed, performing weighted upsampling on the labeling information corresponding to each sample image to obtain labeling information with a weight. Kolouri et al. teach if the labeling information corresponding to each sample image in the sample image set is unevenly distributed, performing weighted upsampling on the labeling information corresponding to each sample image to obtain labeling information with a weight (Separately, the components 314 and 316 localize the object class within the image 304. The responses from the layers of the CNN 308 are up-sampled to create a collection of up-sampled responses 314. The up-sampled responses 314 are combined with the classification weights 312 to generate a linear combination of up-sampled responses 316. The up-sampled responses 314 are combined with the classification weights 312 by weighted averaging (i.e., linear combination) with respect to the classification weights 312. The weighted combination of up-sampled responses 314 results in a localization heatmap 302. Further details regarding these processes are provided below. Specifically, provided below is a description of prior art followed by a detailed description of the machine-vision system for discriminant localization according to the present disclosure, [0048]). Jing et al. and Kolouri et al. are in the same art of neural networks (Jing et al., [0066]; Kolouri et al., [0048]). The combination of Kolouri et al. with Jing et al. and Kim et al. and Wu et al. will enable performing weighted upsampling. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the performing weighted upsampling of Kolouri et al. with the invention of Jing et al. and Kim et al. and Wu et al. as this was known at the time of filing, the combination would have predictable results, and as Kolouri et al. indicate “Further, the system of this disclosure is computationally efficient and capable of achieving localization in one step and hence is orders of magnitude faster than the prior art” ([0043]) demonstrating a computational efficiency benefit to combining inventions. Claim(s) 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jing et al. (IDS: CN114299358A [Machine Translation]) and Kim et al. (US 20200167943 A1) and Wu et al. (IDS: CN109727246A [Machine Translation]) as applied to claim 5 above, further in view of Lin et al. (“Regression Guided by Relative Ranking Using Convolutional Neural Network (R3CNN) for Facial Beauty Prediction”). Regarding claim 10, Jing et al. and Kim et al. and Wu et al. and Wu et al. disclose the method and electronic device according to claim 5. Jing et al. and Kim et al. and Wu et al. do not explicitly disclose the training the model to be trained based on the first quality assessment result, the second quality assessment result, the labeling information corresponding to the first sample image, and the labeling information corresponding to the second sample image to obtain the trained image assessment model, comprises: training the model to be trained by using a joint loss function based on the first quality assessment result, the second quality assessment result, the labeling information corresponding to the first sample image, and the labeling information corresponding to the second sample image to obtain the trained image assessment model, wherein the joint loss function comprises a regression loss function and a rank loss function, the regression loss function being configured for measuring a difference between the first quality assessment result and the labeling information corresponding to the first sample image, and the rank loss function being configured for measuring a relative quality between the first sample image and the second sample image. Lin et al. teach the training the model to be trained based on the first quality assessment result, the second quality assessment result, the labeling information corresponding to the first sample image, and the labeling information corresponding to the second sample image to obtain the trained image assessment model, comprises: training the model to be trained by using a joint loss function based on the first quality assessment result, the second quality assessment result, the labeling information corresponding to the first sample image, and the labeling information corresponding to the second sample image to obtain the trained image assessment model, wherein the joint loss function comprises a regression loss function and a rank loss function, the regression loss function being configured for measuring a difference between the first quality assessment result and the labeling information corresponding to the first sample image, and the rank loss function being configured for measuring a relative quality between the first sample image and the second sample image (To tackle this problem, we propose the following learning schemes for R3CNN: 1) a hard pair sampling strategy that generates challenging to predicted image pairs and pseudo ranking labels from true rating scores; 2) an assemble loss function that combines regression loss and pairwise ranking loss (PR-Loss), where PR-Loss can be a hinge-form loss or a log-sum-exp pairwise loss; 3) a cascaded fine-tuning method that further improves prediction, abstract, In this paper, since Siamese network is used as a ranking component of R3CNN, we propose an offline hard pair sampling strategy for R3CNN by collecting hard pairs from FBP dataset before the model training starts. Further, we also develop two forms of pairwise loss functions based on hinge form loss [41] and log-sum-exp pairwise loss [42], respectively, p124 PNG media_image1.png 810 414 media_image1.png Greyscale p127). Jing et al. and Lin et al. are in the same art of neural networks (Jing et al., [0066]; Lin et al., abstract). The combination of Lin et al. with Jing et al. and Kim et al. and Wu et al. will enable joint loss function comprising a regression loss function and a rank loss function. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the joint loss function of Lin et al. with the invention of Jing et al. and Kim et al. and Wu et al. as this was known at the time of filing, the combination would have predictable results, and as Lin et al. indicate “with related CNN models highlight the effectiveness of the R3CNN architecture for FBP” (abstract) and “Therefore, this paper aims to construct a general regression framework guided by relative ranking with a set of efficient learning schemes for FBP” (p123) demonstrating a discernment effectiveness and efficiency benefit to combining inventions. Allowable Subject Matter Claim 11 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following art is cited as relevant but not sufficient to disclose, teach or fairly suggest the subject matter of claim 11 in entirety even when combined with primary and secondary references: US 20240161318 A1: 3. The method of claim 2, wherein the joint loss includes an unsupervised decomposition loss measuring a difference between an input image and a reconstructed image based on object masks reconstructed by a decoder of the decomposition neural network. 10. The method of claim 7, wherein the decomposition neural network, the alignment neural network, and the transition neural network have been jointly trained to minimize a joint loss, and the joint loss includes an unsupervised alignment loss, the unsupervised alignment loss including a reconstruction loss that measures a difference between an output of the transition neural network for the current time point based on aligned historical feature representations and aligned current feature representations generated by applying the adjacency matrix to a set of current feature representations; “Learning and Transferring Deep Joint Spectral–Spatial Features for Hyperspectral Classification”: PNG media_image2.png 160 408 media_image2.png Greyscale PNG media_image3.png 176 404 media_image3.png Greyscale PNG media_image4.png 276 706 media_image4.png Greyscale PNG media_image5.png 342 406 media_image5.png Greyscale Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHELLE ENTEZARI whose telephone number is (571)270-5084. The examiner can normally be reached 10-7 M-F. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent M Rudolph can be reached at (571) 272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MICHELLE M ENTEZARI HAUSMANN/Primary Examiner, Art Unit 2671
Read full office action

Prosecution Timeline

Oct 16, 2024
Application Filed
Sep 10, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725301
SEMANTIC VISUAL FEATURE SHARING
2y 11m to grant Granted Sep 01, 2026
Patent 12718585
ACCURACY FOR OBJECT DETECTION
2y 11m to grant Granted Aug 25, 2026
Patent 12711781
IDENTIFYING BIDIRECTIONAL CHANNELIZATION ZONES AND LANE DIRECTIONALITY
3y 5m to grant Granted Aug 18, 2026
Patent 12700257
CASCADED DETECTION OF FACIAL ATTRIBUTES
3y 0m to grant Granted Aug 04, 2026
Patent 12688722
SIMULATION OF LABEL DATA TO OPTIMIZE THE VISUAL DOCUMENT UNDERSTANDING BY USING PDFS ANNOTATION AWARE METHODOLOGY
2y 10m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
77%
Grant Probability
98%
With Interview (+21.1%)
2y 12m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 883 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month