Prosecution Insights
Last updated: October 01, 2026
Application No. 18/726,881

Machine Learning Models Featuring Resolution-Flexible Multi-Axis Attention Blocks

Final Rejection §103
Filed
Jul 05, 2024
Priority
Jan 05, 2022 — provisional 63/296,625 +1 more
Examiner
SHIMELES, BEZAWIT NOLAWI
Art Unit
2673
Tech Center
2600 — Communications
Assignee
Google LLC
OA Round
2 (Final)
89%
Grant Probability
Favorable
3-4
OA Rounds
4m
Est. Remaining
89%
With Interview

Examiner Intelligence

Grants 89% — above average
89%
Career Allowance Rate
8 granted / 9 resolved
+26.9% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
26 currently pending
Career history
31
Total Applications
across all art units

Statute-Specific Performance

§101
12.1%
-27.9% vs TC avg
§103
64.9%
+24.9% vs TC avg
§102
4.9%
-35.1% vs TC avg
§112
9.1%
-30.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 9 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 09/14/2026 has been considered by the examiner. Response to Amendment Applicant`s remarks, see page 8, filed 08/05/2026, with respect to rejections of claims 1-14 under 35 U.S.C. 112(b) due to claim language that rendered independent claim 1 as being indefinite, submitted in the non-final office action dated 05/05/2026, have been fully considered and are persuasive due to amendments in accordance with Examiner’s suggested corrections to recite a processor as carrying out the function(s). Thus, these rejections of claims 1-14 under 35 U.S.C. 112(b) have been withdrawn. Applicant`s remarks, see page 8, filed 08/05/2026, with respect to rejections of claims 15-20 under 35 U.S.C. 112(a) and 35 U.S.C. 112(b) for failing to disclose corresponding structure that supports claim limitations that invoked 35 U.S.C. 112(f), submitted in the non-final office action dated 05/05/2026, have been fully considered and are persuasive due to amendments in accordance with Examiner’s suggested corrections of not reciting terms that invoke interpretations under 35 U.S.C. 112(f). Thus, these rejections of claims 15-20 under 35 U.S.C. 112(a) and 35 U.S.C. 112(b) have been withdrawn. Response to Arguments Applicant’s arguments, see remarks, filed 08/05/2026, with respect to claims 1-20, have been fully considered, but are moot because the arguments do not apply to the current references and current combinations of references being used in the current rejection. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-4, 9, and 13-18 are rejected under 35 U.S.C. 103 as being unpatentable over CHEN (US 20230184927 A1), hereinafter referenced as CHEN in view of CRICRÌ (US 20240289590 A1), hereinafter referenced as CRICRÌ in further view of HE (US 20160104056 A1), hereinafter referenced as HE. Regarding claim 1, CHEN teaches a computing system for resolution-flexible image processing (Fig. 2, Paragraph [0068] – CHEN discloses a target detection device [wherein a target detection device is a computing system], including a memory, a processor, and a computer program stored in the memory and operable on the processor. Paragraph [0095] – CHEN further discloses a two-way multi-scale connection operation is enhanced through top-down and bottom-up attention, to guide learning of dynamic attention matrices and enhance feature interaction under different resolutions. See also Paragraphs [0069-0074].), the computing system comprising: one or more processors (Fig. 2, Paragraph [0068] – CHEN discloses a target detection device, including a memory, a processor, and a computer program stored in the memory and operable on the processor.); and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors (Fig. 2, Paragraph [0042] – CHEN discloses the present disclosure further provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program, when executed by a processor, performs steps of the method described above.), cause the computing system (Fig. 2, Paragraph [0068] – CHEN discloses a target detection device [wherein a target detection device is a computing system]) to execute: a machine-learned image processing model configured to process input image data (Fig. 2, Paragraph [0074] – CHEN discloses introducing a cross-resolution attention enhancement neck CAENeck to the model framework CRTransSar [wherein model framework is a machine-learned image processing model]. See also Paragraph [0077].) to generate an output prediction (Fig. 2, Paragraph [0025] – CHEN discloses sending a processing result into a fully connected RCNN network for classification and regression, to position and recognize the target, so as to obtain a final detection result [wherein a final detection result is an output prediction].), wherein the machine-learned image processing model comprises one or more resolution-flexible multi-axis attention blocks (Fig. 2, Paragraph [0085] – CHEN discloses based on multi-head attention, the present disclosure takes into account the CotNet contextual attention mechanism and integrates the self-attention module block into the Swin transformer. Paragraph [0095] – CHEN further discloses a two-way multi-scale connection operation is enhanced through top-down and bottom-up attention, to guide learning of dynamic attention matrices and enhance feature interaction under different resolutions.), each of the one or more resolution-flexible multi-axis attention blocks (Fig. 2, Paragraph [0085]) comprising: and a local processing branch (Fig. 5, Paragraph [0052] – CHEN discloses FIG. 5 shows a self-attention module block.) configured to: perform a second partitioning operation to partition at least a second portion of the input tensor into a plurality of second feature sets (Fig. 4, Paragraph [0081] – CHEN discloses the image is divided into tokens similar to those in NLP. Paragraph [0088] – CHEN further discloses a set of non-overlapping patches [wherein a set of non-overlapping patches are a plurality of feature sets] which are divided from the inputted picture with dimensions of H×W×3 through patch partition processing, so as to reduce the size of the feature map and send the feature map to the Swin transformer block for processing.); Although CHEN further teaches and perform a respective local attention operation on a second axis of each of the plurality of second feature sets (Fig. 4, Paragraph [0081] – CHEN further discloses a layered Transformer is introduced, whose representation is calculated by moving the window [wherein moving the window is performing local attention operation on a second axis]. Paragraph [0082] – CHEN further discloses the Swin transformer performs self-attention calculation in each window.). CHEN fails to explicitly teach a global processing branch configured to: perform a first partitioning operation to partition at least a first portion of an input tensor of the resolution-flexible multi-axis attention block into a plurality of first feature sets, and perform a global attention operation along a first axis of the plurality of first feature sets, wherein the first axis corresponds to the predefined number of the plurality of first feature sets; However, CRICRÌ explicitly teaches a global processing branch (Fig. 14, Paragraph [0371] – CRICRÌ discloses for each group of feature maps, a squeeze-and-attention operation [wherein squeeze-and-attention operation is a global processing branch] may be applied. See also Paragraph [0372]) configured to: perform a first partitioning operation to partition at least a first portion of an input tensor of the resolution-flexible multi-axis attention block into a plurality of first feature sets (Fig. 14, Paragraph [0393] - CRICRÌ discloses the ResNeSt block may comprise a splitting operation, which divides or splits a tensor that is input to the core attention block into K sub-tensors or groups of features.), and perform a global attention operation along a first axis of the plurality of first feature sets (Fig. 14, Paragraph [0371] - CRICRÌ discloses the ResNeSt block is an attention block where the input tensor is divided into K smaller tensors (also referred to as groups of feature maps). For each group of feature maps, a squeeze-and-attention operation may be applied. See also Paragraphs [0366-0368, 0372-0378].), wherein the first axis corresponds to the predefined number of the plurality of first feature sets (Fig. 14, Paragraph [0371] - CRICRÌ discloses the ResNeSt block is an attention block where the input tensor is divided into K smaller tensors (also referred to as groups of feature maps) [wherein K smaller tensors is the predefined number of the plurality of first feature sets]. For each group of feature maps, a squeeze-and-attention operation may be applied. See also Paragraphs [0366-0368, 0372-0378].); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN of having a computing system for resolution-flexible image processing, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to execute: a machine-learned image processing model configured to process input image data to generate an output prediction, wherein the machine-learned image processing model comprises one or more resolution-flexible multi-axis attention blocks, each of the one or more resolution-flexible multi-axis attention blocks comprising: and a local processing branch configured to: perform a second partitioning operation to partition at least a second portion of the input tensor into a plurality of second feature sets; and perform a respective local attention operation on a second axis of each of the plurality of second feature sets, with the teachings of CRICRÌ having a global processing branch configured to: perform a first partitioning operation to partition at least a first portion of an input tensor of the resolution-flexible multi-axis attention block into a plurality of first feature sets, and perform a global attention operation along a first axis of the plurality of first feature sets, wherein the first axis corresponds to the predefined number of the plurality of first feature sets. Wherein CHEN’s computing system for resolution-flexible image processing wherein having a global processing branch configured to: perform a first partitioning operation to partition at least a first portion of an input tensor of the resolution-flexible multi-axis attention block into a plurality of first feature sets, and perform a global attention operation along a first axis of the plurality of first feature sets, wherein the first axis corresponds to the predefined number of the plurality of first feature sets. The motivation behind this modification would have been to provide an improved system for image processing that has the flexibility of modeling at various scales and enhancing visual quality, since both CHEN and CRICRÌ relate to architectures where multiple neural networks are connected in parallel or in series and adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and CRICRÌ relates to a method, apparatus, and computer program product for providing an attention block for neural network-based image and video compression, wherein the proposed block achieves a better rate-distortion performance, multiscale structural similarity (MS-SSIM) is higher (e.g., better visual quality) and bits per pixel (BPP) is lower (e.g., lower bitrate). Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and CRICRÌ (US 20240289590 A1), Paragraphs [0187, 0399]. CHEN in view of CRICRÌ fail to explicitly teach wherein the first partitioning operation comprises a grid operation that partitions the input tensor into a fixed grid having a predefined number of the plurality of first feature sets irrespective of a resolution of the input tensor; However, HE explicitly teaches wherein the first partitioning operation comprises a grid operation that partitions the input tensor into a fixed grid (Fig. 2, Paragraph [0025] – HE discloses convolutional layers 204 may accept arbitrary input image 202 sizes, but they produce outputs of variable sizes. The spatial bins may have sizes proportional to the image size, so the number of bins may be fixed regardless of the image size. See also Fig. 3, Paragraph [0035].) having a predefined number of the plurality of first feature sets irrespective of a resolution of the input tensor (Fig. 3, Paragraph [0027] – HE discloses with spatial pyramid pooling, the input image may be of any size allowing not only arbitrary aspect ratios, but also arbitrary scales. Paragraph [0035] – HE discloses an SPP network 312 of one or more layers that includes spatial bins based on the number of filters a top convolutional layer may pool the extracted features and generate fixed-size outputs for a fully-connected layer 314. See also Paragraph [0036].); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN in view of CRICRÌ of having a computing system for resolution-flexible image processing, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to execute: a machine-learned image processing model configured to process input image data to generate an output prediction, wherein the machine-learned image processing model comprises one or more resolution-flexible multi-axis attention blocks, each of the one or more resolution-flexible multi-axis attention blocks comprising: a global processing branch configured to: perform a first partitioning operation to partition at least a first portion of an input tensor of the resolution-flexible multi-axis attention block into a plurality of first feature sets, and perform a global attention operation along a first axis of the plurality of first feature sets, wherein the first axis corresponds to the predefined number of the plurality of first feature sets; and a local processing branch configured to: perform a second partitioning operation to partition at least a second portion of the input tensor into a plurality of second feature sets; and perform a respective local attention operation on a second axis of each of the plurality of second feature sets, with the teachings of HE having wherein the first partitioning operation comprises a grid operation that partitions the input tensor into a fixed grid having a predefined number of the plurality of first feature sets irrespective of a resolution of the input tensor. Wherein CHEN’s computing system for resolution-flexible image processing wherein the first partitioning operation comprises a grid operation that partitions the input tensor into a fixed grid having a predefined number of the plurality of first feature sets irrespective of a resolution of the input tensor. The motivation behind this modification would have been to provide an improved system for image processing that has the flexibility of modeling at various scales and reduces computational expense, since both CHEN and HE relate to architectures where neural networks are adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and HE relates to methods, devices, and systems to process images using spatial pyramid pooling networks; a system employing SPP-network to process images may advantageously improve usability of object detection in searches, vision systems, and other image analysis implementations, as well as reduce computational expense such as processor load, memory load, and enhance reliability of object detection, for example, in satellite imaging, security monitoring, and comparable systems. Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and HE (US 20160104056 A1), Paragraph [0039]. Regarding claim 2, CHEN and CRICRÌ in view of HE teach the computing system of claim 1, Although CHEN further teaches wherein the input tensor has a height, a width, and a channel depth (Fig. 2, Paragraph [0088] – CHEN discloses each feature map [wherein feature map is the input tensor] has dimensions of 2×3×H×W [wherein 3 is channel depth, H is height, W is width] when sent to the PatchEmbed.), CHEN fails to explicitly teach and wherein the first partitioning operation comprises a grid partitioning operation that partitions the height and width of the input tensor into a grid having the predefined number of the plurality of first feature sets. However, CRICRÌ explicitly teaches and wherein the first partitioning operation comprises a grid partitioning operation (Fig. 14, Paragraph [0393] - CRICRÌ discloses the ResNeSt block may comprise a splitting operation [wherein a splitting operation is a grid partitioning operation], which divides or splits a tensor that is input to the core attention block into K sub-tensors or groups of features.) that partitions the height and width of the input tensor into a grid having the predefined number of the plurality of first feature sets (Fig. 14, Paragraph [0371] - CRICRÌ discloses the ResNeSt block is an attention block where the input tensor is divided into K smaller tensors (also referred to as groups of feature maps) [wherein K smaller tensors is the predefined number of the plurality of first feature sets]. For each group of feature maps, a squeeze-and-attention operation may be applied. See also Paragraphs [0366-0368].). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computing system for resolution-flexible image processing, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to execute: a machine-learned image processing model configured to process input image data to generate an output prediction, wherein the machine-learned image processing model comprises one or more resolution-flexible multi-axis attention blocks, each of the one or more resolution-flexible multi-axis attention blocks comprising: a global processing branch configured to: perform a first partitioning operation to partition at least a first portion of an input tensor of the resolution-flexible multi-axis attention block into a plurality of first feature sets, with the teachings of CRICRÌ having and wherein the first partitioning operation comprises a grid partitioning operation that partitions the height and width of the input tensor into a grid having the predefined number of the plurality of first feature sets. Wherein CHEN’s computing system for resolution-flexible image processing and wherein the first partitioning operation comprises a grid partitioning operation that partitions the height and width of the input tensor into a grid having the predefined number of the plurality of first feature sets. The motivation behind this modification would have been to provide an improved system for image processing that has the flexibility of modeling at various scales and enhancing visual quality, since both CHEN and CRICRÌ relate to architectures where multiple neural networks are connected in parallel or in series and adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and CRICRÌ relates to a method, apparatus, and computer program product for providing an attention block for neural network-based image and video compression, wherein the proposed block achieves a better rate-distortion performance, multiscale structural similarity (MS-SSIM) is higher (e.g., better visual quality) and bits per pixel (BPP) is lower (e.g., lower bitrate). Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and CRICRÌ (US 20240289590 A1), Paragraphs [0187, 0399]. Regarding claim 3, CHEN and CRICRÌ in view of HE teach the computing system of claim 1, CHEN further teaches wherein the second partitioning operation comprises a block partitioning operation (Fig. 6, Paragraph [0088] – CHEN discloses the PatchEmbed module mainly increases the channel dimensions of each patch in a set of non-overlapping patches which are divided from the inputted picture with dimensions of H×W×3 through patch partition processing [wherein patch partition processing is a block partitioning operation]) that partitions the second portion of the input tensor into the plurality of second feature sets such that each of the plurality of second feature sets has a predefined height and width (Fig. 6, Paragraph [0088] – CHEN discloses each feature map has dimensions of 2×3×H×W when sent to the PatchEmbed, and has dimensions of 2×96×H/4×W/4 when finally sent to the next module [wherein H/4×W/4 is a predefined height and width].). Regarding claim 4, CHEN and CRICRÌ in view of HE the computing system of claim 3, PNG media_image1.png 385 883 media_image1.png Greyscale Annotated cropped image of Fig. 6 (CHEN) Although CHEN further teaches wherein the number of the plurality of first feature sets is equal to a multiplication of the predefined height and width of each of the plurality of second feature sets (Fig. 6, Paragraph [0085] – CHEN discloses it is determined whether the width and height of the feature map are each an integer multiple of 4; [See annotated cropped image of Fig. 6 above where CHEN teaches a multiplication of the width and height of each feature set is equal to the number of the plurality of feature sets].) CHEN fails to explicitly teach predefined number. However, CRICRÌ explicitly teaches predefined number (Fig. 14, Paragraph [0371] - CRICRÌ discloses the ResNeSt block is an attention block where the input tensor is divided into K smaller tensors (also referred to as groups of feature maps) [wherein K smaller tensors is the predefined number of the plurality of first feature sets]. For each group of feature maps, a squeeze-and-attention operation may be applied. See also Paragraphs [0366-0368].). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computing system for resolution-flexible image processing, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to execute: a machine-learned image processing model configured to process input image data to generate an output prediction, wherein the machine-learned image processing model comprises one or more resolution-flexible multi-axis attention blocks, each of the one or more resolution-flexible multi-axis attention blocks comprising: a global processing branch configured to: perform a first partitioning operation to partition at least a first portion of an input tensor of the resolution-flexible multi-axis attention block into a plurality of first feature sets, with the teachings of CRICRÌ having the predefined number. Wherein CHEN’s computing system for resolution-flexible image processing wherein having the predefined number of the plurality of first feature sets. The motivation behind this modification would have been to provide an improved system for image processing that has the flexibility of modeling at various scales and enhancing visual quality, since both CHEN and CRICRÌ relate to architectures where multiple neural networks are connected in parallel or in series and adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and CRICRÌ relates to a method, apparatus, and computer program product for providing an attention block for neural network-based image and video compression, wherein the proposed block achieves a better rate-distortion performance, multiscale structural similarity (MS-SSIM) is higher (e.g., better visual quality) and bits per pixel (BPP) is lower (e.g., lower bitrate). Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and CRICRÌ (US 20240289590 A1), Paragraphs [0187, 0399]. Regarding claim 9, CHEN and CRICRÌ in view of HE teach the computing system of claim 1, CHEN fails to explicitly teach wherein the first portion of the input tensor comprises a first half of a plurality of depth channels of the input tensor and the second portion of the input tensor comprises a second half of the plurality of depth channels of the input tensor. However, CRICRÌ explicitly teaches wherein the first portion of the input tensor comprises a first half of a plurality of depth channels of the input tensor (Fig. 16, Paragraph [0371] - CRICRÌ discloses ResNeSt block is an attention block where the input tensor is divided into K smaller tensors (also referred to as groups of feature maps). Paragraph [0380] - CRICRÌ further discloses the proposed variation is a simplification of the ResNeSt block, in which the input tensor is divided into two groups. Paragraph [0393] - CRICRÌ further discloses the output of such operations may be a tensor with a number of channels that is equal to r times the number of channels of each sub-sub-tensor or split.) and the second portion of the input tensor comprises a second half of the plurality of depth channels of the input tensor (Fig. 16, Paragraph [0371] - CRICRÌ discloses ResNeSt block is an attention block where the input tensor is divided into K smaller tensors (also referred to as groups of feature maps). Paragraph [0380] - CRICRÌ further discloses the proposed variation is a simplification of the ResNeSt block, in which the input tensor is divided into two groups. Paragraph [0393] - CRICRÌ further discloses the output of such operations may be a tensor with a number of channels that is equal to r times the number of channels of each sub-sub-tensor or split.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computing system for resolution-flexible image processing, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to execute: a machine-learned image processing model configured to process input image data to generate an output prediction, wherein the machine-learned image processing model comprises one or more resolution-flexible multi-axis attention blocks, each of the one or more resolution-flexible multi-axis attention blocks comprising: a global processing branch configured to: perform a first partitioning operation to partition at least a first portion of an input tensor of the resolution-flexible multi-axis attention block into a plurality of first feature sets, with the teachings of CRICRÌ having wherein the first portion of the input tensor comprises a first half of a plurality of depth channels of the input tensor and the second portion of the input tensor comprises a second half of the plurality of depth channels of the input tensor. Wherein CHEN’s computing system for resolution-flexible image processing wherein the first portion of the input tensor comprises a first half of a plurality of depth channels of the input tensor and the second portion of the input tensor comprises a second half of the plurality of depth channels of the input tensor. The motivation behind this modification would have been to provide an improved system for image processing that has the flexibility of modeling at various scales and enhancing visual quality, since both CHEN and CRICRÌ relate to architectures where multiple neural networks are connected in parallel or in series and adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and CRICRÌ relates to a method, apparatus, and computer program product for providing an attention block for neural network-based image and video compression, wherein the proposed block achieves a better rate-distortion performance, multiscale structural similarity (MS-SSIM) is higher (e.g., better visual quality) and bits per pixel (BPP) is lower (e.g., lower bitrate). Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and CRICRÌ (US 20240289590 A1), Paragraphs [0187, 0399]. Regarding claim 13, CHEN and CRICRÌ in view of HE teach the computing system of claim 1, Although CHEN further teaches an image recognition prediction (Fig. 1, Paragraph [0062] – CHEN discloses Step 2.4: Obtain a set of suitable proposal boxes Proposal after the foregoing steps, send the received feature maps and the suitable proposal boxes Proposal into ROI pooling for unified processing, and then finally send the received feature maps and the suitable proposal boxes Proposal to a fully connected RCNN network for classification and regression, to position and recognize the image, so as to obtain a final detection result.), or an object recognition prediction (Fig. 1, Paragraph [0067] – CHEN discloses image positioning and recognition module is configured to perform classification and regression on the proposal boxes Proposal, to position and recognize the target, so as to obtain a final detection result.), CHEN fails to explicitly teach wherein the output prediction comprises an image classification prediction, However, CRICRÌ explicitly teaches wherein the output prediction comprises an image classification prediction (Fig. 16, Paragraph [0384] - CRICRÌ discloses the DSA block may be used for a different use case than compression, such as for semantic segmentation or for image classification.), Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computing system for resolution-flexible image processing, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to execute: a machine-learned image processing model configured to process input image data to generate an output prediction, with the teachings of CRICRÌ having wherein the output prediction comprises an image classification prediction. Wherein CHEN’s computing system for resolution-flexible image processing wherein the output prediction comprises an image classification prediction. The motivation behind this modification would have been to provide an improved system for image processing that has the flexibility of modeling at various scales and enhancing visual quality, since both CHEN and CRICRÌ relate to architectures where multiple neural networks are connected in parallel or in series and adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and CRICRÌ relates to a method, apparatus, and computer program product for providing an attention block for neural network-based image and video compression, wherein the proposed block achieves a better rate-distortion performance, multiscale structural similarity (MS-SSIM) is higher (e.g., better visual quality) and bits per pixel (BPP) is lower (e.g., lower bitrate). Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and CRICRÌ (US 20240289590 A1), Paragraphs [0187, 0399]. Regarding claim 14, CHEN and CRICRÌ in view of HE the computing system of claim 1, CHEN fails to explicitly teach wherein the output prediction comprises a restored image, the restored image having been one or more of: denoised, deblurred, derained, dehazed, or retouched. However, CRICRÌ explicitly teaches wherein the output prediction comprises a restored image (Fig. 9, Paragraph [0256] – CRICRÌ discloses the encoder takes an image as an input and produces a code, to represent the input image, which requires less bits than the input image. This code may have been obtained by a binarization or quantization process after the encoder. The decoder takes in this code and reconstructs the image [wherein a reconstructed image is a restored image] which was input to the encoder.), the restored image having been one or more of: denoised (Fig. 9, Paragraph [0249] – CRICRÌ discloses after the feature extraction layers there may be one or more layers performing a certain task, for example, classification, semantic segmentation, object detection, denoising, style transfer, super-resolution, and the like.), deblurred (Fig. 9, Paragraph [0356-0357] – CRICRÌ discloses it is assumed that the decoder comprises one or more neural networks. Some examples of such decoder side neural networks may include the following: A NN post-processing filter, for either an end-to-end learned codec, or for a hybrid codec (a non-learned codec that incorporates one or more learned NN tools), or for a completely non-learned codec. Examples of possible types of post-processing are enhancement of visual quality for humans, enhancement of visual quality for machine analysis or processing, super-resolution, denoising, application of visual effects [wherein enhancement of visual quality can comprise deblurring];), derained (Fig. 9, Paragraph [0356-0357] – CRICRÌ discloses it is assumed that the decoder comprises one or more neural networks. Some examples of such decoder side neural networks may include the following: A NN post-processing filter, for either an end-to-end learned codec, or for a hybrid codec (a non-learned codec that incorporates one or more learned NN tools), or for a completely non-learned codec. Examples of possible types of post-processing are enhancement of visual quality for humans, enhancement of visual quality for machine analysis or processing, super-resolution, denoising, application of visual effects [wherein enhancement of visual quality can comprise deraining];), dehazed (Fig. 9, Paragraph [0356-0357] – CRICRÌ discloses it is assumed that the decoder comprises one or more neural networks. Some examples of such decoder side neural networks may include the following: A NN post-processing filter, for either an end-to-end learned codec, or for a hybrid codec (a non-learned codec that incorporates one or more learned NN tools), or for a completely non-learned codec. Examples of possible types of post-processing are enhancement of visual quality for humans, enhancement of visual quality for machine analysis or processing, super-resolution, denoising, application of visual effects [wherein enhancement of visual quality can comprise deraining];), or retouched (Fig. 9, Paragraph [0356-0357] – CRICRÌ discloses it is assumed that the decoder comprises one or more neural networks. Some examples of such decoder side neural networks may include the following: A NN post-processing filter, for either an end-to-end learned codec, or for a hybrid codec (a non-learned codec that incorporates one or more learned NN tools), or for a completely non-learned codec. Examples of possible types of post-processing are enhancement of visual quality for humans, enhancement of visual quality for machine analysis or processing, super-resolution, denoising, application of visual effects [wherein enhancement of visual quality can comprise retouching];). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computing system for resolution-flexible image processing, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to execute: a machine-learned image processing model configured to process input image data to generate an output prediction, with the teachings of CRICRÌ having wherein the output prediction comprises a restored image, the restored image having been one or more of: denoised, deblurred, derained, dehazed, or retouched. Wherein CHEN’s computing system for resolution-flexible image processing wherein the output prediction comprises a restored image, the restored image having been one or more of: denoised, deblurred, derained, dehazed, or retouched. The motivation behind this modification would have been to provide an improved system for image processing that has the flexibility of modeling at various scales and enhancing visual quality, since both CHEN and CRICRÌ relate to architectures where multiple neural networks are connected in parallel or in series and adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and CRICRÌ relates to a method, apparatus, and computer program product for providing an attention block for neural network-based image and video compression, wherein the proposed block achieves a better rate-distortion performance, multiscale structural similarity (MS-SSIM) is higher (e.g., better visual quality) and bits per pixel (BPP) is lower (e.g., lower bitrate). Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and CRICRÌ (US 20240289590 A1), Paragraphs [0187, 0399]. Regarding claim 15, CHEN teaches a computer-implemented method for image processing (Fig. 1, Paragraph [0056] – CHEN discloses the present disclosure provides a contextual visual-based SAR target detection method.), the method comprising: obtaining an input image (Fig. 1, Paragraph [0057] – CHEN discloses Step 1: Obtain an SAR image.); processing the input image by executing, by one or more processors (Fig. 2, Paragraph [0041] – CHEN discloses a target detection device, including a memory, a processor, and a computer program stored in the memory and operable on the processor, where the processor performs steps of the method described above when running the computer program.), a machine-learned image processing model (Fig. 2, Paragraph [0074] – CHEN discloses introducing a cross-resolution attention enhancement neck CAENeck to the model framework CRTransSar [wherein model framework is a machine-learned image processing model]. See also Paragraph [0077].) to generate an output prediction (Fig. 2, Paragraph [0025] – CHEN discloses sending a processing result into a fully connected RCNN network for classification and regression, to position and recognize the target, so as to obtain a final detection result [wherein a final detection result is an output prediction].), wherein executing the machine-learned image processing model (Fig. 2, Paragraph [0074]) comprises, at each of one or more resolution-flexible multi-axis attention blocks of the machine-learned image processing model (Fig. 2, Paragraph [0085] – CHEN discloses based on multi-head attention, the present disclosure takes into account the CotNet contextual attention mechanism and integrates the self-attention module block into the Swin transformer. Paragraph [0095] – CHEN further discloses a two-way multi-scale connection operation is enhanced through top-down and bottom-up attention, to guide learning of dynamic attention matrices and enhance feature interaction under different resolutions.): and executing a local processing branch of the resolution-flexible multi-axis attention block (Fig. 5, Paragraph [0052] – CHEN discloses FIG. 5 shows a self-attention module block.) comprising: performing a second partitioning operation to partition at least a second portion of the input tensor into a plurality of second feature sets (Fig. 4, Paragraph [0081] – CHEN discloses the image is divided into tokens similar to those in NLP. Paragraph [0088] – CHEN further discloses a set of non-overlapping patches [wherein a set of non-overlapping patches are a plurality of feature sets] which are divided from the inputted picture with dimensions of H×W×3 through patch partition processing, so as to reduce the size of the feature map and send the feature map to the Swin transformer block for processing.); and performing a respective local attention operation on a second axis of each of the plurality of second feature sets (Fig. 4, Paragraph [0081] – CHEN further discloses a layered Transformer is introduced, whose representation is calculated by moving the window [wherein moving the window is performing local attention operation on a second axis]. Paragraph [0082] – CHEN further discloses the Swin transformer performs self-attention calculation in each window.); Although CHEN further teaches and providing the output prediction as an output (Fig. 2, Paragraph [0025] – CHEN discloses sending a processing result into a fully connected RCNN network for classification and regression, to position and recognize the target, so as to obtain a final detection result [wherein a final detection result is the output prediction].). CHEN fails to explicitly teach executing a global processing branch of the resolution-flexible multi-axis attention block comprising: performing a first partitioning operation to partition at least a first portion of an input tensor of the resolution-flexible multi-axis attention block into a plurality of first feature sets, and performing a global attention operation along a first axis of the plurality of first feature sets, wherein the first axis corresponds to the predefined number of the plurality of first feature sets; However, CRICRÌ explicitly teaches executing a global processing branch of the resolution-flexible multi-axis attention block (Fig. 14, Paragraph [0371] – CRICRÌ discloses for each group of feature maps, a squeeze-and-attention operation [wherein squeeze-and-attention operation is a global processing branch] may be applied. See also Paragraph [0372]) comprising: performing a first partitioning operation to partition at least a first portion of an input tensor of the resolution-flexible multi-axis attention block into a plurality of first feature sets (Fig. 14, Paragraph [0393] - CRICRÌ discloses the ResNeSt block may comprise a splitting operation, which divides or splits a tensor that is input to the core attention block into K sub-tensors or groups of features.), and performing a global attention operation along a first axis of the plurality of first feature sets (Fig. 14, Paragraph [0371] - CRICRÌ discloses the ResNeSt block is an attention block where the input tensor is divided into K smaller tensors (also referred to as groups of feature maps). For each group of feature maps, a squeeze-and-attention operation may be applied. See also Paragraphs [0366-0368, 0372-0378].), wherein the first axis corresponds to the predefined number of the plurality of first feature sets (Fig. 14, Paragraph [0371] - CRICRÌ discloses the ResNeSt block is an attention block where the input tensor is divided into K smaller tensors (also referred to as groups of feature maps) [wherein K smaller tensors is the predefined number of the plurality of first feature sets]. For each group of feature maps, a squeeze-and-attention operation may be applied. See also Paragraphs [0366-0368, 0372-0378].); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN of having a computer-implemented method for image processing, the method comprising: obtaining an input image; processing the input image by executing, by one or more processors, a machine-learned image processing model to generate an output prediction, wherein executing the machine-learned image processing model comprises, at each of one or more resolution-flexible multi-axis attention blocks of the machine-learned image processing model: and executing a local processing branch of the resolution-flexible multi-axis attention block comprising: performing a second partitioning operation to partition at least a second portion of the input tensor into a plurality of second feature sets; and performing a respective local attention operation on a second axis of each of the plurality of second feature sets; and providing the output prediction as an output, with the teachings of CRICRÌ having executing a global processing branch of the resolution-flexible multi-axis attention block comprising: performing a first partitioning operation to partition at least a first portion of an input tensor of the resolution-flexible multi-axis attention block into a plurality of first feature sets, and performing a global attention operation along a first axis of the plurality of first feature sets, wherein the first axis corresponds to the predefined number of the plurality of first feature sets. Wherein CHEN’s a computer-implemented method for image processing wherein having executing a global processing branch of the resolution-flexible multi-axis attention block comprising: performing a first partitioning operation to partition at least a first portion of an input tensor of the resolution-flexible multi-axis attention block into a plurality of first feature sets, and performing a global attention operation along a first axis of the plurality of first feature sets, wherein the first axis corresponds to the predefined number of the plurality of first feature sets. The motivation behind this modification would have been to provide an improved method for image processing that has the flexibility of modeling at various scales and enhancing visual quality, since both CHEN and CRICRÌ relate to architectures where multiple neural networks are connected in parallel or in series and adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and CRICRÌ relates to a method, apparatus, and computer program product for providing an attention block for neural network-based image and video compression, wherein the proposed block achieves a better rate-distortion performance, multiscale structural similarity (MS-SSIM) is higher (e.g., better visual quality) and bits per pixel (BPP) is lower (e.g., lower bitrate). Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and CRICRÌ (US 20240289590 A1), Paragraphs [0187, 0399]. CHEN in view of CRICRÌ fail to explicitly teach wherein the first partitioning operation comprises a grid operation that partitions the input tensor into a fixed grid having a predefined number of the plurality of first feature sets irrespective of a resolution of the input tensor; However, HE explicitly teaches wherein the first partitioning operation comprises a grid operation that partitions the input tensor into a fixed grid (Fig. 2, Paragraph [0025] – HE discloses convolutional layers 204 may accept arbitrary input image 202 sizes, but they produce outputs of variable sizes. The spatial bins may have sizes proportional to the image size, so the number of bins may be fixed regardless of the image size. See also Fig. 3, Paragraph [0035].) having a predefined number of the plurality of first feature sets irrespective of a resolution of the input tensor (Fig. 3, Paragraph [0027] – HE discloses with spatial pyramid pooling, the input image may be of any size allowing not only arbitrary aspect ratios, but also arbitrary scales. Paragraph [0035] – HE discloses an SPP network 312 of one or more layers that includes spatial bins based on the number of filters a top convolutional layer may pool the extracted features and generate fixed-size outputs for a fully-connected layer 314. See also Paragraph [0036].); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN in view of CRICRÌ of having a computer-implemented method for image processing, the method comprising: obtaining an input image; processing the input image by executing, by one or more processors, a machine-learned image processing model to generate an output prediction, wherein executing the machine-learned image processing model comprises, at each of one or more resolution-flexible multi-axis attention blocks of the machine-learned image processing model: executing a global processing branch of the resolution-flexible multi-axis attention block comprising: performing a first partitioning operation to partition at least a first portion of an input tensor of the resolution-flexible multi-axis attention block into a plurality of first feature sets, and performing a global attention operation along a first axis of the plurality of first feature sets, wherein the first axis corresponds to the predefined number of the plurality of first feature sets; and executing a local processing branch of the resolution-flexible multi-axis attention block comprising: performing a second partitioning operation to partition at least a second portion of the input tensor into a plurality of second feature sets; and performing a respective local attention operation on a second axis of each of the plurality of second feature sets; and providing the output prediction as an output, with the teachings of HE having wherein the first partitioning operation comprises a grid operation that partitions the input tensor into a fixed grid having a predefined number of the plurality of first feature sets irrespective of a resolution of the input tensor. Wherein CHEN’s a computer-implemented method for image processing wherein the first partitioning operation comprises a grid operation that partitions the input tensor into a fixed grid having a predefined number of the plurality of first feature sets irrespective of a resolution of the input tensor. The motivation behind this modification would have been to provide an improved method for image processing that has the flexibility of modeling at various scales and reduces computational expense, since both CHEN and HE relate to architectures where neural networks are adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and HE relates to methods, devices, and systems to process images using spatial pyramid pooling networks; a system employing SPP-network to process images may advantageously improve usability of object detection in searches, vision systems, and other image analysis implementations, as well as reduce computational expense such as processor load, memory load, and enhance reliability of object detection, for example, in satellite imaging, security monitoring, and comparable systems. Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and HE (US 20160104056 A1), Paragraph [0039]. Regarding claim 16, CHEN and CRICRÌ in view of HE teach the computer-implemented method of claim 15, Although CHEN further teaches wherein the input tensor has a height, a width, and a channel depth (Fig. 2, Paragraph [0088] – CHEN discloses each feature map [wherein feature map is the input tensor] has dimensions of 2×3×H×W [wherein 3 is channel depth, H is height, W is width] when sent to the PatchEmbed.), CHEN fails to explicitly teach and wherein the first partitioning operation comprises a grid partitioning operation that partitions the height and width of the input tensor into a grid having the predefined number of the plurality of first feature sets. However, CRICRÌ explicitly teaches and wherein the first partitioning operation comprises a grid partitioning operation (Fig. 14, Paragraph [0393] - CRICRÌ discloses the ResNeSt block may comprise a splitting operation [wherein a splitting operation is a grid partitioning operation], which divides or splits a tensor that is input to the core attention block into K sub-tensors or groups of features.) that partitions the height and width of the input tensor into a grid having the predefined number of the plurality of first feature sets (Fig. 14, Paragraph [0371] - CRICRÌ discloses the ResNeSt block is an attention block where the input tensor is divided into K smaller tensors (also referred to as groups of feature maps) [wherein K smaller tensors is the predefined number of the plurality of first feature sets]. For each group of feature maps, a squeeze-and-attention operation may be applied. See also Paragraphs [0366-0368].). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computer-implemented method for image processing, the method comprising: obtaining an input image; processing the input image by executing, by one or more processors, a machine-learned image processing model to generate an output prediction, wherein executing the machine-learned image processing model comprises, at each of one or more resolution-flexible multi-axis attention blocks of the machine-learned image processing model: executing a global processing branch of the resolution-flexible multi-axis attention block comprising: performing a first partitioning operation to partition at least a first portion of an input tensor of the resolution-flexible multi-axis attention block into a plurality of first feature sets, with the teachings of CRICRÌ of having and wherein the first partitioning operation comprises a grid partitioning operation that partitions the height and width of the input tensor into a grid having the predefined number of the plurality of first feature sets. Wherein CHEN’s computer-implemented method for image processing having and wherein the first partitioning operation comprises a grid partitioning operation that partitions the height and width of the input tensor into a grid having the predefined number of the plurality of first feature sets. The motivation behind this modification would have been to provide an improved method for image processing that has the flexibility of modeling at various scales and enhancing visual quality, since both CHEN and CRICRÌ relate to architectures where multiple neural networks are connected in parallel or in series and adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and CRICRÌ relates to a method, apparatus, and computer program product for providing an attention block for neural network-based image and video compression, wherein the proposed block achieves a better rate-distortion performance, multiscale structural similarity (MS-SSIM) is higher (e.g., better visual quality) and bits per pixel (BPP) is lower (e.g., lower bitrate). Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and CRICRÌ (US 20240289590 A1), Paragraphs [0187, 0399]. Regarding claim 17, CHEN and CRICRÌ in view of HE teach the computer-implemented method of claim 15, CHEN further teaches wherein the second partitioning operation comprises a block partitioning operation (Fig. 6, Paragraph [0088] – CHEN discloses the PatchEmbed module mainly increases the channel dimensions of each patch in a set of non-overlapping patches which are divided from the inputted picture with dimensions of H×W×3 through patch partition processing [wherein patch partition processing is a block partitioning operation]) that partitions the second portion of the input tensor into the plurality of second feature sets such that each of the plurality of second feature sets has a predefined height and width (Fig. 6, Paragraph [0088] – CHEN discloses each feature map has dimensions of 2×3×H×W when sent to the PatchEmbed, and has dimensions of 2×96×H/4×W/4 when finally sent to the next module [wherein H/4×W/4 is a predefined height and width].). Regarding claim 18, CHEN and CRICRÌ in view of HE teach the computer-implemented method of claim 17, PNG media_image1.png 385 883 media_image1.png Greyscale Annotated cropped image of Fig. 6 (CHEN) Although CHEN further teaches wherein the number of the plurality of first feature sets is equal to a multiplication of the predefined height and width of each of the plurality of second feature sets (Fig. 6, Paragraph [0085] – CHEN discloses it is determined whether the width and height of the feature map are each an integer multiple of 4; [See annotated cropped image of Fig. 6 above where CHEN teaches a multiplication of the width and height of each feature set is equal to the number of the plurality of feature sets].). CHEN fails to explicitly teach predefined number. However, CRICRÌ explicitly teaches predefined number (Fig. 14, Paragraph [0371] - CRICRÌ discloses the ResNeSt block is an attention block where the input tensor is divided into K smaller tensors (also referred to as groups of feature maps) [wherein K smaller tensors is the predefined number of the plurality of first feature sets]. For each group of feature maps, a squeeze-and-attention operation may be applied. See also Paragraphs [0366-0368].). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computer-implemented method for image processing, the method comprising: obtaining an input image; processing the input image by executing, by one or more processors, a machine-learned image processing model to generate an output prediction, wherein executing the machine-learned image processing model comprises, at each of one or more resolution-flexible multi-axis attention blocks of the machine-learned image processing model: executing a global processing branch of the resolution-flexible multi-axis attention block comprising: performing a first partitioning operation to partition at least a first portion of an input tensor of the resolution-flexible multi-axis attention block into a plurality of first feature sets, with the teachings of CRICRÌ having the predefined number. Wherein CHEN’s computer-implemented method for image processing wherein having the predefined number of the plurality of first feature sets. The motivation behind this modification would have been to provide an improved method for image processing that has the flexibility of modeling at various scales and enhancing visual quality, since both CHEN and CRICRÌ relate to architectures where multiple neural networks are connected in parallel or in series and adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and CRICRÌ relates to a method, apparatus, and computer program product for providing an attention block for neural network-based image and video compression, wherein the proposed block achieves a better rate-distortion performance, multiscale structural similarity (MS-SSIM) is higher (e.g., better visual quality) and bits per pixel (BPP) is lower (e.g., lower bitrate). Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and CRICRÌ (US 20240289590 A1), Paragraphs [0187, 0399]. Claims 5, 8, 19, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over CHEN (US 20230184927 A1), hereinafter referenced as CHEN in view of CRICRÌ (US 20240289590 A1), hereinafter referenced as CRICRÌ in further view of HE (US 20160104056 A1), hereinafter referenced as HE in further view of TROCKMAN (US 20230096021 A1), hereinafter referenced as TROCKMAN. Regarding claim 5, CHEN and CRICRÌ in view of HE teaches the computing system of claim 1, Although CHEN further teaches and performing the local attention operation comprises, for each of the second feature sets (Fig. 4, Paragraph [0081] – CHEN further discloses a layered Transformer is introduced, whose representation is calculated by moving the window. Paragraph [0082] – CHEN further discloses the Swin transformer performs self-attention calculation in each window [wherein performing self-attention calculation in each window is performing the local attention operation, and wherein each window is each of the second feature sets].). CHEN fails to explicitly teach wherein: performing the global attention operation comprises, for each of a number of positions within each first feature set, However, CRICRÌ explicitly teaches wherein: performing the global attention operation comprises, for each of a number of positions within each first feature set (Fig. 14, Paragraph [0371] - CRICRÌ discloses the ResNeSt block is an attention block where the input tensor is divided into K smaller tensors (also referred to as groups of feature maps). For each group of feature maps [wherein each groups of feature maps is each first feature set], a squeeze-and-attention operation may be applied [wherein a squeeze-and-attention operation is the global attention operation]. See also Paragraphs [0366-0368, 0372-0378].), Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computing system for resolution-flexible image processing, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to execute: a machine-learned image processing model configured to process input image data to generate an output prediction, wherein the machine-learned image processing model comprises one or more resolution-flexible multi-axis attention blocks, each of the one or more resolution-flexible multi-axis attention blocks comprising: a global processing branch configured to: perform a global attention operation along a first axis of the plurality of first feature sets, with the teachings of CRICRÌ having wherein: performing the global attention operation comprises, for each of a number of positions within each first feature set, Wherein CHEN’s computing system for resolution-flexible image processing wherein: performing the global attention operation comprises, for each of a number of positions within each first feature set, The motivation behind this modification would have been to provide an improved system for image processing that has the flexibility of modeling at various scales and enhancing visual quality, since both CHEN and CRICRÌ relate to architectures where multiple neural networks are connected in parallel or in series and adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and CRICRÌ relates to a method, apparatus, and computer program product for providing an attention block for neural network-based image and video compression, wherein the proposed block achieves a better rate-distortion performance, multiscale structural similarity (MS-SSIM) is higher (e.g., better visual quality) and bits per pixel (BPP) is lower (e.g., lower bitrate). Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and CRICRÌ (US 20240289590 A1), Paragraphs [0187, 0399]. CHEN and CRICRÌ in view of HE fail to explicitly teach processing a respective feature value from each of the plurality of first feature sets with a gated multi-layer perceptron block; processing all feature values within the second feature set with the gated multi-layer perceptron block. However, TROCKMAN explicitly teaches processing a respective feature value from each of the plurality of first feature sets (Figs. 6-7, Paragraph [0058] – TROCKMAN discloses the first layer is a linear transformation applied identically to non-overlapping square regions of the input. Then, the model processes the patch embeddings rather than the original image.) with a gated multi-layer perceptron block (Figs. 6-7, Paragraph [0074] – TROCKMAN discloses vision transformers have inspired a new paradigm of isotropic architectures which use patch embeddings for the first layer. These models look similar to repeated transformer-encoder blocks with different operations replacing the self-attention and MLP operations. CycleMLP, gMLP, and vision permutator, replace one or both blocks with various novel operations [wherein gMLP is a gated multi-layer perceptron block].); processing all feature values within the second feature set (Figs. 6-7, Paragraph [0058] – TROCKMAN discloses the first layer is a linear transformation applied identically to non-overlapping square regions of the input. Then, the model processes the patch embeddings rather than the original image.) with the gated multi-layer perceptron block (Figs. 6-7, Paragraph [0074] – TROCKMAN vision transformers have inspired a new paradigm of isotropic architectures which use patch embeddings for the first layer. These models look similar to repeated transformer-encoder blocks with different operations replacing the self-attention and MLP operations. CycleMLP, gMLP, and vision permutator, replace one or both blocks with various novel operations [wherein gMLP is a gated multi-layer perceptron block].). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computing system for resolution-flexible image processing, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to execute: a machine-learned image processing model configured to process input image data to generate an output prediction, wherein the machine-learned image processing model comprises one or more resolution-flexible multi-axis attention blocks, each of the one or more resolution-flexible multi-axis attention blocks comprising: a global processing branch configured to: perform a global attention operation along a first axis of the plurality of first feature sets, with the teachings of TROCKMAN having processing a respective feature value from each of the plurality of first feature sets with a gated multi-layer perceptron block; processing all feature values within the second feature set with the gated multi-layer perceptron block. Wherein CHEN’s computing system for resolution-flexible image processing wherein having processing a respective feature value from each of the plurality of first feature sets with a gated multi-layer perceptron block; processing all feature values within the second feature set with the gated multi-layer perceptron block. The motivation behind this modification would have been to provide an improved system for image processing that is amenable to vision tasks and has the flexibility of modeling at various scales, since both CHEN and TROCKMAN relate to arrangements for image or video recognition and understanding using pattern recognition or machine learning, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and TROCKMAN relates to improvements allowing for reduced parameters in an isotonic convolutional neural network; inductive bias, which includes translation invariance, is amenable to vision tasks and leads to high data efficiency. Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and TROCKMAN (US 20230096021 A1), Paragraphs [0069]. Regarding claim 8, CHEN and CRICRÌ in view of HE teaches the computing system of claim 1, CHEN further teaches and performing the local attention operation (Fig. 4, Paragraph [0081] – CHEN further discloses a layered Transformer is introduced, whose representation is calculated by moving the window [wherein moving the window is performing local attention operation].) comprises, for each of the second feature sets (Fig. 4, Paragraph [0085] – CHEN further discloses feature dimensions and feature channels of the contextual self-attention module remain unchanged, which strengthens the information exchange between different windows on the feature map.), processing all feature values within the second feature set with one of the following: self-attention (Paragraph [0082] – CHEN discloses the Swin transformer performs self-attention calculation in each window [wherein each window comprises feature values].), processing a respective feature value from each of the plurality of first feature sets with one of the following: self-attention (Fig. 4, Paragraph [0081] – CHEN discloses the Swin transformer performs self-attention calculation in each window to obtain an updated window. Next, through a patch merging operation, the windows are merged, and the self-attention calculation is continued in the merged window. Paragraph [0082].), CHEN fails to explicitly teach wherein: performing the global attention operation comprises, for each of a number of positions within each first feature set, processing a respective feature value from each of the plurality of first feature sets with one of the following: or a Fourier transform; However, CRICRÌ explicitly teaches wherein: performing the global attention operation (Fig. 14, Paragraph [0371] - CRICRÌ discloses the ResNeSt block is an attention block where the input tensor is divided into K smaller tensors (also referred to as groups of feature maps). For each group of feature maps, a squeeze-and-attention operation may be applied. See also Paragraphs [0366-0368, 0372-0378].) comprises, for each of a number of positions within each first feature set (Fig. 14, Paragraph [0371] - CRICRÌ discloses for each group of feature maps, a squeeze-and-attention operation may be applied.), processing a respective feature value from each of the plurality of first feature sets with one of the following: or a Fourier transform (Fig. 8, Paragraph [0290] - CRICRÌ discloses transform and quantization block or circuit 806 may perform a transform of input data [wherein input data comprises the plurality of first feature sets] to a different domain, for example, the FFT transform would transform the data to frequency domain.); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computing system for resolution-flexible image processing, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to execute: a machine-learned image processing model configured to process input image data to generate an output prediction, wherein the machine-learned image processing model comprises one or more resolution-flexible multi-axis attention blocks, each of the one or more resolution-flexible multi-axis attention blocks comprising: a global processing branch configured to: perform a global attention operation along a first axis of the plurality of first feature sets, with the teachings of CRICRÌ having wherein: performing the global attention operation comprises, for each of a number of positions within each first feature set, processing a respective feature value from each of the plurality of first feature sets with one of the following: or a Fourier transform; Wherein CHEN’s computing system for resolution-flexible image processing wherein: performing the global attention operation comprises, for each of a number of positions within each first feature set, processing a respective feature value from each of the plurality of first feature sets with one of the following: or a Fourier transform; The motivation behind this modification would have been to provide an improved system for image processing that has the flexibility of modeling at various scales and enhancing visual quality, since both CHEN and CRICRÌ relate to architectures where multiple neural networks are connected in parallel or in series and adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and CRICRÌ relates to a method, apparatus, and computer program product for providing an attention block for neural network-based image and video compression, wherein the proposed block achieves a better rate-distortion performance, multiscale structural similarity (MS-SSIM) is higher (e.g., better visual quality) and bits per pixel (BPP) is lower (e.g., lower bitrate). Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and CRICRÌ (US 20240289590 A1), Paragraphs [0187, 0399]. CHEN and CRICRÌ in view of HE fail to explicitly teach processing a respective feature value from each of the plurality of first feature sets with one of the following: a spatial multi-layer perceptron; processing all feature values within the second feature set with one of the following: a spatial multi-layer perceptron, or a Fourier transform. However, TROCKMAN explicitly teaches processing a respective feature value from each of the plurality of first feature sets with one of the following: a spatial multi-layer perceptron (Figs. 6-7, Paragraph [0074] – TROCKMAN discloses vision transformers have inspired a new paradigm of isotropic architectures which use patch embeddings for the first layer. These models look similar to repeated transformer-encoder blocks with different operations replacing the self-attention and MLP operations. For example, MLP-Mixer replaces them both with MLPs applied across different dimensions (i.e., spatial and channel location mixing); ResMLP is a data-efficient variation on this theme. CycleMLP, gMLP, and vision permutator, replace one or both blocks with various novel operations. See also Paragraph [0058].); processing all feature values within the second feature set with one of the following: a spatial multi-layer perceptron (Figs. 6-7, Paragraph [0074] – TROCKMAN discloses vision transformers have inspired a new paradigm of isotropic architectures which use patch embeddings for the first layer. These models look similar to repeated transformer-encoder blocks with different operations replacing the self-attention and MLP operations. For example, MLP-Mixer replaces them both with MLPs applied across different dimensions (i.e., spatial and channel location mixing); ResMLP is a data-efficient variation on this theme. CycleMLP, gMLP, and vision permutator, replace one or both blocks with various novel operations. See also Paragraph [0058].), or a Fourier transform (Figs. 6-7, Paragraph [0033] – TROCKMAN discloses after the patch embedding step, the internal resolution of the network is always h/p×w/p. Performing convolutions with large kernel sizes on high-resolution internal representations can be expensive. However, in the Fourier domain, the running time of this operation is independent of the kernel size, this could be leveraged in select deep learning frameworks in which the framework automatically switches to FFT processing.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computing system for resolution-flexible image processing, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to execute: a machine-learned image processing model configured to process input image data to generate an output prediction, wherein the machine-learned image processing model comprises one or more resolution-flexible multi-axis attention blocks, each of the one or more resolution-flexible multi-axis attention blocks comprising: a global processing branch configured to: perform a global attention operation along a first axis of the plurality of first feature sets, with the teachings of TROCKMAN having processing a respective feature value from each of the plurality of first feature sets with one of the following: a spatial multi-layer perceptron; processing all feature values within the second feature set with one of the following: a spatial multi-layer perceptron, or a Fourier transform. Wherein CHEN’s computing system for resolution-flexible image processing wherein having processing a respective feature value from each of the plurality of first feature sets with one of the following: a spatial multi-layer perceptron; processing all feature values within the second feature set with one of the following: a spatial multi-layer perceptron, or a Fourier transform. The motivation behind this modification would have been to provide an improved system for image processing that is amenable to vision tasks and has the flexibility of modeling at various scales, since both CHEN and TROCKMAN relate to arrangements for image or video recognition and understanding using pattern recognition or machine learning, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and TROCKMAN relates to improvements allowing for reduced parameters in an isotonic convolutional neural network; inductive bias, which includes translation invariance, is amenable to vision tasks and leads to high data efficiency. Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and TROCKMAN (US 20230096021 A1), Paragraphs [0069]. Regarding claim 19, CHEN and CRICRÌ in view of HE teaches the computer-implemented method of claim 15, Although CHEN further teaches and performing the local attention operation (Fig. 4, Paragraph [0081] – CHEN further discloses a layered Transformer is introduced, whose representation is calculated by moving the window [wherein moving the window is performing local attention operation].) comprises, for each of the second feature sets (Fig. 4, Paragraph [0085] – CHEN further discloses feature dimensions and feature channels of the contextual self-attention module remain unchanged, which strengthens the information exchange between different windows on the feature map.), CHEN fails to explicitly teach wherein: performing the global attention operation comprises, for each of a number of positions within each first feature set, However, CRICRÌ explicitly teaches wherein: performing the global attention operation comprises, for each of a number of positions within each first feature set (Fig. 14, Paragraph [0371] - CRICRÌ discloses the ResNeSt block is an attention block where the input tensor is divided into K smaller tensors (also referred to as groups of feature maps). For each group of feature maps [wherein each groups of feature maps is each first feature set], a squeeze-and-attention operation may be applied [wherein a squeeze-and-attention operation is the global attention operation]. See also Paragraphs [0366-0368, 0372-0378].), Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computer-implemented method for image processing, the method comprising: obtaining an input image; processing the input image by executing, by one or more processors, a machine-learned image processing model to generate an output prediction, wherein executing the machine-learned image processing model comprises, at each of one or more resolution-flexible multi-axis attention blocks of the machine-learned image processing model: performing a global attention operation along a first axis of the plurality of first feature sets, with the teachings of CRICRÌ having wherein: performing the global attention operation comprises, for each of a number of positions within each first feature set, Wherein CHEN’s computer-implemented method for image processing wherein: performing the global attention operation comprises, for each of a number of positions within each first feature set, The motivation behind this modification would have been to provide an improved method for image processing that has the flexibility of modeling at various scales and enhancing visual quality, since both CHEN and CRICRÌ relate to architectures where multiple neural networks are connected in parallel or in series and adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and CRICRÌ relates to a method, apparatus, and computer program product for providing an attention block for neural network-based image and video compression, wherein the proposed block achieves a better rate-distortion performance, multiscale structural similarity (MS-SSIM) is higher (e.g., better visual quality) and bits per pixel (BPP) is lower (e.g., lower bitrate). Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and CRICRÌ (US 20240289590 A1), Paragraphs [0187, 0399]. CHEN and CRICRÌ in view of HE fail to explicitly teach processing a respective feature value from each of the plurality of first feature sets with a gated multi-layer perceptron block; processing all feature values within the second feature set with the gated multi-layer perceptron block. However, TROCKMAN explicitly teaches processing a respective feature value from each of the plurality of first feature sets (Figs. 6-7, Paragraph [0058] – TROCKMAN discloses the first layer is a linear transformation applied identically to non-overlapping square regions of the input. Then, the model processes the patch embeddings rather than the original image.) with a gated multi-layer perceptron block (Figs. 6-7, Paragraph [0074] – TROCKMAN discloses vision transformers have inspired a new paradigm of isotropic architectures which use patch embeddings for the first layer. These models look similar to repeated transformer-encoder blocks with different operations replacing the self-attention and MLP operations. CycleMLP, gMLP, and vision permutator, replace one or both blocks with various novel operations [wherein gMLP is a gated multi-layer perceptron block].); processing all feature values within the second feature set (Figs. 6-7, Paragraph [0058] – TROCKMAN discloses the first layer is a linear transformation applied identically to non-overlapping square regions of the input. Then, the model processes the patch embeddings rather than the original image.) with the gated multi-layer perceptron block (Figs. 6-7, Paragraph [0074] – TROCKMAN vision transformers have inspired a new paradigm of isotropic architectures which use patch embeddings for the first layer. These models look similar to repeated transformer-encoder blocks with different operations replacing the self-attention and MLP operations. CycleMLP, gMLP, and vision permutator, replace one or both blocks with various novel operations [wherein gMLP is a gated multi-layer perceptron block].). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computer-implemented method for image processing, the method comprising: obtaining an input image; processing the input image by executing, by one or more processors, a machine-learned image processing model to generate an output prediction, wherein executing the machine-learned image processing model comprises, at each of one or more resolution-flexible multi-axis attention blocks of the machine-learned image processing model: performing a global attention operation along a first axis of the plurality of first feature sets, with the teachings of TROCKMAN having processing a respective feature value from each of the plurality of first feature sets with a gated multi-layer perceptron block; processing all feature values within the second feature set with the gated multi-layer perceptron block. Wherein CHEN’s computer-implemented method for image processing wherein having processing a respective feature value from each of the plurality of first feature sets with a gated multi-layer perceptron block; processing all feature values within the second feature set with the gated multi-layer perceptron block. The motivation behind this modification would have been to provide an improved method for image processing that is amenable to vision tasks and has the flexibility of modeling at various scales, since both CHEN and TROCKMAN relate to arrangements for image or video recognition and understanding using pattern recognition or machine learning, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and TROCKMAN relates to improvements allowing for reduced parameters in an isotonic convolutional neural network; inductive bias, which includes translation invariance, is amenable to vision tasks and leads to high data efficiency. Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and TROCKMAN (US 20230096021 A1), Paragraphs [0069]. Regarding claim 20, CHEN and CRICRÌ in view of HE teaches the computer-implemented method of claim 15, CHEN further teaches and performing the local attention operation (Fig. 4, Paragraph [0081] – CHEN further discloses a layered Transformer is introduced, whose representation is calculated by moving the window [wherein moving the window is performing local attention operation].) comprises, for each of the second feature sets (Fig. 4, Paragraph [0085] – CHEN further discloses feature dimensions and feature channels of the contextual self-attention module remain unchanged, which strengthens the information exchange between different windows on the feature map.), processing all feature values within the second feature set with one of the following: self-attention (Paragraph [0082] – CHEN discloses the Swin transformer performs self-attention calculation in each window [wherein each window comprises feature values].), processing a respective feature value from each of the plurality of first feature sets with one of the following: self-attention (Fig. 4, Paragraph [0081] – CHEN discloses the Swin transformer performs self-attention calculation in each window to obtain an updated window. Next, through a patch merging operation, the windows are merged, and the self-attention calculation is continued in the merged window. Paragraph [0082].), CHEN fails to explicitly teach wherein: performing the global attention operation comprises, for each of a number of positions within each first feature set, processing a respective feature value from each of the plurality of first feature sets with one of the following: or a Fourier transform; However, CRICRÌ explicitly teaches wherein: performing the global attention operation (Fig. 14, Paragraph [0371] - CRICRÌ discloses the ResNeSt block is an attention block where the input tensor is divided into K smaller tensors (also referred to as groups of feature maps). For each group of feature maps, a squeeze-and-attention operation may be applied. See also Paragraphs [0366-0368, 0372-0378].) comprises, for each of a number of positions within each first feature set (Fig. 14, Paragraph [0371] - CRICRÌ discloses for each group of feature maps, a squeeze-and-attention operation may be applied.), processing a respective feature value from each of the plurality of first feature sets with one of the following: or a Fourier transform (Fig. 8, Paragraph [0290] - CRICRÌ discloses transform and quantization block or circuit 806 may perform a transform of input data [wherein input data comprises the plurality of first feature sets] to a different domain, for example, the FFT transform would transform the data to frequency domain.); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computer-implemented method for image processing, the method comprising: obtaining an input image; processing the input image by executing, by one or more processors, a machine-learned image processing model to generate an output prediction, wherein executing the machine-learned image processing model comprises, at each of one or more resolution-flexible multi-axis attention blocks of the machine-learned image processing model: performing a global attention operation along a first axis of the plurality of first feature sets, with the teachings of CRICRÌ having wherein: performing the global attention operation comprises, for each of a number of positions within each first feature set, processing a respective feature value from each of the plurality of first feature sets with one of the following: or a Fourier transform; Wherein CHEN’s computer-implemented method for image processing wherein: performing the global attention operation comprises, for each of a number of positions within each first feature set, processing a respective feature value from each of the plurality of first feature sets with one of the following: or a Fourier transform; The motivation behind this modification would have been to provide an improved method for image processing that has the flexibility of modeling at various scales and enhancing visual quality, since both CHEN and CRICRÌ relate to architectures where multiple neural networks are connected in parallel or in series and adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and CRICRÌ relates to a method, apparatus, and computer program product for providing an attention block for neural network-based image and video compression, wherein the proposed block achieves a better rate-distortion performance, multiscale structural similarity (MS-SSIM) is higher (e.g., better visual quality) and bits per pixel (BPP) is lower (e.g., lower bitrate). Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and CRICRÌ (US 20240289590 A1), Paragraphs [0187, 0399]. CHEN and CRICRÌ in view of HE fail to explicitly teach processing a respective feature value from each of the plurality of first feature sets with one of the following: a spatial multi-layer perceptron; processing all feature values within the second feature set with one of the following: a spatial multi-layer perceptron, or a Fourier transform. However, TROCKMAN explicitly teaches processing a respective feature value from each of the plurality of first feature sets with one of the following: a spatial multi-layer perceptron (Figs. 6-7, Paragraph [0074] – TROCKMAN discloses vision transformers have inspired a new paradigm of isotropic architectures which use patch embeddings for the first layer. These models look similar to repeated transformer-encoder blocks with different operations replacing the self-attention and MLP operations. For example, MLP-Mixer replaces them both with MLPs applied across different dimensions (i.e., spatial and channel location mixing); ResMLP is a data-efficient variation on this theme. CycleMLP, gMLP, and vision permutator, replace one or both blocks with various novel operations. See also Paragraph [0058].); processing all feature values within the second feature set with one of the following: a spatial multi-layer perceptron (Figs. 6-7, Paragraph [0074] – TROCKMAN discloses vision transformers have inspired a new paradigm of isotropic architectures which use patch embeddings for the first layer. These models look similar to repeated transformer-encoder blocks with different operations replacing the self-attention and MLP operations. For example, MLP-Mixer replaces them both with MLPs applied across different dimensions (i.e., spatial and channel location mixing); ResMLP is a data-efficient variation on this theme. CycleMLP, gMLP, and vision permutator, replace one or both blocks with various novel operations. See also Paragraph [0058].), or a Fourier transform (Figs. 6-7, Paragraph [0033] – TROCKMAN discloses after the patch embedding step, the internal resolution of the network is always h/p×w/p. Performing convolutions with large kernel sizes on high-resolution internal representations can be expensive. However, in the Fourier domain, the running time of this operation is independent of the kernel size, this could be leveraged in select deep learning frameworks in which the framework automatically switches to FFT processing.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computer-implemented method for image processing, the method comprising: obtaining an input image; processing the input image by executing, by one or more processors, a machine-learned image processing model to generate an output prediction, wherein executing the machine-learned image processing model comprises, at each of one or more resolution-flexible multi-axis attention blocks of the machine-learned image processing model: performing a global attention operation along a first axis of the plurality of first feature sets, with the teachings of TROCKMAN having processing a respective feature value from each of the plurality of first feature sets with one of the following: a spatial multi-layer perceptron; processing all feature values within the second feature set with one of the following: a spatial multi-layer perceptron, or a Fourier transform. Wherein CHEN’s computer-implemented method for image processing wherein having processing a respective feature value from each of the plurality of first feature sets with one of the following: a spatial multi-layer perceptron; processing all feature values within the second feature set with one of the following: a spatial multi-layer perceptron, or a Fourier transform. The motivation behind this modification would have been to provide an improved method for image processing that is amenable to vision tasks and has the flexibility of modeling at various scales, since both CHEN and TROCKMAN relate to arrangements for image or video recognition and understanding using pattern recognition or machine learning, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and TROCKMAN relates to improvements allowing for reduced parameters in an isotonic convolutional neural network; inductive bias, which includes translation invariance, is amenable to vision tasks and leads to high data efficiency. Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and TROCKMAN (US 20230096021 A1), Paragraphs [0069]. Claims 6 and 7 are rejected under 35 U.S.C. 103 as being unpatentable over CHEN (US 20230184927 A1), hereinafter referenced as CHEN in view of CRICRÌ (US 20240289590 A1), hereinafter referenced as CRICRÌ in further view of HE (US 20160104056 A1), hereinafter referenced as HE in further view of TROCKMAN (US 20230096021 A1), hereinafter referenced as TROCKMAN in further view of SAHU (US 20220300740 A1), hereinafter referenced as SAHU. Regarding claim 6, CHEN and CRICRÌ in view of HE, in further view of TROCKMAN teach the computing system of claim 5, CHEN and CRICRÌ in view of HE, in further view of TROCKMAN fail to explicitly teach wherein the gated multi-layer perceptron block generates one or more gating weights for input feature values, and wherein the gated multi-layer perceptron block applies the one or more gating weights to gate the input feature values. However, SAHU explicitly teaches wherein the gated multi-layer perceptron block generates one or more gating weights for input feature values (Fig. 3, Paragraph [0070] - SAHU discloses the global representation 320 and the local representation 336 of the input video data 202 and/or the input audio data 204 [wherein input video data and/or input audio data are input feature values] are combined using a gating mechanism, which uses the multi-level output representations as inputs and determines weights to be given to the global and local representations. See also Paragraph [0078] and Algorithm 1.), and wherein the gated multi-layer perceptron block applies the one or more gating weights to gate the input feature values (Fig. 3, Paragraph [0071] – SAHU discloses a combiner 368 adds or otherwise combines the weighted global representation 362 and the weighted local representation 366 to generate a final output representation 370 of the input video data 202 and/or the input audio data 204. The final output representation 370 may be provided to the classifier function 218 or other destination(s) for further processing or use. See also Paragraph [0078] and Algorithm 1.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computing system for resolution-flexible image processing, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to execute: a machine-learned image processing model configured to process input image data to generate an output prediction, wherein the machine-learned image processing model comprises one or more resolution-flexible multi-axis attention blocks, each of the one or more resolution-flexible multi-axis attention blocks comprising: a global processing branch configured to: perform a first partitioning operation to partition at least a first portion of an input tensor of the resolution-flexible multi-axis attention block into a plurality of first feature sets, with the teachings of SAHU having wherein the gated multi-layer perceptron block generates one or more gating weights for input feature values, and wherein the gated multi-layer perceptron block applies the one or more gating weights to gate the input feature values. Wherein CHEN’s computing system for resolution-flexible image processing wherein the gated multi-layer perceptron block generates one or more gating weights for input feature values, and wherein the gated multi-layer perceptron block applies the one or more gating weights to gate the input feature values. The motivation behind this modification would have been to provide an improved system for image processing that has the flexibility of modeling at various scales and output accurate representations, since both CHEN and SAHU relate to architectures where neural networks are adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and SAHU relates to a system and method for enhancing a machine learning model for audio/video understanding using gated multi-level attention and temporal adversarial training; this approach results in machine learning models that are trained to be more robust to adversarial attacks and that improve the generalizability of the machine learning models. Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and SAHU (US 20220300740 A1), Paragraphs [0002, 0030]. Regarding claim 7, CHEN and CRICRÌ in view of HE, in further view of TROCKMAN teach the computing system of claim 5, CHEN and CRICRÌ in view of HE, in further view of TROCKMAN fail to explicitly teach wherein the gated multi-layer perceptron block generates one or more gating weights for input feature values, and wherein the gated multi-layer perceptron block applies the one or more gating weights to gate the other feature values associated with a different feature stream. However, SAHU explicitly teaches wherein the gated multi-layer perceptron block generates one or more gating weights for input feature values (Fig. 3, Paragraph [0070] - SAHU discloses the global representation 320 and the local representation 336 of the input video data 202 and/or the input audio data 204 [wherein input video data and/or input audio data are input feature values] are combined using a gating mechanism, which uses the multi-level output representations as inputs and determines weights to be given to the global and local representations. See also Paragraph [0078] and Algorithm 1.), and wherein the gated multi-layer perceptron block applies the one or more gating weights to gate the other feature values associated with a different feature stream (Fig. 3, Paragraph [0072] – SAHU discloses each time frame of the input audio/video data can be weighted locally as well as globally while making predictions. The outputs of the linear transformation functions 338 and 340 are further processed in order to determine how much weight should be given to global attention characteristics and how much weight should be given to local attention characteristics. See also Algorithm 1.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computing system for resolution-flexible image processing, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to execute: a machine-learned image processing model configured to process input image data to generate an output prediction, wherein the machine-learned image processing model comprises one or more resolution-flexible multi-axis attention blocks, each of the one or more resolution-flexible multi-axis attention blocks comprising: a global processing branch configured to: perform a first partitioning operation to partition at least a first portion of an input tensor of the resolution-flexible multi-axis attention block into a plurality of first feature sets, with the teachings of SAHU having wherein the gated multi-layer perceptron block generates one or more gating weights for input feature values, and wherein the gated multi-layer perceptron block applies the one or more gating weights to gate the other feature values associated with a different feature stream. Wherein CHEN’s computing system for resolution-flexible image processing wherein the gated multi-layer perceptron block generates one or more gating weights for input feature values, and wherein the gated multi-layer perceptron block applies the one or more gating weights to gate the other feature values associated with a different feature stream. The motivation behind this modification would have been to provide an improved system for image processing that has the flexibility of modeling at various scales and output accurate representations, since both CHEN and SAHU relate to architectures where neural networks are adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and SAHU relates to a system and method for enhancing a machine learning model for audio/video understanding using gated multi-level attention and temporal adversarial training; this approach results in machine learning models that are trained to be more robust to adversarial attacks and that improve the generalizability of the machine learning models. Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and SAHU (US 20220300740 A1), Paragraphs [0002, 0030]. Claims 10-12 are rejected under 35 U.S.C. 103 as being unpatentable over CHEN (US 20230184927 A1), hereinafter referenced as CHEN in view of CRICRÌ (US 20240289590 A1), hereinafter referenced as CRICRÌ in further view of HE (US 20160104056 A1), hereinafter referenced as HE in further view of ZHENG (US 20220327657 A1), hereinafter referenced as ZHENG. Regarding claim 10, CHEN and CRICRÌ in view of HE teaches the computing system of claim 1, CHEN further teaches wherein: the one or more resolution-flexible multi-axis attention blocks comprises a plurality of resolution-flexible multi-axis attention blocks; (Fig. 6, Paragraph [0085] – CHEN discloses based on multi-head attention, the present disclosure takes into account the CotNet contextual attention mechanism and integrates the self-attention module block into the Swin transformer. Paragraph [0095] – CHEN further discloses a two-way multi-scale connection operation is enhanced through top-down and bottom-up attention, to guide learning of dynamic attention matrices and enhance feature interaction under different resolutions.), the machine-learned image processing model (Fig. 2, Paragraph [0074] – CHEN discloses introducing a cross-resolution attention enhancement neck CAENeck to the model framework CRTransSar [wherein model framework is a machine-learned image processing model]. See also Paragraph [0077].) comprises one or more backbone blocks (Fig. 6, Paragraph [0071] – CHEN discloses adding, to the model framework CRTransSar, a feature extraction network CRbackbone [wherein CRbackbone is a backbone block] based on contextual joint representation learning Transformer;), CHEN fails to explicitly teach and each of the plurality of encoders and each of the plurality of decoders contains a respective one of the plurality of resolution-flexible multi-axis attention blocks. However, CRICRÌ explicitly teaches and each of the plurality of encoders and each of the plurality of decoders contains a respective one of the plurality of resolution-flexible multi-axis attention blocks (Fig. 19, Paragraph [0396] – CRICRÌ discloses an example use case of a dense split attention block being used within an end-to-end learned codec, in accordance with some embodiments. In this embodiment, the end-to-end learned codec is shown to include an encoder 1902 and a decoder 1904.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computing system for resolution-flexible image processing, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to execute: a machine-learned image processing model configured to process input image data to generate an output prediction, wherein the machine-learned image processing model comprises one or more resolution-flexible multi-axis attention blocks, with the teachings of CRICRÌ having and each of the plurality of encoders and each of the plurality of decoders contains a respective one of the plurality of resolution-flexible multi-axis attention blocks. Wherein CHEN’s computing system for resolution-flexible image processing wherein having and each of the plurality of encoders and each of the plurality of decoders contains a respective one of the plurality of resolution-flexible multi-axis attention blocks. The motivation behind this modification would have been to provide an improved system for image processing that has the flexibility of modeling at various scales and enhancing visual quality, since both CHEN and CRICRÌ relate to architectures where multiple neural networks are connected in parallel or in series and adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and CRICRÌ relates to a method, apparatus, and computer program product for providing an attention block for neural network-based image and video compression, wherein the proposed block achieves a better rate-distortion performance, multiscale structural similarity (MS-SSIM) is higher (e.g., better visual quality) and bits per pixel (BPP) is lower (e.g., lower bitrate). Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and CRICRÌ (US 20240289590 A1), Paragraphs [0187, 0399]. CHEN and CRICRÌ in view of HE fail to explicitly teach each of the one or more backbone blocks comprising a hierarchical structure of a plurality of encoders and a plurality of decoders; However, ZHENG explicitly teaches each of the one or more backbone blocks comprising a hierarchical structure of a plurality of encoders (Fig. 2-3, Paragraph [0063] – ZHENG discloses the digital image layout neural network 320 includes multiple encoders and/or decoders to generate accurate and realistic refined images that match the style of the input image 302 while adhering to the arrangement and structure provided by the edited layout 304.) and a plurality of decoders (Fig. 2-3, Paragraph [0063] – ZHENG discloses the digital image layout neural network 320 includes multiple encoders and/or decoders to generate accurate and realistic refined images that match the style of the input image 302 while adhering to the arrangement and structure provided by the edited layout 304.); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computing system for resolution-flexible image processing, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to execute: a machine-learned image processing model configured to process input image data to generate an output prediction, wherein the machine-learned image processing model comprises one or more resolution-flexible multi-axis attention blocks, with the teachings of ZHENG having each of the one or more backbone blocks comprising a hierarchical structure of a plurality of encoders and a plurality of decoders; Wherein CHEN’s computing system for resolution-flexible image processing wherein each of the one or more backbone blocks comprising a hierarchical structure of a plurality of encoders and a plurality of decoders; The motivation behind this modification would have been to provide an improved system for image processing that has the flexibility of modeling at various scales and accurately transfers visual details to new layouts, since both CHEN and ZHENG relate to neural networks adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and ZHENG relates to implementations of a semantic layout system (e.g., a digital image semantic layout manipulation system) that generates refined digital images resembling the style of one or more input images while aligning to the structure of an edited semantic layout, wherein the semantic layout system improves efficiency relative to conventional systems, and the semantic layout system can also improve flexibility relative to conventional systems. Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and ZHENG (US 20220327657 A1), Paragraphs [0017, 0035, 0037]. Regarding claim 11, CHEN and CRICRÌ in view of HE, in further view of ZHENG teach the computing system of claim 10, CHEN further teaches wherein the machine-learned image processing model (Fig. 2, Paragraph [0074] – CHEN discloses introducing a cross-resolution attention enhancement neck CAENeck to the model framework CRTransSar [wherein model framework is a machine-learned image processing model]. See also Paragraph [0077].) comprises a plurality of backbone blocks arranged in a sequence one after the other (Fig. 6, Paragraph [0071] – CHEN discloses adding, to the model framework CRTransSar, a feature extraction network CRbackbone [wherein CRbackbone is a backbone block] based on contextual joint representation learning Transformer; Paragraph [0088] – CHEN discloses structural diagram of the entire feature extraction network is as shown in FIG. 6.). Regarding claim 12, CHEN and CRICRÌ in view of HE, in further view of ZHENG teach the computing system of claim 10, CHEN and CRICRÌ in view of HE fail to explicitly teach wherein, for each backbone block, the hierarchical structure of the plurality of encoders and the plurality of decoders is trained with loss accumulating across multiple scales of the plurality of encoders and the plurality of decoders. However, ZHENG explicitly teaches wherein, for each backbone block (Fig. 5, Paragraph [0124] – ZHENG discloses the semantic layout system 106 trains a Stage-1 network portion of the digital image layout neural network 320 [wherein the digital image layout neural network 320 is comprised of each backbone block] that includes the dilated convolutional encoder 508 and/or the contextual attention encoder 510 along with the coarse decoder 512.), the hierarchical structure of the plurality of encoders and the plurality of decoders is trained with loss accumulating across multiple scales of the plurality of encoders and the plurality of decoders (Fig. 4A, Paragraph [0063] – ZHENG discloses the digital image layout neural network 320 includes multiple encoders and/or decoders to generate accurate and realistic refined images that match the style of the input image 302 while adhering to the arrangement and structure provided by the edited layout 304. Paragraph [0047] – ZHENG discloses the semantic layout system employs multiple loss functions and minimizes overall loss between multiple networks and models. Paragraph [0101] – ZHENG discloses the semantic layout system 106 compares error loss between low-resolution and high-resolution versions of the sparse correspondence feature map 408. Paragraph [0128] – ZHENG discloses the semantic layout system 106 back propagates the combined loss, as described above, to the fine decoder 514 while also back-propagating the pixel-wise custom-character loss to the coarse decoder 512, the dilated convolutional encoder 508, and/or the contextual attention encoder.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of CHEN and CRICRÌ in view of HE of having a computing system for resolution-flexible image processing, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to execute: a machine-learned image processing model configured to process input image data to generate an output prediction, wherein the machine-learned image processing model comprises one or more resolution-flexible multi-axis attention blocks, with the teachings of ZHENG having wherein, for each backbone block, the hierarchical structure of the plurality of encoders and the plurality of decoders is trained with loss accumulating across multiple scales of the plurality of encoders and the plurality of decoders. Wherein CHEN’s computing system for resolution-flexible image processing wherein, for each backbone block, the hierarchical structure of the plurality of encoders and the plurality of decoders is trained with loss accumulating across multiple scales of the plurality of encoders and the plurality of decoders. The motivation behind this modification would have been to provide an improved system for image processing that has the flexibility of modeling at various scales and accurately transfers visual details to new layouts, since both CHEN and ZHENG relate to neural networks adapted for image or video recognition, wherein CHEN relates to the field of target detection, and in particular, to a contextual visual-based synthetic-aperture radar (SAR) target detection method and apparatus, and a storage medium, where the model can make full use of contextual information to perform joint representation learning, and extract richer contextual feature salient information, thereby improving the feature description of multi-scale SAR targets, and ZHENG relates to implementations of a semantic layout system (e.g., a digital image semantic layout manipulation system) that generates refined digital images resembling the style of one or more input images while aligning to the structure of an edited semantic layout, wherein the semantic layout system improves efficiency relative to conventional systems, and the semantic layout system can also improve flexibility relative to conventional systems. Please see CHEN (US 20230184927 A1), Paragraphs [0002, 0045], and ZHENG (US 20220327657 A1), Paragraphs [0017, 0035, 0037]. Conclusion Listed below are the prior arts made of record and not relied upon but are considered pertinent to applicant’s disclosure. DONG et al. (US 20230086141 A1) - Described herein are computer-implemented systems and methods of creating a labeled image. The computer-implemented systems or methods may comprise tiling an input image comprising a native resolution and a native size to generate a set of tiled images and tiling instructions or an encoding thereof used to generate the set of tile images, labeling the set of tiled images to generate a set of labeled tile images, merging the set of labeled tile images using the tiling instructions or encoding thereof to generate a labeled merged image, wherein the labeled merged image comprises the native resolution, the native size, and one or more merged labels... … Fig. 1, Abstract. QUINTON et al. (US 20220327811 A1) - The present disclosures provides systems and methods for generating composite based data for use in machine learning systems, such as for use in training a machine learning system on the composite based data to identify an object of interest. In an aspect, a method of generating composite based data for use in training machine learning systems comprises: receiving a plurality of images, each of the plurality of images having a corresponding label; generating a composite image comprising the plurality of images, each of the plurality of images occupying a region of the composite image; generating a response map for the composite image, the response map having a plurality of response entries, each response entry encoded with a desired label corresponding to a fragment of the composite image..… Fig. 1, Abstract. WSHAH et al. (US 20170372174 A1) - According to exemplary methods of training a convolutional neural network, input images are received into a computerized device having an image processor. The image processor evaluates the input images using first convolutional layers. The number of first convolutional layers is based on a first size for the input images. Each layer of the first convolutional layers receives layer input signals comprising features of the input images and generates layer output signals that include signals from the input images and ones of the layer output signals from previous layers within the first convolutional layers..… Fig. 1, Abstract. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to BEZAWIT N SHIMELES whose telephone number is (571)272-7663. The examiner can normally be reached M-F 7:30am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /BEZAWIT NOLAWI SHIMELES/Examiner, Art Unit 2673 /CHINEYERE WILLS-BURNS/Supervisory Patent Examiner, Art Unit 2673
Read full office action

Prosecution Timeline

Jul 05, 2024
Application Filed
May 05, 2026
Non-Final Rejection mailed — §103
Jul 27, 2026
Examiner Interview Summary
Jul 27, 2026
Applicant Interview (Telephonic)
Aug 05, 2026
Response Filed
Sep 23, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725390
DEVICE AND METHOD FOR DETECTING RADIOGRAPHIC OBJECT USING EXTREMAL DATA
2y 10m to grant Granted Sep 01, 2026
Patent 12711736
FACE IMAGE CLUSTERING METHOD AND SYSTEM BASED ON LOCALIZED SIMPLE MULTIPLE KERNEL K-MEANS
2y 6m to grant Granted Aug 18, 2026
Patent 12705714
JITTER CORRECTION IMAGE ANALYSIS
2y 7m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 3 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
89%
Grant Probability
89%
With Interview (+0.0%)
2y 7m (~4m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 9 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month