Prosecution Insights
Last updated: October 02, 2026
Application No. 18/876,650

IMAGE SUPER-RESOLUTION METHOD AND APPARATUS

Non-Final OA §103§112
Filed
Dec 18, 2024
Priority
Dec 28, 2022 — CN 202211699926.5 +1 more
Examiner
WOLFSON, ETHAN NOAH
Art Unit
Tech Center
Assignee
Beijing Zitiao Network Technology Co., Ltd.
OA Round
1 (Non-Final)
86%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 86% — above average
86%
Career Allowance Rate
6 granted / 7 resolved
+25.7% vs TC avg
Strong +50% interview lift
Without
With
+50.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
26 currently pending
Career history
30
Total Applications
across all art units

Statute-Specific Performance

§101
4.7%
-35.3% vs TC avg
§103
75.6%
+35.6% vs TC avg
§102
8.7%
-31.3% vs TC avg
§112
8.7%
-31.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 7 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Receipt is acknowledged of certified copies of papers submitted under 35 U.S.C. 119(a)-(d), which papers have been placed of record in the file. Information Disclosure Statement The information disclosure statements (IDS) submitted on 03/10/2025 is being considered by the examiner. Claim Objections Claims 1-2, 5, 8, 11-12, 15, 18, 21, 23 are objected to because of the following informalities: In claim 1, line 4, the term “the first image feature by using a channel attention network” should be changed to “the first image feature by utilizing a channel attention network” in order to enhance and clarify the claim language. In claim 1, line 6-7, the term “self-attention layers is configured to divide” should be changed to “self-attention layers is configured to: divide” in order to avoid typographical issue and avoid a sentence run-on. In claim 1, line 7-8, the term “first feature blocks, separately recalibrate” should be changed to “first feature blocks; separately recalibrate” in order to avoid typographical issue and avoid a sentence run-on. In claim 1, line 9-10, the term “first feature block, combine second feature blocks” should be changed to “first feature block; combine second feature blocks” in order to avoid typographical issue and avoid a sentence run-on. In claim 1, line 10-11, the term “combined feature, and obtain” should be changed to “combined feature; and obtain” in order to avoid typographical issue and avoid a sentence run-on. In claim 2, line 1-2, the term “wherein the separately recalibrating the plurality” should be changed to “wherein the separately recalibrate the plurality” in order to avoid typographical issue and maintain consistency/clarity throughout the claims. In claim 2, line 2, the term “the channel self-attention mechanism, to obtain the second feature block” should be changed to “the channel self-attention mechanismin order to avoid typographical issue and maintain consistency/clarity throughout the claims. In claim 2, line 6, the term “the flattened feature by using a first fully connected layer” should be changed to “the flattened feature by utilizing a first fully connected layer” in order to enhance and clarify the claim language. In claim 2, line 11, the term “channel attention matrix, to obtain a recalibrated feature” should be changed to “channel attention matrixin order to avoid typographical issue and maintain consistency/clarity throughout the claims. In claim 2, line 13, the term “the recalibrated feature, to obtain the second feature” should be changed to “the recalibrated featurein order to avoid typographical issue and maintain consistency/clarity throughout the claims. In claim 5, line 2, the term “wherein the obtaining the output feature” should be changed to “wherein the obtain the output feature” in order to avoid typographical issue and maintain consistency/clarity throughout the claims. In claim 5, line 4, the term “combined feature by using a feedforward network” should be changed to “combined feature by utilizing a feedforward network” in order to enhance and clarify the claim language. In claim 8, line 6, the term “the upsampled feature, to obtain the super- resolution image” should be changed to “the upsampled featurein order to avoid typographical issue and maintain consistency/clarity throughout the claims. In claim 11, line 7, the term “the first image feature by using a channel attention network” should be changed to “the first image feature by utilizing a channel attention network” in order to enhance and clarify the claim language. In claim 11, line 9, the term “self-attention layers is configured to divide” should be changed to “self-attention layers is configured to: divide” in order to avoid typographical issue and avoid a sentence run-on. In claim 11, line 10-11, the term “first feature blocks, separately recalibrate” should be changed to “first feature blocks; separately recalibrate” in order to avoid typographical issue and avoid a sentence run-on. In claim 11, line 12-13, the term “first feature block, combine second feature blocks” should be changed to “first feature block; combine second feature blocks” in order to avoid typographical issue and avoid a sentence run-on. In claim 11, line 13-14, the term “combined feature, and obtain” should be changed to “combined feature; and obtain” in order to avoid typographical issue and avoid a sentence run-on. In claim 12, line 7, the term “the first image feature by using a channel attention network” should be changed to “the first image feature by utilizing a channel attention network” in order to enhance and clarify the claim language. In claim 12, line 9, the term “self-attention layers is configured to divide” should be changed to “self-attention layers is configured to: divide” in order to avoid typographical issue and avoid a sentence run-on. In claim 12, line 10-11, the term “first feature blocks, separately recalibrate” should be changed to “first feature blocks; separately recalibrate” in order to avoid typographical issue and avoid a sentence run-on. In claim 12, line 12-13, the term “first feature block, combine second feature blocks” should be changed to “first feature block; combine second feature blocks” in order to avoid typographical issue and avoid a sentence run-on. In claim 12, line 13-14, the term “combined feature, and obtain” should be changed to “combined feature; and obtain” in order to avoid typographical issue and avoid a sentence run-on. In claim 15, line 2, the term “wherein the separately recalibrating the plurality” should be changed to “wherein the separately recalibrate the plurality” in order to avoid typographical issue and maintain consistency/clarity throughout the claims. In claim 15, line 2-3, the term “the channel self-attention mechanism, to obtain the second feature block” should be changed to “the channel self-attention mechanismin order to avoid typographical issue and maintain consistency/clarity throughout the claims. In claim 15, line 12-13, the term “channel attention matrix, to obtain a recalibrated feature” should be changed to “channel attention matrixin order to avoid typographical issue and maintain consistency/clarity throughout the claims. In claim 15, line 14, the term “the recalibrated feature, to obtain the second feature” should be changed to “the recalibrated featurein order to avoid typographical issue and maintain consistency/clarity throughout the claims. In claim 18, line 2, the term “wherein the obtaining the output feature” should be changed to “wherein the obtain the output feature” in order to avoid typographical issue and maintain consistency/clarity throughout the claims. In claim 21, line 6-7, the term “the upsampled feature, to obtain the super- resolution image” should be changed to “the upsampled featurein order to avoid typographical issue and maintain consistency/clarity throughout the claims. In claim 23, line 2, the term “wherein the separately recalibrating the plurality” should be changed to “wherein the separately recalibrate the plurality” in order to avoid typographical issue and maintain consistency/clarity throughout the claims. In claim 23, line 2-3, the term “the channel self-attention mechanism, to obtain the second feature block” should be changed to “the channel self-attention mechanismin order to avoid typographical issue and maintain consistency/clarity throughout the claims. In claim 23, line 11-12, the term “channel attention matrix, to obtain a recalibrated feature” should be changed to “channel attention matrixin order to avoid typographical issue and maintain consistency/clarity throughout the claims. In claim 23, line 13, the term “the recalibrated feature, to obtain the second feature” should be changed to “the recalibrated featurein order to avoid typographical issue and maintain consistency/clarity throughout the claims. Appropriate correction is required. In claim 23, line 6, the term “flattened feature by using a first fully connected layer” should be changed to “flattened feature by utilizing a first fully connected layer” in order to enhance and clarify the claim language. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that use the word “means” or “step” but are nonetheless not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph because the claim limitation(s) recite(s) sufficient structure, materials, or acts to entirely perform the recited function. Claims 1, 5, 11-12 and 18 recite limitations that use words like “means” (or “step”) or similar terms with functional language but do not invoke 35 U.S.C. 112(f): Claim 1; recites the limitation, “using a channel attention network to……,” [Line 4]. Claim 5; recites the limitation, “processing the combined feature by using a feedforward network to……,” [Line 4]. Claim 11; recites the limitation, “the memory is configured to……,” [Line 2]. Claim 11; recites the limitation, “the processor is configured to……,” [Line 2]. Claim 11; recites the limitation, “using a channel attention network to……,” [Line 7]. Claim 12; recites the limitation, “using a channel attention network to……,” [Line 7]. Claim 18; recites the limitation, “processing the combined feature by using a feedforward network to……,” [Line 4]. Such claim limitation(s) is/are: “memory….” has a structure associated with it a memory. “processor….” has a structure associated with it a processor. “a channel attention network….” has a structure associated with it a network. “a feedforward network….” has a structure associated with it a network. Because this/these claim limitation(s) is/are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are not being interpreted to cover only the corresponding structure, material, or acts described in the specification as performing the claimed function, and equivalents thereof. If applicant intends to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to remove the structure, materials, or acts that performs the claimed function; or (2) present a sufficient showing that the claim limitation(s) does/do not recite sufficient structure, materials, or acts to perform the claimed function. Claims 1 and 11-12 recite limitations that use words like “means” (or “step”) or similar terms with functional language and do invoke 35 U.S.C. 112(f): Claim 1; recites the limitation, “any one of the local channel self-attention layers is configured to…..” [Line 6]. Claim 11; recites the limitation, “any one of the local channel self-attention layers is configured to…..” [Line 9]. Claim 12; recites the limitation, “executed by a computing device…..” [Line 3]. Claim 12; recites the limitation, “causes the computing device to perform…..” [Line 3]. Claim 12; recites the limitation, “any one of the local channel self-attention layers is configured to…..” [Line 9]. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. After a careful analysis, as disclosed above, and a careful review of the specification the following limitations in claims 1 and 11-12; (i) “any one of the local channel self-attention layers” (Fig. 10, #102 processor. Page 19, Line [19-22]- FIG. 10 is a structural diagram of an electronic device according to an embodiment of the present application. As shown in FIG. 10, the electronic device provided in this embodiment includes a memory 101 and a processor 102. The memory 101 is configured to store a computer program. The processor 102 is configured to, when calling the computer program, perform the image super-resolution method provided in the foregoing embodiments. The any one of the local channel self-attention layers is performed by a processor thus has sufficient structure or material wherein is a processor.). (ii) “computing device” (Page 19, Line [23-29] and Page 20, Line [18-24]- Based on the same inventive concept, an embodiment of the present application further provides a computer- readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a computing device, the computing device is caused to perform the image super-resolution method provided in the foregoing embodiments. Based on the same inventive concept, an embodiment of the present application further provides a computer program product. When the computer program product runs on a computing device, the computing device is caused to perform the image super-resolution method provided in the foregoing embodiments. Examples of the computer storage medium include, but are not limited to, a phase change memory (PRAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), another type of random access memory (RAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory or another memory technology, a compact disc read-only memory (CD-ROM), a digital versatile disc (DVD) or another optical storage, a magnetic cassette, a magnetic disk storage, or another magnetic storage device, or any other non-transmission medium, which may be used for storing information that can be accessed by the computing device. The computing device thus does not have sufficient structure or material associated with it.). If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 12 along with its dependent claims is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention. Claim 12 limitations: Claim 12; recites the limitation, “executed by a computing device…..” [Line 3]. Claim 12; recites the limitation, “causes the computing device to perform…..” [Line 3]. Claim 12 respectively invokes 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. The specification is devoid of adequate structure to perform the claimed functions. The specification does not provide sufficient details such that one of the ordinary skill in the art would understand which structure performed(s) the claimed function. Therefore, the claim is indefinite and is rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph. Applicant may: (a) Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph; (b) Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (c) Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)). If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either: (a) Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (b) Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181. The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claim 12 along with its dependent claims is rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for pre-AIA the inventor(s), at the time the application was filed, had possession of the claimed invention. As described above, the disclosure does not provide adequate structure to perform the claimed function in the recited limitation. Claim 12; recites the limitation, “executed by a computing device…..” [Line 3]. Claim 12; recites the limitation, “causes the computing device to perform…..” [Line 3]. The specification does not demonstrate that applicant has made an invention that achieves the claimed function because the invention is not described with sufficient detail such that one of ordinary skill in the art can reasonably conclude that the inventor had possession of the claimed invention. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 9, 11-12, and 22 are rejected under 35 U.S.C. 103 as being unpatentable over XU et al. (US 20240070809 A1), hereinafter referenced as XU, in view of SIDIYA et al. (US 20230169626 A1), hereinafter referenced as SIDIYA. Regarding claim 1, XU explicitly teaches an image super-resolution method, comprising (Fig. 2. Paragraph [0041]-XU discloses the LIT framework 200 can be aimed to produce a high-resolution (HR) 207 from a given low-resolution (LR) image 201.): performing feature extraction on a to-be-super-resolved image to obtain a first image feature (Fig. 2. Paragraph [0041]-XU discloses an encoder 202 can first extract a feature embedding, denoted by Z ∈ R.sup.H×W×C, from the LR image I.sup.LR 201.); separately recalibrate the plurality of first feature blocks (Fig. 4-5. Paragraph [0043]-XU discloses the LIT 400 can first project an extracted feature embedding custom-character by separating it into four convolutional layers 406-409 and then performing upsampling operations to produce four latent embeddings, corresponding to query q, key k, value v, and frequency f (wherein the latent embeddings are first feature blocks).) based on a channel self-attention mechanism (Fig. 5. Paragraph [0045]-XU discloses FIG. 5 shows a cross-scale local attention block (CSLAB) 500 according to embodiments of the disclosure. The LIT exploits the CSLAB 500 to perform a local attention mechanism over a local grid to generate a local latent embedding custom-character for each HR coordinate.) to obtain a second feature block corresponding to each first feature block (Fig. 5. Paragraph [0045]-XU discloses the CSLAB 500 can first calculate an inner product 501 of the query q and the key k, and adds the result with the relative positional bias B to derive an attention matrix via element-wise addition 502. The relative positional bias B can be produced by feeding a set of local relative coordinates δx into a positional encoding block 505 and then a fully-connected layer 506. The attention matrix can then be normalized by a softmax operation σ 503 to generate a local attention map (wherein the recalibration is normalizing by the softmax operation and wherein the second feature block is the local attention map).), generating, based on the second image feature (Fig. 8. Paragraph [0053]-XU discloses LIT.sup.i can estimate the residual image I.sub.r.sup.i from the feature embedding custom-character.sup.i with the coordinate and the corresponding cell cell.sup.i 811-813. Further in paragraph [0053]-XU discloses the final residual image 806 can be produced by adding each individual residual image I.sub.r.sup.i of LIT.sup.i via element-wise addition 814-815 (wherein the residual image is the second image feature).) and the to-be-super-resolved image (Fig. 2. Paragraph [0041]-XU discloses a bilinearly upsampled image 206 can be produced by sending the LR image 201 through a bilinear upsampling operation (e.g., a bilinear upsampling operator) 211 (wherein the to-be-super-resolved image is the LR image #201).), a super-resolution image corresponding to the to-be-super-resolved image (Fig. 2, #207 called high-resolution image. Paragraph [0041]-XU discloses the residual image 205 can be combined with the bilinearly upsampled image 206 via element-wise addition to derive a HR image 207 (wherein a HR image is a super resolution image).). XU fails to explicitly teach processing the first image feature by using a channel attention network to obtain a second image feature, wherein the channel attention network comprises multi-level cascaded local channel self-attention layers, and any one of the local channel self-attention layers is configured to divide an input feature of the local channel self-attention layer into a plurality of first feature blocks, combine second feature blocks corresponding to the plurality of first feature blocks to obtain a combined feature, and obtain an output feature of the local channel self-attention layer based on the combined feature; and. However, SIDIYA explicitly teaches processing the first image feature by using a channel attention network to obtain a second image feature (Fig. 1 and 3, illustrate each encoder block contains a transformer block. Paragraph [0039]-SIDIYA discloses the encoder block 101-2 receives the encoder features f.sub.1 from the encoder block 101-1 and generates the encoder features f.sub.2. Further in paragraph [0060]-SIDIYA discloses FIG. 3 is a block diagram illustrating a transformer block in the neural network system shown in FIG. 1, FIG. 2 or FIG. 5 in accordance with an example of the present disclosure. As shown in FIG. 3, the transformer block 300 includes a self-attention layer 301 with a skip connection, a convolution layer 302, a Leaky Rectified Linear Activation (LReLU) layer 303, and a convolution layer 304.), wherein the channel attention network comprises multi-level cascaded local channel self-attention layers (Fig. 1 and 4, illustrate a channel attention network comprised of multi-level cascaded local channel self-attention layers. Paragraph [0036]-SIDIYA discloses the encoder network, i.e., the encoder, is built using successive transformer blocks, each of which may include self-attention layer and residual block. Further in paragraph [0062]-SIDIYA discloses FIG. 4 is a block diagram illustrating a self-attention layer in the transformer block shown in FIG. 3 in accordance with an example of the present disclosure. The self-attention layer 301 may include a plurality of projection layers, e.g., separable depth-wise convolution layers, each of which respectively learns query, key, and value features.), and any one of the local channel self-attention layers is configured to divide an input feature of the local channel self-attention layer into a plurality of first feature blocks (Fig. 4. Paragraph [0062]-SIDIYA discloses the self-attention layer 301 may include a plurality of projection layers, e.g., separable depth-wise convolution layers, each of which respectively learns query, key, and value features. The query, key, and value features may be embeddings related to inputs of the self-attention layer. The outputs of the projection layers are divided into small patches through a patch division layer 402. K, Q and V may be respectively matrices of a set of key features, query features and value features.), combine second feature blocks corresponding to the plurality of first feature blocks to obtain a combined feature (Fig. 4, #405 called patch merge. Paragraph [0062]-SIDIYA discloses the key features K is transposed using a transpose layer 403, the query features Q and the transpose of key features K are multiplied, and an attention map is obtained through a softmax layer 404. Moreover, the attention map is multiplied by the value features V and the output is merged using an inverse of the patch division operation through a patch merge layer 405 (wherein the attention map is second feature blocks and the output merged is the combined feature).), and obtain an output feature of the local channel self-attention layer based on the combined feature (Fig. 4. Paragraph [0062]-SIDIYA discloses the attention map is multiplied by the value features V and the output is merged using an inverse of the patch division operation through a patch merge layer 405 and a final convolution is applied using a convolution layer 406 to generate the output of the self-attention layer 301 (wherein the output feature is the output of the self-attention layer).); and Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU of an image super-resolution method, comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of SIDIYA of processing the first image feature by using a channel attention network to obtain a second image feature, wherein the channel attention network comprises multi-level cascaded local channel self-attention layers, and any one of the local channel self-attention layers is configured to divide an input feature of the local channel self-attention layer into a plurality of first feature blocks, combine second feature blocks corresponding to the plurality of first feature blocks to obtain a combined feature, and obtain an output feature of the local channel self-attention layer based on the combined feature; and. Wherein having XU’s method of image super-resolution processing the first image feature by using a channel attention network to obtain a second image feature, wherein the channel attention network comprises multi-level cascaded local channel self-attention layers, and any one of the local channel self-attention layers is configured to divide an input feature of the local channel self-attention layer into a plurality of first feature blocks, combine second feature blocks corresponding to the plurality of first feature blocks to obtain a combined feature, and obtain an output feature of the local channel self-attention layer based on the combined feature; and. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and SIDIYA employ transformers in order to generate a super-resolution image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while SIDIYA divides feature maps to patch block so as to reduce the computational cost without losing results performance. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and SIDIYA et al. (US 20230169626 A1), Paragraph [0062]. Regarding claim 9, XU in view of SIDIYA explicitly teach the method according to claim 1, XU further explicitly teaches wherein the performing feature extraction on the to-be-super-resolved image to obtain the first image feature comprises (Fig. 2. Paragraph [0041]-XU discloses an encoder 202 can first extract a feature embedding, denoted by Z ∈ R.sup.H×W×C, from the LR image I.sup.LR 201.): Although XU explicitly teaches convolution processing, XU is silent on performing convolution processing on the to-be-super-resolved image to obtain the first image feature. However, SIDIYA explicitly teaches performing convolution processing on the to-be-super-resolved image to obtain the first image feature (Fig. 1. Paragraph [0038]-SADIYA discloses the encoder block 101-1 receives an input image having a low-resolution and extracts encoder features f.sub.1 from the input image. Further in paragraph [0038]-SIDIYA discloses in the encoder block 101-1, the convolution layer EC 1 and the plurality of transformer blocks T11-T16 are stacked to each other, and the convolution layer EC 1 is followed by the plurality of transformer blocks T11-T16.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA of an image super-resolution method, comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of SIDIYA of performing convolution processing on the to-be-super-resolved image to obtain the first image feature. Wherein having XU’s method of image super-resolution having performing convolution processing on the to-be-super-resolved image to obtain the first image feature. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and SIDIYA employ transformers in order to generate a super-resolution image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while SIDIYA divides feature maps to patch block so as to reduce the computational cost without losing results performance. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and SIDIYA et al. (US 20230169626 A1), Paragraph [0062]. Regarding claim 11, XU explicitly teaches an electronic device, comprising (Fig. 2. Paragraph [0062]-XU discloses the processes and functions described herein can be implemented as a computer program which, when executed by one or more processors, can cause the one or more processors to perform the respective processes and functions. Further in paragraph [0066]-XU discloses the computer readable medium may include any apparatus that stores, communicates, propagates, or transports the computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable medium can be magnetic, optical, electronic, electromagnetic, infrared, or semiconductor system (or apparatus or device).): a memory (Paragraph [0066]-XU discloses the computer-readable medium may include a computer-readable non-transitory storage medium such as a semiconductor or solid state memory, magnetic tape, a removable computer diskette, a random access memory (RAM), a read-only memory (ROM), a magnetic disk and an optical disk, and the like.) and a processor (Paragraph [0065]-XU discloses the processes and functions described herein can be implemented as a computer program which, when executed by one or more processors, can cause the one or more processors to perform the respective processes and functions.), wherein the memory is configured to store a computer program (Paragraph [0065]-XU discloses the computer program may be stored or distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with, or as part of, other hardware.), and the processor is configured to, when calling the computer program, cause the electronic device to perform an image super-resolution method comprising (Paragraph [0065]-XU discloses the processes and functions described herein can be implemented as a computer program which, when executed by one or more processors, can cause the one or more processors to perform the respective processes and functions. The computer program may be stored or distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with, or as part of, other hardware.): performing feature extraction on a to-be-super-resolved image to obtain a first image feature (Fig. 2. Paragraph [0041]-XU discloses an encoder 202 can first extract a feature embedding, denoted by Z ∈ R.sup.H×W×C, from the LR image I.sup.LR 201.); separately recalibrate the plurality of first feature blocks (Fig. 4-5. Paragraph [0043]-XU discloses the LIT 400 can first project an extracted feature embedding custom-character by separating it into four convolutional layers 406-409 and then performing upsampling operations to produce four latent embeddings, corresponding to query q, key k, value v, and frequency f (wherein the latent embeddings are first feature blocks).) based on a channel self-attention mechanism (Fig. 5. Paragraph [0045]-XU discloses FIG. 5 shows a cross-scale local attention block (CSLAB) 500 according to embodiments of the disclosure. The LIT exploits the CSLAB 500 to perform a local attention mechanism over a local grid to generate a local latent embedding custom-character for each HR coordinate.) to obtain a second feature block corresponding to each first feature block (Fig. 5. Paragraph [0045]-XU discloses the CSLAB 500 can first calculate an inner product 501 of the query q and the key k, and adds the result with the relative positional bias B to derive an attention matrix via element-wise addition 502. The relative positional bias B can be produced by feeding a set of local relative coordinates δx into a positional encoding block 505 and then a fully-connected layer 506. The attention matrix can then be normalized by a softmax operation σ 503 to generate a local attention map (wherein the recalibration is normalizing by the softmax operation and wherein the second feature block is the local attention map).), generating, based on the second image feature (Fig. 8. Paragraph [0053]-XU discloses LIT.sup.i can estimate the residual image I.sub.r.sup.i from the feature embedding custom-character.sup.i with the coordinate and the corresponding cell cell.sup.i 811-813. Further in paragraph [0053]-XU discloses the final residual image 806 can be produced by adding each individual residual image I.sub.r.sup.i of LIT.sup.i via element-wise addition 814-815 (wherein the residual image is the second image feature).) and the to-be-super-resolved image (Fig. 2. Paragraph [0041]-XU discloses a bilinearly upsampled image 206 can be produced by sending the LR image 201 through a bilinear upsampling operation (e.g., a bilinear upsampling operator) 211 (wherein the to-be-super-resolved image is the LR image #201).), a super-resolution image corresponding to the to-be-super-resolved image (Fig. 2, #207 called high-resolution image. Paragraph [0041]-XU discloses the residual image 205 can be combined with the bilinearly upsampled image 206 via element-wise addition to derive a HR image 207 (wherein a HR image is a super resolution image).). XU fails to explicitly teach processing the first image feature by using a channel attention network to obtain a second image feature, wherein the channel attention network comprises multi-level cascaded local channel self-attention layers, and any one of the local channel self-attention layers is configured to divide an input feature of the local channel self-attention layer into a plurality of first feature blocks, combine second feature blocks corresponding to the plurality of first feature blocks to obtain a combined feature, and obtain an output feature of the local channel self-attention layer based on the combined feature; and. However, SIDIYA explicitly teaches processing the first image feature by using a channel attention network to obtain a second image feature (Fig. 1 and 3, illustrate each encoder block contains a transformer block. Paragraph [0039]-SIDIYA discloses the encoder block 101-2 receives the encoder features f.sub.1 from the encoder block 101-1 and generates the encoder features f.sub.2. Further in paragraph [0060]-SIDIYA discloses FIG. 3 is a block diagram illustrating a transformer block in the neural network system shown in FIG. 1, FIG. 2 or FIG. 5 in accordance with an example of the present disclosure. As shown in FIG. 3, the transformer block 300 includes a self-attention layer 301 with a skip connection, a convolution layer 302, a Leaky Rectified Linear Activation (LReLU) layer 303, and a convolution layer 304.), wherein the channel attention network comprises multi-level cascaded local channel self-attention layers (Fig. 1 and 4, illustrate a channel attention network comprised of multi-level cascaded local channel self-attention layers. Paragraph [0036]-SIDIYA discloses the encoder network, i.e., the encoder, is built using successive transformer blocks, each of which may include self-attention layer and residual block. Further in paragraph [0062]-SIDIYA discloses FIG. 4 is a block diagram illustrating a self-attention layer in the transformer block shown in FIG. 3 in accordance with an example of the present disclosure. The self-attention layer 301 may include a plurality of projection layers, e.g., separable depth-wise convolution layers, each of which respectively learns query, key, and value features.), and any one of the local channel self-attention layers is configured to divide an input feature of the local channel self-attention layer into a plurality of first feature blocks (Fig. 4. Paragraph [0062]-SIDIYA discloses the self-attention layer 301 may include a plurality of projection layers, e.g., separable depth-wise convolution layers, each of which respectively learns query, key, and value features. The query, key, and value features may be embeddings related to inputs of the self-attention layer. The outputs of the projection layers are divided into small patches through a patch division layer 402. K, Q and V may be respectively matrices of a set of key features, query features and value features.), combine second feature blocks corresponding to the plurality of first feature blocks to obtain a combined feature (Fig. 4, #405 called patch merge. Paragraph [0062]-SIDIYA discloses the key features K is transposed using a transpose layer 403, the query features Q and the transpose of key features K are multiplied, and an attention map is obtained through a softmax layer 404. Moreover, the attention map is multiplied by the value features V and the output is merged using an inverse of the patch division operation through a patch merge layer 405 (wherein the attention map is second feature blocks and the output merged is the combined feature).), and obtain an output feature of the local channel self-attention layer based on the combined feature (Fig. 4. Paragraph [0062]-SIDIYA discloses the attention map is multiplied by the value features V and the output is merged using an inverse of the patch division operation through a patch merge layer 405 and a final convolution is applied using a convolution layer 406 to generate the output of the self-attention layer 301 (wherein the output feature is the output of the self-attention layer).); and Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU of an electronic device, comprising: a memory and a processor, wherein the memory is configured to store a computer program, and the processor is configured to, when calling the computer program, cause the electronic device to perform an image super-resolution method comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of SIDIYA of processing the first image feature by using a channel attention network to obtain a second image feature, wherein the channel attention network comprises multi-level cascaded local channel self-attention layers, and any one of the local channel self-attention layers is configured to divide an input feature of the local channel self-attention layer into a plurality of first feature blocks, combine second feature blocks corresponding to the plurality of first feature blocks to obtain a combined feature, and obtain an output feature of the local channel self-attention layer based on the combined feature; and. Wherein having XU’s method of image super-resolution processing the first image feature by using a channel attention network to obtain a second image feature, wherein the channel attention network comprises multi-level cascaded local channel self-attention layers, and any one of the local channel self-attention layers is configured to divide an input feature of the local channel self-attention layer into a plurality of first feature blocks, combine second feature blocks corresponding to the plurality of first feature blocks to obtain a combined feature, and obtain an output feature of the local channel self-attention layer based on the combined feature; and. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and SIDIYA employ transformers in order to generate a super-resolution image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while SIDIYA divides feature maps to patch block so as to reduce the computational cost without losing results performance. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and SIDIYA et al. (US 20230169626 A1), Paragraph [0062]. Regarding claim 12, XU explicitly teaches a non-transitory computer-readable storage medium (Paragraph [0066]-XU discloses the computer-readable medium may include a computer-readable non-transitory storage medium such as a semiconductor or solid state memory, magnetic tape, a removable computer diskette, a random access memory (RAM), a read-only memory (ROM), a magnetic disk and an optical disk, and the like.), wherein a computer program is stored on the computer-readable storage medium (Paragraph [0065]-XU discloses the computer program may be stored or distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with, or as part of, other hardware.), and the computer program, when executed by a computing device, causes the computing device to perform an image super-resolution method comprising (Paragraph [0065]-XU discloses the processes and functions described herein can be implemented as a computer program which, when executed by one or more processors, can cause the one or more processors to perform the respective processes and functions. The computer program may be stored or distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with, or as part of, other hardware.): performing feature extraction on a to-be-super-resolved image to obtain a first image feature (Fig. 2. Paragraph [0041]-XU discloses an encoder 202 can first extract a feature embedding, denoted by Z ∈ R.sup.H×W×C, from the LR image I.sup.LR 201.); separately recalibrate the plurality of first feature blocks (Fig. 4-5. Paragraph [0043]-XU discloses the LIT 400 can first project an extracted feature embedding custom-character by separating it into four convolutional layers 406-409 and then performing upsampling operations to produce four latent embeddings, corresponding to query q, key k, value v, and frequency f (wherein the latent embeddings are first feature blocks).) based on a channel self-attention mechanism (Fig. 5. Paragraph [0045]-XU discloses FIG. 5 shows a cross-scale local attention block (CSLAB) 500 according to embodiments of the disclosure. The LIT exploits the CSLAB 500 to perform a local attention mechanism over a local grid to generate a local latent embedding custom-character for each HR coordinate.) to obtain a second feature block corresponding to each first feature block (Fig. 5. Paragraph [0045]-XU discloses the CSLAB 500 can first calculate an inner product 501 of the query q and the key k, and adds the result with the relative positional bias B to derive an attention matrix via element-wise addition 502. The relative positional bias B can be produced by feeding a set of local relative coordinates δx into a positional encoding block 505 and then a fully-connected layer 506. The attention matrix can then be normalized by a softmax operation σ 503 to generate a local attention map (wherein the recalibration is normalizing by the softmax operation and wherein the second feature block is the local attention map).), generating, based on the second image feature (Fig. 8. Paragraph [0053]-XU discloses LIT.sup.i can estimate the residual image I.sub.r.sup.i from the feature embedding custom-character.sup.i with the coordinate and the corresponding cell cell.sup.i 811-813. Further in paragraph [0053]-XU discloses the final residual image 806 can be produced by adding each individual residual image I.sub.r.sup.i of LIT.sup.i via element-wise addition 814-815 (wherein the residual image is the second image feature).) and the to-be-super-resolved image (Fig. 2. Paragraph [0041]-XU discloses a bilinearly upsampled image 206 can be produced by sending the LR image 201 through a bilinear upsampling operation (e.g., a bilinear upsampling operator) 211 (wherein the to-be-super-resolved image is the LR image #201).), a super-resolution image corresponding to the to-be-super-resolved image (Fig. 2, #207 called high-resolution image. Paragraph [0041]-XU discloses the residual image 205 can be combined with the bilinearly upsampled image 206 via element-wise addition to derive a HR image 207 (wherein a HR image is a super resolution image).). XU fails to explicitly teach processing the first image feature by using a channel attention network to obtain a second image feature, wherein the channel attention network comprises multi-level cascaded local channel self-attention layers, and any one of the local channel self-attention layers is configured to divide an input feature of the local channel self-attention layer into a plurality of first feature blocks, combine second feature blocks corresponding to the plurality of first feature blocks to obtain a combined feature, and obtain an output feature of the local channel self-attention layer based on the combined feature; and. However, SIDIYA explicitly teaches processing the first image feature by using a channel attention network to obtain a second image feature (Fig. 1 and 3, illustrate each encoder block contains a transformer block. Paragraph [0039]-SIDIYA discloses the encoder block 101-2 receives the encoder features f.sub.1 from the encoder block 101-1 and generates the encoder features f.sub.2. Further in paragraph [0060]-SIDIYA discloses FIG. 3 is a block diagram illustrating a transformer block in the neural network system shown in FIG. 1, FIG. 2 or FIG. 5 in accordance with an example of the present disclosure. As shown in FIG. 3, the transformer block 300 includes a self-attention layer 301 with a skip connection, a convolution layer 302, a Leaky Rectified Linear Activation (LReLU) layer 303, and a convolution layer 304.), wherein the channel attention network comprises multi-level cascaded local channel self-attention layers (Fig. 1 and 4, illustrate a channel attention network comprised of multi-level cascaded local channel self-attention layers. Paragraph [0036]-SIDIYA discloses the encoder network, i.e., the encoder, is built using successive transformer blocks, each of which may include self-attention layer and residual block. Further in paragraph [0062]-SIDIYA discloses FIG. 4 is a block diagram illustrating a self-attention layer in the transformer block shown in FIG. 3 in accordance with an example of the present disclosure. The self-attention layer 301 may include a plurality of projection layers, e.g., separable depth-wise convolution layers, each of which respectively learns query, key, and value features.), and any one of the local channel self-attention layers is configured to divide an input feature of the local channel self-attention layer into a plurality of first feature blocks (Fig. 4. Paragraph [0062]-SIDIYA discloses the self-attention layer 301 may include a plurality of projection layers, e.g., separable depth-wise convolution layers, each of which respectively learns query, key, and value features. The query, key, and value features may be embeddings related to inputs of the self-attention layer. The outputs of the projection layers are divided into small patches through a patch division layer 402. K, Q and V may be respectively matrices of a set of key features, query features and value features.), combine second feature blocks corresponding to the plurality of first feature blocks to obtain a combined feature (Fig. 4, #405 called patch merge. Paragraph [0062]-SIDIYA discloses the key features K is transposed using a transpose layer 403, the query features Q and the transpose of key features K are multiplied, and an attention map is obtained through a softmax layer 404. Moreover, the attention map is multiplied by the value features V and the output is merged using an inverse of the patch division operation through a patch merge layer 405 (wherein the attention map is second feature blocks and the output merged is the combined feature).), and obtain an output feature of the local channel self-attention layer based on the combined feature (Fig. 4. Paragraph [0062]-SIDIYA discloses the attention map is multiplied by the value features V and the output is merged using an inverse of the patch division operation through a patch merge layer 405 and a final convolution is applied using a convolution layer 406 to generate the output of the self-attention layer 301 (wherein the output feature is the output of the self-attention layer).); and Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU of a non-transitory computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program, when executed by a computing device, causes the computing device to perform an image super-resolution method comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of SIDIYA of processing the first image feature by using a channel attention network to obtain a second image feature, wherein the channel attention network comprises multi-level cascaded local channel self-attention layers, and any one of the local channel self-attention layers is configured to divide an input feature of the local channel self-attention layer into a plurality of first feature blocks, combine second feature blocks corresponding to the plurality of first feature blocks to obtain a combined feature, and obtain an output feature of the local channel self-attention layer based on the combined feature; and. Wherein having XU’s method of image super-resolution processing the first image feature by using a channel attention network to obtain a second image feature, wherein the channel attention network comprises multi-level cascaded local channel self-attention layers, and any one of the local channel self-attention layers is configured to divide an input feature of the local channel self-attention layer into a plurality of first feature blocks, combine second feature blocks corresponding to the plurality of first feature blocks to obtain a combined feature, and obtain an output feature of the local channel self-attention layer based on the combined feature; and. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and SIDIYA employ transformers in order to generate a super-resolution image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while SIDIYA divides feature maps to patch block so as to reduce the computational cost without losing results performance. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and SIDIYA et al. (US 20230169626 A1), Paragraph [0062]. Regarding claim 22, XU in view of SIDIYA explicitly teach the non-transitory computer-readable storage medium according to claim 12, XU further explicitly teaches wherein the performing feature extraction on the to-be-super-resolved image to obtain the first image feature comprises (Fig. 2. Paragraph [0041]-XU discloses an encoder 202 can first extract a feature embedding, denoted by Z ∈ R.sup.H×W×C, from the LR image I.sup.LR 201.): Although XU explicitly teaches convolution processing, XU is silent on performing convolution processing on the to-be-super-resolved image to obtain the first image feature. However, SIDIYA explicitly teaches performing convolution processing on the to-be-super-resolved image to obtain the first image feature (Fig. 1. Paragraph [0038]-SADIYA discloses the encoder block 101-1 receives an input image having a low-resolution and extracts encoder features f.sub.1 from the input image. Further in paragraph [0038]-SIDIYA discloses in the encoder block 101-1, the convolution layer EC 1 and the plurality of transformer blocks T11-T16 are stacked to each other, and the convolution layer EC 1 is followed by the plurality of transformer blocks T11-T16.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA of a non-transitory computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program, when executed by a computing device, causes the computing device to perform an image super-resolution method comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of SIDIYA of performing convolution processing on the to-be-super-resolved image to obtain the first image feature. Wherein having XU’s method of image super-resolution having performing convolution processing on the to-be-super-resolved image to obtain the first image feature. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and SIDIYA employ transformers in order to generate a super-resolution image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while SIDIYA divides feature maps to patch block so as to reduce the computational cost without losing results performance. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and SIDIYA et al. (US 20230169626 A1), Paragraph [0062]. Claims 2-3, 15-16, and 23 are rejected under 35 U.S.C. 103 as being unpatentable over XU et al. (US 20240070809 A1), hereinafter referenced as XU, in view of SIDIYA et al. (US 20230169626 A1), hereinafter referenced as SIDIYA, and further in view of HAN et al. (US 20240161237 A1), hereinafter referenced as HAN, and further in view of WANG (US 20210295089 A1), hereinafter referenced as WANG. Regarding claim 2, XU in view of SIDIYA explicitly teach the method according to claim 1, XU further explicitly teaches wherein the separately recalibrating the plurality of first feature blocks (Fig. 4-5. Paragraph [0043]-XU discloses the LIT 400 can first project an extracted feature embedding custom-character by separating it into four convolutional layers 406-409 and then performing upsampling operations to produce four latent embeddings, corresponding to query q, key k, value v, and frequency f (wherein the latent embeddings are first feature blocks).) based on the channel self-attention mechanism (Fig. 5. Paragraph [0045]-XU discloses FIG. 5 shows a cross-scale local attention block (CSLAB) 500 according to embodiments of the disclosure. The LIT exploits the CSLAB 500 to perform a local attention mechanism over a local grid to generate a local latent embedding custom-character for each HR coordinate.), to obtain the second feature block corresponding to each first feature block comprises (Fig. 5. Paragraph [0045]-XU discloses the CSLAB 500 can first calculate an inner product 501 of the query q and the key k, and adds the result with the relative positional bias B to derive an attention matrix via element-wise addition 502. The relative positional bias B can be produced by feeding a set of local relative coordinates δx into a positional encoding block 505 and then a fully-connected layer 506. The attention matrix can then be normalized by a softmax operation σ 503 to generate a local attention map (wherein the recalibration is normalizing by the softmax operation and wherein the second feature block is the local attention map).): obtaining a channel attention matrix based on the first encoded feature and the second encoded feature (Fig. 5. Paragraph [0045]-XU discloses the CSLAB 500 can first calculate an inner product 501 of the query q and the key k, and adds the result with the relative positional bias B to derive an attention matrix via element-wise addition 502 (wherein query q and the key k are first and second encoded features).); XU fails to explicitly teach recalibrating the third encoded feature based on the channel attention matrix, to obtain a recalibrated feature; and unflattening the recalibrated feature, to obtain the second feature block corresponding to the first feature block. However, SIDIYA explicitly teaches recalibrating the third encoded feature based on the channel attention matrix, to obtain a recalibrated feature (Fig. 4. Paragraph [0062]-SIDIYA discloses the attention map is multiplied by the value features V and the output is merged using an inverse of the patch division operation through a patch merge layer 405 (wherein the attention map is the channel attention matrix, the third encoded feature is value features V, and the recalibrated feature is the output).); and unflattening the recalibrated feature (Fig. 3. Paragraph [0062]-SIDIYA discloses the attention map is multiplied by the value features V and the output is merged using an inverse of the patch division operation through a patch merge layer 405 and a final convolution is applied using a convolution layer 406 to generate the output of the self-attention layer 301. The patch division layer 402 divides feature maps to patch block so as to reduce the computational cost without losing results performance (wherein unflattening is patch merging using the inverse of patch division).), to obtain the second feature block corresponding to the first feature block (Fig. 3. Paragraph [0062]-SIDIYA discloses the attention map is multiplied by the value features V and the output is merged using an inverse of the patch division operation through a patch merge layer 405 and a final convolution is applied using a convolution layer 406 to generate the output of the self-attention layer 301. The patch division layer 402 divides feature maps to patch block so as to reduce the computational cost without losing results performance (wherein the second feature block is the feature output following patch merge 405).). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA of an image super-resolution method, comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of SIDIYA of recalibrating the third encoded feature based on the channel attention matrix, to obtain a recalibrated feature; and unflattening the recalibrated feature, to obtain the second feature block corresponding to the first feature block. Wherein having XU’s method of image super-resolution having recalibrating the third encoded feature based on the channel attention matrix, to obtain a recalibrated feature; and unflattening the recalibrated feature, to obtain the second feature block corresponding to the first feature block. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and SIDIYA employ transformers in order to generate a super-resolution image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while SIDIYA divides feature maps to patch block so as to reduce the computational cost without losing results performance. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and SIDIYA et al. (US 20230169626 A1), Paragraph [0062]. XU in view of SIDIYA fails to explicitly teach flattening the first feature block into a two-dimensional feature to obtain a flattened feature. However, HAN explicitly teaches flattening the first feature block into a two-dimensional feature to obtain a flattened feature (Fig. 2. Paragraph [0048]-HAN discloses the electronic device may obtain an initial image feature matrix F.sub.im∈R.sup.M×C of the face image by flattening and cascading the obtained feature maps F1′, F2′, F3′, and F4′ of four levels. In this case, F.sub.im may be a feature of an ith row and an mth column of an initial image feature matrix, R denotes a set of real numbers representing the matrix F.sub.im, M may be the number of rows of the initial image feature matrix, and C may be the number of columns of the initial image feature matrix (wherein the two-dimensional feature is a matrix of M rows and C columns and the feature matrix is the flattened feature).); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA of an image super-resolution method, comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of HAN of flattening the first feature block into a two-dimensional feature to obtain a flattened feature. Wherein having XU’s method of image super-resolution having flattening the first feature block into a two-dimensional feature to obtain a flattened feature. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and HAN employ transformers in order to generate a super-resolution image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while HAN one or more embodiments may include a method and electronic device with image processing, which is able to aggregate multi-level image features, using an FSR model based on a deformable attention mechanism, to explore facial structure intrinsically and improve super-resolution results, without requesting additional facial prior annotations. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and HAN et al. (US 20240161237 A1), Paragraph [0040]. XU in view of SIDIYA and further in view of HAN fail to explicitly teach encoding the flattened feature by using a first fully connected layer, a second fully connected layer, and a third fully connected layer, respectively, to obtain a first encoded feature, a second encoded feature, and a third encoded feature. However, WANG explicitly teaches encoding the flattened feature (Fig. 1B. Paragraph [0070]-WANG discloses the content tagging network includes a residual net (ResNet) 22 configured to receive the first feature map and generate a second feature map having a scale smaller than a scale of the first feature map.) by using a first fully connected layer (Fig. 1B. Paragraph [0079]-WANG discloses a first fully connected layer 27 configured to receive the fourth feature map and generate the second predicted probability of the content tag of the input image.), a second fully connected layer (Fig. 1B. Paragraph [0104]-WANG discloses after the sixth feature map is generated, the sixth feature map is inputted into the second fully connected layer 33, and the second fully connected layer 33 has a plurality of nodes, each node is a binary classifier. For example, the second fully connected layer 33 has 2048 nodes, and each of the 2048 nodes is a binary classifier.), and a third fully connected layer (Fig. 1B, #45 called third fully connected layer. Paragraph [0117]-WANG discloses the third fully connected layer 45 is configured to receive the ninth feature map and generate the predicted probability of the type tag of the input image.), respectively, to obtain a first encoded feature (Fig. 1B. Paragraph [0084]-WANG discloses the first fully connected layer 27 receives the fourth feature map having the size of 1*1*2048 and predicts probability of the content tag based on the fourth feature map to generate the second predicted probability of the content tag (wherein the predicted probability of the content tag is the first encoded feature).), a second encoded feature (Fig. 1B. Paragraph [0103]-WANG discloses the theme tagging network 3 includes the second fully connected layer 33 configured to receive the sixth feature map and generate the predicted probability of the theme tag of the input image (wherein the predicted probability of theme tag is the second encoded feature).), and a third encoded feature (Fig. 1B, #45 called third fully connected layer. Paragraph [0117]-WANG discloses the third fully connected layer 45 is configured to receive the ninth feature map and generate the predicted probability of the type tag of the input image (wherein the predicted probability of type tag is a third encoded feature).); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA and further in view of HAN of an image super-resolution method, comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of WANG of encoding the flattened feature by using a first fully connected layer, a second fully connected layer, and a third fully connected layer, respectively, to obtain a first encoded feature, a second encoded feature, and a third encoded feature. Wherein having XU’s method of image super-resolution having encoding the flattened feature by using a first fully connected layer, a second fully connected layer, and a third fully connected layer, respectively, to obtain a first encoded feature, a second encoded feature, and a third encoded feature. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and WANG employ networks with fully connected layers to process input images, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while WANG enables the convolutional layer to add or extract a feature having a scale different from that of the input image. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and WANG (US 20210295089 A1), Paragraph [0052]. Regarding claim 3, XU in view of SIDIYA and further in view of HAN and further in view of WANG explicitly teach the method according to claim 2, XU further explicitly teaches wherein the obtaining the channel attention matrix based on the first encoded feature and the second encoded feature comprises (Fig. 5. Paragraph [0045]-XU discloses the CSLAB 500 can first calculate an inner product 501 of the query q and the key k, and adds the result with the relative positional bias B to derive an attention matrix via element-wise addition 502 (wherein query q and the key k are first and second encoded features).): XU fails to explicitly teach performing transposition on the second encoded feature to obtain a fourth encoded feature; and obtaining the channel attention matrix based on the first encoded feature, the fourth encoded feature, and a normalization exponential function. However, SIDIYA explicitly teaches performing transposition on the second encoded feature to obtain a fourth encoded feature (Fig. 4. Paragraph [0062]-SIDIYA discloses the key features K is transposed using a transpose layer 403, the query features Q and the transpose of key features K are multiplied (wherein key features K is the second encoded feature and the fourth encoded feature is the transpose of key features k).); and obtaining the channel attention matrix based on the first encoded feature, the fourth encoded feature, and a normalization exponential function (Fig. 4. Paragraph [0062]-SIDIYA discloses the key features K is transposed using a transpose layer 403, the query features Q and the transpose of key features K are multiplied, and an attention map is obtained through a softmax layer 404 (wherein the attention map is the channel attention matrix, the first encoded feature is query features Q, the fourth encoded feature is the transpose of key features K, and the normalization exponential function is a softmax layer).). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU of an image super-resolution method, comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of SIDIYA of performing transposition on the second encoded feature to obtain a fourth encoded feature; and obtaining the channel attention matrix based on the first encoded feature, the fourth encoded feature, and a normalization exponential function. Wherein having XU’s method of image super-resolution having performing transposition on the second encoded feature to obtain a fourth encoded feature; and obtaining the channel attention matrix based on the first encoded feature, the fourth encoded feature, and a normalization exponential function. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and SIDIYA employ transformers in order to generate a super-resolution image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while SIDIYA divides feature maps to patch block so as to reduce the computational cost without losing results performance. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and SIDIYA et al. (US 20230169626 A1), Paragraph [0062]. Regarding claim 15, XU in view of SIDIYA explicitly teach the non-transitory computer-readable storage medium according to claim 12, XU further explicitly teaches wherein the separately recalibrating the plurality of first feature blocks (Fig. 4-5. Paragraph [0043]-XU discloses the LIT 400 can first project an extracted feature embedding custom-character by separating it into four convolutional layers 406-409 and then performing upsampling operations to produce four latent embeddings, corresponding to query q, key k, value v, and frequency f (wherein the latent embeddings are first feature blocks).) based on the channel self-attention mechanism (Fig. 5. Paragraph [0045]-XU discloses FIG. 5 shows a cross-scale local attention block (CSLAB) 500 according to embodiments of the disclosure. The LIT exploits the CSLAB 500 to perform a local attention mechanism over a local grid to generate a local latent embedding custom-character for each HR coordinate.), to obtain the second feature block corresponding to each first feature block comprises (Fig. 5. Paragraph [0045]-XU discloses the CSLAB 500 can first calculate an inner product 501 of the query q and the key k, and adds the result with the relative positional bias B to derive an attention matrix via element-wise addition 502. The relative positional bias B can be produced by feeding a set of local relative coordinates δx into a positional encoding block 505 and then a fully-connected layer 506. The attention matrix can then be normalized by a softmax operation σ 503 to generate a local attention map (wherein the recalibration is normalizing by the softmax operation and wherein the second feature block is the local attention map).): obtaining a channel attention matrix based on the first encoded feature and the second encoded feature (Fig. 5. Paragraph [0045]-XU discloses the CSLAB 500 can first calculate an inner product 501 of the query q and the key k, and adds the result with the relative positional bias B to derive an attention matrix via element-wise addition 502 (wherein query q and the key k are first and second encoded features).); XU fails to explicitly teach recalibrating the third encoded feature based on the channel attention matrix, to obtain a recalibrated feature; and unflattening the recalibrated feature, to obtain the second feature block corresponding to the first feature block. However, SIDIYA explicitly teaches recalibrating the third encoded feature based on the channel attention matrix, to obtain a recalibrated feature (Fig. 4. Paragraph [0062]-SIDIYA discloses the attention map is multiplied by the value features V and the output is merged using an inverse of the patch division operation through a patch merge layer 405 (wherein the attention map is the channel attention matrix, the third encoded feature is value features V, and the recalibrated feature is the output).); and unflattening the recalibrated feature (Fig. 3. Paragraph [0062]-SIDIYA discloses the attention map is multiplied by the value features V and the output is merged using an inverse of the patch division operation through a patch merge layer 405 and a final convolution is applied using a convolution layer 406 to generate the output of the self-attention layer 301. The patch division layer 402 divides feature maps to patch block so as to reduce the computational cost without losing results performance (wherein unflattening is patch merging using the inverse of patch division).), to obtain the second feature block corresponding to the first feature block (Fig. 3. Paragraph [0062]-SIDIYA discloses the attention map is multiplied by the value features V and the output is merged using an inverse of the patch division operation through a patch merge layer 405 and a final convolution is applied using a convolution layer 406 to generate the output of the self-attention layer 301. The patch division layer 402 divides feature maps to patch block so as to reduce the computational cost without losing results performance (wherein the second feature block is the feature output following patch merge 405).). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU of a non-transitory computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program, when executed by a computing device, causes the computing device to perform an image super-resolution method comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of SIDIYA of recalibrating the third encoded feature based on the channel attention matrix, to obtain a recalibrated feature; and unflattening the recalibrated feature, to obtain the second feature block corresponding to the first feature block. Wherein having XU’s method of image super-resolution having recalibrating the third encoded feature based on the channel attention matrix, to obtain a recalibrated feature; and unflattening the recalibrated feature, to obtain the second feature block corresponding to the first feature block. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and SIDIYA employ transformers in order to generate a super-resolution image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while SIDIYA divides feature maps to patch block so as to reduce the computational cost without losing results performance. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and SIDIYA et al. (US 20230169626 A1), Paragraph [0062]. XU in view of SIDIYA fails to explicitly teach flattening the first feature block into a two-dimensional feature to obtain a flattened feature. However, HAN explicitly teaches flattening the first feature block into a two-dimensional feature to obtain a flattened feature (Fig. 2. Paragraph [0048]-HAN discloses the electronic device may obtain an initial image feature matrix F.sub.im∈R.sup.M×C of the face image by flattening and cascading the obtained feature maps F1′, F2′, F3′, and F4′ of four levels. In this case, F.sub.im may be a feature of an ith row and an mth column of an initial image feature matrix, R denotes a set of real numbers representing the matrix F.sub.im, M may be the number of rows of the initial image feature matrix, and C may be the number of columns of the initial image feature matrix (wherein the two-dimensional feature is a matrix of M rows and C columns and the feature matrix is the flattened feature).); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA of a non-transitory computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program, when executed by a computing device, causes the computing device to perform an image super-resolution method comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of HAN of flattening the first feature block into a two-dimensional feature to obtain a flattened feature. Wherein having XU’s method of image super-resolution having flattening the first feature block into a two-dimensional feature to obtain a flattened feature. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and HAN employ transformers in order to generate a super-resolution image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while HAN one or more embodiments may include a method and electronic device with image processing, which is able to aggregate multi-level image features, using an FSR model based on a deformable attention mechanism, to explore facial structure intrinsically and improve super-resolution results, without requesting additional facial prior annotations. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and HAN et al. (US 20240161237 A1), Paragraph [0040]. XU in view of SIDIYA and further in view of HAN fail to explicitly teach encoding the flattened feature by using a first fully connected layer, a second fully connected layer, and a third fully connected layer, respectively, to obtain a first encoded feature, a second encoded feature, and a third encoded feature. However, WANG explicitly teaches encoding the flattened feature (Fig. 1B. Paragraph [0070]-WANG discloses the content tagging network includes a residual net (ResNet) 22 configured to receive the first feature map and generate a second feature map having a scale smaller than a scale of the first feature map.) by using a first fully connected layer (Fig. 1B. Paragraph [0079]-WANG discloses a first fully connected layer 27 configured to receive the fourth feature map and generate the second predicted probability of the content tag of the input image.), a second fully connected layer (Fig. 1B. Paragraph [0104]-WANG discloses after the sixth feature map is generated, the sixth feature map is inputted into the second fully connected layer 33, and the second fully connected layer 33 has a plurality of nodes, each node is a binary classifier. For example, the second fully connected layer 33 has 2048 nodes, and each of the 2048 nodes is a binary classifier.), and a third fully connected layer (Fig. 1B, #45 called third fully connected layer. Paragraph [0117]-WANG discloses the third fully connected layer 45 is configured to receive the ninth feature map and generate the predicted probability of the type tag of the input image.), respectively, to obtain a first encoded feature (Fig. 1B. Paragraph [0084]-WANG discloses the first fully connected layer 27 receives the fourth feature map having the size of 1*1*2048 and predicts probability of the content tag based on the fourth feature map to generate the second predicted probability of the content tag (wherein the predicted probability of the content tag is the first encoded feature).), a second encoded feature (Fig. 1B. Paragraph [0103]-WANG discloses the theme tagging network 3 includes the second fully connected layer 33 configured to receive the sixth feature map and generate the predicted probability of the theme tag of the input image (wherein the predicted probability of theme tag is the second encoded feature).), and a third encoded feature (Fig. 1B, #45 called third fully connected layer. Paragraph [0117]-WANG discloses the third fully connected layer 45 is configured to receive the ninth feature map and generate the predicted probability of the type tag of the input image (wherein the predicted probability of type tag is a third encoded feature).); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA and further in view of HAN of a non-transitory computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program, when executed by a computing device, causes the computing device to perform an image super-resolution method comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of WANG of encoding the flattened feature by using a first fully connected layer, a second fully connected layer, and a third fully connected layer, respectively, to obtain a first encoded feature, a second encoded feature, and a third encoded feature. Wherein having XU’s method of image super-resolution having encoding the flattened feature by using a first fully connected layer, a second fully connected layer, and a third fully connected layer, respectively, to obtain a first encoded feature, a second encoded feature, and a third encoded feature. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and WANG employ networks with fully connected layers to process input images, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while WANG enables the convolutional layer to add or extract a feature having a scale different from that of the input image. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and WANG (US 20210295089 A1), Paragraph [0052]. Regarding claim 16, XU in view of SIDIYA and further in view of HAN and further in view of WANG explicitly teach the non-transitory computer-readable storage medium according to claim 15, XU further explicitly teaches wherein the obtaining the channel attention matrix based on the first encoded feature and the second encoded feature comprises (Fig. 5. Paragraph [0045]-XU discloses the CSLAB 500 can first calculate an inner product 501 of the query q and the key k, and adds the result with the relative positional bias B to derive an attention matrix via element-wise addition 502 (wherein query q and the key k are first and second encoded features).): XU fails to explicitly teach performing transposition on the second encoded feature to obtain a fourth encoded feature; and obtaining the channel attention matrix based on the first encoded feature, the fourth encoded feature, and a normalization exponential function. However, SIDIYA explicitly teaches performing transposition on the second encoded feature to obtain a fourth encoded feature (Fig. 4. Paragraph [0062]-SIDIYA discloses the key features K is transposed using a transpose layer 403, the query features Q and the transpose of key features K are multiplied (wherein key features K is the second encoded feature and the fourth encoded feature is the transpose of key features k).); and obtaining the channel attention matrix based on the first encoded feature, the fourth encoded feature, and a normalization exponential function (Fig. 4. Paragraph [0062]-SIDIYA discloses the key features K is transposed using a transpose layer 403, the query features Q and the transpose of key features K are multiplied, and an attention map is obtained through a softmax layer 404 (wherein the attention map is the channel attention matrix, the first encoded feature is query features Q, the fourth encoded feature is the transpose of key features K, and the normalization exponential function is a softmax layer).). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU of a non-transitory computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program, when executed by a computing device, causes the computing device to perform an image super-resolution method comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of SIDIYA of performing transposition on the second encoded feature to obtain a fourth encoded feature; and obtaining the channel attention matrix based on the first encoded feature, the fourth encoded feature, and a normalization exponential function. Wherein having XU’s method of image super-resolution having performing transposition on the second encoded feature to obtain a fourth encoded feature; and obtaining the channel attention matrix based on the first encoded feature, the fourth encoded feature, and a normalization exponential function. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and SIDIYA employ transformers in order to generate a super-resolution image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while SIDIYA divides feature maps to patch block so as to reduce the computational cost without losing results performance. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and SIDIYA et al. (US 20230169626 A1), Paragraph [0062]. Regarding claim 23, XU in view of SIDIYA explicitly teach the electronic device according to claim 11, XU further explicitly teaches wherein the separately recalibrating the plurality of first feature blocks (Fig. 4-5. Paragraph [0043]-XU discloses the LIT 400 can first project an extracted feature embedding custom-character by separating it into four convolutional layers 406-409 and then performing upsampling operations to produce four latent embeddings, corresponding to query q, key k, value v, and frequency f (wherein the latent embeddings are first feature blocks).) based on the channel self-attention mechanism (Fig. 5. Paragraph [0045]-XU discloses FIG. 5 shows a cross-scale local attention block (CSLAB) 500 according to embodiments of the disclosure. The LIT exploits the CSLAB 500 to perform a local attention mechanism over a local grid to generate a local latent embedding custom-character for each HR coordinate.), to obtain the second feature block corresponding to each first feature block comprises (Fig. 5. Paragraph [0045]-XU discloses the CSLAB 500 can first calculate an inner product 501 of the query q and the key k, and adds the result with the relative positional bias B to derive an attention matrix via element-wise addition 502. The relative positional bias B can be produced by feeding a set of local relative coordinates δx into a positional encoding block 505 and then a fully-connected layer 506. The attention matrix can then be normalized by a softmax operation σ 503 to generate a local attention map (wherein the recalibration is normalizing by the softmax operation and wherein the second feature block is the local attention map).): obtaining a channel attention matrix based on the first encoded feature and the second encoded feature (Fig. 5. Paragraph [0045]-XU discloses the CSLAB 500 can first calculate an inner product 501 of the query q and the key k, and adds the result with the relative positional bias B to derive an attention matrix via element-wise addition 502 (wherein query q and the key k are first and second encoded features).); XU fails to explicitly teach recalibrating the third encoded feature based on the channel attention matrix, to obtain a recalibrated feature; and unflattening the recalibrated feature, to obtain the second feature block corresponding to the first feature block. However, SIDIYA explicitly teaches recalibrating the third encoded feature based on the channel attention matrix, to obtain a recalibrated feature (Fig. 4. Paragraph [0062]-SIDIYA discloses the attention map is multiplied by the value features V and the output is merged using an inverse of the patch division operation through a patch merge layer 405 (wherein the attention map is the channel attention matrix, the third encoded feature is value features V, and the recalibrated feature is the output).); and unflattening the recalibrated feature (Fig. 3. Paragraph [0062]-SIDIYA discloses the attention map is multiplied by the value features V and the output is merged using an inverse of the patch division operation through a patch merge layer 405 and a final convolution is applied using a convolution layer 406 to generate the output of the self-attention layer 301. The patch division layer 402 divides feature maps to patch block so as to reduce the computational cost without losing results performance (wherein unflattening is patch merging using the inverse of patch division).), to obtain the second feature block corresponding to the first feature block (Fig. 3. Paragraph [0062]-SIDIYA discloses the attention map is multiplied by the value features V and the output is merged using an inverse of the patch division operation through a patch merge layer 405 and a final convolution is applied using a convolution layer 406 to generate the output of the self-attention layer 301. The patch division layer 402 divides feature maps to patch block so as to reduce the computational cost without losing results performance (wherein the second feature block is the feature output following patch merge 405).). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU of an electronic device, comprising: a memory and a processor, wherein the memory is configured to store a computer program, and the processor is configured to, when calling the computer program, cause the electronic device to perform an image super-resolution method comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of SIDIYA of recalibrating the third encoded feature based on the channel attention matrix, to obtain a recalibrated feature; and unflattening the recalibrated feature, to obtain the second feature block corresponding to the first feature block. Wherein having XU’s method of image super-resolution having recalibrating the third encoded feature based on the channel attention matrix, to obtain a recalibrated feature; and unflattening the recalibrated feature, to obtain the second feature block corresponding to the first feature block. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and SIDIYA employ transformers in order to generate a super-resolution image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while SIDIYA divides feature maps to patch block so as to reduce the computational cost without losing results performance. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and SIDIYA et al. (US 20230169626 A1), Paragraph [0062]. XU in view of SIDIYA fails to explicitly teach flattening the first feature block into a two-dimensional feature to obtain a flattened feature. However, HAN explicitly teaches flattening the first feature block into a two-dimensional feature to obtain a flattened feature (Fig. 2. Paragraph [0048]-HAN discloses the electronic device may obtain an initial image feature matrix F.sub.im∈R.sup.M×C of the face image by flattening and cascading the obtained feature maps F1′, F2′, F3′, and F4′ of four levels. In this case, F.sub.im may be a feature of an ith row and an mth column of an initial image feature matrix, R denotes a set of real numbers representing the matrix F.sub.im, M may be the number of rows of the initial image feature matrix, and C may be the number of columns of the initial image feature matrix (wherein the two-dimensional feature is a matrix of M rows and C columns and the feature matrix is the flattened feature).); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA of an image super-resolution method, comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of HAN of flattening the first feature block into a two-dimensional feature to obtain a flattened feature. Wherein having XU’s method of image super-resolution having flattening the first feature block into a two-dimensional feature to obtain a flattened feature. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and HAN employ transformers in order to generate a super-resolution image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while HAN one or more embodiments may include a method and electronic device with image processing, which is able to aggregate multi-level image features, using an FSR model based on a deformable attention mechanism, to explore facial structure intrinsically and improve super-resolution results, without requesting additional facial prior annotations. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and HAN et al. (US 20240161237 A1), Paragraph [0040]. XU in view of SIDIYA and further in view of HAN fail to explicitly teach encoding the flattened feature by using a first fully connected layer, a second fully connected layer, and a third fully connected layer, respectively, to obtain a first encoded feature, a second encoded feature, and a third encoded feature. However, WANG explicitly teaches encoding the flattened feature (Fig. 1B. Paragraph [0070]-WANG discloses the content tagging network includes a residual net (ResNet) 22 configured to receive the first feature map and generate a second feature map having a scale smaller than a scale of the first feature map.) by using a first fully connected layer (Fig. 1B. Paragraph [0079]-WANG discloses a first fully connected layer 27 configured to receive the fourth feature map and generate the second predicted probability of the content tag of the input image.), a second fully connected layer (Fig. 1B. Paragraph [0104]-WANG discloses after the sixth feature map is generated, the sixth feature map is inputted into the second fully connected layer 33, and the second fully connected layer 33 has a plurality of nodes, each node is a binary classifier. For example, the second fully connected layer 33 has 2048 nodes, and each of the 2048 nodes is a binary classifier.), and a third fully connected layer (Fig. 1B, #45 called third fully connected layer. Paragraph [0117]-WANG discloses the third fully connected layer 45 is configured to receive the ninth feature map and generate the predicted probability of the type tag of the input image.), respectively, to obtain a first encoded feature (Fig. 1B. Paragraph [0084]-WANG discloses the first fully connected layer 27 receives the fourth feature map having the size of 1*1*2048 and predicts probability of the content tag based on the fourth feature map to generate the second predicted probability of the content tag (wherein the predicted probability of the content tag is the first encoded feature).), a second encoded feature (Fig. 1B. Paragraph [0103]-WANG discloses the theme tagging network 3 includes the second fully connected layer 33 configured to receive the sixth feature map and generate the predicted probability of the theme tag of the input image (wherein the predicted probability of theme tag is the second encoded feature).), and a third encoded feature (Fig. 1B, #45 called third fully connected layer. Paragraph [0117]-WANG discloses the third fully connected layer 45 is configured to receive the ninth feature map and generate the predicted probability of the type tag of the input image (wherein the predicted probability of type tag is a third encoded feature).); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA and further in view of HAN of an image super-resolution method, comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of WANG of encoding the flattened feature by using a first fully connected layer, a second fully connected layer, and a third fully connected layer, respectively, to obtain a first encoded feature, a second encoded feature, and a third encoded feature. Wherein having XU’s method of image super-resolution having encoding the flattened feature by using a first fully connected layer, a second fully connected layer, and a third fully connected layer, respectively, to obtain a first encoded feature, a second encoded feature, and a third encoded feature. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and WANG employ networks with fully connected layers to process input images, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while WANG enables the convolutional layer to add or extract a feature having a scale different from that of the input image. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and WANG (US 20210295089 A1), Paragraph [0052]. Claims 4 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over XU et al. (US 20240070809 A1), hereinafter referenced as XU, in view of SIDIYA et al. (US 20230169626 A1), hereinafter referenced as SIDIYA, and further in view of HAN et al. (US 20240161237 A1), hereinafter referenced as HAN, and further in view of WANG (US 20210295089 A1), hereinafter referenced as WANG, and further in view of LIAN et al. (US 20220391636 A1), hereinafter referenced as LIAN. Regarding claim 4, XU in view of SIDIYA and further in view of HAN and further in view of WANG explicitly teach the method according to claim 2, XU in view of SIDIYA and further in view of HAN and further in view of WANG fail to explicitly teach wherein convolution kernel sizes of the first fully connected layer, the second fully connected layer, and the third fully connected layer are all different. However, LIAN explicitly teaches wherein convolution kernel sizes of the first fully connected layer, the second fully connected layer, and the third fully connected layer are all different (Fig. 4. Paragraph [0039]-LIAN discloses the three depth-wise convolutional layers, C.sub.1 424, C.sub.2 422, C.sub.3 426∈custom-character.sup.rc with different kernel sizes of 3×3, 5×5, 7×7, are imposed on the three parts of the expended feature respectfully.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA and further in view of HAN and further in view of WANG of an image super-resolution method, comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of LIAN of wherein convolution kernel sizes of the first fully connected layer, the second fully connected layer, and the third fully connected layer are all different. Wherein having XU’s method of image super-resolution having wherein convolution kernel sizes of the first fully connected layer, the second fully connected layer, and the third fully connected layer are all different. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and LIAN employ transformers to process image features, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while LIAN a transformer that is highly efficient and can be easily combined with convolutional networks for image and computer vision tasks is described. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and LIAN et al. (US 20220391636 A1), Paragraph [0025]. Regarding claim 17, XU in view of SIDIYA and further in view of HAN and further in view of WANG explicitly teach the non-transitory computer-readable storage medium according to claim 15, XU in view of SIDIYA and further in view of HAN and further in view of WANG fail to explicitly teach wherein convolution kernel sizes of the first fully connected layer, the second fully connected layer, and the third fully connected layer are all different However, LIAN explicitly teaches wherein convolution kernel sizes of the first fully connected layer, the second fully connected layer, and the third fully connected layer are all different (Fig. 4. Paragraph [0039]-LIAN discloses the three depth-wise convolutional layers, C.sub.1 424, C.sub.2 422, C.sub.3 426∈custom-character.sup.rc with different kernel sizes of 3×3, 5×5, 7×7, are imposed on the three parts of the expended feature respectfully.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA and further in view of HAN and further in view of WANG of a non-transitory computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program, when executed by a computing device, causes the computing device to perform an image super-resolution method comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of LIAN of wherein convolution kernel sizes of the first fully connected layer, the second fully connected layer, and the third fully connected layer are all different. Wherein having XU’s method of image super-resolution having wherein convolution kernel sizes of the first fully connected layer, the second fully connected layer, and the third fully connected layer are all different. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and LIAN employ transformers to process image features, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while LIAN a transformer that is highly efficient and can be easily combined with convolutional networks for image and computer vision tasks is described. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and LIAN et al. (US 20220391636 A1), Paragraph [0025]. Claims 5-7 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over XU et al. (US 20240070809 A1), hereinafter referenced as XU, in view of SIDIYA et al. (US 20230169626 A1), hereinafter referenced as SIDIYA, and further in view of ZAMIR et al. (doi: 10.1109/CVPR52688.2022.00564.), hereinafter referenced as ZAMIR. Regarding claim 5, XU in view of SIDIYA explicitly teach the method according to claim 1, XU fails to explicitly teach wherein the obtaining the output feature of the local channel self-attention layer based on the combined feature comprises. However, SIDIYA explicitly teaches wherein the obtaining the output feature of the local channel self-attention layer based on the combined feature comprises (Fig. 4. Paragraph [0062]-SIDIYA discloses the attention map is multiplied by the value features V and the output is merged using an inverse of the patch division operation through a patch merge layer 405 and a final convolution is applied using a convolution layer 406 to generate the output of the self-attention layer 301 (wherein the output feature is the output of the self-attention layer).): Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA of an image super-resolution method, comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of SIDIYA of wherein the obtaining the output feature of the local channel self-attention layer based on the combined feature comprises. Wherein having XU’s method of image super-resolution having wherein the obtaining the output feature of the local channel self-attention layer based on the combined feature comprises. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and SIDIYA employ transformers in order to generate a super-resolution image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while SIDIYA divides feature maps to patch block so as to reduce the computational cost without losing results performance. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and SIDIYA et al. (US 20230169626 A1), Paragraph [0062]. XU in view of SIDIYA fail to explicitly teach processing the combined feature by using a feedforward network to obtain a feedforward feature; and obtaining the output feature based on the feedforward feature. However, ZAMIR explicitly teaches processing the combined feature by using a feedforward network to obtain a feedforward feature (Fig. 2, illustrates using a feedforward network to obtain a feedforward feature. Page 2, Col. 1, Line [14-17]-ZAMIR discloses a new gated-Dconv feed-forward network (GDFN) that performs controlled feature transformation, i.e., suppressing less informative features, and allowing only the useful information to pass further through the network hierarchy (wherein the feedforward feature is the output of features following feature transformation). Further on Page 4, Col. 1, Line [35-40]-ZAMIR discloses to transform features, the regular feed-forward network (FN) [17,77] operates on each pixel location separately and identically. It uses two 1×1 convolutions, one to expand the feature channels (usually by factor γ=4) and second to reduce channels back to the original input dimension. A non-linearity is applied in the hidden layer. Please see annotated Fig. 2 below.); and obtaining the output feature based on the feedforward feature (Fig. 2, illustrates obtaining the output feature based on the feedforward feature. Page 3, Col. 1, Line [30-33] and Col. 2, Line [1-6]-ZAMIR discloses these shallow features F0 pass through a 4-level symmetric encoder-decoder and transformed into deep features F_d∈R^H×W×2C. Each level of encoder-decoder contains multiple Transformer blocks, where the number of blocks are gradually increased from the top to bottom levels to maintain efficiency. Starting from the high-resolution input, the encoder hierarchically reduces spatial size, while expanding channel capacity. The decoder takes low-resolution latent features F_l ∈ R^H/8×W/8×8C as input and progressively recovers the high-resolution representations (wherein the deep feature is the output feature, as the deep feature is obtained following the feedforward features being concatenated with shallow features and output through another transformer block). Please see annotated Fig. 2 below.) PNG media_image1.png 598 1556 media_image1.png Greyscale Annotated diagram of ZAMIR’s Fig. 2 illustrating the feedforward network, the feedforward feature, and the output feature obtained based on the feedforward feature. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA of an image super-resolution method, comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of ZAMIR of processing the combined feature by using a feedforward network to obtain a feedforward feature; and obtaining the output feature based on the feedforward feature. Wherein having XU’s method of image super-resolution having processing the combined feature by using a feedforward network to obtain a feedforward feature; and obtaining the output feature based on the feedforward feature. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and ZAMIR employ transformers in order to increase the resolution of an image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while ZAMIR Table 7-10 show that our contributions yield quality performance improvements and introduce key designs to the core components of the Transformer block for improved feature aggregation and transformation. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and ZAMIR et al. (doi: 10.1109/CVPR52688.2022.00564.), Page 8, Section 4.5. Ablation Studies. Regarding claim 6, XU in view of SIDIYA explicitly teach the method according to claim 1, XU further explicitly teaches wherein generating, based on the second image feature (Fig. 8. Paragraph [0053]-XU discloses LIT.sup.i can estimate the residual image I.sub.r.sup.i from the feature embedding custom-character.sup.i with the coordinate and the corresponding cell cell.sup.i 811-813. Further in paragraph [0053]-XU discloses the final residual image 806 can be produced by adding each individual residual image I.sub.r.sup.i of LIT.sup.i via element-wise addition 814-815 (wherein the residual image is the second image feature).) and the to-be-super-resolved image (Fig. 2. Paragraph [0041]-XU discloses a bilinearly upsampled image 206 can be produced by sending the LR image 201 through a bilinear upsampling operation (e.g., a bilinear upsampling operator) 211 (wherein the to-be-super-resolved image is the LR image #201).), the super-resolution image corresponding to the to-be-super-resolved image comprises (Fig. 2, #207 called high-resolution image. Paragraph [0041]-XU discloses the residual image 205 can be combined with the bilinearly upsampled image 206 via element-wise addition to derive a HR image 207 (wherein a HR image is a super resolution image).): XU in view of SIDIYA fail to explicitly teach upsampling the second image feature to obtain an upsampled feature; and generating, based on the upsampled feature and the to-be-super-resolved image, the super-resolution image corresponding to the to-be-super-resolved image. However, ZAMIR explicitly teaches upsampling the second image feature to obtain an upsampled feature (Fig. 2, illustrates upsampling of image features. Page 3, Col. 2, Line [4-8]-ZAMIR discloses the decoder takes low-resolution latent features F_l ∈ R^H/8×W/8×8C as input and progressively recovers the high-resolution representations. For feature downsampling and upsampling, we apply pixel-unshuffle and pixel-shuffle operations [69], respectively. Please see annotated Fig. 2 below.); and generating, based on the upsampled feature (Fig. 2, illustrates upsampling of image features. Page 3, Col. 2, Line [4-8]-ZAMIR discloses the decoder takes low-resolution latent features F_l ∈ R^H/8×W/8×8C as input and progressively recovers the high-resolution representations. For feature downsampling and upsampling, we apply pixel-unshuffle and pixel-shuffle operations [69], respectively.) and the to-be-super-resolved image (Fig. 2. Page 3, Col. 1, Line [26-29]-ZAMIR discloses given a degraded image I ∈R^H×W×3, Restormer first applies a convolution to obtain low-level feature embeddings F0∈RH×W×C; where H×W denotes the spatial dimension and C is the number of channels (where a degraded image is the to-be-super-resolved image).), the super-resolution image corresponding to the to-be-super-resolved image (Fig. 2. Page 3, Col. 2, Line [20-23]-ZAMIR discloses a convolution layer is applied to the refined features to generate residual image R∈R^H×W×3 to which degraded image is added to obtain the restored image: I ^ =I+R (wherein restored I ^   is the super-resolution image). Please see annotated Fig. 2 below.). PNG media_image2.png 598 1556 media_image2.png Greyscale Annotated diagram of ZAMIR’s Fig. 2 illustrating the second image features, upsampling second image features, generating the super-resolution image based on the upsampled image and the to-be-super-resolved image. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA of an image super-resolution method, comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of ZAMIR of upsampling the second image feature to obtain an upsampled feature; and generating, based on the upsampled feature and the to-be-super-resolved image, the super-resolution image corresponding to the to-be-super-resolved image. Wherein having XU’s method of image super-resolution having upsampling the second image feature to obtain an upsampled feature; and generating, based on the upsampled feature and the to-be-super-resolved image, the super-resolution image corresponding to the to-be-super-resolved image. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and ZAMIR employ transformers in order to increase the resolution of an image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while ZAMIR Table 7-10 show that our contributions yield quality performance improvements and introduce key designs to the core components of the Transformer block for improved feature aggregation and transformation. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and ZAMIR et al. (doi: 10.1109/CVPR52688.2022.00564.), Page 8, Section 4.5. Ablation Studies. Regarding claim 7, XU in view of SIDIYA and further in view of ZAMIR explicitly teach the method according to claim 6, XU in view of SIDIYA fail to explicitly teach wherein the upsampling the second image feature comprises: upsampling the second image feature in a pixel shuffle upsampling manner. However, ZAMIR explicitly teaches wherein the upsampling the second image feature comprises (Fig. 2, illustrates upsampling of image features. Page 3, Col. 2, Line [4-8]-ZAMIR discloses the decoder takes low-resolution latent features F_l ∈ R^H/8×W/8×8C as input and progressively recovers the high-resolution representations. For feature downsampling and upsampling, we apply pixel-unshuffle and pixel-shuffle operations [69], respectively. Please see annotated Fig. 2 above.): upsampling the second image feature in a pixel shuffle upsampling manner (Fig. 2, illustrates upsampling of image features. Page 3, Col. 2, Line [4-8]-ZAMIR discloses the decoder takes low-resolution latent features F_l ∈ R^H/8×W/8×8C as input and progressively recovers the high-resolution representations. For feature downsampling and upsampling, we apply pixel-unshuffle and pixel-shuffle operations [69], respectively.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA of an image super-resolution method, comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of ZAMIR of wherein the upsampling the second image feature comprises: upsampling the second image feature in a pixel shuffle upsampling manner. Wherein having XU’s method of image super-resolution having wherein the upsampling the second image feature comprises: upsampling the second image feature in a pixel shuffle upsampling manner. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and ZAMIR employ transformers in order to increase the resolution of an image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while ZAMIR Table 7-10 show that our contributions yield quality performance improvements and introduce key designs to the core components of the Transformer block for improved feature aggregation and transformation. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and ZAMIR et al. (doi: 10.1109/CVPR52688.2022.00564.), Page 8, Section 4.5. Ablation Studies. Regarding claim 18, XU in view of SIDIYA explicitly teach the non-transitory computer-readable storage medium according to claim 12, XU fails to explicitly teach wherein the obtaining the output feature of the local channel self-attention layer based on the combined feature comprises: However, SIDIYA explicitly teaches wherein the obtaining the output feature of the local channel self-attention layer based on the combined feature comprises (Fig. 4. Paragraph [0062]-SIDIYA discloses the attention map is multiplied by the value features V and the output is merged using an inverse of the patch division operation through a patch merge layer 405 and a final convolution is applied using a convolution layer 406 to generate the output of the self-attention layer 301 (wherein the output feature is the output of the self-attention layer).): Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU of a non-transitory computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program, when executed by a computing device, causes the computing device to perform an image super-resolution method comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of SIDIYA of wherein the obtaining the output feature of the local channel self-attention layer based on the combined feature comprises. Wherein having XU’s method of image super-resolution having wherein the obtaining the output feature of the local channel self-attention layer based on the combined feature comprises. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and SIDIYA employ transformers in order to generate a super-resolution image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while SIDIYA divides feature maps to patch block so as to reduce the computational cost without losing results performance. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and SIDIYA et al. (US 20230169626 A1), Paragraph [0062]. XU in view of SIDIYA fail to explicitly teach processing the combined feature by using a feedforward network to obtain a feedforward feature; and obtaining the output feature based on the feedforward feature. However, ZAMIR explicitly teaches processing the combined feature by using a feedforward network to obtain a feedforward feature (Fig. 2, illustrates using a feedforward network to obtain a feedforward feature. Page 2, Col. 1, Line [14-17]-ZAMIR discloses a new gated-Dconv feed-forward network (GDFN) that performs controlled feature transformation, i.e., suppressing less informative features, and allowing only the useful information to pass further through the network hierarchy (wherein the feedforward feature is the output of features following feature transformation). Further on Page 4, Col. 1, Line [35-40]-ZAMIR discloses to transform features, the regular feed-forward network (FN) [17,77] operates on each pixel location separately and identically. It uses two 1×1 convolutions, one to expand the feature channels (usually by factor γ=4) and second to reduce channels back to the original input dimension. A non-linearity is applied in the hidden layer. Please see annotated Fig. 2 below.); and obtaining the output feature based on the feedforward feature (Fig. 2, illustrates obtaining the output feature based on the feedforward feature. Page 3, Col. 1, Line [30-33] and Col. 2, Line [1-6]-ZAMIR discloses these shallow features F0 pass through a 4-level symmetric encoder-decoder and transformed into deep features F_d∈R^H×W×2C. Each level of encoder-decoder contains multiple Transformer blocks, where the number of blocks are gradually increased from the top to bottom levels to maintain efficiency. Starting from the high-resolution input, the encoder hierarchically reduces spatial size, while expanding channel capacity. The decoder takes low-resolution latent features F_l ∈ R^H/8×W/8×8C as input and progressively recovers the high-resolution representations (wherein the deep feature is the output feature, as the deep feature is obtained following the feedforward features being concatenated with shallow features and output through another transformer block). Please see annotated Fig. 2 below.) PNG media_image1.png 598 1556 media_image1.png Greyscale Annotated diagram of ZAMIR’s Fig. 2 illustrating the feedforward network, the feedforward feature, and the output feature obtained based on the feedforward feature. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA of a non-transitory computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program, when executed by a computing device, causes the computing device to perform an image super-resolution method comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of ZAMIR of processing the combined feature by using a feedforward network to obtain a feedforward feature; and obtaining the output feature based on the feedforward feature. Wherein having XU’s method of image super-resolution having processing the combined feature by using a feedforward network to obtain a feedforward feature; and obtaining the output feature based on the feedforward feature. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and ZAMIR employ transformers in order to increase the resolution of an image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while ZAMIR Table 7-10 show that our contributions yield quality performance improvements and introduce key designs to the core components of the Transformer block for improved feature aggregation and transformation. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and ZAMIR et al. (doi: 10.1109/CVPR52688.2022.00564.), Page 8, Section 4.5. Ablation Studies. Regarding claim 19, XU in view of SIDIYA explicitly teach the non-transitory computer-readable storage medium according to claim 12, XU further explicitly teaches wherein generating, based on the second image feature (Fig. 8. Paragraph [0053]-XU discloses LIT.sup.i can estimate the residual image I.sub.r.sup.i from the feature embedding custom-character.sup.i with the coordinate and the corresponding cell cell.sup.i 811-813. Further in paragraph [0053]-XU discloses the final residual image 806 can be produced by adding each individual residual image I.sub.r.sup.i of LIT.sup.i via element-wise addition 814-815 (wherein the residual image is the second image feature).) and the to-be-super-resolved image (Fig. 2. Paragraph [0041]-XU discloses a bilinearly upsampled image 206 can be produced by sending the LR image 201 through a bilinear upsampling operation (e.g., a bilinear upsampling operator) 211 (wherein the to-be-super-resolved image is the LR image #201).), the super-resolution image corresponding to the to-be-super-resolved image comprises (Fig. 2, #207 called high-resolution image. Paragraph [0041]-XU discloses the residual image 205 can be combined with the bilinearly upsampled image 206 via element-wise addition to derive a HR image 207 (wherein a HR image is a super resolution image).): XU in view of SIDIYA fail to explicitly teach upsampling the second image feature to obtain an upsampled feature; and generating, based on the upsampled feature and the to-be-super-resolved image, the super-resolution image corresponding to the to-be-super-resolved image. However, ZAMIR explicitly teaches upsampling the second image feature to obtain an upsampled feature (Fig. 2, illustrates upsampling of image features. Page 3, Col. 2, Line [4-8]-ZAMIR discloses the decoder takes low-resolution latent features F_l ∈ R^H/8×W/8×8C as input and progressively recovers the high-resolution representations. For feature downsampling and upsampling, we apply pixel-unshuffle and pixel-shuffle operations [69], respectively. Please see annotated Fig. 2 below.); and generating, based on the upsampled feature (Fig. 2, illustrates upsampling of image features. Page 3, Col. 2, Line [4-8]-ZAMIR discloses the decoder takes low-resolution latent features F_l ∈ R^H/8×W/8×8C as input and progressively recovers the high-resolution representations. For feature downsampling and upsampling, we apply pixel-unshuffle and pixel-shuffle operations [69], respectively.) and the to-be-super-resolved image (Fig. 2. Page 3, Col. 1, Line [26-29]-ZAMIR discloses given a degraded image I ∈R^H×W×3, Restormer first applies a convolution to obtain low-level feature embeddings F0∈RH×W×C; where H×W denotes the spatial dimension and C is the number of channels (where a degraded image is the to-be-super-resolved image).), the super-resolution image corresponding to the to-be-super-resolved image (Fig. 2. Page 3, Col. 2, Line [20-23]-ZAMIR discloses a convolution layer is applied to the refined features to generate residual image R∈R^H×W×3 to which degraded image is added to obtain the restored image: I ^ =I+R (wherein restored I ^   is the super-resolution image). Please see annotated Fig. 2 below.). PNG media_image2.png 598 1556 media_image2.png Greyscale Annotated diagram of ZAMIR’s Fig. 2 illustrating the second image features, upsampling second image features, generating the super-resolution image based on the upsampled image and the to-be-super-resolved image. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA of a non-transitory computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program, when executed by a computing device, causes the computing device to perform an image super-resolution method comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of ZAMIR of upsampling the second image feature to obtain an upsampled feature; and generating, based on the upsampled feature and the to-be-super-resolved image, the super-resolution image corresponding to the to-be-super-resolved image. Wherein having XU’s method of image super-resolution having upsampling the second image feature to obtain an upsampled feature; and generating, based on the upsampled feature and the to-be-super-resolved image, the super-resolution image corresponding to the to-be-super-resolved image. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and ZAMIR employ transformers in order to increase the resolution of an image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while ZAMIR Table 7-10 show that our contributions yield quality performance improvements and introduce key designs to the core components of the Transformer block for improved feature aggregation and transformation. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and ZAMIR et al. (doi: 10.1109/CVPR52688.2022.00564.), Page 8, Section 4.5. Ablation Studies. Regarding claim 20, XU in view of SIDIYA and further in view of ZAMIR explicitly teach the non-transitory computer-readable storage medium according to claim 19, XU in view of SIDIYA fail to explicitly teach wherein the upsampling the second image feature comprises: upsampling the second image feature in a pixel shuffle upsampling manner. However, ZAMIR explicitly teaches wherein the upsampling the second image feature comprises (Fig. 2, illustrates upsampling of image features. Page 3, Col. 2, Line [4-8]-ZAMIR discloses the decoder takes low-resolution latent features F_l ∈ R^H/8×W/8×8C as input and progressively recovers the high-resolution representations. For feature downsampling and upsampling, we apply pixel-unshuffle and pixel-shuffle operations [69], respectively. Please see annotated Fig. 2 above.): upsampling the second image feature in a pixel shuffle upsampling manner (Fig. 2, illustrates upsampling of image features. Page 3, Col. 2, Line [4-8]-ZAMIR discloses the decoder takes low-resolution latent features F_l ∈ R^H/8×W/8×8C as input and progressively recovers the high-resolution representations. For feature downsampling and upsampling, we apply pixel-unshuffle and pixel-shuffle operations [69], respectively.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA of a non-transitory computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program, when executed by a computing device, causes the computing device to perform an image super-resolution method comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of ZAMIR of wherein the upsampling the second image feature comprises: upsampling the second image feature in a pixel shuffle upsampling manner. Wherein having XU’s method of image super-resolution having wherein the upsampling the second image feature comprises: upsampling the second image feature in a pixel shuffle upsampling manner. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and ZAMIR employ transformers in order to increase the resolution of an image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while ZAMIR Table 7-10 show that our contributions yield quality performance improvements and introduce key designs to the core components of the Transformer block for improved feature aggregation and transformation. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and ZAMIR et al. (doi: 10.1109/CVPR52688.2022.00564.), Page 8, Section 4.5. Ablation Studies. Claims 8 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over XU et al. (US 20240070809 A1), hereinafter referenced as XU, in view of SIDIYA et al. (US 20230169626 A1), hereinafter referenced as SIDIYA, and further in view of ZAMIR et al. (doi: 10.1109/CVPR52688.2022.00564.), hereinafter referenced as ZAMIR, and further in view of ALFARRAJ et al. (US 20240161235 A1), hereinafter referenced as ALFARRAJ. Regarding claim 8, XU in view of SIDIYA and further in view of ZAMIR explicitly teach the method according to claim 6, XU in view of SIDIYA fail to explicitly teach wherein the generating, based on the upsampled feature and the to-be-super-resolved image, the super-resolution image corresponding to the to-be-super-resolved image comprises. However, ZAMIR explicitly teaches wherein the generating, based on the upsampled feature (Fig. 2, illustrates upsampling of image features. Page 3, Col. 2, Line [4-8]-ZAMIR discloses the decoder takes low-resolution latent features F_l ∈ R^H/8×W/8×8C as input and progressively recovers the high-resolution representations. For feature downsampling and upsampling, we apply pixel-unshuffle and pixel-shuffle operations [69], respectively.) and the to-be-super-resolved image (Fig. 2. Page 3, Col. 1, Line [26-29]-ZAMIR discloses given a degraded image I ∈R^H×W×3, Restormer first applies a convolution to obtain low-level feature embeddings F0∈RH×W×C; where H×W denotes the spatial dimension and C is the number of channels (where a degraded image is the to-be-super-resolved image).), the super-resolution image corresponding to the to-be-super-resolved image comprises (Fig. 2. Page 3, Col. 2, Line [20-23]-ZAMIR discloses a convolution layer is applied to the refined features to generate residual image R∈R^H×W×3 to which degraded image is added to obtain the restored image: I ^ =I+R (wherein restored I ^   is the super-resolution image). Please see annotated Fig. 2 above.): Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA of an image super-resolution method, comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of ZAMIR of wherein the generating, based on the upsampled feature and the to-be-super-resolved image, the super-resolution image corresponding to the to-be-super-resolved image comprises. Wherein having XU’s method of image super-resolution having wherein the generating, based on the upsampled feature and the to-be-super-resolved image, the super-resolution image corresponding to the to-be-super-resolved image comprises. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and ZAMIR employ transformers in order to increase the resolution of an image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while ZAMIR Table 7-10 show that our contributions yield quality performance improvements and introduce key designs to the core components of the Transformer block for improved feature aggregation and transformation. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and ZAMIR et al. (doi: 10.1109/CVPR52688.2022.00564.), Page 8, Section 4.5. Ablation Studies. XU in view of SIDIYA and further in view of ZAMIR fail to explicitly teach performing linear interpolation on the to-be-super-resolved image to obtain an interpolated image; and adding and fusing the interpolated image and the upsampled feature, to obtain the super-resolution image corresponding to the to-be-super-resolved image. However, ALFARRAJ explicitly teaches performing linear interpolation on the to-be-super-resolved image to obtain an interpolated image (Fig. 2. Paragraph [0051]-ALFARRAJ discloses the skip connection 255 is configured with a bicubic interpolation function to upsample the input I.sup.LR to the desired dimension. The skip connection 255 generates an image, known as I.sup.BI using the bicubic interpolation function (wherein the input I.sup.LR is the to-be-super-resolved image and I.sup.BI is the interpolated image).); and adding and fusing the interpolated image and the upsampled feature (Fig. 2. Paragraph [0051]-ALAFARRAJ discloses the element wise addition section 250 is configured to add to the output of the late upsampling section 245 with the output of the skip connection 255 I.sup.BI to produce the output image I.sup.SR (wherein I.sup.BI is the interpolated image and the output of the late upsampling section is the upsampled feature).), to obtain the super-resolution image corresponding to the to-be-super-resolved image (Fig. 2. Paragraph [0051]-ALAFARRAJ discloses the element wise addition section 250 is configured to add to the output of the late upsampling section 245 with the output of the skip connection 255 I.sup.BI to produce the output image I.sup.SR (wherein the output image is the super-resolution image).). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA of an image super-resolution method, comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of ALFARRAJ of performing linear interpolation on the to-be-super-resolved image to obtain an interpolated image; and adding and fusing the interpolated image and the upsampled feature, to obtain the super-resolution image corresponding to the to-be-super-resolved image. Wherein having XU’s method of image super-resolution having performing linear interpolation on the to-be-super-resolved image to obtain an interpolated image; and adding and fusing the interpolated image and the upsampled feature, to obtain the super-resolution image corresponding to the to-be-super-resolved image. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and ALFARRAJ relate to generating super-resolution images, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while ALFARRAJ it is one object of the present disclosure to provide a system for real-time image super-resolution that reduces the number of parameters by using DSC layers throughout the operation and reduces MACs by adopting a late upsampling scheme. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and ALFARRAJ et al. (US 20240161235 A1), Paragraph [0007]. Regarding claim 21, XU in view of SIDIYA and further in view of ZAMIR explicitly teach the non-transitory computer-readable storage medium according to claim 19, XU in view of SIDIYA fail to explicitly teach wherein the generating, based on the upsampled feature and the to-be-super-resolved image, the super-resolution image corresponding to the to-be-super-resolved image comprises: However, ZAMIR explicitly teaches wherein the generating, based on the upsampled feature (Fig. 2, illustrates upsampling of image features. Page 3, Col. 2, Line [4-8]-ZAMIR discloses the decoder takes low-resolution latent features F_l ∈ R^H/8×W/8×8C as input and progressively recovers the high-resolution representations. For feature downsampling and upsampling, we apply pixel-unshuffle and pixel-shuffle operations [69], respectively.) and the to-be-super-resolved image (Fig. 2. Page 3, Col. 1, Line [26-29]-ZAMIR discloses given a degraded image I ∈R^H×W×3, Restormer first applies a convolution to obtain low-level feature embeddings F0∈RH×W×C; where H×W denotes the spatial dimension and C is the number of channels (where a degraded image is the to-be-super-resolved image).), the super-resolution image corresponding to the to-be-super-resolved image comprises (Fig. 2. Page 3, Col. 2, Line [20-23]-ZAMIR discloses a convolution layer is applied to the refined features to generate residual image R∈R^H×W×3 to which degraded image is added to obtain the restored image: I ^ =I+R (wherein restored I ^   is the super-resolution image). Please see annotated Fig. 2 above.): Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA of a non-transitory computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program, when executed by a computing device, causes the computing device to perform an image super-resolution method comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of ZAMIR of wherein the generating, based on the upsampled feature and the to-be-super-resolved image, the super-resolution image corresponding to the to-be-super-resolved image comprises. Wherein having XU’s method of image super-resolution having wherein the generating, based on the upsampled feature and the to-be-super-resolved image, the super-resolution image corresponding to the to-be-super-resolved image comprises. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and ZAMIR employ transformers in order to increase the resolution of an image, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while ZAMIR Table 7-10 show that our contributions yield quality performance improvements and introduce key designs to the core components of the Transformer block for improved feature aggregation and transformation. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and ZAMIR et al. (doi: 10.1109/CVPR52688.2022.00564.), Page 8, Section 4.5. Ablation Studies. XU in view of SIDIYA and further in view of ZAMIR fail to explicitly teach performing linear interpolation on the to-be-super-resolved image to obtain an interpolated image; and adding and fusing the interpolated image and the upsampled feature, to obtain the super-resolution image corresponding to the to-be-super-resolved image. However, ALFARRAJ explicitly teaches performing linear interpolation on the to-be-super-resolved image to obtain an interpolated image (Fig. 2. Paragraph [0051]-ALFARRAJ discloses the skip connection 255 is configured with a bicubic interpolation function to upsample the input I.sup.LR to the desired dimension. The skip connection 255 generates an image, known as I.sup.BI using the bicubic interpolation function (wherein the input I.sup.LR is the to-be-super-resolved image and I.sup.BI is the interpolated image).); and adding and fusing the interpolated image and the upsampled feature (Fig. 2. Paragraph [0051]-ALAFARRAJ discloses the element wise addition section 250 is configured to add to the output of the late upsampling section 245 with the output of the skip connection 255 I.sup.BI to produce the output image I.sup.SR (wherein I.sup.BI is the interpolated image and the output of the late upsampling section is the upsampled feature).), to obtain the super-resolution image corresponding to the to-be-super-resolved image (Fig. 2. Paragraph [0051]-ALAFARRAJ discloses the element wise addition section 250 is configured to add to the output of the late upsampling section 245 with the output of the skip connection 255 I.sup.BI to produce the output image I.sup.SR (wherein the output image is the super-resolution image).). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of XU in view of SIDIYA of a non-transitory computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program, when executed by a computing device, causes the computing device to perform an image super-resolution method comprising: performing feature extraction on a to-be-super-resolved image to obtain a first image feature; separately recalibrate the plurality of first feature blocks based on a channel self-attention mechanism to obtain a second feature block corresponding to each first feature block, generating, based on the second image feature and the to-be-super-resolved image, a super-resolution image corresponding to the to-be-super-resolved image with the teachings of ALFARRAJ of performing linear interpolation on the to-be-super-resolved image to obtain an interpolated image; and adding and fusing the interpolated image and the upsampled feature, to obtain the super-resolution image corresponding to the to-be-super-resolved image. Wherein having XU’s method of image super-resolution having performing linear interpolation on the to-be-super-resolved image to obtain an interpolated image; and adding and fusing the interpolated image and the upsampled feature, to obtain the super-resolution image corresponding to the to-be-super-resolved image. The motivation behind the modification would have been to obtain method of image super-resolution that improves computation performance and efficiency when carrying out the image super-resolution method. Since both XU and ALFARRAJ relate to generating super-resolution images, wherein XU it can be observed that a significant improvement in adopting the cross-scale local attention block and a relatively minor gain with the frequency encoding block, while ALFARRAJ it is one object of the present disclosure to provide a system for real-time image super-resolution that reduces the number of parameters by using DSC layers throughout the operation and reduces MACs by adopting a late upsampling scheme. Please see XU et al. (US 20240070809 A1), Fig. 14 and Paragraph [0062], and ALFARRAJ et al. (US 20240161235 A1), Paragraph [0007]. Conclusion Listed below are the prior arts made of record and not relied upon but are considered pertinent to applicant’s disclosure. DUDHANE et al. (US 20240135496 A1) - A mobile device and mobile application, in which the mobile device includes a camera having an image capture circuit operating in a mode to capture a RAW image burst, and processing circuitry, including a neural network engine, to generate a single enhanced image from the RAW image burst. The neural network engine executing program instructions including an edge boosting feature alignment stage to remove inter-frame spatial and color misalignment from the RAW image burst to obtain aligned burst frames, a pseudo-burst feature fusion stage to create a set of pseudo-burst features that combine complementary information from the aligned burst frames, and an adaptive group upsampling stage to progressively increase spatial resolution while merging the set of pseudo-burst features and output the single enhanced image. The mobile application and mobile device perform super-resolution, low-light image enhancement, and burst denoising using a RAW image burst…Abstract, Fig. 8. SHLENS et al. (US 20220215654 A1) - A system implemented as computer programs on one or more computers in one or more locations that implements a computer vision model is described. The computer vision model includes a positional local self-attention layer that is configured to receive an input feature map and to generate an output feature map. For each input element in the input feature map, the positional local self-attention layer generates a respective output element for the output feature map by generating a memory block including neighboring input elements around the input element, generates a query vector using the input element and a query weight matrix, for each neighboring element in the memory block, performs positional local self-attention operations to generate a temporary output element, and generates the respective output element by summing temporary output elements of the neighboring elements in the memory block…Abstract, Fig. 1. NASROLLAHI et al. (US 20230325974 A1) – An image processing method including acquiring a first image whose spatial resolution and lightness are to be enhanced; generating a residual image from the first image using a multi-scale hierarchical neural network for joint learning of low-light enhancement and super-resolution, the network comprising an encoder stage and a decoder stage forming a plurality of symmetrical encoder-decoder levels, each encoder and decoder in each level comprising a vision transformer block; generating a reconstructed image based on the first and residual images…Abstract, Fig. 1A-2. ARORA et al. (DOI: https://doi.org/10.48550/arXiv.2101.00850) - Images captured under low-light conditions manifest poor visibility, lack contrast and color vividness. Compared to conventional approaches, deep convolutional neural networks (CNNs) perform well in enhancing images. However, being solely reliant on confined fixed primitives to model dependencies, existing data-driven deep models do not exploit the contexts at various spatial scales to address low-light image enhancement. These contexts can be crucial towards inferring several image enhancement tasks, e.g., local and global contrast, brightness and color corrections; which requires cues from both local and global spatial extent. To this end, we introduce a context-aware deep network for low-light image enhancement. First, it features a global context module that models spatial correlations to find complementary cues over full spatial domain. Second, it introduces a dense residual block that captures local context with a relatively large receptive field. We evaluate the proposed approach using three challenging datasets: MIT-Adobe FiveK, LoL, and SID. On all these datasets, our method performs favorably against the state-of-the-arts in terms of standard image fidelity metrics. In particular, compared to the best performing method on the MIT-Adobe FiveK dataset, our algorithm improves PSNR from 23.04 dB to 24.45 dB…Abstract, Fig. 2. ZHANG et al. (DOI: https://doi.org/10.48550/arXiv.2107.00708) - Image super-resolution (SR) research has witnessed impressive progress thanks to the advance of convolutional neural networks (CNNs) in recent years. However, most existing SR methods are non-blind and assume that degradation has a single fixed and known distribution (e.g., bicubic) which struggle while handling degradation in real-world data that usually follows a multi-modal, spatially variant, and unknown distribution. The recent blind SR studies address this issue via degradation estimation, but they do not generalize well to multi-source degradation and cannot handle spatially variant degradation. We design CRL-SR, a contrastive representation learning network that focuses on blind SR of images with multi-modal and spatially variant distributions. CRL-SR addresses the blind SR challenges from two perspectives. The first is contrastive decoupling encoding which introduces contrastive learning to extract resolution-invariant embedding and discard resolution-variant embedding under the guidance of a bidirectional contrastive loss. The second is contrastive feature refinement which generates lost or corrupted high-frequency details under the guidance of a conditional contrastive loss. Extensive experiments on synthetic datasets and real images show that the proposed CRL-SR can handle multi-modal and spatially variant degradation effectively under blind settings and it also outperforms state-of-the-art SR methods qualitatively and quantitatively…Abstract, Fig. 1. DOSOVITSKIY et al. (DOI: https://doi.org/10.48550/arXiv.2010.11929) - While the Transformer architecture has become the de-facto standard for natural language processing tasks, its applications to computer vision remain limited. In vision, attention is either applied in conjunction with convolutional networks, or used to replace certain components of convolutional networks while keeping their overall structure in place. We show that this reliance on CNNs is not necessary and a pure transformer applied directly to sequences of image patches can perform very well on image classification tasks. When pre-trained on large amounts of data and transferred to multiple mid-sized or small image recognition benchmarks (ImageNet, CIFAR-100, VTAB, etc.), Vision Transformer (ViT) attains excellent results compared to state-of-the-art convolutional networks while requiring substantially fewer computational resources to train…Abstract, Fig. 1. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ETHAN N WOLFSON whose telephone number is (571)272-1898. The examiner can normally be reached Monday - Friday 8:00 am - 5:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ETHAN N WOLFSON/Examiner, Art Unit 2673 /CHINEYERE WILLS-BURNS/Supervisory Patent Examiner, Art Unit 2673
Read full office action

Prosecution Timeline

Dec 18, 2024
Application Filed
Aug 12, 2026
Non-Final Rejection mailed — §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
86%
Grant Probability
99%
With Interview (+50.0%)
2y 7m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 7 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month