Prosecution Insights
Last updated: October 02, 2026
Application No. 18/922,568

METHOD, ELECTRONIC DEVICE, AND STORAGE MEDIUM FOR VIDEO-SPECIFIC SUPER-RESOLUTION

Non-Final OA §103
Filed
Oct 22, 2024
Priority
Nov 23, 2022 — CN 202211476937.7 +1 more
Examiner
ANSARI, TAHMINA N
Art Unit
Tech Center
Assignee
Tencent Technology (Shenzhen) Company Limited
OA Round
1 (Non-Final)
85%
Grant Probability
Favorable
1-2
OA Rounds
7m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 85% — above average
85%
Career Allowance Rate
770 granted / 902 resolved
+25.4% vs TC avg
Strong +19% interview lift
Without
With
+18.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
22 currently pending
Career history
918
Total Applications
across all art units

Statute-Specific Performance

§101
12.8%
-27.2% vs TC avg
§103
42.8%
+2.8% vs TC avg
§102
21.8%
-18.2% vs TC avg
§112
10.3%
-29.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 902 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status Claims 1-20 are pending in this application. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. Specification The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made. Claims 1-5, 7-14 and 16-20 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (US PGPub US2020/0372618A1, published November 26, 2020), hereby referred to as “Zhang”, in view of Shenzen et al. (CN Publication CN112184548A, published January 5, 2021), hereby referred to as “Shenzen”. Consider Claims 1, 10 and 19. Zhang teaches: -; 1. A video-specific super-resolution method, performed by a computer device, and comprising: / 10. A computer device, comprising one or more processors and a memory storing at least one instruction, at least one program, a code set, or an instruction set, that when the at least one instruction, the at least one program, the code set, or the instruction set being executed, cause the one or more processors to perform: / 19. A non-transitory computer-readable storage medium containing at least one program that, when being executed, causes at least one processor to perform: (Zhang: abstract, A method of video deblurring by an electronic device is described. The processing circuitry of the electronic device acquires N continuous image frames from a video clip the N being a positive integer, and the N continuous image frames including a blurry image frame to be processed. The processing circuitry of the electronic device performs three-dimensional (3D) convolution processing on the N continuous image frames with a generative adversarial network model, to acquire spatio-temporal information corresponding to the blurry image frame. The spatio-temporal information includes spatial feature information of the blurry image frame, and temporal feature information between the blurry image frame and a neighboring image frame of the N continuous image frames The processing circuitry of the electronic device performs deblurring processing on the blurry image frame by using the spatio-temporal information corresponding to the blurry image frame through the generative adversarial network model, to output a sharp image frame. Figure 1, [0027]-[0028] FIG. 1 is a schematic block flowchart of a video deblurring method according to an embodiment of this application. The video deblurring method can be executed by an electronic device. The electronic device may be a terminal or server. The following is described with the electronic device executing the video deblurring method as an example. Referring to FIG. 1, the method may include the following steps: [0029] In step 101, the electronic device acquires N continuous image frames from a video clip, the N being a positive integer, and the N image frames including a blurry image frame to be processed. [0030] According to this embodiment of this application, the video clip may be shot by the terminal through a camera, or may be downloaded from the Internet by the terminal. [0031] In step 102, the electronic device performs 3D convolution processing on the N image frames with a generative adversarial network model, to acquire spatio-temporal information corresponding to the blurry image frame, the spatio-temporal information including: spatial feature information of the blurry image frame, and temporal feature information between the blurry image frame and a neighboring image frame of the N image frames. [0032] According to this embodiment of this application, a trained generative adversarial network model may be used to perform video deblurring processing. After acquiring the N continuous image frames, the N continuous image frames are inputted into the generative adversarial network model; the 3D convolution operation is performed by using a 3D convolution kernel in the generative adversarial network model; and the spatio-temporal information implicit between the continuous image frames is extracted. The spatio-temporal information includes: the spatial feature information of the blurry image frame, that is, the spatial feature information is hidden in a single frame of blurry image, and the temporal feature information is the temporal information between the blurry image frame and the neighboring image frame. For example, temporal feature information can be extracted through the 3D convolution operation, the information being between a blurry image frame and two neighboring image frames prior to the blurry image frame and two neighboring image frames behind the blurry image frame. According to this embodiment of this application, the spatio-temporal information, that is, the temporal feature information and the spatial feature information can be extracted through the 3D convolution kernel. Therefore, effective utilization of the feature information hidden between the continuous images in the video clip, in combination with the trained generative adversarial network model, can enhance the effect of deblurring processing on a blurry image frame. For details, reference is made to descriptions of video deblurring in a subsequent embodiment. [0033] In step 103, the electronic device performs deblurring processing on the blurry image frame by using the spatio-temporal information corresponding to the blurry image frame through the generative adversarial network model, to output a sharp image frame. [0034] According to this embodiment of this application, after the spatio-temporal information corresponding to the blurry image frame is extracted by performing the 3D convolution operation through the 3D convolution kernel in the generative adversarial network model, the spatio-temporal information corresponding to the blurry image frame can be used as an image feature to perform a predicted output through the generative adversarial network model. An output result of the generative adversarial network model is a sharp image frame obtained after deblurring the blurry image frame. Because the 3D convolution operation is adopted by the generative adversarial network model according to this embodiment of this application, the temporal feature information and the spatial feature information can be extracted. This type of feature information can be used to predict the sharp image frame corresponding to the blurry image frame.) -; 1. obtaining an (i+1)th frame of image from a video, and obtaining image features of an ith frame of image in the video and long time series features before the ith frame of image, the image features of the ith frame of image and the long time series features before the ith frame of image being cached during super-resolution processing of the ith frame of image; / 10. obtaining an (i+1)th frame of image from a video, and obtaining image features of an ith frame of image in the video and long time series features before the ith frame of image, the image features of the ith frame of image and the long time series features before the ith frame of image being cached during super-resolution processing of the ith frame of image; / 19. obtaining an (i+1)th frame of image from a video, and obtaining image features of an ith frame of image in the video and long time series features before the ith frame of image, the image features of the ith frame of image and the long time series features before the ith frame of image being cached during super-resolution processing of the ith frame of image; (Zhang: [0042]-[0045], Figures 1-3, [0046] Step A11 includes performing convolution processing on the N sample image frames with the first 3D convolution kernel, to acquire low-level spatio-temporal features corresponding to the blurry sample image frame. [0047] Step A12 includes performing the convolution processing on the low-level spatio-temporal features with the second 3D convolution kernel, to acquire high-level spatio-temporal features corresponding to the blurry sample image frame. [0048] Step A13 includes fusing the high-level spatio-temporal features corresponding to the blurry sample image frame, to acquire spatio-temporal information corresponding to the blurry sample image frame. [0049] At first, two 3D convolutional layers are set in the generative network model. On each 3D convolutional layer, different 3D convolution kernels may be used. For example, the first 3D convolution kernel and the second 3D convolution kernel have different weight parameters. The first 3D convolution kernel is first used to perform the convolution processing on the N sample image frames, to acquire the low-level spatio-temporal features corresponding to the blurry sample image frame. The low-level spatio-temporal features indicate inconspicuous feature information, such as features of lines. Then taking the low-level spatio-temporal features as an input condition, the convolution processing is performed on the other 3D convolutional layer, to acquire the high-level spatio-temporal features corresponding to the blurry sample image frame. The high-level spatio-temporal features indicate feature information of neighboring image frames. At last, the high-level spatio-temporal features are fused together to acquire the spatio-temporal information corresponding to the blurry sample image frame. The spatio-temporal information can be used as a feature map for training of the generative network model. The following is an illustration with an example: the first 3D convolution kernel is first used to perform the convolution processing on 5 sample image frames, to acquire 3 low-level spatio-temporal features in different dimensions; then the convolution processing is performed on the low-level spatio-temporal features by using the second 3D convolution kernel, to acquire high-level spatio-temporal features corresponding to the blurry sample image frames; and the high-level spatio-temporal features are fused to acquire spatio-temporal information corresponding to the blurry sample image frames. Because 5 frames of images are inputted into the generative network model, after 2 times of 3D convolution processing, one frame of feature map is outputted, that is, after the 2 times of 3D convolution processing, a quantity of channels of time series is changed from 5 to 1. [0050] Further, according to some embodiments of this application, the generative network model further includes M 2D convolution kernels, the M being a positive integer. Step A3 that the electronic device performs deblurring processing on the blurry sample image frame by using the spatio-temporal information corresponding to the blurry sample image frame through the generative network model, to output a sharp sample image frame includes the following steps. [0051] Step A31 includes performing convolution processing on the spatio-temporal information corresponding to the blurry sample image frame by using each 2D convolution kernel of the M 2D convolution kernels in sequence, and acquire the sharp sample image frame after the convolution processing is performed by using the last 2D convolution kernel of the M 2D convolution kernels.) Zhang does not teach: -; 1. performing super-resolution prediction on the image features of the ith frame of image, the long time series features before the ith frame of image, and the (i+1)th frame of image using a generative network, to obtain a super-resolution image of the (i+1)th frame of image, image features of the (i+1)th frame of image, and long time series features before the (i+1)th frame of image; / 10. performing super-resolution prediction on the image features of the ith frame of image, the long time series features before the ith frame of image, and the (i+1)th frame of image using a generative network, to obtain a super-resolution image of the (i+1)th frame of image, image features of the (i+1)th frame of image, and long time series features before the (i+1)th frame of image; / 19. performing super-resolution prediction on the image features of the ith frame of image, the long time series features before the ith frame of image, and the (i+1)th frame of image using a generative network, to obtain a super-resolution image of the (i+1)th frame of image, image features of the (i+1)th frame of image, and long time series features before the (i+1)th frame of image; -; 1. and caching the image features of the (i+1)th frame of image and the long time series features before the (i+1)th frame of image; i being a positive integer greater than 2. / 10. and caching the image features of the (i+1)th frame of image and the long time series features before the (i+1)th frame of image; i being a positive integer greater than 2. / 19. and caching the image features of the (i+1)th frame of image and the long time series features before the (i+1)th frame of image; i being a positive integer greater than 2. Shenzen teaches: -; 1. A video-specific super-resolution method, performed by a computer device, and comprising: / 10. A computer device, comprising one or more processors and a memory storing at least one instruction, at least one program, a code set, or an instruction set, that when the at least one instruction, the at least one program, the code set, or the instruction set being executed, cause the one or more processors to perform: / 19. A non-transitory computer-readable storage medium containing at least one program that, when being executed, causes at least one processor to perform: (Shenzen: abstract, The embodiments of the present application provide an image super-resolution method, apparatus, device, and storage medium, which are applicable to the technical field of image processing. The method includes: acquiring a first image; inputting the first image into a super-resolution network model for processing to obtain a second image, where the resolution of the second image is greater than that of the first image, and the super-resolution network model is A network model using a pixel attention mechanism, which is used to generate a corresponding weight for each pixel in the input feature map. Since the super-resolution network model adopts the pixel attention mechanism, it can make the model pay more attention to key features in the image processing process, and then can use the features more efficiently, greatly reducing the amount of parameters and parameters required by the super-resolution network model. The amount of computation increases the processing speed of the super-resolution network model. And the pixel attention mechanism is more suitable for low-level visual tasks such as super-resolution, which improves the super-resolution performance and provides users with a higher-quality visual experience.) -; 1. obtaining an (i+1)th frame of image from a video, and obtaining image features of an ith frame of image in the video and long time series features before the ith frame of image, the image features of the ith frame of image and the long time series features before the ith frame of image being cached during super-resolution processing of the ith frame of image; / 10. obtaining an (i+1)th frame of image from a video, and obtaining image features of an ith frame of image in the video and long time series features before the ith frame of image, the image features of the ith frame of image and the long time series features before the ith frame of image being cached during super-resolution processing of the ith frame of image; / 19. obtaining an (i+1)th frame of image from a video, and obtaining image features of an ith frame of image in the video and long time series features before the ith frame of image, the image features of the ith frame of image and the long time series features before the ith frame of image being cached during super-resolution processing of the ith frame of image; (Shenzen: Disclosure of Invention The embodiment of the application provides an image super-resolution method, device, equipment and storage medium, which can solve the problems that a super-resolution network model in the related art cannot effectively utilize characteristics and needs large parameter quantity and calculated quantity, improve the super-resolution performance and provide high-quality visual experience for users. In a first aspect, an embodiment of the present application provides an image super-resolution method, including: acquiring a first image, wherein the resolution of the first image is a first resolution; and inputting the first image into a super-resolution network model for processing to obtain a second image, wherein the resolution of the second image is a second resolution which is higher than the first resolution, the super-resolution network model is a network model adopting a pixel attention mechanism, and the pixel attention mechanism is used for generating a corresponding weight for each pixel in the input feature map. Optionally, the pixel attention mechanism is configured to perform convolution processing on an input feature map by specifying a convolution layer, and generate a three-dimensional attention map of the input feature map, where the three-dimensional attention map of the input feature map includes channel information of the input feature map, position information of each pixel in the input feature map, and a weight of each pixel. Optionally, the specified convolutional layer is 1 x 1 convolutional layer. Optionally, the super-resolution network model includes: the pixel attention mechanism comprises a feature extraction module, a nonlinear mapping module and a reconstruction module, wherein at least one of the nonlinear mapping module and the reconstruction module adopts the pixel attention mechanism. Optionally, the inputting the first image into a super-resolution network model for processing to obtain a second image includes: inputting the first image into the feature extraction module, and performing feature extraction on the first image through the feature extraction module to obtain a first feature map; inputting the first feature map into the nonlinear mapping module, and carrying out nonlinear mapping on the first feature map through the nonlinear mapping module to obtain a second feature map; inputting the second feature map into the reconstruction module, reconstructing the second feature map through the reconstruction module to obtain a third feature map, and generating the second image based on the third feature map.) -; 1. performing super-resolution prediction on the image features of the ith frame of image, the long time series features before the ith frame of image, and the (i+1)th frame of image using a generative network, to obtain a super-resolution image of the (i+1)th frame of image, image features of the (i+1)th frame of image, and long time series features before the (i+1)th frame of image; / 10. performing super-resolution prediction on the image features of the ith frame of image, the long time series features before the ith frame of image, and the (i+1)th frame of image using a generative network, to obtain a super-resolution image of the (i+1)th frame of image, image features of the (i+1)th frame of image, and long time series features before the (i+1)th frame of image; / 19. performing super-resolution prediction on the image features of the ith frame of image, the long time series features before the ith frame of image, and the (i+1)th frame of image using a generative network, to obtain a super-resolution image of the (i+1)th frame of image, image features of the (i+1)th frame of image, and long time series features before the (i+1)th frame of image; (Shenzen: Fig. 1 is a flowchart of an image super-resolution method provided in an embodiment of the present application. In the embodiment of the present application, an execution subject of the image super-resolution method is a terminal device, where the terminal device includes, but is not limited to, a mobile device such as a smart phone, a tablet computer, a Personal Digital Assistant (PDA), and the like, and may also include a device such as a desktop computer. As shown in fig. 1, the method comprises the steps of: step 101: a first image is acquired, and the resolution of the first image is a first resolution. It should be noted that the resolution of an image refers to the amount of information stored in the image, and is used to indicate the pixel density in the image, and is usually expressed by how many pixels are per inch of the image. The first image is a low-resolution image to be subjected to super-resolution processing, and a corresponding high-resolution image needs to be obtained through the super-resolution processing. The first resolution may be less than or equal to a first resolution threshold, and the first resolution threshold may be preset according to actual needs. For example, the first resolution threshold is 100DPI (Dots Per Inch pixels) or 300DPI (deep depth Per Inch), or the like. Illustratively, the first resolution is 20DPI or 72DPI, etc. The first image may be a single image, or may be any video frame in one video, and the like. As an example, the embodiment of the present application may perform super-resolution processing on a sequence of video frames according to the method provided by the embodiment of the present application to obtain clearer video and improve the video experience of a user. In addition, the first image may be obtained by shooting with a camera, may be obtained from a storage space of a terminal device, may be obtained by downloading from a network, or may be obtained by sending from another device. Step 102: and inputting the first image into a super-resolution network model for processing to obtain a second image, wherein the resolution of the second image is the second resolution, and the second resolution is higher than the first resolution, the super-resolution network model is a network model adopting a pixel attention mechanism, and the pixel attention mechanism is used for generating a corresponding weight for each pixel in the input feature map. The image resolution of the second image is greater than the image resolution of the first image, that is, the first image is a low-resolution image, and the second image is a high-resolution image having a resolution greater than that of the first image. The super-resolution network model is used to recover a high-resolution image from a low-resolution image, i.e. to increase the resolution of the image. The first image is processed by the super-resolution network model, and a second image with higher resolution can be obtained. For example, a first image with a resolution of 72DPI may be processed by a super-resolution network model to obtain a second image with a resolution of 300DPI or 350 DPI. As an example, the second resolution is greater than or equal to a second resolution threshold, such as the second resolution threshold being 300DPI or 500DPI, or the like. Alternatively, the second resolution is m times the first resolution, m being greater than 1. For example, m is 5 or 10. Alternatively, the resolution difference between the second resolution and the first resolution is greater than or equal to a difference threshold, for example, the difference threshold is 200 or 300 DPI. It should be noted that the pixel attention mechanism is a pixel-dimensional attention mechanism for generating a corresponding weight for each pixel in the input feature map. The input feature map may be a feature map generated at any link in the super-resolution network model, for example, a feature map generated at any link in the super-resolution network model for the first image. That is, in the embodiment of the present application, the pixel attention mechanism may be set in any task link of the super-resolution network model according to actual needs, so as to improve the feature processing efficiency. It should be further noted that the pixel attention mechanism increases the utilization efficiency of parameters and features by generating a corresponding weight for each pixel, so that the pixel attention mechanism is more effective for underlying vision problems and lightweight networks, and is more suitable for an underlying vision task such as super-resolution. Moreover, the pixel attention mechanism can be applied to various deep learning tasks in a plug-and-play manner, and has good universality. Fig. 3 is a flowchart of another super-resolution method provided by an embodiment of the present application, which is applied to the super-resolution network model shown in fig. 2, and as shown in fig. 3, the method includes the following steps: step 301: a first image is acquired, and the resolution of the first image is a first resolution. Step 301 may refer to the description related to step 101 in the embodiment of fig. 1, and is not described herein again in this embodiment of the application. Step 302: and inputting the first image into a feature extraction module of the super-resolution network model, and performing feature extraction on the first image through the feature extraction module to obtain a first feature map. The feature extraction module is used for extracting features of the first image to obtain a feature map of the first image. As one example, the feature extraction module may be a convolutional layer. The convolution layer may be an M × M convolution layer, i.e., the convolution kernel is M × M, and M is a positive integer. In this embodiment, M may be preset, for example, M may be 1, 3, 4, or 5, and this is not limited in this application. As an example, the convolution layer of the feature extraction module is a 3 × 3 convolution layer, and the 3 × 3 convolution layer refers to a convolution layer with a convolution kernel of 3 × 3.) -; 1. and caching the image features of the (i+1)th frame of image and the long time series features before the (i+1)th frame of image; i being a positive integer greater than 2. / 10. and caching the image features of the (i+1)th frame of image and the long time series features before the (i+1)th frame of image; i being a positive integer greater than 2. / 19. and caching the image features of the (i+1)th frame of image and the long time series features before the (i+1)th frame of image; i being a positive integer greater than 2. (Shenzen: As one example, the pixel attention mechanism is used to generate a corresponding three-dimensional attention map for a feature map of the first image, the three-dimensional attention map including channel information for the feature map, location information for each pixel in the feature map, and a weight for each pixel. The channel information is used to indicate the channel where the feature map is located, and may be a channel number or the like. The position information is used to indicate the position of the pixel in the feature map, and may be coordinates or the like. As an example, the pixel attention mechanism may be implemented by a specified convolution layer, i.e., a specified convolution may be used to generate a corresponding weight for each pixel in the input feature map. The pixel attention mechanism is realized by only using one convolution layer, so that the operation is simpler and more effective, the calculated amount of the model is reduced, and the actual running speed of the model is improved. For example, the pixel attention mechanism is used to perform convolution processing on an input feature map by specifying a convolution layer, and generate a three-dimensional attention map of the feature map, where the three-dimensional attention map of the feature map includes a channel of the feature map, position information of each pixel in the feature map, and a weight of each pixel. It should be noted that the designated convolutional layer may be an N × N convolutional layer, that is, the convolutional core of the designated convolutional layer is N × N, where N is a positive integer. N may be preset, for example, N may be 1, 2, or 3, and the like, which is not limited in this embodiment of the application. As an example, the convolutional layer is designated as 1 × 1 convolutional layer, i.e., the convolutional kernel of the convolutional layer is designated as 1 × 1. Because the calculation of the 1 x 1 convolution layer on the hardware is simpler, the calculated amount of the model is further reduced, and the actual running speed of the model is improved. In the embodiment of the application, the super-resolution network model adopting the pixel attention mechanism is used for processing the low-resolution images to generate the corresponding high-resolution images so as to obtain clear images with image quality according with practical application scenes, and therefore cleaner, clear, natural and comfortable image or video experience can be provided for users. Moreover, because the super-resolution network model adopts the pixel attention mechanism, the model can pay more attention to key features in the image processing process, so that the features can be more efficiently utilized, the parameters and the calculated amount required by the super-resolution network model are greatly reduced, the processing speed of the super-resolution network model is improved, the network transmission burden is reduced, and the economic cost is reduced. Further, the pixel attention mechanism can generate corresponding weight for each pixel in the feature map in the network, and utilization rate of parameters and features is increased, so that the pixel attention mechanism is more effective for bottom layer vision problems and lightweight networks, and is more suitable for bottom layer vision tasks such as super-resolution compared with a space attention mechanism or a channel attention mechanism designed for high-level vision tasks, super-resolution performance is improved, the super-resolution effect of the obtained high-resolution images is better, and image or video experience of users is improved. Furthermore, the pixel attention mechanism can be arranged in any task link of the super-resolution network model according to requirements, and the method has good universality. Furthermore, the pixel attention mechanism is realized by only using one convolution layer, so that the operation is simpler and more effective, the operation on hardware is simpler, the calculated amount of the model is reduced, and the actual running speed of the model is improved. Fig. 2 is a schematic structural diagram of a super-resolution network model provided in an embodiment of the present application, and as shown in fig. 2, the super-resolution network model includes a feature extraction module 21, a nonlinear mapping module 22, and a reconstruction module 23. For convenience of description, the following will describe, by taking the pixel attention mechanism as an example, in which at least one of the nonlinear mapping module 22 and the reconstruction module 23 employs the pixel attention mechanism, that is, the pixel attention mechanism may be set in at least one of the nonlinear mapping task and the reconstruction task of the super-resolution network model. Fig. 3 is a flowchart of another super-resolution method provided by an embodiment of the present application, which is applied to the super-resolution network model shown in fig. 2, and as shown in fig. 3, the method includes the following steps: Step 303: inputting the first characteristic diagram into a nonlinear mapping module, carrying out nonlinear mapping on the first characteristic diagram through the nonlinear mapping module to obtain a second characteristic diagram, wherein the nonlinear mapping module adopts a pixel attention mechanism. The nonlinear mapping module is used for fitting the mapping relation between the low-resolution image and the high-resolution image, so that the reconstruction module reconstructs the high-resolution image based on the mapping relation between the low-resolution image and the high-resolution image. As an example, the nonlinear mapping module is a nonlinear mapping module that uses a pixel attention mechanism, and in the process of performing nonlinear mapping on the first feature map, the feature map segmentation may be performed on the first feature map first, and then the pixel attention mechanism is used on the segmented feature map.) It would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to modify the image processing method and system of Zhang with the image super-resolution method and system of Shenzen. The determination of obviousness is predicated upon the following findings: Both are directed towards image processing methods and systems; One skilled in the art would have been motivated to modify Zhang for processing multi-frame image data using a GAN to process the image data with the teachings of Shenzen in order to improved algorithm for super-resolution image processing in order to ensure a higher quality of image processing and reconstruction. Furthermore, the prior art collectively includes each element claimed (though not all in the same reference), and one of ordinary skill in the art could have combined the elements in the manner explained above using known engineering design, interface and/or programming techniques, without changing a “fundamental” operating principle of Zhang, while the teaching of Shenzen continues to perform the same function as originally taught prior to being combined, in order to produce the repeatable and predictable result of applying super-resolution image processing in order to ensure a higher quality of image processing and reconstruction.. It is for at least the aforementioned reasons that the examiner has reached a conclusion of obviousness with respect to the claim in question. Consider Claims 2, 11 and 20. The combination of Zhang and Shenzen teaches: 2. The method according to claim 1, wherein the generative network comprises a feature extraction network, a feature fusion network, and an upsampling network; and performing the super-resolution prediction on the image features of the ith frame of image, the long time series features before the ith frame of image, and the (i+1)th frame of image by using a generative network comprises: performing feature extraction on the (i+1)th frame of image using the feature extraction network, to obtain the image features of the (i+1)th frame of image; fusing the image features of the ith frame of image, the long time series features before the ith frame of image, and the image features of the (i+1)th frame of image using the feature fusion network, to obtain the long time series features before the (i+1)th frame of image; and performing prediction on the image features of the (i+1)th frame of image and the long time series features before the (i+1)th frame of image by using the upsampling network, to obtain the super-resolution image of the (i+1)th frame of image. / 11. The device according to claim 10, wherein the generative network comprises a feature extraction network, a feature fusion network, and an upsampling network; and the one or more processors are further configured to perform: performing feature extraction on the (i+1)th frame of image using the feature extraction network, to obtain the image features of the (i+1)th frame of image; fusing the image features of the ith frame of image, the long time series features before the ith frame of image, and the image features of the (i+1)th frame of image using the feature fusion network, to obtain the long time series features before the (i+1)th frame of image; and performing prediction on the image features of the (i+1)th frame of image and the long time series features before the (i+1)th frame of image by using the upsampling network, to obtain the super-resolution image of the (i+1)th frame of image. / 20. The storage medium according to claim 19, wherein the generative network comprises a feature extraction network, a feature fusion network, and an upsampling network; and the at least one processor is further configured to perform: performing feature extraction on the (i+1)th frame of image using the feature extraction network, to obtain the image features of the (i+1)th frame of image; fusing the image features of the ith frame of image, the long time series features before the ith frame of image, and the image features of the (i+1)th frame of image using the feature fusion network, to obtain the long time series features before the (i+1)th frame of image; and performing prediction on the image features of the (i+1)th frame of image and the long time series features before the (i+1)th frame of image by using the upsampling network, to obtain the super-resolution image of the (i+1)th frame of image. (Zhang: [0048] Step A13 includes fusing the high-level spatio-temporal features corresponding to the blurry sample image frame, to acquire spatio-temporal information corresponding to the blurry sample image frame. [0049] At first, two 3D convolutional layers are set in the generative network model. On each 3D convolutional layer, different 3D convolution kernels may be used. For example, the first 3D convolution kernel and the second 3D convolution kernel have different weight parameters. The first 3D convolution kernel is first used to perform the convolution processing on the N sample image frames, to acquire the low-level spatio-temporal features corresponding to the blurry sample image frame. The low-level spatio-temporal features indicate inconspicuous feature information, such as features of lines. Then taking the low-level spatio-temporal features as an input condition, the convolution processing is performed on the other 3D convolutional layer, to acquire the high-level spatio-temporal features corresponding to the blurry sample image frame. The high-level spatio-temporal features indicate feature information of neighboring image frames. At last, the high-level spatio-temporal features are fused together to acquire the spatio-temporal information corresponding to the blurry sample image frame. The spatio-temporal information can be used as a feature map for training of the generative network model. The following is an illustration with an example: the first 3D convolution kernel is first used to perform the convolution processing on 5 sample image frames, to acquire 3 low-level spatio-temporal features in different dimensions; then the convolution processing is performed on the low-level spatio-temporal features by using the second 3D convolution kernel, to acquire high-level spatio-temporal features corresponding to the blurry sample image frames; and the high-level spatio-temporal features are fused to acquire spatio-temporal information corresponding to the blurry sample image frames. Because 5 frames of images are inputted into the generative network model, after 2 times of 3D convolution processing, one frame of feature map is outputted, that is, after the 2 times of 3D convolution processing, a quantity of channels of time series is changed from 5 to 1. [0050] Further, according to some embodiments of this application, the generative network model further includes M 2D convolution kernels, the M being a positive integer. Step A3 that the electronic device performs deblurring processing on the blurry sample image frame by using the spatio-temporal information corresponding to the blurry sample image frame through the generative network model, to output a sharp sample image frame includes the following steps. [0051] Step A31 includes performing convolution processing on the spatio-temporal information corresponding to the blurry sample image frame by using each 2D convolution kernel of the M 2D convolution kernels in sequence, and acquire the sharp sample image frame after the convolution processing is performed by using the last 2D convolution kernel of the M 2D convolution kernels. / Shenzen: Fig. 3 is a flowchart of another super-resolution method provided by an embodiment of the present application, which is applied to the super-resolution network model shown in fig. 2, and as shown in fig. 3, the method includes the following steps: step 301: a first image is acquired, and the resolution of the first image is a first resolution. Step 301 may refer to the description related to step 101 in the embodiment of fig. 1, and is not described herein again in this embodiment of the application. Step 302: and inputting the first image into a feature extraction module of the super-resolution network model, and performing feature extraction on the first image through the feature extraction module to obtain a first feature map. The feature extraction module is used for extracting features of the first image to obtain a feature map of the first image. As one example, the feature extraction module may be a convolutional layer. The convolution layer may be an M × M convolution layer, i.e., the convolution kernel is M × M, and M is a positive integer. In this embodiment, M may be preset, for example, M may be 1, 3, 4, or 5, and this is not limited in this application. As an example, the convolution layer of the feature extraction module is a 3 × 3 convolution layer, and the 3 × 3 convolution layer refers to a convolution layer with a convolution kernel of 3 × 3. Step 303: inputting the first characteristic diagram into a nonlinear mapping module, carrying out nonlinear mapping on the first characteristic diagram through the nonlinear mapping module to obtain a second characteristic diagram, wherein the nonlinear mapping module adopts a pixel attention mechanism. The nonlinear mapping module is used for fitting the mapping relation between the low-resolution image and the high-resolution image, so that the reconstruction module reconstructs the high-resolution image based on the mapping relation between the low-resolution image and the high-resolution image. As an example, the nonlinear mapping module is a nonlinear mapping module that uses a pixel attention mechanism, and in the process of performing nonlinear mapping on the first feature map, the feature map segmentation may be performed on the first feature map first, and then the pixel attention mechanism is used on the segmented feature map. Referring to fig. 4, fig. 4 is a schematic structural diagram of another super-resolution network model provided in the embodiment of the present application, and as shown in fig. 4, the non-linear mapping module 22 may include a plurality of designated self-correcting volume blocks. Wherein, designating a self-correcting volume block refers to a self-correcting volume block using a pixel attention mechanism. The non-linear mapping module 22 may perform non-linear mapping on the first feature map through a plurality of designated self-correcting volume blocks to obtain a second feature map) Consider Claims 3 and 12. The combination of Zhang and Shenzen teaches: 3. The method according to claim 2, wherein the feature fusion network comprises a first feature fusion layer and a second feature fusion layer; and fusing the image features of the ith frame of image, the long time series features before the ith frame of image, and the image features of the (i+1)th frame of image comprises: fusing the image features of the (i+1)th frame of image and the image features of the ith frame of image by using the first feature fusion layer, to obtain fused time series features; and fusing the fused time series features and the long time series features before the ith frame of image by using the second feature fusion layer, to obtain the long time series features before the (i+1)th frame of image. / 12. The device according to claim 11, wherein the feature fusion network comprises a first feature fusion layer and a second feature fusion layer; and the one or more processors are further configured to perform: fusing the image features of the (i+1)th frame of image and the image features of the ith frame of image by using the first feature fusion layer, to obtain fused time series features; and fusing the fused time series features and the long time series features before the ith frame of image by using the second feature fusion layer, to obtain the long time series features before the (i+1)th frame of image. (Zhang: [0048] Step A13 includes fusing the high-level spatio-temporal features corresponding to the blurry sample image frame, to acquire spatio-temporal information corresponding to the blurry sample image frame. [0049] At first, two 3D convolutional layers are set in the generative network model. On each 3D convolutional layer, different 3D convolution kernels may be used. For example, the first 3D convolution kernel and the second 3D convolution kernel have different weight parameters. The first 3D convolution kernel is first used to perform the convolution processing on the N sample image frames, to acquire the low-level spatio-temporal features corresponding to the blurry sample image frame. The low-level spatio-temporal features indicate inconspicuous feature information, such as features of lines. Then taking the low-level spatio-temporal features as an input condition, the convolution processing is performed on the other 3D convolutional layer, to acquire the high-level spatio-temporal features corresponding to the blurry sample image frame. The high-level spatio-temporal features indicate feature information of neighboring image frames. At last, the high-level spatio-temporal features are fused together to acquire the spatio-temporal information corresponding to the blurry sample image frame. The spatio-temporal information can be used as a feature map for training of the generative network model. The following is an illustration with an example: the first 3D convolution kernel is first used to perform the convolution processing on 5 sample image frames, to acquire 3 low-level spatio-temporal features in different dimensions; then the convolution processing is performed on the low-level spatio-temporal features by using the second 3D convolution kernel, to acquire high-level spatio-temporal features corresponding to the blurry sample image frames; and the high-level spatio-temporal features are fused to acquire spatio-temporal information corresponding to the blurry sample image frames. Because 5 frames of images are inputted into the generative network model, after 2 times of 3D convolution processing, one frame of feature map is outputted, that is, after the 2 times of 3D convolution processing, a quantity of channels of time series is changed from 5 to 1. [0050] Further, according to some embodiments of this application, the generative network model further includes M 2D convolution kernels, the M being a positive integer. Step A3 that the electronic device performs deblurring processing on the blurry sample image frame by using the spatio-temporal information corresponding to the blurry sample image frame through the generative network model, to output a sharp sample image frame includes the following steps. [0051] Step A31 includes performing convolution processing on the spatio-temporal information corresponding to the blurry sample image frame by using each 2D convolution kernel of the M 2D convolution kernels in sequence, and acquire the sharp sample image frame after the convolution processing is performed by using the last 2D convolution kernel of the M 2D convolution kernels. / Shenzen: As one example, a given upsampled convolution block includes an upsampled layer, a pixel attention layer, and a convolution layer. Referring to fig. 7, fig. 7 is a schematic structural diagram of a specified upsampling volume block according to an embodiment of the present application. As shown in fig. 7, the designated upsampling convolutional block includes an upsampling layer 71, a sixth convolutional layer 72, a second pixel attention layer 73, and a seventh convolutional layer 74. The up-sampling layer 71 is configured to perform amplification processing on the second input feature map to obtain a ninth feature map. By way of example, the upsampling layer 71 may be a nearest neighbor amplification layer, such as a nearest neighbor amplification layer for amplifying the second input feature map by a factor of two. The sixth convolution layer 72 is configured to perform convolution processing on the ninth feature map to obtain a tenth feature map. Illustratively, the sixth convolution layer 72 is an M × M convolution layer, such as a 1 × 1 convolution layer, or a 3 × 3 convolution layer, for example. The second pixel attention layer 73 is configured to process the tenth feature map, and generate a weight for each pixel in the tenth feature map. In addition, an eleventh feature map may be generated based on the weight of each pixel in the tenth feature map and the tenth feature map. For example, the eleventh feature map is obtained by performing weighting processing on the corresponding pixel in the tenth feature map based on the weight of each pixel in the tenth feature map.) Consider Claims 4 and 13. The combination of Zhang and Shenzen teaches: 4. The method according to claim 3, further comprising: obtaining a first frame of image from the video; performing super-resolution prediction on the first frame of image using the generative network, to obtain a super-resolution image of the first frame of image and image features of the first frame of image; obtaining a second frame of image from the video; performing super-resolution prediction on the image features of the first frame of image and the second frame of image using the generative network, to obtain a super-resolution image of the second frame of image and image features of the second frame of image; obtaining a third frame of image from the video; performing super-resolution prediction on the image features of the second frame of image and the third frame of image using the generative network, to obtain a super-resolution image of the third frame of image, image features of the third frame of image, and long time series features before the third frame of image; and caching the image features of the third frame of image and the long time series features before the third frame of image. / 13. The device according to claim 12, wherein the one or more processors are further configured to perform: obtaining a first frame of image from the video; performing super-resolution prediction on the first frame of image using the generative network, to obtain a super-resolution image of the first frame of image and image features of the first frame of image; obtaining a second frame of image from the video; performing super-resolution prediction on the image features of the first frame of image and the second frame of image using the generative network, to obtain a super-resolution image of the second frame of image and image features of the second frame of image; obtaining a third frame of image from the video; performing super-resolution prediction on the image features of the second frame of image and the third frame of image using the generative network, to obtain a super-resolution image of the third frame of image, image features of the third frame of image, and long time series features before the third frame of image; and caching the image features of the third frame of image and the long time series features before the third frame of image. (Zhang: [0030] According to this embodiment of this application, the video clip may be shot by the terminal through a camera, or may be downloaded from the Internet by the terminal. As long as there is at least one frame of blurry image in the video clip, a sharp image can be restored by using the video deblurring method according to this embodiment of this application. At first, the N continuous image frames are acquired from the video clip, the N image frames including at least one blurry image frame to be processed. The blurry image frame may be caused by a shake of a shooting device or a movement of an object to be shot. There may be one blurry image frame to be processed of the N continuous image frame acquired at first according to this embodiment of this application. For example, the blurry image frame may be a middle image frame of the N continuous image frame. If the value of N is 3, the blurry image frame may be the second image frame, or if the value of N is 5, the blurry image frame may be the third image frame. The value of N is a positive integer not limited herein. [0031] In step 102, the electronic device performs 3D convolution processing on the N image frames with a generative adversarial network model, to acquire spatio-temporal information corresponding to the blurry image frame, the spatio-temporal information including: spatial feature information of the blurry image frame, and temporal feature information between the blurry image frame and a neighboring image frame of the N image frames. / Shenzen: As can be seen from fig. 5, the designated self-correcting convolution block includes two branches, both of which start with a convolution layer, namely a first convolution layer 51 and a second convolution layer 52. The first convolution layer 51 and the second convolution layer 52 are used for respectively performing feature map segmentation on the input feature map of the designated self-correcting convolution block to obtain a first segmentation map and a second segmentation map, and then respectively inputting the first segmentation map and the second segmentation map into two branches. For example, the first segmentation map and the second segmentation map may be feature maps that have half of the features in the original input feature map after feature map segmentation. In addition, the first convolution layer 51 and the second convolution layer 52 may both be N × N convolution layers, such as 1 × 1 convolution layers. For the upper branch, the upper branch also includes two convolutional layers after the first convolutional layer 51: a third convolutional layer 53 and a fourth convolutional layer 55, and the third convolutional layer 53 is provided with a pixel attention mechanism. The third convolution layer 53 is used to perform convolution processing on the first segmentation map to obtain a fourth feature map. The first pixel attention layer 54 is used to process the first segmentation map to obtain a weight for each pixel in the first segmentation map. The fifth feature map may be generated based on the weight of each pixel in the first segmentation map and the fourth feature map, for example, each pixel in the fourth feature map may be weighted based on the weight of each pixel in the first segmentation map to obtain the fifth feature map. The fourth convolution layer 55 is used to perform convolution processing on the fifth feature map to obtain a sixth feature map. For example, the third convolution layer 53 and the fourth convolution layer 55 may both be M × M convolution layers, such as 3 × 3 convolution layers. Illustratively, the first pixel attention layer 54 includes a specified convolutional layer, or includes a specified convolutional layer and a first logistic regression layer. And designating the convolution layer for performing convolution processing on the first segmentation map to generate a three-dimensional attention map of the first segmentation map, wherein the three-dimensional attention map of the first segmentation map comprises channel information of the first segmentation map, position information of each pixel in the first segmentation map and weight of each pixel. The first logistic regression layer is used for conducting logistic regression processing on the weight of each pixel in the first segmentation graph so as to map the weight of each pixel in the first segmentation graph to be between 0 and 1. Optionally, the processing module 1102 includes: the feature extraction unit is used for inputting the first image into the feature extraction module, and performing feature extraction on the first image through the feature extraction module to obtain a first feature map; the nonlinear mapping unit is used for inputting the first feature map into the nonlinear mapping module, and the nonlinear mapping module is used for carrying out nonlinear mapping on the first feature map to obtain a second feature map; and the reconstruction unit is used for inputting the second feature map into the reconstruction module, reconstructing the second feature map through the reconstruction module to obtain a third feature map, and generating the second image based on the third feature map. Optionally, the non-linear mapping module includes a plurality of designated self-correcting volume blocks, where the designated self-correcting volume blocks are self-correcting volume blocks using a pixel attention mechanism; the nonlinear mapping unit is configured to: and carrying out nonlinear mapping on the first feature map through the plurality of specified self-correction volume blocks to obtain the second feature map. Optionally, a first self-correcting volume block of the plurality of designated self-correcting volume blocks includes at least a first volume layer, a second volume layer, a third volume layer, a first pixel attention layer, a fourth volume layer, a fifth volume layer, and a blend layer, the first self-correcting volume block being any one of the plurality of designated self-correcting volume blocks; the nonlinear mapping unit is configured to: for the first self-correcting convolution block, respectively carrying out feature map segmentation on a first input feature map through the first convolution layer and the second convolution layer to obtain a first segmentation map and a second segmentation map; wherein, if the first self-corrected volume block is a first designated self-corrected volume block in the plurality of designated self-corrected volume blocks, the first input feature map is the first feature map, and if the first self-corrected volume block is another designated self-corrected volume block except the first designated self-corrected volume block in the plurality of designated self-corrected volume blocks, the first input feature map is a feature map output by a last designated self-corrected volume block of the first self-corrected volume block; inputting the first segmentation graph into the third convolution layer for convolution processing to obtain a fourth feature graph, inputting the first segmentation graph into the first pixel attention layer for processing to generate a weight of each pixel in the first segmentation graph, generating a fifth feature graph based on the weight of each pixel in the first segmentation graph and the fourth feature graph, and inputting the fifth feature graph into the fourth convolution layer for convolution processing to obtain a sixth feature graph; inputting the second segmentation graph into the fifth convolution layer for convolution processing to obtain a seventh feature graph; inputting the sixth feature map and the seventh feature map into the fusion layer for feature map fusion to obtain an eighth feature map, and generating an output feature map of the first self-correcting volume block based on the eighth feature map.) Consider Claims 5 and 14. The combination of Zhang and Shenzen teaches: 5. The method according to claim 1, wherein the generative network is trained by performing: caching an ith frame of sample image and an (i+1)th frame of sample image from a sample video, i being a positive integer greater than 2; predicting a super-resolution image of the ith frame of sample image and a super-resolution image of the (i+1)th frame of sample image using the generative network; calculating an inter-frame stability loss between a first change and a second change using an inter-frame stability loss function, the first change being a change between the ith frame of sample image and the (i+1)th frame of sample image, the second change being a change between the super-resolution image of the ith frame of sample image and the super-resolution image of the (i+1)th frame of sample image, and the inter-frame stability loss being configured for constraining super-resolution stability between adjacent frames of images; and training the generative network based on the inter-frame stability loss. / 14. The device according to claim 10, wherein the one or more processors are further configured to train the generative network by performing: caching an ith frame of sample image and an (i+1)th frame of sample image from a sample video, i being a positive integer greater than 2; predicting a super-resolution image of the ith frame of sample image and a super-resolution image of the (i+1)th frame of sample image using the generative network; calculating an inter-frame stability loss between a first change and a second change using an inter-frame stability loss function, the first change being a change between the ith frame of sample image and the (i+1)th frame of sample image, the second change being a change between the super-resolution image of the ith frame of sample image and the super-resolution image of the (i+1)th frame of sample image, and the inter-frame stability loss being configured for constraining super-resolution stability between adjacent frames of images; and training the generative network based on the inter-frame stability loss. (Examiner Note: Zhang teaches both a reconstruction loss and adversarial loss functions, wherein an inter-frame stability loss function is analogous in scope to the adversarial loss function and is obviated by these two functions; Zhang: [0044] After the generative network model outputs the sharp sample image frame, whether the outputted sharp sample image frame is blurry or sharp is discriminated according to the sharp sample image frame and the real sharp image frame by using the discriminative network model. An adversarial loss function is introduced by using the discriminative network model, to alternately train the generative network model and the discriminative network model a plurality of times, thereby better ensuring that a restored sharp video is more real. [0055] According to some embodiments of this application, step A4 that the electronic device trains the generative network model and the discriminative network model alternately according to the sharp sample image frame and the real sharp image frame includes the following steps. [0056] Step A41 includes acquiring a reconstruction loss function according to the sharp sample image frame and the real sharp image frame. [0057] Step A42 includes training the generative network model through the reconstruction loss function. [0058] Step A43 includes training the discriminative network model by using the real sharp image frame and the sharp sample image frame, to acquire an adversarial loss function outputted by the discriminative network model. [0059] Step A44 includes training the generative network model continually through the adversarial loss function. [0060] In order to acquire a more real deblurred video, the discriminative network model may be further introduced when the generative network model is trained. During the training, the generative network model is first trained: inputting a blurry sample image frame into the generative network model, to generate a sharp sample image frame; acquiring the reconstruction loss function by comparing the sharp sample image frame to the real sharp image frame; and adjusting a weight parameter of the generative network model through the reconstruction loss function. Next, the discriminative network model is trained: inputting the real sharp video and the generated sharp sample video into the discriminative network model, to acquire the adversarial loss function; and adjusting the generative network model through the adversarial loss function, to equip the discriminative network model with a capability of discriminating between the real sharp image and the sharp image generated from the blurry image frame. In this way, structures of the two network models are alternately trained. [0061] Further, according to some embodiments of this application, step A4 of training the generative network model and the discriminative network model alternately according to the sharp sample image frame and the real sharp image frame, in addition to including the foregoing step A41 to step A44, may further include the following steps. [0062] Step A45 includes reacquiring a reconstruction loss function through the generative network model, and reacquire an adversarial loss function through the discriminative network model, after the training the generative network model continually through the adversarial loss function. [0063] Step A46 includes performing weighted fusion on the reacquired reconstruction loss function and the reacquired adversarial loss function, to acquire a fused loss function. [0064] Step A47 includes training the generative network model continually through the fused loss function. [0065] After training the generative network model and the discriminative network model through the foregoing step A41 to step A44, step A45 to step A47 are performed based on the two network models after the first training. When training the generative network model again, a structure of the generative network model is adjusted by using the two types of loss functions together, so that an image can be similar to a real sharp image at a pixel level, and appear more like a sharp image as a whole. The two loss functions can be joined with a weight parameter. The weight can be used to control the effect of the two types of loss functions on feedback regulation. The effect of the generative network model is to generate a sharp video from a blurry video, and the effect of the discriminative network model is to discriminate whether an inputted video frame is a real sharp image or a generated sharp image. Through the adversarial learning according to this embodiment of this application, the discriminative capability of the discriminative network model gets increasingly strong, and the video generated by the generative network model gets increasingly real. [0066] According to the above description, the N continuous image frames are first acquired from the video clip, the N image frames including the blurry image frame to be processed; then the 3D convolution processing is performed on the N image frames with the generative adversarial network model, to acquire the spatio-temporal information corresponding to the blurry image frame, the spatio-temporal information including: the spatial feature information of the blurry image frame, and the temporal feature information between the blurry image frame and the neighboring image frame of the N image frames. The deblurring processing is finally performed on the blurry image frame by using the spatio-temporal information corresponding to the blurry image frame through the generative adversarial network model, to output the sharp image frame. According to the embodiments of this application, the generative adversarial network model extracts the spatio-temporal information implicit between the continuous image frames through a 3D convolution operation, so that the deblurring processing on the blurry image frame is completed by using the spatio-temporal information corresponding to the blurry image frames through the generative adversarial network model. Therefore, a more real sharp image can be obtained, and the effect of video deblurring is enhanced. [0072] FIG. 3 is a schematic diagram of a training process of a generative network model and a discriminative network model according to an embodiment of this application. A discriminative network model (briefly referred to as discriminative network) and a generative network model (briefly referred to as generative network) are joined together to form an adversarial network. The two compete with each other. In order to acquire a more real deblurred video, an adversarial network structure is introduced when the generative network model shown in FIG. 2 is trained. The network structure in FIG. 2 is taken as a generator (namely, generative network model), and a discriminator (namely, discriminative network model) is added. During the training, the generative network is trained at first: inputting a blurry video frame into the generative network, to acquire a sharp video frame; acquiring a reconstruction loss function (namely, loss function 1 shown in FIG. 3), by comparing the sharp video frame to a real video frame; and adjusting a weight parameter of the generative network through the loss function. Next, the discriminative network is trained: inputting the real sharp video and the generated sharp video into the discriminative network, to acquire an adversarial loss function (namely, loss function 2 in FIG. 3); and adjusting the generative network structure through the adversarial loss function, to equip the discriminative network with a capability of discriminating between a real sharp image and a sharp image generated from a blurry image. The two network structures are alternately trained. When the generative network is trained later, a structure of the network model is adjusted by using the two types of loss functions together, so that an image can be similar to a real sharp image at a pixel level, and appear more like a sharp image as a whole. The two loss functions are joined with a weight parameter. The weight can be used to control the effect of the two types of loss functions on feedback regulation. The effect of the generative network is to generate a sharp video from a blurry video, and the effect of the discriminative network is to discriminate whether the inputted video frame is a real sharp image or a generated sharp video frame. Through the adversarial learning, the discriminative capability of the discriminative network gets increasingly strong, and the video generated by the generative network gets increasingly real. [0073] Next, weighted fusion of two different types of loss functions is illustrated with an example. Because two networks are used in this embodiment of this application, that is, a generative network and a discriminative network, two loss functions are used in this embodiment of this application, that is, a content loss function based on pixel-value differencing (that is a reconstruction loss function) and an adversarial loss function. [0074]-[0079]) Consider Claims 7 and 13. The combination of Zhang and Shenzen teaches: 7. The method according to claim 5, further comprising: discriminating between the super-resolution images of the (i+1)th frame of sample image and the (i+1)th frame of sample image using a discriminative network, to obtain a discrimination result; calculating a first error loss of the discrimination result based on the discrimination result and an adversarial loss function; and training the generative network and the discriminative network alternately based on the first error loss; the adversarial loss function being configured for constraining consistency between the super-resolution result of the (i+1)th frame of sample image and the discrimination result. 16. The device according to claim 14, wherein the one or more processors are further configured to perform: discriminating between the super-resolution images of the (i+1)th frame of sample image and the (i+1)th frame of sample image using a discriminative network, to obtain a discrimination result; calculating a first error loss of the discrimination result based on the discrimination result and an adversarial loss function; and training the generative network and the discriminative network alternately based on the first error loss; the adversarial loss function being configured for constraining consistency between the super-resolution result of the (i+1)th frame of sample image and the discrimination result. (Examiner Note: Zhang teaches both a reconstruction loss and adversarial loss functions, wherein an inter-frame stability loss function is analogous in scope to the adversarial loss function and is obviated by these two functions; Zhang: [0044] After the generative network model outputs the sharp sample image frame, whether the outputted sharp sample image frame is blurry or sharp is discriminated according to the sharp sample image frame and the real sharp image frame by using the discriminative network model. An adversarial loss function is introduced by using the discriminative network model, to alternately train the generative network model and the discriminative network model a plurality of times, thereby better ensuring that a restored sharp video is more real. [0055] According to some embodiments of this application, step A4 that the electronic device trains the generative network model and the discriminative network model alternately according to the sharp sample image frame and the real sharp image frame includes the following steps. [0056] Step A41 includes acquiring a reconstruction loss function according to the sharp sample image frame and the real sharp image frame. [0057] Step A42 includes training the generative network model through the reconstruction loss function. [0058] Step A43 includes training the discriminative network model by using the real sharp image frame and the sharp sample image frame, to acquire an adversarial loss function outputted by the discriminative network model. [0059] Step A44 includes training the generative network model continually through the adversarial loss function. [0060] In order to acquire a more real deblurred video, the discriminative network model may be further introduced when the generative network model is trained. During the training, the generative network model is first trained: inputting a blurry sample image frame into the generative network model, to generate a sharp sample image frame; acquiring the reconstruction loss function by comparing the sharp sample image frame to the real sharp image frame; and adjusting a weight parameter of the generative network model through the reconstruction loss function. Next, the discriminative network model is trained: inputting the real sharp video and the generated sharp sample video into the discriminative network model, to acquire the adversarial loss function; and adjusting the generative network model through the adversarial loss function, to equip the discriminative network model with a capability of discriminating between the real sharp image and the sharp image generated from the blurry image frame. In this way, structures of the two network models are alternately trained. [0061] Further, according to some embodiments of this application, step A4 of training the generative network model and the discriminative network model alternately according to the sharp sample image frame and the real sharp image frame, in addition to including the foregoing step A41 to step A44, may further include the following steps. [0062] Step A45 includes reacquiring a reconstruction loss function through the generative network model, and reacquire an adversarial loss function through the discriminative network model, after the training the generative network model continually through the adversarial loss function. [0063] Step A46 includes performing weighted fusion on the reacquired reconstruction loss function and the reacquired adversarial loss function, to acquire a fused loss function. [0064] Step A47 includes training the generative network model continually through the fused loss function. [0065] After training the generative network model and the discriminative network model through the foregoing step A41 to step A44, step A45 to step A47 are performed based on the two network models after the first training. When training the generative network model again, a structure of the generative network model is adjusted by using the two types of loss functions together, so that an image can be similar to a real sharp image at a pixel level, and appear more like a sharp image as a whole. The two loss functions can be joined with a weight parameter. The weight can be used to control the effect of the two types of loss functions on feedback regulation. The effect of the generative network model is to generate a sharp video from a blurry video, and the effect of the discriminative network model is to discriminate whether an inputted video frame is a real sharp image or a generated sharp image. Through the adversarial learning according to this embodiment of this application, the discriminative capability of the discriminative network model gets increasingly strong, and the video generated by the generative network model gets increasingly real. [0066] According to the above description, the N continuous image frames are first acquired from the video clip, the N image frames including the blurry image frame to be processed; then the 3D convolution processing is performed on the N image frames with the generative adversarial network model, to acquire the spatio-temporal information corresponding to the blurry image frame, the spatio-temporal information including: the spatial feature information of the blurry image frame, and the temporal feature information between the blurry image frame and the neighboring image frame of the N image frames. The deblurring processing is finally performed on the blurry image frame by using the spatio-temporal information corresponding to the blurry image frame through the generative adversarial network model, to output the sharp image frame. According to the embodiments of this application, the generative adversarial network model extracts the spatio-temporal information implicit between the continuous image frames through a 3D convolution operation, so that the deblurring processing on the blurry image frame is completed by using the spatio-temporal information corresponding to the blurry image frames through the generative adversarial network model. Therefore, a more real sharp image can be obtained, and the effect of video deblurring is enhanced. [0072] FIG. 3 is a schematic diagram of a training process of a generative network model and a discriminative network model according to an embodiment of this application. A discriminative network model (briefly referred to as discriminative network) and a generative network model (briefly referred to as generative network) are joined together to form an adversarial network. The two compete with each other. In order to acquire a more real deblurred video, an adversarial network structure is introduced when the generative network model shown in FIG. 2 is trained. The network structure in FIG. 2 is taken as a generator (namely, generative network model), and a discriminator (namely, discriminative network model) is added. During the training, the generative network is trained at first: inputting a blurry video frame into the generative network, to acquire a sharp video frame; acquiring a reconstruction loss function (namely, loss function 1 shown in FIG. 3), by comparing the sharp video frame to a real video frame; and adjusting a weight parameter of the generative network through the loss function. Next, the discriminative network is trained: inputting the real sharp video and the generated sharp video into the discriminative network, to acquire an adversarial loss function (namely, loss function 2 in FIG. 3); and adjusting the generative network structure through the adversarial loss function, to equip the discriminative network with a capability of discriminating between a real sharp image and a sharp image generated from a blurry image. The two network structures are alternately trained. When the generative network is trained later, a structure of the network model is adjusted by using the two types of loss functions together, so that an image can be similar to a real sharp image at a pixel level, and appear more like a sharp image as a whole. The two loss functions are joined with a weight parameter. The weight can be used to control the effect of the two types of loss functions on feedback regulation. The effect of the generative network is to generate a sharp video from a blurry video, and the effect of the discriminative network is to discriminate whether the inputted video frame is a real sharp image or a generated sharp video frame. Through the adversarial learning, the discriminative capability of the discriminative network gets increasingly strong, and the video generated by the generative network gets increasingly real. [0073] Next, weighted fusion of two different types of loss functions is illustrated with an example. Because two networks are used in this embodiment of this application, that is, a generative network and a discriminative network, two loss functions are used in this embodiment of this application, that is, a content loss function based on pixel-value differencing (that is a reconstruction loss function) and an adversarial loss function. [0074]-[0079]) Consider Claims 8 and 17. The combination of Zhang and Shenzen teaches: 8. The method according to claim 7, further comprising: calculating a second error loss between features of the (i+1)th frame of sample image and features of the super-resolution image of the (i+1)th frame of sample image using a perceptual loss function; and training the generative network based on the second error loss; the perceptual loss function being configured for constraining consistency between the (i+1)th frame of sample image and the super-resolution image of the (i+1)th frame of sample image in terms of eigenspace. / 17. The device according to claim 16, wherein the one or more processors are further configured to perform: calculating a second error loss between features of the (i+1)th frame of sample image and features of the super-resolution image of the (i+1)th frame of sample image using a perceptual loss function; and training the generative network based on the second error loss; the perceptual loss function being configured for constraining consistency between the (i+1)th frame of sample image and the super-resolution image of the (i+1)th frame of sample image in terms of eigenspace. (Examiner Note: Zhang teaches both a reconstruction loss and adversarial loss functions, wherein a perceptual loss function is analogous in scope to the reconstruction loss function and is obviated by these two functions; Zhang: [0044] After the generative network model outputs the sharp sample image frame, whether the outputted sharp sample image frame is blurry or sharp is discriminated according to the sharp sample image frame and the real sharp image frame by using the discriminative network model. An adversarial loss function is introduced by using the discriminative network model, to alternately train the generative network model and the discriminative network model a plurality of times, thereby better ensuring that a restored sharp video is more real. [0055] According to some embodiments of this application, step A4 that the electronic device trains the generative network model and the discriminative network model alternately according to the sharp sample image frame and the real sharp image frame includes the following steps. [0056] Step A41 includes acquiring a reconstruction loss function according to the sharp sample image frame and the real sharp image frame. [0057] Step A42 includes training the generative network model through the reconstruction loss function. [0058] Step A43 includes training the discriminative network model by using the real sharp image frame and the sharp sample image frame, to acquire an adversarial loss function outputted by the discriminative network model. [0059] Step A44 includes training the generative network model continually through the adversarial loss function. [0060] In order to acquire a more real deblurred video, the discriminative network model may be further introduced when the generative network model is trained. During the training, the generative network model is first trained: inputting a blurry sample image frame into the generative network model, to generate a sharp sample image frame; acquiring the reconstruction loss function by comparing the sharp sample image frame to the real sharp image frame; and adjusting a weight parameter of the generative network model through the reconstruction loss function. Next, the discriminative network model is trained: inputting the real sharp video and the generated sharp sample video into the discriminative network model, to acquire the adversarial loss function; and adjusting the generative network model through the adversarial loss function, to equip the discriminative network model with a capability of discriminating between the real sharp image and the sharp image generated from the blurry image frame. In this way, structures of the two network models are alternately trained. [0061] Further, according to some embodiments of this application, step A4 of training the generative network model and the discriminative network model alternately according to the sharp sample image frame and the real sharp image frame, in addition to including the foregoing step A41 to step A44, may further include the following steps. [0062] Step A45 includes reacquiring a reconstruction loss function through the generative network model, and reacquire an adversarial loss function through the discriminative network model, after the training the generative network model continually through the adversarial loss function. [0063] Step A46 includes performing weighted fusion on the reacquired reconstruction loss function and the reacquired adversarial loss function, to acquire a fused loss function. [0064] Step A47 includes training the generative network model continually through the fused loss function. [0065] After training the generative network model and the discriminative network model through the foregoing step A41 to step A44, step A45 to step A47 are performed based on the two network models after the first training. When training the generative network model again, a structure of the generative network model is adjusted by using the two types of loss functions together, so that an image can be similar to a real sharp image at a pixel level, and appear more like a sharp image as a whole. The two loss functions can be joined with a weight parameter. The weight can be used to control the effect of the two types of loss functions on feedback regulation. The effect of the generative network model is to generate a sharp video from a blurry video, and the effect of the discriminative network model is to discriminate whether an inputted video frame is a real sharp image or a generated sharp image. Through the adversarial learning according to this embodiment of this application, the discriminative capability of the discriminative network model gets increasingly strong, and the video generated by the generative network model gets increasingly real. [0066] According to the above description, the N continuous image frames are first acquired from the video clip, the N image frames including the blurry image frame to be processed; then the 3D convolution processing is performed on the N image frames with the generative adversarial network model, to acquire the spatio-temporal information corresponding to the blurry image frame, the spatio-temporal information including: the spatial feature information of the blurry image frame, and the temporal feature information between the blurry image frame and the neighboring image frame of the N image frames. The deblurring processing is finally performed on the blurry image frame by using the spatio-temporal information corresponding to the blurry image frame through the generative adversarial network model, to output the sharp image frame. According to the embodiments of this application, the generative adversarial network model extracts the spatio-temporal information implicit between the continuous image frames through a 3D convolution operation, so that the deblurring processing on the blurry image frame is completed by using the spatio-temporal information corresponding to the blurry image frames through the generative adversarial network model. Therefore, a more real sharp image can be obtained, and the effect of video deblurring is enhanced. [0072] FIG. 3 is a schematic diagram of a training process of a generative network model and a discriminative network model according to an embodiment of this application. A discriminative network model (briefly referred to as discriminative network) and a generative network model (briefly referred to as generative network) are joined together to form an adversarial network. The two compete with each other. In order to acquire a more real deblurred video, an adversarial network structure is introduced when the generative network model shown in FIG. 2 is trained. The network structure in FIG. 2 is taken as a generator (namely, generative network model), and a discriminator (namely, discriminative network model) is added. During the training, the generative network is trained at first: inputting a blurry video frame into the generative network, to acquire a sharp video frame; acquiring a reconstruction loss function (namely, loss function 1 shown in FIG. 3), by comparing the sharp video frame to a real video frame; and adjusting a weight parameter of the generative network through the loss function. Next, the discriminative network is trained: inputting the real sharp video and the generated sharp video into the discriminative network, to acquire an adversarial loss function (namely, loss function 2 in FIG. 3); and adjusting the generative network structure through the adversarial loss function, to equip the discriminative network with a capability of discriminating between a real sharp image and a sharp image generated from a blurry image. The two network structures are alternately trained. When the generative network is trained later, a structure of the network model is adjusted by using the two types of loss functions together, so that an image can be similar to a real sharp image at a pixel level, and appear more like a sharp image as a whole. The two loss functions are joined with a weight parameter. The weight can be used to control the effect of the two types of loss functions on feedback regulation. The effect of the generative network is to generate a sharp video from a blurry video, and the effect of the discriminative network is to discriminate whether the inputted video frame is a real sharp image or a generated sharp video frame. Through the adversarial learning, the discriminative capability of the discriminative network gets increasingly strong, and the video generated by the generative network gets increasingly real. [0073] Next, weighted fusion of two different types of loss functions is illustrated with an example. Because two networks are used in this embodiment of this application, that is, a generative network and a discriminative network, two loss functions are used in this embodiment of this application, that is, a content loss function based on pixel-value differencing (that is a reconstruction loss function) and an adversarial loss function. [0074]-[0079]) Consider Claims 9 and 18. The combination of Zhang and Shenzen teaches: 9. The method according to claim 8, further comprising: calculating a third error loss between the super-resolution image of the (i+1)th frame of sample image and the (i+1)th frame of sample image using a pixel loss function; and training the generative network based on the third error loss; the pixel loss function being configured for constraining consistency between the super-resolution image of the (i+1)th frame of sample image and the (i+1)th frame of sample image in terms of image content. / 18. The device according to claim 17, wherein the one or more processors are further configured to perform: calculating a third error loss between the super-resolution image of the (i+1)th frame of sample image and the (i+1)th frame of sample image using a pixel loss function; and training the generative network based on the third error loss; the pixel loss function being configured for constraining consistency between the super-resolution image of the (i+1)th frame of sample image and the (i+1)th frame of sample image in terms of image content. (Examiner Note: Zhang teaches both a reconstruction loss and adversarial loss functions, wherein the process for calculating these losses would inherently require the calculation of a pixel loss; Zhang: [0044] After the generative network model outputs the sharp sample image frame, whether the outputted sharp sample image frame is blurry or sharp is discriminated according to the sharp sample image frame and the real sharp image frame by using the discriminative network model. An adversarial loss function is introduced by using the discriminative network model, to alternately train the generative network model and the discriminative network model a plurality of times, thereby better ensuring that a restored sharp video is more real. [0055] According to some embodiments of this application, step A4 that the electronic device trains the generative network model and the discriminative network model alternately according to the sharp sample image frame and the real sharp image frame includes the following steps. [0056] Step A41 includes acquiring a reconstruction loss function according to the sharp sample image frame and the real sharp image frame. [0057] Step A42 includes training the generative network model through the reconstruction loss function. [0058] Step A43 includes training the discriminative network model by using the real sharp image frame and the sharp sample image frame, to acquire an adversarial loss function outputted by the discriminative network model. [0059] Step A44 includes training the generative network model continually through the adversarial loss function. [0060] In order to acquire a more real deblurred video, the discriminative network model may be further introduced when the generative network model is trained. During the training, the generative network model is first trained: inputting a blurry sample image frame into the generative network model, to generate a sharp sample image frame; acquiring the reconstruction loss function by comparing the sharp sample image frame to the real sharp image frame; and adjusting a weight parameter of the generative network model through the reconstruction loss function. Next, the discriminative network model is trained: inputting the real sharp video and the generated sharp sample video into the discriminative network model, to acquire the adversarial loss function; and adjusting the generative network model through the adversarial loss function, to equip the discriminative network model with a capability of discriminating between the real sharp image and the sharp image generated from the blurry image frame. In this way, structures of the two network models are alternately trained. [0061] Further, according to some embodiments of this application, step A4 of training the generative network model and the discriminative network model alternately according to the sharp sample image frame and the real sharp image frame, in addition to including the foregoing step A41 to step A44, may further include the following steps. [0062] Step A45 includes reacquiring a reconstruction loss function through the generative network model, and reacquire an adversarial loss function through the discriminative network model, after the training the generative network model continually through the adversarial loss function. [0063] Step A46 includes performing weighted fusion on the reacquired reconstruction loss function and the reacquired adversarial loss function, to acquire a fused loss function. [0064] Step A47 includes training the generative network model continually through the fused loss function. [0065] After training the generative network model and the discriminative network model through the foregoing step A41 to step A44, step A45 to step A47 are performed based on the two network models after the first training. When training the generative network model again, a structure of the generative network model is adjusted by using the two types of loss functions together, so that an image can be similar to a real sharp image at a pixel level, and appear more like a sharp image as a whole. The two loss functions can be joined with a weight parameter. The weight can be used to control the effect of the two types of loss functions on feedback regulation. The effect of the generative network model is to generate a sharp video from a blurry video, and the effect of the discriminative network model is to discriminate whether an inputted video frame is a real sharp image or a generated sharp image. Through the adversarial learning according to this embodiment of this application, the discriminative capability of the discriminative network model gets increasingly strong, and the video generated by the generative network model gets increasingly real. [0066] According to the above description, the N continuous image frames are first acquired from the video clip, the N image frames including the blurry image frame to be processed; then the 3D convolution processing is performed on the N image frames with the generative adversarial network model, to acquire the spatio-temporal information corresponding to the blurry image frame, the spatio-temporal information including: the spatial feature information of the blurry image frame, and the temporal feature information between the blurry image frame and the neighboring image frame of the N image frames. The deblurring processing is finally performed on the blurry image frame by using the spatio-temporal information corresponding to the blurry image frame through the generative adversarial network model, to output the sharp image frame. According to the embodiments of this application, the generative adversarial network model extracts the spatio-temporal information implicit between the continuous image frames through a 3D convolution operation, so that the deblurring processing on the blurry image frame is completed by using the spatio-temporal information corresponding to the blurry image frames through the generative adversarial network model. Therefore, a more real sharp image can be obtained, and the effect of video deblurring is enhanced. [0072] FIG. 3 is a schematic diagram of a training process of a generative network model and a discriminative network model according to an embodiment of this application. A discriminative network model (briefly referred to as discriminative network) and a generative network model (briefly referred to as generative network) are joined together to form an adversarial network. The two compete with each other. In order to acquire a more real deblurred video, an adversarial network structure is introduced when the generative network model shown in FIG. 2 is trained. The network structure in FIG. 2 is taken as a generator (namely, generative network model), and a discriminator (namely, discriminative network model) is added. During the training, the generative network is trained at first: inputting a blurry video frame into the generative network, to acquire a sharp video frame; acquiring a reconstruction loss function (namely, loss function 1 shown in FIG. 3), by comparing the sharp video frame to a real video frame; and adjusting a weight parameter of the generative network through the loss function. Next, the discriminative network is trained: inputting the real sharp video and the generated sharp video into the discriminative network, to acquire an adversarial loss function (namely, loss function 2 in FIG. 3); and adjusting the generative network structure through the adversarial loss function, to equip the discriminative network with a capability of discriminating between a real sharp image and a sharp image generated from a blurry image. The two network structures are alternately trained. When the generative network is trained later, a structure of the network model is adjusted by using the two types of loss functions together, so that an image can be similar to a real sharp image at a pixel level, and appear more like a sharp image as a whole. The two loss functions are joined with a weight parameter. The weight can be used to control the effect of the two types of loss functions on feedback regulation. The effect of the generative network is to generate a sharp video from a blurry video, and the effect of the discriminative network is to discriminate whether the inputted video frame is a real sharp image or a generated sharp video frame. Through the adversarial learning, the discriminative capability of the discriminative network gets increasingly strong, and the video generated by the generative network gets increasingly real. [0073] Next, weighted fusion of two different types of loss functions is illustrated with an example. Because two networks are used in this embodiment of this application, that is, a generative network and a discriminative network, two loss functions are used in this embodiment of this application, that is, a content loss function based on pixel-value differencing (that is a reconstruction loss function) and an adversarial loss function. [0074]-[0079]) Claims 6 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (US PGPub US2020/0372618A1, published November 26, 2020), hereby referred to as “Zhang”, in view of Shenzen et al. (CN Publication CN112184548A, published January 5, 2021), hereby referred to as “Shenzen” further in view of (US2023/0021463A1, filed on July 21, 2021) hereby referred to as “Chee”. Consider Claims 6 and 15. The combination of Zhang and Shenzen teaches: The method of Claim 5 and The device of Claim 14 as presented above. (Examiner Note: Zhang teaches both a reconstruction loss and adversarial loss functions, wherein an inter-frame stability loss function is analogous in scope to the adversarial loss function and is obviated by these two functions; Zhang: [0044] After the generative network model outputs the sharp sample image frame, whether the outputted sharp sample image frame is blurry or sharp is discriminated according to the sharp sample image frame and the real sharp image frame by using the discriminative network model. An adversarial loss function is introduced by using the discriminative network model, to alternately train the generative network model and the discriminative network model a plurality of times, thereby better ensuring that a restored sharp video is more real. [0055] According to some embodiments of this application, step A4 that the electronic device trains the generative network model and the discriminative network model alternately according to the sharp sample image frame and the real sharp image frame includes the following steps. [0056] Step A41 includes acquiring a reconstruction loss function according to the sharp sample image frame and the real sharp image frame. [0057] Step A42 includes training the generative network model through the reconstruction loss function. [0058] Step A43 includes training the discriminative network model by using the real sharp image frame and the sharp sample image frame, to acquire an adversarial loss function outputted by the discriminative network model. [0059] Step A44 includes training the generative network model continually through the adversarial loss function. [0060] In order to acquire a more real deblurred video, the discriminative network model may be further introduced when the generative network model is trained. During the training, the generative network model is first trained: inputting a blurry sample image frame into the generative network model, to generate a sharp sample image frame; acquiring the reconstruction loss function by comparing the sharp sample image frame to the real sharp image frame; and adjusting a weight parameter of the generative network model through the reconstruction loss function. Next, the discriminative network model is trained: inputting the real sharp video and the generated sharp sample video into the discriminative network model, to acquire the adversarial loss function; and adjusting the generative network model through the adversarial loss function, to equip the discriminative network model with a capability of discriminating between the real sharp image and the sharp image generated from the blurry image frame. In this way, structures of the two network models are alternately trained. [0061] Further, according to some embodiments of this application, step A4 of training the generative network model and the discriminative network model alternately according to the sharp sample image frame and the real sharp image frame, in addition to including the foregoing step A41 to step A44, may further include the following steps. [0062] Step A45 includes reacquiring a reconstruction loss function through the generative network model, and reacquire an adversarial loss function through the discriminative network model, after the training the generative network model continually through the adversarial loss function. [0063] Step A46 includes performing weighted fusion on the reacquired reconstruction loss function and the reacquired adversarial loss function, to acquire a fused loss function. [0064] Step A47 includes training the generative network model continually through the fused loss function. [0065] After training the generative network model and the discriminative network model through the foregoing step A41 to step A44, step A45 to step A47 are performed based on the two network models after the first training. When training the generative network model again, a structure of the generative network model is adjusted by using the two types of loss functions together, so that an image can be similar to a real sharp image at a pixel level, and appear more like a sharp image as a whole. The two loss functions can be joined with a weight parameter. The weight can be used to control the effect of the two types of loss functions on feedback regulation. The effect of the generative network model is to generate a sharp video from a blurry video, and the effect of the discriminative network model is to discriminate whether an inputted video frame is a real sharp image or a generated sharp image. Through the adversarial learning according to this embodiment of this application, the discriminative capability of the discriminative network model gets increasingly strong, and the video generated by the generative network model gets increasingly real. [0066] According to the above description, the N continuous image frames are first acquired from the video clip, the N image frames including the blurry image frame to be processed; then the 3D convolution processing is performed on the N image frames with the generative adversarial network model, to acquire the spatio-temporal information corresponding to the blurry image frame, the spatio-temporal information including: the spatial feature information of the blurry image frame, and the temporal feature information between the blurry image frame and the neighboring image frame of the N image frames. The deblurring processing is finally performed on the blurry image frame by using the spatio-temporal information corresponding to the blurry image frame through the generative adversarial network model, to output the sharp image frame. According to the embodiments of this application, the generative adversarial network model extracts the spatio-temporal information implicit between the continuous image frames through a 3D convolution operation, so that the deblurring processing on the blurry image frame is completed by using the spatio-temporal information corresponding to the blurry image frames through the generative adversarial network model. Therefore, a more real sharp image can be obtained, and the effect of video deblurring is enhanced. [0072] FIG. 3 is a schematic diagram of a training process of a generative network model and a discriminative network model according to an embodiment of this application. A discriminative network model (briefly referred to as discriminative network) and a generative network model (briefly referred to as generative network) are joined together to form an adversarial network. The two compete with each other. In order to acquire a more real deblurred video, an adversarial network structure is introduced when the generative network model shown in FIG. 2 is trained. The network structure in FIG. 2 is taken as a generator (namely, generative network model), and a discriminator (namely, discriminative network model) is added. During the training, the generative network is trained at first: inputting a blurry video frame into the generative network, to acquire a sharp video frame; acquiring a reconstruction loss function (namely, loss function 1 shown in FIG. 3), by comparing the sharp video frame to a real video frame; and adjusting a weight parameter of the generative network through the loss function. Next, the discriminative network is trained: inputting the real sharp video and the generated sharp video into the discriminative network, to acquire an adversarial loss function (namely, loss function 2 in FIG. 3); and adjusting the generative network structure through the adversarial loss function, to equip the discriminative network with a capability of discriminating between a real sharp image and a sharp image generated from a blurry image. The two network structures are alternately trained. When the generative network is trained later, a structure of the network model is adjusted by using the two types of loss functions together, so that an image can be similar to a real sharp image at a pixel level, and appear more like a sharp image as a whole. The two loss functions are joined with a weight parameter. The weight can be used to control the effect of the two types of loss functions on feedback regulation. The effect of the generative network is to generate a sharp video from a blurry video, and the effect of the discriminative network is to discriminate whether the inputted video frame is a real sharp image or a generated sharp video frame. Through the adversarial learning, the discriminative capability of the discriminative network gets increasingly strong, and the video generated by the generative network gets increasingly real. [0073] Next, weighted fusion of two different types of loss functions is illustrated with an example. Because two networks are used in this embodiment of this application, that is, a generative network and a discriminative network, two loss functions are used in this embodiment of this application, that is, a content loss function based on pixel-value differencing (that is a reconstruction loss function) and an adversarial loss function. [0074]-[0079]) Even if The combination of Zhang and Shenzen does not specifically teach limitations from claims 6 and 15 for: “optical flow” Chee teaches: The method of Claim and The device of Claim 14 (Chee: abstract, The present invention discloses a multi-frame image super resolution system that utilizes both deep learning models and traditional models of enhancing the resolution of an image so that minimal computational resources are used. A frame alignment module of the invention aligns the frames of the image after which a processing module configured within the system process the Y and the UV channels of the image by using multiple deep and traditional resolution enhancement models. A merging unit merges the output of the processors to produce a super resolution image incorporating the advantages of both of the image enhancement methods.) -; 6. The method according to claim 5, wherein calculating the inter-frame stability loss between the first change and the second change comprises: / 15. The device according to claim 14, wherein the one or more processors are further configured to perform: (Chee: [0032] FIG. 1A illustrates a system for enhancing resolution of an image by combining a number of traditional models with deep learning models. The proposed multi-frame resolution enhancement system 100 integrates traditional super resolution methods and light-weighted deep learning models such that the minimal computational time is used. Before the actual processing of the image, a number of frames of the image are aligned. It is the responsibility of the frame alignment module 102 of the system 100 to align the multiple frames of the image. Instead of considering just a single frame, multiple frames are considered for alignment by the frame alignment module 102 due to several reasons. [0061] FIG. 3B illustrates merging of traditional and deep learning results in Y channel. As per mentioned earlier, Y channel consists of a lot of high frequency information and thus, requires better enhancement approaches to ensure the visual quality of the final image. In this part, we further split the processing steps into two branches. The first branch consists of a lightweight deep learning model which is trained to super-resolve, de-noise and de-blur the given frames. Deep learning models such as CARN and FSRCNN are used. [0062] In order to conform to the actual use case, changes are made to the data preparation such that it includes real noise patterns. Furthermore, additional loss functions are introduced during the model training for detail enhancements. However, considering the restriction in computation complexity, there is a limit to the model performance. More specifically, when the losses are designed such that details are emphasized, its de-noising capability would be affected.) -; 6. calculating a first optical flow of the first change using an optical flow network; calculating a second optical flow of the second change using the optical flow network; / 15. calculating a first optical flow of the first change using an optical flow network; (Chee: [0033] Under frame alignment, the main step is to find similar pixels in each frame. Using these pixels, relationship between all frames with respect to one or more reference frames is calculated. Examples of structures representing the correspondences include, but are not limited tohomography matrix, optical flow field and block matching. Each has their own pros and cons as there is a trade-off between computation complexity and precision. [0034] Homography matrix is a 3×3 matrix with 8 degree of freedom that relates the transformation between two images of the same planer surface in the space. It measures the translation, rotation and scaling between the two images in the 3D space. Optical flow field is a vector field between two images that shows how each pixel in the first image can be moved to form the second image. In other words, it finds the correspondence between pixels of the two images. Block matching represents a set of vectors which indicates the matching blocks between two images. The images are first divided into blocks and the similarity between the blocks of the two images is calculated. The resulting vector for each block shows its movement of from the first image to the second.) -; 6. and substituting the first optical flow and the second optical flow into the inter-frame stability loss function, to calculate the inter-frame stability loss. / 15. calculating a second optical flow of the second change using the optical flow network; and substituting the first optical flow and the second optical flow into the inter-frame stability loss function, to calculate the inter-frame stability loss. (Chee: [0052] In the image frame alignment process, at least one image frame needs to be selected as the reference frame for the alignment process, and other image frames and the reference frame itself are aligned to the reference frame. Examples of structures representing the correspondences include, but are not limited to, homography matrix, optical flow field and block matching. Each has their own pros and cons as there is a trade-off between computation complexity and precision. [0053] Homography matrix is a 3×3 matrix with 8 degree of freedom that relates the transformation between two images of the same planer surface in the space. It measures the translation, rotation and scaling between the two images in the 3D space. Optical flow field is a vector field between two images that shows how each pixel in the first image can be moved to form the second image. In other words, it finds the correspondence between pixels of the two images. Block matching represents a set of vectors which indicate the matching blocks between two images. The images are first divided into blocks and the similarity between the blocks of the two images is calculated. The resulting vector for each block shows its movement from the first image to the second.) It would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to modify the combination of Zhang and Shenzen with the super-resolution image processing method and system with the teachings of Chee for multi-frame super resolution image processing as they are all directed towards the same field of endeavor. The determination of obviousness is predicated upon the following findings: Both are directed towards image processing methods and systems; One skilled in the art would have been motivated to modify the combination of Zhang and Shenzen for processing multi-frame image data using a GAN to process the image data with the teachings of Chee in order to improved algorithm for super-resolution image processing in order to take into account optical flow fields in the image processing and reconstruction techniques. Furthermore, the prior art collectively includes each element claimed (though not all in the same reference), and one of ordinary skill in the art could have combined the elements in the manner explained above using known engineering design, interface and/or programming techniques, without changing a “fundamental” operating principle of Zhang, while the teaching of Shenzen continues to perform the same function as originally taught prior to being combined, in order to produce the repeatable and predictable result of take into account optical flow fields in the image processing and reconstruction techniques. It is for at least the aforementioned reasons that the examiner has reached a conclusion of obviousness with respect to the claim in question. Conclusion The prior art made of record in form PTO-892 and not relied upon is considered pertinent to applicant's disclosure. PNG media_image1.png 247 892 media_image1.png Greyscale Any inquiry concerning this communication or earlier communications from the examiner should be directed to TAHMINA ANSARI whose telephone number is 571-270-3379. The examiner can normally be reached on IFP Flex - Monday through Friday 9 to 5. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, O’NEAL MISTRY can be reached on 313-446-4912. The fax phone numbers for the organization where this application or proceeding is assigned are 571-273-8300 for regular communications and 571-273-8300 for After Final communications. TC 2600’s customer service number is 571-272-2600. Any inquiry of a general nature or relating to the status of this application or proceeding should be directed to the receptionist whose telephone number is 571-272-2600. 2674 /Tahmina Ansari/ September 3, 2026 /TAHMINA N ANSARI/Primary Examiner, Art Unit 2674
Read full office action

Prosecution Timeline

Oct 22, 2024
Application Filed
Sep 09, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749336
SYSTEM, METHOD, AND COMPUTER PROGRAM PRODUCT FOR SENSITIVE DATA OBFUSCATION ON A CAPTURED CREDENTIAL
4y 0m to grant Granted Sep 29, 2026
Patent 12738040
METHOD, SYSTEM AND STORAGE MEDIUM IMPLEMENTING CROSS-MODAL CONTRASTIVE LEARNING TO IMPROVE ITEM CATEGORIZATION BERT MODEL
3y 11m to grant Granted Sep 15, 2026
Patent 12725400
METHOD AND SYSTEM FOR TRAINING A MACHINE LEARNING MODEL WITH A SUBCLASS OF ONE OR MORE PREDEFINED CLASSES OF VISUAL OBJECTS
3y 0m to grant Granted Sep 01, 2026
Patent 12725846
METHOD AND DEVICE FOR DETECTING ABNORMALITY IN BATTERY CELL TYPE ELECTRODES
2y 7m to grant Granted Sep 01, 2026
Patent 12714881
METHODS, SYSTEMS AND COMPUTER READABLE MEDIUMS FOR DETERMINING A REGION-OF-INTEREST IN SURFACE-GUIDED MONITORING
4y 1m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
85%
Grant Probability
99%
With Interview (+18.6%)
2y 6m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 902 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month