DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claims 2, 11-12, and 15 are objected to because of the following informalities:
Regarding claim 2, the limitation “in which a plurality of nodes in the third learning model has a weight value that is determined based on the first control parameter”, should be corrected to “in which a plurality of nodes in the third learning model have a weight value that is determined based on the first control parameter”.
Regarding claim 2, the limitation “a weight value of a corresponding node in the first learning model and a weight value of a corresponding node corresponding in the second learning model.”, should be corrected to “a weight value of a corresponding node in the first learning model and a weight value of a corresponding node
Regarding claim 11, the limitation “A controlling method of an electronic apparatus, the method comprising:”, should be corrected to “A controlling an electronic apparatus, the method comprising:”.
Regarding claim 12, the limitation “a weight value of a corresponding node in the first learning model and a weight value of a corresponding node corresponding in the second learning model.”, should be corrected to “a weight value of a corresponding node in the first learning model and a weight value of a corresponding node in the second learning model.”.
Regarding claim 15, the limitation “A non-transitory computer-readable recording medium storing a program for executing a controlling method of an electronic apparatus, the method comprising:”, should be corrected to “A non-transitory computer-readable recording medium storing a program for executing a controlling an electronic apparatus, the method comprising:”.
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 2-10 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 2 recites the limitation "The device as claimed in claim 1". There is insufficient antecedent basis for this limitation in the claim. For the purposes of examination, the limitation is interpreted as “The electronic apparatus as claimed in claim 1”.
As per claim(s) 3-10, arguments made in rejecting claim(s) 2 are analogous.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-5, 7-9, and 11-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ahn et al. (A Fast 4K Video Frame Interpolation Using a Multi-Scale Optical Flow Reconstruction Network) hereinafter referenced as Ahn, in view of Kong et al. (ClassSR: A General Framework to Accelerate Super-Resolution Networks by Data Characteristic) hereinafter referenced as Kong, and Wang et al. (Deep Network Interpolation for Continuous Imagery Effect Transition) hereinafter referenced as Wang
Regarding claim 1, Ahn discloses: An electronic apparatus comprising: memory storing a first learning model, and configured to estimate a motion between two frames (Ahn: Abstract: “However, these methods demand huge amounts of memory and run time for high-resolution videos, and are unable to process a 4K frame in a single pass. In this paper, we propose a fast 4K video frame interpolation method, based upon a multi-scale optical flow reconstruction scheme.”;
Section: 2. Proposed Method: “The entire network is composed of an optical flow estimation (OFE) network and two multi-scale optical flow reconstruction (OFR) networks… The OFE network predicts the optical flow map in low resolution, and OFR networks reconstruct the map in a higher resolution which is the original size of the input frames.”);
at least one processor, comprising processing circuitry, individually and/or (The limitation’s use of “and/or” is interpreted as disjunctive “or”, thus indicating that only one limitation is required.) collectively configured to obtain a first frame included in an input image and a second frame which is a previous frame of the first frame, and generate an interpolation frame using the obtained first frame and second frame (Ahn: Abstract: “In this paper, we propose a fast 4K video frame interpolation method, based upon a multi-scale optical flow reconstruction scheme.”;
Section: 1. Introduction: “In this paper, we propose a novel fast 4K video frame interpolation method using a multi-scale motion reconstruction network. The proposed network is composed of three sub-networks: An optical flow estimation (OFE) network and two multi-scale optical flow reconstruction (OFR) networks. The OFE network predicts bi-directional optical flow in quarter resolution of input frames.”; Wherein the prediction of bi-directional optical flow for the input images constitutes the presence of an input and a previous image.); and
wherein the at least one processor is individually and/or (The limitation’s use of “and/or” is interpreted as disjunctive “or”, thus indicating that only one limitation is required.) collectively configured to estimate a motion between the first frame and the second frame using the first learning model, and generate the interpolation frame based on the estimated motion (Ahn: Section: 1. Introduction: “In this paper, we propose a novel fast 4K video frame interpolation method using a multi-scale motion reconstruction network. The proposed network is composed of three sub-networks: An optical flow estimation (OFE) network and two multi-scale optical flow reconstruction (OFR) networks. The OFE network predicts bi-directional optical flow in quarter resolution of input frames. The OFR networks reconstruct the intermediate optical flow into half and original resolution, respectively.”; Wherein the prediction of bi-directional optical flow for the input images constitutes the presence of an input and a previous image.).
Ahn does not disclose expressly: memory storing first and second learning models, each with the same network structure, and configured to estimate a motion between two frames; wherein the first learning model is a model configured to be trained with image data having a first characteristic;
wherein the second learning model is a model configured to be trained with image data having a second characteristic which is opposite to the first characteristic;.
Kong discloses: memory storing first and second learning models, and configured to estimate a high resolution image (Kong: Abstract: “we propose a new solution pipeline – ClassSR that combines classification and SR in a unified framework. In particular, it first uses a Class-Module to classify the subimages into different classes according to restoration difficulties, then applies an SR-Module to perform SR for different classes. The Class-Module is a conventional classification network, while the SR-Module is a network container that consists of the to-be-accelerated SR network and its simplified versions.”;
3.4. SRModule: “The SR-Module is designed as a container that consists of several independent branches {fjSR}Mj =1. In general, each branch can be any learning-based SR network. As our goal is to accelerate an existing SR method (e.g., FSRCNN, CARN), we adopt this SR network as the base network, and set it as the most complex branch fMSR. The other branches are obtained by reducing the network complexity of fMSR.”); wherein the first learning model is a model configured to be trained with image data having a first characteristic; and wherein the second learning model is a model configured to be trained with image data having a second characteristic which is opposite to the first characteristic (Kong: Section: 3.7. Training Strategy: “To pre-train the SR-Module, we use the data classified by the PSNR values. Specifically, all sub-images are passed through a well-trained MSRResNet. Then these sub-images are ranked according to their PSNR values. Next, the first 1/3 sub-images are assigned to the hard class, while the last 1/3 belong to the simple class, just as in Sec. 3.1. Then we train the simple/medium/complex SR branch on the corresponding simple/medium/hard data.”).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to implement an image classifier and branch network for the processing of images based on their classifications as taught by Kong for the generation of the reconstruct high resolution optical flow disclosed by Ahn by processing the input images based on a multi-scale motion reconstruction network branch selected based on a classification. The suggestion/motivation for doing so would have been “The key idea is using a Class-Module to classify the sub-images into different classes (e.g., “simple, medium, hard”), each class corresponds to different processing branches with different network capacity. Extensive experiments well demonstrate that ClassSR can accelerate most existing methods on different datasets.” (Kong: Section: 5. Conclusion). Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Ahn in view of Kong does not disclose expressly: An electronic apparatus comprising: memory storing first and second learning models, each with the same network structure, and configured to estimate a motion between two frames; and wherein the at least one processor is individually and/or collectively configured to generate a third learning model using a first control parameter and the first and second learning models, estimate a motion between the first frame and the second frame using the generated third learning model, and generate the interpolation frame based on the estimated motion.
Wang discloses: An electronic apparatus comprising: memory storing first and second learning models, each with the same network structure (Wang: Section: 3.1. Deep Network Interpolation: “Consider two networks GA and GB with the same structure, achieving different effects A and B, respectively. The networks consist of common operations such as convolution, up/down-sampling and non-linear activation…Our aim is to achieve a continuous transition between the effects A and B. We do so by the proposed Deep Network Interpolation (DNI)…It is worth noticing that the choice of the network structure for DNI is flexible, as long as the structures of models to be interpolated are kept the same.”); and wherein at least one processor is individually and/or collectively configured to generate a third learning model using a first control parameter and the first and second learning models, and estimate an image output using the generated third learning model (Wang: Section: 1. Introduction: “In this paper, we address these drawbacks by introducing a more general, simple but effective approach, known as Deep Network Interpolation (DNI). Continuous imagery effect transition is achieved via linear interpolation in the parameter space of existing trained networks. Specifically, provided with a model for a particular effect A, we fine-tune it to realize another relevant effect B. DNI applies linear interpolation for all the corresponding parameters of these two deep networks…Performing feed-forward operations on these interpolated models using the same input allows us to outputs with a continuous transition between the different effects A and B.”
Section: 3.1. Deep Network Interpolation: “Our aim is to achieve a continuous transition between the effects A and B. We do so by the proposed Deep Network Interpolation (DNI). DNI interpolates all the corresponding parameters of these two models to derive a new interpolated model Ginterp, whose parameters are:
PNG
media_image1.png
58
436
media_image1.png
Greyscale
where α ∈ [0,1] is the interpolation coefficient. Indeed, it is a linear interpolation of the two parameter vectors θA and θB. The interpolation coefficient α controls a balance of the effect A and B. By smoothly sliding α, we achieve continuous transition effects without abrupt changes.”).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to incorporate the methods for performing Deep Network Interpolation taught by Wang to interpolate the multi-scale motion reconstruction network branch models disclosed by Ahn in view of Kong. The suggestion/motivation for doing so would have been “With extensive experiments on super-resolution, denoising, image-to-image translation and style transfer, we demonstrate that the proposed method is applicable for a wide range of low-level vision tasks despite its simplicity. Compared with existing methods that achieve continuous transition by task-specific designs, our method is easy to generalize with negligible computational overhead.” (Wang: Section: 5. Conclusion). Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Ahn in view of Kong with Wang to obtain the invention as specified in claim 1.
Regarding claim 2, Ahn in view of Kong and Wang discloses: The device as claimed in claim 1 (Claim limitation is interpreted according to the rejection of claim 2 under 35 U.S.C. 112(b) as disclosed above), wherein the at least one processor is individually and/or (The limitation’s use of “and/or” is interpreted as disjunctive “or”, thus indicating that only one limitation is required.) collectively configured to generate the third learning model which has the same network structure as the first and second learning models, in which a plurality of nodes in the third learning model has a weight value that is determined based on the first control parameter, a weight value of a corresponding node in the first learning model and a weight value of a corresponding node corresponding in the second learning model (Wang: Section: 3.1. Deep Network Interpolation: “Consider two networks GA and GB with the same structure, achieving different effects A and B, respectively…The parameters in CNNs are mainly the weights of convolutional layers, called filters, filtering the input image or the precedent features. We assume that their parameters θA and θB have a “strong correlation” with each other, i.e., the filter orders and filter patterns in the same position of GA and GB are similar…
Our aim is to achieve a continuous transition between the effects A and B. We do so by the proposed Deep Network Interpolation (DNI). DNI interpolates all the corresponding parameters of these two models to derive a new interpolated model Ginterp, whose parameters are:
PNG
media_image1.png
58
436
media_image1.png
Greyscale
where α ∈ [0,1] is the interpolation coefficient. Indeed, it is a linear interpolation of the two parameter vectors θA and θB. The interpolation coefficient α controls a balance of the effect A and B. By smoothly sliding α, we achieve continuous transition effects without abrupt changes.”).
Regarding claim 3, Ahn in view of Kong and Wang discloses: The device as claimed in claim 2 (Claim limitation is interpreted according to the rejection of claim 3 under 35 U.S.C. 112(b) as disclosed above), wherein the first control parameter has a value between 0 and 1; and wherein the at least one processor is individually and/or (The limitation’s use of “and/or” is interpreted as disjunctive “or”, thus indicating that only one limitation is required.) collectively configured to generate the third learning model having a weight value of each of a plurality of nodes in the third learning model as a sum of a value obtained at least by multiplying a weight value of a corresponding node in the second learning model by 1 minus the first control parameter and a value obtained by multiplying a weight value of a corresponding node in the first learning model by the first control parameter (Wang: Section: 3.1. Deep Network Interpolation: “DNI interpolates all the corresponding parameters of these two models to derive a new interpolated model Ginterp, whose parameters are:
PNG
media_image1.png
58
436
media_image1.png
Greyscale
where α ∈ [0,1] is the interpolation coefficient. Indeed, it is a linear interpolation of the two parameter vectors θA and θB. The interpolation coefficient α controls a balance of the effect A and B. By smoothly sliding α, we achieve continuous transition effects without abrupt changes.”).
Regarding claim 4, Ahn in view of Kong and Wang discloses: The device as claimed in claim 1 (Claim limitation is interpreted according to the rejection of claim 4 under 35 U.S.C. 112(b) as disclosed above), wherein the at least one processor is individually and/or (The limitation’s use of “and/or” is interpreted as disjunctive “or”, thus indicating that only one limitation is required.) collectively configured to generate an interpolation frame using the estimated motion (Ahn: 2. Proposed Method: “Figure 2 shows the architecture of the proposed video frame interpolation network. Our method produces the interpolated frame in a pixel blending manner with warped frames using an optical flow map.”).
Ahn in view of Kong and Wang does not disclose expressly: wherein the memory is configured to store a fourth learning model trained to generate an interpolation frame having a third characteristic based on an estimated motion and a fifth learning model which has the same network structure as the fourth learning model and is trained to generate an interpolation frame having a fourth characteristic which is opposite to the third characteristic; and wherein the at least one processor is individually or collectively configured to generate a sixth learning model using a second control parameter, the fourth learning model and the fifth learning model, and generate an interpolation frame using the estimated motion and the sixth learning model.
Wang further discloses: a fourth learning model trained to generate a frame having a third characteristic and a fifth learning model which has the same network structure as the fourth learning model and is trained to generate a frame having a fourth characteristic which is opposite to the third characteristic (Wang: Section: 4.1. Image Restoration: “Balance MSE and GAN effects in super-resolution. A super-resolution model trained with MSE loss [35] tends to produce oversmooth images. We fine-tune it with GAN loss and perceptual loss [20], obtaining results with vivid details but always together with unpleasant artifacts (e.g., the eaves and water waves in Fig. 5).”;
Figure 5: “Balancing the MSE and GAN effects with DNI in super-resolution. The MSE effect is over-smooth while the GAN effect is always accompanied with unpleasant artifacts (e.g., the eaves and water waves). DNI allows smooth transition from one effect to the other and produces visually-pleasing results with largely reduced artifacts while maintaining the textures. In contrast, the pixel interpolation strategy fails to separate the artifacts and textures.”); and generate a sixth learning model using a second control parameter, the fourth learning model and the fifth learning model, and generate a frame using the sixth learning model (Wang: Section: 3.1. Deep Network Interpolation: “DNI interpolates all the corresponding parameters of these two models to derive a new interpolated model Ginterp, whose parameters are:
PNG
media_image1.png
58
436
media_image1.png
Greyscale
where α ∈ [0,1] is the interpolation coefficient. Indeed, it is a linear interpolation of the two parameter vectors θA and θB. The interpolation coefficient α controls a balance of the effect A and B. By smoothly sliding α, we achieve continuous transition effects without abrupt changes.”;
Section: 4.1. Image Restoration: “As presented in Fig. 5, DNI is able to smoothly alter the outputs from the MSE effect to the GAN effect. With appropriate interpolation coefficient, it produces visually pleasing results with largely reduced artifacts while maintaining the textures.”).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to incorporate the super-resolution model as further taught by Wang for the super-resolution processing of the interpolated image disclosed by Ahn in view of Kong and Wang. The suggestion/motivation for doing so would have been “Balancing the MSE and GAN effects with DNI in super-resolution. The MSE effect is over-smooth while the GAN effect is always accompanied with unpleasant artifacts (e.g., the eaves and water waves). DNI allows smooth transition from one effect to the other and produces visually-pleasing results with largely reduced artifacts while maintaining the textures.” (Wang: Figure 5). Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Ahn in view of Kong and Wang with the further teaching of Wang to obtain the invention as specified in claim 4.
Regarding claim 5, Ahn in view of Kong and Wang discloses: The device as claimed in claim 4 (Claim limitation is interpreted according to the rejection of claim 5 under 35 U.S.C. 112(b) as disclosed above), wherein the at least one processor is individually and/or (The limitation’s use of “and/or” is interpreted as disjunctive “or”, thus indicating that only one limitation is required.) collectively configured to generate the sixth learning model which has the same network structure as the fourth and fifth learning models, in which a weight value of a plurality of nodes in the sixth learning model is determined based on the second control parameter, a weight value of a corresponding node in the fourth learning model and a weight value of a corresponding node corresponding in the fifth learning model (Wang: Section: 3.1. Deep Network Interpolation: “DNI interpolates all the corresponding parameters of these two models to derive a new interpolated model Ginterp, whose parameters are:
PNG
media_image1.png
58
436
media_image1.png
Greyscale
where α ∈ [0,1] is the interpolation coefficient. Indeed, it is a linear interpolation of the two parameter vectors θA and θB. The interpolation coefficient α controls a balance of the effect A and B. By smoothly sliding α, we achieve continuous transition effects without abrupt changes.”;
Section: 4.1. Image Restoration: “As presented in Fig. 5, DNI is able to smoothly alter the outputs from the MSE effect to the GAN effect. With appropriate interpolation coefficient, it produces visually pleasing results with largely reduced artifacts while maintaining the textures.”).
Regarding claim 7, Ahn in view of Kong and Wang discloses: The device as claimed in claim 1 (Claim limitation is interpreted according to the rejection of claim 7 under 35 U.S.C. 112(b) as disclosed above).
Ahn in view of Kong and Wang does not disclose expressly: further comprising: an input/output interface, comprising circuitry, configured to input the first control parameter, wherein the at least one processor is individually and/or (The limitation’s use of “and/or” is interpreted as disjunctive “or”, thus indicating that only one limitation is required.) collectively configured to generate the third learning model using the input first parameter.
Wang further discloses: an input/output interface, comprising circuitry, configured to input a first control parameter, wherein the at least one processor is individually or collectively configured to generate a third learning model using the input first parameter (Wang: Section: 4.1. Image Restoration: “most popular image editing softwares (e.g., Photoshop) have controllable options for each tool. For example, the noise reduction tool comes with sliding bars for controlling the denoising strength and the percentage of preserving or sharpening details. We show an example to illustrate the importance of adjustable denoising strength…
our proposed DNI is able to achieve adjustable denoising strength by simply tweaking the interpolation coefficient α of different denoising models for N20, N40 and N60. For the grass, a weaker denoising strength could preserve more details while in the sky region, the stronger denoising strength could obtain an artifact-free result (red frames in Fig. 6). This example demonstrates the flexibility of DNI to customize restoration results based on the task at hand and the specific user preference.”).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to incorporate the interface for adjusting the interpolation coefficient as further taught by Wang for the adjusting of the interpolation coefficient disclosed by Ahn in view of Kong and Wang. The suggestion/motivation for doing so would have been “our proposed DNI is able to achieve adjustable denoising strength by simply tweaking the interpolation coefficient α of different denoising models for N20, N40 and N60. For the grass, a weaker denoising strength could preserve more details while in the sky region, the stronger denoising strength could obtain an artifact-free result (red frames in Fig. 6). This example demonstrates the flexibility of DNI to customize restoration results based on the task at hand and the specific user preference.” (Wang: Section: 4.1. Image Restoration). Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Ahn in view of Kong and Wang with the further teaching of Wang to obtain the invention as specified in claim 7.
Regarding claim 8, Ahn in view of Kong and Wang discloses: The device as claimed in claim 1 (Claim limitation is interpreted according to the rejection of claim 8 under 35 U.S.C. 112(b) as disclosed above), wherein the at least one processor is individually and/or (The limitation’s use of “and/or” is interpreted as disjunctive “or”, thus indicating that only one limitation is required.) collectively configured to identify an image characteristic of the input image (Kong: Abstract: “we propose a new solution pipeline – ClassSR that combines classification and SR in a unified framework. In particular, it first uses a Class-Module to classify the subimages into different classes according to restoration difficulties, then applies an SR-Module to perform SR for different classes.”).
Ahn in view of Kong and Wang does not disclose expressly: wherein the at least one processor is individually or collectively configured to determine the first control parameter corresponding to the identified image characteristic, and generate the third learning model using the determined first control parameter.
Wang further discloses: identifying an image characteristic of an input image, determining a first control parameter corresponding to the identified image characteristic, and generating a third learning model using the determined first control parameter (Wang: Section: 3.2. Understanding Network Interpolation: “we perform linear interpolation between the filters from the N20 and N60 models. With optimal coefficients α, the interpolated filters could visually fit those learned filters (2nd row with red frames, Fig. 3). We further calculate the correlation index for each interpolated filter with the first N20 filter. The correlation curves for learned and interpolated filters are also very close. The optimal α is obtained through the final performance of the interpolated network. Specifically, we perform DNI with α from 0 to 1 with an interval of 0.05. The best α for each noise level is selected based on which that makes the interpolated network to produce the highest PSNR on the test dataset.”; Wherein the optimal interpolation coefficient for each image class is determined and stored based on a test dataset PSNR score).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to incorporate the algorithms for determining an interpolation coefficient for an image class as further taught by Wang for the determination of the optimal interpolation coefficients for each image class disclosed by Ahn in view of Kong and Wang. The suggestion/motivation for doing so would have been “The continuous changes of learned filters suggest that it possible to obtain the intermediate filters by interpolating the two ends…With optimal coefficients α, the interpolated filters could visually fit those learned filters (2nd row with red frames, Fig. 3)…The best α for each noise level is selected based on which that makes the interpolated network to produce the highest PSNR on the test dataset.” (Wang: Section: 3.2. Understanding Network Interpolation). Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Ahn in view of Kong and Wang with the further teaching of Wang to obtain the invention as specified in claim 8.
Regarding claim 9, Ahn in view of Kong and Wang discloses: The device as claimed in claim 1 (Claim limitation is interpreted according to the rejection of claim 9 under 35 U.S.C. 112(b) as disclosed above), wherein the at least one processor is individually and/or (The limitation’s use of “and/or” is interpreted as disjunctive “or”, thus indicating that only one limitation is required.) collectively configured to identify an image characteristic of the generated interpolation frame (Ahn: Section: 3. Experimental Results: “To evaluate and compare the performance of the proposed method, we choose Ultra Video and SJTU Media datasets, because they have 4K image resolution and are suitable for our target. In addition, for various dataset comparisons to the state-of-art methods, we also consider the Vimeo dataset with the same experimental condition introduced in [15]… We report PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity) [30], which are often used to evaluate the performance of video frame interpolation algorithms. We also compare the inference time to show the effectiveness of the proposed method.”).
Ahn in view of Kong and Wang does not disclose expressly: wherein the at least one processor is individually or collectively configured to update the first control parameter to a parameter corresponding to the identified image characteristic.
Wang further discloses: identifying an image characteristic of the generated frame, and update the first control parameter to a parameter corresponding to the identified image characteristic (Wang: Section: 3.2. Understanding Network Interpolation: “we perform linear interpolation between the filters from the N20 and N60 models. With optimal coefficients α, the interpolated filters could visually fit those learned filters (2nd row with red frames, Fig. 3). We further calculate the correlation index for each interpolated filter with the first N20 filter. The correlation curves for learned and interpolated filters are also very close. The optimal α is obtained through the final performance of the interpolated network. Specifically, we perform DNI with α from 0 to 1 with an interval of 0.05. The best α for each noise level is selected based on which that makes the interpolated network to produce the highest PSNR on the test dataset.”; Wherein the optimal interpolation coefficient for each image class is determined based on an iterative analysis of the generated images compared to the test dataset PSNR score).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to incorporate the algorithms for determining an interpolation coefficient for an image class as further taught by Wang for the determination of the optimal interpolation coefficients for each image class disclosed by Ahn in view of Kong and Wang. The suggestion/motivation for doing so would have been “The continuous changes of learned filters suggest that it it possible to obtain the intermediate filters by interpolating the two ends…With optimal coefficients α, the interpolated filters could visually fit those learned filters (2nd row with red frames, Fig. 3)…The best α for each noise level is selected based on which that makes the interpolated network to produce the highest PSNR on the test dataset.” (Wang: Section: 3.2. Understanding Network Interpolation). Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Ahn in view of Kong and Wang with the further teaching of Wang to obtain the invention as specified in claim 9.
As per claim(s) 11, arguments made in rejecting claim(s) 1 are analogous.
As per claim(s) 12, arguments made in rejecting claim(s) 2 are analogous.
As per claim(s) 13, arguments made in rejecting claim(s) 3 are analogous.
As per claim(s) 14, arguments made in rejecting claim(s) 4 are analogous.
As per claim(s) 15, arguments made in rejecting claim(s) 1 are analogous.
Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ahn in view of Kong and Wang, and further in view of Song et al. (The SJTU 4K video sequence dataset) hereinafter referenced as Song.
Regarding claim 6, Ahn in view of Kong and Wang discloses: The device as claimed in claim 4 (Claim limitation is interpreted according to the rejection of claim 6 under 35 U.S.C. 112(b) as disclosed above), wherein the fourth learning model is a model trained to generate an interpolation frame having a blurred characteristic; and wherein the fifth learning model is a model trained to generate an interpolation frame having a grainy characteristic (Wang: Section: 4.1. Image Restoration: “Balance MSE and GAN effects in super-resolution. The aim of super-resolution is to estimate a high-resolution image from its low-resolution counterpart. A super-resolution model trained with MSE loss [35] tends to produce oversmooth images. We fine-tune it with GAN loss and perceptual loss [20], obtaining results with vivid details but always together with unpleasant artifacts (e.g., the eaves and water waves in Fig. 5).”;
Figure 5:
PNG
media_image2.png
510
1204
media_image2.png
Greyscale
;
Wherein the oversmoothed images associated with MSE loss and the artifact images associated with GAN loss, constitute images with blurred and grainy characteristics, respectively.).
Ahn in view of Kong and Wang does not disclose expressly: wherein the first learning model is a model trained with image data having complex movements; and wherein the second learning model is a model trained with image data having simple movements.
Thus, Ahn in view of Kong and Wang does not disclose expressly: the training of the first and the second video frame interpolation networks, which are each trained on distinct images from the training dataset based on image classifications. Wherein the images are classified as having complex or simple movements.
Song discloses: A video sequence dataset comprising 4K resolution ultra-high definition video sequences for the evaluation of video quality assessment algorithms (Song: Abstract). Wherein the video sequences are analyzed based upon a complexity of movements determined between each adjacent frame (Song: Section: 2.2. Sequences characteristics description: “Considering that video content plays a key role in the related researches, we attempt to shoot video sequences which can be representative of a wide variety of content types. All scenes are chosen in Shanghai, China. The factors such as image texture, image detail, movement speed of the object in the image, light intensity, and the camera lens stretching, panning are taken into account…
Generally, the spatial and temporal information were used as representing the video content. The spatial and temporal perceptual information of the scenes are critical parameters. These parameters play a crucial role in determining the amount of video compression and exert an important influence on the corresponding researches to some extent.
The analysis of content classification has been performed by computing the SI and TI indexes on the luminance component of each video sequences according to [7]. And the calculation process can be represented in equation form as:
PNG
media_image3.png
88
679
media_image3.png
Greyscale
”;
Wherein the temporal information, calculating the motion between each adjacent frame, constitutes a measure of motion complexity.).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to incorporate the calculation of temporal information used for the analysis of video sequences taught by Song into the classification of image pairs disclosed by Ahn in view of Kong and Wang. The suggestion/motivation for doing so would have been “Generally, the spatial and temporal information were used as representing the video content. The spatial and temporal perceptual information of the scenes are critical parameters. These parameters play a crucial role in determining the amount of video compression and exert an important influence on the corresponding researches to some extent.” (Song; Section: 2.2. Sequences characteristics description; Wherein the calculation of temporal information serves as a critical measure for the evaluation of video.). Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Ahn in view of Kong and Wang with Song to obtain the invention as specified in claim 6.
Claim(s) 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ahn in view of Kong and Wang, and further in view of Edelson et al. (US-20020180761-A1) hereinafter referenced as Edelson.
Regarding claim 10, Ahn in view of Kong and Wang discloses: The device as claimed in claim 1 (Claim limitation is interpreted according to the rejection of claim 10 under 35 U.S.C. 112(b) as disclosed above).
Ahn in view of Kong and Wang does not disclose expressly: further comprising: a display, wherein the at least one processor is individually and/or (The limitation’s use of “and/or” is interpreted as disjunctive “or”, thus indicating that only one limitation is required.) collectively configured to control the display to display an image in the order of the second frame, the interpolation frame and the first frame.
Edelson discloses: A system comprising a display, wherein the display is controlled to display a sequence of frames in the order of a previous frame, an interpolation frame and a current frame (Edelson: Abstract: “In a system for displaying medical images…The spacing between the images is such that the display of the images as a motion picture would result in the motion being depicted as jerky in the motion picture. A video processor generates dense motion vector fields between adjacent frames of the original set of images and, from the dense motion vector fields, generates interpolated images between the images of the original set. The interpolated images are assembled into a motion picture set of images, which are displayed by a video display device.”).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to incorporate the methods for assembling the interpolated images into an original image set and displaying them as disclosed by Edelson for the displaying of the original and interpolated video frames disclosed by Ahn in view of Kong and Wang. The suggestion/motivation for doing so would have been “when the medical images are sequences of images in time, the resulting motion picture will show changes in internal organs or tissue with time as smooth motion in the displayed motion picture. When the successive images are successive slices through an organ or through tissue in the body, the resulting displayed motion picture will show how the organs or tissue changes from slice to slice, moving through the body, as a smooth motion picture.” (Edelson: 0020). Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Ahn in view of Kong and Wang with Edelson to obtain the invention as specified in claim 10.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANTHONY J RODRIGUEZ whose telephone number is (703)756-5821. The examiner can normally be reached Monday-Friday 10am-7pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sumati Lefkowitz can be reached at (571) 272-3638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ANTHONY J RODRIGUEZ/Examiner, Art Unit 2672
/SUMATI LEFKOWITZ/Supervisory Patent Examiner, Art Unit 2672