DETAILED ACTION
Notice of Pre-AIA or AIA Status
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 7/27/26 has been entered.
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-10, 12, 14, 16 and 21 are pending in the application. Claims 1, 6-7, 14 and 16 have been amended, claims 11, 13, 15 and 17-20 have been canceled, and claim 21 has been added.
Response to Arguments
Applicant’s arguments, filed 7/27/26, with respect to 103 rejection to claim(s) 1 (see page 6-7) have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Objections
Claim 21 last 2 lines “the at least two of the plurality of sets” has no antecedent basis.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(d):
(d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
The following is a quotation of pre-AIA 35 U.S.C. 112, fourth paragraph:
Subject to the following paragraph [i.e., the fifth paragraph of pre-AIA 35 U.S.C. 112], a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
Claim 3 is rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends.
Claim 3 recites “shifting in a temporal direction within a range of 1 to N items”, while claim 1, the parent claim, recites “shifting in a temporal direction with a range of M (M is an integer greater than or equal to 2 and less than N) of the items of input image data”. Therefore the range defined in claim 3 is broader than the range defined in claim 1, resulting in claim 3 does not further limit claim 1.
Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 5-6, 12, 14 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (Chen H, Jin Y, Xu K, Chen Y, Zhu C. Multiframe-to-multiframe network for video denoising. IEEE Transactions on Multimedia. 2021 May 3;24:2164-78. Hereafter Chen), in view of Chen (US 20250272800 A1, hereafter Chen (’800)).
As per claim 1, Chen teaches the invention substantially as claimed including an information processing apparatus (page 2169 left col. A. Experimental Settings 1) Implementation Details: “The proposed network is implemented using the PyTorch 1 framework running on a GPU (Titan RTX)”) comprising:
one or more memories (Chen teaches a computer-implemented method performed by a GPU (Abstract; FIG. 2-5). The recited “one or more memories” is inherently taught.); and
one or more processors (See above cited portion “GPU”), wherein the one or more processors and the one or more memories are configured to:
acquire a plurality of items of input image data included in video data (Abstract “video”; Fig. 2 (c) “MM denoising scheme” shows a plurality of image frames t-n … t … t+n are acquired as input image data);
create a set of N (N is an integer greater than or equal to 3) items of input image data by selecting from the plurality of items of input image data (Fig. 2 (c) t-n … t … t+n, in which at least 3 frames are shown); and
output N items of first image data corresponding to the N items of input image data as a set by performing a restoration process using a first neural network to the N items of input image data in the one set (Fig. 2 (c) “MM denoising scheme” shows a plurality of image frames t-
n
^
… t … t+
n
^
are generated by a Denoiser. Chen further teaches the number of output frames could be the same as input frames (See below picture 1 captured from page 2167 right col.
PNG
media_image1.png
222
606
media_image1.png
Greyscale
(picture 1)
The Denoiser (MMNet) adopts a spatiotemporal convolutional architecture, considering both the interframe similarity and single-frame characteristics. As shown in Fig. 3, the proposed MMNet consists of an interframe denoising module, an intraframe denoising module, and a merging module. Both the interframe denoising module and the intraframe denoising module adopt an encoder-decoder architecture. See page 2167-2168 B. The Architecture of MMNet).
Chen, however, does not teach creating a plurality of sets of N (N is an integer greater than or equal to 3) items of input image data by selecting from the plurality of items of input image data, each of the sets being created by shifting in a temporal direction with a range of M (M is an integer greater than or equal to 2 and less than N) of the items of input image data (emphasis added to show the difference).
Chen (‘800) in an analogous field discloses a video processing method includes: obtaining a plurality of image groups on the basis of a video frame sequence of an initial video; performing motion blur processing on the basis of each frame of image in a target image group, and fusing images which are obtained by performing motion blur processing on each frame of image, so as to obtain a motion-blurred image corresponding to the target image group; on the basis of a specified frame of image in the target image group, determining a main body object area and a background area, which correspond to the target image group; fusing the motion-blurred image with the specified frame of image according to the main body object area and the background area, so as to obtain a target fused image; and generating a target video on the basis of target fused images respectively corresponding to the plurality of image groups (Abstract). When obtaining a plurality of image groups on the basis of a video frame sequence, Chen (‘800) creates a plurality of sets of N (N is an integer greater than or equal to 3) items of input image data by selecting from the plurality of items of input image data, each of the sets being created by shifting in a temporal direction with a range of M (M is an integer greater than or equal to 2 and less than N) of the items of input image data (As described in para. [0034], every 6 frames are selected as one image group for processing, but adjacent image groups have overlapping frames therebetween, that is, P1-P6 are taken as one image group, P4-P9 as one image group, P7-P12 as one image group, P10-P15 as one image group, . . . and so on. That says N=6, and M=3).
It would have been obvious for a person with ordinary skill in the art before the effective filing date of the claimed invention to have modified the teaching of Chen to incorporate the teaching of Chen (‘800) to create a plurality of sets of N (N is an integer greater than or equal to 3) items of input image data by selecting from the plurality of items of input image data, each of the sets being created by shifting in a temporal direction with a range of M (M is an integer greater than or equal to 2 and less than N) of the items of input image data. By doing so, both rationality of the number of the image groups (that is, rationality of a frame rate of subsequently generated video) and an image merging effect of each image group in subsequent processing can be ensured (Chen (‘800) para. [0033]).
As per claim 5, dependent upon claim 1, Chen in view of Chen (‘800) teaches the plurality of items of input image data is a plurality of chronologically consecutive items of input image data (See above Chen picture 1 in claim 1 and Fig. 2(c) t-n … t … t+n).
As per claim 6, dependent upon claim 1, Chen in view of Chen (‘800) teaches the one or more processors and the one or more memories are further configured to acquire a trained model of the neural network (Chen Fig. 2(c) “Denoiser”; Fig. 3; Table 1 last column “MMNet”) and acquire a trained model of a second neural network (Chen Table 1 col. 3-10 listing comparison results from other neural network models).
As per claim 12, dependent upon claim 1, Chen in view of Chen (‘800) teaches degradation to be processed includes at least one of noise ( Chen Abstract; Fig. 2(c); Fig. 3), compression, low resolution, blur, aberration, defect, and contrast reduction due to an influence of weather at a time of shooting.
Claim 14, an dependent method claim, recites similar steps as recited in claim 1. Therefore the recited steps of 14 are mapped to Chen and Chen (‘800) in the same manner as corresponding steps in claim 1. Additionally, the motivation for combining Chen and Chen (‘800) applied in claim 1 is also applicable here.
Claim 16, an dependent medium claim, recites similar elements as recited in claim 1. Therefore the recited elements of 16 are mapped to Chen and Chen (‘800) in the same manner as corresponding elements in claim 1. Additionally, the motivation for combining Chen and Chen (‘800) applied in claim 1 is also applicable here.
Claim(s) 2-4, 7-10 and 21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (Chen H, Jin Y, Xu K, Chen Y, Zhu C. Multiframe-to-multiframe network for video denoising. IEEE Transactions on Multimedia. 2021 May 3;24:2164-78. Hereafter Chen), in view of Chen (US 20250272800 A1, hereafter Chen (’800)), as applied to claim 1 above, and further in view of Wang et al. (US Patent 11,468,543 B1, hereafter Wang).
As per claim 2, Chen in view of Chen (‘800) teaches inputting a plurality of frames (N frame) as a set and outputting corresponding N frames for the set (Chen Fig. 2(c)).
Chen in view of Chen (‘800), however, does not teach concatenating the N items of input image data as a set.
Wang in the same field of endeavor discloses a neural network that receives mono-color pixels in a Bayer pattern from an image sensor and interpolates pixels to generate full-color RGB pixels while also enlightening the image for better detail mages (Abstract). FIG. 5 shows a convolutional neural network for image enlightening. The initial concatenate layer 46 received multiple images, such as three raw images with different scale factors or that were shot at different time slots, and outputs the concatenation of the three raw images. Since the depth of the initial Bayer raw image is 4 with Bayer pattern R, G1, B, G2 , the output of initial concatenate layer 46 has the dimension at [12, H, W]. That is to say the 3 image frames are concatenate along channel direction (FIG. 6; FIG. 12; col. 6 ln 53-61).
It would have been obvious for a person with ordinary skill in the art before the effective filing date of the claimed invention to have modified the teaching of Chen and Chen (‘800) to incorporate the teaching of Wang to concatenate the plurality of input items. Doing so would allow the plurality of items to form a data matrix when performing a series of convolution operations, such as contracting through contracting layers 52, 54, 56, 58 and expanding via expansion layers 72, 74, 76, 78 as recognized by Wang (FIG. 5).
As per claim 3, Chen in view of Chen (‘800) and Wang teaches wherein the one or more processors and the one or more memories are further configured to create a plurality of sets of the N items of input image data by selecting from the plurality of items of input image data, shifting in a temporal direction within a range of 1 to N items (Chen (‘800). As described in para. [0034], every 6 frames are selected as one image group for processing, but adjacent image groups have overlapping frames therebetween, that is, P1-P6 are taken as one image group, P4-P9 as one image group, P7-P12 as one image group, P10-P15 as one image group, . . . and so on. That says N=6 and the number of frame shifted is 3).
As per claim 4, dependent claim 2, Chen in view of Chen (‘800) and Wang teaches the one or more processors and the one or more memories are further configured to concatenate the N items of input image data by overlaying each pixel at same coordinates (Wang FIG. 5 shows a convolutional neural network for image enlightening. The initial concatenate layer 46 received multiple images, such as three raw images with different scale factors or that were shot at different time slots, and outputs the concatenation of the three raw images. Since the depth of the initial Bayer raw image is 4 with Bayer pattern R, G1, B, G2 , the output of initial concatenate layer 46 has the dimension at [12, H, W]. That is to say the 3 image frames are concatenate along channel direction (FIG. 6; FIG. 12; col. 6 ln 53-61), with each pixel of HXW pixels in one frame align each pixel of HXW pixels at the same coordinates in another frame).
.
As per claim 7, dependent upon claim 6, Chen in view of Chen (‘800) and Wang teaches the one or more processors and the one or more memories are further configured to, based on a plurality of items of the first image data output at a same time (Chen page 2167 right col. (see below capture picture 2)), output one item of second image data at that time (Wang FIG. 5 last two layers 80 and 82 are two post-processing layers. Layer 80 converts the n-deep output from the last of BP channel shrink layer 60 to a 12-bit depth, i.e., restores the original dimension [12, H, W] as after concatenate layer 46 and before layer 48. Depth-to-space layer 82 converts input dimensions [12, H, W] to output dimensions [3, 2H, 2 W]. No matter in layer 80 or layer 82, all the frames are converted at the same time, as a whole like one item (matrix). Same motivation applied to claim 2 is applicable here.)
PNG
media_image2.png
139
604
media_image2.png
Greyscale
(picture 2)
As per claim 8, dependent upon claim 7, Chen in view of Chen (‘800) and Wang teaches the one or more processors and the one or more memories are further configured to combine the plurality of items of the first image data at the same time to output one item of the second image data (As analyzed above in rejections applied to claim 2, Wang FIG 5 the initial concatenate layer 46 received multiple images, such as three raw images with different scale factors or that were shot at different time slots, and outputs the concatenation of the three raw images. Since the depth of the initial Bayer raw image is 4 with Bayer pattern R, G1, B, G2 , the output of initial concatenate layer 46 has the dimension at [12, H, W]. That is to say the 3 image frames are concatenate along channel direction (FIG. 6; FIG. 12; col. 6 ln 53-61). That says the concatenate layer 46 combines the plurality of input frames at the same time. Outputting one item of the second image data is analyzed above in claim 7).
As per claim 9, dependent upon claim 7, Chen in view of Chen (‘800) and Wang teaches the one or more processors and the one or more memories are further configured to combine the plurality of items of the first image data at the same time using a neural network to output one item of the second image data (Wang FIG. 5 layers 46, 48, …, 80, 82; See rejection applied to claims 7-8 above).
As per claim 10, dependent upon claim 7, Chen in view of Chen (‘800) and Wang teaches the one or more processors and the one or more memories are further configured to iteratively output N items of first image data corresponding to the N items of input image data and iteratively output one item of second image data at that time (Wang FIG. 5 shows a neural network for image enlightening. Wang further teaches training the neural network (FIG. 12). The training is performed in an iterative manner (col. 7 ln 63-66). The combination of Chen and Wang renders outputting N items of first image data corresponding to the N items of input image data and outputting one item of second image data at that time in an iterative manner.
It would have been obvious for a person with ordinary skill in the art before the effective filing date of the claimed invention to have modified the teaching of Chen and Chen (‘800) to incorporate the teaching of Wang to train the neural network in an iterative manner. Doing so would allow the optimal parameters of the neural network to be obtained (Wang col. 7 ln 63-66).
As per claim 21, dependent upon claim 1, Chen in view of Chen (‘800) and Wang further teaches wherein the one or more processors and the one or more memories are further configured to output one second image data at a time corresponding to the N items of input image data in the one set by combining the plurality of items of first image data at the same time among the output the plurality of items of the first image data corresponding to the plurality of items of the input image data included in the at least two of the plurality of sets (Wang FIG. 5 last two layers 80 and 82 are two post-processing layers. Layer 80 converts the n-deep output from the last of BP channel shrink layer 60 to a 12-bit depth, i.e., restores the original dimension [12, H, W] as after concatenate layer 46 and before layer 48. Depth-to-space layer 82 converts input dimensions [12, H, W] to output dimensions [3, 2H, 2 W]. No matter in layer 80 or layer 82, all the frames are converted at the same time, as a whole like one item (matrix)).
Conclusion
Prior art searched but not cited is recorded in PTO-892.
Additional prior art Ferrés et al. (US Patent 11,900,566 B1) discloses apparatus and method for denoising video frames (Abstract). As shown in FIG. 5, a video frame sequence at times t-2, t-1, t, t+1 and t+2 is obtained. The system creates 3 sets of frames, each with 3 frames selected from the obtained video frame sequence by shifting in a temporal direction within one frame. As shown in FIG. 5, the 1st set comprises frames at t-2, t-1, and t, the 2nd set comprises frames at t-1, t and t+1, and the 3rd set includes frames at t, t+1 and t+2.
Contact
Any inquiry concerning this communication or earlier communications from the examiner should be directed to XUEMEI G CHEN whose telephone number is (571)270-3480. The examiner can normally be reached Monday-Friday 9am-6pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, John M Villecco can be reached at (571) 272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/XUEMEI G CHEN/Primary Examiner, Art Unit 2661