DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of the Claims
Claims 38-57, as amended, are currently pending and have been considered below.
Specification
The disclosure is objected to because of the following informalities:
Page 26, line 33 of the specification recites, “The input to the processing module 500 can be provided through various input modules as indicated in block 531 already described in relation to FIG. 5D.” This sentence appears to contain a typographical error and should recite, “The input to the processing module 500 can be provided through various input modules as indicated in block 531 already described in relation to FIG. 5B” as there is no Figure 5D.
Page 34, lines 19-20 of the specification recites, “the receptive field value r provided by the syntax element nn_receptive_filed.” This sentence appears to contain a typographical error and should recite, “the receptive field value r provided by the syntax element nn_receptive_field.”
Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 38-57 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hannuksela, et al., "AHG9: NNR post-filter SEI message." The Joint Video Experts Team (JVET) of ITU-T SG16 WP3 and ISO/IEC JTC 1/SC29 26th Meeting by teleconference, 20-29 April 2022, hereinafter, “Hannuksela”, and further in view of Esenlik et al., International Publication No. WO 2022/128105, hereinafter, “Esenlik”.
As per claim 38, Hannuksela discloses a method comprising:
obtaining a video stream (Hannuksela, Abstract, The NNR post-filter SEI message may carry an ISO/IEC 15938-17 bitstream that specifies a neural-network-based post-filter; Hannuksela, page 12, 4 Background, An SEI message specifying the filter itself. Such an SEI message could for example contain or reference (through a URI) an MPEG Neural Network Representation (NNR, ISO/IEC 15938-17) bitstream);
obtaining metadata associated with the video stream representative of margins around a patch for an inference process of a neural network based image processing tool (Hannuksela, page 6, 2.2 NNR post-filter SEI message, nnrpf_overlap specifies the overlapping horizontal and vertical sample counts of adjacent input tensors of the post-processing filter. The value of nnrpf_overlap shall be in the range of 0 to 16383, inclusive. The variables inpPatchWidth, inpPatchHeight, outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows: ... overlapSize = nnrpf_overlap; Hannuksela, page 6, Table XYZ2B – Process for deriving the input tensors inputTensor for a given vertical sample coordinate cTop and a horizontal sample coordinate cLeft specifying the top-left sample location for the patch of samples included in the input tensors); and
decoding the video stream applying the neural network based image processing tool using the metadata (Hannuksela, page 6, Table XYZ2B – Process for deriving the input tensors inputTensor for a given vertical sample coordinate cTop and a horizontal sample coordinate cLeft specifying the top-left sample location for the patch of samples included in the input tensors; Hannuksela, page 8, Table XYZ3B – Process for deriving sample values in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensors outputTensor for a given vertical sample coordinate cTop and a horizontal sample coordinate cLeft specifying the top-left sample location for the patch of samples included in the input tensors [Process for deriving the input tensors makes use of the variable overlapSize which is equal to nnrpf_overlap]).
Hannuksela does not explicitly disclose the following limitation as further recited however Esenlik discloses
wherein the margins depend on a receptive field depending on a neural network used in the neural network based image processing tool (Esenlik, page 45, lines 8-13, the present disclosure can be applied to a whole or a part of a video compression and decompression process, if at least part of the process includes NN; Esenlik, page 47, line 3 - page 48, line 14, a method is provided for reconstructing a picture from a bitstream. The method may comprise obtaining, based on the bitstream, an input set of samples (L, illustrated in Fig. 17) representing the picture. The method includes dividing the input set L into two or more input sub-sets (e.g. L1, L2 shown in Fig. 17). Then, the method may include determining a size ((hi, w1); (h2, w2)) for each of the two or more input sub-sets (L1, L2) and/or a size ((H1, W1); (H2, W2)) for each of two or more output sub-sets (R1, R2). The determination may be based on side information from the bitstream ... Fig. 17A illustrates the exemplary implementation. First, the input is divided into 4 overlapping regions, Li, L2, L3, and L4. The Li and L2 are shown in the figure. One can notice that Li and L2 are overlapping ... The total receptive field consists of the union set of input samples that are all used in the calculation of a set of output samples; Esenlik, page 70, line 30 - page 71, line 8, As discussed before for the picture reconstructing, the input sub-sets (L1, L2) may not comprise the total receptive field (TRF) ... As a result of using a sub-field of the TRF, the encoding side does per se not know how to partition the picture. Therefore, the encoding side may need to determine side information indicating how to partition the picture. The respective side information may include parameters e.g. related to the picture or other parameters. The side information may then be signaled to the decoder side in the bitstream ... the encoder may decide on the acceptable selection of overlapping parameters in the input and output spaces to control the reconstruction quality. For example, the encoder may use the size (h, w) of the input picture and determine the overlap amount y of the input sub-sets Li and/or the overlap amount X of the output sub-sets Ri, based on the input-output characteristics of the neural network).
It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine the teachings of Esenlik and Hannuksela because they are in the same field of endeavor. One skilled in the art would have been motivated to include the receptive field as taught by Esenlik in the system of Hannuksela in order to improve quality and accuracy in the reconstructed output (Esenlik, Abstract).
As per claim 39, Hannuksela and Esenlik disclose the method of claim 38, wherein the metadata comprise at least one syntax element representative of the receptive field (Hannnuksela, page 6, 2.2 NNR post-filter SEI message, nnrpf_overlap specifies the overlapping horizontal and vertical sample counts of adjacent input tensors of the post-processing filter; Esenlik, page 70, line 30 - page 71, line 8, As discussed before for the picture reconstructing, the input sub-sets (L1, L2) may not comprise the total receptive field (TRF) ... As a result of using a sub-field of the TRF, the encoding side does per se not know how to partition the picture. Therefore, the encoding side may need to determine side information indicating how to partition the picture. The respective side information may include parameters e.g. related to the picture or other parameters. The side information may then be signaled to the decoder side in the bitstream ... the encoder may decide on the acceptable selection of overlapping parameters in the input and output spaces to control the reconstruction quality. For example, the encoder may use the size (h, w) of the input picture and determine the overlap amount y of the input sub-sets Li and/or the overlap amount X of the output sub-sets Ri, based on the input-output characteristics of the neural network).
As per claim 40, Hannuksela and Esenlik disclose the method of claim 39, wherein the at least one syntax element comprises a first syntax element defining the receptive field vertically and a second syntax element defining the receptive field horizontally (Hannuksela, page 6, 2.2 NNR post-filter SEI message, The variables inpPatchWidth, inpPatchHeight, outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows: inpPatchWidth = nnrpf_patch_size_minus1 + 1; inpPatchHeight = nnrpf_patch_size_minus1 + 1; outPatchWidth = ( nnrpf_pic_width_in_luma_samples * inpPatchWidth ) / InpPicWidthInLumaSamples; outPatchHeight = ( nnrpf_pic_height_in_luma_samples * inpPatchHeight ) / InpPicHeightInLumaSamples; Esenlik, page 4, lines 18-28, According to an implementation example of the method, the side information includes an indication of one or more of: a number of the input sub-sets, a size of the input set, a size (hi, w1) of each of the two or more input sub-sets, a size (H, W) of the reconstructed picture (R), a size (H1, W1 ) of each of the two or more output sub-set, an amount of overlap between the two or more input sub-sets (L1, L2), an amount of overlap between the two or more output sub-sets (R1, R2). Hence, the signaling of a variety of parameters through side information may be performed; Esenlik, page 47, line 3 - page 48, line 14, a method is provided for reconstructing a picture from a bitstream. The method may comprise obtaining, based on the bitstream, an input set of samples (L, illustrated in Fig. 17) representing the picture. The method includes dividing the input set L into two or more input sub-sets (e.g. L1, L2 shown in Fig. 17). Then, the method may include determining a size ((hi, w1); (h2, w2)) for each of the two or more input sub-sets (L1, L2) and/or a size ((H1, W1); (H2, W2)) for each of two or more output sub-sets (R1, R2). The determination may be based on side information from the bitstream ... Fig. 17A illustrates the exemplary implementation. First, the input is divided into 4 overlapping regions, Li, L2, L3, and L4. The Li and L2 are shown in the figure. One can notice that Li and L2 are overlapping … The total receptive field consists of the union set of input samples that are all used in the calculation of a set of output samples).
As per claim 41, Hannuksela and Esenlik disclose the method of claim 38, wherein a capacity of the inference process to process patches larger than a patch size considered during a definition of the neural network used in the neural network based image processing tool is specified in the metadata by a syntax element (Hannuksela, page 2, 2.1 Overview, A flag (nnrpf_constant_patch_size_flag) has been added to indicate, when equal to 0, that the patch size used as input to the filter can be any positive integer multiple of the signalled size).
As per claim 42, Hannuksela and Esenlik disclose the method of claim 38, wherein the metadata comprises at least one syntax element representative of a position of an output patch of the inference process of the neural network based image processing tool in an output tensor generated by the inference process (Hannuksela, page 8, Table XYZ3B – Process for deriving sample values in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensors outputTensor for a given vertical sample coordinate cTop and a horizontal sample coordinate cLeft specifying the top-left sample location for the patch of samples included in the input tensors).
As per claim 43, Hannuksela discloses a method comprising:
obtaining a video stream (Hannuksela, Abstract, The NNR post-filter SEI message may carry an ISO/IEC 15938-17 bitstream that specifies a neural-network-based post-filter; Hannuksela, page 12, 4 Background, An SEI message specifying the filter itself. Such an SEI message could for example contain or reference (through a URI) an MPEG Neural Network Representation (NNR, ISO/IEC 15938-17) bitstream); and
signaling information representative of margins around a patch for an inference process of a neural network based image processing tool in the form of metadata associated to the video stream (Hannuksela, page 6, 2.2 NNR post-filter SEI message, nnrpf_overlap specifies the overlapping horizontal and vertical sample counts of adjacent input tensors of the post-processing filter. The value of nnrpf_overlap shall be in the range of 0 to 16383, inclusive. The variables inpPatchWidth, inpPatchHeight, outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows: ... overlapSize = nnrpf_overlap; Hannuksela, page 6, Table XYZ2B – Process for deriving the input tensors inputTensor for a given vertical sample coordinate cTop and a horizontal sample coordinate cLeft specifying the top-left sample location for the patch of samples included in the input tensors).
Hannuksela does not explicitly disclose the following limitation as further recited however Esenlik discloses
wherein the margins depend on a receptive field depending on a neural network used in the neural network based image processing tool (Esenlik, page 45, lines 8-13, the present disclosure can be applied to a whole or a part of a video compression and decompression process, if at least part of the process includes NN; Esenlik, page 47, line 3 - page 48, line 14, a method is provided for reconstructing a picture from a bitstream. The method may comprise obtaining, based on the bitstream, an input set of samples (L, illustrated in Fig. 17) representing the picture. The method includes dividing the input set L into two or more input sub-sets (e.g. L1, L2 shown in Fig. 17). Then, the method may include determining a size ((hi, w1); (h2, w2)) for each of the two or more input sub-sets (L1, L2) and/or a size ((H1, W1); (H2, W2)) for each of two or more output sub-sets (R1, R2). The determination may be based on side information from the bitstream ... Fig. 17A illustrates the exemplary implementation. First, the input is divided into 4 overlapping regions, Li, L2, L3, and L4. The Li and L2 are shown in the figure. One can notice that Li and L2 are overlapping ... The total receptive field consists of the union set of input samples that are all used in the calculation of a set of output samples; Esenlik, page 70, line 30 - page 71, line 8, As discussed before for the picture reconstructing, the input sub-sets (L1, L2) may not comprise the total receptive field (TRF) ... As a result of using a sub-field of the TRF, the encoding side does per se not know how to partition the picture. Therefore, the encoding side may need to determine side information indicating how to partition the picture. The respective side information may include parameters e.g. related to the picture or other parameters. The side information may then be signaled to the decoder side in the bitstream ... the encoder may decide on the acceptable selection of overlapping parameters in the input and output spaces to control the reconstruction quality. For example, the encoder may use the size (h, w) of the input picture and determine the overlap amount y of the input sub-sets Li and/or the overlap amount X of the output sub-sets Ri, based on the input-output characteristics of the neural network).
It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine the teachings of Esenlik and Hannuksela because they are in the same field of endeavor. One skilled in the art would have been motivated to include the receptive field as taught by Esenlik in the system of Hannuksela in order to improve quality and accuracy in the reconstructed output (Esenlik, Abstract).
As per claim 44, Hannuksela and Esenlik disclose the method of claim 43, wherein the metadata comprise at least one syntax element representative of the receptive field (Hannnuksela, page 6, 2.2 NNR post-filter SEI message, nnrpf_overlap specifies the overlapping horizontal and vertical sample counts of adjacent input tensors of the post-processing filter; Esenlik, page 70, line 30 - page 71, line 8, As discussed before for the picture reconstructing, the input sub-sets (L1, L2) may not comprise the total receptive field (TRF) ... As a result of using a sub-field of the TRF, the encoding side does per se not know how to partition the picture. Therefore, the encoding side may need to determine side information indicating how to partition the picture. The respective side information may include parameters e.g. related to the picture or other parameters. The side information may then be signaled to the decoder side in the bitstream ... the encoder may decide on the acceptable selection of overlapping parameters in the input and output spaces to control the reconstruction quality. For example, the encoder may use the size (h, w) of the input picture and determine the overlap amount y of the input sub-sets Li and/or the overlap amount X of the output sub-sets Ri, based on the input-output characteristics of the neural network).
As per claim 45, Hannuksela and Esenlik disclose the method of claim 44, wherein the at least one syntax element comprises a first syntax element defining the receptive field vertically and a second syntax element defining the receptive field horizontally (Hannuksela, page 6, 2.2 NNR post-filter SEI message, The variables inpPatchWidth, inpPatchHeight, outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows: inpPatchWidth = nnrpf_patch_size_minus1 + 1; inpPatchHeight = nnrpf_patch_size_minus1 + 1; outPatchWidth = ( nnrpf_pic_width_in_luma_samples * inpPatchWidth ) / InpPicWidthInLumaSamples; outPatchHeight = ( nnrpf_pic_height_in_luma_samples * inpPatchHeight ) / InpPicHeightInLumaSamples; Esenlik, page 4, lines 18-28, According to an implementation example of the method, the side information includes an indication of one or more of: a number of the input sub-sets, a size of the input set, a size (hi, w1) of each of the two or more input sub-sets, a size (H, W) of the reconstructed picture (R), a size (H1, W1 ) of each of the two or more output sub-set, an amount of overlap between the two or more input sub-sets (L1, L2), an amount of overlap between the two or more output sub-sets (R1, R2). Hence, the signaling of a variety of parameters through side information may be performed; Esenlik, page 47, line 3 - page 48, line 14, a method is provided for reconstructing a picture from a bitstream. The method may comprise obtaining, based on the bitstream, an input set of samples (L, illustrated in Fig. 17) representing the picture. The method includes dividing the input set L into two or more input sub-sets (e.g. L1, L2 shown in Fig. 17). Then, the method may include determining a size ((hi, w1); (h2, w2)) for each of the two or more input sub-sets (L1, L2) and/or a size ((H1, W1); (H2, W2)) for each of two or more output sub-sets (R1, R2). The determination may be based on side information from the bitstream ... Fig. 17A illustrates the exemplary implementation. First, the input is divided into 4 overlapping regions, Li, L2, L3, and L4. The Li and L2 are shown in the figure. One can notice that Li and L2 are overlapping … The total receptive field consists of the union set of input samples that are all used in the calculation of a set of output samples).
As per claim 46, Hannuksela and Esenlik disclose the method of claim 43, wherein a capacity of the inference process to process patches larger than a patch size considered during a definition of the neural network used in the neural network based image processing tool is specified in the metadata by a syntax element (Hannuksela, page 2, 2.1 Overview, A flag (nnrpf_constant_patch_size_flag) has been added to indicate, when equal to 0, that the patch size used as input to the filter can be any positive integer multiple of the signalled size).
As per claim 47, Hannuksela and Esenlik disclose the method of claim 43, wherein the metadata comprises at least one syntax element representative of a position of an output patch of the inference process of the neural network based image processing tool in an output tensor generated by the inference process (Hannuksela, page 8, Table XYZ3B – Process for deriving sample values in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensors outputTensor for a given vertical sample coordinate cTop and a horizontal sample coordinate cLeft specifying the top-left sample location for the patch of samples included in the input tensors).
As per claim 48, Hannuksela and Esenlik disclose the method of claim 43, wherein the video stream is obtained by applying a video compression process to an original video, the video compression process comprising the neural network based image processing tool in a prediction loop of the video compression process or the neural network based image processing tool being a post-processing tool (Hannuksela, Abstract, The NNR post-filter SEI message may carry an ISO/IEC 15938-17 bitstream that specifies a neural-network-based post-filter; Esenlik, page 45, lines 8-13, the present disclosure can be applied to a whole or a part of a video compression and decompression process, if at least part of the process includes NN and if such NN includes convolution or transposed convolution operations. For example, the present disclosure is applicable to individual processing tasks as being performed as a processing part by the encoder and/or decoder, including in-loop filtering, post-filtering, and/or prefiltering).
As per claim 49, Hannuksela discloses a device comprising electronic circuitry configured for:
obtaining a video stream (Hannuksela, Abstract, The NNR post-filter SEI message may carry an ISO/IEC 15938-17 bitstream that specifies a neural-network-based post-filter; Hannuksela, page 12, 4 Background, An SEI message specifying the filter itself. Such an SEI message could for example contain or reference (through a URI) an MPEG Neural Network Representation (NNR, ISO/IEC 15938-17) bitstream);
obtaining metadata associated with the video stream representative of margins around a patch for an inference process of a neural network based image processing tool (Hannuksela, page 6, 2.2 NNR post-filter SEI message, nnrpf_overlap specifies the overlapping horizontal and vertical sample counts of adjacent input tensors of the post-processing filter. The value of nnrpf_overlap shall be in the range of 0 to 16383, inclusive. The variables inpPatchWidth, inpPatchHeight, outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows: ... overlapSize = nnrpf_overlap; Hannuksela, page 6, Table XYZ2B – Process for deriving the input tensors inputTensor for a given vertical sample coordinate cTop and a horizontal sample coordinate cLeft specifying the top-left sample location for the patch of samples included in the input tensors); and
decoding the video stream applying the neural network based image processing tool using the metadata (Hannuksela, page 6, Table XYZ2B – Process for deriving the input tensors inputTensor for a given vertical sample coordinate cTop and a horizontal sample coordinate cLeft specifying the top-left sample location for the patch of samples included in the input tensors; Hannuksela, page 8, Table XYZ3B – Process for deriving sample values in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensors outputTensor for a given vertical sample coordinate cTop and a horizontal sample coordinate cLeft specifying the top-left sample location for the patch of samples included in the input tensors [Process for deriving the input tensors makes use of the variable overlapSize which is equal to nnrpf_overlap]).
Hannuksela does not explicitly disclose the following limitation as further recited however Esenlik discloses
wherein the margins depend on a receptive field depending on a neural network used in the neural network based image processing tool (Esenlik, page 45, lines 8-13, the present disclosure can be applied to a whole or a part of a video compression and decompression process, if at least part of the process includes NN; Esenlik, page 47, line 3 - page 48, line 14, a method is provided for reconstructing a picture from a bitstream. The method may comprise obtaining, based on the bitstream, an input set of samples (L, illustrated in Fig. 17) representing the picture. The method includes dividing the input set L into two or more input sub-sets (e.g. L1, L2 shown in Fig. 17). Then, the method may include determining a size ((hi, w1); (h2, w2)) for each of the two or more input sub-sets (L1, L2) and/or a size ((H1, W1); (H2, W2)) for each of two or more output sub-sets (R1, R2). The determination may be based on side information from the bitstream ... Fig. 17A illustrates the exemplary implementation. First, the input is divided into 4 overlapping regions, Li, L2, L3, and L4. The Li and L2 are shown in the figure. One can notice that Li and L2 are overlapping ... The total receptive field consists of the union set of input samples that are all used in the calculation of a set of output samples; Esenlik, page 70, line 30 - page 71, line 8, As discussed before for the picture reconstructing, the input sub-sets (L1, L2) may not comprise the total receptive field (TRF) ... As a result of using a sub-field of the TRF, the encoding side does per se not know how to partition the picture. Therefore, the encoding side may need to determine side information indicating how to partition the picture. The respective side information may include parameters e.g. related to the picture or other parameters. The side information may then be signaled to the decoder side in the bitstream ... the encoder may decide on the acceptable selection of overlapping parameters in the input and output spaces to control the reconstruction quality. For example, the encoder may use the size (h, w) of the input picture and determine the overlap amount y of the input sub-sets Li and/or the overlap amount X of the output sub-sets Ri, based on the input-output characteristics of the neural network).
It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine the teachings of Esenlik and Hannuksela because they are in the same field of endeavor. One skilled in the art would have been motivated to include the receptive field as taught by Esenlik in the system of Hannuksela in order to improve quality and accuracy in the reconstructed output (Esenlik, Abstract).
As per claim 50, Hannuksela and Esenlik disclose disclose the device of claim 49, wherein the metadata comprise at least one syntax element representative of the receptive field (Hannnuksela, page 6, 2.2 NNR post-filter SEI message, nnrpf_overlap specifies the overlapping horizontal and vertical sample counts of adjacent input tensors of the post-processing filter; Esenlik, page 70, line 30 - page 71, line 8, As discussed before for the picture reconstructing, the input sub-sets (L1, L2) may not comprise the total receptive field (TRF) ... As a result of using a sub-field of the TRF, the encoding side does per se not know how to partition the picture. Therefore, the encoding side may need to determine side information indicating how to partition the picture. The respective side information may include parameters e.g. related to the picture or other parameters. The side information may then be signaled to the decoder side in the bitstream ... the encoder may decide on the acceptable selection of overlapping parameters in the input and output spaces to control the reconstruction quality. For example, the encoder may use the size (h, w) of the input picture and determine the overlap amount y of the input sub-sets Li and/or the overlap amount X of the output sub-sets Ri, based on the input-output characteristics of the neural network).
As per claim 51, Hannuksela and Esenlik disclose the device of claim 50, wherein the at least one syntax element comprises a first syntax element defining the receptive field vertically and a second syntax element defining the receptive field horizontally (Hannuksela, page 6, 2.2 NNR post-filter SEI message, The variables inpPatchWidth, inpPatchHeight, outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows: inpPatchWidth = nnrpf_patch_size_minus1 + 1; inpPatchHeight = nnrpf_patch_size_minus1 + 1; outPatchWidth = ( nnrpf_pic_width_in_luma_samples * inpPatchWidth ) / InpPicWidthInLumaSamples; outPatchHeight = ( nnrpf_pic_height_in_luma_samples * inpPatchHeight ) / InpPicHeightInLumaSamples; Esenlik, page 4, lines 18-28, According to an implementation example of the method, the side information includes an indication of one or more of: a number of the input sub-sets, a size of the input set, a size (hi, w1) of each of the two or more input sub-sets, a size (H, W) of the reconstructed picture (R), a size (H1, W1 ) of each of the two or more output sub-set, an amount of overlap between the two or more input sub-sets (L1, L2), an amount of overlap between the two or more output sub-sets (R1, R2). Hence, the signaling of a variety of parameters through side information may be performed; Esenlik, page 47, line 3 - page 48, line 14, a method is provided for reconstructing a picture from a bitstream. The method may comprise obtaining, based on the bitstream, an input set of samples (L, illustrated in Fig. 17) representing the picture. The method includes dividing the input set L into two or more input sub-sets (e.g. L1, L2 shown in Fig. 17). Then, the method may include determining a size ((hi, w1); (h2, w2)) for each of the two or more input sub-sets (L1, L2) and/or a size ((H1, W1); (H2, W2)) for each of two or more output sub-sets (R1, R2). The determination may be based on side information from the bitstream ... Fig. 17A illustrates the exemplary implementation. First, the input is divided into 4 overlapping regions, Li, L2, L3, and L4. The Li and L2 are shown in the figure. One can notice that Li and L2 are overlapping … The total receptive field consists of the union set of input samples that are all used in the calculation of a set of output samples).
As per claim 52, Hannuksela and Esenlik disclose the device of claim 49, wherein a capacity of the inference process to process patches larger than a patch size considered during a definition of the neural network used in the neural network based image processing tool is specified in the metadata by a syntax element (Hannuksela, page 2, 2.1 Overview, A flag (nnrpf_constant_patch_size_flag) has been added to indicate, when equal to 0, that the patch size used as input to the filter can be any positive integer multiple of the signalled size).
As per claim 53, Hannuksela and Esenlik disclose the device of claim 49, wherein the metadata comprise at least one syntax element representative of a position of an output patch of the inference process of the neural network based image processing tool in an output tensor generated by the inference process (Hannuksela, page 8, Table XYZ3B – Process for deriving sample values in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensors outputTensor for a given vertical sample coordinate cTop and a horizontal sample coordinate cLeft specifying the top-left sample location for the patch of samples included in the input tensors).
As per claim 54, Hannuksela discloses a device comprising electronic circuitry configured for:
obtaining a video stream (Hannuksela, Abstract, The NNR post-filter SEI message may carry an ISO/IEC 15938-17 bitstream that specifies a neural-network-based post-filter; Hannuksela, page 12, 4 Background, An SEI message specifying the filter itself. Such an SEI message could for example contain or reference (through a URI) an MPEG Neural Network Representation (NNR, ISO/IEC 15938-17) bitstream); and
signaling information representative of margins around a patch for an inference process of a neural network based image processing tool in the form of metadata associated to the video stream (Hannuksela, page 6, 2.2 NNR post-filter SEI message, nnrpf_overlap specifies the overlapping horizontal and vertical sample counts of adjacent input tensors of the post-processing filter. The value of nnrpf_overlap shall be in the range of 0 to 16383, inclusive. The variables inpPatchWidth, inpPatchHeight, outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows: ... overlapSize = nnrpf_overlap; Hannuksela, page 6, Table XYZ2B – Process for deriving the input tensors inputTensor for a given vertical sample coordinate cTop and a horizontal sample coordinate cLeft specifying the top-left sample location for the patch of samples included in the input tensors).
Hannuksela does not explicitly disclose the following limitation as further recited however Esenlik discloses
wherein the margins depend on a receptive field depending on a neural network used in the neural network based image processing tool (Esenlik, page 45, lines 8-13, the present disclosure can be applied to a whole or a part of a video compression and decompression process, if at least part of the process includes NN; Esenlik, page 47, line 3 - page 48, line 14, a method is provided for reconstructing a picture from a bitstream. The method may comprise obtaining, based on the bitstream, an input set of samples (L, illustrated in Fig. 17) representing the picture. The method includes dividing the input set L into two or more input sub-sets (e.g. L1, L2 shown in Fig. 17). Then, the method may include determining a size ((hi, w1); (h2, w2)) for each of the two or more input sub-sets (L1, L2) and/or a size ((H1, W1); (H2, W2)) for each of two or more output sub-sets (R1, R2). The determination may be based on side information from the bitstream ... Fig. 17A illustrates the exemplary implementation. First, the input is divided into 4 overlapping regions, Li, L2, L3, and L4. The Li and L2 are shown in the figure. One can notice that Li and L2 are overlapping ... The total receptive field consists of the union set of input samples that are all used in the calculation of a set of output samples; Esenlik, page 70, line 30 - page 71, line 8, As discussed before for the picture reconstructing, the input sub-sets (L1, L2) may not comprise the total receptive field (TRF) ... As a result of using a sub-field of the TRF, the encoding side does per se not know how to partition the picture. Therefore, the encoding side may need to determine side information indicating how to partition the picture. The respective side information may include parameters e.g. related to the picture or other parameters. The side information may then be signaled to the decoder side in the bitstream ... the encoder may decide on the acceptable selection of overlapping parameters in the input and output spaces to control the reconstruction quality. For example, the encoder may use the size (h, w) of the input picture and determine the overlap amount y of the input sub-sets Li and/or the overlap amount X of the output sub-sets Ri, based on the input-output characteristics of the neural network).
It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine the teachings of Esenlik and Hannuksela because they are in the same field of endeavor. One skilled in the art would have been motivated to include the receptive field as taught by Esenlik in the system of Hannuksela in order to improve quality and accuracy in the reconstructed output (Esenlik, Abstract).
As per claim 55, Hannuksela and Esenlik disclose the device of claim 54, wherein the metadata comprise at least one syntax element representative of the receptive field (Hannnuksela, page 6, 2.2 NNR post-filter SEI message, nnrpf_overlap specifies the overlapping horizontal and vertical sample counts of adjacent input tensors of the post-processing filter; Esenlik, page 70, line 30 - page 71, line 8, As discussed before for the picture reconstructing, the input sub-sets (L1, L2) may not comprise the total receptive field (TRF) ... As a result of using a sub-field of the TRF, the encoding side does per se not know how to partition the picture. Therefore, the encoding side may need to determine side information indicating how to partition the picture. The respective side information may include parameters e.g. related to the picture or other parameters. The side information may then be signaled to the decoder side in the bitstream ... the encoder may decide on the acceptable selection of overlapping parameters in the input and output spaces to control the reconstruction quality. For example, the encoder may use the size (h, w) of the input picture and determine the overlap amount y of the input sub-sets Li and/or the overlap amount X of the output sub-sets Ri, based on the input-output characteristics of the neural network).
As per claim 56, Hannuksela and Esenlik disclose the device of claim 55, wherein the at least one syntax element comprises a first syntax element defining the receptive field vertically and a second syntax element defining the receptive field horizontally (Hannuksela, page 6, 2.2 NNR post-filter SEI message, The variables inpPatchWidth, inpPatchHeight, outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows: inpPatchWidth = nnrpf_patch_size_minus1 + 1; inpPatchHeight = nnrpf_patch_size_minus1 + 1; outPatchWidth = ( nnrpf_pic_width_in_luma_samples * inpPatchWidth ) / InpPicWidthInLumaSamples; outPatchHeight = ( nnrpf_pic_height_in_luma_samples * inpPatchHeight ) / InpPicHeightInLumaSamples; Esenlik, page 4, lines 18-28, According to an implementation example of the method, the side information includes an indication of one or more of: a number of the input sub-sets, a size of the input set, a size (hi, w1) of each of the two or more input sub-sets, a size (H, W) of the reconstructed picture (R), a size (H1, W1 ) of each of the two or more output sub-set, an amount of overlap between the two or more input sub-sets (L1, L2), an amount of overlap between the two or more output sub-sets (R1, R2). Hence, the signaling of a variety of parameters through side information may be performed; Esenlik, page 47, line 3 - page 48, line 14, a method is provided for reconstructing a picture from a bitstream. The method may comprise obtaining, based on the bitstream, an input set of samples (L, illustrated in Fig. 17) representing the picture. The method includes dividing the input set L into two or more input sub-sets (e.g. L1, L2 shown in Fig. 17). Then, the method may include determining a size ((hi, w1); (h2, w2)) for each of the two or more input sub-sets (L1, L2) and/or a size ((H1, W1); (H2, W2)) for each of two or more output sub-sets (R1, R2). The determination may be based on side information from the bitstream ... Fig. 17A illustrates the exemplary implementation. First, the input is divided into 4 overlapping regions, Li, L2, L3, and L4. The Li and L2 are shown in the figure. One can notice that Li and L2 are overlapping … The total receptive field consists of the union set of input samples that are all used in the calculation of a set of output samples).
As per claim 57, Hannuksela and Esenlik disclose the device of claim 54, wherein a capacity of the inference process to process patches larger than a patch size considered during a definition of the neural network used in the neural network based image processing tool is specified in the metadata by a syntax element (Hannuksela, page 2, 2.1 Overview, A flag (nnrpf_constant_patch_size_flag) has been added to indicate, when equal to 0, that the patch size used as input to the filter can be any positive integer multiple of the signalled size).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TRACY MANGIALASCHI whose telephone number is (571)270-5189. The examiner can normally be reached M-F, 9:30AM TO 6:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be reached at (571) 272-7332. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/TRACY MANGIALASCHI/Primary Examiner, Art Unit 2668