DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
The following title is suggested: METHOD, APPARATUS, AND MEDIUM FOR VIDEO PROCESSING INVOLVING BI-DIRECTIONAL OPTICAL FLOW.
Applicant is reminded of the proper language and format for an abstract of the disclosure.
The abstract should be in narrative form and generally limited to a single paragraph on a separate sheet within the range of 50 to 150 words in length. The abstract should describe the disclosure sufficiently to assist readers in deciding whether there is a need for consulting the full patent text for details.
The language should be clear and concise and should not repeat information given in the title. It should avoid using phrases which can be implied, such as, “The disclosure concerns,” “The disclosure defined by this invention,” “The disclosure describes,” etc. In addition, the form and legal phraseology often used in patent claims, such as “means” and “said,” should be avoided.
The abstract of the disclosure is objected to because it contains legal phraseology such as “comprises”. A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b).
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-4, 9, 14, 16-19 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by US 2019/0349589 A1 (“Lee”).
Regarding claim 1, Lee discloses a method for video processing, comprising:
obtaining, for a conversion between a current video block of a video and a bitstream of the video (e.g. see encoder, e.g. 100 in Fig. 1, and/or decoder, e.g. 200 in Fig. 2, for a conversion of input video, e.g. current block such as shown 1201 in Fig. 12, and bitstream), a plurality of weights, wherein the plurality of weights are used to weight a plurality of values for a metric at respective samples in a target region (e.g. see applying a weighting function g as shown in Equation 25, e.g. see weights in Figs. 17-18, paragraphs [0323]-[0340]) for a bi-directional optical flow (BDOF) process applied on the current video block (e.g. see bi-directional optical (BIO) flow, paragraphs [0006]-[0009]), the metric comprises at least one of a gradient or a difference of sample values (e.g. see Gx and/or Gy, paragraphs [0323]-[0340]), and the plurality of values for the metric are determined based on a plurality of reference video blocks of the current video block (e.g. see Gx and/or Gy determined based on plurality of reference pictures, e.g. Ref0 1020 and Ref1 1030 in Fig. 10, paragraphs [0201]-[0217], Abstract); and
performing the conversion based on the plurality of weights (e.g. see encoder, e.g. 100 in Fig. 1, and/or decoder, e.g. 200 in Fig. 2, for a conversion of input video, e.g. current block 1201 in Fig. 12, and bitstream; thus, the encoder and/or decoder perform the conversion based on the applied weighting function g as shown in Equation 25).
Regarding claim 2, Lee further discloses wherein performing the conversion comprises:
determining a set of parameters for the BDOF process based on the plurality of weights (e.g. see Equation 18 described above (e.g. paragraphs [0218]-[0224]) may be expressed as shown in Equation 25, paragraphs [0323]-[0340]);
determining, based on the set of parameters, at least one offset for refining a motion vector (MV) of the current video block or adjusting a current sample in the current video block (e.g. see Vx and/or Vy in Equation 19 or 21 for predicting the current pixel in Equation 20, paragraphs [0218]-[0224]); and
performing the conversion based on the at least one offset (e.g. see encoder, e.g. 100 in Fig. 1, and/or decoder, e.g. 200 in Fig. 2, for a conversion of input video, e.g. current block 1201 in Fig. 12, and bitstream; thus, the encoder and/or decoder perform the conversion based on Vx and/or Vy in Equation 10 or 21 for predicting the current pixel in Equation 20).
Regarding claim 3, Lee further discloses wherein the at least one offset is used for refining the motion vector of the current video block, a size of the current video block is MxN, the target region comprises a region around the current video block of a size (M+K1)x(N+K2), and each of M, N, K1 and K2 is an integer, or
wherein the at least one offset is used for adjusting the current sample, the target region comprises a region around the current sample with a size of K3 xK4, and each of K3 and K4 is an integer, or
wherein the set of parameters for the BDOF process is determined based on the plurality of weights and a complete linear equation formula, or
wherein the set of parameters comprises a first parameter, a second parameter, a third parameter, a fourth parameter, and a fifth parameter that are determined in a predetermined manner (e.g. see at least s1, s2, s3, s5 and s6 in Equation 25, paragraphs [0323]-[0340]).
Regarding claim 4, Lee further discloses wherein the gradient comprises at least one of a horizontal gradient or a vertical gradient, and the set of parameters is determined based on the following:
PNG
media_image1.png
136
110
media_image1.png
Greyscale
wherein s1 represents the first parameter, s2 represents the second parameter, s3 represents the third parameter, s5 represents the fourth parameter, s6 represents the fifth parameter, Gx represents a summation of values for the horizontal gradient determined for each of the plurality of reference video blocks, Gy represents a summation of values for the vertical gradient determined for each of the plurality of reference video blocks, dI represents the difference of sample values between the plurality of reference video blocks, and
PNG
media_image2.png
20
28
media_image2.png
Greyscale
represents a weighted sum inside the target region based on the plurality of weights, or
wherein the gradient comprises at least one of a horizontal gradient or a vertical gradient, and the set of parameters is determined based on the following:
PNG
media_image3.png
816
654
media_image3.png
Greyscale
(e.g. see Vx and/or Vy in Equation 21 for predicting the current pixel in Equation 20, paragraphs [0218]-[0224]).
Regarding claim 9, Lee further discloses wherein each of the plurality of weights is equal to a same predetermined value, or
wherein a first weight of the plurality of weights that corresponds to a first sample in the target region is dependent on a position of the first sample in the target region, or
wherein the plurality of weights are determined based on a predetermined probability distribution, or wherein the plurality of weights are implemented with shift operations, or
wherein the plurality of weights are dependent on at least one of the following: a block size, a block shape, a block characteristic, or a sequence resolution, or
wherein the plurality of weights are indicated in one of the following: a sequence parameter set (SPS), a picture parameter set (PPS), or a slice header (SH), or
wherein whether to apply an BDOF process for MV refinement on the current video block is dependent on a first condition, or
whether to apply a BDOF process for sample adjustment on the current video block is dependent on a second condition (e.g. see at least in the existing BIO method, the same weight is granted to the gradient included in the window area, paragraph [0023], or see granting the weight according to the distance from the median value of the window… grant small weight to a pixel value positioned far away from the median value … grant a large weight to a pixel value positioned closer to the median value… the median value means gradient component positioned at the center of the window, e.g. see Figs. 17-18, paragraphs [0323]-[0340]).
Regarding claim 14, Lee further discloses wherein the gradient at a sample is determined based on a difference between two neighboring samples of the sample, or
wherein the gradient at a sample is determined based on a difference between two shifted neighboring samples of the sample, or
wherein the gradient at a sample is determined based on a first number of samples before the sample and a second number of samples after the sample (e.g. see Gx and/or Gy determined based on difference of pixel values, paragraphs [0201]-[0217], Abstract).
Regarding claim 16, Lee further discloses wherein the conversion includes encoding the current video block into the bitstream (e.g. see encoder, e.g. 100 in Fig. 1, for a conversion of input video, e.g. current block such as shown 1201 in Fig. 12, and bitstream).
Regarding claim 17, Lee further discloses wherein the conversion includes decoding the current video block from the bitstream (e.g. see decoder, e.g. 200 in Fig. 2, for a conversion of input video, e.g. current block such as shown 1201 in Fig. 12, and bitstream).
Regarding claims 18-19, the claims recite analogous limitations to the claims above and are therefore rejected on the same premise.
Claim(s) 20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by US 2015/0016537 A1 (“Karczewicz”).
Regarding claim 20, Karczewicz discloses a non-transitory computer-readable recording medium storing a bitstream of a video which is generated by a method performed by an apparatus for video processing, wherein the method comprises:
obtaining a plurality of weights, wherein the plurality of weights are used to weight a plurality of values for a metric at respective samples in a target region for a bi-directional optical flow (BDOF) process applied on a current video block of the video, the metric comprises at least one of a gradient or a difference of sample values, and the plurality of values for the metric are determined based on a plurality of reference video blocks of the current video block; and
generating the bitstream based on the plurality of weights (e.g. see at least 34 in Fig. 1, paragraphs [0033], [0038]; note: the non-transitory computer-readable medium merely serves as a support for the bitstream, see MPEP 2111.05. No patentable weight is given to the bitstream).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 5-6, 10-11, 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lee in view of US 2022/0201313 A1 (“Zhang”).
Regarding claim 5, although Lee discloses wherein a operation is used for determining the at least one offset (e.g. see Vx and/or Vy in Equation 19 or 21 for predicting the current pixel in Equation 20, paragraphs [0218]-[0224]), it is noted Lee differs from the present invention in that it fails to particularly disclose wherein a shift operation or a clip operation is used for determining the at least one offset. Zhang however, teaches wherein a shift operation or a clip operation is used for determining the at least one offset (e.g. see Vx and/or Vy in Equation 1-6-4, paragraphs [0200]-[0206]).
Therefore, given the teachings as a whole, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention, having the references of Lee and Zhang before him/her, to modify the Image processing method based on inter prediction mode, and apparatus of Lee with Zhang in order to promote selection between coding modes that may provide better coding performance.
Regarding claim 6, Lee in view of Zhang further teaches wherein the shift operation is applied on at least one of a numerator or a denominator of a term in an equation for determining the at least one offset, or
wherein the shift operation comprises left shifting by a predetermined number of bits, or
wherein the shift operation is applied in a predetermined order, or wherein the clip operation is applied on the at least one offset to update the at least one offset (Zhang: e.g. see Vx and/or Vy in Equation 1-6-4, paragraphs [0200]-[0206]). The motivation above in the rejection of claim 5 applies here.
Regarding claim 10, although Lee discloses a size used as a BDOF MV refinement (e.g. see BIO applied to the current block, paragraph [0224], e.g. see width and height of the current block, paragraphs [0229]-[0231]), it is noted Lee differs from the present invention in that it fails to particularly disclose wherein a subblock size used as a BDOF MV refinement subblock size is dependent on a condition. Zhang however, teaches wherein a subblock size used as a BDOF MV refinement subblock size is dependent on a condition (e.g. see dimension of sub-block, paragraphs [0212]-[0219], [0226]-[0234], [0238]-[0239]). The motivation above in the rejection of claim 5 applies here.
Regarding claim 11, Lee in view of Zhang further teaches wherein the subblock size is fixed, or
wherein the subblock size is dependent on at least one of the following:
a size of a current prediction unit (PU) comprising the current video block,
a size of a current coding unit (CU) comprising the current video block,
a characteristic of the plurality of reference video blocks,
a similarity of a plurality of predictors from the plurality of reference video blocks,
a distribution of difference between a plurality of predictors from the plurality of reference video blocks,
a temporal gradient of the plurality of reference video blocks, a spatial gradient of the plurality of reference video blocks, a prediction type,
an adjustment value determined in a first pass of a multi-pass decoder side motion vector refinement (DMVR),
an adjustment value determined in a second pass of the multi-pass DMVR, or a sequence resolution, or
wherein a height of the subblock size is dependent on a height or a width of the current video block, or a width of the subblock size is dependent on a height or a width of the current video block (Zhang: e.g. see at least sizes may be fixed, or CU has more than 64 luma samples and CU height and width larger than or equal to 8 luma samples, paragraphs [0212]-[0219], [0226]-[0234], [0238]-[0239]). The motivation above in the rejection of claim 5 applies here.
Regarding claim 15, although Lee discloses calculating a predictor for the current pixel (e.g. see paragraphs [0218]-[0224]), it is noted Lee differs from the present invention in that it fails to particularly disclose wherein a division operation is implemented with at least one non-division operation. Zhang however, teaches wherein a division operation is implemented with at least one non-division operation (e.g. see at least shifting in Equation 1-6-6, paragraphs [0200]-[0206], [0266]). The motivation above in the rejection of claim 5 applies here.
Claim(s) 7-8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lee in view of US 2022/0417522 A1 (“Huang”).
Regarding claim 7, although Lee discloses wherein performing the conversion based on the at least one offset comprises: performing the conversion based on the at least one offset (e.g. see encoder, e.g. 100 in Fig. 1, and/or decoder, e.g. 200 in Fig. 2, for a conversion of input video, e.g. current block 1201 in Fig. 12, and bitstream; thus, the encoder and/or decoder perform the conversion based on Vx and/or Vy in Equation 10 or 21 for predicting the current pixel in Equation 20), it is noted Lee differs from the present invention in that it fails to particularly disclose adjusting the at least one offset with at least one scaling factor; and performing the conversion based on the at least one adjusted offset. Huang however, teaches adjusting the at least one offset with at least one scaling factor; and performing the conversion based on the at least one adjusted offset (e.g. see a third pass can include performing sub-block BDOF MV refinement… a refined MV can be derived by applying BDOF to an 8x8 (or other size) grid sub-block… BDOF refinement can be applied to derive a scaled Vx and Vy without clipping, starting from the refined MV of the parent sub-block of the second pass, paragraph [0179]).
Therefore, given the teachings as a whole, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention, having the references of Lee and Huang before him/her, to modify the Image processing method based on inter prediction mode, and apparatus of Lee with Huang in order to improve refinement of motion vectors.
Regarding claim 8, Lee in view of Huang further teaches wherein a first offset in the at least one offset is adjusted by multiplying the first offset by a first scaling factor in the at least one scaling factor, and the first scaling factor is a real number, or
wherein a second offset in the at least one offset is adjusted by dividing the second offset by a second scaling factor in the at least one scaling factor, and the second scaling factor is a real number, or
wherein different offsets in the at least one offset are adjusted with different scaling factors in the at least one scaling factor, or
wherein the at least one scaling factor is dependent on at least one of the following: a block size, a sequence resolution, or a block characteristic (Huang: e.g. see a third pass can include performing sub-block BDOF MV refinement… a refined MV can be derived by applying BDOF to an 8x8 (or other size) grid sub-block… BDOF refinement can be applied to derive a scaled Vx and Vy without clipping, starting from the refined MV of the parent sub-block of the second pass, paragraph [0179]). The motivation above in the rejection of claim 7 applies here.
Claim(s) 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lee in view of US 2022/0368943 A1 (“Ye”).
Regarding claim 12, although Lee discloses the BDOF process (e.g. see bi-directional optical (BIO) flow, paragraphs [0006]-[0009]), it is noted Lee differs from the present invention in that it fails to particularly disclose wherein a first cost for evaluating a condition for the BDOF process is dependent on a second cost between the plurality of reference video blocks. Ye however, teaches wherein a first cost for evaluating a condition for the BDOF process is dependent on a second cost between the plurality of reference video blocks (e.g. see determining to skip BIO or not based on determining whether reference blocks are similar or dissimilar, paragraphs [0003]-[0004], [0112], [0117]-[0118], and see sub-block distortion that depend on the CU-level distortion, paragraph [0121]).
Therefore, given the teachings as a whole, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention, having the references of Lee and Ye before him/her, to modify the Image processing method based on inter prediction mode, and apparatus of Lee with Huang in order to selectively decide whether to enable BIO or not.
Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lee in view of US 2018/0262773 A1 (“Chuang”).
Regarding claim 13, although Lee discloses the at least one offset and MV determined by the BDOF process for a subblock of the current video block (e.g. see Vx and/or Vy in Equation 19 or 21 for predicting the current pixel in Equation 20, paragraphs [0218]-[0224] and Fig. 10), it is noted Lee differs from the present invention in that it fails to particularly disclose wherein a filter is applied on the at least one offset, or wherein a filter with a predetermined shape is applied on a MV determined by the BDOF process for a subblock of the current video block. Chuang however, teaches wherein a filter with a predetermined shape is applied on a MV determined by the BDOF process for a subblock of the current video block (e.g. see interpolation filters, paragraphs [0110]-[0121]).
Therefore, given the teachings as a whole, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention, having the references of Lee and Chuang before him/her, to modify the Image processing method based on inter prediction mode, and apparatus of Lee with Chuang in order to improve BIO video coding techniques.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 20220201328 A1, Galpin et al., METHOD AND APPARATUS FOR VIDEO ENCODING AND DECODING WITH OPTICAL FLOW BASED ON BOUNDARY SMOOTHED MOTION COMPENSATION
US 20220109871 A1, Galpin et al., METHOD AND APPARATUS FOR VIDEO ENCODING AND DECODING WITH BI-DIRECTIONAL OPTICAL FLOW ADAPTED TO WEIGHTED PREDICTION
US 20220116648 A1, Sethuraman et al., ENCODER, A DECODER AND CORRESPONDING METHODS
Any inquiry concerning this communication or earlier communications from the examiner should be directed to FRANCIS G GEROLEO whose telephone number is (571)270-7206. The examiner can normally be reached M-F 7:00 am - 3:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Anna M Momper can be reached at (571) 270-5788. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Francis Geroleo/Primary Examiner, Art Unit 3619