DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
Claims 1, 3-13, and 15-21 are pending. Claims 2 and 14 are canceled, and Claim 21 is new.
Response to Arguments
Applicant’s arguments and amendments, see p.9, filed 07/07/2026, with respect to the interpretation of Claims 1-12 under 35 U.S.C. 112(f) have been fully considered and are persuasive. Therefore, the interpretation of Claims 1-12 under 35 U.S.C. 112(f) has been withdrawn.
Applicant’s arguments, see p.9-11, filed 07/07/2026, with respect to the rejections of Claims 1-9 and 12-20 under 35 U.S.C. 101 have been fully considered but are not persuasive. Applicant argues that the Specification clearly articulates an improvement to the technologies of video compression and submits that any potential judicial exception is integrated into a practical application which includes an improvement to the technology which is properly reflected in the claims. Examiner respectfully disagrees because the limitation “estimate an initial search position by providing a first frame and a second frame of a video as input to a neural network that is trained to output an affine matrix value which represents the initial search position” is only the output of a search position as an affine matrix by a neural network which simply receives two image frames as input and is thus considered to be an insignificant extra-solution activity. Therefore, the judicial exception is not integrated into a practical application because the claim only recites these insignificant extra-solution activities wherein other additional recited elements in certain other claims are just only generic computer components. Accordingly, these additional elements do not integrate the abstract idea into a practical application because it is a field-of-use limitation that does not impose any meaningful limits on practicing the abstract idea due to the claim limitation only reciting simply performing motion estimation based on the search position output from the neural network as opposed to properly imposing meaningful limits on practicing the idea by providing further details on performing the motion estimation comprising real-time analysis and control of the search range based on the affine matrix value as recited in Claim 21. Furthermore, Examiner respectfully disagrees that the claims recite additional elements that amount to significantly more than the judicial exception such as through improvements to the functioning of a computer, technology, or technical field. The additional elements/steps amount to no more than insignificant extra-solution activities and are therefore not sufficient to amount to significantly more than the judicial exception. Therefore, the claim as a whole, recites an abstract idea.
Applicant’s arguments, see p.12-15, filed 07/07/2026, with respect to the rejections of Claims 1-20 under 35 U.S.C. 103 have been fully considered but are moot because Applicant’s amendments of the independent claims 1 and 13 has altered the scope of the claims, and therefore, necessitated new grounds of rejection which are presented below. Examiner has considered applicants arguments with respect to the new claim 21. However, arguments are moot due to new claims being presented and are therefore being analyzed as presented below. However, independent Claim 20 was not amended in which Applicant argues that Goswami does not use the affine transformation to indicate an initial search position, and, instead, the affine is applied to the results of the search to determine a quality of the predicted block. Examiner respectfully disagrees, and for further clarification, Goswami, FIGs. 20 and 23B and Paras. 133 and 135, teaches an inter prediction technique wherein an affine transform is applied on the predicted block of samples to generate an affine transformation of the predicted block within the search range given the motion vector and wherein the affine transformation may be described using equation 16 x´=ax+by+c and y´=dx+ey+f in which a, b, c, d, e, and f are six affine parameters and (x, y), (x′, y′) are the coordinates of the same pixel before and after the affine transform, i.e., the output of the affine transformation includes parameters indicative of a matrix representing position before and after affine transformation to represent the position of the initial search position being the predicted blocks. Accordingly, this action is made FINAL.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1, 3-9 and 12-13, 15-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more, and the claimed invention is directed to non-statutory subject matter as follows. The claims recite estimating a search position represented as an affine matrix value using a first and second frame of a video and estimating motion based on the search position.
Step 1:
With regard to Step 1, the instant claims are directed to a method, which is among the statutory categories of invention.
Step 2A – Prong 1:
With regard to Step 2A – Prong 1, for example in Claim 1, the limitations of "…perform motion estimation based on the initial search position", as drafted only involves mental processes or mathematical calculations, such as estimating motion based on the search position. That is, nothing in the above-described claim elements preclude the steps from practically being performed in the mind or on a piece of paper. If a claim limitation, under its broadest reasonably interpretation covers performance of the limitation in the mind or through mathematical calculations, but for the recitation of a generic apparatus components, such as a processor, computer program, or machine-readable media, then it falls within the "mental processes", which include concepts performed in the human mind, including an observation, evaluation, judgement, opinion, or mathematical calculations groupings of the abstract idea. Accordingly, the claim recites an abstract idea.
Step 2A – Prong 2:
The 2019 PEG defines the phrase “integration into a practical application” to require an additional element or a combination of additional elements in the claim to apply, rely on, or use the judicial exception. In the instant case, the additional elements in the claims do not apply, rely on, or use the judicial exception.
This judicial exception is not integrated into a practical application because the claim only recites the following additional steps "An image processing apparatus comprising: at least one processor; and a memory configured to store instructions which, when executed by the at least one processor, cause the image processing apparatus to: estimate an initial search position by providing a first frame and a second frame of a video as input to a neural network that is trained to output an affine matrix value which represents the initial search position…", i.e., applying the abstract idea on a generic computer and lacking a technological improvement. The other additional recited elements in certain other claims are just a non-transitory computer-readable storage medium with a processor, which are generic computer components. Accordingly, these additional elements do not integrate the abstract idea into a practical application because it is a field-of-use limitation that does not impose any meaningful limits on practicing the abstract idea. Therefore, the claim as a whole, recites an abstract idea.
Step 2B:
Because the claim fails under Step 2A, the claims are further evaluated under Step 2B. The claim herein does not include additional steps that are sufficient to amount to significantly more than the judicial exception because as discussed above with respect to integration of the abstract idea into practical application, the claim does not recite any additional elements/steps. Mere instructions to apply an exception using generic apparatus component, such as a processor or a neural network, cannot provide an inventive concept. The claim is not patent eligible. It should be noted that a similar analysis may be performed with respect to independent Claims 13 and 20.
Further, with regard to dependent Claims 3-9, 12, and 15-19 viewed individually, these additional steps are under their broadest reasonable interpretation, cover performance of the limitation in the mind and do not provide meaningful limitations to transform the abstract idea into a patent eligible application of the abstract idea such that the claims limitations amount to significantly more than the abstract idea itself. For example, the neural network comprising a convolutional neural network as recited in 3 2 or converting the size of the image frames to a given input size as recited in Claim 4 are only examples of routine and conventional image processing steps and do not amount to significantly more to consider as inventive steps. Accordingly, Claims 1, 3-9 and 12-13, and 15-20 are rejected under 35 U.S.C. 101.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 3, 5, 13, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Choi et al. ("Affine transformation-based deep frame prediction") in view of Yang et al. (US 20230343140 A1).
Regarding Claim 1, Choi teaches "An image processing apparatus comprising: at least one processor; and a memory configured to store instructions which, when executed by the at least one processor, cause the image processing apparatus to: estimate an initial search position by providing a first frame and a second frame of a video as input to a neural network that is trained to output an affine matrix value which represents the initial search position"; (Choi, Abstract, Section II, Section III, and Section IV-C, teaches a neural network model to estimate the current frame from two reference frames using affine transformation wherein a 2 X 3 affine transformation matrix can be formed for each x,y position and each ti using translation and isotropic scaling parameters for the affine transformation wherein spatial affine coordinate transformation is computed by applying the affine transformation matrix to the location x,y and show how the location is warped according to the estimated affine parameters with a new horizontal and vertical coordinate and wherein the affine motion parameters are estimated using the Spatial Transformer Network and mapped to the motion between the desired output frame and the input reference frames indicating positions where to apply the computed filters in which an operating system and GPU with memory are used for the experiment , i.e. an image processing apparatus comprises a processor and memory to estimate a search position as an affine matrix value being the indicated positions resulting from the affine transformation matrix output by a neural network which receives a first and second frame as input).
However, Choi does not explicitly teach "perform motion estimation based on the initial search position".
In an analogous field of endeavor, Yang teaches "perform motion estimation based on the initial search position"; (Yang, FIG. 3 and Paras. 11-13, 20-22, and 37, teaches searching for a block most similar to a macroblock to be matched in a previous frame based on matching criterion in a current frame wherein the matching criterion determines a motion vector corresponding to two macroblocks that are being matched and wherein the macroblock to be selected is determined using the search template a motion estimation search algorithm used to a three-step search method, a diamond search method, or a four-step search method, i.e., motion estimation module configured to perform motion estimation based on search position being the search algorithm and blocks).
It would have been obvious to one having ordinary skill in the art before the effective filing date to modify the invention of Choi by including the performance of motion estimation based on a search position taught by Yang. One of ordinary skill in the art would be motivated to combine the references since it improves estimation accuracy (Yang, Abstract, teaches the motivation of combination to be to correct cumulative error and improve estimation accuracy).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date.
Regarding Claim 3, the combination of references of Choi in view of Yang teaches "The image processing apparatus of claim 1, wherein the neural network comprises a convolutional neural network (CNN)"; (Choi, Para. 76, teaches the first network includes convolutional layers, i.e., the neural network comprises a CNN).
Regarding Claim 5, the combination of references of Choi in view of Yang teaches "The image processing apparatus of claim 1, wherein the instructions, when executed by the at least one processor, further cause the image processing apparatus to perform the motion estimation by moving a search range toward the initial search position"; (Yang, Paras. 22-26, teaches using a motion estimation search algorithm such as a three-step search method wherein a larger region that completely contains the macroblock is set as the search window and the center of the macroblock is a center point of the search window and then selecting a point with an optimal index as a center point of a next search and with the point obtained as a center, reducing a currently searched step length to half of a previously searched step length, then performing a similar search, and obtaining an optimal matching point, i.e., moving a search range toward the initial search position being the optimal index as the center point of the next search to perform motion estimation).
The proposed combination as well as the motivation for combining the Choi and Yang references presented in the rejection of Claim 1, applies to claim 5. Thus, the apparatus recited in claim 5 is met by Choi in view of Yang.
Claim 13 recites a method with steps corresponding to the elements of the system recited in Claim 1. Therefore, the recited steps of this claim are mapped to the proposed combination in the same manner as the corresponding elements in its corresponding system claim. Additionally, the rationale and motivation to combine the Choi and Yang references, presented in rejection of Claim 1, apply to this claim.
Claim 16 recites a method with steps corresponding to the elements of the system recited in Claim 5. Therefore, the recited steps of this claim are mapped to the proposed combination in the same manner as the corresponding elements in its corresponding system claim. Additionally, the rationale and motivation to combine the Choi and Yang references, presented in rejection of Claim 1, apply to this claim.
Claims 4 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Choi in view of Yang and Wu et al. (US 20210256290 A1).
Regarding Claim 4, the combination of references of Choi in view of Yang does not explicitly teach "The image processing apparatus of claim 1, wherein the first frame and the second frame have an original size, and wherein the instructions, when executed by the at least one processor, further cause the image processing apparatus to: convert the first frame into a converted first frame having an input size associated with the neural network, and convert the second frame into a converted second frame having the input size".
In an analogous field of endeavor, Wu teaches "The image processing apparatus of claim 1, wherein the first frame and the second frame have an original size, and wherein the instructions, when executed by the at least one processor, further cause the image processing apparatus to: convert the first frame into a converted first frame having an input size associated with the neural network, and convert the second frame into a converted second frame having the input size"; (Wu, Claim 3, teaches converting the first image and the second image included in the training data to a size enlarged or reduced by using the magnification wherein a plurality of the first images and second images generated by the size image generator are applied to the convolutional neural network as the input images, i.e., first and second frames have an original size which are converted to an input size associated with the neural network).
It would have been obvious to one having ordinary skill in the art before the effective filing date to modify the invention of Choi and Yang by including the conversion of frame sizes for input to a neural network taught by Wu. One of ordinary skill in the art would be motivated to combine the references since it improves performance of the CNN (Wu, Para. 161, teaches the motivation of combination to be to improve the performance of the CNN).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date.
Claim 15 recites a method with steps corresponding to the elements of the system recited in Claim 4. Therefore, the recited steps of this claim are mapped to the proposed combination in the same manner as the corresponding elements in its corresponding system claim. Additionally, the rationale and motivation to combine the Choi, Yang, and Wu references, presented in rejection of Claim 4, apply to this claim.
Claims 6 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Choi in view of Yang and Bernal et al. (US 20160142727 A1).
Regarding Claim 6, the combination of references of Choi in view of Yang does not explicitly teach "The image processing apparatus of claim 1, wherein the instructions, when executed by the at least one processor, further cause the image processing apparatus to: convert a value of the initial search position into a converted value having a target image size associated with the image processing apparatus".
In an analogous field of endeavor, Bernal teaches "The image processing apparatus of claim 1, wherein the instructions, when executed by the at least one processor, further cause the image processing apparatus to: convert a value of the initial search position into a converted value having a target image size associated with the image processing apparatus"; (Bernal, Para. 44, teaches the motion estimation module in the proposed system greatly reduces the size of the search neighborhood by centering it at the location indicated by the pixel-wise displacement values estimated by the motion direction and magnitude estimation module, i.e., search position is converted to a target size associated with the motion estimation module).
It would have been obvious to one having ordinary skill in the art before the effective filing date to modify the invention of Choi and Yang by including the converting the search position to a size for the motion estimation taught by Bernal. One of ordinary skill in the art would be motivated to combine the references since it allows for adapting to motion (Bernal, Para. 44, teaches the motivation of combination to be to adapt to wide ranges of inter-frame motion).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date.
Claim 17 recites a method with steps corresponding to the elements of the system recited in Claim 6. Therefore, the recited steps of this claim are mapped to the proposed combination in the same manner as the corresponding elements in its corresponding system claim. Additionally, the rationale and motivation to combine the Choi, Yang, and Bernal references, presented in rejection of Claim 6, apply to this claim.
Claims 7-8 and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Choi in view of Yang, Goswami, and Molchanov et al. (US 20180365532 A1).
Regarding Claim 7, the combination of references of Choi in view of Yang and Goswami does not explicitly teach "The image processing apparatus of claim 1, wherein the instructions, when executed by the at least one processor, further cause the image processing apparatus to perform a differentiable training process to train the neural network using a back-propagation method".
In an analogous field of endeavor, Molchanov teaches "The image processing apparatus of claim 1, wherein the instructions, when executed by the at least one processor, further cause the image processing apparatus to perform a differentiable training process to train the neural network using a back-propagation method"; (Molchanov, Para. 38, teaches the use of the soft-argmax layer to extract landmark locations from pixel-level predictions makes the entire sequential multi-tasking system differentiable and trainable end-to-end through back-propagation, i.e., training module is differentiable and trains the neural network using a back-propagation method).
It would have been obvious to one having ordinary skill in the art before the effective filing date to modify the invention of Choi, Yang, and Goswami by including the training being differentiable and using back-propagation taught by Molchanov. One of ordinary skill in the art would be motivated to combine the references since it enhances learning (Molchanov, Para. 38, teaches the motivation of combination to be to enhance landmark localization learning).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date.
Regarding Claim 8, the combination of references of Choi in view of Yang, Goswami, and Molchanov teaches "The image processing apparatus of claim 7, wherein the instructions, when executed by the at least one processor, further cause the image processing apparatus to train the neural network using unsupervised learning"; (Molchanov, Para. 59, teaches an unsupervised learning technique may be employed to train the neural network model to generate predictions that are consistent when different transformations are applied to the input image, i.e., training module configured to train the neural network using unsupervised learning).
The proposed combination as well as the motivation for combining the Choi, Yang, Goswami, and Molchanov references presented in the rejection of Claim 7, applies to claim 8. Thus, the apparatus recited in claim 8 is met by Choi in view of Yang, Goswami, and Molchanov.
Claim 18 recites a method with steps corresponding to the elements of the system recited in Claim 7. Therefore, the recited steps of this claim are mapped to the proposed combination in the same manner as the corresponding elements in its corresponding system claim. Additionally, the rationale and motivation to combine the Choi, Yang, Goswami, and Molchanov references, presented in rejection of Claim 7, apply to this claim.
Claim 19 recites a method with steps corresponding to the elements of the system recited in Claim 8. Therefore, the recited steps of this claim are mapped to the proposed combination in the same manner as the corresponding elements in its corresponding system claim. Additionally, the rationale and motivation to combine the Choi, Yang, Goswami, and Molchanov references, presented in rejection of Claim 7, apply to this claim.
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Choi in view of Yang, Goswami, Molchanov, and Gordon et al. (US 20210279511 A1).
Regarding Claim 9, the combination of references of Choi in view of Yang, Goswami, and Molchanov does not explicitly teach "The image processing apparatus of claim 7, wherein the instructions, when executed by the at least one processor, further cause the image processing apparatus to train the neural network so that a loss function between a first training frame and a predicted image is minimized, wherein the predicted image is predicted from a second training frame using the neural network”.
In an analogous field of endeavor, Gordon teaches "The image processing apparatus of claim 7, wherein the instructions, when executed by the at least one processor, further cause the image processing apparatus to train the neural network so that a loss function between a first training frame and a predicted image is minimized"; (Gordon, Paras. 30-32 and Para. 41, teaches a pixelwise photometric loss between a first image and the predicted image at the first time point wherein the task loss engine corresponding to the neural network can determine a supervised parameter update according to the error, i.e., training neural network so that a loss between a first training frame and predicted image is minimized);
"wherein the predicted image is predicted from a second training frame using the neural network"; (Gordon, Paras. 30-32, teaches the predicted image at the first time point generated by interpolating the pixels of the second image using the predicted motion, i.e., predicted image is predicted from a second training frame using the neural network).
It would have been obvious to one having ordinary skill in the art before the effective filing date to modify the invention of Choi, Yang, Goswami, and Molchanov by including the training minimizing a loss between images taught by Gordon. One of ordinary skill in the art would be motivated to combine the references since it determines a penalty for the differences (Gordon, Para. 32, teaches the motivation of combination to be to determine a penalty on the difference in RGB space and structural similarity).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date.
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Choi in view of Yang, Goswami, Molchanov, Gordon, Chen et al. (US 20210166396 A1), Park et al. (US 20130113950 A1), and Bleyer et al. (US 20240380913 A1).
Regarding Claim 10, the combination of references of Choi in view of Yang, Goswami, Molchanov, and Gordon does not explicitly teach "The image processing apparatus of claim 9, instructions, when executed by the at least one processor, further cause the image processing apparatus to: receive the first training frame and the second training frame as input; perform affine transformation on the second training frame based on the affine matrix value; perform the motion estimation on the affine-transformed second training frame; output a motion kernel; and perform motion compensation based on a result of the motion estimation; and output the predicted image".
In an analogous field of endeavor, Chen teaches "The image processing apparatus of claim 9, instructions, when executed by the at least one processor, further cause the image processing apparatus to: receive the first training frame and the second training frame as input; perform affine transformation on the second training frame based on the affine matrix value"; (Chen, Abstract and Para. 61, teaches selecting a current image frame and determining a reference image frame before the current image frame and performing an affine transformation on the current image frame with reference to an affine transformation relationship between the first location information and a target object key point template to obtain a target object diagram wherein the affine transformation on the current image frame is performed according to the transformation matrix, i.e., receive two image frames as input and perform affine transformation on the second frame based on the matrix).
It would have been obvious to one having ordinary skill in the art before the effective filing date to modify the invention of Choi in view of Yang, Goswami, Molchanov, and Gordon by including the receiving of a first and second frame and performing affine transformation on the second frame based on the affine matrix value taught by Chen. One of ordinary skill in the art would be motivated to combine the references since it improves accuracy (Chen, Para. 100, teaches the motivation of combination to be to improve accuracy).
However, the combination of references of Choi in view of Yang, Goswami, Molchanov, Gordon, and Chen does not explicitly teach “perform the motion estimation on the affine-transformed second training frame; output a motion kernel; and perform motion compensation based on a result of the motion estimation; and output the predicted image".
In an analogous field of endeavor, Park teaches "perform the motion estimation on the affine-transformed second training frame; output a motion kernel"; (Park, Para. 76, teaches estimating a motion vector of each block between the first and second images and estimating a motion blur kernel of each block using the motion vector of each block, i.e., perform motion estimation of the second frame and output a motion kernel).
It would have been obvious to one having ordinary skill in the art before the effective filing date to modify the invention of Choi in view of Yang, Goswami, Molchanov, Gordon, and Chen wherein the second image is an affine-transformed image by including the motion estimation of the second image to output a motion kernel taught by Park. One of ordinary skill in the art would be motivated to combine the references since it prevents motion halting (Park, Para. 8, teaches the motivation of combination to be to prevent motion halting during video shooting).
However, the combination of references of Choi in view of Yang, Goswami, Molchanov, Gordon, Chen, and Park does not explicitly teach "and perform motion compensation based on a result of the motion estimation; and output the predicted image".
In an analogous field of endeavor, Bleyer teaches "and perform motion compensation based on a result of the motion estimation; and output the predicted image"; (Bleyer, Abstract and Paras. 60-61, teaches optical flow-based motion compensation utilizes inertial tracking data, the affine transformation-compensated image and a current image as inputs wherein an optical flow between the affine transformation-compensated image and the current image is determined and applying the optical flow to the affine transformation-compensated image to obtain the motion-compensated image , i.e., perform motion compensation based on a result of motion estimation being the optical flow to output the predicted image being the motion-compensated image).
It would have been obvious to one having ordinary skill in the art before the effective filing date to modify the invention of Choi in view of Yang, Goswami, Molchanov, Gordon, Chen, and Park by including the motion compensation based on motion estimation to output a predicted image taught by Bleyer. One of ordinary skill in the art would be motivated to combine the references since it mitigates artifacts (Bleyer, Para. 27, teaches the motivation of combination to be to mitigate artifacts from low signal input images).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date.
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Choi in view of Yang and Seo (US 20170201751 A1).
Regarding Claim 12, the combination of references of Choi in view of Yang does not explicitly teach "The image processing apparatus of claim 1, wherein the image processing apparatus comprises at least one of a video codec device and an image signal processing (ISP) device".
In an analogous field of endeavor, Seo teaches "The image processing apparatus of claim 1, wherein the image processing apparatus comprises at least one of a video codec device and an image signal processing (ISP) device"; (Seo, Para. 57, teaches an ISP module of the application processor and an external video codec communicative coupled to the application processor).
It would have been obvious to one having ordinary skill in the art before the effective filing date to modify the invention of Choi and Yang by including the apparatus comprising a video codec device and an ISP device taught by Seo. One of ordinary skill in the art would be motivated to combine the references since it improves coding efficiency (Seo, Para. 75, teaches the motivation of combination to be to improve coding efficiency).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date.
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Seo in view of Na et al. (US 20210021823 A1), Lee et al. (US 20210295532 A1), and Goswami.
Regarding Claim 20, Seo teaches "An apparatus for optimizing motion estimation of a video in a video codec device or an image signal processing (ISP) device, the apparatus comprising: one or more processors configured to: read a current frame and a reference frame of the video from a memory before the motion estimation of the video is performed by the video codec device or the ISP device"; (Seo, Abstract, teaches an application processor including a processor, video coder, a memory, and a GPU to modify , based on a modification parameter, a reference image to generate a modified reference image, the reference image being configured to be stored in the memory, and determining motion information associated with a coding block of a current image, the current image and the reference image being temporally different wherein the motion information is associated with at least one of the modification parameter and the modified reference image, i.e., processors read a current image frame and a reference image frame from memory before motion estimation of the video is performed).
However, Seo does not explicitly teach “obtain an initial search position for the motion estimation using a neural network, and provide the initial search position to the video codec device or the ISP device, wherein the neural network is trained to receive the current frame and the reference frame as input and to output an affine matrix value which represents the initial search position”.
In an analogous field of endeavor, Na teaches "obtain an initial search position for the motion estimation using a neural network"; (Na, Para. 18, teaches a video decoding apparatus including a CNN setting unit configured to set input data containing a search region in at least one reference picture and a CNN execution unit configured to generate a motion vector of a current block or predicted pixels of the current block by applying a CNN having predetermined filter coefficients to the input data, i.e., obtain an initial search position being the search region for motion estimation using a neural network);
"and provide the initial search position to the video codec device or the ISP device"; (Na, Para. 18, teaches a video decoding apparatus including a CNN setting unit configured to set input data containing a search region in at least one reference picture and a CNN execution unit configured to generate a motion vector of a current block or predicted pixels of the current block by applying a CNN having predetermined filter coefficients to the input data, i.e., search position being the search region is provided to the video codec device being the video decoding apparatus).
It would have been obvious to one having ordinary skill in the art before the effective filing date to modify the invention of Seo by including the obtaining of a search position for motion estimation and providing the search position to a codec device taught by Na. One of ordinary skill in the art would be motivated to combine the references since it improves prediction accuracy (Na, Para. 15, teaches the motivation of combination to be to improve prediction accuracy).
However, the combination of references of Seo in view of Na does not explicitly teach "wherein the neural network is trained to receive the current frame and the reference frame as input and to output an affine matrix value which represents the initial search position”.
In an analogous field of endeavor, Lee teaches "wherein the neural network is trained to receive the current frame and the reference frame as input"; (Lee, FIG. 5 and Abstract and Paragraphs 52-53, teaches an apparatus including a memory and processor wherein a first input image and second input image are provided to a neural network in which a searching region in a second input image is extracted using the neural network and estimate a position of the target in the searching region from the score matrix, i.e., a first and second frame of video is input to a neural network trained to output a search position).
It would have been obvious to one having ordinary skill in the art before the effective filing date to modify the invention of Seo and Na by including the current and reference frames input to the neural network taught by Lee. One of ordinary skill in the art would be motivated to combine the references since it improves tracking accuracy (Lee, Para. 42, teaches the motivation of combination to be to improve accuracy in tracking targets).
However, the combination of references of Seo in view of Na and Lee does not explicitly teach "and to output an affine matrix value which represents the initial search position".
In an analogous field of endeavor, Goswami teaches "and to output an affine matrix value which represents the initial search position"; (Goswami, Paras. 129-130 and 135, teaches determining the best matching reference block in the search range based on a prediction error wherein an affine transform is applied on the predicted block of samples in which affine transform parameters may be selected from among a plurality of affine transform parameters used to generate an affine transformation of predicted block and a prediction error determined for each of the affine transformations, i.e., search position output as an affine matrix value being the search range block undergoing affine transformation. For further clarification, Goswami, FIGs. 20 and 23B and Paras. 133 and 135, teaches an inter prediction technique wherein an affine transform is applied on the predicted block of samples to generate an affine transformation of the predicted block within the search range given the motion vector and wherein the affine transformation may be described using equation 16 x´=ax+by+c and y´=dx+ey+f in which a, b, c, d, e, and f are six affine parameters and (x, y), (x′, y′) are the coordinates of the same pixel before and after the affine transform, i.e., the output of the affine transformation includes parameters indicative of a matrix representing position before and after affine transformation to represent the position of the initial search position being the predicted blocks).
It would have been obvious to one having ordinary skill in the art before the effective filing date to modify the invention of Seo, Na, and Lee by including the search position being output as an affine matrix value taught by Goswami. One of ordinary skill in the art would be motivated to combine the references since it allows parameters that correspond to the lowest prediction (Goswami, Para. 135, teaches the motivation of combination to be to select affine parameters that correspond to the lowest prediction).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date.
Claim 21 is rejected under 35 U.S.C. 103 as being unpatentable over Choi in view of Yang, and Van Leuven et al. (US 20180124425 A1).
Regarding Claim 21, the combination of references of Choi in view of Yang does not explicitly teach "The image processing apparatus of claim 1, wherein, during the performing of the motion estimation, a search range of the motion estimation is controlled in real time based on the affine matrix value".
In an analogous field of endeavor, Van Leuven teaches "The image processing apparatus of claim 1, wherein, during the performing of the motion estimation, a search range of the motion estimation is controlled in real time based on the affine matrix value"; (Van Leuven, Paras. 84-86, teaches the displacement of the input block relative to the reference block can be any of an affine transformation in which the estimation of dense motion fields can be used as an additional input to a blockmatching algorithm to guide the search operation by reducing the search space they need to explore or as augmenting information to improve accuracy wherein the trainable module would be data adaptive, i.e., search range of the motion estimation is controlled in real time based on the affine value being the reduction of search space from estimated motion fields resulting from the affine transformation displacements to guide the search operation).
It would have been obvious to one having ordinary skill in the art before the effective filing date to modify the invention of Choi and Yang wherein the affine transformation indicates an affine matrix value by including the motion estimation including controlling a search range of the motion estimation based on the affine value taught by Van Leuven. One of ordinary skill in the art would be motivated to combine the references since it improves efficiency and accuracy (Van Leuven, Para. 86, teaches the motivation of combination to be improve efficiency and improve accuracy of the search operation).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date.
Allowable Subject Matter
Claim 11 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The examiner’s stated reason for indication of allowable subject matter can be found in the Non-Final Rejection Office Action dated 04/07/2026.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANDREW STEVEN BUDISALICH whose telephone number is (703)756-5568. The examiner can normally be reached Monday - Friday 8:30am-5:00pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached on (571) 272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx
for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ANDREW S BUDISALICH/Examiner, Art Unit 2662
/AMANDEEP SAINI/Supervisory Patent Examiner, Art Unit 2662