Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-6, 8, 9, 13-15, 18 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Sezer et al. (US PGPUB 20160028967) in view of Li et al. (US PGPUB 20230252608).
[Claim 1]
Sezer teaches a method of image processing based on global motion estimation, the method comprising:
estimating global motion parameters corresponding to components of a global motion between a current image frame and a reference image frame by executing a global motion estimation model that input the current image frame and the reference image frame (Paragraph 44, The system 300 then performs feature point selection (315) to select certain of the originally identified feature points to use for image alignment. For example, the system 300 may process movement of the feature points over the image sequence to identify feature points that are indicative of global camera movement throughout the image capture sequence and avoid feature points that are indicative of object movement in the scene. In particular, the system 300 may estimate a transformation (e.g., an affine matrix transformation or geometric transformation matrix) of the feature points between images to classify the movement between the images. For example, during the capture sequence, the camera may move towards or away from the scene, rotate, move up or down. Each of these motions can be accounted for using an affine matrix transformation (e.g., linear translation for up/down or left/right movements, scaling for moving towards/away movements, and/or rotational affine transformations for camera rotation.)
generating a geometric transformation matrix by combining the global motion parameters (Paragraph 44, In particular, the system 300 may estimate a transformation (e.g., an affine matrix transformation) of the feature points between images to classify the movement between the images. For example, during the capture sequence, the camera may move towards or away from the scene, rotate, move up or down. Each of these motions can be accounted for using an affine matrix transformation (e.g., linear translation for up/down or left/right movements, scaling for moving towards/away movements, and/or rotational affine transformations for camera rotation.); and
generating at least one of an output image and an output video using the geometric transformation matrix (Paragraph 49, Further, the system 300 performs a global affine transformation estimation (320) for the remaining set of images using the selected feature points to compute a compute affine alignment matrices, M.sub.1 to M.sub.N, to warp, zoom, rotate, move, and/or shift each of the images Img.sub.i to Img.sub.N to align with Img.sub.0. In particular, the transformation estimation corrects alignment errors to register the images for combination by accounting for global movement of the camera across the sequence of images while ignoring independent movement of objects in the scene. The system 300 provides a single image capture of a scene, using multi-frame capture and blending to enhance the quality of the output provided to the user).
Sezer teaches a global estimation model but fails to teach estimating global motion comprising one or more neural networks. However Li teaches motion vectors 312a-312q can be the same as or similar to the motion vectors 310 of FIG. 3A. Each motion vector 312a-312q represents global motion between two image frames of the different set of image frames. The global motion includes both rotation motion and translation motion (Paragraph 77). The image frames 410 are provided to a small motion optical flow network (OfNet) operation 420, which generally operates to estimate small camera motion due to handshake when the image frames 410 are captured. In some embodiments, the small motion OfNet operation 420 includes a neural network. The neural network architecture 420a of FIG. 4B illustrates one specific example architecture that may be used by the small motion OfNet operation 420. Note, however, that any other trained machine learning model may be used here (Paragraph 86).
Therefore taking the combined teachings of Sezer and Li, it would be obvious to one skilled in the art before the effective filing date of the invention to have been motivated to have estimated global motion comprising one or more neural networks in order to automatically learn the most relevant features and representations directly from the raw data eliminating the need for any manual feature extraction.
[Claim 2]
Sezer teaches wherein the generating least one of the output image and the output video further comprises: generating the output video by encoding the current image frame and the reference image frame using the geometric transformation matrix (Paragraph 49, Further, the system 300 performs a global affine transformation estimation (320) for the remaining set of images using the selected feature points to compute a compute affine alignment matrices, M.sub.1 to M.sub.N, to warp, zoom, rotate, move, and/or shift each of the images Img.sub.i to Img.sub.N to align with Img.sub.0. In particular, the transformation estimation corrects alignment errors to register the images for combination by accounting for global movement of the camera across the sequence of images while ignoring independent movement of objects in the scene. The system 300 provides a single image capture of a scene, using multi-frame capture and blending to enhance the quality of the output provided to the user and Paragraph 50, Embodiments of the present disclosure enable seamless fusion of multiple frames to form a single image (or a sequence of frame for video data) with improved resolution and reduced noise levels).
[Claim 3]
Sezer teaches wherein the generating the at least one of the output image and the output video further comprises: inputting the current image frame, the reference image frame, and the geometric transformation matrix into a video codec (fig. 6, blending 615) configured to execute one or more operations using the geometric transformation matrix (Paragraphs 49 and 54).
[Claim 4]
Sezer teaches wherein the video codec is configured to execute at least one of a translation mode using a global translation motion (Paragraph 49, Further, the system 300 performs a global affine transformation estimation (320) for the remaining set of images using the selected feature points to compute a compute affine alignment matrices, M.sub.1 to M.sub.N, to warp, zoom, rotate, move, and/or shift each of the images Img.sub.i to Img.sub.N to align with Img.sub.0. In particular, the transformation estimation corrects alignment errors to register the images for combination by accounting for global movement of the camera across the sequence of images while ignoring independent movement of objects in the scene.)).
[Claim 5]
Sezer teaches wherein the global motion estimation model comprises one or more sub-models corresponding to the affine mode (Paragraph 49, Further, the system 300 performs a global affine transformation estimation (320) for the remaining set of images using the selected feature points to compute a compute affine alignment matrices, M.sub.1 to M.sub.N, to warp, zoom, rotate, move, and/or shift each of the images Img.sub.i to Img.sub.N to align with Img.sub.0. In particular, the transformation estimation corrects alignment errors to register the images for combination by accounting for global movement of the camera across the sequence of images while ignoring independent movement of objects in the scene).
[Claim 6]
Sezer teaches wherein the geometric transformation matrix is an affine transformation matrix (Paragraph 49).
[Claim 8]
Sezer teaches wherein the generating of the at least one of the output image and the output video comprises: generating the output image by driving an image signal processor (ISP, e.g. controller 210) using the geometric transformation matrix (Paragraph 51, FIGS. 6A to 6F illustrate block diagrams of image alignment and combination systems according to illustrative embodiments of this disclosure. For example, the image alignment and combination systems may be implemented by the controller 210 in cooperation with the memory 230 of the image processing device 200 in FIG. 2. As illustrated in FIG. 6A, system 600 combines global/local registration 605 with motion artifact processing 610, for example, such as the image registration techniques as discussed above with regard to FIG. 3, to improve the blending 615 of images together to produce a final result 620 having improved image detail).
[Claim 9]
Li teaches wherein the geometric transformation matrix is a homography transformation matrix (Paragraph 88, The output from the small motion OfNet operation 420 is provided to a homograph matrix operation 430. The homograph matrix operation 430 identifies motion using the optical flow maps provided by the small motion OfNet operation 420. In some cases, the homograph matrix operation 430 can generate a homograph matrix representing the motion. The output from the homograph matrix operation 430 is provided to a motion vector generator operation 440, which generally decomposes the homograph matrix to generate a motion vector representing the global motion of each pixel from one frame to another frame.) in order to have a compact matrix to represent all transformation including rotation, translation, scaling, and perspective distortion. This allows multiple transformations to be chained into one operation, saving processor time.
[Claim 13]
Sezer teaches an electronic device comprising:
a camera (fig. 2 , camera 240) configured to generate a current image frame and a reference image frame (Paragraph 41, several key components are utilized by the system 300 to compute the best alignment between a currently processed image frame and a reference image (Img.sub.0) 325 from the captured sequence);
a memory storing one or more instructions (Paragraph 13);
a video codec (fig. 6, blending 615); and
at least one processor operatively coupled to the memory, the camera, and the video codec, wherein the one or more instructions, when executed by the at least one processor (Paragraph 13), cause the electronic device to:
estimate global motion parameters corresponding to components of a global motion between the current image frame and the reference image frame by executing the global motion estimation model comprising one or more neural networks that input the reference image frame (Paragraph 44, The system 300 then performs feature point selection (315) to select certain of the originally identified feature points to use for image alignment. For example, the system 300 may process movement of the feature points over the image sequence to identify feature points that are indicative of global camera movement throughout the image capture sequence and avoid feature points that are indicative of object movement in the scene. In particular, the system 300 may estimate a transformation (e.g., an affine matrix transformation or geometric transformation matrix) of the feature points between images to classify the movement between the images. For example, during the capture sequence, the camera may move towards or away from the scene, rotate, move up or down. Each of these motions can be accounted for using an affine matrix transformation (e.g., linear translation for up/down or left/right movements, scaling for moving towards/away movements, and/or rotational affine transformations for camera rotation.);
generate a geometric transformation matrix by combining the global motion parameters (Paragraph 44, In particular, the system 300 may estimate a transformation (e.g., an affine matrix transformation) of the feature points between images to classify the movement between the images. For example, during the capture sequence, the camera may move towards or away from the scene, rotate, move up or down. Each of these motions can be accounted for using an affine matrix transformation (e.g., linear translation for up/down or left/right movements, scaling for moving towards/away movements, and/or rotational affine transformations for camera rotation.); and
control the video codec to generate an output video using the geometric transformation matrix (Paragraphs 49 and 54).
Sezer teaches a global estimation model but fails to teach storing a global motion estimation model based on a neural network and estimating global motion comprising one or more neural networks. However Li teaches The memory 130 can include a volatile and/or non-volatile memory. For example, the memory 130 can store commands or data related to at least one other component of the electronic device 101. In accordance with this disclosure, the memory 130 can store software and/or a program 140 (Note, ofnet is a software program used for estimating motion, Paragraph 47). The program 140 includes, for example, a kernel 141, middleware 143, an application programming interface (API) 145, and/or an application program (or “application”) 147. At least a portion of the kernel 141, middleware 143, or API 145 may be denoted an operating system (OS). motion vectors 312a-312q can be the same as or similar to the motion vectors 310 of FIG. 3A. Each motion vector 312a-312q represents global motion between two image frames of the different set of image frames. The global motion includes both rotation motion and translation motion (Paragraph 77). The image frames 410 are provided to a small motion optical flow network (OfNet) operation 420, which generally operates to estimate small camera motion due to handshake when the image frames 410 are captured. In some embodiments, the small motion OfNet operation 420 includes a neural network. The neural network architecture 420a of FIG. 4B illustrates one specific example architecture that may be used by the small motion OfNet operation 420. Note, however, that any other trained machine learning model may be used here (Paragraph 86).
Therefore taking the combined teachings of Sezer and Li, it would be obvious to one skilled in the art before the effective filing date of the invention to have been motivated to have stored a global motion estimation model based on a neural network and estimating global motion comprising one or more neural networks in order to automatically learn the most relevant features and representations eliminating the need for any manual feature extraction.
[Claim 14]
Sezer teaches wherein the video codec is configured to execute at least one of a translation mode using a global translation motion (Paragraph 49, Further, the system 300 performs a global affine transformation estimation (320) for the remaining set of images using the selected feature points to compute a compute affine alignment matrices, M.sub.1 to M.sub.N, to warp, zoom, rotate, move, and/or shift each of the images Img.sub.i to Img.sub.N to align with Img.sub.0. In particular, the transformation estimation corrects alignment errors to register the images for combination by accounting for global movement of the camera across the sequence of images while ignoring independent movement of objects in the scene.)).
[Claim 15]
Sezer teaches wherein the geometric transformation matrix is an affine transformation matrix (Paragraph 49).
[Claim 18]
Sezer teaches an electronic device comprising:
a camera (fig. 2, camera 240) configured to generate a current image frame and a reference image frame (Paragraph 41, several key components are utilized by the system 300 to compute the best alignment between a currently processed image frame and a reference image (Img.sub.0) 325 from the captured sequence);
a memory storing one or more instructions (Paragraph 33);
an image signal processor (615);
at least one processor (210) operatively coupled to the memory, the camera, and the ISP; wherein the one or more instructions, when executed by the at least one processor (Paragraph 33), cause the electronic device to:
estimate global motion parameters corresponding to components of a global motion between the current image frame and the reference image frame by executing the global motion estimation model based on the current image frame and the reference image frame (Paragraph 44, The system 300 then performs feature point selection (315) to select certain of the originally identified feature points to use for image alignment. For example, the system 300 may process movement of the feature points over the image sequence to identify feature points that are indicative of global camera movement throughout the image capture sequence and avoid feature points that are indicative of object movement in the scene. In particular, the system 300 may estimate a transformation (e.g., an affine matrix transformation or geometric transformation matrix) of the feature points between images to classify the movement between the images. For example, during the capture sequence, the camera may move towards or away from the scene, rotate, move up or down. Each of these motions can be accounted for using an affine matrix transformation (e.g., linear translation for up/down or left/right movements, scaling for moving towards/away movements, and/or rotational affine transformations for camera rotation);
generate a geometric transformation matrix by combining the global motion parameters (Paragraph 44, In particular, the system 300 may estimate a transformation (e.g., an affine matrix transformation) of the feature points between images to classify the movement between the images. For example, during the capture sequence, the camera may move towards or away from the scene, rotate, move up or down. Each of these motions can be accounted for using an affine matrix transformation (e.g., linear translation for up/down or left/right movements, scaling for moving towards/away movements, and/or rotational affine transformations for camera rotation.); and
control the ISP to generate an output image using the geometric transformation matrix Paragraphs 49 and 54).
Sezer teaches a global estimation model but fails to teach storing a global motion estimation model based on a neural network and estimating global motion comprising one or more neural networks. However Li teaches a memory 130 can include a volatile and/or non-volatile memory. For example, the memory 130 can store commands or data related to at least one other component of the electronic device 101. In accordance with this disclosure, the memory 130 can store software and/or a program 140 (Note, ofnet is a software program used for estimating motion, Paragraph 47). The program 140 includes, for example, a kernel 141, middleware 143, an application programming interface (API) 145, and/or an application program (or “application”) 147. At least a portion of the kernel 141, middleware 143, or API 145 may be denoted an operating system (OS). motion vectors 312a-312q can be the same as or similar to the motion vectors 310 of FIG. 3A. Each motion vector 312a-312q represents global motion between two image frames of the different set of image frames. The global motion includes both rotation motion and translation motion (Paragraph 77). The image frames 410 are provided to a small motion optical flow network (OfNet) operation 420, which generally operates to estimate small camera motion due to handshake when the image frames 410 are captured. In some embodiments, the small motion OfNet operation 420 includes a neural network. The neural network architecture 420a of FIG. 4B illustrates one specific example architecture that may be used by the small motion OfNet operation 420. Note, however, that any other trained machine learning model may be used here (Paragraph 86).
Therefore taking the combined teachings of Sezer and Li, it would be obvious to one skilled in the art before the effective filing date of the invention to have been motivated to have stored a global motion estimation model based on a neural network and estimating global motion comprising one or more neural networks in order to automatically learn the most relevant features and representations eliminating the need for any manual feature extraction.
[Claim 20]
Li teaches wherein the geometric transformation matrix is a homography transformation matrix (Paragraph 88, The output from the small motion OfNet operation 420 is provided to a homograph matrix operation 430. The homograph matrix operation 430 identifies motion using the optical flow maps provided by the small motion OfNet operation 420. In some cases, the homograph matrix operation 430 can generate a homograph matrix representing the motion. The output from the homograph matrix operation 430 is provided to a motion vector generator operation 440, which generally decomposes the homograph matrix to generate a motion vector representing the global motion of each pixel from one frame to another frame.) in order to have a compact matrix to represent all transformation including rotation, translation, scaling, and perspective distortion. This allows multiple transformations to be chained into one operation, saving processor time.
Allowable Subject Matter
Claims 7, 10-12, 16, 17 and 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The prior art fails to teach or suggest as in claims 7, 16 and 19, “determining one or more function values by substituting one or more global motion
parameters into one or more functions; and determining one or more elements of the geometric transformation matrix by combining the global motion parameters based on (i) operations between the global motion parameters, (ii) operations between the global motion parameters and the one or more function values, (iii) operations between a plurality of function values of the one or more function values, or (iv) a combination thereof” and
claims 10 and 17, “generating a scaled current image frame by scaling the current image frame to a target size; and generating a scaled reference image frame by scaling the reference image frame to the target size, wherein the global motion estimation model is executed based on the scaled current image frame and the scaled reference image frame”, and
claim 12, “a first estimation model configured to estimate first global motion parameters and a second estimation model configured to estimate second global motion parameters, and the geometric transformation matrix comprises an affine transformation matrix determined by a combination of the first global motion parameters and a homography transformation matrix determined by a combination of the second global motion parameters”.
Claim 11 is dependent upon claim 10.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to YOGESH K AGGARWAL whose telephone number is (571)272-7360. The examiner can normally be reached Monday - Friday 9:30-6.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sinh Tran can be reached at 5712727564. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/YOGESH K AGGARWAL/Primary Examiner, Art Unit 2637