Prosecution Insights
Last updated: August 18, 2026
Application No. 19/097,449

IMAGE PROCESSING METHOD BASED ON GLOBAL MOTION ESTIMATION AND DEVICE USING THE SAME

Non-Final OA §103
Filed
Apr 01, 2025
Priority
Nov 06, 2024 — RE 10-2024-0156411
Examiner
AGGARWAL, YOGESH K
Art Unit
2637
Tech Center
2600 — Communications
Assignee
Samsung Electronics Co., Ltd.
OA Round
1 (Non-Final)
90%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
96%
With Interview

Examiner Intelligence

Grants 90% — above average
90%
Career Allowance Rate
1020 granted / 1135 resolved
+27.9% vs TC avg
Moderate +7% lift
Without
With
+6.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
26 currently pending
Career history
1160
Total Applications
across all art units

Statute-Specific Performance

§101
4.6%
-35.4% vs TC avg
§103
52.1%
+12.1% vs TC avg
§102
36.9%
-3.1% vs TC avg
§112
3.9%
-36.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1135 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-6, 8, 9, 13-15, 18 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Sezer et al. (US PGPUB 20160028967) in view of Li et al. (US PGPUB 20230252608). [Claim 1] Sezer teaches a method of image processing based on global motion estimation, the method comprising: estimating global motion parameters corresponding to components of a global motion between a current image frame and a reference image frame by executing a global motion estimation model that input the current image frame and the reference image frame (Paragraph 44, The system 300 then performs feature point selection (315) to select certain of the originally identified feature points to use for image alignment. For example, the system 300 may process movement of the feature points over the image sequence to identify feature points that are indicative of global camera movement throughout the image capture sequence and avoid feature points that are indicative of object movement in the scene. In particular, the system 300 may estimate a transformation (e.g., an affine matrix transformation or geometric transformation matrix) of the feature points between images to classify the movement between the images. For example, during the capture sequence, the camera may move towards or away from the scene, rotate, move up or down. Each of these motions can be accounted for using an affine matrix transformation (e.g., linear translation for up/down or left/right movements, scaling for moving towards/away movements, and/or rotational affine transformations for camera rotation.) generating a geometric transformation matrix by combining the global motion parameters (Paragraph 44, In particular, the system 300 may estimate a transformation (e.g., an affine matrix transformation) of the feature points between images to classify the movement between the images. For example, during the capture sequence, the camera may move towards or away from the scene, rotate, move up or down. Each of these motions can be accounted for using an affine matrix transformation (e.g., linear translation for up/down or left/right movements, scaling for moving towards/away movements, and/or rotational affine transformations for camera rotation.); and generating at least one of an output image and an output video using the geometric transformation matrix (Paragraph 49, Further, the system 300 performs a global affine transformation estimation (320) for the remaining set of images using the selected feature points to compute a compute affine alignment matrices, M.sub.1 to M.sub.N, to warp, zoom, rotate, move, and/or shift each of the images Img.sub.i to Img.sub.N to align with Img.sub.0. In particular, the transformation estimation corrects alignment errors to register the images for combination by accounting for global movement of the camera across the sequence of images while ignoring independent movement of objects in the scene. The system 300 provides a single image capture of a scene, using multi-frame capture and blending to enhance the quality of the output provided to the user). Sezer teaches a global estimation model but fails to teach estimating global motion comprising one or more neural networks. However Li teaches motion vectors 312a-312q can be the same as or similar to the motion vectors 310 of FIG. 3A. Each motion vector 312a-312q represents global motion between two image frames of the different set of image frames. The global motion includes both rotation motion and translation motion (Paragraph 77). The image frames 410 are provided to a small motion optical flow network (OfNet) operation 420, which generally operates to estimate small camera motion due to handshake when the image frames 410 are captured. In some embodiments, the small motion OfNet operation 420 includes a neural network. The neural network architecture 420a of FIG. 4B illustrates one specific example architecture that may be used by the small motion OfNet operation 420. Note, however, that any other trained machine learning model may be used here (Paragraph 86). Therefore taking the combined teachings of Sezer and Li, it would be obvious to one skilled in the art before the effective filing date of the invention to have been motivated to have estimated global motion comprising one or more neural networks in order to automatically learn the most relevant features and representations directly from the raw data eliminating the need for any manual feature extraction. [Claim 2] Sezer teaches wherein the generating least one of the output image and the output video further comprises: generating the output video by encoding the current image frame and the reference image frame using the geometric transformation matrix (Paragraph 49, Further, the system 300 performs a global affine transformation estimation (320) for the remaining set of images using the selected feature points to compute a compute affine alignment matrices, M.sub.1 to M.sub.N, to warp, zoom, rotate, move, and/or shift each of the images Img.sub.i to Img.sub.N to align with Img.sub.0. In particular, the transformation estimation corrects alignment errors to register the images for combination by accounting for global movement of the camera across the sequence of images while ignoring independent movement of objects in the scene. The system 300 provides a single image capture of a scene, using multi-frame capture and blending to enhance the quality of the output provided to the user and Paragraph 50, Embodiments of the present disclosure enable seamless fusion of multiple frames to form a single image (or a sequence of frame for video data) with improved resolution and reduced noise levels). [Claim 3] Sezer teaches wherein the generating the at least one of the output image and the output video further comprises: inputting the current image frame, the reference image frame, and the geometric transformation matrix into a video codec (fig. 6, blending 615) configured to execute one or more operations using the geometric transformation matrix (Paragraphs 49 and 54). [Claim 4] Sezer teaches wherein the video codec is configured to execute at least one of a translation mode using a global translation motion (Paragraph 49, Further, the system 300 performs a global affine transformation estimation (320) for the remaining set of images using the selected feature points to compute a compute affine alignment matrices, M.sub.1 to M.sub.N, to warp, zoom, rotate, move, and/or shift each of the images Img.sub.i to Img.sub.N to align with Img.sub.0. In particular, the transformation estimation corrects alignment errors to register the images for combination by accounting for global movement of the camera across the sequence of images while ignoring independent movement of objects in the scene.)). [Claim 5] Sezer teaches wherein the global motion estimation model comprises one or more sub-models corresponding to the affine mode (Paragraph 49, Further, the system 300 performs a global affine transformation estimation (320) for the remaining set of images using the selected feature points to compute a compute affine alignment matrices, M.sub.1 to M.sub.N, to warp, zoom, rotate, move, and/or shift each of the images Img.sub.i to Img.sub.N to align with Img.sub.0. In particular, the transformation estimation corrects alignment errors to register the images for combination by accounting for global movement of the camera across the sequence of images while ignoring independent movement of objects in the scene). [Claim 6] Sezer teaches wherein the geometric transformation matrix is an affine transformation matrix (Paragraph 49). [Claim 8] Sezer teaches wherein the generating of the at least one of the output image and the output video comprises: generating the output image by driving an image signal processor (ISP, e.g. controller 210) using the geometric transformation matrix (Paragraph 51, FIGS. 6A to 6F illustrate block diagrams of image alignment and combination systems according to illustrative embodiments of this disclosure. For example, the image alignment and combination systems may be implemented by the controller 210 in cooperation with the memory 230 of the image processing device 200 in FIG. 2. As illustrated in FIG. 6A, system 600 combines global/local registration 605 with motion artifact processing 610, for example, such as the image registration techniques as discussed above with regard to FIG. 3, to improve the blending 615 of images together to produce a final result 620 having improved image detail). [Claim 9] Li teaches wherein the geometric transformation matrix is a homography transformation matrix (Paragraph 88, The output from the small motion OfNet operation 420 is provided to a homograph matrix operation 430. The homograph matrix operation 430 identifies motion using the optical flow maps provided by the small motion OfNet operation 420. In some cases, the homograph matrix operation 430 can generate a homograph matrix representing the motion. The output from the homograph matrix operation 430 is provided to a motion vector generator operation 440, which generally decomposes the homograph matrix to generate a motion vector representing the global motion of each pixel from one frame to another frame.) in order to have a compact matrix to represent all transformation including rotation, translation, scaling, and perspective distortion. This allows multiple transformations to be chained into one operation, saving processor time. [Claim 13] Sezer teaches an electronic device comprising: a camera (fig. 2 , camera 240) configured to generate a current image frame and a reference image frame (Paragraph 41, several key components are utilized by the system 300 to compute the best alignment between a currently processed image frame and a reference image (Img.sub.0) 325 from the captured sequence); a memory storing one or more instructions (Paragraph 13); a video codec (fig. 6, blending 615); and at least one processor operatively coupled to the memory, the camera, and the video codec, wherein the one or more instructions, when executed by the at least one processor (Paragraph 13), cause the electronic device to: estimate global motion parameters corresponding to components of a global motion between the current image frame and the reference image frame by executing the global motion estimation model comprising one or more neural networks that input the reference image frame (Paragraph 44, The system 300 then performs feature point selection (315) to select certain of the originally identified feature points to use for image alignment. For example, the system 300 may process movement of the feature points over the image sequence to identify feature points that are indicative of global camera movement throughout the image capture sequence and avoid feature points that are indicative of object movement in the scene. In particular, the system 300 may estimate a transformation (e.g., an affine matrix transformation or geometric transformation matrix) of the feature points between images to classify the movement between the images. For example, during the capture sequence, the camera may move towards or away from the scene, rotate, move up or down. Each of these motions can be accounted for using an affine matrix transformation (e.g., linear translation for up/down or left/right movements, scaling for moving towards/away movements, and/or rotational affine transformations for camera rotation.); generate a geometric transformation matrix by combining the global motion parameters (Paragraph 44, In particular, the system 300 may estimate a transformation (e.g., an affine matrix transformation) of the feature points between images to classify the movement between the images. For example, during the capture sequence, the camera may move towards or away from the scene, rotate, move up or down. Each of these motions can be accounted for using an affine matrix transformation (e.g., linear translation for up/down or left/right movements, scaling for moving towards/away movements, and/or rotational affine transformations for camera rotation.); and control the video codec to generate an output video using the geometric transformation matrix (Paragraphs 49 and 54). Sezer teaches a global estimation model but fails to teach storing a global motion estimation model based on a neural network and estimating global motion comprising one or more neural networks. However Li teaches The memory 130 can include a volatile and/or non-volatile memory. For example, the memory 130 can store commands or data related to at least one other component of the electronic device 101. In accordance with this disclosure, the memory 130 can store software and/or a program 140 (Note, ofnet is a software program used for estimating motion, Paragraph 47). The program 140 includes, for example, a kernel 141, middleware 143, an application programming interface (API) 145, and/or an application program (or “application”) 147. At least a portion of the kernel 141, middleware 143, or API 145 may be denoted an operating system (OS). motion vectors 312a-312q can be the same as or similar to the motion vectors 310 of FIG. 3A. Each motion vector 312a-312q represents global motion between two image frames of the different set of image frames. The global motion includes both rotation motion and translation motion (Paragraph 77). The image frames 410 are provided to a small motion optical flow network (OfNet) operation 420, which generally operates to estimate small camera motion due to handshake when the image frames 410 are captured. In some embodiments, the small motion OfNet operation 420 includes a neural network. The neural network architecture 420a of FIG. 4B illustrates one specific example architecture that may be used by the small motion OfNet operation 420. Note, however, that any other trained machine learning model may be used here (Paragraph 86). Therefore taking the combined teachings of Sezer and Li, it would be obvious to one skilled in the art before the effective filing date of the invention to have been motivated to have stored a global motion estimation model based on a neural network and estimating global motion comprising one or more neural networks in order to automatically learn the most relevant features and representations eliminating the need for any manual feature extraction. [Claim 14] Sezer teaches wherein the video codec is configured to execute at least one of a translation mode using a global translation motion (Paragraph 49, Further, the system 300 performs a global affine transformation estimation (320) for the remaining set of images using the selected feature points to compute a compute affine alignment matrices, M.sub.1 to M.sub.N, to warp, zoom, rotate, move, and/or shift each of the images Img.sub.i to Img.sub.N to align with Img.sub.0. In particular, the transformation estimation corrects alignment errors to register the images for combination by accounting for global movement of the camera across the sequence of images while ignoring independent movement of objects in the scene.)). [Claim 15] Sezer teaches wherein the geometric transformation matrix is an affine transformation matrix (Paragraph 49). [Claim 18] Sezer teaches an electronic device comprising: a camera (fig. 2, camera 240) configured to generate a current image frame and a reference image frame (Paragraph 41, several key components are utilized by the system 300 to compute the best alignment between a currently processed image frame and a reference image (Img.sub.0) 325 from the captured sequence); a memory storing one or more instructions (Paragraph 33); an image signal processor (615); at least one processor (210) operatively coupled to the memory, the camera, and the ISP; wherein the one or more instructions, when executed by the at least one processor (Paragraph 33), cause the electronic device to: estimate global motion parameters corresponding to components of a global motion between the current image frame and the reference image frame by executing the global motion estimation model based on the current image frame and the reference image frame (Paragraph 44, The system 300 then performs feature point selection (315) to select certain of the originally identified feature points to use for image alignment. For example, the system 300 may process movement of the feature points over the image sequence to identify feature points that are indicative of global camera movement throughout the image capture sequence and avoid feature points that are indicative of object movement in the scene. In particular, the system 300 may estimate a transformation (e.g., an affine matrix transformation or geometric transformation matrix) of the feature points between images to classify the movement between the images. For example, during the capture sequence, the camera may move towards or away from the scene, rotate, move up or down. Each of these motions can be accounted for using an affine matrix transformation (e.g., linear translation for up/down or left/right movements, scaling for moving towards/away movements, and/or rotational affine transformations for camera rotation); generate a geometric transformation matrix by combining the global motion parameters (Paragraph 44, In particular, the system 300 may estimate a transformation (e.g., an affine matrix transformation) of the feature points between images to classify the movement between the images. For example, during the capture sequence, the camera may move towards or away from the scene, rotate, move up or down. Each of these motions can be accounted for using an affine matrix transformation (e.g., linear translation for up/down or left/right movements, scaling for moving towards/away movements, and/or rotational affine transformations for camera rotation.); and control the ISP to generate an output image using the geometric transformation matrix Paragraphs 49 and 54). Sezer teaches a global estimation model but fails to teach storing a global motion estimation model based on a neural network and estimating global motion comprising one or more neural networks. However Li teaches a memory 130 can include a volatile and/or non-volatile memory. For example, the memory 130 can store commands or data related to at least one other component of the electronic device 101. In accordance with this disclosure, the memory 130 can store software and/or a program 140 (Note, ofnet is a software program used for estimating motion, Paragraph 47). The program 140 includes, for example, a kernel 141, middleware 143, an application programming interface (API) 145, and/or an application program (or “application”) 147. At least a portion of the kernel 141, middleware 143, or API 145 may be denoted an operating system (OS). motion vectors 312a-312q can be the same as or similar to the motion vectors 310 of FIG. 3A. Each motion vector 312a-312q represents global motion between two image frames of the different set of image frames. The global motion includes both rotation motion and translation motion (Paragraph 77). The image frames 410 are provided to a small motion optical flow network (OfNet) operation 420, which generally operates to estimate small camera motion due to handshake when the image frames 410 are captured. In some embodiments, the small motion OfNet operation 420 includes a neural network. The neural network architecture 420a of FIG. 4B illustrates one specific example architecture that may be used by the small motion OfNet operation 420. Note, however, that any other trained machine learning model may be used here (Paragraph 86). Therefore taking the combined teachings of Sezer and Li, it would be obvious to one skilled in the art before the effective filing date of the invention to have been motivated to have stored a global motion estimation model based on a neural network and estimating global motion comprising one or more neural networks in order to automatically learn the most relevant features and representations eliminating the need for any manual feature extraction. [Claim 20] Li teaches wherein the geometric transformation matrix is a homography transformation matrix (Paragraph 88, The output from the small motion OfNet operation 420 is provided to a homograph matrix operation 430. The homograph matrix operation 430 identifies motion using the optical flow maps provided by the small motion OfNet operation 420. In some cases, the homograph matrix operation 430 can generate a homograph matrix representing the motion. The output from the homograph matrix operation 430 is provided to a motion vector generator operation 440, which generally decomposes the homograph matrix to generate a motion vector representing the global motion of each pixel from one frame to another frame.) in order to have a compact matrix to represent all transformation including rotation, translation, scaling, and perspective distortion. This allows multiple transformations to be chained into one operation, saving processor time. Allowable Subject Matter Claims 7, 10-12, 16, 17 and 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The prior art fails to teach or suggest as in claims 7, 16 and 19, “determining one or more function values by substituting one or more global motion parameters into one or more functions; and determining one or more elements of the geometric transformation matrix by combining the global motion parameters based on (i) operations between the global motion parameters, (ii) operations between the global motion parameters and the one or more function values, (iii) operations between a plurality of function values of the one or more function values, or (iv) a combination thereof” and claims 10 and 17, “generating a scaled current image frame by scaling the current image frame to a target size; and generating a scaled reference image frame by scaling the reference image frame to the target size, wherein the global motion estimation model is executed based on the scaled current image frame and the scaled reference image frame”, and claim 12, “a first estimation model configured to estimate first global motion parameters and a second estimation model configured to estimate second global motion parameters, and the geometric transformation matrix comprises an affine transformation matrix determined by a combination of the first global motion parameters and a homography transformation matrix determined by a combination of the second global motion parameters”. Claim 11 is dependent upon claim 10. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to YOGESH K AGGARWAL whose telephone number is (571)272-7360. The examiner can normally be reached Monday - Friday 9:30-6. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sinh Tran can be reached at 5712727564. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YOGESH K AGGARWAL/Primary Examiner, Art Unit 2637
Read full office action

Prosecution Timeline

Apr 01, 2025
Application Filed
Jul 15, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12707135
VIDEO EQUIPMENT PLATFORMS AND RELATED METHODS
2y 4m to grant Granted Aug 11, 2026
Patent 12707143
ROBUST OPTICAL MONITORING DEVICE AND OPERATING METHOD THEREOF
2y 4m to grant Granted Aug 11, 2026
Patent 12707151
ELECTRONIC DEVICE FOR MEASURING EFFECTIVE DYNAMIC RANGE LENGTH AND METHOD OF OPERATING THE ELECTRONIC DEVICE
1y 11m to grant Granted Aug 11, 2026
Patent 12701642
LIGHT DRIVER CALIBRATION
2y 10m to grant Granted Aug 04, 2026
Patent 12701325
METHOD AND APPARATUS FOR CONTROLLING AERIAL VEHICLE TO SHOOT BASED ON PORTRAIT MODEL, DEVICE, AND MEDIUM
2y 4m to grant Granted Aug 04, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
90%
Grant Probability
96%
With Interview (+6.6%)
2y 5m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1135 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month