Prosecution Insights
Last updated: October 02, 2026
Application No. 18/341,979

REPLACEMENT OF SOURCE MOVING OBJECTS WITH TARGET MOVING OBJECTS IN A VIDEO USING GAN

Non-Final OA §102§103
Filed
Jun 27, 2023
Examiner
ELLIOTT, JORDAN MCKENZIE
Art Unit
Tech Center
Assignee
International Business Machines Corporation
OA Round
1 (Non-Final)
47%
Grant Probability
Moderate
1-2
OA Rounds
0m
Est. Remaining
51%
With Interview

Examiner Intelligence

Grants 47% of resolved cases
47%
Career Allowance Rate
15 granted / 32 resolved
-13.1% vs TC avg
Minimal +4% lift
Without
With
+4.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
25 currently pending
Career history
69
Total Applications
across all art units

Statute-Specific Performance

§101
7.9%
-32.1% vs TC avg
§103
55.8%
+15.8% vs TC avg
§102
25.2%
-14.8% vs TC avg
§112
11.2%
-28.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 32 resolved cases

Office Action

§102 §103
DETAILED ACTION Claims 1-20 are pending in this application. Information Disclosure Statement The information disclosure statement (IDS) submitted on 06/27/23 in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are: generator component in claims 1, 2, 3, and 4. classifier component in claims 1, 5 and 8. shadow and reflection extraction component of claim 6 and 7. a consistency generator component of claim 8. consistency classifier component of claim 8. cycle consistency component of claim 8. system of claims 10, 11, 12, 13, 14 and 15. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-5, 8-13 and 16 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Mann (US 20230132243 A1). Regarding claim 1 Mann discloses; A system comprising: a memory that stores computer executable components (Mann, [0031]-[0032] the system is executed on a computer); and a processor that executes computer executable components stored in the memory, wherein the computer executable components comprise (Mann, [0031]-[0032] the system has machine readable instructions which are stored on a memory and executed by a processor or processors): a generative adversarial network comprising (Mann, [0070] the machine learning models used may comprise a generative adversarial network with a generator and a discriminator): a generator component (Mann, [0070] the machine learning models used may comprise a generative adversarial network with a generator (generator component) and a discriminator) that replaces a moving source object of a first video with a moving target object of a second video to produce an output video (Mann, [0010] a first sequence of image frames (first video) is determined, where there is detected first event or object (moving source object) over the first sequence of frames [0022] a second sequence of video frames (second video) containing a second object (moving target object) is detected, where the object may be the same object or a different object from the first object in the first sequence of frames, [0028] a third sequence of frames is generated (output video) where at least part of the first object in the first set of frames is transitioned/replaced by the object in the second set of frames, where [0100] optical flow may be used to render and adjust the object based on movement, indicating the objects are moving); PNG media_image1.png 244 272 media_image1.png Greyscale (Mann, [0010]) PNG media_image2.png 250 274 media_image2.png Greyscale (Mann, [0022]) PNG media_image3.png 84 282 media_image3.png Greyscale PNG media_image4.png 124 282 media_image4.png Greyscale (Mann, [0028]) and a classifier component that classifies the output video as real or fake (Mann, [0070] the system has a discriminator network (classifier component) to determine whether the a given image is a genuine instance of the object (real) or was generated by the model (fake), [0078] the discriminator (classifier component) takes the reconstructed video instance (output video) as input as shown in figure 4). PNG media_image5.png 298 278 media_image5.png Greyscale (Mann, [0070]) PNG media_image6.png 300 274 media_image6.png Greyscale (Mann, [0078]) PNG media_image7.png 422 646 media_image7.png Greyscale (Mann, figure 4, emphasis added) Regarding claim 2 Mann discloses; The system of claim 1, wherein generator component performs instance segmentation on the first video and the second video produce a segmented moving source object and a segmented moving target object (Mann, [0012] the machine learning model may generate a segmentation mask of the object in each instance video frame to separate it from the background, [0072] a “synthetic image” is generated from each isolated instance of an object or objects and then each isolated instance is masked ([0022] and [0052] notes that there may be one or more objects) [0073] each isolated instance of the objects may be segmented to generate a binary segmentation mask of each instance as shown in figure 4). PNG media_image8.png 300 286 media_image8.png Greyscale (Mann, [0012], emphasis added) PNG media_image9.png 310 274 media_image9.png Greyscale (Mann, [0072]) PNG media_image10.png 248 274 media_image10.png Greyscale PNG media_image11.png 102 280 media_image11.png Greyscale (Mann, [0073]) PNG media_image12.png 422 646 media_image12.png Greyscale (Mann, figure 4, emphasis added) Regarding claim 3 Mann discloses; The system of claim 1, wherein the generator component removes the moving source object from the first video and the moving target source from the second video via a binary mask to generate a masked first video and a removed moving target object (Mann, [0006] the instances of the objects in the sequences are isolated in order to generate a synthetic or replacement portion of the object instance, and then the first object (moving source object) is replaced with a synthetic instance of the object from another set of frames (moving target object), which indicates that the first object (moving source object) is removed from the video and the synthetic object (moving target object) is removed from the second video, where [0008] notes an example of removing and replacing a region of an actors face in a first video with a modified or second replacement facial region, [0072] the video frames/images (first video) may have a binary segmentation mask in which the masked area is transparent designating a removed/transparent object region, [0073] the region of the synthetic image (removed moving target region) is prepared via masking such that it may be overlaid onto to the masked region/removed region of the first object instance sequence of frames, [0074] and [0080] Figure 5 B shows the attention mask corresponding to the removed object from figure 5A, where the region on the first video instance frame (504) has the first object/region removed (removed moving source object) and a masked region is generated (masked first video) for the second instance video object to replace the masked region). PNG media_image13.png 764 486 media_image13.png Greyscale (Mann, Figure 5A-5C) Regarding claim 4 Mann discloses; The system of claim 3, wherein the generator component pastes the removed moving target object in the masked first video (Mann, [0073] the isolated instance frame (first video) has the corresponding synthetic image (second video region which has the target object in it) overlaid on top of it, resulting in a composite frame image (output video) which has the region of the first video frame replaced with the target object, the overlaying/pasting process is done by masking regions in both frames to isolate the target and source objects, [0074] figure 5A shows this final output where the first video frame data (504) has the target object content (502) overlaid on top of it, resulting frame is a first video frame with the target object in it). PNG media_image14.png 250 270 media_image14.png Greyscale PNG media_image15.png 200 276 media_image15.png Greyscale (Mann, [0073]-[0074]) PNG media_image16.png 275 485 media_image16.png Greyscale (Mann, figure 5 A) Regarding claim 5 Mann discloses; The system of claim 1, wherein the classifier component comprises a moving source object classifier and a moving target object classifier (Mann, [0013], and [0078]-[0079] the system may use an adversarial loss function in combination with another loss function such as a photometric loss function (moving source object classified and target object classifier) which detect and compare differences between the first object instance (source object) and the synthetic or second object instance (target object)). Regarding claim 8 Mann discloses; The system of claim 1, further comprising: a consistency generator component that replaces the moving target object of the output video with a removed moving source object of the first video to produce a consistency output video (Mann, [0027] the system trains the model to seamlessly transition between frames in which an object has been replaced/modified with the original frame content pasted back in [0080] the system may use a masking operation to apply an attention mask generated from the isolated instance (removed moving source data) to the composite frame (output video) in order to compute a loss, this is input into the discriminator network); a consistency classifier component that classifies the consistency output video as real or fake (Mann, [0070] the system has a discriminator network (classifier component) to determine whether the input reconstructed frame (constancy output video) is a genuine instance of the object (real) or was generated by the model (fake)); and a cycle consistency component that compares the first video to the consistency output video to generate a cycle consistency loss (Mann,[0078] the reconstructed frame (consistency output) is compared with the isolated instance frame (first video) to determine the adversarial loss). Regarding claim 9 Mann discloses; The system of claim 8, wherein the classifier component classifies the output video as real or fake based at least in part on the cycle consistency loss (Mann, [0070] the loss function is used to train the model to determine whether or not the output is identical or close to identical to the ground truth frame). Regarding claim 10 Mann discloses; A computer-implemented method, comprising: replacing, by a system coupled to a processor (Mann, [0031]-[0032] the system has machine readable instructions which are stored on a memory and executed by a processor or processors), a moving source object of a first video with a moving target object of a second video to produce an output video using a generative adversarial network (Mann, [0010] a first sequence of image frames (first video) is determined, where there is detected first event or object (moving source object) over the first sequence of frames [0022] a second sequence of video frames (second video) containing a second object (moving target object) is detected, where the object may be the same object or a different object from the first object in the first sequence of frames, [0028] a third sequence of frames is generated (output video) where at least part of the first object in the first set of frames is transitioned/replaced by the object in the second set of frames, where [0100] optical flow may be used to render and adjust the object based on movement, indicating the objects are moving); and classifying, by the system, the output video as real or fake (Mann, [0070] the system has a discriminator network (classifier component) to determine whether the a given image is a genuine instance of the object (real) or was generated by the model (fake), [0078] the discriminator (classifier component) takes the reconstructed video instance (output video) as input as shown in figure 4). Regarding claim 11 Mann discloses; The computer-implemented method of claim 10, further comprising: performing, by the system, instance segmentation on the first video and the second video produce a segmented moving source object and a segmented moving target object (Mann, [0012] the machine learning model may generate a segmentation mask of the object in each instance video frame to separate it from the background, [0072] a “synthetic image” is generated from each isolated instance of an object or objects and then each isolated instance is masked ([0022] and [0052] notes that there may be one or more objects) [0073] each isolated instance of the objects may be segmented to generate a binary segmentation mask of each instance as shown in figure 4). Regarding claim 12 Mann discloses; The computer-implemented method of claim 10, further comprising: removing, by the system, the moving source object from the first video and the moving target source from the second video via a binary mask to generate a masked first video and a removed moving target object (Mann, [0006] the instances of the objects in the sequences are isolated in order to generate a synthetic or replacement portion of the object instance, and then the first object (moving source object) is replaced with a synthetic instance of the object from another set of frames (moving target object), which indicates that the first object (moving source object) is removed from the video and the synthetic object (moving target object) is removed from the second video, where [0008] notes an example of removing and replacing a region of an actors face in a first video with a modified or second replacement facial region, [0072] the video frames/images (first video) may have a binary segmentation mask in which the masked area is transparent designating a removed/transparent object region, [0073] the region of the synthetic image (removed moving target region) is prepared via masking such that it may be overlaid onto to the masked region/removed region of the first object instance sequence of frames, [0074] and [0080] Figure 5 B shows the attention mask corresponding to the removed object from figure 5A, where the region on the first video instance frame (504) has the first object/region removed (removed moving source object) and a masked region is generated (masked first video) for the second instance video object to replace the masked region). Regarding claim 13 Mann discloses; The computer-implemented method of claim 12, further comprising: pasting, by the system, the removed moving target object in the masked first video (Mann, [0073] the isolated instance frame (first video) has the corresponding synthetic image (second video region which has the target object in it) overlaid on top of it, resulting in a composite frame image (output video) which has the region of the first video frame replaced with the target object, the overlaying/pasting process is done by masking regions in both frames to isolate the target and source objects, [0074] figure 5A shows this final output where the first video frame data (504) has the target object content (502) overlaid on top of it, resulting frame is a first video frame with the target object in it). Regarding claim 16 Mann discloses; A computer program product facilitating replacement of a source moving object with a target moving object in a video using a generative adversarial network (Mann, [0070] the machine learning models used may comprise a generative adversarial network with a generator and a discriminator), the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to (Mann, [0031]-[0032] the system has machine readable instructions which are stored on a memory and executed by a processor or processors): replace a moving source object of a first video with a moving target object of a second video to produce an output video (Mann, [0010] a first sequence of image frames (first video) is determined, where there is detected first event or object (moving source object) over the first sequence of frames [0022] a second sequence of video frames (second video) containing a second object (moving target object) is detected, where the object may be the same object or a different object from the first object in the first sequence of frames, [0028] a third sequence of frames is generated (output video) where at least part of the first object in the first set of frames is transitioned/replaced by the object in the second set of frames, where [0100] optical flow may be used to render and adjust the object based on movement, indicating the objects are moving); and classify the output video as real or fake (Mann, [0070] the system has a discriminator network (classifier component) to determine whether the a given image is a genuine instance of the object (real) or was generated by the model (fake), [0078] the discriminator (classifier component) takes the reconstructed video instance (output video) as input as shown in figure 4). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. 2. Claims 6-7, 14-15 and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over Mann (US 20230132243 A1) in view of Joachim (US 20240378832 A1). Regarding claim 6 Mann fails to disclose; The system of claim 1, further comprising: a shadow and reflection extraction component that determines whether at least one of a reflection and a shadow of the moving source object exist in the first video, and in response to a determination that at least one of the reflection and the shadow of the moving object exist in the first video, removes the at least one of the reflection and the shadow from the first video. However, in the same field of endeavor of video content generation and processing, teaches; a shadow and reflection extraction component that determines whether at least one of a reflection and a shadow of the moving source object exist in the first video (Joachim, [0114] the system detects a shadow of a distracting object in the video frame), and in response to a determination that at least one of the reflection and the shadow of the moving object exist in the first video, removes the at least one of the reflection and the shadow from the first video (Joachim, [0114] shadows are detected and then the shadows are removed from the video upon detection). The combination of Mann and Joachim would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for combination lies in that when removing or editing objects in a video, a shadow caused by that object may remain in the video which can be distracting to the viewer, the removal of a shadow corresponding to a removed object allows the quality of the generated or edited image to be improved. (Joachim, [0456]-[0464]) Regarding claim 7 the combination of Mann and Joachim teaches; The system of claim 6, further comprising: a shadow and reflection rendering component that, in response to a determination that at least one of a reflection and a shadow of the moving object exist in the first video, renders corresponding reflections and shadows of the moving target object in the first video (Joachim, [0539]-[0541] the system generates a shadow for an object being edited into the frame or moved into the frame in response to lighting and shadow feature detections in the first video). The combination of Mann and Joachim would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for combination lies in that when removing or editing objects in a video, the shadow and lighting of the objects must be consistent with the placement of the objects in order to output an image that looks correct visually. (Joachim, [0456]-[0464] and [0537]-[0543]) Regarding claim 14 the combination of Mann and Joachim teaches; The computer-implemented method of claim 10, further comprising: determining, by the system, whether at least one of a reflection and a shadow of the moving source object exist in the first video (Joachim, [0114] the system detects a shadow of a distracting object in the video frame), and in response to a determination that at least one of the reflection and the shadow of the moving object exist in the first video, removes the at least one of the reflection and the shadow from the first video (Joachim, [0114] shadows are detected and then the shadows are removed from the video upon detection). The combination of Mann and Joachim would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for combination lies in that when removing or editing objects in a video, a shadow caused by that object may remain in the video which can be distracting to the viewer, the removal of a shadow corresponding to a removed object allows the quality of the generated or edited image to be improved. (Joachim, [0456]-[0464]) Regarding claim 15 the combination of Mann and Joachim teaches; The computer-implemented method of claim 14, further comprising: in response to a determination that at least one of a reflection and a shadow of the moving object exist in the first video, rendering, by the system, corresponding reflections and shadows of the moving target object in the first video(Joachim, [0539]-[0541] the system generates a shadow for an object being edited into the frame or moved into the frame in response to lighting and shadow feature detections in the first video). The combination of Mann and Joachim would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for combination lies in that when removing or editing objects in a video, the shadow and lighting of the objects must be consistent with the placement of the objects in order to output an image that looks correct visually. (Joachim, [0456]-[0464] and [0537]-[0543]) Regarding claim 17 the combination of Mann and Joachim teaches; The computer program product of claim 16, wherein the program instructions are further executable by the processor to cause the processor to (Mann, [0031]-[0032] the system has machine readable instructions which are stored on a memory and executed by a processor or processors): determine whether at least one of a reflection and a shadow of the moving source object exist in the first video (Joachim, [0114] the system detects a shadow of a distracting object in the video frame), and in response to a determination that at least one of the reflection and the shadow of the moving object exist in the first video, removes the at least one of the reflection and the shadow from the first video (Joachim, [0114] shadows are detected and then the shadows are removed from the video upon detection). The combination of Mann and Joachim would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for combination lies in that when removing or editing objects in a video, a shadow caused by that object may remain in the video which can be distracting to the viewer, the removal of a shadow corresponding to a removed object allows the quality of the generated or edited image to be improved. (Joachim, [0456]-[0464]) Regarding claim 18 the combination of Mann and Joachim teaches; The computer program product of claim 17, wherein the program instructions are further executable by the processor to cause the processor to (Mann, [0031]-[0032] the system has machine readable instructions which are stored on a memory and executed by a processor or processors): in response to a determination that at least one of a reflection and a shadow of the moving object exist in the first video, render corresponding reflections and shadows of the moving target object in the first video(Joachim, [0539]-[0541] the system generates a shadow for an object being edited into the frame or moved into the frame in response to lighting and shadow feature detections in the first video). The combination of Mann and Joachim would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for combination lies in that when removing or editing objects in a video, the shadow and lighting of the objects must be consistent with the placement of the objects in order to output an image that looks correct visually. (Joachim, [0456]-[0464] and [0537]-[0543]) Regarding claim 19 the combination of Mann and Joachim teaches; The computer program product of claim 17, wherein the program instructions are further executable by the processor to cause the processor to (Mann, [0031]-[0032] the system has machine readable instructions which are stored on a memory and executed by a processor or processors): replace the moving target object of the output video with a removed moving source object of the first video to produce a consistency output video (Mann, [0027] the system trains the model to seamlessly transition between frames in which an object has been replaced/modified with the original frame content pasted back in [0080] the system may use a masking operation to apply an attention mask generated from the isolated instance (removed moving source data) to the composite frame (output video) in order to compute a loss, this is input into the discriminator network); classify the consistency output video as real or fake (Mann, [0070] the system has a discriminator network (classifier component) to determine whether the input reconstructed frame (constancy output video) is a genuine instance of the object (real) or was generated by the model (fake)); and compare the first video to the consistency output video to generate a cycle consistency loss (Mann,[0078] the reconstructed frame (consistency output) is compared with the isolated instance frame (first video) to determine the adversarial loss). Regarding claim 20 the combination of Mann and Joachim teaches; The computer program product of claim 19, wherein the program instructions are further executable by the processor to cause the processor to (Mann, [0031]-[0032] the system has machine readable instructions which are stored on a memory and executed by a processor or processors): classify the output video as real or fake based at least in part on the cycle consistency loss (Mann, [0070] the loss function is used to train the model to determine whether or not the output is identical or close to identical to the ground truth frame). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. For a full listing of analogous prior art as determined by the examiner, please see the attached PTO-892 Notice of References Cited form. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JORDAN M ELLIOTT whose telephone number is (703)756-5463. The examiner can normally be reached M-F 8AM-5PM ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emily Terrell can be reached at (571) 270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /J.M.E./Examiner, Art Unit 2666 /EMILY C TERRELL/Supervisory Patent Examiner, Art Unit 2666
Read full office action

Prosecution Timeline

Jun 27, 2023
Application Filed
Nov 24, 2023
Response after Non-Final Action
Aug 18, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705715
Tone Mapping for Preserving Contrast of Fine Features in an Image
3y 6m to grant Granted Aug 11, 2026
Patent 12682611
IMAGE ACQUISITION MODEL TRAINING METHOD AND APPARATUS, IMAGE DETECTION METHOD AND APPARATUS, AND DEVICE
3y 1m to grant Granted Jul 14, 2026
Patent 12573117
METHOD AND DEVICE FOR DEEP LEARNING-BASED PATCHWISE RECONSTRUCTION FROM CLINICAL CT SCAN DATA
3y 1m to grant Granted Mar 10, 2026
Patent 12475998
SYSTEMS AND METHODS OF ADAPTIVELY GENERATING FACIAL DEVICE SELECTIONS BASED ON VISUALLY DETERMINED ANATOMICAL DIMENSION DATA
2y 2m to grant Granted Nov 18, 2025
Patent 12450918
AUTOMATIC LANE MARKING EXTRACTION AND CLASSIFICATION FROM LIDAR SCANS
3y 1m to grant Granted Oct 21, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
47%
Grant Probability
51%
With Interview (+4.2%)
2y 11m (~0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 32 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month