DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
The amendments, filed 7/2/2026, have been entered and made of record. Claims 1, 11, and 20 have been amended. Claims 1-20 are pending.
Response to Arguments
Applicant's arguments filed 7/2/2026 have been fully considered but they are not persuasive.
In re page 11, the applicant states “the cited references fail to teach or suggest at least "selecting, by the computing system, the second layer to be off to omit the background of the video data" or "providing, by the computing system, the first layer for display without the background of the video data via a graphical user interface" as recited by amended claim 11. As such, Applicant respectfully submits that independent claim 11, at least as amended, patentably defines over all the cited references, read alone or in any combination”.
In response, the examiner respectfully disagrees. Park teaches a computer implemented method comprising: obtaining, by a computing system, video data comprising a plurality of image frames(“A user terminal 800 is a terminal used by a user or a security official who manages the video search system 1, and may be a personal computer (PC) or a mobile terminal. The user may control the video search system 1 through the user terminal 800. The user terminal 800 includes the input device 600 that is a user interface capable of inputting a query (search condition) to the video search system 1.” in Para.[0081], “The video analysis engine 101 analyzes the original video, classifies the original video according to a predefined condition such as a predefined category, and extracts attributes of an object detected from the original video, for example, a type, a color, a size, a form, a motion, and a trajectory of the object” in Para.[0044], );
automatically, by the computing system, recognizing one or more objects within a first image frame of the plurality of image frames(“The video analysis engine 101 analyzes the original video, classifies the original video according to a predefined condition such as a predefined category, and extracts attributes of an object detected from the original video, for example, a type, a color, a size, a form, a motion, and a trajectory of the object” in Para.[0044]);
generating, by the computing system, a first layer for a first object of the one or more objects;(“In summarized video I (b) and summarized video II (c), video data in which the object A and the object B appear is extracted to reduce the play time of the original video into a time T2 and a time T3, respectively. The summarized video I (b) and the summarized video II (c) have different overlapping degrees between the object A and the object B. Under a condition that the object B having a temporally later appearing order than the object A does not appear before the object A, the degree of overlapping between the object A and the object B in the summarized video may be adjusted by adjusting the complexity of the summarized video” in Para.[0061], Fig. 3),
generating, by the computing system, a second layer for a background of the video data(“The video analysis engine 101 performs background region detection, foreground and object detection, object counting, camera tampering detection,” in Para.[0046], “The browsing engine 505 renders a background model” in Para.[0053]);
selecting, by the computing system, the first layer for display; selecting, by the computing system, the second layer to be off to omit the background of the video data; and providing, by the computing system, the first layer for display without the background of the video data via a graphical user interface (“In summarized video I (b) and summarized video II (c), video data in which the object A and the object B appear is extracted to reduce the play time of the original video into a time T2 and a time T3, respectively. The summarized video I (b) and the summarized video II (c) have different overlapping degrees between the object A and the object B. Under a condition that the object B having a temporally later appearing order than the object A does not appear before the object A, the degree of overlapping between the object A and the object B in the summarized video may be adjusted by adjusting the complexity of the summarized video” in Para.[0061], “The summarized video generation unit 510 renders the summarized video by limiting the number of appearing objects. It may be confusing if too many objects initially appear at the same time. Thus, the summarized video generation unit 510 may configure the summarized video by limiting the number of objects to” in Para.[0062], Fig. 3, Any object layer can be selected to be displayed so the background also is able to be selected to be displayed because it can be considered one of layers based on video analysis engine. The background model is omitted in Fig. 3.)
Therefore, Park discloses selecting, by the computing system, the second layer to be off to omit the background of the video data … providing, by the computing system, the first layer for display without the background of the video data via a graphical user interface as recited by amended claim 11.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Park
Claim 11 is rejected under 35 U.S.C. 102(a)(1)/(a)(2) as being anticipated by Park et al.(USPubN 2015/0127626; hereinafter Park).
As per claim 11, Park teaches a computer implemented method comprising: obtaining, by a computing system, video data comprising a plurality of image frames(“A user terminal 800 is a terminal used by a user or a security official who manages the video search system 1, and may be a personal computer (PC) or a mobile terminal. The user may control the video search system 1 through the user terminal 800. The user terminal 800 includes the input device 600 that is a user interface capable of inputting a query (search condition) to the video search system 1.” in Para.[0081], “The video analysis engine 101 analyzes the original video, classifies the original video according to a predefined condition such as a predefined category, and extracts attributes of an object detected from the original video, for example, a type, a color, a size, a form, a motion, and a trajectory of the object” in Para.[0044], );
automatically, by the computing system, recognizing one or more objects within a first image frame of the plurality of image frames(“The video analysis engine 101 analyzes the original video, classifies the original video according to a predefined condition such as a predefined category, and extracts attributes of an object detected from the original video, for example, a type, a color, a size, a form, a motion, and a trajectory of the object” in Para.[0044]);
generating, by the computing system, a first layer for a first object of the one or more objects;(“In summarized video I (b) and summarized video II (c), video data in which the object A and the object B appear is extracted to reduce the play time of the original video into a time T2 and a time T3, respectively. The summarized video I (b) and the summarized video II (c) have different overlapping degrees between the object A and the object B. Under a condition that the object B having a temporally later appearing order than the object A does not appear before the object A, the degree of overlapping between the object A and the object B in the summarized video may be adjusted by adjusting the complexity of the summarized video” in Para.[0061], Fig. 3),
generating, by the computing system, a second layer for a background of the video data(“The video analysis engine 101 performs background region detection, foreground and object detection, object counting, camera tampering detection,” in Para.[0046], “The browsing engine 505 renders a background model” in Para.[0053]);
selecting, by the computing system, the first layer for display; selecting, by the computing system, the second layer to be off to omit the background of the video data; and providing, by the computing system, the first layer for display without the background of the video data via a graphical user interface (“In summarized video I (b) and summarized video II (c), video data in which the object A and the object B appear is extracted to reduce the play time of the original video into a time T2 and a time T3, respectively. The summarized video I (b) and the summarized video II (c) have different overlapping degrees between the object A and the object B. Under a condition that the object B having a temporally later appearing order than the object A does not appear before the object A, the degree of overlapping between the object A and the object B in the summarized video may be adjusted by adjusting the complexity of the summarized video” in Para.[0061], “The summarized video generation unit 510 renders the summarized video by limiting the number of appearing objects. It may be confusing if too many objects initially appear at the same time. Thus, the summarized video generation unit 510 may configure the summarized video by limiting the number of objects to” in Para.[0062], Fig. 3, Any object layer can be selected to be displayed so the background also is able to be selected to be displayed because it can be considered one of layers based on video analysis engine. The background model is omitted in Fig. 3.)
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Park in view of Tsai
Claims 12-16, 18, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Park et al.(USPubN 2015/0127626; hereinafter Park) in view of Tsai et al.(USPubN 2021/0004962; hereinafter Tsai).
As per claim 12, Park teaches all of limitation of claim 11.
Park is silent about the operations comprising: transferring, by the computing system, high resolution details of the video data in a post processing step subsequent to receiving the background layer and the one or more object layers.
Tsai teaches the operations comprising: transferring, by the computing system, high resolution details of the video data in a post processing step subsequent to receiving the background layer and the one or more object layers (“when performing another iteration of steps 706 through 710, the binarized version of the saliency map generated at step 710 can help improve the quality or accuracy of the foreground queries calculated at step 706. This in turn can also help improve the quality or accuracy of the results or calculations at steps 708 and 710. Thus, in some cases, the additional iteration(s) of steps 706 through 710 can produce a saliency map of progressively higher quality or accuracy, which can be used to generate an output image of higher quality or accuracy (e.g., better field-of-view effect, etc.)” in Para.[0171], Higher quality can be interpreted as high resolution details.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings Park with the above teachings of Tsai in order to improve user experience.
As per claim 13, Park teaches all of limitation of claim 11.
Park is silent about comprising: generating one or more object maps, wherein each of the one or more object maps is descriptive of a respective location of at least one object of the one or more objects within the image frame.
Tsai teaches comprising: generating one or more object maps, wherein each of the one or more object maps is descriptive of a respective location of at least one object of the one or more objects within the image frame (“the image processing system 100 can generate a saliency map (S.sub.crf). The saliency map (S.sub.crf) can be generated based on the saliency refinement of the ranking map S at block 220. In some examples, the image processing system 100 can generate the saliency map (S.sub.crf) using the respective probability of each pixel being salient, which can be determined based on the refined ranking map S” in Para.[0088], “the location of the region of interest (e.g., the target or object that the user wants to separate from the background) in the image can be represented or illustrated using a probability map as shown below in item 312 of FIG. 3” in Para.[0075]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings Park with the above teachings of Tsai in order to improve user experience.
As per claim 14, Park and Tsai teach all of limitation of claim 13.
Park is silent about comprising: inputting the image frame and the one or more object maps into a machine-learned layer renderer model; and receiving, as output from the machine-learned layer renderer model, a background layer illustrative of a background of the video data and one or more object layers respectively associated with one of the one or more object maps.
Tsai teaches comprising: inputting the image frame and the one or more object maps into a machine-learned layer renderer model; and receiving, as output from the machine-learned layer renderer model, a background layer illustrative of a background of the video data and one or more object layers respectively associated with one of the one or more object maps (“the compute components 110 can include a central processing unit (CPU) 112, a graphics processing unit (GPU) 114, a digital signal processor (DSP) 116, and an image signal processor (ISP) 118. The compute components 110 can perform various operations such as image enhancement, object or image segmentation, computer vision, graphics rendering, augmented reality, image/video processing, sensor processing, recognition (e.g., text recognition, object recognition, feature recognition, tracking or pattern recognition, scene change recognition, etc.), disparity detection, machine learning, filtering, depth-of-field effect calculations or renderings, and any of the various operations described herein. In some examples, the compute components 110 can implement the image processing engine 120, the neural network 122, and the rendering engine 124. In other examples, the compute components 110 can also implement one or more other processing engines” in Para.[0052], “the image processing system 100 can obtain a disparity map or “depth map” for the image and binarize the disparity map to values of 0 or 1 (e.g., [0, 1]), which can provide an indication of the potential location of a target or object of interest in the image or FOV. In some cases, the disparity map can represent apparent pixel differences, motion, or depth (the disparity). Typically, objects that are close to the image sensor that captured the image will have greater separation or motion (e.g., will appear to move a significant distance) while objects that are further away will have less separation or motion. Such separation or motion can be captured by the disparity values in the disparity map. Thus, the disparity map can provide an indication of which objects are likely within a region of interest (e.g., the foreground), and which are likely not within the region of interest” in Para.[0152], “the disparity information for the disparity map can be obtained from hardware (e.g., an image sensor or camera device). For example, the disparity map can be generated based on auto-focus information from hardware (e.g., image sensor, camera, etc.) used to produce the image. The auto-focus information can help identify where a target or object of interest (e.g., a foreground object) is likely to be in the FOV. To illustrate, an auto-focus function can be leveraged to help the image processing system 100 identify where the target or object of interest is likely to be in the FOV. In some examples, an auto-focus function on hardware can automatically adjust a lens setting to set the optical focal points on the target or object of interest. When the image processing system 100 checks the disparity map, the scene behind the target or object of interest can have a negative disparity value, while the scene before the target or object of interest can have a positive disparity value and areas around the target or object of interest can contain a disparity value closer to zero” in Para.[0153]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings Park with the above teachings of Tsai in order to improve user experience.
As per claim 15, Park and Tsai teach all of limitation of claim 14.
Park is silent about wherein the machine-learned layer renderer model comprises a neural network.
Tsai teaches wherein the machine-learned layer renderer model comprises a neural network (“The compute components 110 can perform various operations such as image enhancement, object or image segmentation, computer vision, graphics rendering, augmented reality, image/video processing, sensor processing, recognition (e.g., text recognition, object recognition, feature recognition, tracking or pattern recognition, scene change recognition, etc.), disparity detection, machine learning, filtering, depth-of-field effect calculations or renderings, and any of the various operations described herein. In some examples, the compute components 110 can implement the image processing engine 120, the neural network 122, and the rendering engine 124. In other examples, the compute components 110 can also implement one or more other processing engines” in Para.[0052]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings Park with the above teachings of Tsai in order to improve user experience.
As per claim 16, Park and Tsai teach all of limitation of claim 15.
Park is silent about wherein the machine-learned layer renderer model has been trained based at least in part on a reconstruction loss, a mask loss, and a regularization loss.
Tsai teaches wherein the machine-learned layer renderer model has been trained based at least in part on a reconstruction loss, a mask loss, and a regularization loss (“the neural network 122 can adjust the weights of the nodes using a training process such as backpropagation. Backpropagation can include a forward pass, a loss function, a backward pass, and a weight update. The forward pass, loss function, backward pass, and parameter update is performed for one training iteration. The process can be repeated for a certain number of iterations for each set of training data (e.g., image data) until the weights of the layers 502, 504, 506 in the neural network 122 are accurately tuned” in Para.[0127]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings Park with the above teachings of Tsai in order to improve user experience.
As per claim 18, Park teaches all of limitation of claim 11.
Park is silent about wherein the one or more object layers comprise one or more object maps.
Tsai teaches wherein the one or more object layers comprise one or more object maps(“the image processing system 100 can generate a saliency map (S.sub.crf). The saliency map (S.sub.crf) can be generated based on the saliency refinement of the ranking map S at block 220. In some examples, the image processing system 100 can generate the saliency map (S.sub.crf) using the respective probability of each pixel being salient, which can be determined based on the refined ranking map S” in Para.[0088]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings Park with the above teachings of Tsai in order to improve user experience.
As per claim 19, Park and Tsai teach all of limitation of claim 18.
Park is silent about wherein the one or more object maps comprise one or more texture maps.
Tsai teaches wherein the one or more object maps comprise one or more texture maps (“the set of features is detected using a trained network, and wherein the set of features comprises at least one of semantic features, texture information, and color components” in Claim 3).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings Park with the above teachings of Tsai in order to improve user experience.
Park in view of Popov
Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Park et al.(USPubN 2015/0127626; hereinafter Park) in view of Popov et al.(USPubN 2021/0156960; hereinafter Popov).
As per claim 17, Park teaches all of limitation of claim 11.
Park teaches wherein, for each image frame, each of the one or more object layers comprises image data illustrative of the first object such that the one or more object layers and the background layer can be re-combined with modified relative timings (“In summarized video I (b) and summarized video II (c), video data in which the object A and the object B appear is extracted to reduce the play time of the original video into a time T2 and a time T3, respectively. The summarized video I (b) and the summarized video II (c) have different overlapping degrees between the object A and the object B. Under a condition that the object B having a temporally later appearing order than the object A does not appear before the object A, the degree of overlapping between the object A and the object B in the summarized video may be adjusted by adjusting the complexity of the summarized video” in Para.[0061], “a single summarized video that is being reproduced, in which a plurality of vehicles appearing for a predetermined period of time is rendered on a background model, may be changed into a plurality of summarized video layers including a first layer (a) that is a summarized video of an event corresponding to appearance of a first vehicle at 1:37 AM, a second layer (b) that is a summarized video of an event corresponding to appearance of a second vehicle at 6:08 AM, and a third layer (c) that is a summarized video of an event corresponding to appearance of a third vehicle at 1:24 PM, in response to the image change request. The lowest layer of the screen is a background model BG and a summarized video indicating a temporally preceding event is situated at a higher layer.” in Para.[0067], Fig. 3 and 5).
Park is silent about wherein, for each image frame, each of the one or more object layers comprises one or more trace effects at least partially attributable to the first object.
Popov teaches wherein, for each image frame, each of the one or more object layers comprises one or more trace effects at least partially attributable to the first object (“the projection image and corresponding reflection characteristics may be stored in multiple layers of a tensor, with pixel values for the different layers storing different reflection characteristics.” in Para.[0091]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings Park with the above teachings of Popov in order to improve user experience.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the claims at issue are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); and In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on a nonstatutory double patenting ground provided the reference application or patent either is shown to be commonly owned with this application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The USPTO internet Web site contains terminal disclaimer forms which may be used. Please visit http://www.uspto.gov/forms/. The filing date of the application will determine what form should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to http://www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp.
Claims 1-20 are rejected on the ground of nonstatutory obviousness-type double patenting as being unpatentable over claims 1-22 of U.S. Patent No. 12,243,145. Although the conflicting claims at issue are not identical, they are not patentably distinct from each other. See the reasons sets forth below:
Instance Application No. 19/046,007
U.S. Patent No. 12,243,145
1. A computing system configured to decompose video data into a plurality of layers, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that store instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising: obtaining, by the computing system, video data comprising a plurality of image frames; processing, by the computing system, the video data to generate one or more object layers comprising at least a first object layer; obtaining, by the computing system, data comprising one or more user-defined temporal alignment points for the first object layer; retiming, by the computing system, a background layer of the video data and the first object layer based at least in part on the one or more temporal alignment points to alter a playback speed of the first object layer relative to the background layer to generate an updated video; and providing, by the computing system, the updated video for display via a user interface.
2. The computing system of claim 1, wherein the first object layer comprises a first object and one or more trace effects associated with the first object.
3. The computing system of claim 2, wherein, for each image frame, each of the one or more object layers comprises image data illustrative of the first object and one or more trace effects at least partially attributable to the first object such that the one or more object layers and the background layer can be re-combined with modified relative timings.
4. The computing system of claim 1, wherein the background layer and the one or more object layers comprise one or more color channels and an opacity matte.
5. The computing system of claim 1, the operations comprising: inputting, by the computing system, a first image frame and one or more object maps into a machine-learned layer renderer model; and receiving, by the computing system as output from the machine-learned layer renderer model, a background layer illustrative of a background of the video data and one or more object layers respectively associated with one of the one or more object maps.
6. The computing system of claim 1, wherein the one or more object layers comprise one or more object maps.
7. The computing system of claim 6, wherein the one or more object maps comprise one or more re-sampled texture maps.
8. The computing system of claim 6, wherein the one or more object maps comprise one or more texture maps.
9. The computing system of claim 8, wherein obtaining, by the computing system, one or more object maps comprises: obtaining, by the computing system, one or more UV maps, each of the UV maps indicative of the first object of one or more objects depicted within the one or more frames; obtaining, by the computing system, a background deep texture map and one or more object deep texture maps; and resampling, by the computing system, the one or more object deep texture maps based at least in part on the one or more UV maps.
10. The computing system of claim 1, the operations comprising: transferring, by the computing system, high resolution details of the video data in a post processing step subsequent to receiving the background layer and the one or more object layers.
11. A computer implemented method comprising: obtaining, by a computing system, video data comprising a plurality of image frames; automatically, by the computing system, recognizing one or more objects within a first image frame of the plurality of image frames; generating, by the computing system, a first layer for a first object of the one or more objects; generating, by the computing system, a second layer for a background of the video data; selecting, by the computing system, the first layer for display; selecting, by the computing system, the second layer to be offto omit the background of the video data; and providing, by the computing system, the first layer for display without the background of the video data via a graphical user interface.
12. The computer implemented method of claim 11, comprising: transferring high resolution details of the video data in a post processing step subsequent to receiving the background layer and the one or more object layers.
13. The computer implemented method of claim 11, comprising: generating one or more object maps, wherein each of the one or more object maps is descriptive of a respective location of at least one object of the one or more objects within the image frame.
14. The computer implemented method of claim 13, comprising: inputting the image frame and the one or more object maps into a machine-learned layer renderer model; and receiving, as output from the machine-learned layer renderer model, a background layer illustrative of a background of the video data and one or more object layers respectively associated with one of the one or more object maps.
15. The computer implemented method of claim 14, wherein the machine- learned layer renderer model comprises a neural network.
16. The computer implemented method of claim 15, wherein the machine- learned layer renderer model has been trained based at least in part on a reconstruction loss, a mask loss, and a regularization loss.
17. The computer implemented method of claim 11, wherein, for each image frame, each of the one or more object layers comprises image data illustrative of at least one object and one or more trace effects at least partially attributable to the at least one object such that the one or more object layers and the background layer can be re-combined with modified relative timings.
18. The computer implemented method of claim 11, the first layer comprises one or more object maps.
19. The computer implemented method of claim 18, wherein the one or more object maps comprise one or more texture maps.
20. One or more non-transitory computer readable media storing instructions that are executable by one or more processors to perform operations comprising: obtaining video data comprising a plurality of video frames; processing the video data to generate one or more object layers comprising at least a first object layer; obtaining data comprising one or more user-defined temporal alignment points for the first object layer; retiming a background layer of the video data and the first object layer based at least in part on the one or more temporal alignment points to alter a playback speed of the first object layer relative to the background layer to generate an updated video; and providing the updated video for display via a user interface.
1. A computer-implemented method for decomposing videos into multiple layers that can be individually retimed and re-combined with modified relative timings, the computer-implemented method comprising: obtaining, by a computing system comprising one or more computing devices, video data, the video data comprising a plurality of image frames depicting one or more objects; and for each of the plurality of image frames: generating, by the computing system, one or more object maps, wherein each of the one or more object maps is descriptive of a respective location of at least one object of the one or more objects within the image frame; inputting, by the computing system, the image frame and the one or more object maps into a machine-learned layer renderer model, comprising iteratively individually inputting each of the one or more object maps into the machine-learned layer renderer model; receiving, by the computing system as output from the machine-learned layer renderer model, a background layer illustrative of a background of the video data and one or more object layers respectively associated with one of the one or more object maps, wherein each of the one or more object layers comprises image data illustrative of the at least one object and one or more trace effects at least partially attributable to the at least one object; and generating, by the computing system, a retimed video by: retiming at least one of the background layer or the one or more object layers; and re-combining the one or more retimed lavers.
2. The computer-implemented method of claim 1, wherein inputting, by the computing system, the image frame and the one or more object maps into the machine-learned layer renderer model comprises iteratively individually receiving, as output from the machine-learned layer renderer model and by the computing system, each of the one or more object layers respective to the one or more object maps.
3. The computer-implemented method of claim 1, wherein the background layer and the one or more object layers comprise one or more color channels and an opacity matte.
4. The computer-implemented method of claim 1, wherein the machine-learned layer renderer model comprises a neural network.
5. The computer-implemented method of claim 1, wherein the machine-learned layer renderer model has been trained based at least in part on a reconstruction loss, a mask loss, and a regularization loss.
6. The computer-implemented method of claim 5, wherein the training was performed on downsampled video and then upsampled.
7. The computer-implemented method of claim 1, wherein the one or more object maps comprise one or more texture maps.
8. The computer-implemented method of claim 1, wherein the one or more object maps comprise one or more re-sampled texture maps.
9. The computer-implemented method of claim 8, wherein obtaining, by the computing system, one or more object maps comprises: obtaining, by the computing system, one or more UV maps, each of the UV maps indicative of the at least one object of the one or more objects depicted within the one or more frames; obtaining, by the computing system, a background deep texture map and one or more object deep texture maps; and resampling, by the computing system, the one or more object deep texture maps based at least in part on the one or more UV maps.
10. The computer-implemented method of claim 9, wherein generating, by the computing system, the one or more UV maps comprises: identifying, by the computing system, one or more keypoints; and obtaining, by the computing system, one or more UV maps based on the one or more keypoints.
11. The computer-implemented method of claim 1, further comprising: transferring, by the computing system, high resolution details of the video data in a post processing step subsequent to receiving the background layer and the one or more object layers.
12. A computing system configured to decompose video data into a plurality of layers, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that store instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising: obtaining video data, the video data comprising a plurality of image frames depicting one or more objects; and for each of the plurality of image frames: generating one or more object maps, wherein each of the one or more object maps is descriptive of a respective location of at least one object of the one or more objects within the image frame; inputting the image frame and the one or more object maps into a machine-learned layer renderer model, comprising iteratively individually inputting each of the one or more object maps into the machine-learned layer renderer model; receiving, as output from the machine-learned layer renderer model, a background layer illustrative of a background of the video data and one or more object layers respectively associated with one of the one or more object maps, wherein each of the one or more object layers comprises image data illustrative of the at least one object and one or more trace effects at least partially attributable to the at least one object; and generating, by the computing system, a retimed video by: retiming at least one of the background layer or the one or more object layers; and re-combining the one or more retimed layers.
13. The computing system of claim 12, wherein inputting the image frame and the one or more object maps into the machine-learned layer renderer model comprises iteratively individually receiving, as output from the machine-learned layer renderer model and by the computing system, each of the one or more object layers respective to the one or more object maps.
14. The computing system of claim 12, wherein the background layer and the one or more object layers comprise one or more color channels and an opacity matte.
15. The computing system of claim 12, wherein the machine-learned layer renderer model comprises a neural network.
16. The computing system of claim 12, wherein the machine-learned layer renderer model has been trained based at least in part on a reconstruction loss, a mask loss, and a regularization loss.
17. The computing system of claim 16, wherein the training was performed on downsampled video and then upsampled.
18. The computing system of claim 12, wherein the one or more object maps comprise one or more texture maps.
19. The computing system of claim 12, wherein obtaining one or more object maps comprises: obtaining one or more UV maps, each of the UV maps indicative of the at least one object of the one or more objects depicted within the one or more frames; obtaining a background deep texture map and one or more object deep texture maps; and resampling the one or more object deep texture maps based at least in part on the one or more UV maps.
20. The computing system of claim 19, wherein obtaining the one or more UV maps comprises: identifying one or more keypoints; and generating one or more UV maps based on the one or more keypoints.
21. The computing system of claim 12, wherein the instructions further comprise: transferring high resolution details of the video data in a post processing step subsequent to receiving the background layer and the one or more object layers.
22. The computing system of claim 12, wherein the one or more trace effects comprise at least one of: shadows, reflections, splashes, or motion of loose clothing.
Claims 1-20 are anticipated by U.S. Patent No. 12,243,145 claims 1-22 as show in the table above.
Allowable Subject Matter
Claim 1-10, and 20 would be allowable if rewritten to overcome the rejection(s) under the ground of nonstatutory obviousness-type double patenting, set forth in this office action and to include all of the limitations of the base claim and any intervening claims.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SUNGHYOUN PARK whose telephone number is (571)270-1333. The examiner can normally be reached M - Thur 6:00 am - 4 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, THAI Q TRAN can be reached at (571)272-7382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SUNGHYOUN PARK/Examiner, Art Unit 2484