Prosecution Insights
Last updated: August 18, 2026
Application No. 18/332,939

METHOD AND APPARATUS WITH FRAME IMAGE RECONSTRUCTION

Non-Final OA §103
Filed
Jun 12, 2023
Priority
Dec 09, 2022 — RE 10-2022-0171259
Examiner
BONANSINGA, AARON TIMOTHY
Art Unit
2673
Tech Center
2600 — Communications
Assignee
Samsung Electronics Co., Ltd.
OA Round
3 (Non-Final)
77%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
27 granted / 35 resolved
+15.1% vs TC avg
Strong +35% interview lift
Without
With
+34.8%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
17 currently pending
Career history
58
Total Applications
across all art units

Statute-Specific Performance

§101
5.5%
-34.5% vs TC avg
§103
76.2%
+36.2% vs TC avg
§102
10.4%
-29.6% vs TC avg
§112
7.9%
-32.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 35 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 05/12/2026 has been entered. Response to Arguments Applicant’s arguments (see remarks), filed 05/12/2026, with respect to the claim 1-7, 10-17 and 19-20 have been fully considered but are moot because they do not apply to the current references and current combinations of references being used in the current rejection. In addition, the prior art by FLEISHMAN (US 20190043203 A1) was used to teach "a previous reconstruction image of a previous frame having a different second resolution" as recited in claim 1. This change was made in light of further consideration of the prior art, claimed invention and amendments. Based on the breadth of the claim language, FLEISHMAN (US 20190043203 A1) explicitly teaches obtaining a previous reconstruction image of a previous frame having a different second resolution (Fig. 4. Paragraph [0045]-FLEISHMAN disclosed process 400 may include “obtain a video sequence of frames of image data and comprising a current frame” 402 (wherein the current frame is reconstructed from previously reconstructed frame/semantic maps and the current frame may undergo pre-processing operations). The pre-processing could include resolution reduction. Therefore, it would have been obvious to a person of ordinary skill in the art to obtain a previous reconstruction image having a different second resolution. FLEISHMAN explicitly teaches obtaining a current reconstructed frame from a previous reconstructed image where the current frame may undergo resolution reduction, which can result in previous/current frames having different resolutions and a previous frame having a greater resolution. As FLEISHMAN states at paragraph [0068], “pre-processing operations converts the rendered semantic map input to an expected format for input to a neural network”. Thus, it would be obvious to a person of ordinary skill to use a different resolution and/or greater resolution for the reconstructed frame because it ensures frames/models/maps are converted into the correct format, allows for reconstructed frames to be generated from different resolutions, and improves both computational efficiency and frame rendering by rescaling frames to either a higher or lower resolution. Please also read paragraph [0059 and 0068]). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2, 4-7, 11-13, 15-17 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over FLEISHMAN et al. (US 20190043203 A1), hereinafter referenced as FLEISHMAN in view of MALLYA et al. (US 20210374552 A1), hereinafter referenced as MALLYA in further view of POTTORFF et al. (US 20240098216 A1), hereinafter referenced as POTTORFF. Regarding claim 1, FLEISHMAN explicitly teaches a processor-implemented method (Fig. 4. Paragraph [0044]-FLEISHMAN discloses referring to FIG. 4, a process 400 is provided for a method and system of recurrent semantic segmentation for imaging processing. In paragraph [0083]-FLEISHMAN discloses any one or more of the operations of FIGS. 4, and 5A-5B may be undertaken in response to instructions provided by one or more computer program products. Such program products may include signal bearing media providing instructions that, when executed by a processor, may provide the functionality described herein. The computer program products may be provided in any form of one or more machine-readable media. Thus, a processor including one or more processor core(s) may undertake one or more of the operations of the example processes herein in response to program code and/or instructions or instruction sets conveyed to the processor by one or more computer or machine-readable media. Please also see Fig. 1 read paragraph [0030, 0034, 0056]), the method comprising: generating a semantic map (Fig. 6, #622, #700 and #804 called a semantic map. Paragraph [0047-0052, 0056 and 0065]. In paragraph [0047]-FLEISHMAN discloses process 400 may optionally include “recurrently generate a semantic segmentation map. Please also see Fig. 4-8) indicating a visualization property assigned to a first object (Fig. 3. Paragraph [0039]-FLEISHMAN discloses referring to FIG. 3, an image 300 shows the room 202 (now 301) from image 200 except now with the disclosed example 3D semantic segmentation applied. Each voxel color or shade represents a class of an object that it belongs to. In paragraph [0065]-FLEISHMAN discloses an example segmentation map 700 is provided at the current pose (or k-pose) of the current frame, and where the walls 702, chairs 704, and floor 706 shown in the map 700 are segmented from each other and each have an initial, historically-based (or influenced or based on information of previous frames) semantic label (wall, chair, floor for example). Please also read paragraph [0068]) of an obtained frame image (Fig. 4. Paragraph [0045]-FLEISHMAN discloses process 400 may include “obtain a video sequence of frames of image data and comprising a current frame” 402. Please also see Fig. 1 and 5-6 and read) having a first resolution (Fig. 4. Paragraph [0045]-FLEISHMAN discloses the pre-processing could include resolution reduction. Please also read paragraph [0042 and 0056]), wherein the obtained frame image is a current frame image and the semantic map is a current semantic map of the current frame image (Fig. 4. Paragraph [0047]-FLEISHMAN discloses individual semantic segmentation maps are each associated with a different current frame from the video sequence” 404. This may include generating a 3D semantic segmentation model, based on a 3D geometric model (generated by use of RGB-SLAM for example) with semantic labels registered to the model. Once established, the 3D semantic segmentation model may be projected to an image plane to form a segmentation map with the semantic labels from the model that have pixels or voxels on that plane. The image plane may be the plane formed by the camera pose of the current frame being analyzed. The 3D semantic segmentation model may be updated with semantic segment labels each current frame being semantically analyzed so that the 3D semantic model reflects or represents the history of the semantic segmentation of the 3D space represented by the 3D semantic model up to a current point in time. Please also see Fig. 1 and 5-6); obtaining a previous reconstruction image of a previous frame (Fig. 4. Paragraph [0050]-FLEISHMAN discloses process 400 may include “generate a current and historical semantically segmented frame comprising using both the current semantic features and the historically-influenced semantic features as input to a neural network that indicates semantic labels for areas of the current historical semantically segmented frame” 410. Please also read paragraph [0030, 0061]) having a different second resolution (Fig. 4. Paragraph [0045]-FLEISHMAN disclosed process 400 may include “obtain a video sequence of frames of image data and comprising a current frame” 402 (wherein the current frame is reconstructed from previously reconstructed frame/semantic maps and the current frame may undergo pre-processing operations). The pre-processing could include resolution reduction. Therefore, it would have been obvious to a person of ordinary skill in the art to obtain a previous reconstruction image having a different second resolution. FLEISHMAN explicitly teaches obtaining a current reconstructed frame from a previous reconstructed image where the current frame may undergo resolution reduction, which can result in a previous/current frames having different resolutions and a previous frame having a greater resolution. As FLEISHMAN states at paragraph [0068], “pre-processing operations converts the rendered semantic map input to an expected format for input to a neural network”. Thus, it would be obvious to a person of ordinary skill to use a different resolution and/or greater resolution for the reconstructed frame because it ensures frames/models/maps are converted into the correct format, allows for reconstructed frames to be generated from different resolutions, and improves both computational efficiency and frame rendering by rescaling frames to either a higher or lower resolution. Please also read paragraph [0059 and 0068]); and Although FLEISHMAN explicitly teaches and generating a reconstruction image using an image reconstruction machine learning model provided input based on the obtained frame image, and the semantic map (Fig. 4. Paragraph [0048]-FLEISHMAN disclosed process 400 may include “extract historically-influenced semantically semantic features of the semantic segmentation map” 406. The result of such extraction may be considered historically-influenced features that represent the semantic labeling in the segmentation map. In paragraph [0049]-FLIESHMAN discloses process 400 may include “extract current semantic features of the current frame” 408 (wherein the extraction of both historical and current semantic features is performed by a neural network). In paragraph [0050]-FLEISHMAN discloses process 400 may include “generate a current and historical semantically segmented frame comprising using both the current semantic features and the historically-influenced semantic features as input to a neural network that indicates semantic labels for areas of the current historical semantically segmented frame” 410 (wherein the model may be generated by concatenating current semantic features and historically-influenced semantic features and inputting them into a neural network, such as a CNN). Please also read paragraph [0052]), having the different second resolution and including a second object having a visualization property indicated by the semantic map (Fig. 5. Paragraph [0065]-FLEISHMAN discloses referring to FIG. 7, process 500 may include “render segmentation map from 3D semantic model” 516, and this may involve obtaining the k-pose of the current frame being analyzed, and then projecting the 3D semantic model to an image plane formed by a camera at the k-pose. An example segmentation map 700 is provided at the current pose (or k-pose) of the current frame, and where the walls 702, chairs 704, and floor 706 shown in the map 700 are segmented from each other and each have an initial, historically-based (or influenced or based on information of previous frames) semantic label (wall, chair, floor for example). Objects that are adjacent each other and have the same label may not show as separate components on the segmentation map. Please also read paragraph [0045]). FLIESHMAN fails to explicitly teach obtaining a warped image by warping the previous reconstruction image of the previous frame to the current frame image based on a motion vector map between the current frame image and the previous reconstruction image of the previous frame, wherein warping the previous reconstruction image comprises forward warping; and generating a reconstruction image using an image reconstruction machine learning model provided input based on the obtained warped image. However, MALLAYA explicitly teaches obtaining a warped image (Fig. 3, #320 called an output frame and a flow-warped prior output. Paragraph [0063]) by warping the previous reconstruction image (Fig. 3, #308 called Flow-Warped Prior Output. Paragraph [0063]) of the previous frame (Fig. 3, #312 called a prior output or previously synthesized output frame. Paragraph [0063]) to the current frame image (Fig. 3. Paragraph [0047]-MALLAYA discloses video synthesis system 200 can generate high-level semantic inputs from images or video captured by one or more cameras 202, or otherwise obtained. System 200 can utilize these inputs with one or more neural networks to synthesize or otherwise generate photorealistic images, image sequences, or video (wherein video synthesis system includes at least an input label embedding network 302, an image encoder network 314, a flow embedding network 306, and an image generator network 318). In paragraph [0063]-MALLAYA discloses input video is captured 402 for a scene using one or more physical cameras or other such devices. This video data can be analyzed 404 to generate input labels or embeddings, as may relate to segmentation maps, depth maps, edge maps, or pose data that can be concatenated. An optical flow-modified version of a prior input frame can be generated 406 and processed by a flow embedding network. Guidance images can be generated 408 from a point cloud for this scene, where points of that point cloud can have appearance data (e.g., color, texture, or reflectivity) assigned from when those points appeared in one or more previously generated video frames. A prior frame can be provided 410 as input to an image encoder network. Input labels or embeddings, flow-modified frame data, and a guidance image can be provided 412 as input at sequential layers of an image generator network to generate an output frame, where those layers can be in a sequence with a set of upscaling layers (wherein image generator network 318 may initially generate a small image such as 16 x 32, where each layer double the size of the image until reaching a final output frame 320 that has been upscaled to a greater resolution, such as high definition, 4K, or 8K resolution). Please also read paragraph [0056]) based on a motion vector map between the current frame image and the previous reconstruction image of the previous frame (Fig. 3. Paragraph [0048]-MALLAYA discloses an optical flow module 216 can analyze recent video frames to determine optical flow data that can be used to warp and generate an image conditioned on recent generated images [0055]-MALLAYA discloses guidance can include use of a motion vector. In paragraph [0060]-MALLAY discloses γ.sub.flow and +β.sub.flow are generated using flow-embedding network 306 applied on an optical flow-warped previous frame. This provides additional constraints that a generated output frame 320 should be consistent even in dynamic regions (wherein a transformation using optical flow between a previous and current image is forward warping). Please also read claim 2 and paragraph [0063]); and generating a reconstruction image using an image reconstruction machine learning model (Fig. 3. Paragraph [0047]-MALLAYA discloses video synthesis system 200 can generate high-level semantic inputs from images or video captured by one or more cameras 202, or otherwise obtained. System 200 can utilize these inputs with one or more neural networks to synthesize or otherwise generate photorealistic images, image sequences, or video. In paragraph [0057]-MALLAYA discloses a video synthesis system includes at least four networks, or sub-networks. These networks can include an input label embedding network 302, an image encoder network 314, a flow embedding network 306, and an image generator network 318) provided input based on the obtained warped image (Fig. 3. Paragraph [0048]-MALLAYA discloses an optical flow module 216 can analyze recent video frames to determine optical flow data that can be used to warp and generate an image conditioned on recent generated images (wherein the synthesized images/sequences/video may be based on either one, several or all previously frames/reconstructed frames). Please also read paragraph [0062-0063]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of FLEISHMAN of having a processor-implemented method, the method comprising: generating a semantic map indicating a visualization property assigned to a first object of an obtained frame image having a first resolution, wherein the obtained frame image is a current frame image and the semantic map is a current semantic map of the current frame image; obtaining a previous reconstruction image of a previous frame having a different second resolution, with the teachings of MALLYA of having obtaining a warped image by warping the previous reconstruction image of the previous frame to the current frame image based on a motion vector map between the current frame image and the previous reconstruction image of the previous frame; and generating a reconstruction image using an image reconstruction machine learning model provided input based on the obtained warped image. Wherein FLEISHMAN’s method having obtaining a warped image by warping the previous reconstruction image of the previous frame to the current frame image based on a motion vector map between the current frame image and the previous reconstruction image of the previous frame, wherein warping the previous reconstruction image comprises forward warping or backwards warping; and generating a reconstruction image using an image reconstruction machine learning model provided input based on the obtained warped image, the obtained frame image, and the semantic map, having the different second resolution and including a second object having a visualization property indicated by the semantic map. The motivation behind the modification would have been to obtain a method that improves the efficiency, quality and efficacy of image segmentation and reconstruction, since both FLEISHMAN and MALLYA concern the use of neural networks for image reconstruction and semantic map generation. Wherein FLEISHMAN’s provides improves the accuracy of the semantic labels by using historical data along with the efficiency of semantic segmentation, which, in turn, permits the sematic segmentation to be performed on smaller devices, while MALLYA’s systems and methods that improve neural network training speed, memory performance and the ability to generate images or video that are consistent in appearance over time, viewer, camera, session, or other such variants. Please see FLEISHMAN et al. (US 20190043203 A1), Paragraph [0032, 0040, and 0056] and MALLYA et al. (US 20210374552 A1), Abstract and Paragraph [0002, 0172, 0335, 0340]. FLEISHMAN in view of MALLAYA fails to explicitly teach wherein warping the previous reconstruction image comprises backwards warping. However, POTTORFF explicitly teaches wherein warping the previous reconstruction image (Fig. 1. Paragraph [0091]-POTTORFF discloses FIG. 1 illustrates an example diagram 100 where blending factors for frame motion are generated using a neural network. A processor 102 executes or otherwise performs one or more instructions to use a neural network 110 to generate blending factors of frame motion. Processor 102 uses neural network 110 to generate blending factors of frame motion that are used in frame interpolation. Processor 102 uses neural network 110 to generate blending factors used in frame motion to be used to perform deep-learning based frame interpolation. Inputs to neural network 110 comprise one or more frames (e.g., a previous frame 104 and/or a current frame 106) and additional frame information including motion information of pixels of previous frame 104 and/or current frame 106. In paragraph [0100]-POTTORFF discloses previous frame 104 and/or current frame 106 are generated by spatial upsampling (e.g. by spatial super sampling such as, for example, DLSS, XeSS (or XeSS) from Intel®, FidelityFX™ Super Resolution from AMD®, etc.). Please also see Fig. 3) comprises backwards warping (Fig. 3. Paragraph [0174]-POTTORFF discloses one or more flow warped intermediate images are generated based on, for example, forward optical flow vectors, reverse optical flow vectors, or other such flow vectors. In paragraph [0176]-POTTORFF discloses one or more intermediate images (e.g., generated using blending factors at step 1012) are blended together to generate an intermediate result such as, for example, blended previous to current intermediate frame 902 or blended current to previous intermediate frame 904. Please also read paragraph [01). more flow warped intermediate images are generated using systems and methods such as those described herein. In at least one embodiment, at step 1010, one or more flow warped intermediate images are generated based on, for example, forward optical flow vectors, reverse optical flow vectors, or other such flow vectors. Please also read paragraph [0153-0163, 0643-0646 and 0651-0652]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of FLEISHMAN in view of MALLYA of having a processor-implemented method, the method comprising: generating a semantic map indicating a visualization property assigned to a first object of an obtained frame image having a first resolution, wherein the obtained frame image is a current frame image and the semantic map is a current semantic map of the current frame image, with the teachings of POTTORFF of having wherein warping the previous reconstruction image comprises backwards warping. Wherein FLEISHMAN’s method having wherein warping the previous reconstruction image comprises forward warping or backwards warping. The motivation behind the modification would have been to obtain a method that improves the efficiency, quality and efficacy of image segmentation and reconstruction, since both FLEISHMAN and POTTORFF concern the use of neural networks for image reconstruction. Wherein FLEISHMAN’s provides improves the accuracy of the semantic labels by using historical data along with the efficiency of semantic segmentation, which, in turn, permits the sematic segmentation to be performed on smaller devices, while POTTORFF’s systems and methods improve neural network training and the efficiency, accuracy, and efficacy of image processing, image reconstruction and segmentation. Please see FLEISHMAN et al. (US 20190043203 A1), Paragraph [0032, 0040, and 0056] and POTTORFF et al. (US 20240098216 A1), Abstract and Paragraph [0124 and 0518]. Regarding claim 2, FLEISHMAN in view of MALLYA and in further view of POTTORFF explicitly teach the method of claim 1, FLEISHMAN further teaches wherein the generating of the semantic map (Fig. 4. Paragraph [0044]-FLEISHMAN discloses referring to FIG. 4, a process 400 is provided for a method and system of recurrent semantic segmentation for imaging processing) comprises: obtaining semantic data for the visualization property (Fig. 4. Paragraph [0047]-FLEISHMAN discloses process 400 may include “recurrently generate a semantic segmentation map in a view of a current pose of the current frame and comprising obtaining data to form the semantic segmentation map from a 3D semantic segmentation model, wherein individual semantic segmentation maps are each associated with a different current frame from the video sequence” 404. Once established, the 3D semantic segmentation model may be projected to an image plane to form a segmentation map with the semantic labels from the model that have pixels or voxels on that plane); and generating, based on an object identifier map comprising the semantic data and regions classified by plural objects of the obtained frame image, a semantic map indicating a corresponding visualization property of a corresponding object of the plural objects (Fig. 5. Paragraph [0065]-FLEISHMAN discloses referring to FIG. 7, process 500 may include “render segmentation map from 3D semantic model” 516, and this may involve obtaining the k-pose of the current frame being analyzed, and then projecting the 3D semantic model to an image plane formed by a camera at the k-pose. An example segmentation map 700 is provided at the current pose (or k-pose) of the current frame, and where the walls 702, chairs 704, and floor 706 shown in the map 700 are segmented from each other and each have an initial, historically-based (or influenced or based on information of previous frames) semantic label (wall, chair, floor for example). Objects that are adjacent each other and have the same label may not show as separate components on the segmentation map). Regarding claim 4, FLEISHMAN in view of MALLYA and in further view of POTTORFF explicitly teach the method of claim 1, FLEISHMAN further teaches wherein the generating of the semantic map comprises generating a semantic map indicating one or more of a type, pattern, material, or shape of the first object (Fig. 2A. Paragraph [0057]-FLEISHMAN discloses the segmentation output unit 822 outputs semantic labels, or class or probabilities for the labels or classes, and provides them as part of the current-historical (C-H) semantically segmented (or just segmented) frame 824). In paragraph [0065] Referring to FIG. 7, process 500 may include “render segmentation map from 3D semantic model” 516. An example segmentation map 700 is provided at the current pose (or k-pose) of the current frame, and where the walls 702, chairs 704, and floor 706 shown in the map 700 are segmented from each other and each have an initial, historically-based (or influenced or based on information of previous frames) semantic label (wall, chair, floor for example). Please also read paragraph [0068]). Regarding claim 5, FLEISHMAN in view of MALLYA and in further view of POTTORFF explicitly teach the method of claim 1, FLEISHMAN further teaches wherein the generating of the semantic map (Fig. 3. Paragraph [0044]-FLEISHMAN discloses referring to FIG. 4, a process 400 is provided for a method and system of recurrent semantic segmentation for imaging processing. In paragraph [0047]-FLEISHMAN discloses process 400 may optionally include “recurrently generate a semantic segmentation map in a view of a current pose of the current frame and comprising obtaining data to form the semantic segmentation map from a 3D semantic segmentation model, wherein individual semantic segmentation maps are each associated with a different current frame from the video sequence” 404. Please also see Fig. 5A-B) comprises indicating rendering information including one or more of a color, a diffuse color, a depth, a normal line, a specular reflection, or an albedo of the obtained frame image together with the visualization property (Fig. 3. Paragraph [0039]-FLEISHMAN discloses referring to FIG. 3, an image 300 shows the room 202 (now 301) from image 200 except now with the disclosed example 3D semantic segmentation applied. Each voxel color or shade represents a class of an object that it belongs to. Please also read paragraph [0045-0046, 0056-0058, 0060 and 0065]). Regarding claim 6, FLEISHMAN in view of MALLYA and in further view of POTTORFF explicitly teach the method of claim 1, FLEISHMAN further teaches wherein the generating of the semantic map comprises indicating, for a plurality of objects in the obtained frame image, a visualization property for each object of the plurality of objects (Fig. 4. Paragraph [0057]-FLEISHMAN discloses the segmentation output unit 822 outputs semantic labels, or class or probabilities for the labels or classes, and provides them as part of the current-historical (C-H) semantically segmented (or just segmented) frame 824. Further in paragraph [0065]-FLEISHMAN discloses referring to FIG. 7, process 500 may include “render segmentation map from 3D semantic model” 516, and this may involve obtaining the k-pose of the current frame being analyzed, and then projecting the 3D semantic model to an image plane formed by a camera at the k-pose. An example segmentation map 700 is provided at the current pose (or k-pose) of the current frame, and where the walls 702, chairs 704, and floor 706 shown in the map 700 are segmented from each other and each have an initial, historically-based (or influenced or based on information of previous frames) semantic label (wall, chair, floor for example)). Regarding claim 7, FLEISHMAN in view of MALLYA and in further view of POTTORFF explicitly teach the method of claim 1, FLEISHMAN further teaches wherein the image reconstruction machine learning model (Fig. 4. Paragraph [0037]-FLEISHMAN discloses in the present solution, the recurrent segmentation operation (usage of the semantic information from previous frames as reflected in the 3D semantic model) is learned from the data and tailored to specific scenarios. In paragraph [0044]-FLEISHMAN discloses referring to FIG. 4, a process 400 is provided for a method and system of recurrent semantic segmentation for imaging processing (wherein the operations may be performed by a neural network). Please also read paragraph [0036 and 0051]) is a machine learning model trained using an objective function (Fig. 4. Paragraph [0078]-FLEISHMAN discloses the training of the architecture in a supervised-learning settings may include a training set of RGBD video-sequences, where the frames in each sequence have semantic information. Such a video can be obtained using either (i) a labor intensive method manually segmenting each frame, (ii) segmenting a reconstructed 3D model, or (iii) using synthetic data. See, Dai at el., “Richly-annotated 3D Reconstructions of Indoor Scenes”, Computer Vision and Pattern Recognition (CVPR) (2017). In paragraph [0079]-FLEISHMAN discloses training a recurrent network requires rendered semantic maps of the 3D semantic model. The training may be performed in several operations) calculated based on a second visualization property of a third object of a temporary output image and a third visualization property indicated by a training semantic map together with a difference between a temporary output image and a true value output image obtained from a training input image and the training semantic map (Fig. 4. Paragraph [0080]-FLEISHMAN discloses the first training operation may involve initialization by training a standard semantic-segmentation network. First, a standard single frame CNN-based semantic segmentation algorithm is trained. This resulting initial network may be denoted as n.sub.1 for example. In paragraph [0081]-FLEISHMAN discloses the next training operation may involve data preparation, which refers to generating training data for the recurrent architecture. Given the current network, training data was generated for the next recurrent phase in the form of a triplet (RGBD frame, rendered semantic map of the 3D semantic model, ground truth semantic segmentation). The system runs as shown in FIGS. 6 and 8 with the current network on short sequences of N frames, where N is a tunable parameter. A matching semantic map was rendered for the last frame from the last camera pose in each sequence, and then saved with the frame as training data for next stage. The semantic map was represented as an image of H*W pixels (the size of the frame) with C (the number of classes that the system supports) channels. Since only X<C probabilities are remembered in each voxel, lower C-X probabilities are truncated to zero, and the remaining X probabilities are renormalized to be a proper distribution). Regarding claim 11, FLEISHMAN in view of MALLYA and in further view of POTTORFF explicitly teach the method of claim 1, FLEISHMAN further teaches wherein the different second resolution is higher than the first resolution (Fig. 4. Paragraph [0045]-FLEISHMAN disclosed process 400 may include “obtain a video sequence of frames of image data and comprising a current frame” 402 (wherein the current frame is reconstructed from previously reconstructed frame/semantic maps and the current frame may undergo pre-processing operations). The pre-processing could include resolution reduction. Therefore, it would have been obvious to a person of ordinary skill in the art to obtain a previous reconstruction image having a different second resolution. FLEISHMAN explicitly teaches obtaining a current reconstructed frame from a previous reconstructed image where the current frame may undergo resolution reduction, which can result in a previous/current frames having different resolutions and a previous frame having a greater resolution. As FLEISHMAN states at paragraph [0068], “pre-processing operations converts the rendered semantic map input to an expected format for input to a neural network”. Thus, it would be obvious to a person of ordinary skill to use a different resolution and/or greater resolution for the reconstructed frame because it ensures frames/models/maps are converted into the correct format, allows for reconstructed frames to be generated from different resolutions, and improves both computational efficiency and frame rendering by rescaling frames to either a higher or lower resolution. Please also read paragraph [0059 and 0068]). Regarding claim 12, FLEISHMAN explicitly teaches an apparatus (Fig. 9, #900 called an image processing system. Paragraph [0087]. In paragraph [0044]-FLEISHMAN discloses referring to FIG. 4, a process 400 is provided for a method and system of recurrent semantic segmentation for imaging processing. In paragraph [0083]-FLEISHMAN discloses any one or more of the operations of FIGS. 4, and 5A-5B may be undertaken in response to instructions provided by one or more computer program products. Such program products may include signal bearing media providing instructions that, when executed by a processor, may provide the functionality described herein. The computer program products may be provided in any form of one or more machine-readable media. Thus, a processor including one or more processor core(s) may undertake one or more of the operations of the example processes herein in response to program code and/or instructions or instruction sets conveyed to the processor by one or more computer or machine-readable media. Please also see Fig. 1 and 10 and read paragraph [0030, 0034, 0056])), comprising: a processor (Fig. 9, #920 called processors. Paragraph [0093]-FLEISHMAN discloses the image processing system 900 may have one or more processors 920. Please also read paragraph [0083]) configured to: generate a semantic map (Fig. 6, #622, #700 and #804 called a semantic map. Paragraph [0047-0052, 0056 and 0065]. In paragraph [0047]-FLEISHMAN discloses process 400 may optionally include “recurrently generate a semantic segmentation map. Please also see Fig. 4-5 and 7-8) indicating a first visualization property assigned to a first object (Fig. 3. Paragraph [0039]-FLEISHMAN discloses referring to FIG. 3, an image 300 shows the room 202 (now 301) from image 200 except now with the disclosed example 3D semantic segmentation applied. Each voxel color or shade represents a class of an object that it belongs to. In paragraph [0065]-FLEISHMAN discloses an example segmentation map 700 is provided at the current pose (or k-pose) of the current frame, and where the walls 702, chairs 704, and floor 706 shown in the map 700 are segmented from each other and each have an initial, historically-based (or influenced or based on information of previous frames) semantic label (wall, chair, floor for example). Please also read paragraph [0068]) within a frame image having a first resolution (Fig. 4. Paragraph [0045]-FLEISHMAN discloses process 400 may include “obtain a video sequence of frames of image data and comprising a current frame” 402. This operation may include obtaining pre-processed raw image data. The pre-processing could include resolution reduction), wherein the frame image is a current frame image and the semantic map is a current semantic map of the current frame image (Fig. 4. Paragraph [0047]-FLEISHMAN discloses individual semantic segmentation maps are each associated with a different current frame from the video sequence” 404. This may include generating a 3D semantic segmentation model, based on a 3D geometric model (generated by use of RGB-SLAM for example) with semantic labels registered to the model. Once established, the 3D semantic segmentation model may be projected to an image plane to form a segmentation map with the semantic labels from the model that have pixels or voxels on that plane. The image plane may be the plane formed by the camera pose of the current frame being analyzed. The 3D semantic segmentation model may be updated with semantic segment labels each current frame being semantically analyzed so that the 3D semantic model reflects or represents the history of the semantic segmentation of the 3D space represented by the 3D semantic model up to a current point in time. Please also see Fig. 1 and 5-6); obtaining a previous reconstruction image of a previous frame having a different second resolution (Fig. 4. Paragraph [0045]-FLEISHMAN disclosed the pre-processing could include resolution reduction. Therefore, it would have been obvious to a person of ordinary skill in the art to obtain a previous reconstruction image having a different second resolution. FLEISHMAN explicitly teaches obtaining a current reconstructed frame from a previous reconstructed image where the current frame may undergo resolution reduction, which can result in a previous/current frames having different resolutions and a previous frame having a greater resolution. As FLEISHMAN states at paragraph [0068], “pre-processing operations converts the rendered semantic map input to an expected format for input to a neural network”. Thus, it would be obvious to a person of ordinary skill to use a different resolution and/or greater resolution for the reconstructed frame because it ensures frames/models/maps are converted into the correct format, allows for reconstructed frames to be generated from different resolutions, and improves both computational efficiency and frame rendering by rescaling frames to either a higher or lower resolution. Please also read paragraph [0059 and 0068]); and Although FLEISHMAN explicitly teaches and generating a reconstruction image, by using an image reconstruction machine learning model provided the frame image, and the semantic map (Fig. 4. Paragraph [0048]-FLEISHMAN disclosed process 400 may include “extract historically-influenced semantically semantic features of the semantic segmentation map” 406. The result of such extraction may be considered historically-influenced features that represent the semantic labeling in the segmentation map. In paragraph [0049]-FLIESHMAN discloses process 400 may include “extract current semantic features of the current frame” 408 (wherein the extraction of both historical and current semantic features is performed by a neural network). In paragraph [0050]-FLEISHMAN discloses process 400 may include “generate a current and historical semantically segmented frame comprising using both the current semantic features and the historically-influenced semantic features as input to a neural network that indicates semantic labels for areas of the current historical semantically segmented frame” 410 (wherein the model may be generated by concatenating current semantic features and historically-influenced semantic features and inputting them into a neural network, such as a CNN). Further in paragraph [0052]-FLEISHMAN discloses process 400 may include “semantically update the 3D semantic segmentation model comprising using the current and historical semantically segmented frame” 412, which refers to registering the semantic labels or probabilities of the segmentation frame to the 3D semantic model), having the different second resolution and including a second object having a second visualization property indicated by the semantic map (Fig. 5. Paragraph [0065]-FLEISHMAN discloses referring to FIG. 7, process 500 may include “render segmentation map from 3D semantic model” 516, and this may involve obtaining the k-pose of the current frame being analyzed, and then projecting the 3D semantic model to an image plane formed by a camera at the k-pose. An example segmentation map 700 is provided at the current pose (or k-pose) of the current frame, and where the walls 702, chairs 704, and floor 706 shown in the map 700 are segmented from each other and each have an initial, historically-based (or influenced or based on information of previous frames) semantic label (wall, chair, floor for example). Objects that are adjacent each other and have the same label may not show as separate components on the segmentation map. Please also read paragraph [0045]). FLEISHMAN fails to explicitly teach obtaining a warped image by warping the previous reconstruction image of the previous frame to the current frame image based on a motion vector map between the current frame image and the previous reconstruction image of the previous frame, wherein warping the previous reconstruction image comprises forward warping or backwards warping; and generating a reconstruction image, by using an image reconstruction machine learning model provided the obtained warped image. However, MALLAYA explicitly teaches obtaining a warped image (Fig. 3, #320 called an output frame and a flow-warped prior output. Paragraph [0063]) by warping the previous reconstruction image (Fig. 3, #308 called Flow-Warped Prior Output. Paragraph [0063]) of the previous frame (Fig. 3, #312 called a prior output or previously synthesized output frame. Paragraph [0063]) to the current frame image (Fig. 3. Paragraph [0047]-MALLAYA discloses video synthesis system 200 can generate high-level semantic inputs from images or video captured by one or more cameras 202, or otherwise obtained. System 200 can utilize these inputs with one or more neural networks to synthesize or otherwise generate photorealistic images, image sequences, or video (wherein video synthesis system includes at least an input label embedding network 302, an image encoder network 314, a flow embedding network 306, and an image generator network 318). In paragraph [0063]-MALLAYA discloses input video is captured 402 for a scene using one or more physical cameras or other such devices. This video data can be analyzed 404 to generate input labels or embeddings, as may relate to segmentation maps, depth maps, edge maps, or pose data that can be concatenated. An optical flow-modified version of a prior input frame can be generated 406 and processed by a flow embedding network. Guidance images can be generated 408 from a point cloud for this scene, where points of that point cloud can have appearance data (e.g., color, texture, or reflectivity) assigned from when those points appeared in one or more previously generated video frames. A prior frame can be provided 410 as input to an image encoder network. Input labels or embeddings, flow-modified frame data, and a guidance image can be provided 412 as input at sequential layers of an image generator network to generate an output frame, where those layers can be in a sequence with a set of upscaling layers. Please also read paragraph [0056]) based on a motion vector map between the current frame image and the previous reconstruction image of the previous frame (Fig. 3. Paragraph [0048]-MALLAYA discloses an optical flow module 216 can analyze recent video frames to determine optical flow data that can be used to warp and generate an image conditioned on recent generated images [0055]-MALLAYA discloses guidance can include use of a motion vector. In paragraph [0060]-MALLAY discloses γ.sub.flow and +β.sub.flow are generated using flow-embedding network 306 applied on an optical flow-warped previous frame. This provides additional constraints that a generated output frame 320 should be consistent even in dynamic regions (wherein a transformation using optical flow between a previous and current image is forward warping). Please also read claim 2 and paragraph [0063]). and generating a reconstruction image, by using an image reconstruction machine learning model (Fig. 3. Paragraph [0047]-MALLAYA discloses video synthesis system 200 can generate high-level semantic inputs from images or video captured by one or more cameras 202, or otherwise obtained. System 200 can utilize these inputs with one or more neural networks to synthesize or otherwise generate photorealistic images, image sequences, or video. In paragraph [0057]-MALLAYA discloses a video synthesis system includes at least four networks, or sub-networks. These networks can include an input label embedding network 302, an image encoder network 314, a flow embedding network 306, and an image generator network 318) provided the obtained warped image (Fig. 3. Paragraph [0048]-MALLAYA discloses an optical flow module 216 can analyze recent video frames to determine optical flow data that can be used to warp and generate an image conditioned on recent generated images (wherein the synthesized images/sequences/video may be based on either one, several or all previously frames/reconstructed frames). Please also read paragraph [0062-0063]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of FLEISHMAN of having an apparatus, comprising: a processor configured to: generate a semantic map indicating a first visualization property assigned to a first object within a frame image having a first resolution, wherein the frame image is a current frame image and the semantic map is a current semantic map of the current frame image; obtaining a previous reconstruction image of a previous frame having a different second resolution; and generating a reconstruction image, by using an image reconstruction machine learning model provided the obtained warped image, the frame image, and the semantic map, having the different second resolution and including a second object having a second visualization property indicated by the semantic map with the teachings of MAKKYA of having obtaining a warped image by warping the previous reconstruction image of the previous frame to the current frame image based on a motion vector map between the current frame image and the previous reconstruction image of the previous frame, wherein warping the previous reconstruction image comprises forward warping. Wherein FLEISHMAN’s apparatus having obtaining a warped image by warping the previous reconstruction image of the previous frame to the current frame image based on a motion vector map between the current frame image and the previous reconstruction image of the previous frame, wherein warping the previous reconstruction image comprises forward warping; and generating a reconstruction image, by using an image reconstruction machine learning model provided the obtained warped image. The motivation behind the modification would have been to obtain an apparatus that improves the efficiency, quality and efficacy of image segmentation and reconstruction, since both FLEISHMAN and MALLYA concern the use of neural networks for image reconstruction and semantic map generation. Wherein FLEISHMAN’s provides improves the accuracy of the semantic labels by using historical data along with the efficiency of semantic segmentation, which, in turn, permits the sematic segmentation to be performed on smaller devices, while MALLYA’s systems and methods that improve neural network training speed, memory performance and the ability to generate images or video that are consistent in appearance over time, viewer, camera, session, or other such variants. Please see FLEISHMAN et al. (US 20190043203 A1), Paragraph [0032, 0040, and 0056] and MALLYA et al. (US 20210374552 A1), Abstract and Paragraph [0002, 0172, 0335, 0340]. FLEISHMAN in view of MALLAYA fails to explicitly teach wherein warping the previous reconstruction image comprises backwards warping. However, POTTORFF explicitly teaches wherein warping the previous reconstruction image (Fig. 1. Paragraph [0091]-POTTORFF discloses FIG. 1 illustrates an example diagram 100 where blending factors for frame motion are generated using a neural network. A processor 102 executes or otherwise performs one or more instructions to use a neural network 110 to generate blending factors of frame motion. Processor 102 uses neural network 110 to generate blending factors of frame motion that are used in frame interpolation. Processor 102 uses neural network 110 to generate blending factors used in frame motion to be used to perform deep-learning based frame interpolation. Inputs to neural network 110 comprise one or more frames (e.g., a previous frame 104 and/or a current frame 106) and additional frame information including motion information of pixels of previous frame 104 and/or current frame 106. In paragraph [0100]-POTTORFF discloses previous frame 104 and/or current frame 106 are generated by spatial upsampling (e.g. by spatial super sampling such as, for example, DLSS, XeSS (or XeSS) from Intel®, FidelityFX™ Super Resolution from AMD®, etc.). Please also see Fig. 3) comprises backwards warping (Fig. 3. Paragraph [0174]-POTTORFF discloses one or more flow warped intermediate images are generated based on, for example, forward optical flow vectors, reverse optical flow vectors, or other such flow vectors. In paragraph [0176]-POTTORFF discloses one or more intermediate images (e.g., generated using blending factors at step 1012) are blended together to generate an intermediate result such as, for example, blended previous to current intermediate frame 902 or blended current to previous intermediate frame 904. Please also read paragraph [01). more flow warped intermediate images are generated using systems and methods such as those described herein. In at least one embodiment, at step 1010, one or more flow warped intermediate images are generated based on, for example, forward optical flow vectors, reverse optical flow vectors, or other such flow vectors. Please also read paragraph [0153-0163, 0643-0646 and 0651-0652]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of FLEISHMAN in view of MALLYA of having an apparatus, comprising: a processor configured to: generate a semantic map indicating a first visualization property assigned to a first object within a frame image having a first resolution, wherein the frame image is a current frame image and the semantic map is a current semantic map of the current frame image; obtaining a previous reconstruction image of a previous frame having a different second resolution; and generating a reconstruction image, by using an image reconstruction machine learning model provided the obtained warped image, the frame image, and the semantic map, having the different second resolution and including a second object having a second visualization property indicated by the semantic map with the teachings of POTTORFF of having wherein warping the previous reconstruction image comprises backwards warping. Wherein FLEISHMAN’s apparatus having wherein warping the previous reconstruction image comprises forward warping or backwards warping. The motivation behind the modification would have been to obtain an apparatus that improves the efficiency, quality and efficacy of image segmentation and reconstruction, since both FLEISHMAN and POTTORFF concern the use of neural networks for image reconstruction. Wherein FLEISHMAN’s provides improves the accuracy of the semantic labels by using historical data along with the efficiency of semantic segmentation, which, in turn, permits the sematic segmentation to be performed on smaller devices, while POTTORFF’s systems and methods improve neural network training and the efficiency, accuracy, and efficacy of image processing, image reconstruction and segmentation. Please see FLEISHMAN et al. (US 20190043203 A1), Paragraph [0032, 0040, and 0056] and POTTORFF et al. (US 20240098216 A1), Abstract and Paragraph [0124 and 0518]. Regarding claim 13, FLEISHMAN in view of MALLYA and in further view of POTTORFF explicitly teach the apparatus of claim 12, FLEISHMAN further teaches wherein the processor (Fig. 9, #920 called processors. Paragraph [0093]-FLEISHMAN discloses the image processing system 900 may have one or more processors 920. Please also read paragraph [0083]) is further configured to: obtain semantic data for the first visualization property (Fig. 4. Paragraph [0044]-FLEISHMAN discloses referring to FIG. 4, a process 400 is provided for a method and system of recurrent semantic segmentation for imaging processing. In paragraph [0047]-FLEISHMAN discloses process 400 may include “recurrently generate a semantic segmentation map in a view of a current pose of the current frame and comprising obtaining data to form the semantic segmentation map from a 3D semantic segmentation model, wherein individual semantic segmentation maps are each associated with a different current frame from the video sequence” 404. Once established, the 3D semantic segmentation model may be projected to an image plane to form a segmentation map with the semantic labels from the model that have pixels or voxels on that plane); and generate, based on an object identifier map comprising the obtained semantic data and regions classified by plural objects within the frame image, a semantic map indicating a corresponding visualization property assigned to a corresponding object of the plural objects through a region corresponding to each object (Fig. 5. Paragraph [0065]-FLEISHMAN discloses referring to FIG. 7, process 500 may include “render segmentation map from 3D semantic model” 516, and this may involve obtaining the k-pose of the current frame being analyzed, and then projecting the 3D semantic model to an image plane formed by a camera at the k-pose. An example segmentation map 700 is provided at the current pose (or k-pose) of the current frame, and where the walls 702, chairs 704, and floor 706 shown in the map 700 are segmented from each other and each have an initial, historically-based (or influenced or based on information of previous frames) semantic label (wall, chair, floor for example). Objects that are adjacent each other and have the same label may not show as separate components on the segmentation map). Regarding claim 15, FLEISHMAN in view of MALLYA and in further view of POTTORFF explicitly teach the apparatus of claim 12, FLEISHMAN explicitly teaches wherein the processor is further configured to generate the semantic map to indicate one or more of a type, pattern, material, or shape of the first object (Fig. 2A. Paragraph [0057]-FLEISHMAN discloses the segmentation output unit 822 outputs semantic labels, or class or probabilities for the labels or classes, and provides them as part of the current-historical (C-H) semantically segmented (or just segmented) frame 824). In paragraph [0065] Referring to FIG. 7, process 500 may include “render segmentation map from 3D semantic model” 516. An example segmentation map 700 is provided at the current pose (or k-pose) of the current frame, and where the walls 702, chairs 704, and floor 706 shown in the map 700 are segmented from each other and each have an initial, historically-based (or influenced or based on information of previous frames) semantic label (wall, chair, floor for example). Please also read paragraph [0068]), and wherein a value of the different second resolution is greater than a value of the first resolution (Fig. 4. Paragraph [0045]-FLEISHMAN disclosed process 400 may include “obtain a video sequence of frames of image data and comprising a current frame” 402. (wherein the current frame is reconstructed from previously reconstructed frame/semantic maps, each previous frame is a past current frame and the current frame may undergo pre-processing operations). The pre-processing could include resolution reduction. Therefore, it would have been obvious to a person of ordinary skill in the art to obtain a previous reconstruction image having a different second resolution. FLEISHMAN explicitly teaches obtaining a current reconstructed frame from a previous reconstructed image where the current frame may undergo resolution reduction, which can result in a previous/current frames having different resolutions and a previous frame having a greater resolution. As FLEISHMAN states at paragraph [0068], “pre-processing operations converts the rendered semantic map input to an expected format for input to a neural network”. Thus, it would be obvious to a person of ordinary skill to use a different resolution and/or greater resolution for the reconstructed frame because it ensures frames/models/maps are converted into the correct format, allows for reconstructed frames to be generated from different resolutions, and improves both computational efficiency and frame rendering by rescaling frames to either a higher or lower resolution. Please also read paragraph [0059 and 0068]). Regarding claim 16, FLEISHMAN in view of MALLYA and in further view of POTTORFF explicitly teach the apparatus of claim 12, FLEISHMAN further teaches wherein the processor is further configured to generate the semantic map (Fig. 3. Paragraph [0044]-FLEISHMAN discloses referring to FIG. 4, a process 400 is provided for a method and system of recurrent semantic segmentation for imaging processing. In paragraph [0047]-FLEISHMAN discloses process 400 may optionally include “recurrently generate a semantic segmentation map in a view of a current pose of the current frame and comprising obtaining data to form the semantic segmentation map from a 3D semantic segmentation model, wherein individual semantic segmentation maps are each associated with a different current frame from the video sequence” 404. Please also see Fig. 5A-B) to indicate rendering information including one or more of a color, a diffuse color, a depth, a normal line, a specular reflection, or an albedo of the frame image together with the first visualization property (Fig. 3. Paragraph [0039]-FLEISHMAN discloses referring to FIG. 3, an image 300 shows the room 202 (now 301) from image 200 except now with the disclosed example 3D semantic segmentation applied. Each voxel color or shade represents a class of an object that it belongs to. With such semantic segmentation, actions can be taken depending on the semantic label of the segment whether for computer vision or other applications such as with virtual or augmented reality for example. In paragraph [0045]-FLEISHMAN discloses process 400 may include “obtain a video sequence of frames of image data and comprising a current frame” 402. This operation may include obtaining pre-processed raw image data with RGB, YUV, or other color space values in addition to luminance values for a number of frames of a video sequence. The color and luminance values may be provided in many different additional forms such as gradients, histograms, and so forth. In paragraph [0046]-FLEISHMAN discloses this operation also may include obtaining depth data when the depth data is used for segmentation analysis. In paragraph [0060]-FLEISHMAN discloses process 500 may include “generate depth map” 504, where a depth map for the current image may be formed to establish a 3D space for the video sequence being analyzed, and eventually used to generate a 3D geometric model. Please also read paragraph [0056]). Regarding claim 17, FLEISHMAN in view of MALLYA and in further view of POTTORFF explicitly teach the apparatus of claim 12, FLEISHMAN further teaches wherein the image reconstruction machine learning model (Fig. 4. Paragraph [0037]-FLEISHMAN discloses in the present solution, the recurrent segmentation operation (usage of the semantic information from previous frames as reflected in the 3D semantic model) is learned from the data and tailored to specific scenarios. In paragraph [0044]-FLEISHMAN discloses referring to FIG. 4, a process 400 is provided for a method and system of recurrent semantic segmentation for imaging processing (wherein the operations may be performed by a neural network). Please also read paragraph [0036 and 0051]) is machine learning model trained using an objective function (Fig. 4. Paragraph [0078]-FLEISHMAN discloses the training of the architecture in a supervised-learning settings may include a training set of RGBD video-sequences, where the frames in each sequence have semantic information. Such a video can be obtained using either (i) a labor intensive method manually segmenting each frame, (ii) segmenting a reconstructed 3D model, or (iii) using synthetic data. See, Dai at el., “Richly-annotated 3D Reconstructions of Indoor Scenes”, Computer Vision and Pattern Recognition (CVPR) (2017)) calculated based on a second visualization property of a third object of a temporary output image and a third visualization property indicated by a training semantic map together with a difference between a temporary output image obtained from a training input image and the training semantic map and a true value output image (Fig. 4. Paragraph [0079]-FLEISHMAN discloses training a recurrent network requires rendered semantic maps of the 3D semantic model. The training may be performed in several operations. In paragraph [0080]-FLEISHMAN discloses the first training operation may involve initialization by training a standard semantic-segmentation network. First, a standard single frame CNN-based semantic segmentation algorithm is trained. This resulting initial network may be denoted as n.sub.1 for example. In paragraph [0081]-FLEISHMAN discloses the next training operation may involve data preparation, which refers to generating training data for the recurrent architecture. Given the current network, training data was generated for the next recurrent phase in the form of a triplet (RGBD frame, rendered semantic map of the 3D semantic model, ground truth semantic segmentation). The system runs as shown in FIGS. 6 and 8 with the current network on short sequences of N frames, where N is a tunable parameter. A matching semantic map was rendered for the last frame from the last camera pose in each sequence, and then saved with the frame as training data for next stage. The semantic map was represented as an image of H*W pixels (the size of the frame) with C (the number of classes that the system supports) channels. Since only X<C probabilities are remembered in each voxel, lower C-X probabilities are truncated to zero, and the remaining X probabilities are renormalized to be a proper distribution). Regarding claim 20, FLEISHMAN explicitly teaches a processor-implemented method (Fig. 4. Paragraph [0034]-FLEISHMAN discloses a system and method is disclosed herein that recurrently uses historical semantic data to perform semantic segmentation of a current frame and to be used to update a 3D semantic model. In paragraph [0044]-FLEISHMAN discloses referring to FIG. 4, a process 400 is provided for a method and system of recurrent semantic segmentation for imaging processing [0083]-FLEISHMAN discloses the operations of FIGS. 4, and 5A-5B may be undertaken in response to instructions provided by one or more computer program products. Such program products may include signal bearing media providing instructions that, when executed by a processor may provide the functionality described herein. Please also see Fig. 9-10 and read paragraph [0105]), the method comprising: identifying objects within a frame image (Fig. 4. Paragraph [0065]-FLEISHMAN discloses an example segmentation map 700 is provided at the current pose (or k-pose) of the current frame, and where the walls 702, chairs 704, and floor 706 shown in the map 700 are segmented from each other and each have an initial, historically-based (or influenced or based on information of previous frames) semantic label (wall, chair, floor for example). Please also read paragraph [0035]-FLEISHMAN discloses a the recurrent 3D semantic segmentation algorithm may include CNN-based architecture that receives this paired input to synergistically analyze the distribution of the image data on the input together. For example, using the whole frame enables the system to learn what classes appear together. When the system recognizes a table, it may easily recognize a chair, since it is expected to appear with a table, while the system will eliminate other objects more easily (for example, there is probably no horse in the image). The output of the system is an updated 3D representation (model) of the world with individual voxels of the model being semantically classified. Further in paragraph); generating a semantic map (Fig. 6, #622, #700 and #804 called a semantic map. Paragraph [0047-0052, 0056 and 0065]. Please also see Fig. 4-5 and 7-8) of the frame image (Fig. 4. Paragraph [0045]-FLEISHMAN discloses process 400 may include “obtain a video sequence of frames of image data and comprising a current frame” 402. Please also see Fig. 1 and 5-6), the semantic map including regions for corresponding objects, each object having a visualization property assigned thereto (Fig. 3. Paragraph [0039]-FLEISHMAN discloses referring to FIG. 3, an image 300 shows the room 202 (now 301) from image 200 except now with the disclosed example 3D semantic segmentation applied. Each voxel color or shade represents a class of an object that it belongs to. In paragraph [0065]-FLEISHMAN discloses an example segmentation map 700 is provided at the current pose (or k-pose) of the current frame, and where the walls 702, chairs 704, and floor 706 shown in the map 700 are segmented from each other and each have an initial, historically-based (or influenced or based on information of previous frames) semantic label (wall, chair, floor for example). Please also read paragraph [0035 and 0068]), wherein the frame image is a current frame image and the semantic map is a current semantic map of the current frame image (Fig. 4. Paragraph [0047]-FLEISHMAN discloses process 400 may optionally include “recurrently generate a semantic segmentation map in a view of a current pose of the current frame and comprising obtaining data to form the semantic segmentation map from a 3D semantic segmentation model, wherein individual semantic segmentation maps are each associated with a different current frame from the video sequence” 404. This may include generating a 3D semantic segmentation model, based on a 3D geometric model (generated by use of RGB-SLAM for example) with semantic labels registered to the model. Once established, the 3D semantic segmentation model may be projected to an image plane to form a segmentation map with the semantic labels from the model that have pixels or voxels on that plane. The image plane may be the plane formed by the camera pose of the current frame being analyzed. The 3D semantic segmentation model may be updated with semantic segment labels each current frame being semantically analyzed so that the 3D semantic model reflects or represents the history of the semantic segmentation of the 3D space represented by the 3D semantic model up to a current point in time. Please also see Fig. 1 and 5-6 and read paragraph [0048-0052, 0056-0059 and 0064-0069]); obtaining a previous reconstruction image of a previous frame (Fig. 4. Paragraph [0050]-FLEISHMAN discloses process 400 may include “generate a current and historical semantically segmented frame comprising using both the current semantic features and the historically-influenced semantic features as input to a neural network that indicates semantic labels for areas of the current historical semantically segmented frame” 410. Please and read paragraph [0048-0052, 0056-0059 and 0061]) having a second resolution (Fig. 4. Paragraph [0045]-FLEISHMAN disclosed process 400 may include “obtain a video sequence of frames of image data and comprising a current frame” 402 (wherein the current frame is reconstructed from previously reconstructed frame/semantic maps and the current frame may undergo pre-processing operations). The pre-processing could include resolution reduction. Therefore, it would have been obvious to a person of ordinary skill in the art to obtain a previous reconstruction image having a different second resolution. FLEISHMAN explicitly teaches obtaining a current reconstructed frame from a previous reconstructed image where the current frame may undergo resolution reduction, which can result in a previous/current frames having different resolutions and a previous frame having a greater resolution. As FLEISHMAN states at paragraph [0068], “pre-processing operations converts the rendered semantic map input to an expected format for input to a neural network”. Thus, it would be obvious to a person of ordinary skill to use a different resolution and/or greater resolution for the reconstructed frame because it ensures frames/models/maps are converted into the correct format, allows for reconstructed frames to be generated from different resolutions, and improves both computational efficiency and frame rendering by rescaling frames to either a higher or lower resolution. Please also read paragraph [0059 and 0068]). Although FLEISHMAN explicitly teaches generating, using an image reconstruction machine learning model, a reconstruction image from the frame image, and the semantic map (Fig. 4. Paragraph [0048]-FLEISHMAN disclosed process 400 may include “extract historically-influenced semantically semantic features of the semantic segmentation map” 406. In paragraph [0049]-FLIESHMAN discloses process 400 may include “extract current semantic features of the current frame” 408 (wherein the extraction of both historical and current semantic features is performed by a neural network). In paragraph [0050]-FLEISHMAN discloses process 400 may include “generate a current and historical semantically segmented frame comprising using both the current semantic features and the historically-influenced semantic features as input to a neural network that indicates semantic labels for areas of the current historical semantically segmented frame” 410 (wherein the model may be generated by concatenating current semantic features and historically-influenced semantic features and inputting them into a neural network, such as a CNN). Please also read paragraph [0045, 0052 and 0065]]), wherein the frame image has a first resolution, and wherein the reconstruction image has the second resolution which is greater than the first resolution (Fig. 4. Paragraph [0045]-FLEISHMAN disclosed process 400 may include “obtain a video sequence of frames of image data and comprising a current frame” 402 (wherein the current frame is reconstructed from previously reconstructed frame/semantic maps and the current frame may undergo pre-processing operations). The pre-processing could include resolution reduction. Therefore, it would have been obvious to a person of ordinary skill in the art to obtain a previous reconstruction image having a different second resolution. FLEISHMAN explicitly teaches obtaining a current reconstructed frame from a previous reconstructed image where the current frame may undergo resolution reduction, which can result in a previous/current frames having different resolutions and a previous frame having a greater resolution. As FLEISHMAN states at paragraph [0068], “pre-processing operations converts the rendered semantic map input to an expected format for input to a neural network”. Thus, it would be obvious to a person of ordinary skill to use a different resolution and/or greater resolution for the reconstructed frame because it ensures frames/models/maps are converted into the correct format, allows for reconstructed frames to be generated from different resolutions, and improves both computational efficiency and frame rendering by rescaling frames to either a higher or lower resolution. Please also read paragraph [0059 and 0068]). FLIESHMAN fails to explicitly teach obtaining a warped image by warping the previous reconstruction image of the previous frame to the current frame image based on a motion vector map between the current frame image and the previous reconstruction image of the previous frame, wherein warping the previous reconstruction image comprises forward warping; and generating, using an image reconstruction machine learning model, a reconstruction image from the obtained warped image. However, MALLAYA explicitly teaches obtaining a warped image (Fig. 3, #320 called an output frame and a flow-warped prior output. Paragraph [0063]) by warping the previous reconstruction image (Fig. 3, #308 called Flow-Warped Prior Output. Paragraph [0063]) of the previous frame (Fig. 3, #312 called a prior output or previously synthesized output frame. Paragraph [0063]) to the current frame image (Fig. 3. Paragraph [0047]-MALLAYA discloses video synthesis system 200 can generate high-level semantic inputs from images or video captured by one or more cameras 202, or otherwise obtained. System 200 can utilize these inputs with one or more neural networks to synthesize or otherwise generate photorealistic images, image sequences, or video (wherein video synthesis system includes at least an input label embedding network 302, an image encoder network 314, a flow embedding network 306, and an image generator network 318). In paragraph [0063]-MALLAYA discloses input video is captured 402 for a scene using one or more physical cameras or other such devices. This video data can be analyzed 404 to generate input labels or embeddings, as may relate to segmentation maps, depth maps, edge maps, or pose data that can be concatenated. An optical flow-modified version of a prior input frame can be generated 406 and processed by a flow embedding network. Guidance images can be generated 408 from a point cloud for this scene, where points of that point cloud can have appearance data (e.g., color, texture, or reflectivity) assigned from when those points appeared in one or more previously generated video frames. A prior frame can be provided 410 as input to an image encoder network. Input labels or embeddings, flow-modified frame data, and a guidance image can be provided 412 as input at sequential layers of an image generator network to generate an output frame, where those layers can be in a sequence with a set of upscaling layers. Please also read paragraph [0056]) based on a motion vector map between the current frame image and the previous reconstruction image of the previous frame, wherein warping the previous reconstruction image comprises forward warping (Fig. 3. Paragraph [0048]-MALLAYA discloses an optical flow module 216 can analyze recent video frames to determine optical flow data that can be used to warp and generate an image conditioned on recent generated images [0055]-MALLAYA discloses guidance can include use of a motion vector. In paragraph [0060]-MALLAY discloses γ.sub.flow and +β.sub.flow are generated using flow-embedding network 306 applied on an optical flow-warped previous frame. This provides additional constraints that a generated output frame 320 should be consistent even in dynamic regions (wherein a transformation using optical flow between a previous and current image is forward warping). Please also read claim 2 and paragraph [0063]); and generating, using an image reconstruction machine learning model (Fig. 3. Paragraph [0047]-MALLAYA discloses video synthesis system 200 can generate high-level semantic inputs from images or video captured by one or more cameras 202, or otherwise obtained. System 200 can utilize these inputs with one or more neural networks to synthesize or otherwise generate photorealistic images, image sequences, or video. In paragraph [0057]-MALLAYA discloses a video synthesis system includes at least four networks, or sub-networks. These networks can include an input label embedding network 302, an image encoder network 314, a flow embedding network 306, and an image generator network 318), a reconstruction image from the obtained warped image (Fig. 3. Paragraph [0048]-MALLAYA discloses an optical flow module 216 can analyze recent video frames to determine optical flow data that can be used to warp and generate an image conditioned on recent generated images (wherein the synthesized images/sequences/video may be based on either one, several or all previously frames/reconstructed frames). In paragraph [0063]-MALLYA discloses input labels or embeddings, flow-modified frame data, and a guidance image can be provided 412 as input at sequential layers of an image generator network to generate an output frame, where those layers can be in a sequence with a set of upscaling layers. Please also read paragraph Please also read paragraph [0056 and 0062]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of FLEISHMAN of having a processor-implemented method, the method comprising: identifying objects within a frame image; generating a semantic map of the frame image, the semantic map including regions for corresponding objects, each object having a visualization property assigned thereto, wherein the frame image is a current frame image and the semantic map is a current semantic map of the current frame image; obtaining a previous reconstruction image of a previous frame having a second resolution, with the teachings of MALLYA of having obtaining a warped image by warping the previous reconstruction image of the previous frame to the current frame image based on a motion vector map between the current frame image and the previous reconstruction image of the previous frame, wherein warping the previous reconstruction image comprises forward warping; and generating, using an image reconstruction machine learning model, a reconstruction image from the obtained warped image. Wherein FLEISHMAN’s method having obtaining a warped image by warping the previous reconstruction image of the previous frame to the current frame image based on a motion vector map between the current frame image and the previous reconstruction image of the previous frame, wherein warping the previous reconstruction image comprises forward warping; and generating, using an image reconstruction machine learning model, a reconstruction image from the obtained warped image, the frame image, and the semantic map, wherein the frame image has a first resolution, and wherein the reconstruction image has the second resolution which is greater than the first resolution. The motivation behind the modification would have been to obtain a method that improves the efficiency, quality and efficacy of image segmentation and reconstruction, since both FLEISHMAN and MALLYA concern the use of neural networks for image reconstruction and semantic map generation. Wherein FLEISHMAN’s provides improves the accuracy of the semantic labels by using historical data along with the efficiency of semantic segmentation, which, in turn, permits the sematic segmentation to be performed on smaller devices, while MALLYA’s systems and methods that improve neural network training speed, memory performance and the ability to generate images or video that are consistent in appearance over time, viewer, camera, session, or other such variants. Please see FLEISHMAN et al. (US 20190043203 A1), Paragraph [0032, 0040, and 0056] and MALLYA et al. (US 20210374552 A1), Abstract and Paragraph [0002, 0172, 0335, 0340]. FLEISHMAN in view of MALLAYA fails to explicitly teach wherein warping the previous reconstruction image comprises backwards warping. However, POTTORFF explicitly teaches wherein warping the previous reconstruction image (Fig. 1. Paragraph [0091]-POTTORFF discloses FIG. 1 illustrates an example diagram 100 where blending factors for frame motion are generated using a neural network. A processor 102 executes or otherwise performs one or more instructions to use a neural network 110 to generate blending factors of frame motion. Processor 102 uses neural network 110 to generate blending factors of frame motion that are used in frame interpolation. Processor 102 uses neural network 110 to generate blending factors used in frame motion to be used to perform deep-learning based frame interpolation. Inputs to neural network 110 comprise one or more frames (e.g., a previous frame 104 and/or a current frame 106) and additional frame information including motion information of pixels of previous frame 104 and/or current frame 106. In paragraph [0100]-POTTORFF discloses previous frame 104 and/or current frame 106 are generated by spatial upsampling (e.g. by spatial super sampling such as, for example, DLSS, XeSS (or XeSS) from Intel®, FidelityFX™ Super Resolution from AMD®, etc.). Please also see Fig. 3) comprises backwards warping (Fig. 3. Paragraph [0174]-POTTORFF discloses one or more flow warped intermediate images are generated based on, for example, forward optical flow vectors, reverse optical flow vectors, or other such flow vectors. In paragraph [0176]-POTTORFF discloses one or more intermediate images (e.g., generated using blending factors at step 1012) are blended together to generate an intermediate result such as, for example, blended previous to current intermediate frame 902 or blended current to previous intermediate frame 904. Please also read paragraph [01). more flow warped intermediate images are generated using systems and methods such as those described herein. In at least one embodiment, at step 1010, one or more flow warped intermediate images are generated based on, for example, forward optical flow vectors, reverse optical flow vectors, or other such flow vectors. Please also read paragraph [0153-0163, 0643-0646 and 0651-0652]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of FLEISHMAN in view of MALLYA of having a processor-implemented method, the method comprising: identifying objects within a frame image; generating a semantic map of the frame image, the semantic map including regions for corresponding objects, each object having a visualization property assigned thereto, wherein the frame image is a current frame image and the semantic map is a current semantic map of the current frame image, with the teachings of POTTORFF of having wherein warping the previous reconstruction image comprises backwards warping; Wherein FLEISHMAN’s method having wherein warping the previous reconstruction image comprises forward warping or backwards warping. The motivation behind the modification would have been to obtain a method that improves the efficiency, quality and efficacy of image segmentation and reconstruction, since both FLEISHMAN and POTTORFF concern the use of neural networks for image reconstruction. Wherein FLEISHMAN’s provides improves the accuracy of the semantic labels by using historical data along with the efficiency of semantic segmentation, which, in turn, permits the sematic segmentation to be performed on smaller devices, while POTTORFF’s systems and methods improve neural network training and the efficiency, accuracy, and efficacy of image processing, image reconstruction and segmentation. Please see FLEISHMAN et al. (US 20190043203 A1), Paragraph [0032, 0040, and 0056] and POTTORFF et al. (US 20240098216 A1), Abstract and Paragraph [0124 and 0518]. Claims 3 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over FLEISHMAN et al. (US 20190043203 A1), hereinafter referenced as FLEISHMAN in view of MALLYA et al. (US 20210374552 A1), hereinafter referenced as MALLYA in further view of POTTORFF et al. (US 20240098216 A1), hereinafter referenced as POTTORFF and in further view of BAE et al. (US 20210352307 A1), hereinafter referenced as BAE. Regarding claim 3, FLEISHMAN in view of MALLYA and in further view of POTTORFF explicitly teach the method of claim 2, FLEISHMAN in view of POTTORFF fails to explicitly teach wherein the obtaining of the semantic data comprises receiving a user input to assign, for the plural objects of the obtained frame image, a corresponding visualization property for one or more corresponding objects. However, BAE explicitly teaches wherein the obtaining of the semantic data (Fig. 7. Paragraph [0093]-BAE discloses FIG. 7 illustrates a schematic diagram illustrating an example process 700 of video processing. In FIG. 7, input region 702 of a picture is fed to stage 704, where the spatial importance of input region 702 can be determined. In paragraph [0095]-BAE discloses the object detection technique can identify a bounding region (e.g., a rectangular box) in the picture, which encloses an identified object. Based on whether input region 702 is in the bounding region, a spatial importance level can be assigned to input region 702. In paragraph [0096]-BAE discloses if the semantic segmentation technique is used at stage 704, each pixel of the picture can be labeled with a class or label (e.g., a vehicle, an individual, a building, a tree, or any classification of visual contents) of what is represented. In paragraph [0097]-BAE discloses if the instance segmentation technique is used at stage 704, each pixel of an image can be further associated with a label of an instance of objects of the same class. For a class of “individuals,” the instance segmentation technique can differentiate and associate each pixel in the class with labels of “person 1,” “person 2,” and so on (wherein the semantic segmentation technique can be used to determine spatial importance levels of different classes, and the instance segmentation technique can be applied to each class to determine spatial importance levels of different instances in the same class)) comprises receiving a user input to assign, for the plural objects of the obtained frame image, a corresponding visualization property for one or more corresponding objects (Fig. 7. Paragraph [0096]-BAE discloses different classes can be predetermined with different spatial importance levels based on how interested a viewer can be of each class. The higher the value of the spatial importance level of a class, the more interested the viewer can be of the class. For example, a class of “background” can be associated with a spatial importance level of 0, a class of “buildings” can be associated with a spatial importance level of 1, a class of “vehicle” can be associated with a spatial importance level of 2, a class of “individuals” can be associated with a spatial importance level of 3, or the like. In paragraph [0098]-BAE discloses the associations between the classes (or objects) and spatial importance levels can be assigned by a user before performing stage 704. Further in paragraph [0103]-BAE discloses scenes of fast actions (e.g., fighting scenes), close-up shots, or stunning visual effects can have higher temporal importance levels than other scenes. The associations between the pictures and the temporal importance levels can be assigned by a user before performing stage 804). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of FLEISHMAN in view of MALLYA and in further view of POTTORFF of having processor-implemented method, the method comprising: identifying objects within a frame image; generating a semantic map of the frame image, the semantic map including regions for corresponding objects, each object having a visualization property assigned thereto, wherein the frame image is a current frame image and the semantic map is a current semantic map of the current frame image, with the teachings of BAE of having wherein the obtaining of the semantic data comprises receiving a user input to assign, for the plural objects of the obtained frame image, a corresponding visualization property for one or more corresponding objects. Wherein FLEISHMAN’s method having wherein the obtaining of the semantic data comprises receiving a user input to assign, for the plural objects of the obtained frame image, a corresponding visualization property for one or more corresponding objects. The motivation behind the modification would have been to obtain a method that improves the efficiency of semantic segmentation as well as the image quality for important regions, since both FLEISHMAN and BAE concern semantic segmentation and image analysis. Wherein FLEISHMAN’s provides improves the accuracy of the semantic labels by using historical data along with the efficiency of semantic segmentation, which, in turn, permits the sematic segmentation to be performed on smaller devices, while BAE’s systems and methods greatly improve the image quality for the more important portions after upscaling while also not greatly increasing overall computational costs for resolution enhancement and transcoding. Please see FLEISHMAN et al. (US 20190043203 A1), Paragraph [0032, 0040, and 0056] and BAE et al. (US 20210352307 A1), Abstract and Paragraph [0040]. Regarding claim 14, FLEISHMAN in view of MALLYA and in further view of POTTORFF explicitly teach the apparatus of claim 13, FLEISHMAN in view of POTTORFF fails to explicitly teach wherein the processor is further configured to obtain the semantic data by receiving, for each object of the frame image, an input visualization property based on a user input as the corresponding visualization property of the corresponding object. However, BAE explicitly teaches wherein the processor is further configured to obtain the semantic data by receiving, for each object of the frame image (Fig. 7. Paragraph [0093]-BAE discloses FIG. 7 illustrates a schematic diagram illustrating an example process 700 of video processing. In FIG. 7, input region 702 of a picture is fed to stage 704, where the spatial importance of input region 702 can be determined. In paragraph [0095]-BAE discloses the object detection technique can identify a bounding region (e.g., a rectangular box) in the picture, which encloses an identified object. Based on whether input region 702 is in the bounding region, a spatial importance level can be assigned to input region 702. In paragraph [0096]-BAE discloses if the semantic segmentation technique is used at stage 704, each pixel of the picture can be labeled with a class or label (e.g., a vehicle, an individual, a building, a tree, or any classification of visual contents) of what is represented. In paragraph [0097]-BAE discloses if the instance segmentation technique is used at stage 704, each pixel of an image can be further associated with a label of an instance of objects of the same class. For a class of “individuals,” the instance segmentation technique can differentiate and associate each pixel in the class with labels of “person 1,” “person 2,” and so on (wherein the semantic segmentation technique can be used to determine spatial importance levels of different classes, and the instance segmentation technique can be applied to each class to determine spatial importance levels of different instances in the same class)), an input visualization property based on a user input as the corresponding visualization property of the corresponding object (Fig. 7. Paragraph [0096]-BAE discloses different classes can be predetermined with different spatial importance levels based on how interested a viewer can be of each class. The higher the value of the spatial importance level of a class, the more interested the viewer can be of the class. For example, a class of “background” can be associated with a spatial importance level of 0, a class of “buildings” can be associated with a spatial importance level of 1, a class of “vehicle” can be associated with a spatial importance level of 2, a class of “individuals” can be associated with a spatial importance level of 3, or the like. In paragraph [0098]-BAE discloses the associations between the classes (or objects) and spatial importance levels can be assigned by a user before performing stage 704. Further in paragraph [0103]-BAE discloses scenes of fast actions (e.g., fighting scenes), close-up shots, or stunning visual effects can have higher temporal importance levels than other scenes. The associations between the pictures and the temporal importance levels can be assigned by a user before performing stage 804). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of FLEISHMAN in view of MALLYA and in further view of POTTORFF of having an apparatus, comprising: a processor configured to: generate a semantic map indicating a first visualization property assigned to a first object within a frame image having a first resolution, wherein the frame image is a current frame image and the semantic map is a current semantic map of the current frame image; obtaining a previous reconstruction image of a previous frame having a different second resolution; and generating a reconstruction image, by using an image reconstruction machine learning model provided the obtained warped image, the frame image, and the semantic map, having the different second resolution and including a second object having a second visualization property indicated by the semantic map with the teachings of BAE of having wherein the processor is further configured to obtain the semantic data by receiving, for each object of the frame image, an input visualization property based on a user input as the corresponding visualization property of the corresponding object. Wherein FLEISHMAN’s apparatus having wherein the processor is further configured to obtain the semantic data by receiving, for each object of the frame image, an input visualization property based on a user input as the corresponding visualization property of the corresponding object. The motivation behind the modification would have been to obtain an apparatus that improves the efficiency of semantic segmentation as well as the image quality for important regions, since both FLEISHMAN and BAE concern semantic segmentation and image analysis. Wherein FLEISHMAN’s provides improves the accuracy of the semantic labels by using historical data along with the efficiency of semantic segmentation, which, in turn, permits the sematic segmentation to be performed on smaller devices, while BAE’s systems and methods greatly improve the image quality for the more important portions after upscaling while also not greatly increasing overall computational costs for resolution enhancement and transcoding. Please see FLEISHMAN et al. (US 20190043203 A1), Paragraph [0032, 0040, and 0056] and BAE et al. (US 20210352307 A1), Abstract and Paragraph [0040]. Claims 10 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over FLEISHMAN et al. (US 20190043203 A1), hereinafter referenced as FLEISHMAN in view of MALLYA et al. (US 20210374552 A1), hereinafter referenced as MALLYA in further view of POTTORFF et al. (US 20240098216 A1), hereinafter referenced as POTTORFF and in further view of TOVEY et al. (US 20230298133 A1), hereinafter referenced as TOVEY. Regarding claim 10, FLEISHMAN in view of MALLYA and in further view of POTTORFF explicitly teach the method of claim 1, FLEISHMAN fails to explicitly teach wherein the reconstructing of the current frame image into the reconstruction image of the current frame comprises: generating a disocclusion map indicating whether a corresponding object is in a previous frame image through a region corresponding to each object of the current frame image; and reconstructing the current frame image into the reconstruction image of the current frame based on a previous frame image being masked based on the generated disocclusion map. However, TOVEY explicitly teaches wherein the reconstructing of the current frame image into the reconstruction image of the current frame (Fig. 1. Paragraph [0024]-TOVEY discloses FIG. 1 illustrates an example device 100 in which one or more features described herein, such as a super resolution upscaler 332 (FIG. 3), can be implemented (wherein device 100 includes, for example, a gaming device). In paragraph [0037]-TOVEY discloses the Accelerated Processing Device 116 is configured to implement features of the present disclosure by executing a plurality of functions. The APD 116 is configured to implement a super resolution upscaler 332 that receives a low-resolution rendered frame 502 of video stream. The super resolution upscaler 332 spatially upscales the low-resolution rendered frame 502 by using temporal feedback (e.g., a previously upscaled frame(s) of the video stream) to reconstruct a high-resolution frame 508 representing the rendered frame. In paragraph [0062]-TOVEY discloses the upscaler 332 uses the current frame input and the previous frame input to generate a super resolution upscaled (and anti-aliased) frame 508 (also referred to herein as “upscaled frame 508” for brevity), which corresponds to the rendered frame 502, at the target presentation resolution) comprises: generating a disocclusion map indicating whether a corresponding object is in a previous frame image through a region corresponding to each object of the current frame image (Fig. 6. Paragraph [0074]-TOVEY discloses the depth clip component 518 processes this input to produce a disocclusion mask/map 618 indicating disoccluded areas of the current rendered frame 502. As the camera moves from an initial position (previous frame) to a new position (current frame), a pixel that was initially occluded from the viewpoint of the camera's previous position can become visible (disoccluded) from the viewpoint of the camera's current position. The disocclusion mask 618 is a texture including a value indicating how much a corresponding pixel of the current frame 502 has been disoccluded. A value of 0 indicates that the pixel was entirely occluded in the previous frame and is now disoccluded, and a value of 1 indicates the pixel was fully visible in the previous frame and is fully visible in the current frame 502. Values between 0 and 1 indicate that the pixel was visible in the previous frame to an extent proportional to the value. Please also read paragraph [0071 and 0076]); and reconstructing the current frame image into the reconstruction image of the current frame based on a previous frame image being masked based on the generated disocclusion map (Fig. 6. Paragraph [0082]-TOVEY discloses in the reproject and accumulate stage 611, the reproject and accumulate component 522 takes as input the disocclusion mask 618, the dilated motion vector buffer 614, the reactivity mask 602, the output buffer 506-1 of the previous frame. The reproject and accumulate component 522 processes this input to generate an output buffer (texture) 506-2 for the current frame 502 at the target presentation resolution/size and to also generate reprojected pixel locks (texture) 622 from the previous frame that are mappable to the current frame 502). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of FLEISHMAN in view of MALLYA and in further view of POTTORFF of having a processor-implemented method, the method comprising: identifying objects within a frame image; generating a semantic map of the frame image, the semantic map including regions for corresponding objects, each object having a visualization property assigned thereto, wherein the frame image is a current frame image and the semantic map is a current semantic map of the current frame image, with the teachings of TOVEY of having wherein the reconstructing of the current frame image into the reconstruction image of the current frame comprises: generating a disocclusion map indicating whether a corresponding object is in a previous frame image through a region corresponding to each object of the current frame image; and reconstructing the current frame image into the reconstruction image of the current frame based on a previous frame image being masked based on the generated disocclusion map. Wherein FLEISHMAN’s method having wherein the reconstructing of the current frame image into the reconstruction image of the current frame comprises: generating a disocclusion map indicating whether a corresponding object is in a previous frame image through a region corresponding to each object of the current frame image; and reconstructing the current frame image into the reconstruction image of the current frame based on a previous frame image being masked based on the generated disocclusion map. The motivation behind the modification would have been to obtain a method that improves semantic segmentation and upscaling of images, since both FLEISHMAN and TOVEY concern image analysis and image reconstruction. Wherein FLEISHMAN’s provides improves the accuracy of the semantic labels by using historical data along with the efficiency of semantic segmentation, which, in turn, permits the sematic segmentation to be performed on smaller devices, while LIU’s systems and methods improve upscaling by using temporal feedback to reconstruct high-resolution images while maintaining and improving image quality compared to native rendering. Please see FLEISHMAN et al. (US 20190043203 A1), Paragraph [0032, 0040, and 0056] and TOVEY et al. (US 20230298133 A1), Abstract and Paragraph [0023]. Regarding claim 19, FLEISHMAN in view of MALLYA and in further view of POTTORFF explicitly teach the apparatus of claim 12, although FLEISHMAN explicitly teaches the image reconstruction machine learning model (Fig. 4. Paragraph [0036]-FLIESHMAN discloses in the method and system disclosed herein, the recurrent 3D semantic segmentation algorithm may merge efficient geometric segmentation with the high performance 3D semantic segmentation. Using both dense SLAM (simultaneous localization and mapping) based on RGB-D data (data from RGB and depth cameras, e.g., Intel RealSense depth sensors) and the semantic segmentation with convolutional neural networks (CNN) in a recurrent way as described herein. Thus, the 3D semantic segmentation may include: (i) dense RGBD-SLAM for 3D reconstruction of geometry; (ii) CNN-based recurrent segmentation which receives as an input the current frame and 3D semantic information from previous frames; and (iii) a copy of the results of (ii) to the 3D semantic model. It can be stated that operation (ii) uses the past frames and performs both segmentation and update of the model). FLEISHMAN fails to explicitly teach wherein the processor is further configured to: generate a disocclusion map indicating whether a corresponding object is in a previous frame image through a region corresponding to each object of the current frame image, wherein an input provided to the image reconstruction machine learning model is further based on a masking of the previous frame image based on the generated disocclusion map. However, TOVEY explicitly teaches wherein the processor is further configured to: generate a disocclusion map indicating whether a corresponding object is in a previous frame image through a region corresponding to each object of the current frame image (Fig. 6. Paragraph [0074]-TOVEY discloses the depth clip component 518 processes this input to produce a disocclusion mask/map 618 indicating disoccluded areas of the current rendered frame 502. As the camera moves from an initial position (previous frame) to a new position (current frame), a pixel that was initially occluded from the viewpoint of the camera's previous position can become visible (disoccluded) from the viewpoint of the camera's current position. The disocclusion mask 618 is a texture including a value indicating how much a corresponding pixel of the current frame 502 has been disoccluded. A value of 0 indicates that the pixel was entirely occluded in the previous frame and is now disoccluded, and a value of 1 indicates the pixel was fully visible in the previous frame and is fully visible in the current frame 502. Values between 0 and 1 indicate that the pixel was visible in the previous frame to an extent proportional to the value. Please also read paragraph [0071 and 0076]), wherein an input provided to the image reconstruction machine learning model is further based on a masking of the previous frame image based on the generated disocclusion map (Fig. 6. Paragraph [0082]-TOVEY discloses in the reproject and accumulate stage 611, the reproject and accumulate component 522 takes as input the disocclusion mask 618, the dilated motion vector buffer 614, the reactivity mask 602, the output buffer 506-1 of the previous frame. The reproject and accumulate component 522 processes this input to generate an output buffer (texture) 506-2 for the current frame 502 at the target presentation resolution/size and to also generate reprojected pixel locks (texture) 622 from the previous frame that are mappable to the current frame 502). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of FLEISHMAN in view of MALLYA and in further view of POTTORFF of having an apparatus, comprising: a processor configured to: generate a semantic map indicating a first visualization property assigned to a first object within a frame image having a first resolution, wherein the frame image is a current frame image and the semantic map is a current semantic map of the current frame image; obtaining a previous reconstruction image of a previous frame having a different second resolution; and generating a reconstruction image, by using an image reconstruction machine learning model provided the obtained warped image, the frame image, and the semantic map, having the different second resolution and including a second object having a second visualization property indicated by the semantic map, with the teachings of TOVEY of having wherein the processor is further configured to: generate a disocclusion map indicating whether a corresponding object is in a previous frame image through a region corresponding to each object of the current frame image, wherein an input provided to the image reconstruction machine learning model is further based on a masking of the previous frame image based on the generated disocclusion map. Wherein FLEISHMAN’s apparatus having wherein the processor is further configured to: generate a disocclusion map indicating whether a corresponding object is in a previous frame image through a region corresponding to each object of the current frame image, wherein an input provided to the image reconstruction machine learning model is further based on a masking of the previous frame image based on the generated disocclusion map. The motivation behind the modification would have been to obtain an apparatus that improves semantic segmentation and upscaling of images, since both FLEISHMAN and TOVEY concern image analysis and image reconstruction. Wherein FLEISHMAN’s provides improves the accuracy of the semantic labels by using historical data along with the efficiency of semantic segmentation, which, in turn, permits the sematic segmentation to be performed on smaller devices, while LIU’s systems and methods improve upscaling by using temporal feedback to reconstruct high-resolution images while maintaining and improving image quality compared to native rendering. Please see FLEISHMAN et al. (US 20190043203 A1), Paragraph [0032, 0040, and 0056] and TOVEY et al. (US 20230298133 A1), Abstract and Paragraph [0023]. Conclusion Listed below are the prior arts made of record and not relied upon but are considered pertinent to applicant`s disclosure. SUN et al. (US 20200084427 A1)- Scene flow represents the three-dimensional (3D) structure and movement of objects in a video sequence in three dimensions from frame-to-frame and is used to track objects and estimate speeds for autonomous driving applications. Scene flow is recovered by a neural network system from a video sequence captured from at least two viewpoints (e.g., cameras), such as a left-eye and right-eye of a viewer. An encoder portion of the system extracts features from frames of the video sequence. The features are input to a first decoder to predict optical flow and a second decoder to predict disparity. The optical flow represents pixel movement in (x,y) and the disparity represents pixel movement in z (depth). When combined, the optical flow and disparity represent the scene flow...................... Please see Fig. 1-2 and 5 and para. [0048-0053]. Abstract. ZHOU et al. (US 20100231593 A1)-The present invention relates to methods and systems for the exhibition of a motion picture with enhanced perceived resolution and visual quality. The enhancement of perceived resolution is achieved both spatially and temporally. Spatial resolution enhancement creates image details using both temporal-based methods and learning-based methods. Temporal resolution enhancement creates synthesized new image frames that enable a motion picture to be displayed at a higher frame rate. The digitally enhanced motion picture is to be exhibited using a projection system or a display device that supports a higher frame rate and/or a higher display resolution than what is required for the original motion picture..…....................... Please see Fig. 2-6 and para. [0036-0043]. Abstract. LIU et al. (US 20220092795 A1)- Methods, systems, and storage media are described for motion estimation in video frame interpolation. Disclosed embodiments use feature pyramids as image representations for motion estimation and seamlessly integrates them into a deep neural network for frame interpolation. A feature pyramid is extracted for each of two input frames. These feature pyramids are wrapped together with the input frames to the target temporal position according to the inter-frame motion estimated via optical flow. A frame synthesis network is used to predict interpolation results from the pre-warped feature pyramids and input frames. The feature pyramid extractor and the frame synthesis network are jointly trained for the task of frame interpolation. An extensive quantitative and qualitative evaluation demonstrates that the described embodiments utilizing feature pyramids enables robust, high-quality video frame interpolation. Other embodiments may be described and/or claimed..…....................... Please see Fig. 2. Abstract. TRAN et al. (US 20230344962 A1)- A method includes receiving an input video stream and providing, to a convolutional neural network (CNN), multiple image frames of the video stream including a target pair of consecutive frames, a frame immediately preceding the target pair, and a frame immediately following the target pair. The method includes generating, by the CNN, multiple interpolated image frames by performing 3D space-time convolution on the multiple image frames and outputting a video stream in which the interpolated image frames are inserted between the frames of the target pair. The convolution may include passing a 3D filter over the multiple image frames in common width and height dimensions, and in a depth dimension representing the number of frames. Generating the interpolated image frames may include generating image data for multiple color channels in respective convolutional layers. The CNN may be trained to predict non-linear movements that occur over multiple image frames…....................... Please see Fig. 5-6 and para. [0043]. Abstract. Weinzaepfel (US 20200160065 A1)- A method for training a convolutional recurrent neural network for semantic segmentation in videos, includes (a) training, using a set of semantically segmented training images, a first convolutional neural network;(b) training, using a set of semantically segmented training videos, a convolutional recurrent neural network, corresponding to the first convolutional neural network, wherein a convolutional layer has been replaced by a recurrent module having a hidden state. The training of the convolutional recurrent neural network, for each pair of successive frames (t−1, t ∈ custom-character1; Tcustom-character.sup.2) of a video of the set of semantically segmented training videos includes warping an internal state of a recurrent layer according to an estimated optical flow between the frames of the pair of successive frames, so as to adapt the internal state to the motion of pixels between the frames of the pair and learning parameters of at least the recurrent module....................... Please see Fig. 4-6. Abstract. Palmaro et al. (US 20210327112 A1)- A method of populating a digital environment with digital content is disclosed. Environment data describing the digital environment is accessed. Populator data describing a populator digital object is accessed. The populator data includes semantic data describing the populator digital object. The populator digital object is placed within the digital environment. A semantic map representation of the populator digital object is generated. The semantic map representation is divided into a plurality of cells. A target cell of the plurality of cells is selected as a placeholder in the digital environment for a digital object that is optionally subsequently instantiated. The selecting of the target cell is based on an analysis of the environment data, the populator data, and the semantic map representation. Placeholder data is recorded in the semantic map representation. The placeholder data includes properties corresponding to the digital object that is optionally subsequently instantiated........................... Please see Fig. 3-4. Abstract. KULKARNI et al. (US 20240005587 A1)- Systems and methods for machine learning based controllable animation of still images is provided. In one embodiment, a still image including a fluid element is obtained. Using a flow refinement machine learning model, a refined dense optical flow is generated for the still image based on a selection mask that includes the fluid element and a dense optical flow generated from a motion hint that indicates a direction of animation. The refined dense optical flow indicates a pattern of apparent motion for the at least one fluid element. Thereafter, a plurality of video frames is generated by projecting a plurality of pixels of the still image using the refined dense optical flow..................... Please see Fig. 5-8. Abstract. GOLINSKI et al. (US 20210281867 A1)- Techniques are described herein for coding video content using recurrent-based machine learning tools. A device can include a neural network system including encoder and decoder portions. The encoder portion can generate output data for the current time step of operation of the neural network system based on an input video frame for a current time step of operation of the neural network system, reconstructed motion estimation data from a previous time step of operation, reconstructed residual data from the previous time step of operation, and recurrent state data from at least one recurrent layer of a decoder portion of the neural network system from the previous time step of operation. A decoder portion of the neural network system can generate, based on the output data and recurrent state data from the previous time step of operation, a reconstructed video frame for the current time step of operation.......................... Please see Fig. 3-5. Abstract. WANG et al. (US 20220012536 A1)- A method, computer readable medium, and system are disclosed for creating an image utilizing a map representing different classes of specific pixels within a scene. One or more computing systems use the map to create a preliminary image. This preliminary image is then compared to an original image that was used to create the map. A determination is made whether the preliminary image matches the original image, and results of the determination are used to adjust the computing systems that created the preliminary image, which improves a performance of such computing systems. The adjusted computing systems are then used to create images based on different input maps representing various object classes of specific pixels within a scene...................... Please see Fig. 1 and 5-6. Abstract. KANAMORI et al. (US 20100290713 A1)- According to the present invention, a polarized image is captured, a variation in its light intensity is approximated with a sinusoidal function, and then the object is spatially divided into a specular reflection area (S-area) and a diffuse reflection area (D-area) in Step S402 of dividing a reflection area. Information about the object's refractive index is entered in Step S405, thereby obtaining surface normals by mutually different techniques in Steps S406 and S407, respectively. Finally, in Steps S410 and S411, the two normals are matched to each other in the vicinity of the boundary between the S- and D-areas......................... Please see Fig. 5. Abstract. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Aaron Bonansinga whose telephone number is (703) 756-5380 The examiner can normally be reached on Monday-Friday, 9:00 a.m. - 6:00 p.m. ET. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached by phone at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is (571) 273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /AARON TIMOTHY BONANSINGA/Examiner, Art Unit 2673 /CHINEYERE WILLS-BURNS/Supervisory Patent Examiner, Art Unit 2673
Read full office action

Prosecution Timeline

Show 6 earlier events
Jan 12, 2026
Final Rejection mailed — §103
Feb 01, 2026
Interview Requested
Feb 09, 2026
Applicant Interview (Telephonic)
Feb 10, 2026
Examiner Interview Summary
Apr 09, 2026
Response after Non-Final Action
May 12, 2026
Request for Continued Examination
May 14, 2026
Response after Non-Final Action
Jul 27, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705305
SYSTEM AND METHOD FOR MULTI ATTRIBUTE BASED DATA SYNTHESIS
3y 5m to grant Granted Aug 11, 2026
Patent 12700248
ASSESSING HETEROGENEITY OF FEATURES IN DIGITAL PATHOLOGY IMAGES USING MACHINE LEARNING TECHNIQUES
3y 6m to grant Granted Aug 04, 2026
Patent 12694641
ANALYSIS SYSTEM AND PRODUCTION METHOD OF ANALYSIS IMAGE
3y 7m to grant Granted Jul 28, 2026
Patent 12670598
ANATOMICALLY-INFORMED DEEP LEARNING ON CONTRAST-ENHANCED CARDIAC MRI FOR SCAR SEGMENTATION AND CLINICAL FEATURE EXTRACTION
3y 2m to grant Granted Jun 30, 2026
Patent 12642493
Context-aware volumetric style transfer for estimating single volume surrogates of lung function
3y 9m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
77%
Grant Probability
99%
With Interview (+34.8%)
3y 0m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 35 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month