Prosecution Insights
Last updated: October 02, 2026
Application No. 18/192,471

METHOD FOR SWITCHING ATLAS ACCORDING TO USER'S WATCHING POINT AND DEVICE THEREFOR

Non-Final OA §103
Filed
Mar 29, 2023
Priority
Mar 29, 2022 — RE 10-2022-0038560 +1 more
Examiner
CLOTHIER, MATTHEW MORRIS
Art Unit
2614
Tech Center
2600 — Communications
Assignee
Electronics and Telecommunications Research Institute
OA Round
3 (Non-Final)
83%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
5 granted / 6 resolved
+21.3% vs TC avg
Strong +20% interview lift
Without
With
+20.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 3m
Avg Prosecution
19 currently pending
Career history
37
Total Applications
across all art units

Statute-Specific Performance

§101
5.0%
-35.0% vs TC avg
§103
76.3%
+36.3% vs TC avg
§102
14.4%
-25.6% vs TC avg
§112
3.6%
-36.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 6 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 1. A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 2/4/2026 has been entered. Information Disclosure Statement 2. The information disclosure statement (IDS) submitted on 2/11/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement has been considered by the examiner. Response to Amendment 3. This action is in response to the amendment filed on 2/4/2026. Claims 1, 5, 8, 10, 14-15, and 17 have been amended. Claims 4 and 13 have been cancelled. Claims 1, 5-6, 8-10, 14-15 and 17-18 are pending. Claims 1, 5-6, 8-10, 14-15 and 17-18 remain rejected in the application. Applicant’s amendments to the claims have overcome each and every objection previously set forth in the Non-Final Office Action mailed 11/5/2025. Response to Arguments 4. Applicant’s arguments with respect to independent claims 1 and 10 and dependent claims 4, 8 13, and 17, filed on 2/4/2026, with respect to the rejection under 35 U.S.C. 103 regarding that the prior art does not teach the limitation(s): “wherein in response to the request for the transmission of the first sub-bitstream including the first atlas, a second sub-bitstream including a second atlas is also received along with the first sub-bitstream, the second atlas not being used to render the viewport image, wherein the first atlas comprises data on views included in a first group and the second atlas comprises data on views included in a second group, wherein in response to a switch request triggered by a movement of the watching position that remains close to a view belonging to the first group while moving away from views belonging to the second group, a third sub-bitstream including a third atlas, which comprises data on views included in a third group, is received along with the first sub-bitstream, and reception of the second sub-bitstream is stopped, and wherein the second group comprises a view which is spatially neighboring to a first view included in the first group and the second group comprises a view which is spatially neighboring to a second view included in the first group.” have been fully considered, but are moot because of new grounds for rejection. The claims are now disclosed by Yun and Kroon. 5. Regarding arguments to claims 5-6, 9, 14-15, and 18, they are dependent on independent claims 1 and 10 respectively. Applicant does not argue anything other than independent claims 1 and 10 and dependent claims 4, 8, 13 and 17. Claim Objections 6. Claims 8, 10, 14-15, and 17 are objected to because of the following informalities: In claim 8, line 4, "is received periodically" should read "is received periodically." In claim 18, line 2, "the bitstream reception unit " should read "the transceiver" Appropriate correction is required. Claim Rejections - 35 USC § 103 7. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 8. Claims 1, 5, 10, and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Yun et al. (US-2021/0067757-A1, hereinafter "Yun") in view of Kroon et al. (US-2024/0107110-A1, hereinafter "Kroon"). 9. As per claim 1, Yun discloses: A method of [[switching]] an atlas according to a watching position comprising: (Yun, [0089], “For example, when a user's movement information is input into an immersive video output apparatus, the metadata processor 230 may determine an atlas necessary for video synthesis ...”) obtaining information on a view required to render a viewport image corresponding to the watching position; (Yun, [0089], “For example, when a user's movement information is input into an immersive video output apparatus, the metadata processor 230 may determine an atlas necessary for video synthesis ...” and [0011], “In an immersive video processing method according to the present disclosure, the metadata may include information indicating a type of the atlas, and the type may indicate whether or not the atlas is used for view rendering.”) determining a first atlas which comprises data on the view required to render the viewport image; (Yun, [0014], “An immersive video synthesizing method according to the present disclosure may include: parsing video data and metadata from a bitstream; obtaining at least one atlas by decoding the video data; and extracting patches required for viewport video synthesis according to a user movement from the atlas based on metadata.” and [0089], “For example, when a user's movement information is input into an immersive video output apparatus, the metadata processor 230 may determine an atlas necessary for video synthesis ...”) requesting a transmission of a first [[sub-]]bitstream including the first atlas; (Yun, [0082], “The bitstream generator 150 generates a bitstream based on encoded video data and metadata. A bitstream thus generated may be transmitted to an immersive video output apparatus.” and [0085], “The bitstream parsing unit 210 parses video data and metadata from a bitstream. Video data may include data of an encoded atlas.”) decoding the first atlas from the first [[sub-]]bitstream; and (Yun, “[0096] An immersive video output apparatus may extract atlas data by parsing a bitstream, which is received from an immersive video processing apparatus, and decode an atlas based on extracted data (S511).” and [0105], “An immersive video output apparatus may decode a plurality of atlases and synthesize a video corresponding to a viewport.”) rendering the viewport image based on the first atlas, (Yun, [0194], “In order to render a viewport video corresponding to a user's movement based on a patch, internal and external parameter information of a camera capturing a source video, from which the patch is extracted, is required. For example, information on at least one of a focal distance of a camera, a principal point, a position, and a pose may be required for viewport video rendering.” and [0230], “Generally, an atlas is constructed by using a basic video and a patch video from which redundant data have been removed through a pruning process. For viewport video rendering according to a user's position, it is necessary to extract an atlas including source videos that are adjacent to the user's viewpoint.”) wherein in response to the request for the transmission of the first [[sub-]]bitstream including the first atlas, [[a second sub-bitstream including]] a second atlas is also received along with the first [[sub-]]bitstream, [[the second atlas not being used to render the viewport image,]] (Yun, [0105], “An immersive video output apparatus may decode a plurality of atlases and synthesize a video corresponding to a viewport.” and [0085], “The bitstream parsing unit 210 parses video data and metadata from a bitstream. Video data may include data of an encoded atlas.” and [0082], “The bitstream generator 150 generates a bitstream based on encoded video data and metadata. A bitstream thus generated may be transmitted to an immersive video output apparatus.” and [0070], “A complete video patch and a segment video patch may be set to be allocated in different atlases. Alternatively, an atlas may be constructed by mixing a complete video patch and a segment video patch.” and [0062], “A patch video includes a valid region and/or an invalid region. A valid region means a region excluding an overlapping region between an additional video and a reference video. In other words, a valid region is a region including data that are included in an additional video but not in a reference video.”; Examiner’s note: As described in [0070] of Yun, a complete video patch and a segment video patch could either be allocated together in a single atlas or divided into two separate atlases. Thus, a scenario where a single atlas consisting of a complete video patch and a segment video patch is used for view synthesis could potentially also be accomplished with a first atlas with a complete video patch and a second atlas with a segment video patch.) wherein the first atlas comprises data on views included in a first group and the second atlas comprises data on views included in a second group, (Yun, [0204], “An immersive video processing apparatus may classify source videos into a plurality of groups and independently generate an atlas video of each group. For example, after 24 source videos are classified into three groups including 8 source videos each, an atlas video may be independently generated in each group. In other words, selection of a basic video and pruning may be independently performed for each group.” and [0041], “An immersive video means a video that enables a viewport to dynamically change when a viewing position of a user changes. A plurality of input videos is required to realize an immersive video. Each of the plurality of input videos maybe referred to as a source video or a source view.” and [0014], “An immersive video synthesizing method according to the present disclosure may include: parsing video data and metadata from a bitstream; obtaining at least one atlas by decoding the video data; and extracting patches required for viewport video synthesis according to a user movement from the atlas based on metadata.”) wherein [[in response to a switch request triggered by a]] movement of the watching position that remains close to a view belonging to the first group (Yun, [0230], “For viewport video rendering according to a user's position, it is necessary to extract an atlas including source videos that are adjacent to the user's viewpoint.” and [0097], “In addition, when a user's movement occurs, an atlas required for synthesizing a viewport video according to the user's movement may be determined based on metadata, and patches included in the atlas may be extracted (S512).” and [0041], “An immersive video means a video that enables a viewport to dynamically change when a viewing position of a user changes. A plurality of input videos is required to realize an immersive video. Each of the plurality of input videos maybe referred to as a source video or a source view.”) [[while moving away from]] views belonging to the second group, [[a third sub-bitstream including a third atlas,]] which comprises data on views included in a third group, [[is received along with the first sub-bitstream, and reception of the second sub-bitstream is stopped,]] and (Yun, [0230]-[0231]. “Generally, an atlas is constructed by using a basic video and a patch video from which redundant data have been removed through a pruning process. For viewport video rendering according to a user's position, it is necessary to extract an atlas including source videos that are adjacent to the user's viewpoint. In other words, for transmission of an immersive video and random access, information on a source video and a basic video included in an atlas needs to be delivered. For example, when 16 source videos are divided into two groups including 8 source videos respectively, an atlas for a first group may be generated through a pruning process for the source videos 1 to 8, and an atlas for a second group may be generated through a pruning process for the source videos 9 to 16. When only random/spatial access or an atlas based on user location is to be used for generating a viewport video, source information of patches within an atlas and information indicating a basic video, which is to be decoded first within an atlas, are required.” and [0041], “An immersive video means a video that enables a viewport to dynamically change when a viewing position of a user changes. A plurality of input videos is required to realize an immersive video. Each of the plurality of input videos maybe referred to as a source video or a source view.” and [0090], “The video synthesizer 240 may dynamically synthesize a viewport video according to a user's movement. Specifically, the video synthesizer 240 may extract patches necessary to synthesize a viewport video from an atlas by using information that is determined according to a user' s movement in the metadata processor 230.” and [0204], “An immersive video processing apparatus may classify source videos into a plurality of groups and independently generate an atlas video of each group. For example, after 24 source videos are classified into three groups including 8 source videos each, an atlas video may be independently generated in each group. In other words, selection of a basic video and pruning may be independently performed for each group.”) wherein the second group comprises a view which is spatially neighboring to a first view included in the first group (Yun, [0230]-[0231], “Generally, an atlas is constructed by using a basic video and a patch video from which redundant data have been removed through a pruning process. For viewport video rendering according to a user's position, it is necessary to extract an atlas including source videos that are adjacent to the user's viewpoint. In other words, for transmission of an immersive video and random access, information on a source video and a basic video included in an atlas needs to be delivered. For example, when 16 source videos are divided into two groups including 8 source videos respectively, an atlas for a first group may be generated through a pruning process for the source videos 1 to 8, and an atlas for a second group may be generated through a pruning process for the source videos 9 to 16. When only random/spatial access or an atlas based on user location is to be used for generating a viewport video, source information of patches within an atlas and information indicating a basic video, which is to be decoded first within an atlas, are required.” and [0041], “An immersive video means a video that enables a viewport to dynamically change when a viewing position of a user changes. A plurality of input videos is required to realize an immersive video. Each of the plurality of input videos maybe referred to as a source video or a source view.”; Examiner’s note: As disclosed by Yun in [0230], “it is necessary to extract an atlas including source videos that are adjacent to the user's viewpoint.” These source videos (or “views”) are neighboring each other and can be from different groups, as outlined in [0231].) and the second group comprises a view which is spatially neighboring to a second view included in the first group. (See Yun, [0230]-[0231] and [0041] above.) 10. Yun doesn't explicitly disclose but Kroon discloses: [[A method of]] switching [[an atlas according to a watching position comprising:]] (Kroon, [0026], “When a viewer moves around a virtual scene, the viewport pose gradually changes and thus, so does the rendering priority. This invention further modifies that rendering priority in the process of switching of video tracks.” and [0115], “Conventionally the rendering priority is based on proximity to the viewport. It is proposed to also adapt the rendering priority (i.e. increase or decrease) when the video tracks are changes. For instance, when three views [v1, v2, v3] are rendered, but for the next segment, views [v2, v3, v4] would be available, then the contribution of v1 in the weighted sum can be gradually modified towards zero. When v4 comes in, the contribution (i.e. rendering priority) may be gradually raised starting from zero.” and [0087], “The atlas system contains multiple camera views and through range sensors, depth estimation or otherwise, can also create depth maps. The combination of all of these forms a video track per atlas 202 (in this case).”) [[requesting a transmission of a first]] sub-[[bitstream including the first atlas;]] (See Kroon, [0008] below.) [[decoding the first atlas from the first]] sub-[[bitstream; and]] (Kroon, [0008], “The idea of sub-bitstream access is that part of the units in a bitstream can be removed resulting in a smaller bitstream that is still valid. A common form of sub-bitstream access is temporal access (i.e. lowering the frame rate) but in MIV there is also support for spatial access, which is implemented by allowing the removal of a subset of the atlases. Each atlas provides a video track and combinations of video tracks are used to render the MIV.”) [[wherein in response to the request for the transmission of the first]] sub-[[bitstream including the first atlas,]] (Kroon, [0008], “The idea of sub-bitstream access is that part of the units in a bitstream can be removed resulting in a smaller bitstream that is still valid. A common form of sub-bitstream access is temporal access (i.e. lowering the frame rate) but in MIV there is also support for spatial access, which is implemented by allowing the removal of a subset of the atlases. Each atlas provides a video track and combinations of video tracks are used to render the MIV.”) a second sub-bitstream including [[a second atlas is also received along with the first]] sub-[[bitstream,]] the second atlas not being used to render the viewport image, (Kroon, [0088]-[0089], “FIG. 3 shows a server system 304 sending video tracks 302 to a client system 306. The server system 304 receives an output from each one of the atlases 202 and encodes each one. In this example, the encoded output of a single atlas 202 forms a video track 302. The server system 304 then determines which video tracks 302 to send to the client system 306 (e.g. based on the viewport 104 requested by the client system 306). In FIG. 3, only two video tracks 302 are sent to the client system 306 from the three potential video tracks 302 received by the server system 304 from the atlases 202. For streaming, the server system 304 can encode all of the outputs from each of the atlases 202. However, as it is infeasible to transmit all video tracks 302, only a subset of video tracks 302 are transmitted completely to the client system 306. Other video tracks 302 may also be partly transmitted if certain areas of a viewport 104 are not predictable by the video tracks 302 which were fully transmitted.” and [0029], “The video tracks need not be strictly “removed” (e.g. not requested from the server) when the rendering priority reaches a low value (e.g. 0, 0.1 etc.) or after some time has passed. For example, the video track could be kept whilst at a rendering priority of 0 and thus effectively not being used to render the immersive video.” and [0099], “When a video track 302 has a low rendering priority 402, it will only be used when there is no other video track 302 that has the required information. When a video track 302 has zero rendering priority (i.e. 0%), it is as if the video track 302 is not there.” and [0026], “In general, the rendering priority will be based on the characteristics of the viewport and the video tracks (pose, field of view, etc.). When a viewer moves around a virtual scene, the viewport pose gradually changes and thus, so does the rendering priority. This invention further modifies that rendering priority in the process of switching of video tracks.” and [0008], “The idea of sub-bitstream access is that part of the units in a bitstream can be removed resulting in a smaller bitstream that is still valid. A common form of sub-bitstream access is temporal access (i.e. lowering the frame rate) but in MIV there is also support for spatial access, which is implemented by allowing the removal of a subset of the atlases. Each atlas provides a video track and combinations of video tracks are used to render the MIV.” and [0087], “The atlas system contains multiple camera views and through range sensors, depth estimation or otherwise, can also create depth maps. The combination of all of these forms a video track per atlas 202 (in this case).”; Examiner’s note: As disclosed by Kroon in [0029], video tracks (each provided by atlases, disclosed in [0008]) “need not be strictly ‘removed’ (e.g. not requested from the server) when the rendering priority reaches a low value.” Thus, an atlas with a low rendering priority may be received but may not be part of the rendering process.) [[wherein]] in response to a switch request triggered by a [[movement of the watching position that remains close to a view belonging to the first group]] while moving away from [[views belonging to the second group, (Kroon, [0026], “When a viewer moves around a virtual scene, the viewport pose gradually changes and thus, so does the rendering priority. This invention further modifies that rendering priority in the process of switching of video tracks.” and [0114]-[0115], “For example, view blending is based on a weighted sum of contributions, whereby each view has a weight (uniform or varying per pixel). Each view may be rendered into a separate buffer and the buffers may be blended based on the proximity of the viewport to each of the views. In this example, the proximity of the viewport to each of the views is used to determine the rendering priority. Conventionally the rendering priority is based on proximity to the viewport.”; Examiner’s note: As disclosed by Kroon in [0114], the rendering priority uses “the proximity of the viewport to each of the views is used to determine the rendering priority.” The priority to other views not only changes based on a viewer’s location but also will determine when a switch in video tracks in necessary, as disclosed in [0026].) a third sub-bitstream including a third atlas, [[which comprises data on views included in a third group,]] is received along with the first sub-bitstream, and reception of the second sub-bitstream is stopped, [[and]] (Kroon, [0026], “When a viewer moves around a virtual scene, the viewport pose gradually changes and thus, so does the rendering priority. This invention further modifies that rendering priority in the process of switching of video tracks.” and [0098]-[0099], “FIG. 4 shows a first example of a change of video tracks 302. Two video tracks 302a and 302b are initially used to render the immersive video. However, the previous video track 302a is changed to a new video track 302c between segment two 404 and segment three 406. The rendering priority 402 of the previous video track 302a is gradually lowered before the end of segment two 404 and the rendering priority 402 of the new video track 302c is gradually increased from the start of segment three 406. In this case, the video track 302b is not removed and thus is also used to render the video. When a video track 302 has a low rendering priority 402, it will only be used when there is no other video track 302 that has the required information. When a video track 302 has zero rendering priority (i.e. 0%), it is as if the video track 302 is not there. However, when the information is missing it can be inpainted with the video track 302 with zero rendering priority 402. By gradually changing the rendering priority 402 of a video track 302, the sudden appearance, disappearance or replacement of rendering artefacts is reduced.” and [0102], “The client 306 could also (stepwise) switch to a lower resolution version of the previous video track 302a before dropping it entirely and/or first request a low resolution version of the new video track 302c. This would allow the client 306 to keep the three video tracks 302a, 302b and 302c before switching to a high resolution and dropping the previous video track 302a.” and [0114]-[0115], “For example, view blending is based on a weighted sum of contributions, whereby each view has a weight (uniform or varying per pixel). Each view may be rendered into a separate buffer and the buffers may be blended based on the proximity of the viewport to each of the views. In this example, the proximity of the viewport to each of the views is used to determine the rendering priority. Conventionally the rendering priority is based on proximity to the viewport. It is proposed to also adapt the rendering priority (i.e. increase or decrease) when the video tracks are changes. For instance, when three views [v1, v2, v3] are rendered, but for the next segment, views [v2, v3, v4] would be available, then the contribution of v1 in the weighted sum can be gradually modified towards zero. When v4 comes in, the contribution (i.e. rendering priority) may be gradually raised starting from zero.” and [0028], “In this case, a video track could be removed when the corresponding rendering priority reaches a pre-defined value.” and [0087], “The atlas system contains multiple camera views and through range sensors, depth estimation or otherwise, can also create depth maps. The combination of all of these forms a video track per atlas 202 (in this case).” and [0008], “The idea of sub-bitstream access is that part of the units in a bitstream can be removed resulting in a smaller bitstream that is still valid. A common form of sub-bitstream access is temporal access (i.e. lowering the frame rate) but in MIV there is also support for spatial access, which is implemented by allowing the removal of a subset of the atlases. Each atlas provides a video track and combinations of video tracks are used to render the MIV.” and [0086], “MPEG Immersive Video (MIV) specifies a bitstream format with multiple atlases 202. Preferably, each atlas 202 has at least a geometry and texture attribute video. In the case of adaptive streaming, it is expected that each atlas 202 outputs multiple types of video data forming separate video tracks (e.g. a video track per atlas 202). Each video track is divided into short segments (e.g. in the order of a second) to allow a client to respond to a change in viewport pose 103 by changing the subset of video tracks that is requested from a server.”; Examiner’s note: In [0098], Kroon discloses the process of switching one of two video tracks to a third video track. As part of the process: “The rendering priority 402 of the previous video track 302a is gradually lowered before the end of segment two 404 and the rendering priority 402 of the new video track 302c is gradually increased from the start of segment three 406.” When the previous video track’s rendering priority is low enough, the video track can be removed (and thus stopped as a result), as disclosed in [0028].) 11. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of Yun to include the disclosure of a process of switching an atlas according to a watching position, making use of sub-bitstreams for transmitting video data, atlases that may not be used to render a viewport image, and a switch request triggered by movement of a watching position away from a group of views where a new sub-bitstream of new views is received and an existing sub-bitstream of views being moved away from are stopped, of Kroon. The motivation for this modification could have been to allow for a user’s rendered viewport to dynamically update by switching atlases as a user’s watching position changes. The dynamic update can make the virtual environment feel more engaging with new content or allow for a change in resolution to automatically adjust for bandwidth. In addition, sub-bitstreams can be used to help remove or modify aspects of a bitstream while still transmitting valid data. This can be useful when an atlas switches as existing sub-bitstreams can be updated to remove an old atlas while adding a new atlas. Atlases received but not used for viewport rendering may still be useful as the atlas may contain neighboring views that may be used for rendering later in a immersive video session. Lastly, the ability for sub-bitstreams to dynamically adjust by transmitting a new sub-bitstream while ending another as a user moves throughout a virtual environment helps to reduce latency and quickly update the rendered viewport. Overall, these features should reduce view transitions and help keep a user engaged as they navigate a virtual environment. 12. As per claim 5, Yun in view of Kroon discloses: The method of claim 1, wherein the switch request occurs in consideration of at least one of moving speed of the watching position, rotational speed of the watching position, a spatial interval between views (Kroon, [0026], “In general, the rendering priority will be based on the characteristics of the viewport and the video tracks (pose, field of view, etc.). When a viewer moves around a virtual scene, the viewport pose gradually changes and thus, so does the rendering priority. This invention further modifies that rendering priority in the process of switching of video tracks.” and Kroon, [0115], “Conventionally the rendering priority is based on proximity to the viewport.” and Kroon, [0114], “For example, view blending is based on a weighted sum of contributions, whereby each view has a weight (uniform or varying per pixel). Each view may be rendered into a separate buffer and the buffers may be blended based on the proximity of the viewport to each of the views. In this example, the proximity of the viewport to each of the views is used to determine the rendering priority.” and Kroon, [0116], “Parameters which may be used to determine the rendering priority are proximity to output view and difference in depth value between rendered depth and source depth.”; Examiner’s note: As disclosed by Kroon in [0026], a “switching of video tracks” may be required when a viewer moves around a virtual scene and as rendering priority changes. In [0115], rendering priority is determined based “proximity to the viewport,” which would assign a weight accounting for the distance between the viewer’s current location and other virtual views.) or a number of views in an atlas. (Yun, [0151], “Accordingly, when the syntax atlas_type has a value of 0, it indicates that an atlas includes a basic video or at least one patch necessary for view rendering. On the other hand, when the syntax atlas_type has a value of 1, it indicates that an atlas does not include a basic video or patches necessary for view rendering.” and Yun, [0107], “Accordingly, the present disclosure suggests that a priority order is set among atlases and then a priority order of decoding is determined according the set priority order. Specifically, when the number of atlases is M and the number of video decoders is N, top N videos according to a priority order among M atlases may be decoded through a video decoder.” and [0109]-[0111], “As in the example illustrated in FIG. 8, when an immersive video output apparatus includes two video decoders, only top two atlases in a priority order may be selected as decoding targets. A priority order among atlases may be determined based on whether or not an atlas includes information on a ROI video. Specifically, an atlas including a patch derived from a ROI video (hereinafter, referred to as a ROI patch) maybe set to have higher priority than other atlases. When ROI patches are dispersedly stored in a plurality of atlases, a priority order among atlases may be determined based on the number of ROI patches included in each atlas. Specifically, when the number of ROI patches included in a first atlas is larger than that of ROI patches included in a second atlas, the first atlas may be set to have higher priority than the second atlas. In addition, an atlas including a complete video patch may be set to have higher priority than an atlas including only segment video patches.” and Yun, [0230], “For viewport video rendering according to a user's position, it is necessary to extract an atlas including source videos that are adjacent to the user's viewpoint.” and Yun, [0079], “Atlas-related data may include at least one of information on the number of atlases, information on a priority order among atlases, a flag indicating whether or not an atlas includes a complete video, and information related to scaling of an atlas.” and Yun, [0233], “In Table 19, the syntax num_view_minus1 is used to determine the number of source videos. For example, the syntax num_view_minus1 may represent a value that is obtained by subtracting 1 from the number of source videos.” and Yun, [0239], “An immersive video processing apparatus may determine an atlas and a basic video, which are necessary to synthesize a viewport video, based on information included in the source video parameter list ‘view_params_list’ of Table 19.”; Examiner’s note: As disclosed by Yun in [0151], metadata “atlas_type” stores an attribute that determines if “an atlas includes a basic video or at least one patch necessary for view rendering.” When the value is 1, “it indicates that an atlas does not include a basic video or patches necessary for view rendering.” In the situation, for the number of view(s) included in the atlas, it is known to not contain enough views for view rendering which could prompt a request for an atlas switch. In addition, Yun also discloses in [0107]-[0112] that view videos can be prioritized. For instance, in Yun [0110]: “Specifically, when the number of ROI patches included in a first atlas is larger than that of ROI patches included in a second atlas, the first atlas may be set to have higher priority than the second atlas.” This could prompt a switch from a lower priority atlas with less patches to a higher priority atlas with more patches, especially when a viewport video is updated based on a user’s viewpoint.) 13. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 1 of Yun to include the disclosure of a switch request occurring due to a spatial interval between views, of Kroon. The motivation for this modification could have been to give priority to views depending on their proximity to a user’s watching position. As a watching position changes, the distance interval between watching positions changes, possibly prompting an atlas switch. This would allow a user’s rendered viewport to be appropriately updated depending on the position of the user. 14. Claim 10 is similar in scope to claim 1 except for additional limitations that Yun in view of Kroon discloses: A device of switching an atlas according to a watching position comprising: (Kroon, [0040], “The invention also provides a computer program product comprising computer program code means which, when executed on a computing device having a processing system, cause the processing system to perform all of the steps of the method defined above for transitioning from a first set of video tracks, VT1, to a second set of video tracks, VT2.” and Kroon, [0026], “When a viewer moves around a virtual scene, the viewport pose gradually changes and thus, so does the rendering priority. This invention further modifies that rendering priority in the process of switching of video tracks.”) a MIV decoder configured to: (Yun, [0096], “An immersive video output apparatus may extract atlas data by parsing a bitstream, which is received from an immersive video processing apparatus, and decode an atlas based on extracted data (S511).” and Kroon, [0086], “MPEG Immersive Video (MIV) specifies a bitstream format with multiple atlases 202.” and Kroon, [0007]-[0008], “A new standard (ISO/IEC 23090-12) for MPEG Immersive Video (MIV) describes how to prune and pack multiview+depth video into multiple texture and depth atlases for coding by 2D video codecs. … The idea of sub-bitstream access is that part of the units in a bitstream can be removed resulting in a smaller bitstream that is still valid. A common form of sub-bitstream access is temporal access (i.e. lowering the frame rate) but in MIV there is also support for spatial access, which is implemented by allowing the removal of a subset of the atlases. Each atlas provides a video track and combinations of video tracks are used to render the MIV.” and Kroon, [0104], “However, when the client 306 receives only one video track 302, then a different strategy may be needed. … Some examples on how to halve the resources required to transmit/decode/render a video track 302 are ...”) [[obtain information on a view image to render a viewport image corresponding to the watching position, and determining a first atlas which comprises data on the view required to render the viewport image; and]] (See rejection for claim 1.) a transceiver configured to: (Yun, [0096], “An immersive video output apparatus may extract atlas data by parsing a bitstream, which is received from an immersive video processing apparatus, and decode an atlas based on extracted data (S511).” and Kroon, [0030]-[0031], “Obtaining the video tracks VT2 may be based on requesting the video tracks VT2 from a server or receiving the video tracks VT2 from a server. The video tracks can be requested based on the viewport wanted by the user. For example, the user may require a particular viewport of the immersive video and thus request the video tracks needed to render the wanted viewport.”) [[request a transmission of a first sub-bitstream including the first atlas; and]] (See rejection for claim 1.) receive the first sub-bitstream of the first atlas, (Yun, [0082], “The bitstream generator 150 generates a bitstream based on encoded video data and metadata. A bitstream thus generated may be transmitted to an immersive video output apparatus.” and Yun, [0085], “The bitstream parsing unit 210 parses video data and metadata from a bitstream. Video data may include data of an encoded atlas.”) [[wherein the first atlas is decoded from the first sub-bitstream, wherein the viewport image is rendered based on the first atlas, wherein in response to the request for the transmission of the first sub-bitstream including the first atlas, a second sub-bitstream including a second atlas is also received along with the first sub-bitstream, the second atlas not being used to render the viewport image, wherein the first atlas comprises data on views included in a first group and the second atlas comprises data on views included in a second group,]] (See rejection for claim 1.) wherein views belonging to each of the first group and the second group are classified into a basic view and an additional view, (Kroon, [0085], “In an alternative example, the cameras 204 could provide the complete (basic) views and the sensors 206 could provide additional views. Furthermore, the views from all of the sensors are grouped into atlases 202 whereby atlases 202 may or may not intersect.” and Yun, [0046], “The view optimizer 110 classifies source videos into basic videos and additional videos.” and Yun, [0041], “Each of the plurality of input videos maybe referred to as a source video or a source view.” and [0070], “A complete video patch and a segment video patch may be set to be allocated in different atlases. Alternatively, an atlas may be constructed by mixing a complete video patch and a segment video patch.” and Yun, [0231], “For example, when 16 source videos are divided into two groups including 8 source videos respectively, an atlas for a first group may be generated through a pruning process for the source videos 1 to 8, and an atlas for a second group may be generated through a pruning process for the source videos 9 to 16. When only random/spatial access or an atlas based on user location is to be used for generating a viewport video, source information of patches within an atlas and information indicating a basic video, which is to be decoded first within an atlas, are required.”) [[wherein in response to a switch request triggered by a movement of the watching position that remains close to a view belonging to the first group while moving away from views belonging to the second group, a third sub-bitstream including a third atlas, which comprises data on views included in a third group, is received along with the first sub-bitstream, and reception of the second sub-bitstream is stopped, and wherein the second group comprises a view which is spatially neighboring to a first view included in the first group and the second group comprises a view which is spatially neighboring to a second view included in the first group.]] (See rejection for claim 1.) 15. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the device of Yun to include the disclosure of using a device with a transceiver and MIV decoder for immersive video applications, of Kroon. The motivation for this modification could have been to have a dedicated hardware device that is designed for immersive video applications. Such a device would allow a user to easily and conveniently access an immersive video application without any additional hardware or software. 16. Claim 14, which is similar in scope to dependent claim 5 and independent claim 10, is thus rejected under the same rationale as described above. 17. Claims 6 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Yun et al. (US-2021/0067757-A1, hereinafter "Yun") in view of Kroon et al. (US-2024/0107110-A1, hereinafter "Kroon"), and further in view of Han et al. (US-11381817-B2, hereinafter "Han"). 18. As per claim 6, Yun in view of Kroon discloses: The method of claim 1, wherein each of the first sub-bitstream, the second sub-bitstream and the third sub-bitstream is received from an [[edge]] server, (Kroon, Fig. 4; [0098], “FIG. 4 shows a first example of a change of video tracks 302. Two video tracks 302a and 302b are initially used to render the immersive video. However, the previous video track 302a is changed to a new video track 302c between segment two 404 and segment three 406. The rendering priority 402 of the previous video track 302a is gradually lowered before the end of segment two 404 and the rendering priority 402 of the new video track 302c is gradually increased from the start of segment three 406. In this case, the video track 302b is not removed and thus is also used to render the video.” and [0102], “The client 306 could also (stepwise) switch to a lower resolution version of the previous video track 302a before dropping it entirely and/or first request a low resolution version of the new video track 302c. This would allow the client 306 to keep the three video tracks 302a, 302b and 302c before switching to a high resolution and dropping the previous video track 302a.” and [0025], “For example, the color of a pixel on the rendered immersive video may be based on the color of the corresponding pixels in three video tracks provided.” and [0087], “The atlas system contains multiple camera views and through range sensors, depth estimation or otherwise, can also create depth maps. The combination of all of these forms a video track per atlas 202 (in this case).” and [0008], “The idea of sub-bitstream access is that part of the units in a bitstream can be removed resulting in a smaller bitstream that is still valid. A common form of sub-bitstream access is temporal access (i.e. lowering the frame rate) but in MIV there is also support for spatial access, which is implemented by allowing the removal of a subset of the atlases. Each atlas provides a video track and combinations of video tracks are used to render the MIV.”; Examiner’s note: Fig. 4 in Kroon discloses an example of three video tracks (each provided by atlases, disclosed in [0008]) demonstrating a change of one of the two existing tracks to a third, different track.) [[and a number of sub-bitstreams received by the edge server from a transmission server]] is greater than the number of sub-bitstreams provided by the [[edge]] server to a terminal. (Kroon, [0088]-[0089], “In FIG. 3, only two video tracks 302 are sent to the client system 306 from the three potential video tracks 302 received by the server system 304 from the atlases 202. For streaming, the server system 304 can encode all of the outputs from each of the atlases 202. However, as it is infeasible to transmit all video tracks 302, only a subset of video tracks 302 are transmitted completely to the client system 306. Other video tracks 302 may also be partly transmitted if certain areas of a viewport 104 are not predictable by the video tracks 302 which were fully transmitted.”) 19. Yun in view of Kroon doesn't explicitly disclose but Han discloses: [[The method of claim 1, wherein each of the first sub-bitstream, the second sub-bitstream and the third sub-bitstream is received from an]] edge [[server,]] (Han, col. 4, lines 32-40, “In one example, viewport-guided transcoding is applied at the network-edge, e.g., in an edge server. To illustrate, in one example, the edge server may collect the viewport movement traces from a client device periodically, or according to another schedule. At the client device-side, the video player may collect actual viewport data, e.g., via motion sensors for 360-degree video streaming or volumetric video streaming, or using gaze tracking for regular video streaming or non-360-degree panoramic video streaming.”) and a number of sub-bitstreams received by the edge server from a transmission server [[is greater than the number of sub-bitstreams provided by the]] edge [[server to a terminal.]] (Han, col. 12, lines 56-61, “At optional step 310, the processing system (e.g., of an edge server) may obtain at least a portion of an immersive visual stream, the at least the portion including at least one frame. For instance, the portion of the immersive visual stream may be obtained from a centralized server for distributing immersive visual streams.” and col. 16, lines 12-30, “In another example, the method 300 may include storing the frame (and additional frames and/or chunks of the immersive visual stream) at the edge server. The storing may be prior to performing the operations of steps 320-380, or may be after step 380. For instance, the immersive visual stream, or at least a portion thereof, may be stored for other users who may be interested in experiencing the immersive visual stream via the respective mobile computing devices that may be served by the processing system. In still another example, the method 300 may include performing the steps 320-380 for a plurality of different users and/or mobile computing devices. For instance, the immersive visual stream may be a live or near-live stream that may be experienced simultaneously by multiple users via the processing system (e.g., of an edge server) and/or via other edge servers. Since each of these users may have a unique viewport, the processing system may perform separate viewport predictions and may apply unique viewport-adaptive encodings for each mobile computing device.” and col. 8, lines 50-65, “In the example of FIG. 1, device 132 of user 192 may establish a session with edge server 108 for obtaining an immersive visual stream, which may be obtained as a sequence of frames and/or in chunks comprising a sequence of frames. ... In one example, the edge server 108 may store a copy of the immersive visual stream (e.g., for a recorded video program). In another example, the edge server 108 may obtain the immersive visual stream (e.g., the frames thereof) from a centralized server for distributing immersive visual streams. For instance, AS 104 or server 106 may comprise such a centralized server.” and col. 11, lines 10-16, “Continuing with the present example in reference to FIG. 1, the edge server 108 may transmit the frame 170 containing the applicable encoding to device 132.”; Examiner’s note: As disclosed by Han in col. 8, lines 50-65, an “edge server 108 may store a copy of the immersive visual stream (e.g., for a recorded video program)” obtained “from a centralized server.” In addition, as disclosed in col. 16, lines 12-30, an edge server can store the “immersive visual stream, or at least a portion thereof, … for other users who may be interested in experiencing the immersive visual stream.” These disclosures outline an edge server that can store an entire copy of the immersive visual stream and transmit one or more video frames to more than one user. This would match the scenario of “a number of sub-bitstreams received by the edge server from a transmission server is greater than the number of sub-bitstreams provided by the edge server to a terminal,” especially if the edge server needs to send data to more than one user. It would be necessary for the edge server to store a greater amount of the immersive visual stream data to serve multiple users, especially if each user has a “unique viewport” in which separate processing and viewport predictions are required.) 20. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 1 of Yun in view of Kroon to include the disclosure of utilizing an edge server for immersive video and that the number of sub-bitstreams received by the edge server from a transmission server is greater than the number of sub-bitstreams provided by the edge server to a terminal, of Han. The motivation for this modification could have been to make use of edge servers in order to help reduce latency of an immersive video, especially as a user moves throughout the virtual environment. In addition, an edge server that has received most or all of an immersive video is able to more quickly transition between serving atlases/video tracks to a user (as well as serve multiple users the same immersive video). Due to an edge server likely being nearby, this saves transmission time that would normally be between a central server and a user. 21. Claim 15, which is similar in scope to dependent claim 6 and independent claim 10, is thus rejected under the same rationale as described above. The motivation for this modification is the same as claim 6. 22. Claims 8-9 and 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Yun et al. (US-2021/0067757-A1, hereinafter "Yun") in view of Kroon et al. (US-2024/0107110-A1, hereinafter "Kroon"), and further in view of Roimela et al. (US-2021/0281879-A1, hereinafter "Roimela"). 23. As per claim 8, Yun in view of Kroon discloses: The method of claim 1, wherein the method further comprises: receiving metadata comprising mapping relationships between views and atlases, and (Yun, [0241], “The syntax view_enabled_in_atlas_flag[a][v] and the syntax view_complete_in_atlas_flag[a] [v], which represent a mapping relationship between each source video and each atlas, may be encoded/decoded, only when the syntax spatial_access_flag has a value of 1.” and [0041], “Each of the plurality of input videos maybe referred to as a source video or a source view.” and [0014], “Herein, the metadata may include a first flag …” and [0237], “When the a-th atlas includes a patch that is extracted from the v-th source video, the syntax view_complete_in_atlas_flag[a][v] may be encoded/decoded which indicates whether or not the v-th source video is included in the a-th atlas as a complete video patch.” and [0095]-[0096], “The atlas may be encoded (S416), and the metadata and the encoded atlas may be transmitted to an immersive video output apparatus. An immersive video output apparatus may extract atlas data by parsing a bitstream, which is received from an immersive video processing apparatus, and decode an atlas based on extracted data (S511).”; Examiner’s note: As disclosed by Yun in [0014], metadata uses flags to indicate certain properties about the immersive video. The “view_complete_in_atlas_flag” disclosed in [0241] represents “a mapping relationship between each source video and each atlas.” Metadata, such as this, is then transmitted to an “immersive video output apparatus.”) [[wherein the metadata is received periodically]] 24. Yun in view of Kroon doesn't explicitly disclose but Roimela discloses: wherein the metadata is received periodically (Roimela, [0204]-[0205], “The apparatus may further include wherein the view-to-atlas mapping metadata comprises a temporal update of an atlas map together with a camera extrinsic in a view parameter extrinsic substructure of an adaptation parameter structure. The apparatus may further include wherein the view-to-atlas mapping metadata comprises a temporal update as an atlas map update substructure of an adaptation parameter structure.”) 25. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 1 of Yun in view of Kroon to include the disclosure of metadata being received periodically, of Roimela. The motivation for this modification could have been to send updated metadata to a user that accounts for the user’s position. This would help keep track of which atlases have relevant views to the user and which can be filtered. By keeping track and filtering unneeded atlases, this reduces the bandwidth necessary to send the relevant view data and could provide additional bandwidth for higher quality views. 26. As per claim 9, Yun in view of Kroon, and further in view of Roimela discloses: The method of claim 8, wherein mapping information on more atlases than a number of atlases decoded by a terminal is stored in the terminal. (Roimela, [0122]-[0123], “1. View-to-atlas mapping metadata. In one embodiment, the adaptation_params_rbsp structure that contains MIV related view metadata is contained in the universally accessible “master” atlas (i.e. the atlas with vuh_atlas_id equal to 0x3F). New elements in the adaptation_params_rbsp structure are added to provide information about mapping from views to atlases. This mapping may indicate, for every view, the atlas that contains patches referring back to the view in question. The renderer may apply view frustum culling to each view first. All views that are deemed potentially visible may then be queried for the atlas mapping metadata, and the combined atlas mapping metadata may indicate the atlases that must be accessed in order to render the visible views.” and [0146]-[0149], “In this embodiment, the renderer may cull all patches against the current rendering viewing frustum, and decode only the atlas sub-bitstreams that contain potentially visible patches. This can be implemented in several ways, of which two examples are: loop over all patch atlases, detect potentially visible patches, and once a first potentially visible patch is found, mark that atlas as required and move to the next one, or perform the view culling of Embodiment 1 (1. View-to-atlas mapping metadata) first, then process only patches referring to a potentially visible view, and mark the relevant atlases as required After finding the required atlases, access to those may continue as in Embodiment 1 (1. View-to-atlas mapping metadata), potentially via a network request before decoding the relevant atlas sub-bitstream.” and [0217], “The apparatus may further include wherein the at least one volumetric video atlas is culled without having to access every atlas metadata bitstream.” and [0222], “Other aspects of the apparatus may include the following. The apparatus may further include means for rendering a view frustum corresponding to one or more sets of components of the volumetric video bitstream that have not been culled. The atlas-to-view mapping metadata may be received as a supplemental enhancement information message comprising a payload size and bitmask indicating mapping information between views and atlases. The atlas-to-view mapping metadata may specify a persistence of a previous atlas view supplemental enhancement information message. The persistence may be specified using a flag, wherein the flag being equal to zero specifies that the atlas view supplemental enhancement information message applies to a current atlas frame; and the flag being equal to one specifies that the atlas view supplemental enhancement information message applies to the current atlas frame and persists for subsequent atlas frames in decoding order until meeting at least one condition comprising a beginning of a new sequence, an ending of the at least one volumetric video bitstream, or an atlas frame having a supplemental enhancement information message present. ... The atlas-to-view mapping metadata or the atlas-to-object mapping metadata may be received together with a camera extrinsic in a view parameter extrinsic substructure of an adaptation parameter structure.”; Examiner’s note: As disclosed by Roimela in [0122]-[0123] and [0146]-[0149], a “master” atlas and “View-to-atlas mapping metadata” provide mapping information so that a rendering apparatus may mark and “detect potentially visible patches” ([0147]) to perform view culling ([0148]). [0149] further discloses: “After finding the required atlases, access to those may continue as in Embodiment 1 (1. View-to-atlas mapping metadata), potentially via a network request before decoding the relevant atlas sub-bitstream.” Thus, the culling process makes use of the mapping metadata and performs the process prior to receiving or decoding any atlases. This discloses that the mapping information would have been received by the rendering apparatus prior to receiving or decoding any atlases. In addition, as disclosed in [0222], “atlas-to-view mapping metadata may specify a persistence of a previous atlas view supplemental enhancement information message. The persistence may be specified using a flag, wherein ... the flag being equal to one specifies that the atlas view supplemental enhancement information message applies to the current atlas frame and persists for subsequent atlas frames in decoding order until meeting at least one condition ...” This describes that the mapping metadata can persist and is available beyond an initial use.) 27. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 8 of Yun in view of Kroon to include the disclosure of mapping information on more atlases than a number of atlases received or decoded by a terminal/device is stored in the terminal/device, of Roimela. The motivation for this modification could have been to allow a user device to be aware of available atlases that correspond to views without needing to download any specific atlas. By having additional mapping metadata information, a device can analyze these “atlas-to-view” mappings and determine which mapping would produce a view relevant to a user’s position. This saves bandwidth and time as only atlases that are necessary to produce a view would need to be transmitted to the user device. 28. As per claim 17, Yun in view of Kroon, and further in view of Roimela discloses: The device of claim 10, wherein the transceiver further receives metadata comprising mapping relationships between views and atlases, and (Yun, [0241], “The syntax view_enabled_in_atlas_flag[a][v] and the syntax view_complete_in_atlas_flag[a] [v], which represent a mapping relationship between each source video and each atlas, may be encoded/decoded, only when the syntax spatial_access_flag has a value of 1.” and [0041], “Each of the plurality of input videos maybe referred to as a source video or a source view.” and [0014], “Herein, the metadata may include a first flag …” and [0237], “When the a-th atlas includes a patch that is extracted from the v-th source video, the syntax view_complete_in_atlas_flag[a][v] may be encoded/decoded which indicates whether or not the v-th source video is included in the a-th atlas as a complete video patch.” and [0095]-[0096], “The atlas may be encoded (S416), and the metadata and the encoded atlas may be transmitted to an immersive video output apparatus. An immersive video output apparatus may extract atlas data by parsing a bitstream, which is received from an immersive video processing apparatus, and decode an atlas based on extracted data (S511).”; Examiner’s note: As disclosed by Yun in [0014], metadata uses flags to indicate certain properties about the immersive video. The “view_complete_in_atlas_flag” disclosed in [0241] represents “a mapping relationship between each source video and each atlas.” Metadata, such as this, is then transmitted to an “immersive video output apparatus.”) wherein the metadata is received periodically. (Roimela, [0204]-[0205], “The apparatus may further include wherein the view-to-atlas mapping metadata comprises a temporal update of an atlas map together with a camera extrinsic in a view parameter extrinsic substructure of an adaptation parameter structure. The apparatus may further include wherein the view-to-atlas mapping metadata comprises a temporal update as an atlas map update substructure of an adaptation parameter structure.”) The motivation for this modification is the same as claim 8. 29. As per claim 18, Yun in view of Kroon, and further in view of Roimela discloses: The device of claim 17, wherein mapping information on more atlases than a number of atlases received by the bitstream reception unit is stored in the device. (Roimela, [0122]-[0123], “1. View-to-atlas mapping metadata. In one embodiment, the adaptation_params_rbsp structure that contains MIV related view metadata is contained in the universally accessible “master” atlas (i.e. the atlas with vuh_atlas_id equal to 0x3F). New elements in the adaptation_params_rbsp structure are added to provide information about mapping from views to atlases. This mapping may indicate, for every view, the atlas that contains patches referring back to the view in question. The renderer may apply view frustum culling to each view first. All views that are deemed potentially visible may then be queried for the atlas mapping metadata, and the combined atlas mapping metadata may indicate the atlases that must be accessed in order to render the visible views.” and [0146]-[0149], “In this embodiment, the renderer may cull all patches against the current rendering viewing frustum, and decode only the atlas sub-bitstreams that contain potentially visible patches. This can be implemented in several ways, of which two examples are: loop over all patch atlases, detect potentially visible patches, and once a first potentially visible patch is found, mark that atlas as required and move to the next one, or perform the view culling of Embodiment 1 (1. View-to-atlas mapping metadata) first, then process only patches referring to a potentially visible view, and mark the relevant atlases as required After finding the required atlases, access to those may continue as in Embodiment 1 (1. View-to-atlas mapping metadata), potentially via a network request before decoding the relevant atlas sub-bitstream.” and [0217], “The apparatus may further include wherein the at least one volumetric video atlas is culled without having to access every atlas metadata bitstream.” and [0222], “Other aspects of the apparatus may include the following. The apparatus may further include means for rendering a view frustum corresponding to one or more sets of components of the volumetric video bitstream that have not been culled. The atlas-to-view mapping metadata may be received as a supplemental enhancement information message comprising a payload size and bitmask indicating mapping information between views and atlases. The atlas-to-view mapping metadata may specify a persistence of a previous atlas view supplemental enhancement information message. The persistence may be specified using a flag, wherein the flag being equal to zero specifies that the atlas view supplemental enhancement information message applies to a current atlas frame; and the flag being equal to one specifies that the atlas view supplemental enhancement information message applies to the current atlas frame and persists for subsequent atlas frames in decoding order until meeting at least one condition comprising a beginning of a new sequence, an ending of the at least one volumetric video bitstream, or an atlas frame having a supplemental enhancement information message present. ... The atlas-to-view mapping metadata or the atlas-to-object mapping metadata may be received together with a camera extrinsic in a view parameter extrinsic substructure of an adaptation parameter structure.”; Examiner’s note: As disclosed by Roimela in [0122]-[0123] and [0146]-[0149], a “master” atlas and “View-to-atlas mapping metadata” provide mapping information so that a rendering apparatus may mark and “detect potentially visible patches” ([0147]) to perform view culling ([0148]). Also, [0149] further discloses: “After finding the required atlases, access to those may continue as in Embodiment 1 (1. View-to-atlas mapping metadata), potentially via a network request before decoding the relevant atlas sub-bitstream.” Thus, the culling process makes use of the mapping metadata and performs the process prior to receiving or decoding any atlases. This discloses that the mapping information would have been received by the rendering apparatus prior to receiving or decoding any atlases. In addition, as disclosed in [0222], “atlas-to-view mapping metadata may specify a persistence of a previous atlas view supplemental enhancement information message. The persistence may be specified using a flag, wherein ... the flag being equal to one specifies that the atlas view supplemental enhancement information message applies to the current atlas frame and persists for subsequent atlas frames in decoding order until meeting at least one condition ...” This describes that the mapping metadata can persist and is available beyond an initial use.) The motivation for this modification is the same as claim 9. Conclusion 30. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. These are as follows: Berger et al. (US-2020/0396493-A1) discloses a method to change versions of an immersive video when a user changes their point of view. Also, Berger discloses that the viewing equipment obtains information representing a speed of change of point of view of the user. Boyce et al. (US-2023/0156229-A1) discloses a method that determines if an atlas' metadata identifies at least part of a view of interest corresponding to an atlas and filters atlases that are not part of a view of interest. Shen et al. (US-2024/0236337-A1) discloses a method to reduce latency during viewport switching in immersive video. 31. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW CLOTHIER whose telephone number is (571)272-4667. The examiner can normally be reached Mon-Fri 8:00am-4:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached at (571)272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MATTHEW CLOTHIER/Examiner, Art Unit 2614 /KENT W CHANG/Supervisory Patent Examiner, Art Unit 2614
Read full office action

Prosecution Timeline

Mar 29, 2023
Application Filed
Apr 11, 2025
Non-Final Rejection mailed — §103
Jul 11, 2025
Response Filed
Nov 05, 2025
Final Rejection mailed — §103
Feb 04, 2026
Request for Continued Examination
Feb 18, 2026
Response after Non-Final Action
Aug 18, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731343
UNIFYING DENSEPOSE AND 3D BODY MESH RECONSTRUCTION
3y 2m to grant Granted Sep 08, 2026
Patent 12530842
AIRBORNE LiDAR POINT CLOUD FILTERING METHOD DEVICE BASED ON SUPER-VOXEL GROUND SALIENCY
1y 11m to grant Granted Jan 20, 2026
Patent 12499800
IN-VEHICLE DISPLAY DEVICE
1y 12m to grant Granted Dec 16, 2025
Study what changed to get past this examiner. Based on 3 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
83%
Grant Probability
99%
With Interview (+20.0%)
2y 3m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 6 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month