Prosecution Insights
Last updated: October 02, 2026
Application No. 19/176,082

MULTI-DIRECTIONAL 2D SNAPSHOT IMAGE TRACK FOR V3C CONTENT

Non-Final OA §102§103
Filed
Apr 10, 2025
Priority
Apr 15, 2024 — provisional 63/634,034
Examiner
USSERY, CAIDEN ALEXANDER
Art Unit
Tech Center
Assignee
Samsung Electronics Co., Ltd.
OA Round
1 (Non-Final)
50%
Grant Probability
Moderate
1-2
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 50% of resolved cases
50%
Career Allowance Rate
1 granted / 2 resolved
-10.0% vs TC avg
Strong +100% interview lift
Without
With
+100.0%
Interview Lift
resolved cases with interview
Fast prosecutor
1y 3m
Avg Prosecution
17 currently pending
Career history
21
Total Applications
across all art units

Statute-Specific Performance

§101
1.1%
-38.9% vs TC avg
§103
75.0%
+35.0% vs TC avg
§102
19.6%
-20.4% vs TC avg
§112
3.3%
-36.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 2 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claims 4, 5, 7, 13, 14 and 15 are objected to because of the following informalities: Claims 4, 5, 13, and 14 recite the number of the number of subsamples. The Examiner is requesting clarity if the claims are directed to the number of subsamples or a count of the number of subsamples. Claims 7 and 15 recite a two-dimensional projected images. The Examiner is requesting the typographical error be resolved, for example "the two-dimensional projected images" or "a two-dimensional projected image". Appropriate correction is required. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1 & 18 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Thuong Canh et al. (Pat. Pub. US-20240040148-A1, herein after “Canh”). In regard to claim 1, Canh teaches [a]n apparatus comprising: a communication interface configured to receive a track including one or more samples “The second terminal 102 may receive the coded video data of the other terminal from the network 105” (Canh, ¶ [0043]) where the terminals are communication interfaces, wherein a respective one of the one or more samples includes a plurality of coded two-dimensional projected images of a coded volumetric frame as a plurality of sub-samples “the volumetric data may be a 3D data set of 2D images, such as slices from which a 2D projection of the 3D data set may be projected for example” (Canh, ¶ [0089]) where the received volumetric data is read as a track, and the data set is read as samples of 2D images of a volumetric frame or point cloud. Additionally, “The diagram 600 illustrates exemplary embodiments for streaming of coded point cloud data according to V-PCC” (Canh, ¶ [0088]); and a processor operably coupled to the communication interface “The techniques described above, can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media or by a specifically configured one or more hardware processors” (Canh, ¶ [0178]), the processor configured to: decode at least one of the plurality of coded two-dimensional projected images to generate at least one two-dimensional projected image “The second terminal 102 may receive the coded video data of the other terminal from the network 105, decode the coded data” (Canh, ¶ [0043]) where the video data is decoded, and present the at least one two-dimensional projected image “decode the coded data and may display the recovered video data at a local display device” (Canh, ¶ [0044]) after being decoded the video can be displayed or presented. In regard to claim 18, Canh teaches [a] method performed by an apparatus comprising: receiving a track including one or more samples, wherein a respective one of the one or more samples includes a plurality of coded two-dimensional projected images of a coded volumetric frame as a plurality of sub-samples “the volumetric data may be a 3D data set of 2D images, such as slices from which a 2D projection of the 3D data set may be projected for example” (Canh, ¶ [0089]) where the received volumetric data is read as a track, and the data set is read as samples of 2D images of a volumetric frame or point cloud. Additionally, “The diagram 600 illustrates exemplary embodiments for streaming of coded point cloud data according to V-PCC” (Canh, ¶ [0088]); decoding at least one of the plurality of coded two-dimensional projected images to generate at least one two-dimensional projected image “The second terminal 102 may receive the coded video data of the other terminal from the network 105, decode the coded data” (Canh, ¶ [0043]) where the video data is decoded; and presenting the at least one two-dimensional projected image “decode the coded data and may display the recovered video data at a local display device” (Canh, ¶ [0044]) after being decoded the video can be displayed or presented. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 2-3, 6-8, & 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Canh in view of Hamza Ahmed et al. (Pat. Pub. WO-2023114464-A1, herein after “Ahmed”). In regard to claim 2, Canh teaches [t]he apparatus of claim 1. Canh does not explicitly teach wherein the communication interface is further configured to receive a multi-dimensional snapshot camera information box, wherein the multi-dimensional snapshot camera information box includes a plurality of viewport information elements, each of the plurality of viewport information elements is associated with a respective one of the plurality of sub-samples, and a respective one of the plurality of viewport information elements provides camera information for an associated sub-sample. Ahmed teaches wherein the communication interface is further configured to receive a multi-dimensional snapshot camera information box, wherein the multi-dimensional snapshot camera information box includes a plurality of viewport information elements “The recommended viewport information may provide (e.g., dynamically) changing information, which may include translation and/or rotation of the node that includes the camera object and/or the intrinsic camera parameter of the camera object” (Ahmed, ¶ [0032]) where information is provided to the client regarding the viewport such as parameters or object information. Additionally, the media content received is 3 dimensional (Ahmed, ¶ [0010]), each of the plurality of viewport information elements is associated with a respective one of the plurality of sub-samples “Information about the camera used to provide a viewer’s viewport of the scene (e.g., including intrinsic and extrinsic camera parameters) may be required by the MAF to identify and/or to request the appropriate media for a (e.g., each) viewer of the scene” (Ahmed, ¶ [0037]) where information elements are for each viewer of the scene, additionally “the presentation engine may send an update request comprising updated information associated with the first camera object and updated information associated with the second camera object” (Ahmed, ¶ [0074]) where the information is associated for each camera object, based on the camera, and a respective one of the plurality of viewport information elements provides camera information for an associated sub-sample “Information on the intersection of the bounding box with the view frustum may be used to determine which subset(s) of the media is/are in the camera’s view” (Ahmed, ¶ [0076]) where information can be about a specific sample or portion, which is read as a sub-sample, and can be used to determine what is in view or out of view. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of encoding and decoding volumetric images taught by Canh, with the method of viewport information regarding areas inside the view taught by Ahmed to have viewport information for the encoded volumetric images. The motivation to do so would be to provide the user information about the observed volumetric object “provide information on a number of cameras being used for different viewers of the scene” (Ahmed, ¶ [0075]). In regard to claim 3, Canh in view of Ahmed teach [t]he apparatus of claim 2, wherein the multi-dimensional snapshot camera information box includes number information indicating the number of sub-samples in the respective one of the one or more samples “The MAF may consume the Camera and Viewinfo arguments (e.g., as shown in Table 9) provided in the updateViewQ function. The Camera argument may provide information on a number of cameras being used for different viewers of the scene” (Ahmed, ¶ [0075]) where the number of cameras used varies based on viewers, and each camera view can be read as a sub-sample of the sample views. The motivation to combine Ahmed with Canh would be to include a count of cameras. In regard to claim 6, Canh in view of Ahmed teach [t]he apparatus of claim 2, wherein the at least one two-dimensional projected image is presented based on the multi-dimensional snapshot camera information box “A media content may be enclosed in a Bounding Box. Information on the intersection of the bounding box with the view frustum may be used to determine which subset(s) of the media is/are in the camera’s view” (Ahmed, ¶ [0076]) where information based on what is in the view of the projected image is taught, and “A Media Client may use the inference regarding the relevant subset of the media information which may be visible in the camera viewport and may make an informed request(s) of relevant subset of the media information” (Ahmed, ¶ [0079]) where the information may be visible in the camera viewport. The motivation to combine Ahmed with Canh would be to include the information box in the viewport. In regard to claim 7, Canh in view of Ahmed teach [t]he apparatus of claim 2, wherein the respective one of the plurality of viewport information elements includes location information and direction information of a camera used to render a two-dimensional projected images in an associated sub-sample of the plurality of subsamples “Information described herein (e.g., the positions of the object, pose information of the viewer, and/or the transform of the position and orientation of the object) may use the scene’s coordinate system” (Ahmed, ¶ [0029]) and, “A Media Client may use the inference regarding the relevant subset of the media information which may be visible in the camera viewport and may make an informed request(s) of relevant subset of the media information” (Ahmed, ¶ [0034]) where the camera viewport information includes position, orientation, and pose as in a coordinate system. The motivation to combine Ahmed with Canh would be the same as in claim 2. In regard to claim 8, Canh in view of Ahmed teach [t]he apparatus of claim 7, wherein location information and direction information of a camera for a sub-sample of the plurality of subsamples are different from location information and direction information of a camera for another sub-sample of the plurality of subsamples “A pose of a viewer may be given (e.g., directly or indirectly) by the extrinsic properties of a rendering camera, which may be associated with the viewer and may be used by a presentation client. In many applications there may be multiple rendering cameras used by a rendering client (e.g., for different viewers). For example, each viewer may have a different set of intrinsic properties for their respective camera” (Ahmed, ¶ [0041]) and “Multiple cameras may be used to render and/or compose different viewpoints for different viewers of the same scene” (Ahmed, ¶ [0036]) where different views or cameras may have different viewpoints of the scene, which is read as having a different location information for the viewing of the sub-samples. The motivation to combine Ahmed with Canh would be to have multiple views. In regard to claim 19, Canh teaches [t]he method of claim 18. Canh fails to explicitly teach wherein the method further comprises: receiving a multi-dimensional snapshot camera information box, wherein the multi-dimensional snapshot camera information box includes a plurality of viewport information elements, wherein each of the plurality of viewport information elements is associated with a respective one of the plurality of sub-samples, and a respective one of the plurality of viewport information elements provides camera information for an associated sub-sample. Ahmed teaches receiving a multi-dimensional snapshot camera information box, wherein the multi-dimensional snapshot camera information box includes a plurality of viewport information elements “The recommended viewport information may provide (e.g., dynamically) changing information, which may include translation and/or rotation of the node that includes the camera object and/or the intrinsic camera parameter of the camera object” (Ahmed, ¶ [0032]) where information is provided to the client regarding the viewport such as parameters or object information. Additionally, the media content received is 3 dimensional (Ahmed, ¶ [0010]), wherein each of the plurality of viewport information elements is associated with a respective one of the plurality of sub-samples “Information about the camera used to provide a viewer’s viewport of the scene (e.g., including intrinsic and extrinsic camera parameters) may be required by the MAF to identify and/or to request the appropriate media for a (e.g., each) viewer of the scene” (Ahmed, ¶ [0037]), additionally “the presentation engine may send an update request comprising updated information associated with the first camera object and updated information associated with the second camera object” (Ahmed, ¶ [0074]) where the information is associated for each camera object, based on the camera and information elements are for each viewer of the scene, and a respective one of the plurality of viewport information elements provides camera information for an associated sub-sample “Information on the intersection of the bounding box with the view frustum may be used to determine which subset(s) of the media is/are in the camera’s view” (Ahmed, ¶ [0076]) where information can be about a specific sample or portion, which is read as a sub-sample, and can be used to determine what is in view or out of view. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of encoding and decoding volumetric images taught by Canh, with the method of viewport information regarding areas inside the view taught by Ahmed to have viewport information for the encoded volumetric images. The motivation to do so would be to provide the user information about the observed volumetric object “provide information on a number of cameras being used for different viewers of the scene” (Ahmed, ¶ [0075]). In regard to claim 20, Canh in view of Ahmed teaches[t]he method of claim 19, wherein the multi-dimensional snapshot camera information box includes number information indicating the number of sub-samples in the respective one of the one or more samples “The MAF may consume the Camera and Viewinfo arguments (e.g., as shown in Table 9) provided in the updateViewQ function. The Camera argument may provide information on a number of cameras being used for different viewers of the scene” (Ahmed, ¶ [0075]) where the number of cameras used varies based on viewers, and each camera view can be read as a sub-sample of the sample views. The motivation to combine Ahmed with Canh would be to include the number of cameras. Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Canh in view of Ying Hu (Pat. Pub. WO-2023024843-A1, herein after “Hu”). In regard to claim 10, Canh teaches [a]n apparatus comprising: a communication interface “The communication system 100 may include at least two terminals 102 and 103 interconnected via a network 105” (Canh, ¶ [0043]); each of the one or more samples is associated with a respective one of the one or more volumetric frames and includes at least two coded two-dimensional projected images associated with a volumetric frame as a plurality of sub-samples “ According to exemplary embodiments, the volumetric data may be a 3D data set of 2D images” (Canh, ¶ [0089]), and “At projection to images block 1103, the acquired point cloud data may be projected onto 2D images and encoded as image/video pictures with video-based point cloud coding (V-PCC)” (Canh, ¶ [0090]) where volumetric frames are taught and the point cloud is projected onto 2D images and is encoded; and a processor operably coupled to the communication interface “the embodiments may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program that is stored in a non-transitory computer-readable medium” (Canh, ¶ [0042]); the processor configured to: encode a plurality of two-dimensional projected images of one or more volumetric frames to generate a plurality of coded two-dimensional projected images “At projection to images block 1103, the acquired point cloud data may be projected onto 2D images and encoded as image/video pictures with video-based point cloud coding (V-PCC)” (Canh, ¶ [0090]) where the point cloud is projected onto 2D images and is encoded. Canh teaches of a plurality of 2D images associated with a volumetric frame but does not explicitly teach generate a track including one or more samples and transmit the track. Hu teaches generate a track including one or more samples “the encapsulation unit in the media file encapsulation process, a media track consists of many samples. For example, a sample of a video track is usually a video frame” (Hu, Page 6) where a media track is made of many samples and a sample is a video frame, or 2D image; and transmit the track “After encoding the point cloud media, the encoded data stream needs to be encapsulated and transmitted to the user” (Hu, Page 7) where the media file includes the track. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of encoding and decoding volumetric images taught by Canh, with the method of a media track to manage the samples taught by Hu to have encoded volumetric images handled in a track. The motivation to do so would be to transmit and encode the volumetric images based on the track contents. Claims 4-5 are rejected under 35 U.S.C. 103 as being unpatentable over Canh in view of Ahmed, Hu, and further in view of Soojin Hwang (Pat. Pub. US-20200234499-A1, herein after “Hwang”). In regard to claim 4, Canh in view of Ahmed teach [t]he apparatus of claim 3, wherein the multi-dimensional snapshot camera information box includes a plurality of camera extrinsic information, wherein each of the plurality of camera extrinsic information is associated with a respective one of the plurality of viewport information elements and indicates whether extrinsic camera information is present in an associated viewport information element “In a Presentation Engine, a view of the scene for a viewer may be determined by a camera used for rendering the scene, e.g., from the viewer’s viewpoint. In a gITF file, there may be configuration(s)/indication(s) (e.g., provision(s)) for signaling the specific extrinsic (e.g., a pose (e.g., position and orientation) of the camera) and/or intrinsic camera properties (e.g., projection matrix information of the camera)” (Ahmed, ¶ [0034]) where extrinsic properties are indicated in an information matrix of the camera, if a respective one of the plurality of camera extrinsic properties indicates that extrinsic camera information is present in the respective one of the plurality of viewport information elements associated with the respective one of the plurality of camera extrinsic properties, the multi-dimensional snapshot camera information box further includes extrinsic camera information associated with the respective one camera extrinsic property “A scene may be rendered from the camera object, e.g., after composition. One or more parameters associated with the camera object may provide detailed information about the pose of the camera (e.g., using extrinsic parameters) and the viewing volume (e.g., using intrinsic parameters)” (Ahmed, ¶ [0035]) where the pose of the camera may be included in the camera information and is read as extrinsic information. The motivation to combine Ahmed with Canh would be to include the extrinsic information detected in the information box for the user. Canh in view of Ahmed fail to explicitly teach the number of the plurality of camera extrinsic properties is the same as the number of the number of subsamples. Hu teaches the number of the plurality of camera extrinsic properties is the same as the number of the number of subsamples “the extrinsic camera information parameters described by ExtCameraInfoStruct are expected to be present in every sample” (Hu, Page 15) where extrinsic information is present in each sample. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of encoding and decoding volumetric images and extrinsic parameters taught by Canh in view of Ahmed, with the use of extrinsic parameters in each sample taught by Hu to have an equal number of samples and extrinsic parameters. The motivation to do so would be to note that each camera has extrinsic information. Canh in view of Ahmed and Hu fail to teach the use of extrinsic flags. Hwang teaches the use of extrinsic flags “The ECF (Extrinsic parameters Control Flag) field included in the bit 2 of the byte #14 of AR display mode InfoFrame may include information on whether an image displayed through AR glass is an image subjected to camera calibration based on an extrinsic parameter of at least one camera” (Hwang, ¶ [0339]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of encoding and decoding volumetric images and extrinsic parameters taught by Canh in view of Ahmed and Hu, with the use flags taught by Hwang to have an identifier for the extrinsic or intrinsic presence. The motivation to do so would be to have a positive or negative value to output to a user. In regard to claim 5, Canh in view of Ahmed teach [t]he apparatus of claim 3, wherein the multi-dimensional snapshot camera information box includes a plurality of camera intrinsic properties, wherein each of the plurality of camera intrinsic properties is associated with a respective one of the plurality of viewport information elements and indicates whether intrinsic camera information is present in an associated viewport information element “In a Presentation Engine, a view of the scene for a viewer may be determined by a camera used for rendering the scene, e.g., from the viewer’s viewpoint. In a gITF file, there may be configuration(s)/indication(s) (e.g., provision(s)) for signaling the specific extrinsic (e.g., a pose (e.g., position and orientation) of the camera) and/or intrinsic camera properties (e.g., projection matrix information of the camera)” (Ahmed, ¶ [0034]) where intrinsic properties are indicated in an information matrix of the camera, if a respective one of the plurality of camera intrinsic properties indicates that intrinsic camera information is present in the respective one of the plurality of viewport information elements associated with the respective one of the plurality of camera intrinsic properties, the multi-dimensional snapshot camera information box further includes intrinsic camera information associated with the respective one camera intrinsic property “A scene may be rendered from the camera object, e.g., after composition. One or more parameters associated with the camera object may provide detailed information about the pose of the camera (e.g., using extrinsic parameters) and the viewing volume (e.g., using intrinsic parameters)” (Ahmed, ¶ [0035]) where the viewing volume of the camera may be included in the camera information and is read as intrinsic information. The motivation to combine Ahmed with Canh would be to include the intrinsic information detected in the information box for the user. Canh in view of Ahmed fail to explicitly teach the number of the plurality of camera intrinsic properties is the same as the number of the number of subsamples. Hu teaches the number of the plurality of camera intrinsic properties is the same as the number of the number of subsamples “camera_intrinsic_flag[i]: Equal to 1 indicates that the intrinsic camera parameters exist in the ith viewport of the current sample. It should be equal to 0 if dynamic_int_camera_flag[i] is equal to 0” (Hu, Page 15) where each sample has a flag value for intrinsic information. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of encoding and decoding volumetric images and extrinsic parameters taught by Canh in view of Ahmed, with the use of intrinsic flags in each sample taught by Hu to have an equal number of samples and intrinsic flags. The motivation to do so would be to allow for each sample to have a positive or negative value for intrinsic information. Canh in view of Ahmed and Hu fail to teach the use of intrinsic flags. Hwang teaches the use of intrinsic flags “The ICF (Intrinsic parameters Control Flag) field included in a bit 1 of the byte #14 of AR display mode InfoFrame may include information on whether an image displayed through AR glass is an image subjected to camera calibration based on an intrinsic parameter of at least one camera” (Hwang, ¶ [0338]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of encoding and decoding volumetric images and extrinsic parameters taught by Canh in view of Ahmed and Hu, with the use flags taught by Hwang to have an identifier for the extrinsic or intrinsic presence. The motivation to do so would be to have a positive or negative value to output to a user. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Canh in view of Ahmed, and further in view of Hu. In regard to claim 9, Canh in view of Ahmed teach [t]he apparatus of claim 7. Canh in view of Ahmed fail to explicitly teach wherein the set of the location information and the direction information of cameras for the plurality of sub-samples remain the same in a single track. Hu teaches wherein the set of the location information and the direction information of cameras for the plurality of sub-samples remain the same in a single track “For viewport type equal to 3, the timing metadata indicates the recommended initial viewport information when playing the associated V3C media track, consisting of the initial viewport position and rotation. When intending to use another viewport to start playing the media track, the initial viewport position (cam_pos_x, cam_pos_y, cam_pos_z) is equal to (0,0,0)” (Hu, Page 15) where each sample in a track should begin with the initial viewport position and rotation. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of encoding and decoding volumetric images and location information taught by Canh in view of Ahmed, with the location information remaining the same for a single track taught by Hu to keep each sample camera location in the track identical. The motivation to do so would be to compare the camera views on the object or to start from the same location. Claims 11-12 & 15-17 are rejected under 35 U.S.C. 103 as being unpatentable over Canh in view of Hu, and further in view of Ahmed. In regard to claim 11, Canh in view of Hu teach [t]he apparatus of claim 10. Canh in view of Hu fail to teach wherein the processor is further configured to: generate a multi-dimensional snapshot camera information box, wherein the multi-dimensional snapshot camera information box includes a plurality of viewport information elements, each of the plurality of viewport information elements is associated with a respective one of the plurality of sub-samples, and a respective one of the plurality of viewport information elements provides camera information for an associated sub-sample; and transmit the multi-dimensional snapshot camera information box. Ahmed teaches wherein the processor is further configured to: generate a multi-dimensional snapshot camera information box, wherein the multi-dimensional snapshot camera information box includes a plurality of viewport information elements “The recommended viewport information may provide (e.g., dynamically) changing information, which may include translation and/or rotation of the node that includes the camera object and/or the intrinsic camera parameter of the camera object” (Ahmed, ¶ [0032]) where information is provided to the client regarding the viewport such as parameters or object information. Additionally, the media content received is 3 dimensional (Ahmed, ¶ [0010]), each of the plurality of viewport information elements is associated with a respective one of the plurality of sub-samples “Information about the camera used to provide a viewer’s viewport of the scene (e.g., including intrinsic and extrinsic camera parameters) may be required by the MAF to identify and/or to request the appropriate media for a (e.g., each) viewer of the scene” (Ahmed, ¶ [0037]) where information elements are for each viewer of the scene, which is read as a sub-sample, additionally “the presentation engine may send an update request comprising updated information associated with the first camera object and updated information associated with the second camera object” (Ahmed, ¶ [0074]) where the information is associated for each camera object, based on the camera, and a respective one of the plurality of viewport information elements provides camera information for an associated sub-sample “Information on the intersection of the bounding box with the view frustum may be used to determine which subset(s) of the media is/are in the camera’s view” (Ahmed, ¶ [0076]) where information can be about a specific sample or portion, which is read as a sub-sample, and can be used to determine what is in view or out of view; and transmit the multi-dimensional snapshot camera information box “A media content may be enclosed in a Bounding Box. Information on the intersection of the bounding box with the view frustum may be used to determine which subset(s) of the media is/are in the camera’s view” (Ahmed, ¶ [0076]) where information based on what is in the view of the projected image is taught, and “A Media Client may use the inference regarding the relevant subset of the media information which may be visible in the camera viewport and may make an informed request(s) of relevant subset of the media information” (Ahmed, ¶ [0079]) where the information may be transmitted to the client via the camera viewport. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of encoding and decoding volumetric images and a media track to manage the samples taught by Canh in view of Hu, with the method of information counting the number of sub-samples in the sample taught by Ahmed to allow the user to automatically see the sub-sample count. The motivation to do so would be to save time counting the sub-samples. In regard to claim 12, Canh in view of Hu and Ahmed teach [t]he apparatus of claim 11. Canh in view of Hu fail to teach wherein the multi-dimensional snapshot camera information box includes number information indicating the number of sub-samples in the respective one of the one or more samples. Ahmed teaches wherein the multi-dimensional snapshot camera information box includes number information indicating the number of sub-samples in the respective one of the one or more samples “The MAF may consume the Camera and Viewinfo arguments (e.g., as shown in Table 9) provided in the updateViewQ function. The Camera argument may provide information on a number of cameras being used for different viewers of the scene” (Ahmed, ¶ [0075]) where the number of cameras used varies based on viewers, and each camera view can be read as a sub-sample of the sample views. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of encoding and decoding volumetric images and a media track to manage the samples taught by Canh in view of Hu, with the method of information counting the number of sub-samples in the sample taught by Ahmed to allow the user to automatically see the sub-sample count. The motivation to do so would be to save time counting the sub-samples. In regard to claim 15, Canh in view of Hu teach [t]he apparatus of claim 10. Canh in view of Hu fail to teach wherein the respective one of the plurality of viewport information elements includes location information and direction information of a camera used to render a two-dimensional projected images in an associated sub-sample of the plurality of subsamples. Ahmed teaches wherein the respective one of the plurality of viewport information elements includes location information and direction information of a camera used to render a two-dimensional projected images in an associated sub-sample of the plurality of subsamples “Information described herein (e.g., the positions of the object, pose information of the viewer, and/or the transform of the position and orientation of the object) may use the scene’s coordinate system” (Ahmed, ¶ [0029]) and, “A Media Client may use the inference regarding the relevant subset of the media information which may be visible in the camera viewport and may make an informed request(s) of relevant subset of the media information” (Ahmed, ¶ [0034]) where the camera viewport information includes position, orientation, and pose in a coordinate system. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of encoding and decoding volumetric images and a media track to manage the samples taught by Canh in view of Hu, with the method of location information taught by Ahmed to include the camera position in the viewport information. The motivation to do so would be to track and aid in development with the photo based on its information. In regard to claim 16, Canh in view of Hu and Ahmed teach [t]he apparatus of claim 15, wherein location information and direction information of a camera for a sub-sample of the plurality of subsamples are different from location information and direction information of a camera for another sub-sample of the plurality of subsamples “A pose of a viewer may be given (e.g., directly or indirectly) by the extrinsic properties of a rendering camera, which may be associated with the viewer and may be used by a presentation client. In many applications there may be multiple rendering cameras used by a rendering client (e.g., for different viewers). For example, each viewer may have a different set of intrinsic properties for their respective camera” (Ahmed, ¶ [0041]) and “Multiple cameras may be used to render and/or compose different viewpoints for different viewers of the same scene” (Ahmed, ¶ [0036]) where different views or cameras may have different viewpoints of the scene, which is read as having a different location information for the viewing of the sub-samples. In regard to claim 17, Canh in view of Hu and Ahmed teach [t]he apparatus of claim 15, wherein the set of the location information and the direction information of cameras for the plurality of sub-samples remain the same in a single track “For viewport_type equal to 3, the timing metadata indicates the recommended initial viewport information when playing the associated V3C media track, consisting of the initial viewport position and rotation. When intending to use another viewport to start playing the media track, the initial viewport position (cam_pos_x, cam_pos_y, cam_pos_z) is equal to (0,0,0)” (Hu, Page 15) where each sample in a track should begin with the initial viewport position and rotation. Claims 13 & 14 are rejected under 35 U.S.C. 103 as being unpatentable over Canh in view of Hu, Ahmed, and further in view of Hwang. In regard to claim 13, Canh in view of Hu and Ahmed teach [t]he apparatus of claim 12, wherein the multi-dimensional snapshot camera information box includes a plurality of camera extrinsic properties, wherein each of the plurality of camera extrinsic properties is associated with a respective one of the plurality of viewport information elements and indicates whether extrinsic camera information is present in an associated viewport information element “In a Presentation Engine, a view of the scene for a viewer may be determined by a camera used for rendering the scene, e.g., from the viewer’s viewpoint. In a gITF file, there may be configuration(s)/indication(s) (e.g., provision(s)) for signaling the specific extrinsic (e.g., a pose (e.g., position and orientation) of the camera) and/or intrinsic camera properties (e.g., projection matrix information of the camera)” (Ahmed, ¶ [0034]) where extrinsic properties are indicated in an information matrix of the camera, if a respective one of the plurality of camera extrinsic property indicates that extrinsic camera information is present in the respective one of the plurality of viewport information elements associated with the respective one of the plurality of camera extrinsic properties, the multi-dimensional snapshot camera information box further includes extrinsic camera information associated with the respective one camera extrinsic property “A scene may be rendered from the camera object, e.g., after composition. One or more parameters associated with the camera object may provide detailed information about the pose of the camera (e.g., using extrinsic parameters) and the viewing volume (e.g., using intrinsic parameters)” (Ahmed, ¶ [0035]) where the pose of the camera may be included in the camera information and is read as extrinsic information. Canh in view of Ahmed fail to explicitly teach the number of the plurality of camera extrinsic properties is the same as the number of the number of subsamples. Hu teaches the number of the plurality of camera extrinsic properties is the same as the number of the number of subsamples “the extrinsic camera information parameters described by ExtCameraInfoStruct are expected to be present in every sample” (Hu, Page 15) where extrinsic information is present in each sample. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of encoding and decoding volumetric images and extrinsic parameters taught by Canh in view of Ahmed, with the use of extrinsic parameters in each sample taught by Hu to have an equal number of samples and extrinsic parameters. The motivation to do so would be to note that each camera has extrinsic information. Canh in view of Hu and Ahmed fail to teach the use of extrinsic flags. Hwang teaches the use of extrinsic flags “The ECF (Extrinsic parameters Control Flag) field included in the bit 2 of the byte #14 of AR display mode InfoFrame may include information on whether an image displayed through AR glass is an image subjected to camera calibration based on an extrinsic parameter of at least one camera” (Hwang, ¶ [0339]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of encoding and decoding volumetric images and extrinsic parameters taught by Canh in view of Hu and Ahmed, with the use flags taught by Hwang to have an identifier for the extrinsic or intrinsic presence. The motivation to do so would be to have a positive or negative value to output to a user. In regard to claim 14, Canh in view of Hu and Ahmed teach [t]he apparatus of claim 12, wherein the multi-dimensional snapshot camera information box includes a plurality of camera intrinsic properties, wherein each of the plurality of camera intrinsic f properties lags is associated with a respective one of the plurality of viewport information elements and indicates whether intrinsic camera information is present in an associated viewport information element “In a Presentation Engine, a view of the scene for a viewer may be determined by a camera used for rendering the scene, e.g., from the viewer’s viewpoint. In a gITF file, there may be configuration(s)/indication(s) (e.g., provision(s)) for signaling the specific extrinsic (e.g., a pose (e.g., position and orientation) of the camera) and/or intrinsic camera properties (e.g., projection matrix information of the camera)” (Ahmed, ¶ [0034]) where intrinsic properties are indicated in an information matrix of the camera, if a respective one of the plurality of camera intrinsic property indicates that intrinsic camera information is present in the respective one of the plurality of viewport information elements associated with the respective one of the plurality of camera intrinsic properties, the multi-dimensional snapshot camera information box further includes intrinsic camera information associated with the respective one camera intrinsic property “A scene may be rendered from the camera object, e.g., after composition. One or more parameters associated with the camera object may provide detailed information about the pose of the camera (e.g., using extrinsic parameters) and the viewing volume (e.g., using intrinsic parameters)” (Ahmed, ¶ [0035]) where the viewing volume of the camera may be included in the camera information and is read as intrinsic information. Canh in view of Ahmed fail to explicitly teach the number of the plurality of camera intrinsic properties is the same as the number of the number of subsamples. Hu teaches the number of the plurality of camera intrinsic properties is the same as the number of the number of subsamples “camera_intrinsic_flag[i]: Equal to 1 indicates that the intrinsic camera parameters exist in the ith viewport of the current sample. It should be equal to 0 if dynamic_int_camera_flag[i] is equal to 0” (Hu, Page 15) where each sample has a flag value for intrinsic information. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of encoding and decoding volumetric images and extrinsic parameters taught by Canh in view of Ahmed, with the use of intrinsic flags in each sample taught by Hu to have an equal number of samples and intrinsic flags. The motivation to do so would be to allow for each sample to have a positive or negative value for intrinsic information. Canh in view of Hu and Ahmed fail to teach the use of intrinsic flags. Hwang teaches the use of intrinsic flags “The ICF (Intrinsic parameters Control Flag) field included in a bit 1 of the byte #14 of AR display mode InfoFrame may include information on whether an image displayed through AR glass is an image subjected to camera calibration based on an intrinsic parameter of at least one camera” (Hwang, ¶ [0338]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of encoding and decoding volumetric images and extrinsic parameters taught by Canh in view of Hu and Ahmed, with the use flags taught by Hwang to have an identifier for the extrinsic or intrinsic presence. The motivation to do so would be to have a positive or negative value to output to a user. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Ilola Lauri et al. (Pat. Pub. WO-2023175243-A1) teaches encoding and decoding of volumetric video and a track of sample images with atlas data (abstract). Any inquiry concerning this communication or earlier communications from the examiner should be directed to CAIDEN ALEXANDER USSERY whose telephone number is (571)272-1192. The examiner can normally be reached Monday - Friday* 7:30AM - 5PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tammy Goddard can be reached at (571) 272-7773. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /C.A.U./Examiner, Art Unit 2611 /DAVID T WELCH/Primary Examiner, Art Unit 2613
Read full office action

Prosecution Timeline

Apr 10, 2025
Application Filed
Sep 15, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749265
INVARIANT SPECTRAL MARKERS PROVIDING STEADY REFERENCE AND METHOD FOR USING THE SAME
2y 2m to grant Granted Sep 29, 2026
Study what changed to get past this examiner. Based on 1 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
50%
Grant Probability
99%
With Interview (+100.0%)
1y 3m (~0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 2 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month