Prosecution Insights
Last updated: October 02, 2026
Application No. 19/164,516

TRANSPORTING MULTIMEDIA IMMERSION AND INTERACTION DATA IN A WIRELESS COMMUNICATION SYSTEM

Non-Final OA §102§103§112
Filed
Sep 11, 2025
Priority
Mar 15, 2023 — GR 20230100212 +1 more
Examiner
SIANGCHIN, KEVIN
Art Unit
2486
Tech Center
2400 — Computer Networks
Assignee
Lenovo (United States) Inc.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-58.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
7 currently pending
Career history
10
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§102 §103 §112
DETAILED ACTION The communication is in response to the application received September 11, 2025, wherein: claims 1-3, 5-13, 15, 17-19, 21, and 31 are amended; claims 16, 20, and 22-30 are cancelled; and Accordingly, claims 1-15, 17-19, 21, and 31 are pending and are examined as follows. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Acknowledgment is made of applicant's claim for foreign priority based on an application filed in Greece on March 15, 2023. Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. Information Disclosure Statement The information disclosure statement (IDS) was submitted on September 11, 2025. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Objections Claims 7, 9, 17, and 19 are objected to because of the following informalities. In claims 7 and 17, the Applicant recites transport protocols, including: a real-time protocol (RTP) and a secure real-time protocol (SRTP). Typically, RTP stands for Real-time Transport Protocol and SRTP stands for Secure Real-time Transport Protocol. The Applicant makes use of the standard terminology for RTP and SRTP in the disclosure. Please refer to ¶ [0071] of the Applicant’s specification, for example. Claims 7 and 17 should be amended to reflect this usage. Claim 9, as amended, recites: “peripherals selected peripherals” – i.e. “peripherals selected from a of peripherals”. For the purposes of examination, the Examiner assumes that the Applicant intended this portion of the claim to instead read as follows: “peripherals selected from a list of peripherals”. In claim 19, ‘RGB’, ‘RGBD’, and ‘IR’ should be replaced with (RGB), (RGBD), and (IR), respectively. In doing so, the language of claim 19 would match the corresponding recitation of these elements in claim 9. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention Claim 6 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. REGARDING CLAIM 6, note that claim 6 relies on a selection of at least one of the video codec specifications recited in claim 5. However, after a particular video codec is selected from the recited list, claim 6 nevertheless sets forth limitations directed at both the listed H.26x family of video codec specifications and the AV1 video codec specifications. Consequently, it is unclear which limitations govern the selected video codec specification. Moreover, the codec-specific limitations of claim 6 may be incompatible with the video codec specification(s) selected from claim 5. Therefore, the scope of claim 6 cannot be ascertained with sufficient certainty. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1, 5, 7, 9, 11, 15, 17, 19, 21, and 31 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Oyman et al. (U.S. Patent Application Publication No. US 2022/0021723 A1). REGARDING CLAIM 1, Oyman et al. discloses an apparatus for wireless communication (e.g. communication device 200 in Oyman et al. FIG. 2 and/or user equipment [UE] 101 and UE 102; see also ¶ [0027] lines 1-9 and ¶ [0054] lines 1-10) in a wireless communication system (e.g. network 140A in Oyman et al. FIG. 1, which is wireless – cf. ¶ [0026] and ¶ [0063]), comprising: at least one memory (e.g. Main Memory 204, Static Memory 206, and/or Storage Device 216 in Oyman et al. FIG. 2); and at least one processor coupled with the at least one memory and (e.g. Processor 202 in Oyman et al. FIG. 2, noting that Processor 202 is depicted as coupled to Main Memory 204, Static Memory 206, and/or Storage Device 216), configured to cause the apparatus to: generate, using one or more media sources, one or more data units of multimedia immersion and interaction data (cf. ¶ [0082], ¶ [0067], ¶ [0096], ¶ [0138], and ¶ [0212]-[0217], noting that viewport information [ ¶ [0096] lines 20-21] – including, or in addition to, pose information such as head, eye, and gaze positions and orientation [ ¶ [0212]-[0217] ] – constitutes one or more data units of multimedia immersion [cf. ¶ [0138] line 5] and interaction data, obtained or generated by one or more media sources, such as sensor(s) [ ¶ [0096] line 20 and ¶ [0212] ], head-mounted displays, mobile phones, and/or cameras [cf. ¶ [0067] ]) ; encode, using a video codec, an encoded video stream, wherein the encoded video streamincludes the one or more data units of multimedia immersion and interaction data as non-video coded embedded metadata (cf. Oyman et al. ¶ [0076], ¶ [0096], ¶ [0098], ¶ [0108], ¶ [0110], ¶ [0129], ¶ [0138], ¶ [0200], ¶ [0204]-[0205], and ¶ [0212]-[0217], noting that encoding [¶ [0076] lines 1-3] and decoding of video is performed in accordance with the H.265, HEVC, video codec standard. See, for example, ¶ [0098], ¶ [0108], ¶ [0110], ¶ [0129], ¶ [0138], ¶ [0200], and ¶ [0204]-[0205]. Further note that the one or more data units of multimedia immersion and interaction data – i.e. viewport information [ ¶ [0096] lines 20-21], including, or in addition to, pose information such as head, eye, and gaze positions and orientation [ ¶ [0212]-[0217] ] – are carried in the HEVC-encoded video stream as Supplemental Enhancement Information [SEI] messages. See, for example, ¶ [0096] lines 12-15, ¶ [0098], ¶ [0129] lines 4-5, ¶ [0138] lines 5-6, ¶ [0204] lines 3-6, and ¶ [0212]-[0217], noting that the contents of such SEI messages – “immersive media metadata” [¶ [0138] lines 4-6], “Decoder Metadata” [¶ [0096] lines 12-15 and FIG. 6], and/or “embedded viewport metadata” [ ¶ [0082] ] – constitute non-video coded embedded metadata representing the aforesaid one or more data units of multimedia immersion and interaction data.). and transmit, using a real-time transport protocol, the encoded video stream (cf. Oyman et al. FIG. 6, ¶ [0054], and ¶ [0096], noting that transmission [¶ [0054] lines 10-13] of the HEVC video bitstream occurs, as part of a Real-Time Transport Protocol [RTP] stream [cf. FIG. 6], transported over the network using RTP. See, for example, ¶ [0096].). REGARDING CLAIM 5, note that, for the purposes of examination, the Examiner interprets “wherein the video codec includes a video codec selected from a list of video codecs, including … ” as setting forth a disjunctive list of limitations pertaining to alternative video codecs, such that the scope of the claim encompasses any video codec that satisfies at least one of the listed limitations. As shown above, Oyman et al. teaches all limitations of claim 1. Oyman et al. further teaches that the video codec is: the H.265 video codec specification (cf. Oyman et al. ¶ [0098], ¶ [0108], ¶ [0110], ¶ [0129], ¶ [0138], ¶ [0200], and ¶ [0204]-[0205], noting that Oyman et al. uses H.265 and HEVC, interchangeably throughout the disclosure.). REGARDING CLAIM 7, note that, for the purposes of examination, the Examiner interprets “wherein the real-time transport protocol includes a transport protocol selected from a list of transport protocols, including” as setting forth a disjunctive list of limitations pertaining to alternative transport protocols, such that the scope of the claim encompasses any transport protocol that satisfies at least one of the listed limitations. As shown above, Oyman et al. teaches all limitations of claim 1. Oyman et al. further teaches that the transport protocol is: a real-time protocol(RTP) (cf. Oyman et al. FIG. 6, Abstract, ¶ [0096] lines 1-10, and ¶ [0098]). REGARDING CLAIM 9, please refer to the discussion above in the Claim Objections, regarding how the Examiner interprets the apparent typographical error in the currently amended claim. Further note that, for the purposes of examination, the Examiner interprets “wherein the one or more media sources comprise peripherals selected from a list of peripherals, including” as setting forth a disjunctive list of limitations pertaining to alternative peripherals, such that the scope of the claim encompasses any peripheral(s) that satisfies at least one of the listed limitations. As shown above, Oyman et al. teaches all limitations of claim 1. Oyman et al. further teaches that the one or more media sources comprise: one or more red green blue cf. Oyman et al. FIG. 3 and ¶ [0067], where she “[uses] a Head Mounted Display (HMD) and a camera that captures her video … a 360-degree view of the conference room” and “[h]e also has a 360-degree view of the conference room on his mobile screen and uses his mobile camera for capturing his own video”.). REGARDING CLAIM 11, Oyman et al. discloses an apparatus for wireless communication e.g. communication device 200 in Oyman et al. FIG. 2 and/or user equipment [UE] 101 and UE 102; see also ¶ [0027] lines 1-9 and ¶ [0054] lines 1-10) in a wireless communication system (e.g. network 140A in Oyman et al. FIG. 1, which is wireless – cf. ¶ [0026] and ¶ [0063]), comprising: at least one memory (e.g. Main Memory 204, Static Memory 206, and/or Storage Device 216 in Oyman et al. FIG. 2); andat least one processor coupled with the at least one memory (e.g. Processor 202 in Oyman et al. FIG. 2, noting that Processor 202 is depicted as coupled to Main Memory 204, Static Memory 206, and/or Storage Device 216) and configured to cause the apparatus to: receive, using a real-time transport protocol, an encoded video stream, wherein the encoded video stream cf. Oyman et al. FIG. 6, ¶ [0012], and ¶ [0098], noting that FIG. 6 illustrates a “receiver architecture” [ cf. ¶ [0012] ], which receives an encoded video stream [e.g. RTP stream and/or Elementary stream in FIG. 6] containing an HEVC [i.e. video codec encoded] bitstream. See, for example, ¶ [0098]), andincludes one or more data units of multimedia immersion and interaction data generated by one or more media sources as non-video coded embedded metadata (cf. ¶ [0082], ¶ [0067], ¶ [0096], ¶ [0138], and ¶ [0212]-[0217], noting that viewport information [ ¶ [0096] lines 20-21] – including, or in addition to, pose information such as head, eye, and gaze positions and orientation [ ¶ [0212]-[0217] ] – constitutes one or more data units of multimedia immersion [cf. ¶ [0138] line 5] and interaction data, obtained or generated by one or media sources, such as sensor(s) [ ¶ [0096] line 20 and ¶ [0212] ], head-mounted displays, mobile phones, and/or cameras [cf. ¶ [0067] ]) ; decode, using the video codec, the encoded video streamto extract the non-video coded embedded metadata (cf. Oyman et al. FIG. 6 and ¶ [0204], noting that the HEVC-encoded Elementary stream is fed to the HEVC Decoder, where Decoder Metadata [i.e. non-video coded embedded metadata], containing Supplemental Enhancement Information, is decoded and extracted. See, for example, ¶ [0204].); and consumecf. Oyman et al. FIG. 6 and ¶ [0218], noting that “[t]he VR Renderer uses”, or consumes, “the decoded signals and rendering metadata, together with the pose and the knowledge of the horizontal/vertical field of view, to determine a viewport and render the appropriate part of the video and audio signals”, where “pose and the knowledge of the horizontal/vertical field of view”, indicated by the extracted metadata, constitute the one or more data units of multimedia immersion and interaction data, as discussed above. See, for example, ¶ [0218] lines 7-11.). REGARDING CLAIM 15, as shown above, Oyman et al. teaches all limitations of claim 11. As shown above, with respect to the rejection of claim 5, Oyman et al. further discloses limitations that are substantially identical to those recited in claim 15. As such, the rationales provided above in the rejection of claim 5 are applicable to the corresponding limitations of claim 15. Therefore, Oyman et al. anticipates the apparatus set forth in claim 15, for the same reasons articulated above with respect to claim 5. REGARDING CLAIM 17, as shown above, Oyman et al. teaches all limitations of claim 11. As shown above, with respect to the rejection of claim 7, Oyman et al. further discloses limitations that are substantially identical to those recited in claim 17. As such, the rationales provided above in the rejection of claim 7 are applicable to the corresponding limitations of claim 17. Therefore, Oyman et al. anticipates the apparatus set forth in claim 17, for the same reasons articulated above with respect to claim 7. REGARDING CLAIM 19, as shown above, Oyman et al. teaches all limitations of claim 11. As shown above, with respect to the rejection of claim 9, Oyman et al. further discloses limitations that are substantially identical to those recited in claim 19. As such, the rationales provided above in the rejection of claim 9 are applicable to the corresponding limitations of claim 19. Therefore, Oyman et al. anticipates the method set forth in claim 19, for the same reasons articulated above with respect to claim 9. REGARDING CLAIM 21, note that apparatus set forth in claim 1 is configured to perform functions that are essentially identical to the method steps recited in claim 21. As such, the rationales provided above in the rejection of claim 1 are applicable to the corresponding limitations of claim 21. Therefore, Oyman et al. anticipates the apparatus set forth in claim 21, for the same reasons articulated above with respect to claim 1. REGARDING CLAIM 31, note that apparatus set forth in claim 11 is configured to perform functions that are essentially identical to the method steps recited in claim 31. As such, the rationales provided above in the rejection of claim 11 are applicable to the corresponding limitations of claim 31. Therefore, Oyman et al. anticipates the method set forth in claim 31, for the same reasons articulated above with respect to claim 11. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: Determining the scope and contents of the prior art. Ascertaining the differences between the prior art and the claims at issue. Resolving the level of ordinary skill in the pertinent art. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 8 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Oyman et al. (U.S. Patent Application Publication No. US 2022/0021723 A1), in view of Yang et al. (U.S. Patent Application Publication No. US 2024/0414586 A1). REGARDING CLAIM 8, as shown above, Oyman et al. teaches all limitations of claim 1. Oyman et al. further teaches: the real-time transport protocol includes encryption (cf. Oyman et al. ¶ [0037], ¶ [0044]-[0045], ¶ [0047], ¶ [0096], and ¶ [0200], noting that “the UE parses, possibly decrypts and feeds the elementary stream to the HEVC decoder” [¶ [0096] lines 7-10] and that, in the case that the stream is decrypted, a prior encryption is necessitated.). However, Oyman et al. does not expressly teach that: the real-time transport protocol includes authentication. In contrast, Yang et al., from a similar field of endeavor (i.e. HEVC-based extended-reality [XR] media transport, using a real-time transport protocol, over 5G radio access networks [RAN] – cf. Yang et al. ¶ [0001] and ¶ [0064]) teaches: the real-time transport protocol includes authentication (cf. Yang et al. ¶ [0026] lines 7-17 and ¶ [0046]. Yang et al. teaches that RTP transport for HEVC and/or XR traffic may employ Secure RTP [SRTP] – cf. Yang et al. ¶ [0046] lines 1-9 and lines 12-13 – which is an extension of RTP, well-known in the art to provide encryption and authentication for transported media packets.). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to use SRTP transport, as taught by Yang et al., in the RTP-based immersive media system of Oyman et al., in order to provide the encryption and authentication of the transmitted HEVC-encoded media supported by SRTP (cf. Yang et al. ¶ [0046]). REGARDING CLAIM 18, as shown above, Oyman et al. teaches all limitations of claim 11. Note that the additional limitations set forth in claim 18 are substantially identical to those recited in claim 8. As such, the rationales provided above in the rejection of claim 8 are applicable to the corresponding limitations of claim 18. Therefore, the apparatus of claim 18 would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, in view of the teachings of Oyman et al. and Yang et al., for the same reasons articulated above with respect to claim 8. Claims 2-4, 6, and 12-14 are rejected under 35 U.S.C. 103 as being unpatentable over Oyman et al. (U.S. Patent Application Publication No. US 2022/0021723 A1), in view of Reitel et al. (U.S. Patent Application Publication No. US 2016/0212375 A1). REGARDING CLAIM 2, as shown above, Oyman et al. teaches all limitations of claim 1. However, Oyman et al. does not expressly teach that the non-video coded embedded metadata ncludes: a first field comprising an identifier for a syntax and semantics representation format of the one or more data units of the multimedia immersion and interaction data; and a second field comprising the one or more data units of the multimedia immersion and interaction data[[,]]and encoded according to the syntax and semantics representation format corresponding to the identifier of the first field. In contrast, Reitel et al., from a similar field of endeavor (i.e. encoded mixed-reality [MR] media communication architectures) teaches using: non-video coded embedded metadata (cf. Reitel et al. FIG. 9a, ¶ [0049]-[0050], and ¶ [0053], where the contents of the disclosed Supplemental Enhancement Information [SEI] message data [cf. ¶ [0049] lines 14-15 and ¶ [0050] lines 1-5] – e.g. “camera intrinsic and extrinsic values as custom attributes” [¶ [0053] line 8], as indicated by TLV tuples, uuid_iso_iec_11578 field, payloadSize, etc. [¶ [0053] – constitute non-video coded embedded metadata) that includes: a first field comprising an identifier for a syntax and semantics representation format of the one or more data units of the multimedia immersion and interaction data (cf. Reitel et al. ¶ [0053]. The SEI message shown in ¶ [0053] is structured in accordance with the syntax for User Data Unregistered SEI type messages defined in section 7.3.2.3.1 of the H.264 specification [ISO/IEC14496-10:2010]. See Reitel et al. ¶ [0053] lines 28-33. The first field, uuid_iso_11578, is a 128-bit “universally unique identifier (UUID)”, which is, in turn, associated with a payload – T, L, and V, indicating type, length, and value, respectively – that indicates pose and/or projection matrix information [i.e. one or more data units of the multimedia immersion and interaction data]. See Reitel et al. ¶ [0053] lines 36-47. This UUID [e.g. {0F5DD509-CF7E-4AC4-9E9A-406B68973C42} – ¶ [0053] lines 36-39] is distinguished from other uuid_iso_11578 in that it is associated with a specific syntax [i.e. the UUID identifying the payload format, indicated by T, L, and V] and semantics specifying how to interpret that syntax [cf. ¶ [0053] lines 40-47]); a second field comprising the one or more data units of the multimedia immersion and interaction data[[,]]and encoded according to the syntax and semantics representation format corresponding to the identifier of the first field (cf. Reitel et al. ¶ [0053]. The payload [i.e. T, L, and V] described above constitutes a second field comprising the one or more data units of the multimedia immersion and interaction data, encoded according to syntax and semantics associated with the UUID [cf. ¶ [0053] lines 40-47]). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the SEI messaging utilized in Oyman et al. to “encode the camera intrinsic and extrinsic data in the video channel and carry it in-band as SEI messages” (cf. Reitel et al. ¶ [0051]), as taught by Reitel et al., in order to facilitate the “synchronisation of objects within a shared scene, such as generated in collaborative mixed reality applications” (cf. Reitel et al. ¶ [0004] lines 1-4 and ¶ [0050] lines 10-12) and to “produce a reliable and effective mixed reality experience (cf. Reitel et al. ¶ [0054]). REGARDING CLAIM 3, Oyman et al. and Reitel et al., as combined in the manner discussed above, have been shown to teach or suggest all limitations of claim 2. Reitel et al. further teaches: the identifier is a universally unique identifier (UUID) cf. Reitel et al. ¶ [0053] lines 16-39. See also the discussion above with respect to claim 2). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the SEI messaging utilized in Oyman et al. to “encode the camera intrinsic and extrinsic data in the video channel and carry it in-band as SEI messages” (cf. Reitel et al. ¶ [0051]), as taught by Reitel et al., in order to facilitate the “synchronisation of objects within a shared scene, such as generated in collaborative mixed reality applications” (cf. Reitel et al. ¶ [0004] lines 1-4 and ¶ [0050] lines 10-12) and to “produce a reliable and effective mixed reality experience (cf. Reitel et al. ¶ [0054]). REGARDING CLAIM 4, Oyman et al. and Reitel et al., as combined in the manner discussed above, have been shown to teach or suggest all limitations of claim 3. Reitel et al. further teaches that: the UUID is unique to a specific application or session, or the UUID is globally unique (cf. Reitel et al. ¶ [0053] lines 36-39, noting that the UUID is a “universally unique identifier” which, in some embodiments, is set to a static value of {0F5DD509-CF7E-4AC4-9E9A-406B68973C42}, suggesting that the UUID is globally unique, as opposed to application or session-specific.) . Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the SEI messaging utilized in Oyman et al. to “encode the camera intrinsic and extrinsic data in the video channel and carry it in-band as SEI messages” (cf. Reitel et al. ¶ [0051]), as taught by Reitel et al., in order to facilitate the “synchronisation of objects within a shared scene, such as generated in collaborative mixed reality applications” (cf. Reitel et al. ¶ [0004] lines 1-4 and ¶ [0050] lines 10-12) and to “produce a reliable and effective mixed reality experience (cf. Reitel et al. ¶ [0054]). REGARDING CLAIM 6, as shown above, Oyman et al. teaches all limitations of claim 5. Oyman et al. further teaches: the video codec includes codec specifications (cf. Oyman et al. ¶ [0098], noting that the video codec used is H.265, HEVC) and wherein the at least one processor is configured to cause the apparatus to: encode the (SEI)messages (cf. ¶ [0138] lines 3-6, noting that “360 video transmission could be based on the RTP payload formats for HEVC that carry SEI messages describing immersive media metadata”); However, Oyman et al. does not expressly teach: encoding the video stream, by causing the apparatus to encapsulate the non-video coded embedded metadata as a payload of user data in one or more supplemental enhancement information (SEI) messages of the type 'user data unregistered'. In contrast, Reitel et al., from a similar field of endeavor (i.e. encoded mixed-reality [MR] media communication architectures) teaches: encoding the video stream, by causing the apparatus to encapsulate the non-video coded embedded metadata as a payload of user data in one or more supplemental enhancement information (SEI) messages of the type 'user data unregistered' (cf. Reitel et al. ¶ [0053] and ¶ [0096]. Note that pose information and camera intrinsic parameters [¶ [0053] lines 6-15] are indicated by the tuples of T, L, and V [i.e. non-video coded embedded metadata], which indicate type, length, and value, respectively, of the pose information or intrinsic parameters. See Reitel et al. ¶ [0053] lines 36-47. As shown in the table in ¶ [0053], these are the payload of payloadType 5, or of the type “User Data Unregistered”, associated with the uuid_iso_iec_11578 field. See Reitel et al. ¶ [0053] lines 16-35.). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the SEI messaging utilized in Oyman et al. to “encode the camera intrinsic and extrinsic data in the video channel and carry it in-band as SEI messages” (cf. Reitel et al. ¶ [0051]), as taught by Reitel et al., in order to facilitate the “synchronisation of objects within a shared scene, such as generated in collaborative mixed reality applications” (cf. Reitel et al. ¶ [0004] lines 1-4 and ¶ [0050] lines 10-12) and to “produce a reliable and effective mixed reality experience (cf. Reitel et al. ¶ [0054]). Combining the teachings of Oyman et al. and Reitel et al., in the manner just described, yields: a video codec including the H.264, H.265, or H.266 codec specifications, and wherein the at least one processor is configured to cause the apparatus to encode video stream, by causing the apparatus to encapsulate the non-video coded embedded metadata as a payload of user data in one or more supplemental enhancement information (SEI) messages of the type 'user data unregistered'. as required by claim 6. REGARDING CLAIM 12, as shown above, Oyman et al. teaches all limitations of claim 11. Note that the additional limitations set forth in claim 12 are substantially identical to those recited in claim 2. As such, the rationales provided above in the rejection of claim 2 are applicable to the corresponding limitations of claim 12. Therefore, the apparatus of claim 12 would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, in view of the teachings of Oyman et al. and Reitel et al., for the same reasons articulated above with respect to claim 2. REGARDING CLAIM 13, as shown above, Oyman et al. and Reitel et al., as combined in the manner discussed above, have been shown to teach or suggest all limitations of claim 12. Note that the additional limitations set forth in claim 13 are substantially identical to those recited in claim 3. As such, the rationales provided above in the rejection of claim 3 are applicable to the corresponding limitations of claim 13. Therefore, the apparatus of claim 13 would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, in view of the teachings of Oyman et al. and Reitel et al., for the same reasons articulated above with respect to claim 3. REGARDING CLAIM 14, as shown above, Oyman et al. and Reitel et al., as combined in the manner discussed above, have been shown to teach or suggest all limitations of claim 13. Note that the additional limitations set forth in claim 14 are substantially identical to those recited in claim 4. As such, the rationales provided above in the rejection of claim 4 are applicable to the corresponding limitations of claim 14. Therefore, the apparatus of claim 14 would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, in view of the teachings of Oyman et al. and Reitel et al., for the same reasons articulated above with respect to claim 4. Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Oyman et al. (U.S. Patent Application Publication No. US 2022/0021723 A1), in view of Thomas et al. (Thomas, E. and Teniou, G. "MeCAR Permanent Document v5.0," 3GPP TSG SA WG4#122; S4-230307; February 24, 2023). REGARDING CLAIM 10, as shown above, Oyman et al. teaches all limitations of claim 1. Oyman et al. further teaches that: the one or more data units of multimedia immersion and interaction data ncludes immersion and interaction data, including: a user viewpoint data (cf. Oyman et al. FIG. 3 through FIG. 5, ¶ [0096] lines 20-21, ¶ [0138] line 5, and ¶ [0218] lines 9-11,noting that a “viewport” constitutes a user viewpoint data). a user field of view data (cf. Oyman et al. FIG. 5, ¶ [0096] lines 20-21 and ¶ [0218] lines 9-11, and ¶ [0268] Table 9.3.4-1); a user pose/orientation data (cf. Oyman et al. ¶ [0212]-[0215] and ¶ [0218] lines 7-11, noting that head pose [ ¶ [0214] ] and gaze direction [ ¶ [0215] ] each or in conjunction constitute user pose/orientation data); However, Oyman et al. does not expressly teach that: the one or more data units of multimedia immersion and interaction data includes immersion and interaction data, including a user gesture tracking data; a user body tracking data; a user facial feature tracking data; a user action and/or user input data; a split rendering pose and spatial information; and an augmented reality object representationincluding at least one of a graphical description of an object and an object positional anchor. In contrast, Thomas et al., from a similar field of endeavor (i.e. XR/AR/MR and immersive media communication) teaches that: the one or more data units of multimedia immersion and interaction data (cf. Table 2 and Table 4 on Thomas et al. pages 18-19 and page 21, respectively, noting that the various interactions shown and defined correspond to “metadata representing the user inputs or haptics such as user pose, viewport, gesture, body action, facial expression, FOV and AR anchor point” [cf. Thomas et al. page 18 ¶ 1]. These constitute one or more data units of multimedia immersion and interaction data.) includes immersion and interaction data, including: a user viewpoint data (cf. Table 2 and Table 4: “Viewport” AR/MR Data Type); a user field of view data (cf. Table 2 and Table 4: “FOV” AR/MR Data Type); a user pose/orientation data (cf. Table 2 and Table 4: “User Pose” AR/MR Data Type); a user gesture tracking data (cf. Table 2 and Table 4: “Gesture” AR/MR Data Type); a user body tracking data (cf. Table 2 and Table 4: “Body Action” AR/MR Data Type); a user facial feature tracking data (cf. Table 2 and Table 4: “Facial Expression” AR/MR Data Type); a user action and/or user input data (cf. Table 2 and Table 4: “Sensor Information” AR/MR Data Type, noting that these entail extensions for “new controllers”, such as the hp_mixed_reality_controller indicated in the URL shown in column 4, row 14 of Table 2); a split rendering pose and spatial information (cf. Table 2 and Table 4: “Split Rendered 2D projection video” AR/MR Data Type, “Split Rendered 2D depth map video” AR/MR Data Type, and “Split Rendered 2D alpha channel video” AR/MR Data Type); an augmented reality object representation including at least one of a graphical description of an object (cf. Table 2 and Table 4: “3D Visual Model” AR/MR Data Type) and an object positional anchor (cf. Table 2 and Table 4: “AR Anchor” AR/MR Data Type). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the SEI/metadata messaging utilized in Oyman et al. to incorporate a broader array of metadata representing the multitude of XR/AR/MR user inputs, shown in Thomas et al., in order to capture a richer set of user interaction data for XR/AR/MR applications in the context of 5G system-based delivery, where “a set of well-defined media capabilities are essential to create the conditions of a successful ecosystem” (cf. Thomas et al. page 7, Introduction, ¶ 3). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Please see PTO 892 for additional references. Yip et al. U.S. Patent Application Publication No. US 2022/0366641 A1, METHOD AND APPARATUS FOR AR REMOTE RENDERING PROCESSES. Published November 17, 2022. Relevance: Yip et al. discloses an augmented reality (AR) remote rendering system that, inter alia, generates and transmits metadata associated with rendered video frames to support augmented reality rendering and synchronization between the remote renderer and the AR device. Yip et al. further discloses that the metadata associated with the rendered frame can include pose information, where the associated metadata may comprise video codec parameters typically carried with SEI messages with the context of RTP transport. Yip et al. additionally suggest metadata that indicate AR object anchors. Bouazizi et al. U.S. Patent Application Publication No. US 2024/0273833 A1, SIGNALING POSE INFORMATION TO A SPLIT RENDERING SERVER FOR AUGMENTED REALITY COMMUNICATION SESSIONS. Published November 17, 2022 Relevance: Bouazizi et al. discloses techniques for split rendering of XR data over 5G networks in which an XR server and client cooperatively render an XR scene. Bouazizi et al. describe the use of video codecs, such as AVC and HEVC and the processing of data, such as pose information, from XR-related sensors including cameras, microphones, and tracking systems. Oh et al. U.S. Patent No. US 11,528,509 B2. Video Transmission Method, Video Transmission Device, Video Receiving Method And Video Receiving Device. Published December 13, 2022. Relevance: Oh et al. teaches embedding immersive viewing metadata into the coded video bitstream for later extraction and use during rendering in XR applications, including metadata associated with the user’s viewing position and orientation. Additionally, Oh et al. teaches embedding head motion information and viewing position/orientation (pose) information in SEI messages, and describes immersive media generated using a variety of acquisition modalities, including multiple cameras, depth cameras, inertial sensors, and eye-tracking sensors. International Telecommunication Union. May 2022. “Recommendation ITU-T H.274 (V2): Versatile Supplemental Enhancement Information Messages for Coded Video Bitstreams. ITU-T Recommendation H.274 (05/2022). International Telecommunication Union, Geneva, Switzerland. Relevance: Section 8.4 defines the User Data Unregistered SEI message, including the syntax and semantic of the 16-byte UUID followed by user-defined payload data. E. Carrara, K. Norrman, D. McGrew, M. Naslund, and M. Baugher, “The Secure Real-time Transport Protocol (SRTP),” Request for Comments, no. 3711, 2004, doi: 10.17487/RFC3711. Relevance: This document describes the Secure Real-time Transport Protocol (SRTP), a profile of the Real-time Transport Protocol (RTP), which can provide confidentiality, message authentication, encryption, and replay protection to the RTP traffic and to the control traffic for RTP, the Real-time Transport Control Protocol (RTCP). 3rd Generation Partnership Project (3GPP). 2022. “Support of 5G Glass-Type Augmented Reality / Mixed Reality (AR/MR) Devices. TR 26.998 Version 17.0.0, Release 17. Technical Report. ETSI. Sophia Antipolis, France. Relevance: 3GPP describes a 5G architecture and media framework for supporting glass-type AR and MR devices including support for various AR/MR media formats, metadata, XR spatial descriptions, and network delivery of immersive content over 5G. 3GPP additionally identifies user pose information as metadata, as well as discussing support for WebRTC-based communication, immersive media formats such as 2D+Depth, and XR spatial descriptions which may include spatial anchors of AR/MR assets. Qualcomm Inc. and Lenovo. "Real-time metadata transport over RTP," 3GPP TSG SA WG4#121; Tdoc S4-221256; November 18, 2022. Relevance: Qualcomm Inc. and Lenovo discloses transporting real-time XR interaction metadata over RTP, utilizing RTP metadata structures, payload formats, synchronization attributes, and security functionality for real-time XR applications. Qualcomm Inc. and Lenovo expressly identifies numerous XR/AR/MR interaction metadata, including FoV, user pose, viewport, gestures, body actions, facial expressions, AR anchor points, hand tracking, palm pose, face tracking, and keyboard and controller inputs. This metadata is transported over RTP using SRTP, where SRTP provides integrity/authentication of RTP header extensions and RFC6904 provides confidentiality through encryption of selected RTP header extensions. Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEVIN SIANGCHIN whose telephone number is (571)270-0982. The examiner can normally be reached M, W-F 10:00am-06:00pm and Tu 09:00am-05:00pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jamie Atala can be reached at (571) 272-7384. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /K S/ Examiner Art Unit 2486 /JAMIE J ATALA/ Supervisory Patent Examiner, Art Unit 2486
Read full office action

Prosecution Timeline

Sep 11, 2025
Application Filed
Aug 10, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month