Prosecution Insights
Last updated: October 02, 2026
Application No. 18/698,598

INTERACTIVE ANCHORS IN AUGMENTED REALITY SCENE GRAPHS

Non-Final OA §103
Filed
Apr 04, 2024
Priority
Oct 06, 2021 — EU 21306409.0 +2 more
Examiner
GUO, XILIN
Art Unit
2616
Tech Center
2600 — Communications
Assignee
InterDigital Inc.
OA Round
3 (Non-Final)
82%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
387 granted / 474 resolved
+19.6% vs TC avg
Strong +19% interview lift
Without
With
+18.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 4m
Avg Prosecution
13 currently pending
Career history
489
Total Applications
across all art units

Statute-Specific Performance

§101
8.1%
-31.9% vs TC avg
§103
61.8%
+21.8% vs TC avg
§102
8.8%
-31.2% vs TC avg
§112
17.3%
-22.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 474 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on July 13, 2026 has been entered. Response to Amendment The amendment filed on July 13, 2026 has been entered. In view of the amendment to the claims, the amendment of claims 1, 6 and 15 have been acknowledged. Claims 9 and 18-21 have been canceled. New claims 22-23 have been added. Response to Arguments Applicant’s arguments, see pages 7-10 of Remarks, filed July 13, 2026 have been fully considered. But they are not persuasive. Regarding Claim 1, Applicants state in page 7 of Remarks that “Applicant respectfully traverses the obviousness rejection of Claims 1-8 and 10-17 over Bouazizi on the grounds that no prima facie evidence of obviousness was established for all limitations. Applicant appreciates the Examiner's response to previous arguments on pages 6-8 of the Final Office Action. Unfortunately, the Office's interpretation is too broad to be reasonable. For example, the Office Action states: a process to be performed by an augmented reality engine". The claim just simple recite "the action comprises a description of a process" and """an augmented reality engine" for performing "a description of a process". Bouazizi discloses the anchor point includes various actions that a user may perform, such as movements received via a controller or real-world repositioning of a device worn by the user (Paragraph [0157]). Under the broadest reasonable interpretation, "an augmented reality engine" can be interpreted as a generic computer component to perform "a description of a process". Thus, Bouazizi discloses "action, wherein the action comprises a description of a process to be performed by an augmented reality engine" recited in claim 1. Final Office Action at page 8. The Office gives absolutely no weight to "wherein the action comprises a description of a process to be performed by an augmented reality engine" (Claims 1 and 10, emphasis added) and yet all limitations must be interpreted, given weight, and demonstrated to be present in the reference”. Examiner replies: The examiner disagrees with Applicant’s premises and conclusion. Under the broadest reasonable interpretation, the examiner respectfully maintains that the prior art rejections in this case are proper for the following reasons. In respond to the applicant’s arguments, the examiner recites Bouazizi in order to disclose the issue. The claim recites “wherein the action comprises a description of a process to be performed by an augmented reality engine”. However, the claim just simply recite “an augmented reality engine” for performing the action and does not set forth any elements to describe the configuration, structure or definition of the augmented reality engine. Bouazizi discloses a client device (As shown in FIG. 3) to determine an anchor point from the scene description; anchor the virtual scene to the real-world presentation environment at the determined real-world anchor point and transform the virtual scene to align with the real-world presentation environment. Thus, the function performed by the client device taught by Bouazizi can be considered equivalent to “augmented reality engine”. Therefore, Bouazizi discloses “wherein the action comprises a description of a process to be performed by an augmented reality engine” and the above arguments (claim 10 has the same reasons). Regarding Claim 1, Applicants state in pages 7-8 of Remarks that “Applicant respectfully notes that there are other failures to demonstrate, by evidence, all limitations of the claims. For example, independent claims 1 and 10 recite "one or more anchors, wherein each anchor of the one or more anchors is associated with one or more nodes of the scene graph." Emphasis added. The Final Office Action states: one or more nodes of the scene graph". The prior art reference Bouazizi describes techniques for anchoring a scene description to a user environment (augmented reality environment) and the client device receives a scene description; the scene description includes a scene graph (Paragraphs [0092]-[0094] and [0156]) and anchor point data representing a correspondence between a virtual scene represented by the media data and a real-world presentation environment (Paragraph [0068]). Thus, Bouazizi discloses Final Office Action at pages 6-7 (rejecting at pages 12-13). Actually, Bouazizi's "scene anchor" is a singular, global construct for the initial placement of the entire scene - a mechanism for a one-time, global alignment. Bouazizi's stated purpose is to "describe the anchoring of the scene to a real-world XR space." Id. at para. [0134]; emphasis added. No rationale has been provided for Bouazizi disclosing a system of one or more anchors associated with individual nodes to provide object-level interactivity. A person of ordinary skill in the art would fail to identify any teaching in Bouazizi that would motivate them to use one or more anchors for node-specific control”. Examiner replies: The examiner disagrees with Applicant’s premises and conclusion. Under the broadest reasonable interpretation, the examiner respectfully maintains that the prior art rejections in this case are proper for the following reasons. In respond to the applicant’s arguments, the examiner recites Bouazizi in order to disclose the issue. The claim recites “one or more anchors, wherein each anchor of the one or more anchors is associated with one or more nodes of the scene graph”. However, the above limitation can be considered as “one ”. Bouazizi describes techniques for anchoring a scene description to a user environment (augmented reality environment) and the client device receives a scene description; the scene description includes a scene graph (Paragraphs [0092]-[0094] and [0156]) and anchor point data representing a correspondence between a virtual scene represented by the media data and a real-world presentation environment (Paragraph [0068]). Thus, Bouazizi discloses “one or more anchors, wherein each anchor of the one or more anchors is associated with one or more nodes of the scene graph” recited in claim 1 and the above arguments (claim 10 has the same reasons). Regarding Claim 1, Applicants state in pages 8-9 of Remarks that “Similarly, independent claims 1 and 10 require that "the trigger is activated when a condition... is detected in the real environment." Emphasis added. A person of ordinary skill in the art appreciates that this requires a dynamic, runtime detection of an event. However, the Office states: In additional, the claim 1 recites "a trigger, wherein the trigger is a description of at least one condition: wherein the at least one condition is a detection of a visual or audio or environment-based marker or property, and wherein the triqqer is activated when a condition of the at least one condition is detected in the real environment". Thus, the claim describes "a trigger is a description of at least one condition". Bouazizi discloses the scene description may include data for an anchor point and the data includes whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space (Paragraph [0157]). Thus, the reference space type is a condition for different environment (a view, local, stage, or application) and the data of an anchor is a description of at least one condition. Therefore, under broadest reasonable interpretation, "the data" described by Bouazizi can be considered equivalent to "a trigger". Thus, Bouazizi discloses "a trigger, Final Office Action at page 7 (rejecting at page 13); emphasis added. Again, the Office robs the claim of its meaning by an overly broad interpretation. The rejection incorrectly equates the claimed dynamic "trigger" with the "data" in Bouazizi that indicates "whether the reference space type... is to be a view, local, stage,or application type..." citing Bouazizi para. [0157]. This arises to the level of clear factual error. The "data" is merely a static configuration parameter within the scene description file - it is a pre-set value read once during initialization to determine how the scene should be placed. Bouazizi teaches that the presentation unit "may determine a type of anchor point ... from the data of scene graph 324." Id. at para. [0148]; emphasis added. A static setting in a file is not "detected" from the environment. It is a one-time lookup of a pre-set value. Also, the plain and ordinary meaning of a trigger goes against the Examiner's interpretation. Bouazizi's data is simply not a dynamic condition that is "detected in the real environment" at runtime to activate an event. A broadest reasonable interpretation does not permit an interpretation that is contrary to reason - a static configuration parameter and an event-driven trigger are fundamentally different concepts”. Examiner replies: The examiner disagrees with Applicant’s premises and conclusion. Under the broadest reasonable interpretation, the examiner respectfully maintains that the prior art rejections in this case are proper for the following reasons. In respond to the applicant’s arguments, the examiner recites Bouazizi in order to disclose the issue. The claim recites “a trigger, wherein the trigger is a description of at least one condition; wherein the at least one condition is a detection of a visual or audio or environment-based marker or property, and wherein the trigger is activated when a condition of the at least one condition is detected in the real environment”. Thus, the claim describes “a trigger is a description of at least one condition”. The examiner interpreted “a trigger” as a description of one condition in the previous office action. Bouazizi discloses the scene description may include data for an anchor point and the data includes whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space (Paragraph [0157]). Thus, the reference space type is a condition for different environment (a view, local, stage, or application) and the data of an anchor is a description of at least one condition. Therefore, under broadest reasonable interpretation, “the data” described by Bouazizi can be considered equivalent to “a trigger” and “the data” is activated when one of the reference space types is detected in the real-world environment (Paragraph [0157], ... such as movements received via a controller or real-world repositioning of a device worn by the user ...). As shown in FIG. 9 of Bouazizi, the anchor XR space is of type “stage” corresponding to the floor of the viewer's living room. Accordingly, Bouazizi discloses “a trigger, wherein the trigger is a description of at least one condition; wherein the at least one condition is a detection of a visual or audio or environment-based marker or property, and wherein the trigger is activated when a condition of the at least one condition is detected in the real environment” recited in claim 1 and the above arguments (claim 10 has the same reasons). Regarding Claim 1, Applicants state in pages 9-10 of Remarks that “Finally, independent claims 1 and 10 recite "applying the action of the one anchor... to the one or more nodes associated with the one anchor." Emphasis added. This recites a specific process applied to a specific node (or nodes) in response to a trigger. In contrast, the rejection equates this to Furthermore, Bouazizi discloses the presentation unit of client device determines an anchor point from the scene description based on the user movement received via real-world repositioning of a device worn by the user and determines the reference space is "stage" (Paragraph [0157]); an then the presentation unit anchors the virtual scene to the real-world presentation environment at the determined real-world anchor point (Paragraphs [0133] and [0159]). Thus, Bouazizi discloses "on condition that the trigger of one of the one anchor or more anchors is activated, applying the action of the one anchor of the one or more anchors to the one or more nodes associated with the one anchor of the one or more anchors" recited in claim 1. Final Office Action at page 8 (rejecting at page 14). Bouazizi's disclosure of globally transforming the entire scene during initial setup (para. [0133], [0159]) is not evidence of obviousness. Bouazizi discloses to "transform the virtual scene to align with the real-world presentation environment." Id. at para. [0159]. A person of ordinary skill in the art appreciates that this is a global operation performed once on the entire scene. It is not an "action" applied to a specific node in response to a trigger, such as play a bell sound in a headset (Applicant's published specification at para. [0052]) or stop or pause of a media processing (Applicant's published specification at para. [0088]). A proper rejection requires evidence in Bouazizi of a discrete process being applied to a specific node as claimed”. Examiner replies: The examiner disagrees with Applicant’s premises and conclusion. Under the broadest reasonable interpretation, the examiner respectfully maintains that the prior art rejections in this case are proper for the following reasons. In respond to the applicant’s arguments, the examiner recites Bouazizi in order to disclose the issue. The claim recites “an action, wherein the action comprises a description of a process to be performed by an augmented reality engine” and “on condition that the trigger of one anchor of the one or more anchors is activated, applying the action of the one anchor of the one or more anchors to the one or more nodes associated with the one anchor of the one or more anchors”. Thus, the claim describes “an action” is a description of a process. In additional, the above limitation can be considered as “on condition that the trigger of one anchor of the one ”. As discussed above, the examiner considered the function performed by the client device equivalent to “augmented reality engine” and “the data” described by Bouazizi equivalent to “a trigger”. More specifically, Bouazizi discloses a client device (As shown in FIG. 3) to determine an anchor point from the scene description; anchor the virtual scene to the real-world presentation environment at the determined real-world anchor point and transform the virtual scene to align with the real-world presentation environment when the condition that “the data” is activated (Paragraphs [0130]-[0133], [0157] and [0159]). Therefore, Bouazizi discloses “an action, wherein the action comprises a description of a process to be performed by an augmented reality engine” and “on condition that the trigger of one anchor of the one or more anchors is activated, applying the action of the one anchor of the one or more anchors to the one or more nodes associated with the one anchor of the one or more anchors” recited in claim 1 and the above arguments (claim 10 has the same reasons). Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to understand Bouazizi disclose the invention as specified in claim 1 and above arguments (claim 10 has the same reasons). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2, 6-8, 10-11, 15-17 and 22-23 are rejected under 35 U.S.C. 103 as being unpatentable over Bouazizi et al (U.S. Patent Application Publication 2022/0335694 A1). Regarding claim 1, Bouazizi discloses a method for rendering an augmented reality scene for a user in a real environment (Paragraph [0006], this disclosure describes techniques for streaming immersive media content, e.g., for extended reality (XR) content, such as augmented reality (AR) ... For example, the user may be able to navigate the virtual scene using controllers and/or real-world movement. In order to allow for proper movement in a real-world presentation environment ...), the method comprising: obtaining a description of the augmented reality scene (FIGS. 10 and 11; paragraph [0156], client device 300 may receive a bitstream including a scene description (350); paragraph [0134], according to the techniques of this disclosure, an MPEG_scene_anchor extension may be added as a glTF 2.0 extension to the scene node. An AR scene may contain the MPEG_scene_anchor extension to describe the anchoring of the scene to a real-world XR space), the description comprising: a scene graph (Paragraph [0156], the scene description may include a scene graph ... For example, the scene description may correspond to the example scene graph of FIG. 5, scene description and updates 200 of FIG. 6, scene graph 262 and scene graph updates 264 of FIG. 8, or scene graph 324 of FIG. 10); and one or more anchors, wherein each anchor of the one or more anchors is associated with one or more nodes of the scene graph (Paragraphs [0092]-[0094], FIG. 5 shows a scene graph ... Each node in the graph holds pointers to its children. The child nodes can, among others, be a group of other nodes, a geometry element, a transformation matrix, accessors to media data buffers, camera information for the rendering ... Spatial transformations are represented as nodes of the graph and represented by a transformation matrix. Typical usage of transform nodes is to describe rotation, translation or scaling of the objects in its child nodes ...; FIG. 1; paragraph [0068], client device 40 may be configured to perform the various techniques of this disclosure alone or in any combination. In general, retrieval unit 52 may be configured to retrieve a bitstream including media data (e.g., scene data), as discussed above, as well as a scene description. The scene description may include anchor point data representing a correspondence between a virtual scene represented by the media data and a real-world presentation environment. Client device 40 may be configured to anchor the virtual scene to the real-world presentation environment using the anchor point data, and also transform the virtual scene as needed, e.g., through rotation, translation, and/or scaling) and comprises: a trigger, wherein the trigger is a description of at least one condition (Paragraph [0157], the scene description may include data for an anchor point, such as the MPEG_scene_anchor as discussed above with respect to Table 1. The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space); wherein the at least one condition is a detection of a visual or audio or environment-based marker or property (Paragraph [0158], presentation unit 330 may automatically detect the real-world presentation environment using camera 308. In some examples, client device 300 may receive image and/or video data captured by camera 308 and upload this data via a sceneUnderstandingStream as indicated by the scene description), and wherein the trigger is activated when a condition of the at least one condition is detected in the real environment (Paragraph [0157], the data for the anchor point may further indicate whether the scene data is to be transformed (e.g., rotated, translated, and/or scaled) to match the real-world presentation environment ... In general, the scene description may include data that relates the scene anchor point to a real-world anchor point, such as a particular location on the floor (e.g., a midpoint of the floor)); and an action, wherein the action comprises a description of a process to be performed by an augmented reality (Paragraph [0157], presentation unit 330 of client device 300 may determine an anchor point from the scene description (352). According to the techniques of this disclosure, the scene description may include data for an anchor point, such as the MPEG_scene_anchor as discussed above with respect to Table 1. The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space ... the anchor point may further include various actions that a user may perform, such as movements received via a controller or real-world repositioning of a device worn by the user ...); and on condition that the trigger of one anchor of the one or more anchors is activated (Paragraph [0157], presentation unit 330 of client device 300 may determine an anchor point from the scene description (352). According to the techniques of this disclosure, the scene description may include data for an anchor point, such as the MPEG_scene_anchor as discussed above with respect to Table 1. The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space; paragraph [0133], in the example of FIG. 9, the anchor XR space is of type “stage,” corresponding to the floor of the viewer's living room), applying the action of the one anchor of the one or more anchors to the one or more nodes associated with the one anchor of the one or more anchors (Paragraph [0159], presentation unit 330 may anchor the virtual scene to the real-world presentation environment at the determined real-world anchor point (358). Presentation unit 330 may also transform the virtual scene to align with the real-world presentation environment (360), e.g., using rotation, translation, and/or scaling; paragraphs [0130]-[0133], FIG. 9 is a conceptual diagram illustrating an example anchor XR space indicated by a scene description. According to the techniques of this disclosure, a scene description node may contain a reference to an XR space, which indicates that the scene is anchored to that space. The anchor XR space may be a reference space, e.g., local, view, or stage ... XR runtime systems, such as OpenXR, allow querying of the bounding space for an XR space. The scene description may request that the presentation engine aligns the scene extents, i.e. the bounding box of the scene, to the bounding box of the anchor XR space ... In the example of FIG. 9, the anchor XR space is of type “stage,” corresponding to the floor of the viewer's living room). It's noted that Bouazizi does not use term of “a trigger” by the techniques described in the disclosure. However, the claim recites “a trigger, wherein the trigger is a description of at least one condition”. Paragraph [0157] of Bouazizi describes “The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space” and the data describes at least one type “a view, local, stage, or application” of reference space. Thus, the “type” described by Bouazizi can be considered equivalent to a “condition”. Accordingly, under broadest reasonable interpretation, “the data” described by Bouazizi can be considered equivalent to “a trigger” recited in the claim. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to understand Bouazizi disclose the invention as specified in claim. Regarding claim 2, Bouazizi discloses everything claimed as applied above (see claim 1), and Bouazizi further disclose wherein the trigger of one of the one or more anchors comprises one or more limit conditions (FIGS. 10 and 11; paragraph [0157], the scene description may include data for an anchor point, such as the MPEG_scene_anchor as discussed above with respect to Table 1. The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space) and wherein the media content items are linked to the at least one of the one or more nodes and loaded only (Paragraph [0151], after aligning the virtual scene anchor point with the real-world anchor point and making any necessary transformations, presentation unit 330 may present media data 322 via display 314. For example, media data 322 may include data defining virtual objects, textures, colors, and locations for the virtual objects) when the one or more limit conditions are observed in the augmented reality scene (Paragraph [0133], in the example of FIG. 9, the anchor XR space is of type “stage,” corresponding to the floor of the viewer's living room). Regarding claim 6, Bouazizi discloses everything claimed as applied above (see claim 1), and Bouazizi further disclose wherein the at least one condition is a member of a group of conditions comprising detection of: one or more visual 2D markers, one or more visual 3D markers, one or more visual signatures, one or more visual geometric properties, one or more visual semantic properties; one or more audio markers, one or more audio properties; one or more temperature conditions, one or more movement of real or virtual objects, one or more hygrometry conditions, one or more lighting changes, and one or more wind conditions (FIGS. 10 and 11; paragraph [0148], anchor point detection unit 332 may determine a type of anchor point to be used from the data of scene graph 324, and identify a corresponding real-world anchor point using image data from camera 308. Presentation unit 330 may use the identified anchor point in the real-world presentation environment to align a virtual scene with the real-world presentation environment. For example, the real-world anchor point may be a point on the floor, a surface (e.g., a table), or the like. In some examples, a visual marker on the real-world object may be used, such as a quick response (QR) code, to represent the real-world anchor point). Regarding claim 7, Bouazizi discloses everything claimed as applied above (see claim 1), and Bouazizi further disclose the trigger of one of the one or more anchors relies on a detection of an object in the real environment (FIGS. 10 and 11; paragraph [0146], camera 308 represents a camera used to capture images or video data of the real-world presentation environment ... Additionally or alternatively, camera 308 may detect real-world objects, such as the floor ...) and wherein the trigger is associated with a model of the object or with a semantic description of the object (Paragraph [0157], the scene description may include data for an anchor point, such as the MPEG_scene_anchor as discussed above with respect to Table 1. The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space; paragraphs [0132]-[0133], the scene description may request that the presentation engine aligns the scene extents, i.e. the bounding box of the scene, to the bounding box of the anchor XR space ... If no scaling is applied, the presentation engine aligns the long edge of the scene bounding box to that of the XR space and then centers the scene bounding box to be collocated with the center of the XR space bounding box ... In the example of FIG. 9, the anchor XR space is of type “stage,” corresponding to the floor of the viewer's living room). Regarding claim 8, Bouazizi discloses everything claimed as applied above (see claim 1), and Bouazizi further disclose wherein the action of one of the one or more anchors is a member of a group of actions comprising: playing, pausing or stopping a media content item; modifying the description of the augmented reality scene (FIGS. 10 and 11; paragraph [0151], after aligning the virtual scene anchor point with the real-world anchor point and making any necessary transformations, presentation unit 330 may present media data 322 via display 314 ... Presentation unit 330 may update the presentation according to user movements detected from user interface devices 306, camera 308, and/or sensors 310, and/or based on updated to scene graph 324); and connecting a remote device or service. Regarding claim 10, Bouazizi discloses a device for rendering an augmented reality scene for a user in a real environment (Paragraph [0006], this disclosure describes techniques for streaming immersive media content, e.g., for extended reality (XR) content, such as augmented reality (AR) ... For example, the user may be able to navigate the virtual scene using controllers and/or real-world movement. In order to allow for proper movement in a real-world presentation environment ...), the device comprising a memory associated with a processor (Paragraph [0008], a device for presenting media data includes a memory configured to store media data defining one or more virtual objects in a virtual scene; and one or more processors implemented in circuitry and configured to ...) configured for: obtaining a description of the augmented reality scene (FIGS. 10 and 11; paragraph [0156], client device 300 may receive a bitstream including a scene description (350); paragraph [0134], according to the techniques of this disclosure, an MPEG_scene_anchor extension may be added as a glTF 2.0 extension to the scene node. An AR scene may contain the MPEG_scene_anchor extension to describe the anchoring of the scene to a real-world XR space), the description comprising: a scene graph (Paragraph [0156], he scene description may include a scene graph ... For example, the scene description may correspond to the example scene graph of FIG. 5, scene description and updates 200 of FIG. 6, scene graph 262 and scene graph updates 264 of FIG. 8, or scene graph 324 of FIG. 10); and one or more anchors, wherein each anchor of the one or more anchors is associated with one or more nodes of the scene graph (Paragraphs [0092]-[0094], FIG. 5 shows a scene graph ... Each node in the graph holds pointers to its children. The child nodes can, among others, be a group of other nodes, a geometry element, a transformation matrix, accessors to media data buffers, camera information for the rendering ... Spatial transformations are represented as nodes of the graph and represented by a transformation matrix. Typical usage of transform nodes is to describe rotation, translation or scaling of the objects in its child nodes ...; FIG. 1; paragraph [0068], client device 40 may be configured to perform the various techniques of this disclosure alone or in any combination. In general, retrieval unit 52 may be configured to retrieve a bitstream including media data (e.g., scene data), as discussed above, as well as a scene description. The scene description may include anchor point data representing a correspondence between a virtual scene represented by the media data and a real-world presentation environment. Client device 40 may be configured to anchor the virtual scene to the real-world presentation environment using the anchor point data, and also transform the virtual scene as needed, e.g., through rotation, translation, and/or scaling) and comprises: a trigger, wherein the trigger is a description of at least one condition (Paragraph [0157], the scene description may include data for an anchor point, such as the MPEG_scene_anchor as discussed above with respect to Table 1. The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space); wherein the at least one condition is a detection of a visual or audio or environment-based marker or property (Paragraph [0158], presentation unit 330 may automatically detect the real-world presentation environment using camera 308. In some examples, client device 300 may receive image and/or video data captured by camera 308 and upload this data via a sceneUnderstandingStream as indicated by the scene description), and wherein the trigger is activated when a condition of the at least one condition is detected in the real environment (Paragraph [0157], the data for the anchor point may further indicate whether the scene data is to be transformed (e.g., rotated, translated, and/or scaled) to match the real-world presentation environment ... In general, the scene description may include data that relates the scene anchor point to a real-world anchor point, such as a particular location on the floor (e.g., a midpoint of the floor)); and an action, wherein the action comprises a description of a process to be performed by an augmented reality engine (Paragraph [0157], presentation unit 330 of client device 300 may determine an anchor point from the scene description (352). According to the techniques of this disclosure, the scene description may include data for an anchor point, such as the MPEG_scene_anchor as discussed above with respect to Table 1. The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space ... the anchor point may further include various actions that a user may perform, such as movements received via a controller or real-world repositioning of a device worn by the user ...); and on condition that the trigger of one anchor of the one or more anchors is activated (Paragraph [0157], presentation unit 330 of client device 300 may determine an anchor point from the scene description (352). According to the techniques of this disclosure, the scene description may include data for an anchor point, such as the MPEG_scene_anchor as discussed above with respect to Table 1. The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space; paragraph [0133], in the example of FIG. 9, the anchor XR space is of type “stage,” corresponding to the floor of the viewer's living room), applying the action of the one anchor of the one or more anchors to the one or more nodes associated with the one anchor of the one or more anchors (Paragraph [0159], presentation unit 330 may anchor the virtual scene to the real-world presentation environment at the determined real-world anchor point (358). Presentation unit 330 may also transform the virtual scene to align with the real-world presentation environment (360), e.g., using rotation, translation, and/or scaling; paragraphs [0130]-[0133], FIG. 9 is a conceptual diagram illustrating an example anchor XR space indicated by a scene description. According to the techniques of this disclosure, a scene description node may contain a reference to an XR space, which indicates that the scene is anchored to that space. The anchor XR space may be a reference space, e.g., local, view, or stage ... XR runtime systems, such as OpenXR, allow querying of the bounding space for an XR space. The scene description may request that the presentation engine aligns the scene extents, i.e. the bounding box of the scene, to the bounding box of the anchor XR space ... In the example of FIG. 9, the anchor XR space is of type “stage,” corresponding to the floor of the viewer's living room). It's noted that Bouazizi does not use term of “a trigger” by the techniques described in the disclosure. However, the claim recites “a trigger, wherein the trigger is a description of at least one condition”. Paragraph [0157] of Bouazizi describes “The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space” and the data describes at least one type “a view, local, stage, or application” of reference space. Thus, the “type” described by Bouazizi can be considered equivalent to a “condition”. Accordingly, under broadest reasonable interpretation, “the data” described by Bouazizi can be considered equivalent to “a trigger” recited in the claim. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to understand Bouazizi disclose the invention as specified in claim. Regarding claim 11, Bouazizi discloses everything claimed as applied above (see claim 10), and Bouazizi further disclose wherein the trigger of one of the one or more anchors comprises one or more limit conditions (FIGS. 10 and 11; paragraph [0157], the scene description may include data for an anchor point, such as the MPEG_scene_anchor as discussed above with respect to Table 1. The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space) and wherein the media content items are linked to the one or more nodes and loaded only (Paragraph [0151], after aligning the virtual scene anchor point with the real-world anchor point and making any necessary transformations, presentation unit 330 may present media data 322 via display 314. For example, media data 322 may include data defining virtual objects, textures, colors, and locations for the virtual objects) when the one or more limit conditions are observed in the augmented reality scene (Paragraph [0133], in the example of FIG. 9, the anchor XR space is of type “stage,” corresponding to the floor of the viewer's living room). Regarding claim 15, Bouazizi discloses everything claimed as applied above (see claim 10), and Bouazizi further disclose wherein the at least one condition is a member of a group of conditions comprising detection of: one or more visual 2D markers, one or more visual 3D markers, one or more visual signatures, one or more visual geometric properties; one or more visual semantic properties; one or more audio markers, one or more audio properties; one or more temperature conditions, one or more movement of real or virtual objects, one or more hygrometry conditions, one or more lighting changes, and one or more wind conditions (FIGS. 10 and 11; paragraph [0148], anchor point detection unit 332 may determine a type of anchor point to be used from the data of scene graph 324, and identify a corresponding real-world anchor point using image data from camera 308. Presentation unit 330 may use the identified anchor point in the real-world presentation environment to align a virtual scene with the real-world presentation environment. For example, the real-world anchor point may be a point on the floor, a surface (e.g., a table), or the like. In some examples, a visual marker on the real-world object may be used, such as a quick response (QR) code, to represent the real-world anchor point). Regarding claim 16, Bouazizi discloses everything claimed as applied above (see claim 10), and Bouazizi further disclose the trigger of one of the one or more anchors relies on a detection of an object in the real environment (FIGS. 10 and 11; paragraph [0146], camera 308 represents a camera used to capture images or video data of the real-world presentation environment ... Additionally or alternatively, camera 308 may detect real-world objects, such as the floor ...) and wherein the trigger is associated with a model of the object or with a semantic description of the object (Paragraph [0157], the scene description may include data for an anchor point, such as the MPEG_scene_anchor as discussed above with respect to Table 1. The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space; paragraphs [0132]-[0133], the scene description may request that the presentation engine aligns the scene extents, i.e. the bounding box of the scene, to the bounding box of the anchor XR space ... If no scaling is applied, the presentation engine aligns the long edge of the scene bounding box to that of the XR space and then centers the scene bounding box to be collocated with the center of the XR space bounding box ... In the example of FIG. 9, the anchor XR space is of type “stage,” corresponding to the floor of the viewer's living room). Regarding claim 17, Bouazizi discloses everything claimed as applied above (see claim 10), and Bouazizi further disclose wherein the at least an action of one of the one or more anchors is a member of a group of actions comprising: playing, pausing or stopping a media content item; modifying the description of the augmented reality scene (FIGS. 10 and 11; paragraph [0151], after aligning the virtual scene anchor point with the real-world anchor point and making any necessary transformations, presentation unit 330 may present media data 322 via display 314 ... Presentation unit 330 may update the presentation according to user movements detected from user interface devices 306, camera 308, and/or sensors 310, and/or based on updated to scene graph 324); and connecting a remote device or service. Regarding claim 22, Bouazizi discloses a method for rendering an augmented reality scene for a user in a real environment (Paragraph [0006], this disclosure describes techniques for streaming immersive media content, e.g., for extended reality (XR) content, such as augmented reality (AR) ... For example, the user may be able to navigate the virtual scene using controllers and/or real-world movement. In order to allow for proper movement in a real-world presentation environment ...), the method comprising: obtaining a description of the augmented reality scene (FIGS. 10 and 11; paragraph [0156], client device 300 may receive a bitstream including a scene description (350); paragraph [0134], according to the techniques of this disclosure, an MPEG_scene_anchor extension may be added as a glTF 2.0 extension to the scene node. An AR scene may contain the MPEG_scene_anchor extension to describe the anchoring of the scene to a real-world XR space), the description comprising: a scene graph (Paragraph [0156], the scene description may include a scene graph ... For example, the scene description may correspond to the example scene graph of FIG. 5, scene description and updates 200 of FIG. 6, scene graph 262 and scene graph updates 264 of FIG. 8, or scene graph 324 of FIG. 10); and one or more anchors, wherein each anchor of the one or more anchors is associated with one or more nodes of the scene graph (Paragraphs [0092]-[0094], FIG. 5 shows a scene graph ... Each node in the graph holds pointers to its children. The child nodes can, among others, be a group of other nodes, a geometry element, a transformation matrix, accessors to media data buffers, camera information for the rendering ... Spatial transformations are represented as nodes of the graph and represented by a transformation matrix. Typical usage of transform nodes is to describe rotation, translation or scaling of the objects in its child nodes ...; FIG. 1; paragraph [0068], client device 40 may be configured to perform the various techniques of this disclosure alone or in any combination. In general, retrieval unit 52 may be configured to retrieve a bitstream including media data (e.g., scene data), as discussed above, as well as a scene description. The scene description may include anchor point data representing a correspondence between a virtual scene represented by the media data and a real-world presentation environment. Client device 40 may be configured to anchor the virtual scene to the real-world presentation environment using the anchor point data, and also transform the virtual scene as needed, e.g., through rotation, translation, and/or scaling) and comprises: a trigger, wherein the trigger is a description of at least one condition (Paragraph [0157], the scene description may include data for an anchor point, such as the MPEG_scene_anchor as discussed above with respect to Table 1. The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space); wherein the at least one condition is a detection of a visual or audio or environment-based marker or property (Paragraph [0158], presentation unit 330 may automatically detect the real-world presentation environment using camera 308. In some examples, client device 300 may receive image and/or video data captured by camera 308 and upload this data via a sceneUnderstandingStream as indicated by the scene description), and wherein the trigger is activated when a condition of the at least one condition is detected in the real environment (Paragraph [0157], the data for the anchor point may further indicate whether the scene data is to be transformed (e.g., rotated, translated, and/or scaled) to match the real-world presentation environment ... In general, the scene description may include data that relates the scene anchor point to a real-world anchor point, such as a particular location on the floor (e.g., a midpoint of the floor)); and an action, wherein the action comprises a description of a process (Paragraph [0157], presentation unit 330 of client device 300 may determine an anchor point from the scene description (352). According to the techniques of this disclosure, the scene description may include data for an anchor point, such as the MPEG_scene_anchor as discussed above with respect to Table 1. The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space ... the anchor point may further include various actions that a user may perform, such as movements received via a controller or real-world repositioning of a device worn by the user ...), comprising at least one of playing, pausing, or stopping a media content item, to be performed by an augmented reality engine (FIG. 11; paragraph [0159], presentation unit 330 may then receive data for the virtual scene (356). For example, the bitstream may include data for one or more virtual objects, including data defining the objects themselves, locations of the objects within the virtual scene, textures for the objects, and colors for the objects. Presentation unit 330 may anchor the virtual scene to the real-world presentation environment at the determined real-world anchor point (358). Presentation unit 330 may also transform the virtual scene to align with the real-world presentation environment (360), e.g., using rotation, translation, and/or scaling); and on condition that the trigger of one of the one anchor or more anchors is activated (Paragraph [0157], presentation unit 330 of client device 300 may determine an anchor point from the scene description (352). According to the techniques of this disclosure, the scene description may include data for an anchor point, such as the MPEG_scene_anchor as discussed above with respect to Table 1. The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space; paragraph [0133], in the example of FIG. 9, the anchor XR space is of type “stage,” corresponding to the floor of the viewer's living room), applying the action of the one anchor of the one or more anchors to the one or more nodes associated with the one anchor of the one or more anchors (Paragraph [0159], presentation unit 330 may anchor the virtual scene to the real-world presentation environment at the determined real-world anchor point (358). Presentation unit 330 may also transform the virtual scene to align with the real-world presentation environment (360), e.g., using rotation, translation, and/or scaling; paragraphs [0130]-[0133], FIG. 9 is a conceptual diagram illustrating an example anchor XR space indicated by a scene description. According to the techniques of this disclosure, a scene description node may contain a reference to an XR space, which indicates that the scene is anchored to that space. The anchor XR space may be a reference space, e.g., local, view, or stage ... XR runtime systems, such as OpenXR, allow querying of the bounding space for an XR space. The scene description may request that the presentation engine aligns the scene extents, i.e. the bounding box of the scene, to the bounding box of the anchor XR space ... In the example of FIG. 9, the anchor XR space is of type “stage,” corresponding to the floor of the viewer's living room). It's noted that Bouazizi does not use term of “a trigger” by the techniques described in the disclosure. However, the claim recites “a trigger, wherein the trigger is a description of at least one condition”. Paragraph [0157] of Bouazizi describes “The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space” and the data describes at least one type “a view, local, stage, or application” of reference space. Thus, the “type” described by Bouazizi can be considered equivalent to a “condition”. Accordingly, under broadest reasonable interpretation, “the data” described by Bouazizi can be considered equivalent to “a trigger” recited in the claim. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to understand Bouazizi disclose the invention as specified in claim. Regarding claim 23, Bouazizi discloses a device for rendering an augmented reality scene for a user in a real environment (Paragraph [0006], this disclosure describes techniques for streaming immersive media content, e.g., for extended reality (XR) content, such as augmented reality (AR) ... For example, the user may be able to navigate the virtual scene using controllers and/or real-world movement. In order to allow for proper movement in a real-world presentation environment ...), the device comprising a memory associated with a processor (Paragraph [0008], a device for presenting media data includes a memory configured to store media data defining one or more virtual objects in a virtual scene; and one or more processors implemented in circuitry and configured to ...) configured for: obtaining a description of the augmented reality scene (FIGS. 10 and 11; paragraph [0156], client device 300 may receive a bitstream including a scene description (350); paragraph [0134], according to the techniques of this disclosure, an MPEG_scene_anchor extension may be added as a glTF 2.0 extension to the scene node. An AR scene may contain the MPEG_scene_anchor extension to describe the anchoring of the scene to a real-world XR space), the description comprising: a scene graph (Paragraph [0156], the scene description may include a scene graph ... For example, the scene description may correspond to the example scene graph of FIG. 5, scene description and updates 200 of FIG. 6, scene graph 262 and scene graph updates 264 of FIG. 8, or scene graph 324 of FIG. 10); and one or more anchors, wherein each anchor of the one or more anchors is associated with one or more nodes of the scene graph (Paragraphs [0092]-[0094], FIG. 5 shows a scene graph ... Each node in the graph holds pointers to its children. The child nodes can, among others, be a group of other nodes, a geometry element, a transformation matrix, accessors to media data buffers, camera information for the rendering ... Spatial transformations are represented as nodes of the graph and represented by a transformation matrix. Typical usage of transform nodes is to describe rotation, translation or scaling of the objects in its child nodes ...; FIG. 1; paragraph [0068], client device 40 may be configured to perform the various techniques of this disclosure alone or in any combination. In general, retrieval unit 52 may be configured to retrieve a bitstream including media data (e.g., scene data), as discussed above, as well as a scene description. The scene description may include anchor point data representing a correspondence between a virtual scene represented by the media data and a real-world presentation environment. Client device 40 may be configured to anchor the virtual scene to the real-world presentation environment using the anchor point data, and also transform the virtual scene as needed, e.g., through rotation, translation, and/or scaling) and comprises: a trigger, wherein the trigger is a description of at least one condition (Paragraph [0157], the scene description may include data for an anchor point, such as the MPEG_scene_anchor as discussed above with respect to Table 1. The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space); wherein the at least one condition is a detection of a visual or audio or environment-based marker or property (Paragraph [0158], presentation unit 330 may automatically detect the real-world presentation environment using camera 308. In some examples, client device 300 may receive image and/or video data captured by camera 308 and upload this data via a sceneUnderstandingStream as indicated by the scene description), and wherein the trigger is activated when a condition of the at least one condition is detected in the real environment (Paragraph [0157], the data for the anchor point may further indicate whether the scene data is to be transformed (e.g., rotated, translated, and/or scaled) to match the real-world presentation environment ... In general, the scene description may include data that relates the scene anchor point to a real-world anchor point, such as a particular location on the floor (e.g., a midpoint of the floor)); and an action, wherein the action comprises a description of a process (Paragraph [0157], presentation unit 330 of client device 300 may determine an anchor point from the scene description (352). According to the techniques of this disclosure, the scene description may include data for an anchor point, such as the MPEG_scene_anchor as discussed above with respect to Table 1. The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space ... the anchor point may further include various actions that a user may perform, such as movements received via a controller or real-world repositioning of a device worn by the user ...), comprising at least one of playing, pausing, or stopping a media content item, to be performed by an augmented reality engine (Paragraph [0159], presentation unit 330 may then receive data for the virtual scene (356). For example, the bitstream may include data for one or more virtual objects, including data defining the objects themselves, locations of the objects within the virtual scene, textures for the objects, and colors for the objects. Presentation unit 330 may anchor the virtual scene to the real-world presentation environment at the determined real-world anchor point (358). Presentation unit 330 may also transform the virtual scene to align with the real-world presentation environment (360), e.g., using rotation, translation, and/or scaling); and on condition that the trigger of one anchor of the one or more anchors is activated (Paragraph [0157], presentation unit 330 of client device 300 may determine an anchor point from the scene description (352). According to the techniques of this disclosure, the scene description may include data for an anchor point, such as the MPEG_scene_anchor as discussed above with respect to Table 1. The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space; paragraph [0133], in the example of FIG. 9, the anchor XR space is of type “stage,” corresponding to the floor of the viewer's living room), applying the action of the one anchor of the one or more anchors to the one or more nodes associated with the one anchor of the one or more anchors (Paragraph [0159], presentation unit 330 may anchor the virtual scene to the real-world presentation environment at the determined real-world anchor point (358). Presentation unit 330 may also transform the virtual scene to align with the real-world presentation environment (360), e.g., using rotation, translation, and/or scaling; paragraphs [0130]-[0133], FIG. 9 is a conceptual diagram illustrating an example anchor XR space indicated by a scene description. According to the techniques of this disclosure, a scene description node may contain a reference to an XR space, which indicates that the scene is anchored to that space. The anchor XR space may be a reference space, e.g., local, view, or stage ... XR runtime systems, such as OpenXR, allow querying of the bounding space for an XR space. The scene description may request that the presentation engine aligns the scene extents, i.e. the bounding box of the scene, to the bounding box of the anchor XR space ... In the example of FIG. 9, the anchor XR space is of type “stage,” corresponding to the floor of the viewer's living room). It's noted that Bouazizi does not use term of “a trigger” by the techniques described in the disclosure. However, the claim recites “a trigger, wherein the trigger is a description of at least one condition”. Paragraph [0157] of Bouazizi describes “The data for the anchor point may include, for example, data indicating whether the reference space type (e.g., a real-world presentation environment) is to be a view, local, stage, or application type of reference space” and the data describes at least one type “a view, local, stage, or application” of reference space. Thus, the “type” described by Bouazizi can be considered equivalent to a “condition”. Accordingly, under broadest reasonable interpretation, “the data” described by Bouazizi can be considered equivalent to “a trigger” recited in the claim. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to understand Bouazizi disclose the invention as specified in claim. Claims 3-4 and 12-13 are rejected under 35 U.S.C. 103 as being unpatentable over Bouazizi et al (U.S. Patent Application Publication 2022/0335694 A1) in view of Smet et al (U.S. Patent No. 10,825,258 B1). Regarding claim 3, Bouazizi discloses everything claimed as applied above (see claim 1). However, Bouazizi does not specifically disclose further comprising, when the one or more limit conditions are no longer observed in the augmented reality scene, unloading the media content items linked to the at least one of the one or more nodes. In additional, Smet discloses (Abstract, a method includes by a computing device, displaying a user interface for designing augmented-reality effects. The method includes receiving user input through the user interface. The method includes displaying a graph generated based on the user input ...) further comprising, when the one or more limit conditions are no longer observed in the augmented reality scene (Col 14, lines 24-67, FIGS. 5A-5D illustrate example scene graphs associated with a variety of augmented-reality effects. FIG. 5A illustrates a scene graph 500a corresponding to an augmented-reality effect wherein a virtual statue object is render in an open, palm-up hand in the scene ...; Col 15, lines 3-33, FIG. 5B illustrates a scene graph 500b corresponding to an augmented-reality effect wherein gestures detected in association with face object instances affect the visibility and animation of assets in the effect. The scene graph 500b comprises a collective node 510b labeled “Facetracker” that corresponds to a module for identifying and tracking the first two recognized female faces in the scene. Thus, palm-up hand in the scene is no longer observed in the augmented reality scene after detecting face object), unloading the media content items linked to the at least one of the one or more nodes (Col 15, lines 3-33, the scene graph 500b comprises a collective node 510b labeled “Facetracker” that corresponds to a module for identifying and tracking the first two recognized female faces in the scene ... Thus, the scene graph 500a of “Handtracker” is unloaded after the scene graph 500b of “Facetracker” is loaded). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the scene description taught by Bouazizi incorporate the teachings of Smet, and applying the graph-based design of augmented-reality effects taught by Smet to collect nodes shown in the graph based on the object detection and provide the object type appearing in a scene to the user. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify Bouazizi according to the relied-upon teachings of Smet to obtain the invention as specified in claim. Regarding claim 4, Bouazizi discloses everything claimed as applied above (see claim 1). However, Bouazizi does not specifically disclose wherein the trigger of one of the one or more anchors comprises a descriptor indicating whether the action of the one or more anchors continues once the trigger is no longer activated. In additional, Smet discloses (Abstract, a method includes by a computing device, displaying a user interface for designing augmented-reality effects. The method includes receiving user input through the user interface. The method includes displaying a graph generated based on the user input ...) wherein the trigger of one of the one or more anchors (Col 14, lines 24-67, FIGS. 5A-5D illustrate example scene graphs associated with a variety of augmented-reality effects. FIG. 5A illustrates a scene graph 500a corresponding to an augmented-reality effect wherein a virtual statue object is render in an open, palm-up hand in the scene ...; Col 15, lines 3-33, FIG. 5B illustrates a scene graph 500b corresponding to an augmented-reality effect wherein gestures detected in association with face object instances affect the visibility and animation of assets in the effect. The scene graph 500b comprises a collective node 510b labeled “Facetracker” that corresponds to a module for identifying and tracking the first two recognized female faces in the scene) comprises a descriptor indicating whether the action of the one or more anchors continues once the trigger is no longer activated (Col 15, lines 3-33, the scene graph 500b comprises a collective node 510b labeled “Facetracker” that corresponds to a module for identifying and tracking the first two recognized female faces in the scene ... Thus, the scene graph 500a of “Handtracker” is no longer activated after the scene graph 500b of “Facetracker” is loaded). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the scene description taught by Bouazizi incorporate the teachings of Smet, and applying the graph-based design of augmented-reality effects taught by Smet to collect nodes shown in the graph based on the object detection and provide the object type appearing in a scene to the user. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify Bouazizi according to the relied-upon teachings of Smet to obtain the invention as specified in claim. Regarding claim 12, Bouazizi discloses everything claimed as applied above (see claim 11). However, Bouazizi does not specifically disclose wherein the processor is further configured for, when the one or more limit conditions are no longer observed in the augmented reality scene, unloading the media content items linked to the at least one of the one or more nodes. In additional, Smet discloses (Abstract, a method includes by a computing device, displaying a user interface for designing augmented-reality effects. The method includes receiving user input through the user interface. The method includes displaying a graph generated based on the user input ...) wherein the processor is further configured for, when the one or more limit conditions are no longer observed in the augmented reality scene (Col 14, lines 24-67, FIGS. 5A-5D illustrate example scene graphs associated with a variety of augmented-reality effects. FIG. 5A illustrates a scene graph 500a corresponding to an augmented-reality effect wherein a virtual statue object is render in an open, palm-up hand in the scene ...; Col 15, lines 3-33, FIG. 5B illustrates a scene graph 500b corresponding to an augmented-reality effect wherein gestures detected in association with face object instances affect the visibility and animation of assets in the effect. The scene graph 500b comprises a collective node 510b labeled “Facetracker” that corresponds to a module for identifying and tracking the first two recognized female faces in the scene. Thus, palm-up hand in the scene is no longer observed in the augmented reality scene after detecting face object), unloading the media content items linked to the at least one of the one or more nodes (Col 15, lines 3-33, the scene graph 500b comprises a collective node 510b labeled “Facetracker” that corresponds to a module for identifying and tracking the first two recognized female faces in the scene ... Thus, the scene graph 500a of “Handtracker” is unloaded after the scene graph 500b of “Facetracker” is loaded). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the scene description taught by Bouazizi incorporate the teachings of Smet, and applying the graph-based design of augmented-reality effects taught by Smet to collect nodes shown in the graph based on the object detection and provide the object type appearing in a scene to the user. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify Bouazizi according to the relied-upon teachings of Smet to obtain the invention as specified in claim. Regarding claim 13, Bouazizi discloses everything claimed as applied above (see claim 10). However, Bouazizi does not specifically disclose wherein the trigger of one of the one or more anchors comprises a descriptor indicating whether the action of the one or more anchors continues once the trigger is no longer activated. In additional, Smet discloses (Abstract, a method includes by a computing device, displaying a user interface for designing augmented-reality effects. The method includes receiving user input through the user interface. The method includes displaying a graph generated based on the user input ...) wherein the trigger of one of the one or more anchors (Col 14, lines 24-67, FIGS. 5A-5D illustrate example scene graphs associated with a variety of augmented-reality effects. FIG. 5A illustrates a scene graph 500a corresponding to an augmented-reality effect wherein a virtual statue object is render in an open, palm-up hand in the scene ...; Col 15, lines 3-33, FIG. 5B illustrates a scene graph 500b corresponding to an augmented-reality effect wherein gestures detected in association with face object instances affect the visibility and animation of assets in the effect. The scene graph 500b comprises a collective node 510b labeled “Facetracker” that corresponds to a module for identifying and tracking the first two recognized female faces in the scene) comprises a descriptor indicating whether the action of the one or more anchors continues once the trigger is no longer activated (Col 15, lines 3-33, the scene graph 500b comprises a collective node 510b labeled “Facetracker” that corresponds to a module for identifying and tracking the first two recognized female faces in the scene ... Thus, the scene graph 500a of “Handtracker” is no longer activated after the scene graph 500b of “Facetracker” is loaded). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the scene description taught by Bouazizi incorporate the teachings of Smet, and applying the graph-based design of augmented-reality effects taught by Smet to collect nodes shown in the graph based on the object detection and provide the object type appearing in a scene to the user. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify Bouazizi according to the relied-upon teachings of Smet to obtain the invention as specified in claim. Claims 5 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Bouazizi et al (U.S. Patent Application Publication 2022/0335694 A1) in view of Kottur (U.S. Patent No. 12,198,430 B1). Regarding claim 5, Bouazizi discloses everything claimed as applied above (see claim 1). However, Bouazizi does not specifically disclose wherein the trigger of one of the one or more anchors comprises a description of at least two conditions and a descriptor indicating how to combine the at least two conditions. In additional, Kottur discloses (FIG. 6 shows a simplified scene graph generator) wherein the trigger of one of the one or more anchors (Col 50, lines 53-67; a scene graph may be updated dynamically over time as additional data concerning object(s) of interest is received. FIG. 10A illustrates an example image of a sporting scene 1000 captured by a camera of a client system 130. The image of the scene 1000 portrays a batter 1001, a catcher 1002, and a coach 1003 having various relationships with objects and spatial relationships with one another. FIGS. 10B-10D illustrate the incremental generation of scene graphs concerning the scene of FIG. 10A over time ...; Col 51, lines 41-67, FIG. 10C illustrates such an example final scene graph 1030 at time t. As an example and not by way of limitation, the final scene graph 1030 may include the node N1 corresponding to the batter 1001, as well as the various attributes A1-A4 and their respective values that were incrementally generated over the course of the dialog ...) comprises a description of at least two conditions (Col 52, lines 1-15, the stored final scene graph 1030 may be updated dynamically as information on the attributes A2 and A4 that were assigned values of “unknown” become available at a later time t+1. FIG. 10D illustrates an example final scene graph 1040 derived from final scene graph 1030 at time t+1 ... the assistant system 1020 may update these values (e.g., the value of the “facial expression” attribute A2 may be updated to “smiling”), thus creating a final scene graph 1040 having these updated values ...) and a descriptor indicating how to combine the at least two conditions (Col 52, lines 16-26, the dynamic updating of the scene graph may occur automatically and repeatedly over a period of time. As an example and not by way of limitation, the “facial expression” attribute A2 of the batter 1001 may be updated each time the batter 1001 moves in a way that his facial expression can be seen). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the scene description taught by Bouazizi incorporate the teachings of Kottur, and applying the scene graphs for assistant systems taught by Kottur to provide the multiple conditions into “the trigger” taught by Bouazizi and combine the at least two conditions to generate a set of content objects to display to a user. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify Bouazizi according to the relied-upon teachings of Kottur to obtain the invention as specified in claim. Regarding claim 14, Bouazizi discloses everything claimed as applied above (see claim 10). However, Bouazizi does not specifically disclose wherein a trigger of one of the one or more anchors comprises a description of at least two conditions and a descriptor indicating how to combine the at least two conditions. In additional, Kottur discloses (FIG. 6 shows a simplified scene graph generator) wherein a trigger of one of the one or more anchors (Col 50, lines 53-67; a scene graph may be updated dynamically over time as additional data concerning object(s) of interest is received. FIG. 10A illustrates an example image of a sporting scene 1000 captured by a camera of a client system 130. The image of the scene 1000 portrays a batter 1001, a catcher 1002, and a coach 1003 having various relationships with objects and spatial relationships with one another. FIGS. 10B-10D illustrate the incremental generation of scene graphs concerning the scene of FIG. 10A over time ...; Col 51, lines 41-67, FIG. 10C illustrates such an example final scene graph 1030 at time t. As an example and not by way of limitation, the final scene graph 1030 may include the node N1 corresponding to the batter 1001, as well as the various attributes A1-A4 and their respective values that were incrementally generated over the course of the dialog ...) comprises a description of at least two conditions (Col 52, lines 1-15, the stored final scene graph 1030 may be updated dynamically as information on the attributes A2 and A4 that were assigned values of “unknown” become available at a later time t+1. FIG. 10D illustrates an example final scene graph 1040 derived from final scene graph 1030 at time t+1 ... the assistant system 1020 may update these values (e.g., the value of the “facial expression” attribute A2 may be updated to “smiling”), thus creating a final scene graph 1040 having these updated values ...) and a descriptor indicating how to combine the at least two conditions (Col 52, lines 16-26, the dynamic updating of the scene graph may occur automatically and repeatedly over a period of time. As an example and not by way of limitation, the “facial expression” attribute A2 of the batter 1001 may be updated each time the batter 1001 moves in a way that his facial expression can be seen). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the scene description taught by Bouazizi incorporate the teachings of Kottur, and applying the scene graphs for assistant systems taught by Kottur to provide the multiple conditions into “the trigger” taught by Bouazizi and combine the at least two conditions to generate a set of content objects to display to a user. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify Bouazizi according to the relied-upon teachings of Kottur to obtain the invention as specified in claim. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Xilin Guo whose telephone number is (571)272-5786. The examiner can normally be reached Monday - Friday 9:00 AM-5:30 PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Hajnik can be reached at 571-272-7642. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /XILIN GUO/Primary Examiner, Art Unit 2616
Read full office action

Prosecution Timeline

Apr 04, 2024
Application Filed
Dec 08, 2025
Non-Final Rejection mailed — §103
Apr 17, 2026
Response Filed
May 05, 2026
Final Rejection mailed — §103
Jul 13, 2026
Request for Continued Examination
Jul 14, 2026
Response after Non-Final Action
Aug 13, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749264
DEVICES, METHODS AND GRAPHICAL USER INTERFACES FOR PREVIEW OF COMPUTER-GENERATED VIRTUAL OBJECTS FOR EXTENDED REALITY APPLICATIONS
2y 3m to grant Granted Sep 29, 2026
Patent 12737836
DATA TRANSMISSION METHOD AND RELATED DEVICE FOR VIRTUAL REALITY SYSTEM
2y 2m to grant Granted Sep 15, 2026
Patent 12731208
DATA TRANSMISSION METHOD AND RELATED DEVICE FOR VIRTUAL REALITY SYSTEM
2y 4m to grant Granted Sep 08, 2026
Patent 12725379
METHOD AND APPARATUS FOR CONTROLLING VIRTUAL OBJECT, ELECTRONIC DEVICE, STORAGE MEDIUM, AND PROGRAM PRODUCT
2y 10m to grant Granted Sep 01, 2026
Patent 12725350
LIGHT ESTIMATION USING NEURAL NETWORKS
2y 1m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
82%
Grant Probability
99%
With Interview (+18.6%)
2y 4m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 474 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month