Prosecution Insights
Last updated: October 01, 2026
Application No. 19/009,478

Artificial Intelligence-Assisted Virtual Object Builder

Non-Final OA §103§DOUBLEPATENT
Filed
Jan 03, 2025
Priority
Feb 14, 2022 — provisional 63/309,760 +1 more
Examiner
SHENG, XIN
Art Unit
Tech Center
Assignee
Meta Platforms Inc.
OA Round
1 (Non-Final)
73%
Grant Probability
Favorable
1-2
OA Rounds
7m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 73% — above average
73%
Career Allowance Rate
301 granted / 412 resolved
+13.1% vs TC avg
Strong +17% interview lift
Without
With
+16.8%
Interview Lift
resolved cases with interview
Typical timeline
2y 4m
Avg Prosecution
21 currently pending
Career history
431
Total Applications
across all art units

Statute-Specific Performance

§101
5.9%
-34.1% vs TC avg
§103
79.8%
+39.8% vs TC avg
§102
1.7%
-38.3% vs TC avg
§112
5.4%
-34.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 412 resolved cases

Office Action

§103 §DOUBLEPATENT
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 21, 24, 26, 28, 35-38 is rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-2, 7, 12-13 of U.S. Patent No. 12254564. US App #19009478 21 21 24 26 28 35 36 37 38 US Patent #12254564 1 13 13 2 1 1 7 12 13 US App #19009478 Claim 21 US Patent #12254564 Claim 1 21. A method for building a virtual object in an artificial reality (XR) environment, the method comprising: receiving, by an application, input from a user, wherein the input comprises a verbal component; determining that the input relates to a request to build an object; identifying, based at least in part on the verbal component, an object type associated with the virtual object; building the virtual object based at least in part on the object type; identifying a location associated with the virtual object; and placing the virtual object in the XR environment according to the identified location. 1. A method for building a virtual object in an XR world, the method comprising: receiving, by an artificial intelligence ("AI"), a command from a user, wherein the command is associated with one or more images and the artificial intelligence is represented by a non-player character (NPC) in the XR world; determining that the command is an object build command, wherein the determining is based on A) a determination that the user's attention is directed at the NPC, B) that the command includes the one or more images, and C) that the command does not indicate an existing virtual object in the XR world; parsing a textual representation, of part of the command, for object type information and object location information; building a 3D virtual object based on the one or more images and the object type information; identifying a location, in the XR world, based on the object location information from the command and a direction determined for the user's attention; and placing the built 3D virtual object in the XR world according to the identified location. US App #19009478 Claim 21 US Patent #12254564 Claim 13 21. A method for building a virtual object in an artificial reality (XR) environment, the method comprising: receiving, by an application, input from a user, wherein the input comprises a verbal component; determining that the input relates to a request to build an object; identifying, based at least in part on the verbal component, an object type associated with the virtual object; building the virtual object based at least in part on the object type; identifying a location associated with the virtual object; and placing the virtual object in the XR environment according to the identified location. 13. A computer-readable storage medium storing instructions that, when executed by a computing system, cause the computing system to perform a process for building a virtual object in an XR world, the process comprising: receiving, by an artificial intelligence ("AI"), a command from a user, wherein the artificial intelligence is represented by a non-player character (NPC) in the XR world; determining that the command is an object build command, wherein the determining is based on A) a determination that the user's attention is directed at the NPC and B) that the command does not indicate an existing virtual object in the XR world; parsing a textual representation, of part of the command, for object type information and object location information; building a 3D virtual object using a template, from a 3D model library, matching the object type information identifying a location, in the XR world, based on the object location information from the command and a direction determined for the user's attention and placing the built 3D virtual object in the XR world according to the identified location. US App #19009478 Claim 24 US Patent #12254564 Claim 13 24. (New) The method of claim 21, wherein the building the virtual object includes using a template, from a model library, matching the object type. 13. A computer-readable storage medium storing instructions that, when executed by a computing system, cause the computing system to perform a process for building a virtual object in an XR world, the process comprising: receiving, by an artificial intelligence ("AI"), a command from a user, wherein the artificial intelligence is represented by a non-player character (NPC) in the XR world; determining that the command is an object build command, wherein the determining is based on A) a determination that the user's attention is directed at the NPC and B) that the command does not indicate an existing virtual object in the XR world; parsing a textual representation, of part of the command, for object type information and object location information; building a 3D virtual object using a template, from a 3D model library, matching the object type information identifying a location, in the XR world, based on the object location information from the command and a direction determined for the user's attention and placing the built 3D virtual object in the XR world according to the identified location. US App #19009478 Claim 26 US Patent #12254564 Claim 2 26. The method of claim 25, wherein the one or more images are obtained through a process that guides the user in capturing multiple images of a real-world object to import the real-world object into the XR environment. 2. The method of claim 1, wherein the one or more images are provided through a process that guides the user in capturing multiple images of a real-world object to import into the XR world. US App #19009478 Claim 28 US Patent #12254564 Claim 1 28. The method of claim 21, wherein the virtual object is a 3D virtual object. 1. A method for building a virtual object in an XR world, the method comprising: receiving, by an artificial intelligence ("AI"), a command from a user, wherein the command is associated with one or more images and the artificial intelligence is represented by a non-player character (NPC) in the XR world; determining that the command is an object build command, wherein the determining is based on A) a determination that the user's attention is directed at the NPC, B) that the command includes the one or more images, and C) that the command does not indicate an existing virtual object in the XR world; parsing a textual representation, of part of the command, for object type information and object location information; building a 3D virtual object based on the one or more images and the object type information; identifying a location, in the XR world, based on the object location information from the command and a direction determined for the user's attention; and placing the built 3D virtual object in the XR world according to the identified location. US App #19009478 Claim 35 US Patent #12254564 Claim 1 35. The method of claim 21, wherein the determining that the input relates to the request to build the object is based on a determination that the input does not indicate an existing virtual object in the XR environment. 1. A method for building a virtual object in an XR world, the method comprising: receiving, by an artificial intelligence ("AI"), a command from a user, wherein the command is associated with one or more images and the artificial intelligence is represented by a non-player character (NPC) in the XR world; determining that the command is an object build command, wherein the determining is based on A) a determination that the user's attention is directed at the NPC, B) that the command includes the one or more images, and C) that the command does not indicate an existing virtual object in the XR world; parsing a textual representation, of part of the command, for object type information and object location information; building a 3D virtual object based on the one or more images and the object type information; identifying a location, in the XR world, based on the object location information from the command and a direction determined for the user's attention; and placing the built 3D virtual object in the XR world according to the identified location. US App #19009478 Claim 36 US Patent #12254564 Claim 7 36. The method of claim 35, wherein the determination that the input does not indicate an existing virtual object is based on a determined gaze direction of the user. 7. The method of claim 1, wherein the determination that the command does not indicate an existing virtual object in the XR world is based on a determined gaze of the user. US App #19009478 Claim 37 US Patent #12254564 Claim 12 37. The method of claim 21, wherein the identifying the location associated with the virtual object includes: applying a set of ranking rules, to particular locations, to generate a rank sore for locations based on X) an estimation for where the user was looking or gesturing when the user provided the input, Y) a relevance between the virtual object being built and other objects near the particular location, and Z) a match between location information identified from the input and the particular location; and selecting the highest ranked location. 12. The method of claim 1, wherein the identifying the location in the XR world comprises: applying a set of ranking rules, to particular locations, to generate a rank sore for locations based on X) an estimation for where the user was looking or gesturing when the user made the command, Y) a relevance between the virtual object being built and other objects near the particular location, and Z) a match between the object location information and the particular location; and selecting the highest ranked location. US App #19009478 Claim 38 US Patent #12254564 Claim 13 38. A non-transitory computer-readable storage medium storing instructions, for building a virtual object in an artificial reality (XR) environment, the instructions, when executed by a computing system, cause the computing system to: receiving, by an application, an input from a user, wherein the input comprises a verbal component; determining that the input relates to a request to build an object and, in response: identifying, based at least in part on the verbal component, an object type associated with the virtual object; and building the virtual object based at least in part on the object type; identifying a location associated with the virtual object; and placing the virtual object in the XR environment according to the identified location. 13. A computer-readable storage medium storing instructions that, when executed by a computing system, cause the computing system to perform a process for building a virtual object in an XR world, the process comprising: receiving, by an artificial intelligence ("AI"), a command from a user, wherein the artificial intelligence is represented by a non-player character (NPC) in the XR world; determining that the command is an object build command, wherein the determining is based on A) a determination that the user's attention is directed at the NPC and B) that the command does not indicate an existing virtual object in the XR world; parsing a textual representation, of part of the command, for object type information and object location information; building a 3D virtual object using a template, from a 3D model library, matching the object type information identifying a location, in the XR world, based on the object location information from the command and a direction determined for the user's attention and placing the built 3D virtual object in the XR world according to the identified location. Although the claims at issue are not identical, they are not patentably distinct from each other. For example, Claim 1 of Patent 12254564 discloses "receiving, by an artificial intelligence ("AI"), a command from a user" while Application 19009478 Claim 21 discloses "receiving, by an application, input from a user. Artificial intelligence is implemented as computer application. Also command from a user is an input to computer system. Therefore, Patent 12254564 Claim 1 discloses all limitations of Application 19009478 Claim 21. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 21-24, 38 are rejected under 35 U.S.C. 103 as being unpatentable over Petill et al (US20210012113) in view of Marzorati et al (US20220398401) further in view of Scavezze et al (US20130335405). Regarding Claim 21. Petill teaches A method for building a virtual object in an artificial reality (XR) environment (Petill, abstract, the invention describes a head mounted display device including a display device, a camera device, an input device, and a processor. The processor is configured to store a database of physical objects and virtual objects that have been associated with one or more semantic tags. The processor is further configured to receive a natural language input from a user via the input device and perform semantic processing on the natural language input to determine a user specified operation and identify one or more semantic tags indicated by the natural language input. The processor is further configured to select a target virtual object and a target physical object based on the identified one or more semantic tags, perform the determined user specified operation on the target virtual object based on the target physical object, and display the target virtual object at a physical location associated with the target physical object.), the method comprising: receiving, by an application, input from a user, wherein the input comprises a verbal component (Petill, [0101] The following paragraphs provide additional support for the claims of the subject application. One aspect provides a head mounted display device comprising a display device configured to display virtual objects at locations in a physical environment, a camera device configured to capture images of the physical environment, an input device configured to receive a user input. and a processor. The processor is configured to store a database of physical objects and virtual objects that have been associated with one or more semantic tags. The processor is further configured to receive a natural language input from a user via the input device, and perform semantic processing on the natural language input to determine a user specified operation and identify one or more semantic tags indicated by the natural language input. … In this aspect, additionally or alternatively, the natural language input may be a voice input received via the input device. In this aspect, additionally or alternatively, the determined user specified operation may be a move operation, and to perform the determined user specified operation, the processor may be further configured to update a location of the target virtual object in the physical environment based on the physical location associated with the target physical object, and display the target virtual object at the updated location via the display device. In this aspect, additionally or alternatively, the determined user specified operation may be an application start operation, and to perform the determined user specified operation, the processor may be further configured to select a target application program from a plurality of application programs executable by the processor based on the identified one or more semantic tags, generate a target virtual object associated with the target application program, and display the generated target virtual object at a virtual location based on the physical location associated with the target physical object.); determining that the input relates to a request to build an object (Petill, [0101] The following paragraphs provide additional support for the claims of the subject application. One aspect provides a head mounted display device comprising a display device configured to display virtual objects at locations in a physical environment, a camera device configured to capture images of the physical environment, an input device configured to receive a user input. and a processor. The processor is configured to store a database of physical objects and virtual objects that have been associated with one or more semantic tags. The processor is further configured to receive a natural language input from a user via the input device, and perform semantic processing on the natural language input to determine a user specified operation and identify one or more semantic tags indicated by the natural language input. … In this aspect, additionally or alternatively, the determined user specified operation may be an application start operation, and to perform the determined user specified operation, the processor may be further configured to select a target application program from a plurality of application programs executable by the processor based on the identified one or more semantic tags, generate a target virtual object associated with the target application program, and display the generated target virtual object at a virtual location based on the physical location associated with the target physical object.); Petill fails to explicitly teach, however, Marzorati teaches identifying, based at least in part on the verbal component, an object type associated with the virtual object (Marzorati, abstract, the invention describes method for creating a template for user experience by segmenting a visual surrounding. The embodiment may include receiving real-time and historical data relating to one or more content interactions of a user wearing an augmented reality (AR) device. The embodiment may also include analyzing one or more contextual situations of the one or more content interactions. The embodiment may further include identifying one or more objects of interest in a visual surrounding environment of the user. The embodiment may also include in response to determining the identification of the object type is confident, predicting a contextual need for each object of interest. The embodiment may further include creating one or more information display templates. The embodiment may also include populating the one or more information display templates with information and displaying the one or more populated information display templates to the user. [0046] Then, at 210, the template creation program 110A, 110B receives feedback from the user about the object type for the at least one object of interest whose identification is not confident. According to at least one embodiment, the feedback may be a voice command from the user stating the object type for the at least one object of interest whose identification is not confident. For example, the template creation program 110A, 110B may prompt the user for feedback, and in response the user may audibly state, "This object is a consumer product" or "This object is a food item." The user may also be more specific in identifying an object type. For example, rather than stating, "This object is a consumer product," the user may state, "This object is a stereo." According to at least one other embodiment, the template creation program 110A, 110B may have a low confidence prediction for the object type, e.g., 30%. In such embodiments, a template may be presented to the user and the user would be able to annotate the template with the object type. Either embodiment allows for high quality labeled data for semi-supervised ML to increase the number and quality of identifiable objects in a template library.); Petill and Marzorati are analogous art because they both teach method of receive user’s voice command as input to further trigger operation in AR/VR environment. Marzorati further teaches user’s verbal command contains information of object type data. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AR/VR environment interaction system (taught in Petill), to further identify object type through user verbal command (taught in Marzorati), so as to improve user’s interaction experience in AR/VR environment (Marzorati, [0001-0002]). The combination of Petill and Marzorati fails to explicitly teach, however, Scavezze teaches building the virtual object based at least in part on the object type (Scavezze, abstract, the invention describes a system and method for building and experiencing three-dimensional virtual objects from within a virtual environment in which they will be viewed upon completion. A virtual object may be created, edited and animated using a natural user interface while the object is displayed to the user in a three-dimensional virtual environment. [0027] As explained below, aspects of the present system allow users to generate virtual objects that are displayed three-dimensionally to the user as they are being created. The hub computing system may execute a content-generation software application, which constructs virtual objects within the virtual environment in accordance with input received from the user. As utilized herein, the term "user" may refer to a content creator using a mixed reality system to create, edit and animate virtual objects. The term "end user" may refer to those who thereafter experience the completed virtual objects using a mixed reality system. Page 37, claim 2, The system of claim 1, wherein the computing system generates a virtual object by creating the virtual object in the virtual environment in response to gestures from the user indicating the type of virtual object to be created in the virtual environment.); Petill, Marzorati and Scavezze are analogous art because they all teach method of receive user’s voice command as input to further trigger operation in AR/VR environment. Scavezze further teaches creating virtual object based on object type indicated in user’s command. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AR/VR environment interaction system (taught in Petill and Marzorati), to further creating virtual object based on object type indicated in user’s command (taught in Scavezze), so as to provide an intuitive system to build and experience 3D virtual objects from within a virtual environment (Scavezze, [0001-0003]). The combination of Petill, Marzorati and Scavezze further teaches identifying a location associated with the virtual object; and placing the virtual object in the XR environment according to the identified location (Petill, [0101] The following paragraphs provide additional support for the claims of the subject application. One aspect provides a head mounted display device comprising a display device configured to display virtual objects at locations in a physical environment, a camera device configured to capture images of the physical environment, an input device configured to receive a user input. and a processor. The processor is configured to store a database of physical objects and virtual objects that have been associated with one or more semantic tags. The processor is further configured to receive a natural language input from a user via the input device, and perform semantic processing on the natural language input to determine a user specified operation and identify one or more semantic tags indicated by the natural language input. The processor is further configured to select a target virtual object and a target physical object from the physical objects and virtual objects in the database based on the identified one or more semantic tags, perform the determined user specified operation on the target virtual object based on the target physical object, and display the target virtual object at a physical location associated with the target physical object.). Regarding Claim 22. The combination of Petill, Marzorati and Scavezze further teaches The method of claim 21 further comprising: identifying a gesture performed by the user; wherein the determination that the input relates to the request to build the object is based, at least in part, on the gesture (Scavezze, page 37, claim 2, The system of claim 1, wherein the computing system generates a virtual object by creating the virtual object in the virtual environment in response to gestures from the user indicating the type of virtual object to be created in the virtual environment.). The reasoning for combination of Petill, Marzorati and Scavezze is the same as described in Claim 21. Regarding Claim 23. The combination of Petill, Marzorati and Scavezze further teaches The method of claim 21 further comprising: identifying a gesture performed by the user; wherein the identifying the location associated with the virtual object is based, at least in part, on the gesture (Petill, Petill, [0101] The following paragraphs provide additional support for the claims of the subject application. One aspect provides a head mounted display device comprising a display device configured to display virtual objects at locations in a physical environment, a camera device configured to capture images of the physical environment, an input device configured to receive a user input. and a processor. … In this aspect, additionally or alternatively, the determined user specified operation may be a move operation, and to perform the determined user specified operation, the processor may be further configured to update a location of the target virtual object in the physical environment based on the physical location associated with the target physical object, and display the target virtual object at the updated location via the display device. … In this aspect, additionally or alternatively, the processor may be further configured to determine a user indicated direction for a user of the head mounted display device, and select the target virtual object or the target physical object further based on the determined user indicated direction. In this aspect, additionally or alter natively, the user indicated direction may be determined based on a detected gaze direction of the user or a detected hand gesture of the user.). Regarding Claim 24. The combination of Petill, Marzorati and Scavezze further teaches The method of claim 21, wherein the building the virtual object includes using a template, from a model library, matching the object type (Petill, [0101] The following paragraphs provide additional support for the claims of the subject application. One aspect provides a head mounted display device comprising a display device configured to display virtual objects at locations in a physical environment, a camera device configured to capture images of the physical environment, an input device configured to receive a user input. and a processor. The processor is configured to store a database of physical objects and virtual objects that have been associated with one or more semantic tags. The processor is further configured to receive a natural language input from a user via the input device, and perform semantic processing on the natural language input to determine a user specified operation and identify one or more semantic tags indicated by the natural language input. The processor is further configured to select a target virtual object and a target physical object from the physical objects and virtual objects in the database based on the identified one or more semantic tags, perform the determined user specified operation on the target virtual object based on the target physical object, and display the target virtual object at a physical location associated with the target physical object.). Claim 38 is similar in scope as Claim 21, and thus is rejected under same rationale. Claim 38 further requires: A non-transitory computer-readable storage medium (Petill, [0090] Computing system 1700 includes a logic processor 1702 volatile memory 1704, and a non-volatile storage device 1706. Computing system 1700 may optionally include a display subsystem 1708, input subsystem 1710, communication subsystem 1712, and/or other components not shown in FIG. 17.). Claims 25-26, 28 are rejected under 35 U.S.C. 103 as being unpatentable over Petill et al in view of Marzorati et al, Scavezze et al further in view of Thomas et al (US20160189426). Regarding Claim 25. The combination of Petill, Marzorati and Scavezze fails to explicitly teach, however, Thomas teaches The method of claim 21 further comprising: performing an analysis of one or more images associated with the input; wherein the building the virtual object is based on the analysis of the one or more images (Thomas, abstract, the invention describes methods for generating virtual proxy objects and controlling the location of the virtual proxy objects within an augmented reality environment. In some embodiments, a head-mounted display device (HMO) may identify a real world object for which to generate a virtual proxy object, generate the virtual proxy object corresponding with the real world object, and display the virtual proxy object using the HMD such that the virtual proxy object is perceived to exist within an augmented reality environment displayed to an end user of the HMD. In some cases, image processing techniques may be applied to depth images derived from a depth camera embedded within the HMD in order to identify boundary points for the real-world object and to determine the dimensions of the virtual proxy object corresponding with the real world object. [0079] In step 522, a real-world object is identified using a mobile device. The mobile device may comprise an HMD. The real-world object may comprise, for example, a chair, a table, a couch, a piece of furniture, a painting, or a picture frame. In step 524, a height, a length, and a width of the real-world object is determined. [0081] In step 526, a maximum number of surfaces is determined. In some cases, the maximum number of surfaces may be set by an end user of the mobile device. In step 528, a three-dimensional model is determined based on the maximum number of surfaces. [0082] In step 530, a virtual proxy object is generated using the three-dimensional model and the height, the length, and the width of the real-world object determined in step 524.). Petill, Marzorati, Scavezze and Thomas are analogous art because they all teach method of receive user’s voice command as input to further trigger operation in AR/VR environment. Thomas further teaches creating virtual object based on detected real world object. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AR/VR environment interaction system (taught in Petill, Marzorati and Scavezze), to further creating virtual object based on detected real world object type (taught in Thomas), so as to realistically integrate virtual objects into AR/VR environment (Thomas, [0001-0002]). Regarding Claim 26. The combination of Petill, Marzorati, Scavezze and Thomas further teaches The method of claim 25, wherein the one or more images are obtained through a process that guides the user in capturing multiple images of a real world object to import the real-world object into the XR environment (Thomas, [0081] In step 526, a maximum number of surfaces is determined. In some cases, the maximum number of surfaces may be set by an end user of the mobile device. In step 528, a three-dimensional model is determined based on the maximum number of surfaces. In one example, if the maximum number of surfaces comprises six surfaces, then the three dimensional model may comprise a rectangular prism. In another example, if the maximum number of surfaces comprises eight surfaces, then the three-dimensional model may comprise a hexagonal prism. Therefore, in order to accurately detect the true dimension, shape and number of surfaces of the real world object, it is obvious to a person with ordinary skill in the art to provide guidance for the user to capture one or more images of such object.). The reasoning for combination of Petill, Marzorati, Scavezze and Thomas is the same as described in Claim 24. Regarding Claim 28. The combination of Petill, Marzorati, Scavezze and Thomas further teaches The method of claim 21, wherein the virtual object is a 3D virtual object (Thomas, [0082] In step 530, a virtual proxy object is generated using the three-dimensional model and the height, the length, and the width of the real-world object determined in step 524.). The reasoning for combination of Petill, Marzorati, Scavezze and Thomas is the same as described in Claim 24. Claim 27 is rejected under 35 U.S.C. 103 as being unpatentable over Petill et al in view of Marzorati et al, Scavezze et al further in view of Santhar et al (US20230154126). Regarding Claim 27. The combination of Petill, Marzorati and Scavezze fails to explicitly teach, however, Santhar teaches The method of claim 21, wherein the building the virtual object includes applying a generative machine learning model, trained to produce virtual objects based on user verbal input (Santhar, abstract, the invention describes method for moving a virtual object within virtual space in response to an external input supplied by a user. A machine learning model may predict a movement of the virtual object and implement such movement in a next frame of the virtual space. An example operation may include one or more of receiving a measurement of an external input of a user with respect to a vi1tual object displayed in virtual space, predicting, via execution of a machine learning model, a movement of the virtual object in the virtual space in response to the external input of the user based on the measurement of the external input of the user, and moving the virtual object in the virtual space based on the predicted movement of the virtual object by the machine learning model. [0064] In some embodiments, the machine learning model may include a convolutional neural network (CNN) layer configured to identify a bounding box corresponding to the virtual object after the external stimulus of the user created by the user interaction. In some embodiments, the machine learning model may further include a generative adversarial network (GAN) which receives the bounding box from the CNN and determines a location of the virtual object in virtual space based on the external stimulus of the user created by the user interaction and the bounding box.). Petill, Marzorati, Scavezze and Santhar are analogous art because they all teach method of receive user’s voice command as input to further trigger operation in AR/VR environment. Santhar further teaches using generative machine learning model in manipulating virtual object. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AR/VR environment interaction system (taught in Petill, Marzorati and Scavezze), to further use generative machine learning model in manipulating virtual object (taught in Santhar), so as to predict the movement of the created virtual object in the virtual space (Santhar, [0001-0006]). Claims 29-32 are rejected under 35 U.S.C. 103 as being unpatentable over Petill et al in view of Marzorati et al, Scavezze et al further in view of Faulkner et al (US20220091723). Regarding Claim 29. The combination of Petill, Marzorati and Scavezze fails to explicitly teach, however, Faulkner teaches The method of claim 21, wherein the input is received by an artificial intelligence ("Al") agent represented by an Al Agent (Faulkner, abstract, the invention describes a computing system, while displaying a first view of a first computer-generated three-dimensional environment including a representation of a respective portion of a physical environment, and a first representation of one or more projections of light in a first portion of the first computer generated three-dimensional environment, detects, from a first user, a query directed to a virtual assistant. In response, the computer system displays animated changes of the first representation of the one or more projections of light in the first portion of the first computer-generated three-dimensional environment, including displaying a second representation of the one or more projections of light that is focused on a first sub-portion of the first portion, and then displays content responding to the query at a position corresponding to the first sub-portion of the first portion. [0093] As described herein, a virtual assistant is an embodiment of a function of the computer system that assists the user in a variety of situations and tasks based on contextual information (e.g., location, time, schedule, past interactions, social contacts, currently displayed application or experience, recently accessed applications and experiences, recent interactions with the virtual assistant, etc.) and the user's request. In some embodiments, the virtual assistant has an identity that is associated with a persona, such as an animated character, a virtual animal or person, a robot, etc. In some embodiments, the inputs directed to the virtual assistant include natural language inputs and/or speech from the user, as well as other types of inputs, such as gesture, touch, controller inputs, etc. In some embodiments, the responses provided by the virtual assistant includes natural language responses and/or speech, along with visual content responsive to the user's request. In some embodiments, the virtual assistant includes an artificial intelligence component that learns from past interactions with the user, or with a large number of users, to provide more accurate responses to the user.). Petill, Marzorati, Scavezze and Faulkner are analogous art because they all teach method of receive user’s voice command as input to further trigger operation in AR/VR environment. Faulkner further teaches using AI virtual assistant to interact with user in a variety of situations and tasks. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AR/VR environment interaction system (taught in Petill, Marzorati and Scavezze), to further use teaches using AI virtual assistant to interact with user in a variety of situations and tasks (taught in Faulkner), so as to provide user with support in performing actions associated with virtual objects (Faulkner, [0004]). Regarding Claim 30. The combination of Petill, Marzorati, Scavezze and Faulkner further teaches The method of claim 29, wherein the Al Agent is controlled by multiple users, such that commands from a later user are implemented, in relation to objects built by the Al, based on commands from an earlier user (Faulkner, [0093] As described herein, a virtual assistant is an embodiment of a function of the computer system that assists the user in a variety of situations and tasks based on contextual information (e.g., location, time, schedule, past interactions, social contacts, currently displayed application or experience, recently accessed applications and experiences, recent interactions with the virtual assistant, etc.) and the user's request. In some embodiments, the virtual assistant has an identity that is associated with a persona, such as an animated character, a virtual animal or person, a robot, etc. In some embodiments, the inputs directed to the virtual assistant include natural language inputs and/or speech from the user, as well as other types of inputs, such as gesture, touch, controller inputs, etc. In some embodiments, the responses provided by the virtual assistant includes natural language responses and/or speech, along with visual content responsive to the user's request. In some embodiments, the virtual assistant includes an artificial intelligence component that learns from past interactions with the user, or with a large number of users, to provide more accurate responses to the user.). The reasoning for combination of Petill, Marzorati, Scavezze and Faulkner is the same as described in Claim 29. Regarding Claim 31. The combination of Petill, Marzorati, Scavezze and Faulkner further teaches The method of claim 29, wherein the Al Agent is a voice interface provided by an XR system (Faulkner, [0093] As described herein, a virtual assistant is an embodiment of a function of the computer system that assists the user in a variety of situations and tasks based on contextual information (e.g., location, time, schedule, past interactions, social contacts, currently displayed application or experience, recently accessed applications and experiences, recent interactions with the virtual assistant, etc.) and the user's request. In some embodiments, the virtual assistant has an identity that is associated with a persona, such as an animated character, a virtual animal or person, a robot, etc. In some embodiments, the inputs directed to the virtual assistant include natural language inputs and/or speech from the user, as well as other types of inputs, such as gesture, touch, controller inputs, etc. In some embodiments, the responses provided by the virtual assistant includes natural language responses and/or speech, along with visual content responsive to the user's request. In some embodiments, the virtual assistant includes an artificial intelligence component that learns from past interactions with the user, or with a large number of users, to provide more accurate responses to the user.). The reasoning for combination of Petill, Marzorati, Scavezze and Faulkner is the same as described in Claim 29. Regarding Claim 32. The combination of Petill, Marzorati, Scavezze and Faulkner further teaches The method of claim 29, wherein the Al Agent is a character representation in the XR environment (Faulkner, [0093] As described herein, a virtual assistant is an embodiment of a function of the computer system that assists the user in a variety of situations and tasks based on contextual information (e.g., location, time, schedule, past interactions, social contacts, currently displayed application or experience, recently accessed applications and experiences, recent interactions with the virtual assistant, etc.) and the user's request. In some embodiments, the virtual assistant has an identity that is associated with a persona, such as an animated character, a virtual animal or person, a robot, etc. In some embodiments, the inputs directed to the virtual assistant include natural language inputs and/or speech from the user, as well as other types of inputs, such as gesture, touch, controller inputs, etc. In some embodiments, the responses provided by the virtual assistant includes natural language responses and/or speech, along with visual content responsive to the user's request. In some embodiments, the virtual assistant includes an artificial intelligence component that learns from past interactions with the user, or with a large number of users, to provide more accurate responses to the user.). The reasoning for combination of Petill, Marzorati, Scavezze and Faulkner is the same as described in Claim 29. Claim 33 is rejected under 35 U.S.C. 103 as being unpatentable over Petill et al in view of Marzorati et al, Scavezze et al further in view of Maher et al ("Designworld: an augmented 3D virtual world for multidisciplinary, collaborative design." Proceedings of CAADRIA 2006 (2006): 133-142.). Regarding Claim 33. The combination of Petill, Marzorati and Scavezze fails to explicitly teach, however, Maher teaches The method of claim 21, wherein the building the virtual object is further based on a second input, comprising a second verbal component, from a second user different from the user (Maher, abstract, the paper introduces DesignWorld, a prototype system for enabling collaboration between designers from different disciplines who may be in different physical locations. DesignWorld consists of a 3D virtual world augmented with a number of web-based communication and design tools. DesignWorld uses agent technology to maintain different views of a single design in order to support multidisciplinary collaboration and address issues such as multiple representations of objects, versioning, ownership and relationships between objects from different disciplines. Page 137, par 1, A virtual world is a distributed, persistent, virtual space. People can interact with other people, objects or computer controlled agents using an avatar controlled by the mouse and keyboard. DesignWorld uses the Second Life (www.secondlife.com) virtual environment as the platform for design and collaboration. Second Life allows collaborative manipulation and visualization of shared objects, both synchronously and asynchronously. Designers can select from a range of primitive objects with which to build, which can then be further modified using a number of built-in tools to achieve more complex objects.). Petill, Marzorati, Scavezze and Maher are analogous art because they all teach method of receive user’s command as input to further trigger operation in AR/VR environment. Maher further teaches collaborative environment which allows multiple users to modify the virtual object/world synchronously and asynchronously. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AR/VR environment interaction system (taught in Petill, Marzorati and Scavezze), to further use collaborative environment which allows multiple users to modify the virtual object/world (taught in Maher), so as to provide designer user to work with professional architects in creating complex virtual environment (Maher, page 141, par 3). Claim 39 is rejected under 35 U.S.C. 103 as being unpatentable over Petill et al in view of Marzorati et al, Scavezze et al further in view of Thomas et al (US20160189426), Santhar et al (US20230154126). Claim 39 is similar in scope as Claim 24, 25, 27, and thus is rejected under same rationale. Petill, Marzorati, Scavezze, Thomas and Santhar are analogous art because they all teach method of receive user’s voice command as input to further trigger operation in AR/VR environment. Santhar further teaches using generative machine learning model in manipulating virtual object. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AR/VR environment interaction system (taught in Petill, Marzorati, Scavezze and Thomas), to further use generative machine learning model in manipulating virtual object (taught in Santhar), so as to predict the movement of the created virtual object in the virtual space (Santhar, [0001-0006]). Claim 40 is rejected under 35 U.S.C. 103 as being unpatentable over Petill et al in view of Faulkner et al (US20220091723) further in view of Yilanci et al (US20220148268) Regarding Claim 40. Petill teaches A computing system for building a virtual object in an XR environment (Petill, abstract, the invention describes a head mounted display device including a display device, a camera device, an input device, and a processor. The processor is configured to store a database of physical objects and virtual objects that have been associated with one or more semantic tags. The processor is further configured to receive a natural language input from a user via the input device and perform semantic processing on the natural language input to determine a user specified operation and identify one or more semantic tags indicated by the natural language input. The processor is further configured to select a target virtual object and a target physical object based on the identified one or more semantic tags, perform the determined user specified operation on the target virtual object based on the target physical object, and display the target virtual object at a physical location associated with the target physical object.), the computing system comprising: one or more processors (Petill, [0090] Computing system 1700 includes a logic processor 1702 volatile memory 1704, and a non-volatile storage device 1706. Computing system 1700 may optionally include a display subsystem 1708, input subsystem 1710, communication subsystem 1712, and/or other components not shown in FIG. 17.); and one or more memories storing instructions that, when executed by the one or more processors (Petill, [0090] Computing system 1700 includes a logic processor 1702 volatile memory 1704, and a non-volatile storage device 1706. Computing system 1700 may optionally include a display subsystem 1708, input subsystem 1710, communication subsystem 1712, and/or other components not shown in FIG. 17.), cause the computing system to: Petill fails to explicitly teach, however, Faulkner teaches receive, by an artificial intelligence ("Al"), a command from a user, wherein the artificial intelligence is represented by an Al Agent (Faulkner, abstract, the invention describes a computing system, while displaying a first view of a first computer-generated three-dimensional environment including a representation of a respective portion of a physical environment, and a first representation of one or more projections of light in a first portion of the first computer generated three-dimensional environment, detects, from a first user, a query directed to a virtual assistant. In response, the computer system displays animated changes of the first representation of the one or more projections of light in the first portion of the first computer-generated three-dimensional environment, including displaying a second representation of the one or more projections of light that is focused on a first sub-portion of the first portion, and then displays content responding to the query at a position corresponding to the first sub-portion of the first portion. [0093] As described herein, a virtual assistant is an embodiment of a function of the computer system that assists the user in a variety of situations and tasks based on contextual information (e.g., location, time, schedule, past interactions, social contacts, currently displayed application or experience, recently accessed applications and experiences, recent interactions with the virtual assistant, etc.) and the user's request. In some embodiments, the virtual assistant has an identity that is associated with a persona, such as an animated character, a virtual animal or person, a robot, etc. In some embodiments, the inputs directed to the virtual assistant include natural language inputs and/or speech from the user, as well as other types of inputs, such as gesture, touch, controller inputs, etc. In some embodiments, the responses provided by the virtual assistant includes natural language responses and/or speech, along with visual content responsive to the user's request. In some embodiments, the virtual assistant includes an artificial intelligence component that learns from past interactions with the user, or with a large number of users, to provide more accurate responses to the user.). Petill and Faulkner are analogous art because they both teach method of receive user’s voice command as input to further trigger operation in AR/VR environment. Faulkner further teaches using AI virtual assistant to interact with user in a variety of situations and tasks. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AR/VR environment interaction system (taught in Petill), to further use teaches using AI virtual assistant to interact with user in a variety of situations and tasks (taught in Faulkner), so as to provide user with support in performing actions associated with virtual objects (Faulkner, [0004]).; The combination of Petill and Faulkner fails to explicitly teach, however, Yilanci teaches determine that the command is for an object edit command (Yilanci, abstract, the invention teaches methods for generating and providing a virtual experience. The virtual experience includes digital elements that are generated and provided according to settings as well as user feedback and feedback from sensors in an environment that is being represented by the virtual experience. The virtual experience includes a real world image of the environment and virtual objects overlapping the real world image.); identify a type of the object edit command; map i) portions of the command and/or an identified context of the command to ii) parameters of the object edit command; and edit the existing virtual object by executing the identified type of the object edit command, with the mapped parameters (Yilanci, [0049] Virtual experience creation module 220, initial spawn location, and further movements of the virtual objects are demonstrated in FIG. 5. As it is illustrated further in FIGS. 6 to 10, virtual experience includes a medium 244 constructed by two components: virtual object pool 240 and object parameters 242. Object pool includes every element that can be used, and object parameters dictate the roles of said objects on the medium. Specifically, in the embodiment shown in FIG. 7, the medium uses three fish models, a bubble model, and an ambient soundtrack from pool 240. The route, animation, location, size, number of the fishes and bubbles, the boundary they can move in are denoted by parameters 242. In the next step, the positioning module 246 determines the initial spawn 248 positions of the medium. After spawning, positional parameters of the medium update in response to user inputs 250 obtained by positioning module 246. In an embodiment, examples of the initial position of user 238 and virtual objects 302 are shown in FIG. 7. In this embodiment, fish are in a flock movement circling around user 238, and following the user 238, as the user 238 moves around the house, and also continuing to circle around the user, and avoiding the boundaries of room. For example, some methods of AR utilize SLAM, or LIDAR scan. These methods can detect physical boundaries of the real world for virtual objects to fill in. This process is realized by the positioning module 246. In an embodiment, as shown in FIG. 8, the virtual experience object 302 can be spawned and positioned within detected physical boundaries in front of the user 238 by gaze sensing and eye tracking utilizing HMD 502.). Petill, Faulkner and Yilanci are analogous art because they all teach method of receive user’s voice command as input to further trigger operation in AR/VR environment. Yilanci further teaches update/edit virtual object based on corresponding associated movement parameters. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AR/VR environment interaction system (taught in Petill and Faulkner), to further update/edit virtual object based on corresponding associated movement parameters (taught in Yilanci), so as to provide user with seamless creation method of virtual experiences (Yilanci, [0050]).; Allowable Subject Matter Claims 34-37 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Regarding Claim 34, it recites “The method of claim 21, wherein the determining that the input relates to the request to build the object is based on A) a determination that the user's attention is directed at an Al Agent and B) that the input does not indicate an existing virtual object in the XR environment” in the context of Claim 34. The prior arts of record either alone or in combination fails to teach or suggest the above quoted limitation of Claim 34. Therefore, Claim 34 is allowable over prior art. Regarding Claim 35, it recites “The method of claim 21, wherein the determining that the input relates to the request to build the object is based on a determination that the input does not indicate an existing virtual object in the XR environment” in the context of Claim 35. The prior arts of record either alone or in combination fails to teach or suggest the above quoted limitation of Claim 35. Therefore, Claim 35 is allowable over prior art. Claim 36 depend from Claim 35 with respective additional limitations. Therefore, Claim 36 is allowable over prior art. Regarding Claim 37, it recites “The method of claim 21, wherein the identifying the location associated with the virtual object includes: applying a set of ranking rules, to particular locations, to generate a rank sore for locations based on X) an estimation for where the user was looking or gesturing when the user provided the input, Y) a relevance between the virtual object being built and other objects near the particular location, and Z) a match between location information identified from the input and the particular location; and selecting the highest ranked location” in the context of Claim 37. The prior arts of record either alone or in combination fails to teach or suggest the above quoted limitation of Claim 37. Therefore, Claim 37 is allowable over prior art. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Choi et al (US20200410770), abstract, the invention describes an augmented reality (AR) providing method for recognizing a context using a neural network includes acquiring, by processing circuitry, a video; analyzing, by the processing circuitry, the video and rendering the video to arrange a virtual object on a plane included in the video; determining whether a scene change is present in a current frame by comparing the current frame included in the video with a previous frame; determining a context recognition processing status for the video based on the determining of whether the scene change is present in the current frame; and in response to determining that the context recognition processing status is true, analyzing at least one of the video or a sensing value received from a sensor using the neural network and calculating at least one piece of context information, and generating additional content to which the context information is applied and providing the additional content. Any inquiry concerning this communication or earlier communications from the examiner should be directed to XIN SHENG whose telephone number is (571)272-5734. The examiner can normally be reached M-F 9:30AM-3:30PM 6:00PM-8:30PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jason Chan can be reached at 5712723022. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Xin Sheng/Primary Examiner, Art Unit 2619
Read full office action

Prosecution Timeline

Jan 03, 2025
Application Filed
Aug 11, 2026
Non-Final Rejection mailed — §103, §DOUBLEPATENT (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743859
TECHNIQUES FOR INTERACTING WITH VIRTUAL AVATARS AND/OR USER REPRESENTATIONS
2y 9m to grant Granted Sep 22, 2026
Patent 12730514
METHOD AND SYSTEM TO PROVIDE MULTISENSORY DIGITAL INTERACTION EXPERIENCES
2y 3m to grant Granted Sep 08, 2026
Patent 12725312
REPRODUCTION APPARATUS, GENERATION APPARATUS, CONTROL METHOD, AND RECORDING MEDIUM
2y 4m to grant Granted Sep 01, 2026
Patent 12718502
WAYPOINT CREATION IN MAP DETECTION
2y 0m to grant Granted Aug 25, 2026
Patent 12705848
METHOD AND APPARATUS FOR STYLIZING THREE-DIMENSIONAL MODEL, ELECTRONIC DEVICE, AND STORAGE MEDIUM
2y 5m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
73%
Grant Probability
90%
With Interview (+16.8%)
2y 4m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 412 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month