Prosecution Insights
Last updated: October 02, 2026
Application No. 18/652,201

SCENE CREATION USING LANGUAGE MODELS

Non-Final OA §102§103
Filed
May 01, 2024
Priority
Nov 07, 2023 — provisional 63/596,729
Examiner
TRAN, JENNY NGAN
Art Unit
2615
Tech Center
2600 — Communications
Assignee
Roblox Corporation
OA Round
3 (Non-Final)
44%
Grant Probability
Moderate
3-4
OA Rounds
2m
Est. Remaining
78%
With Interview

Examiner Intelligence

Grants 44% of resolved cases
44%
Career Allowance Rate
4 granted / 9 resolved
-17.6% vs TC avg
Strong +33% interview lift
Without
With
+33.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
25 currently pending
Career history
41
Total Applications
across all art units

Statute-Specific Performance

§101
6.6%
-33.4% vs TC avg
§103
59.0%
+19.0% vs TC avg
§102
16.3%
-23.7% vs TC avg
§112
16.3%
-23.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 9 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of the Claims Claims 1-20 are currently pending in the present application, with claims 1, 9, and 16 being independent. Response to Amendments / Arguments Applicant’s arguments, see Pg. 7-11, filed 05/07/2026, with respect to the rejection(s) of claim(s) 1-20 under 35 U.S.C. § 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of newly found prior art. Regarding the remaining arguments: Applicant argues with respect to the amended claim language, which is fully addressed in the prior art rejections set forth below. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1-20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Kouzelis et. al. Synthesizing play-ready VR scenes with natural language prompts through GPT API. InInternational Symposium on Visual Computing 2023 Oct 16 (pp. 15-26). Cham: Springer Nature Switzerland, hereinafter referred to as “Kouzelis”. Regarding claim 1, Kouzelis discloses a computer-implemented method, the method comprising: initializing a large language model with an initialization prompt that includes instructions to the large language model to generate a virtual experience based on one or more subsequent prompts (Fig. 2 and Pg. 18, Section 3; Our methodology employs a streamlined process using an advanced language model for selecting and positioning 3D objects in Unity3D scenes. Utilizing the GPTAPI, this approach simplifies and improves virtual 3D object manipulation. Figure 2 outlines the method, which we detail here, focusing on the generation and interpretation of spatial and object data), wherein the instructions in the initialization prompt specify that overlap is impermissible when placing objects in the virtual experience (Pg. 19, Section 3.1; real-time parsing and prompt generating…extracting real-time dimensions for each furniture item from their mesh renderer component's bounds. This prevents object overlap by recording width, depth, and height. Fig. 3; You must take into account the dimensions of the furniture and the area, ensuring items do not overlap or float, and that they are oriented in a realistic manner…The furniture must be correctly positioned on the floor, not floating or overlapping…Be mindful of the furniture sizes so as they do not overlap…The door is at (12, 49, 0, 4) and should not be obstructed…)); receiving a user prompt, the user prompt comprising text criteria that specify criteria for generation or modification of the virtual experience, wherein the user prompt includes at least one of text data in natural language (Pg. 15-16, Section 1; we leverage the power of natural language processing (NLP) via a widely adopted large language model (LLM) which is the PGT API…translates simple natural language text into 3D scenes, as shown in Fig. 1), audio data, or video data (Fig. 3 and Pg. 20-21, Section 3.1; User input, shown in red, is brief but triggers extensive operations, as seen in the example "Create a living room". This blend of pre-determined, hard-coded elements, and adaptable components generated in real-time, facilitates the seamless adaptability to changes in the scene); identifying one or more objects in the virtual experience having one or more attributes that correspond to the text criteria, the one or more objects being identified by a large language model (Fig. 4 and Pg. 21-22, Section 3.2; This object contains an array of furniture data, where each entry corresponds to an individual piece. The "object_name" field acts as a key to access the corresponding prefab from our object data); determining spatial placement information in the virtual experience for the one or more objects using the large language model to interpret the text criteria to determine locations for the one or more objects in the virtual experience (Fig. 4 and Pg. 21, Section 3.2; Once retrieved, the prefab is instantiated within the Unity scene at the position specified by the X, Y, and Z values. Section 4.3; "A coffee table is in front of a sofa. A TV stand is opposite the coffee table"…"A dressing table is in front of a chair. A bed is next to the dressing table"…), wherein the large language model determines the locations based on the user prompt and the initialization prompt such that overlap between the one or more objects is avoided (Fig. 3 and Pg. 21-22, Section 3.2; Leveraging the information derived from the spatial layout and orientation of existing furniture within a scene, our system forms a comprehensive request to the GPT API, incorporating new elements…parses all objects present in the Unity3D scene, collecting their positions and orientations…); and placing the one or more objects in the virtual experience based on the spatial placement information (Pg. 22-25; Case Study 1-4 and Fig. 5-7). Regarding claim 2, Kouzelis discloses the computer-implemented method of claim 1, and further discloses further comprising modifying the virtual experience by changing an attribute of a specified object of the one or more objects in the virtual experience based on the text criteria, wherein the attribute comprises at least one of an appearance, a behavior, an orientation (Fig. 4 and Pg. 21, Section 3.2; In terms of object orientation in the JSON code, we used the "facing" key from the GPT API output. The facing direction of each object is determined by the vector between a _FRONT component and its position relative to the object pivot…if the GPT model's output indicates that a given furniture item should face another object, our system orients the object appropriately by rotating them based on this direction vector…Pg. 22, Section 3.2; parses all objects present in the Unity3D scene, collecting their positions and orientations…), a style, a material, a texture, a cost, a property, or another modifiable aspect of the specified object. Regarding claim 3, Kouzelis discloses the computer-implemented method of claim 1, and further discloses wherein the identifying of the one or more objects in the virtual experience comprises: generating one or more keywords using the large language model (Pg. 18, Section 3.1; we create two key data sets: one for the Unity3D scene plane and another for the available 3D objects. These data sets are the backbone of our GPT API-driven method for object selection and placement, ensuring system adaptability across varying scene dimensions and object types...flexible database of prefabs - reusable game objects that include 3D models along with their associated properties and behaviors…Pg. 22, Section 4; The complete list of objects available to the PGT API is the following: Armchair, basket, bed, bookshelf, candles, chair, coffee table, desk lamp, dresser, floor lamp, laptop, nightstand, notepad, office chair, sofa, table, tv, tv stand, vase); and performing a keyword search based on the keywords (Pg. 20-21, Section 3.2; we dynamically parse the database prefabs…array of furniture data, where each entry corresponds to an individual piece. The "object_name" field acts as a key to access the corresponding prefab from our object database…Pg. 22, Section 4.1; It is evident that the GPT API was able to select multiple pieces of furniture that are appropriate for the living room setup). Regarding claim 4, Kouzelis discloses the computer-implemented method of claim 1, and further discloses wherein the placing comprises placing objects such that overlap between the objects is avoided based on using object dimensions (Pg. 19, Section 3.1; real-time parsing and prompt generating…extracting real-time dimensions for each furniture item from their mesh renderer component's bounds. This prevents object overlap by recording width, depth, and height. Fig. 3; Be mindful of the furniture sizes so as they do not overlap…The door is at (12, 49, 0, 4) and should not be obstructed…). Regarding claim 5, Kouzelis discloses the computer-implemented method of claim 1, and further discloses wherein the user prompt comprises an updated prompt (Pg. 22-25; Case Study 1. Fig. 5; additional prompt "Try to fit in a dining area". Fig. 6; "Create a bedroom. Also, incorporate an office area". Fig. 7; "Create a living room. A coffee table is in front of a sofa. A TV stand is opposite the coffee table"). Regarding claim 6, Kouzelis discloses the computer-implemented method of claim 1, and further discloses providing, to a user, at least one of a view of the virtual experience including the one or more objects as placed or a summary of changes made to the virtual experience (Fig. 5-7). Regarding claim 7, Kouzelis discloses the computer-implemented method of claim 1, and further discloses wherein the large language model uses at least one of scene context and a history of user prompts to perform at least one of identifying the one or more objects or determining the spatial placement information (Pg. 22, Section 3.2; we input real-time parsed context from the existing scene, eliminating the need to resend past prompts. Section 4.1, Case Study 2; we use an already generated room…the previously generated living room, and parse the existing objects in the scene at runtime…). Regarding claim 8, Kouzelis discloses the computer-implemented method of claim 1, and further discloses wherein the large language model uses at least one macro obtained from the user prompt to perform at least one of identifying the one or more objects or determining the spatial placement information (Fig. 2; Inputs. User Prompt: "Create a living room"…Run-time parsing. Request -> GPT API -> Response…Construction. Parse JSON code. Instantiate Prefab armchair at (8.5, 0, 3.8) Make armchair face TV). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 9-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kouzelis et. al. Synthesizing play-ready VR scenes with natural language prompts through GPT API. InInternational Symposium on Visual Computing 2023 Oct 16 (pp. 15-26). Cham: Springer Nature Switzerland, hereinafter referred to as “Kouzelis”. Regarding claim 9, Kouzelis discloses initializing a large language model with an initialization prompt that includes instructions to the large language model to generate a virtual experience based on one or more subsequent prompts (Fig. 2 and Pg. 18, Section 3; Our methodology employs a streamlined process using an advanced language model for selecting and positioning 3D objects in Unity3D scenes. Utilizing the GPTAPI, this approach simplifies and improves virtual 3D object manipulation. Figure 2 outlines the method, which we detail here, focusing on the generation and interpretation of spatial and object data), wherein the instructions in the initialization prompt specify that overlap is impermissible when placing objects in the virtual experience (Pg. 19, Section 3.1; real-time parsing and prompt generating…extracting real-time dimensions for each furniture item from their mesh renderer component's bounds. This prevents object overlap by recording width, depth, and height. Fig. 3; You must take into account the dimensions of the furniture and the area, ensuring items do not overlap or float, and that they are oriented in a realistic manner…The furniture must be correctly positioned on the floor, not floating or overlapping…Be mindful of the furniture sizes so as they do not overlap…The door is at (12, 49, 0, 4) and should not be obstructed…)); receiving a user prompt, the user prompt comprising text criteria that specify criteria for generation or modification of the virtual experience, wherein the user prompt includes at least one of text data in natural language (Pg. 15-16, Section 1; we leverage the power of natural language processing (NLP) via a widely adopted large language model (LLM) which is the PGT API…translates simple natural language text into 3D scenes, as shown in Fig. 1), audio data, or video data (Fig. 3 and Pg. 20-21, Section 3.1; User input, shown in red, is brief but triggers extensive operations, as seen in the example "Create a living room". This blend of pre-determined, hard-coded elements, and adaptable components generated in real-time, facilitates the seamless adaptability to changes in the scene); identifying one or more objects in the virtual experience having one or more attributes that correspond to the text criteria, the one or more objects being identified by a large language model (Fig. 4 and Pg. 21-22, Section 3.2; This object contains an array of furniture data, where each entry corresponds to an individual piece. The "object_name" field acts as a key to access the corresponding prefab from our object data); determining spatial placement information in the virtual experience for the one or more objects using the large language model to interpret the text criteria to determine locations for the one or more objects in the virtual experience (Fig. 4 and Pg. 21, Section 3.2; Once retrieved, the prefab is instantiated within the Unity scene at the position specified by the X, Y, and Z values. Section 4.3; "A coffee table is in front of a sofa. A TV stand is opposite the coffee table"…"A dressing table is in front of a chair. A bed is next to the dressing table"…), wherein the large language model determines the locations based on the user prompt and the initialization prompt such that overlap between the one or more objects is avoided (Fig. 3 and Pg. 21-22, Section 3.2; Leveraging the information derived from the spatial layout and orientation of existing furniture within a scene, our system forms a comprehensive request to the GPT API, incorporating new elements…parses all objects present in the Unity3D scene, collecting their positions and orientations…); and placing the one or more objects in the virtual experience based on the spatial placement information (Pg. 22-25; Case Study 1-4 and Fig. 5-7). Kouzelis does not expressly disclose a non-transitory computer-readable medium comprising instructions that, responsive to execution by a processing device, causes the processing device to perform the recited operations. Official Notice is taken, however, that it is well known and conventional, at the time of the invention, that computer-implemented machine-learning and large language model methods are routinely embodied as computer-executable instructions stored on a non-transitory computer-readable storage medium and executed by one or more processors. Kouzelis further teaches that its disclosed method is implemented using conventional software components executing on standard computer platforms (Pg. 22, Section 4; incorporation of the GPT-4 API into the Unity3D platform (2021.3.8.f1 LTS). The VR avatar used in our case studies was sourced from the UltimateXR toolkit. All 3D models utilized in this study were modeled by the author in Blender 3.20 and textured using Adobe Substance 3D Painter. The complete list of objects available to the GPT API is the following…). Accordingly, such implementation would have been an implied or otherwise obvious computer implementation of the disclosed model operations according to standard programming and machine-learning techniques, yielding predictable result of causing a computer processor to perform the disclosed operations (KSR Int’l Co. v. Teleflex Inc., 550 U.S. 398,417 (2007)). Regarding claim 16, Kouzelis discloses initializing a large language model with an initialization prompt that includes instructions to the large language model to generate a virtual experience based on one or more subsequent prompts (Fig. 2 and Pg. 18, Section 3; Our methodology employs a streamlined process using an advanced language model for selecting and positioning 3D objects in Unity3D scenes. Utilizing the GPTAPI, this approach simplifies and improves virtual 3D object manipulation. Figure 2 outlines the method, which we detail here, focusing on the generation and interpretation of spatial and object data), wherein the instructions in the initialization prompt specify that overlap is impermissible when placing objects in the virtual experience (Pg. 19, Section 3.1; real-time parsing and prompt generating…extracting real-time dimensions for each furniture item from their mesh renderer component's bounds. This prevents object overlap by recording width, depth, and height. Fig. 3; You must take into account the dimensions of the furniture and the area, ensuring items do not overlap or float, and that they are oriented in a realistic manner…The furniture must be correctly positioned on the floor, not floating or overlapping…Be mindful of the furniture sizes so as they do not overlap…The door is at (12, 49, 0, 4) and should not be obstructed…)); receiving a user prompt, the user prompt comprising text criteria that specify criteria for generation or modification of the virtual experience, wherein the user prompt includes at least one of text data in natural language (Pg. 15-16, Section 1; we leverage the power of natural language processing (NLP) via a widely adopted large language model (LLM) which is the PGT API…translates simple natural language text into 3D scenes, as shown in Fig. 1), audio data, or video data (Fig. 3 and Pg. 20-21, Section 3.1; User input, shown in red, is brief but triggers extensive operations, as seen in the example "Create a living room". This blend of pre-determined, hard-coded elements, and adaptable components generated in real-time, facilitates the seamless adaptability to changes in the scene); identifying one or more objects in the virtual experience having one or more attributes that correspond to the text criteria, the one or more objects being identified by a large language model (Fig. 4 and Pg. 21-22, Section 3.2; This object contains an array of furniture data, where each entry corresponds to an individual piece. The "object_name" field acts as a key to access the corresponding prefab from our object data); determining spatial placement information in the virtual experience for the one or more objects using the large language model to interpret the text criteria to determine locations for the one or more objects in the virtual experience (Fig. 4 and Pg. 21, Section 3.2; Once retrieved, the prefab is instantiated within the Unity scene at the position specified by the X, Y, and Z values. Section 4.3; "A coffee table is in front of a sofa. A TV stand is opposite the coffee table"…"A dressing table is in front of a chair. A bed is next to the dressing table"…), wherein the large language model determines the locations based on the user prompt and the initialization prompt such that overlap between the one or more objects is avoided (Fig. 3 and Pg. 21-22, Section 3.2; Leveraging the information derived from the spatial layout and orientation of existing furniture within a scene, our system forms a comprehensive request to the GPT API, incorporating new elements…parses all objects present in the Unity3D scene, collecting their positions and orientations…); and placing the one or more objects in the virtual experience based on the spatial placement information (Pg. 22-25; Case Study 1-4 and Fig. 5-7). Kouzelis does not expressly disclose a system comprising a memory with instructions stored thereon, and a processing device coupled to the memory, the processing device configured to access the memory and execute the instructions, wherein the instructions cause the processing device to perform the recited operations. Official Notice is taken, however, that it is well known, at the time of the invention, that computer-implemented machine-learning are conventionally executed by processors accessing instructions stored in memory. Kouzelis further describes implementation of its disclosed techniques using conventional computing software (Pg. 22, Section 4; incorporation of the GPT-4 API into the Unity3D platform (2021.3.8.f1 LTS). The VR avatar used in our case studies was sourced from the UltimateXR toolkit. All 3D models utilized in this study were modeled by the author in Blender 3.20 and textured using Adobe Substance 3D Painter. The complete list of objects available to the GPT API is the following…), evidencing execution on ordinary computer hardware rather than any specialized architecture. Accordingly, such implementation would have been an implied or otherwise obvious computer implementation of the disclosed model operations according to standard programming and machine-learning techniques, yielding predictable result of causing a computer processor to perform the disclosed operations (KSR Int’l Co. v. Teleflex Inc., 550 U.S. 398,417 (2007)). Regarding claim 10, claim 10 has similar limitations as of claim 2, except it is a CRM claim (see Official Notice above), therefore it is rejected under the same rationale as claim 2. Regarding claims 11 and 18, claims 11 and 18 has similar limitations as of claim 3, except claim 11 is the CRM claim and claim 18 is the system claim (see Official Notice above), therefore it is rejected under the same rationale as claim 3. Regarding claims 12, claims 12 has similar limitations as of claim 4, except claim 12 is the CRM claim (see Official Notice above), therefore it is rejected under the same rationale as claim 4. Regarding claims 14 and 20, claims 14 and 20 has similar limitations as of claim 6, except claim 14 is the CRM claim and claim 20 is the system claim (see Official Notice above), therefore it is rejected under the same rationale as claim 6. Regarding claim 15, claim 15 has similar limitations as of claim 7, except it is a CRM claim (see Official Notice above), therefore it is rejected under the same rationale as claim 7. Regarding claims 13 and 17, claims 13 and 17 has similar limitations as of claim 8, except claim 13 is the CRM claim and claim 17 is the system claim (see Official Notice above), therefore it is rejected under the same rationale as claim 8. Regarding claim 19, Kouzelis discloses the system of claim 16, and further discloses wherein the placing comprises placing objects such that overlap between the objects is avoided by detecting whether any of the placed objects would overlap, and correcting object placement if potential overlap would occur, in response to receiving a supplemental prompt comprising instructions to remove the potential overlap (Pg. 19, Section 3.1; real-time parsing and prompt generating…extracting real-time dimensions for each furniture item from their mesh renderer component's bounds. This prevents object overlap by recording width, depth, and height. Fig. 3; You must take into account the dimensions of the furniture and the area, ensuring items do not overlap or float, and that they are oriented in a realistic manner…The furniture must be correctly positioned on the floor, not floating or overlapping…Be mindful of the furniture sizes so as they do not overlap…The door is at (12, 49, 0, 4) and should not be obstructed…)) Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JENNY NGAN TRAN whose telephone number is (571) 272-6888. The examiner can normally be reached Mon-Thurs 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alicia Harrington can be reached at (571) 272-2330. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JENNY N TRAN/Examiner, Art Unit 2615 /ALICIA M HARRINGTON/Supervisory Patent Examiner, Art Unit 2615
Read full office action

Prosecution Timeline

Show 3 earlier events
Dec 03, 2025
Examiner Interview Summary
Dec 03, 2025
Applicant Interview (Telephonic)
Jan 20, 2026
Response Filed
Mar 09, 2026
Final Rejection mailed — §102, §103
May 07, 2026
Response after Non-Final Action
Jun 02, 2026
Request for Continued Examination
Jun 08, 2026
Response after Non-Final Action
Aug 10, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12718383
METHODS AND SYSTEMS FOR MOTION VECTOR CALCULATION AND PROCESSING
2y 9m to grant Granted Aug 25, 2026
Patent 12499589
SYSTEMS AND METHODS FOR IMAGE GENERATION VIA DIFFUSION
2y 6m to grant Granted Dec 16, 2025
Study what changed to get past this examiner. Based on 2 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
44%
Grant Probability
78%
With Interview (+33.3%)
2y 7m (~2m remaining)
Median Time to Grant
High
PTA Risk
Based on 9 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month