DETAILED ACTION
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1 – 20 are rejected under 35 U.S.C. 103 as being unpatentable over Bradford (Publication: US 2024/0321034 A1) in view of Shah et al. (Publication: US 2023/0334773 A1).
Regarding claim 1, see rejection on claim 11.
Regarding claim 2, see rejection on claim 12.
Regarding claim 3, see rejection on claim 13.
Regarding claim 4, see rejection on claim 14.
Regarding claim 5, see rejection on claim 15.
Regarding claim 6, see rejection on claim 16.
Regarding claim 7, see rejection on claim 17.
Regarding claim 8, see rejection on claim 18.
Regarding claim 9, see rejection on claim 19.
Regarding claim 10, see rejection on claim 20.
Regarding claim 11, Bradford discloses a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform operations comprising ([0185], [0009], [0114], Fig. 10- a system of a head-wearable apparatus with memory include software processed by the processor to perform the following methods: ):
accessing an image captured by a camera of a client device, wherein the image depicts an area of a physical world around the client device ([0168] - The image capturing device 910, ”client”, enables the user 992 to provide gesture 958 input via the UI module 954, where the body tracking module 952 processes or analyzes the images 918 to determine the gesture 958 “physical world around the client device” and the user intent 964 based on the analysis of the images 918, “accessing an image”.);
identifying a set of objects depicted in the image ([0168] - where the body tracking module 952 processes or analyzes the images 918 to determine “identifying” the gesture 958 and the user intent 964 based on the analysis of the images 918, “identifying a set of objects”.);
generating an [[LLM prompt]] based on the accessed image and the set of objects (
[0190] - The user 992 selects a user interface item 960 by moving their body to cause the hand 1216 of their avatar 951 to remain over a user interface item 960 for a threshold period of time. That is the user sees the image, captured by the image capturing device, and items on display then by moving their body to cause the hand to make a selection “generating”.),
wherein the [[LLM prompt]] comprises a set of candidate actions performable by an AR character and instructions for an [[LLM]] to select a subset of the candidate actions for the AR character to perform based on the accessed image and the set of objects (
[0176] - using an AR gesture 958 to select an item 960 from a menu of items 960 to perform the function of selecting a vending item 925 , “prompt”.
[0200] - In display F 1812 the user 992 selects icon 1702 to “select a drink”. The 3D presentation module 989 indicates that to select an item 960 of a menu the user 992 is to “select menu with a long tap.” Here the items 960 are sub-menus, “subset of the candidate actions”.
[0188] - The UI module 954 will determine the user interface item 960 was selected by the user 992 if the user 992 moves his hand 967 so that the hand of the avatar 951 is at the same location as the user interface item 960 and the hand of the avatar 951 is held at the same location for a threshold period of time, “select a subset of the candidate actions for the AR character to perform”.
[0168], [0190] - enables the user 992 to provide gesture 958 input via the UI module 954, where the body tracking module 952 processes or analyzes the images 918 to determine the gesture 958 and the user intent 964 based on the analysis of the images 918 thus “access image” can be read on.The user 992 selects a user interface item 960 by moving their body to cause the hand 1216 of their avatar 951 to remain over a user interface item 960 for a threshold period of time, “set of objects”, thus “AR character to perform based on the accessed image and the set of objects” can be read on.);
receiving a response, wherein the response comprises a series of selected actions, wherein the series of selected actions comprises a selected subset of the candidate actions ([0190] - The user 992 selects a user interface item 960 by moving their body to cause the hand 1216 of their avatar 951 to remain over a user interface item 960 for a threshold period of time.
[0200] - In display F 1812 the user 992 selects icon 1702 to “select a drink”. The 3D presentation module 989 indicates that to select an item 960 of a menu the user 992 is to “select menu with a long tap.” Here after receiving a response “select a drink”, the items 960 are sub-menus, “subset of the candidate actions”.);
generating augmented reality content based on the series of selected actions, wherein the augmented reality content comprises a series of animations of the AR character corresponding to the series of selected actions ([0074] - generate augmented content and augmented reality experiences, such as animations to real-world images.
Fig. 17 contains a series of animations of the AR character.
[0209] In display A 2250 the users 992 have selected “Soda Studio” 1704 of FIG. 17. The 3D presentation module 989 is displaying engaging augmented graphics 1104 and avatars 951 that indicate that the users 992 may create a music video, play a game, or play music. For example, a keyboard 2204 is displayed within a bubble. Icon 2202 with label “Start Game” is an item 960.
PNG
media_image1.png
394
788
media_image1.png
Greyscale
); and
displaying the augmented reality content on the client device ( [0188] – 1104 Augmented Graphics is displayed on the vending machine.
The augmented graphics 1104 is a content item 943 that the rendering module 957 rendered.
PNG
media_image2.png
542
598
media_image2.png
Greyscale
).
Bradford does not however Shah discloses
LLM Prompt ([0007] - a scripted prompt that describes the contextual information and the method further includes providing the scripted prompt as input to a large language model (LLM) and outputting, with the LLM, the simulation scenario based on the contextual information in the scripted prompt.)
LLM ([0007] - a scripted prompt that describes the contextual information and the method further includes providing the scripted prompt as input to a large language model (LLM) and outputting, with the LLM, the simulation scenario based on the contextual information in the scripted prompt.);
receiving a response to the LLM, perform output ([0007] - providing the LLM scripted prompt as input to a large language model (LLM), receives the input to the LLM, and response with outputting stylized digital media);
transmitting the LLM prompt to the LLM ([0007] - providing the LLM scripted prompt as input to a large language model (LLM) and outputting, “transmitting” .).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to modify Bradford with LLM Prompt; LLM, receiving a response to the LLM, perform output and transmitting the LLM prompt to the LLM as taught by Shah. The motivation for doing is to enable relevant to its intended purpose.
Regarding claim 12, Bradford in view of Shah disclose all the limitations of claim 11 including LLM prompt and LLM.
Bradford discloses identifying a target object of the set of objects ([0030] - determining the user performed a gesture to select the vending item, one of the vending items.); and
generating the [[LLM prompt]] with instructions for the [[LLM]] to select the subset of the candidate actions based on the identified target object (
[0190] - The 3D presentation module 989 uses the images 918, content items 943, items 960, and/or the avatar 951 to generate frames 977 that are displayed on the display 908.; user interface items 960 that include icon 1202, icon 1204, icon 1206, icon 1208, icon 1210, icon 1212, and icon 1214, “prompt”. The user 992 selects a user interface item 960 by moving their body to cause the hand 1216 of their avatar 951 to remain over a user interface item 960 for a threshold period of time. ).
Regarding claim 13, Bradford in view of Shah disclose all the limitations of claim 11 including LLM prompt and LLM.
Bradford discloses to generate computer-executable code for animating the AR character ([0190] - The 3D presentation module 989 uses the images 918, content items 943, items 960, and/or the avatar 951 to generate frames 977 that are displayed on the display 908.; user interface items 960 that include icon 1202, icon 1204, icon 1206, icon 1208, icon 1210, icon 1212, and icon 1214. The user 992 selects a user interface item 960 by moving their body to cause the hand 1216 of their avatar 951 “AR character” to remain over a user interface item 960 for a threshold period of time.).
Regarding claim 14, Bradford in view of Shah disclose all the limitations of claim 11 including LLM prompt and LLM.
Bradford discloses to generate computer-executable code in a markup language ([0040] – markup-language.).
Regarding claim 15, Bradford in view of Shah disclose all the limitations of claim 11 including LLM prompt and LLM.
Bradford discloses a tag for each of the set of candidate actions ([0094] - tag for the content items.).
Regarding claim 16, Bradford in view of Shah disclose all the limitations of claim 11 including LLM prompt and LLM.
Bradford discloses comprises a description corresponding to each of the set of candidate actions ([0190] - The 3D presentation module 989 uses the images 918, content items 943, items 960, and/or the avatar 951 to generate frames 977 that are displayed on the display 908.; user interface items 960 that include icon 1202, icon 1204, icon 1206, icon 1208, icon 1210, icon 1212, and icon 1214. The user 992 selects a user interface item 960 by moving their body to cause the hand 1216 of their avatar 951 to remain over a user interface item 960 for a threshold period of time “description corresponding to each of the set of candidate actions”.).
Shah discloses text description ([0028] - a scripted prompt (e.g., human understandable text describing a scenario).).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to modify Bradford in view of Shah with text description as taught by Shah. The motivation for doing is to enable relevant to its intended purpose.
Regarding claim 17, Bradford in view of Shah disclose all the limitations of claim 11 including LLM prompt and LLM.
Bradford discloses identifying an animation for the AR character corresponding to each of the series of selected actions ([0201] In display G 1814, the icon 1802 for drink “A” and icon 1804 for drink “B” are displayed within bubbles. Additionally, a return to previous menu icon 1806 is displayed. The 3D presentation module 989 displays animations (not illustrated) and additional product information to induce the user 992 to purchase the drink and to educate the user 992 about the drinks. In display H 1816, the user 992 has selected drink “A” 1804. The 3D presentation module 989 generates frames 977 to provide an AR experience for the user 992. In display I 1818, the AR experience has ended and the user 992 is provided with a code 1810 to access the video. The video may be encrypted for privacy and the code 1810 may be sent to a mobile device 982 of the user 992 and/or a social media account of the user 992, an email account, and so forth.).
Regarding claim 18, Bradford in view of Shah disclose all the limitations of claim 11 including LLM prompt and LLM.
Bradford discloses transmitting the AR content to the client device ([0002], [0072] - As shown in Fig. 1, Augment reality interactions between systems
PNG
media_image3.png
788
530
media_image3.png
Greyscale
).
Regarding claim 19, Bradford in view of Shah disclose all the limitations of claim 11 including LLM prompt and LLM.
Bradford discloses augmenting the accessed image to include the augmented reality content ([0053] A camera system 204 includes control software (e.g., in a camera application) that interacts with and controls hardware camera hardware (e.g., directly or via operating system controls) of the user system 102 to modify and augment real-time images captured and displayed via the interaction client 104.
[0168], [0190] - enables the user 992 to provide gesture 958 input via the UI module 954, where the body tracking module 952 processes or analyzes the images 918 to determine the gesture 958 and the user intent 964 based on the analysis of the images 918 thus “access image” can be read on.The user 992 selects a user interface item 960 by moving their body to cause the hand 1216 of their avatar 951 to remain over a user interface item 960 for a threshold period of time.) .
Regarding claim 20, Bradford in view of Shah disclose all the limitations of claim 11 including LLM prompt and LLM.
Bradford discloses augmenting a set of images captured after the accessed image to include the AR content ([0053] A camera system 204 includes control software (e.g., in a camera application) that interacts with and controls hardware camera hardware (e.g., directly or via operating system controls) of the user system 102 to modify and augment real-time images captured and displayed via the interaction client 104.
[0168], [0190] - enables the user 992 to provide gesture 958 input via the UI module 954, where the body tracking module 952 processes or analyzes the images 918 to determine the gesture 958 and the user intent 964 based on the analysis of the images 918 thus “access image” can be read on.The user 992 selects a user interface item 960 by moving their body to cause the hand 1216 of their avatar 951 to remain over a user interface item 960 for a threshold period of time.).
Response to Arguments
Examiner suggests to amend a specific element in the claim that when reading a claim in light of the invention, it directs to a unique technology. The examiner can be reached at 571-270-0724 for further discussion.
Claim Rejection Under 35 U.S.C. 103
Applicant asserts “The Office Action relies primarily on Bradford for the claimed limitations, citing Shah only for the presence of an LLM. Office Action at 8-9. Bradford is directed to an augmented reality activated dispensing machine. For the limitation requiring an LLM prompt with instructions to select a subset of candidate actions for an AR character to perform, the Office Action cites Bradford at [0176], [0188], [0190], and [0200]. Office Action at 5-6. These passages describe a user interface in which a human user selects menu items through gestures.
For example, Bradford at [0190] states that "[t]he user 992 selects a user interface item 960 by moving their body to cause the hand 1216 of their avatar 951 to remain over a user interface item 960 for a threshold period of time." Bradford at [0200] describes that "the user 992 selects icon 1702 to 'select a drink."' In other words, Bradford describes a user manually selecting items from a vending machine menu through gestures, not a system that uses an LLM to select actions for an AR character to perform.
Therefore, Bradford does not teach "generating an LLM prompt based on the accessed image and the set of objects, wherein the LLM prompt comprises a set of candidate actions performable by an AR character and instructions for an LLM to select a subset of the candidate actions for the AR character to perform based on the accessed image and the set of objects." In Bradford, the user performs the selection, not an LLM. Bradford at [0190]. Bradford's items are vending machine menu options, not candidate actions for an AR character to perform. Bradford also does not teach "generating augmented reality content based on the series of selected actions, wherein the augmented reality content comprises a series of animations of the AR character corresponding to the series of selected actions" as recited in the claims. Bradford's selected items are vending machine products, not actions to be performed by an AR character that result in a series of character animations. See, e.g., Bradford at [0201] ("The 3D presentation module 989 displays animations (not illustrated) and additional product information to induce the user 992 to purchase the drink").
Examiner disagrees. In response to Applicant's arguments against the references individually, one cannot show nonobviousness by referencing references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986).
It is the combination of Bradford in view of Shah that disclose the language above.
It is the Shah that discloses the LLM to perform the operations and Bradford disclose the selecting items from a vending machine menu.
Bradford discloses [0176] - using an AR gesture 958 to select an item 960 from a menu of items 960 to perform the function of selecting a vending item 925 , “prompt”.
[0200] - In display F 1812 the user 992 selects icon 1702 to “select a drink”. The 3D presentation module 989 indicates that to select an item 960 of a menu the user 992 is to “select menu with a long tap.” Here the items 960 are sub-menus, “subset of the candidate actions”.
[0188] - The UI module 954 will determine the user interface item 960 was selected by the user 992 if the user 992 moves his hand 967 so that the hand of the avatar 951 is at the same location as the user interface item 960 and the hand of the avatar 951 is held at the same location for a threshold period of time, “select a subset of the candidate actions for the AR character to perform”.
[0168], [0190] - enables the user 992 to provide gesture 958 input via the UI module 954, where the body tracking module 952 processes or analyzes the images 918 to determine the gesture 958 and the user intent 964 based on the analysis of the images 918 thus “access image” can be read on.The user 992 selects a user interface item 960 by moving their body to cause the hand 1216 of their avatar 951 to remain over a user interface item 960 for a threshold period of time, “set of objects”, thus “AR character to perform based on the accessed image and the set of objects” can be read on.
[0190] - The 3D presentation module 989 uses the images 918, content items 943, items 960, and/or the avatar 951 to generate frames 977 that are displayed on the display 908.; user interface items 960 that include icon 1202, icon 1204, icon 1206, icon 1208, icon 1210, icon 1212, and icon 1214. The user 992 selects a user interface item 960 by moving their body to cause the hand 1216 of their avatar 951 to remain over a user interface item 960 for a threshold period of time “description corresponding to each of the set of candidate actions”.
Shah discloses [0007] - a scripted prompt that describes the contextual information and the method further includes providing the scripted prompt as input to a large language model (LLM) and outputting, with the LLM, the simulation scenario based on the contextual information in the scripted prompt, “selection”.
[0028] - a scripted prompt (e.g., human understandable text describing a scenario).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to modify Bradford with LLM Prompt; LLM, receiving a response to the LLM, perform output and transmitting the LLM prompt to the LLM as taught by Shah. The motivation for doing is to enable relevant to its intended purpose.
Regarding claims 2 – 10 and 12 – 20 the Applicant asserts that they are not obvious over based on their dependency from independent claims 1 and 11 respectively. The examiner cannot concur with the Applicant respectfully from same reason noted in the examiner’s response to argument asserted from claims 1 and 11 respectively.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Ming Wu whose telephone number is (571) 270-0724. The examiner can normally be reached on Monday-Thursday and alternate Fridays (9:30am - 6:00pm) EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Devona Faulk can be reached on 571-272-7515. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Ming Wu/
Primary Examiner, Art Unit 2618