DETAILED ACTION
Status of Claims
Applicant's amendments filed on 29 July 2026 have been entered. Claims 1-7, 9-17, and 19 have been amended. Claim 8 has been canceled. Claim 21 has been added. Claims 1-7 and 9-21 are still pending in this application, with claims 1, 10, and 19 being independent.
Response to Arguments
Applicant's arguments filed 29 have been fully considered but they are not persuasive.
With respect to independent claims 1 and 10, Applicant argues for their allowance. It is noted that these claims are in fact allowed.
No mention of any rejection of independent claim 19 is presented. This claim stands rejected for the reasons below.
Reasons for Allowance
Claims 1-7 and 9-18 are allowed.
The following is an examiner’s statement of reasons for allowance:
Claim 1 is allowable over the prior art of record since the cited references taken individually or in combination fails to particularly disclose or suggest a method comprising: generating third data that associates the one or more shapes with the one or more first captions and the one or more second captions, as presented in the environment of the remaining limitations of claim 1. It is noted that the closest prior art, Menashof et al. (US Pub. 2025/0148753), hereinafter Menashof, shows generating, based at least on first data representative of one or more poses, one or more images to represent one or more shapes associated with the one or more poses from one or more perspectives of one or more canonical viewpoints; generating, based at least on one or more language models processing second data associated with the one or more images: one or more first captions describing the one or more shapes; and one or more second captions describing the one or more shapes. However, Menashof fails to disclose or suggest the one or more first captions associated with one or more first numbers of words that are less than a threshold number of words, the one or more second captions being associated with one or more second numbers of words that are equal to or greater than the threshold number of words; and generating third data that associates the one or more shapes with the one or more first captions and the one or more second captions.
Claim 10 is allowable over the prior art of record since the cited references taken individually or in combination fails to particularly disclose or suggest a system comprising: generate third data that associates the shape with the first caption and the second caption, as presented in the environment of the remaining limitations of claim 10. It is noted that the closest prior art, Menashof, shows one or more processors to: generate first data representing a format for generating captions associated with at least a shape; generate second data associated with one or more images of the shape. However, Menashof fails to disclose or suggest the format indicating to generate at least a first caption that includes a first length and a second caption that includes a second length that is greater than the first length; generate, using one or more language models and based at least on the first data and the second data: the first caption that describes the shape using the first length; and the second caption that describes the shape using the second length; and generate third data that associates the shape with the first caption and the second caption.
The remaining claims 2-7, 9, and 11-18 depend from one of the above independent claims, either directly or indirectly, and are accordingly allowable.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 19 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Menashof et al. (US Pub. 2025/0148753), hereinafter Menashof, in view of Greasley (US Pub. 2021/0397842).
Regarding claim 19, Menashof discloses one or more processors (Fig. 7) comprising processing circuitry to: receive first data representative of one or more poses associated with one or more shapes (Paragraph [0024]: the encoder extracts every tenth frame of the video and causes a first generative machine model to generate a descriptor for the extracted frames (e.g., a natural language description of the soccer ball and the location). In this example, the extracted frames are selected based on an interval of time (e.g., every tenth frame), although, as described in greater detail below, other algorithms can be used for selecting frames of the video to be extracted; Paragraph [0071]: the system implementing the method 300 generates descriptors based at least in part on the pivot images and/or video. In one example, an LLM or other machine learning model generates descriptors based at least in part on the pivot images. The descriptors, in an embodiment, include natural language descriptions of the pivot images. In other embodiments, the descriptors include structured data (e.g., a JavaScript Object Notation [JSON] file) that indicates attributes (e.g., location, movement, size, shape, etc.) of objects within the pivot images. For example, the descriptors include text-based data (e.g., data, pseudo-code, source code, etc.) that is provided to a generative model of an encoder to enable the generative model to reconstruct the video. In an embodiment, descriptors are generated for the video and/or portions thereof); generate, based at least on the one or more poses, one or more images to represent the one or more shapes (Paragraph [0005]: the encoder includes a large language model (LLM) or other natural language model to generate descriptors of the pivot images. For example, the descriptors can include natural language descriptions of the objects, backgrounds, interactions, and concepts included in the pivot images; Paragraph [0024]: the encoder extracts every tenth frame of the video and causes a first generative machine model to generate a descriptor for the extracted frames (e.g., a natural language description of the soccer ball and the location). In this example, the extracted frames are selected based on an interval of time (e.g., every tenth frame), although, as described in greater detail below, other algorithms can be used for selecting frames of the video to be extracted; Paragraph [0071]: the system implementing the method 300 generates descriptors based at least in part on the pivot images and/or video. In one example, an LLM or other machine learning model generates descriptors based at least in part on the pivot images. The descriptors, in an embodiment, include natural language descriptions of the pivot images. In other embodiments, the descriptors include structured data (e.g., a JavaScript Object Notation [JSON] file) that indicates attributes (e.g., location, movement, size, shape, etc.) of objects within the pivot images. For example, the descriptors include text-based data (e.g., data, pseudo-code, source code, etc.) that is provided to a generative model of an encoder to enable the generative model to reconstruct the video. In an embodiment, descriptors are generated for the video and/or portions thereof); generate, based at least on one or more language models processing second data associated with one or more images, one or more captions describing the one or more shapes (Paragraph [0005]: the encoder includes a large language model (LLM) or other natural language model to generate descriptors of the pivot images. For example, the descriptors can include natural language descriptions of the objects, backgrounds, interactions, and concepts included in the pivot images; Paragraph [0024]: the encoder extracts every tenth frame of the video and causes a first generative machine model to generate a descriptor for the extracted frames (e.g., a natural language description of the soccer ball and the location). In this example, the extracted frames are selected based on an interval of time (e.g., every tenth frame), although, as described in greater detail below, other algorithms can be used for selecting frames of the video to be extracted; Paragraph [0071]: the system implementing the method 300 generates descriptors based at least in part on the pivot images and/or video. In one example, an LLM or other machine learning model generates descriptors based at least in part on the pivot images. The descriptors, in an embodiment, include natural language descriptions of the pivot images. In other embodiments, the descriptors include structured data (e.g., a JavaScript Object Notation [JSON] file) that indicates attributes (e.g., location, movement, size, shape, etc.) of objects within the pivot images. For example, the descriptors include text-based data (e.g., data, pseudo-code, source code, etc.) that is provided to a generative model of an encoder to enable the generative model to reconstruct the video. In an embodiment, descriptors are generated for the video and/or portions thereof).
Menashof does not explicitly disclose the one or more shapes from one or more canonical viewpoints; and generate third data that associates the one or more shapes with the one or more captions.
However, Greasley teaches generating a text description for images (Paragraphs [0038]-[0039]), further comprising: the one or more shapes from one or more canonical viewpoints (Paragraph [0014]: the XR system may detect head movement and, in response, adjust graphical content and an acoustic field presented to the person in a manner similar to how such views and sounds would change in a physical environment. As another example, the XR system may detect movement of the electronic device presenting the XR environment (e.g., a mobile phone, a tablet, a laptop, or the like) and, in response, adjust graphical content and an acoustic field presented to the person in a manner similar to how such views and sounds would change in a physical environment. In some situations (e.g., for accessibility reasons), the XR system may adjust characteristic(s) of graphical content in the XR environment in response to representations of physical motions; Paragraph [0026]: the environment is an XR environment, the objects 112 are XR representations of objective-effectuators. In some implementations, the objective-effectuators model characters from fictional materials such as movies, video games, comics, and novels. For example, the character object 112a may represent and model the behavior of a character from a fictional comic, and the character object 112b represents and models the behavior of a character from a fictional video game. In some implementations, the XR environment includes XR representations of objective-effectuators that represent and model the behavior of characters from different fictional materials (e.g., from different movies/games/comics/novels). In various implementations, the objective-effectuators represent and model the behavior of physical elements (e.g., tangible objects). For example, in some implementations, the objective-effectuators model the behavior of equipment (e.g., machinery such as planes, tanks, robots, cars, etc.). In the example of FIG. 1, the robot object 112c represents and models the behavior of a robot, and the drone object 112d represents and models the behavior of a drone); and generate third data that associates the one or more shapes with the one or more captions (Paragraph [0038]: if the user 106 indicates a preference for a short listening experience, the electronic device 102 and/or the controller 104 may generate a terse text description, e.g., with a low word count, for example, less than a threshold number of words. On the other hand, if the user 106 indicates a preference for a longer listening experience, the electronic device 102 and/or the controller 104 may generate a more verbose text description, e.g., with a high word count, for example, greater than a threshold number of words; Paragraph [0039]: if the user 106 is moving at a low velocity or is stationary, the electronic device 102 and/or the controller 104 may generate a text description with a high word count (e.g., greater than a threshold number of words). If the user 106 is moving at a high velocity, the electronic device 102 and/or the controller 104 may generate a text description with a low word count (e.g., less than a threshold number of words) or may increase the speed (e.g., words per minute) of the narration to keep pace with the movement of the user 106; Paragraph [0053]: length of the text description may be determined based on the velocity of the user. For example, if the user is stationary or moving at a low velocity, the audio output generator 224 may generate a verbose text description with a high word count (e.g., greater than a threshold number of words). If the user is moving at a high velocity, the audio output generator 224 may generate a terse text description with a low word count (e.g., less than a threshold number of words). In some implementations, the audio output generator 224 selects a depth to which the ontology 226 is traversed). Greasley teaches that this will allow for user preference of descriptions to be realized (Paragraph [0038]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Menashof with the features of above as taught by Greasley so as to allow for user preference of descriptions to be realized as presented by Greasley.
Regarding claim 20, Menashof, in view of Greasley teaches the one or more processors of claim 19, Menashof discloses wherein the one or more processors are comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using one or more large language models (LLMs) (Paragraph [0005]: the encoder includes a large language model (LLM) or other natural language model to generate descriptors of the pivot images); a system for performing operations using one or more visual language models (VLMs); a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
Allowable Subject Matter
Claim 21 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim 21 would be allowable over the prior art of record since the cited references taken individually or in combination fails to particularly disclose or suggest one or more processors comprising processing circuitry wherein: the one or more captions include at least the one or more first captions that includes the one or more first lengths and the one or more second captions that includes the one or more second lengths, as presented in the environment of the remaining limitations of claim 19. It is noted that the closest prior art, Menashof, shows obtain fourth data representative of a format for generating captions associated with the one or more shapes. However, Menashof fails to disclose or suggest the format indicating to generate one or more first captions that include one or more first lengths and one or more second captions that include one or more second lengths that is greater than the one or more first lengths, wherein: the one or more captions are further generated based at least on the fourth data; and the one or more captions include at least the one or more first captions that includes the one or more first lengths and the one or more second captions that includes the one or more second lengths.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW D SALVUCCI whose telephone number is (571)270-5748. The examiner can normally be reached M-F: 7:30-4:00PT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, XIAO WU can be reached at (571) 272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MATTHEW SALVUCCI/Primary Examiner, Art Unit 2613