DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
Applicant is reminded of the proper language and format for an abstract of the disclosure.
The abstract should be in narrative form and generally limited to a single paragraph on a separate sheet within the range of 50 to 150 words in length. The abstract should describe the disclosure sufficiently to assist readers in deciding whether there is a need for consulting the full patent text for details.
The language should be clear and concise and should not repeat information given in the title. It should avoid using phrases which can be implied, such as, “The disclosure concerns,” “The disclosure defined by this invention,” “The disclosure describes,” etc. In addition, the form and legal phraseology often used in patent claims, such as “means” and “said,” should be avoided.
The abstract of the disclosure is objected to because the abstract should be within the range of 50 to 150 words in length. A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b).
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-3, 7-15, 18-20, 25-26, 28-31 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Wang et al (US Pub 2023/0147641 A1).
Regarding Claim 1, Wang et al teaches a method performed by one or more computers and for generating a controllable video (Paragraph 0052 for “generate user-controllable video”),
the video comprising a respective video frame corresponding to each of a sequence of time
points (Paragraph 0047, 0052), and the method comprising, for each of the sequence of time points:
processing at least the video frame corresponding to the time point using a video encoder (Paragraph 0294, 0296) neural network (Paragraph 0045) to generate a first set of tokens representing the video as of the time point (Paragraph 0045-0047);
obtaining data selecting an action from a set of actions (Paragraph 0047; (Page 48, see Claims 3, 9, 15); (Page 49, see Claims 21, 27);
processing a dynamics input comprising the first set of tokens and the selected action using a dynamics neural network (Paragraph 0045-0047) to generate a second set of tokens representing a video frame at a next time point in the sequence in the sequence given that the selected action is performed at the time point (Paragraph 0045-0047, 0049); and
processing at least the second set of tokens using a video decoder (Paragraph 0294, 0296) neural network to generate the video frame at the next time point (Paragraph 0045-0047).
Regarding Claims 2, 3, Wang et al teaches the method wherein each token in the first set of tokens and in the second set of tokens is a respective token from a discrete set of tokens; wherein the video encoder neural network is a causally masked neural network that generates the first set of tokens conditioned on the video frame corresponding to the time point and any video frames at any preceding time points in the sequence (Paragraph 0045-0048).
Regarding Claim 7, Wang et al teaches the method wherein the set of actions is a discrete set of learned, latent actions (Paragraph 0047).
Regarding Claims 8, 9, Wang et al teaches the method wherein obtaining data selecting an action from a set of actions comprises: receiving a user input selecting an action from the set of actions; wherein obtaining data selecting an action from a set of actions comprises: sampling an action from a distribution over the set of actions (Paragraph 0047-0052).
Regarding Claims 10-12, Wang et al teaches the method wherein the video further comprises one or more initial video frames preceding the video frames corresponding to the sequence of time points; obtaining the one or more initial video frames; wherein obtaining the one or more initial video frames comprises: receiving a user input identifying the one or more initial video frames (Paragraph 0047-0052).
Regarding Claim 13, Wang et al teaches the method wherein obtaining the one or more initial video frames comprises: receiving, as user input, one or more context inputs; and processing the one or more context inputs using an image generation neural network to generate the one or more initial video frames (Paragraph 0047-0052).
Regarding Claim 14, Wang et al teaches the method in which the one or more initial video frames are images of a real-world environment captured by a camera device (Paragraph 0329, 0356).
Regarding Claim 15, Wang et al teaches the method wherein the dynamics input further comprises, for each of one or more video frames at one or more preceding time steps, (i) a respective first set of tokens representing the video frame and (ii) a selected action at the preceding time step (Paragraph 0049-0052).
Regarding Claims 18-20, Wang et al teaches the method wherein the dynamics neural network is a masked generative image Transformer; wherein the video encoder and video decoder neural networks have been jointly trained on a video reconstruction objective on an unsupervised video data set; wherein the video reconstruction objective is a VQ-VAE objective. (Paragraph 0045-0048).
Regarding Claims 25-26, Wang et al teaches the method wherein, after the video encoder and video decoder neural networks have been trained, the dynamics neural network is trained on the unsupervised video data set on a video token prediction task; wherein the unsupervised video data set does not include any text or action labels (Paragraph 0049-0052).
Regarding Claim 28, Wang et al teaches a method performed by one or more computers, the method comprising, at each of a plurality of time steps: obtaining an image of an environment being interacted with by an agent, wherein the agent is controllable by a set of control inputs (Paragraph 0045-0047);
processing an input comprising the image at the time step using a policy neural network to generate a policy output that assigns a respective probability to each of a set of latent actions; selecting a latent action using the policy output (Paragraph 0045-0052);
mapping the selected latent action to a particular control input from the set of control inputs; and
controlling the agent by submitting the particular control input (Paragraph 0045-0052).
Regarding Claim 29, Wang et al teaches the method wherein the learned, latent actions have been learned by training a latent action model on an unsupervised video data set. (Paragraph 0047).
Regarding Claim 30, Wang et al teaches the method wherein the policy neural network has been trained through imitation learning on trajectories generated by processing sequences of images using the trained latent action model to identify a latent action performed by the agent in each image in the sequence (Paragraph 0045-0047).
Regarding Claim 31, the method Claim 31 is rejected for same reason as the apparatus Claim 1, since claim limitations are same in both claims.
Allowable Subject Matter
Claims 4-6, 16-17, 21-24, 27 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Examiner cites particular columns and line numbers in the references as applied to the claims below for the convenience of the applicant. Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested that, in preparing responses, the applicant fully consider the references in entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the examiner.
It is noted that any citation to specific pages, columns, figures, or lines in the prior art references any interpretation of the references should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. In re Heck, 699 F.2d 1331-33, 216 USPQ 1038-39 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 USPQ 275, 277 (CCPA 1968)).
Examiner’s Note
Examiner has cited particular paragraphs/columns and line numbers or figures in the references as applied to the claims below for the convenience of the applicant. Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested from the applicant, in preparing the responses, to fully consider the references in their entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the examiner.
Applicant is reminded that the Examiner is entitled to give the broadest reasonable interpretation to the language of the claims. Furthermore, the Examiner is not limited to Applicant’s definition which is not specifically set forth in the claims.
In the case of amending the claimed invention, Applicant is respectfully requested to indicate the portion(s) of the specification which dictate(s) the structure relied on for proper interpretation and also to verify and ascertain the metes and bounds of the claimed invention.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to VIJAY SHANKAR whose telephone number is (571)272-7682. The examiner can normally be reached M-F 9 am- 6 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matthew Eason can be reached at 571-270-7230. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
VIJAY SHANKAR
Primary Examiner
Art Unit 2624
/VIJAY SHANKAR/Primary Examiner, Art Unit 2624