DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
1. The information disclosure statement (IDS) submitted on 4 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement has been considered by the examiner.
Response to Amendment
2. This action is in response to the amendment filed on June 26th, 2026. Claims 1, 4, 7, 11, and 15-18 have been amended. Claims 5-6, 9-10, and 13-14 have been cancelled. Claims 21-25 have been added. Claims 1-4, 7-8, 11-12 and 15-25 are pending. Claims 1-4, 7-8, 11-12 and 15-25 remain rejected in the application. The Drawings are still objected to as outlined in the non-final office action mailed on March 26th, 2026.
Response to Arguments
3. Applicant’s arguments with respect to claim 1, and similarly claims 4 and 18, filed on 6/26/2026, have been fully considered. Applicant’s arguments are directed against the prior art of Gao et al. (WO-2026/001197-A1) with respect to the rejection under 35 U.S.C. 102.
4. Applicant’s arguments regarding claim 1 that the prior art of Gao et al. (WO-2026/001197-A1) does not teach the limitation(s): "wherein the camera motion input is one of a set of camera motion labels that each represent a different camera motion" and “tokenizing the camera motion label into a set of one or more tokens and conditioning the video generation neural network on the set of one or more tokens when generating the output video” have been fully considered. These limitations are from claims 9 and 10 and are incorporated into amended claim 1 respectively. Applicant argues that Gao does not teach these limitations, however, Gao is not relied upon to disclose these limitations in the non-final rejection and the arguments are moot. Specifically, claim 9 was rejected by Han et al. (US-2021/0144418-A1) and claim 10 was rejected by Osotsi et al. (US-2024/0202604-A1). Thus, claim 1 is now rejected by the combination of Gao, Han, and Osotsi.
5. Applicant’s arguments regarding claim 4 that the prior art of Gao et al. (WO-2026/001197-A1) does not teach the limitation(s): "wherein the camera motion input comprises a respective score for each of a set of a plurality of camera motion labels that each represent a different camera motion" and “tokenizing the respective scores for the camera motion labels into a set of one or more tokens and conditioning the video generation neural network on the set of one or more tokens” have been fully considered. These limitations are from claims 13 and 14 and are incorporated into amended claim 4 respectively. Applicant argues that Gao does not teach these limitations, however, Gao is not relied upon to disclose these limitations in the non-final rejection and the arguments are moot. Specifically, claim 13 was rejected by Han et al. (US-2021/0144418-A1) and claim 14 was rejected by Osotsi et al. (US-2024/0202604-A1). Thus, claim 4 is now rejected by the combination of Gao, Han, and Osotsi.
6. Applicant’s arguments regarding claim 18 that the prior art of Gao et al. (WO-2026/001197-A1) does not teach the limitation(s): "wherein the camera motion output comprises a respective score for each of a set of a plurality of camera motion labels that each represent a different camera motion” has been fully considered. This limitation is from claims 13 and is incorporated into amended claim 18. Applicant argues that Gao does not teach this limitation, however, Gao is not relied upon to disclose this limitation in the non-final rejection and the argument is moot. Specifically, claim 13 was rejected by Han et al. (US-2021/0144418-A1). Thus, claim 18 is now rejected by the combination of Gao and Han.
7. Regarding arguments to claims 2-3, 7-8, 11-12, 15-17, and 19-25, they are dependent on independent claims 1, 4, and 18 respectively. Applicant does not argue anything other than independent claim 1, and similarly claims 4 and 18. The limitations in those claims, in conjunction with their combination, has previously been established and explained.
Drawings
8. The drawings are objected to as failing to comply with 37 CFR 1.84(p)(4) because reference character “106” has been used to designate both "Camera Motion Input" and "Subject Motion Input" in Fig. 1. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
9. The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they do not include the following reference sign(s) mentioned in the description: 104, 600, 602, 612, 614, and 622. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Objections
10. Claims 15 and 2 objected to because of the following informalities:
In claim 15, line 1 and claim 21, line 1, "The method claim" should read "The method of claim".
Appropriate correction is required.
Claim Rejections - 35 USC § 103
11. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
12. Claims 18 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Gao et al (WO-2026/001197-A1, hereinafter "Gao") in view of Han et al. (US-2021/0144418-A1, hereinafter "Han"). (Examiner’s note: Citations to Gao use the original WO-2026/001197-A1 document locations.)
13. As per claim 18, Gao discloses: A method performed by one or more computers and for training a video generation neural network that generates output videos conditioned on respective conditioning inputs, the method comprising: (Gao, page 5, lines 32-37, “This disclosure provides a video generation method, and also relates to a method for generating motion videos of virtual objects, a video editing method, a video generation model training method, ... a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.”)
obtaining a plurality of initial training examples, each training example comprising (i) an initial conditioning input (Gao, page 10, lines 33-36, “In one optional embodiment of this disclosure, the training method of the video generation model is described. That is, before inputting the video generation data into the video generation model to obtain the target video of the target object, the following steps may be included: Obtain sample videos and corresponding sample object images, sample action sequences, and sample camera parameters …”) and (ii) a target video that spans a corresponding time window and is characterized by the initial conditioning input; (Gao, page 10, lines 37-39, “Input the sample object image, sample action sequence and sample camera parameters into the pre-trained generative model to obtain the first predicted feature output by the generative unit in the pre-trained generative model; Input the sample video into the image encoding unit in the pre-trained generative model to obtain the first sample features …” and page 5, lines 27-31, “Through the parameter cross-attention unit, the model can associate virtual camera parameters and video frames in the feature dimension, learn the control and generation of camera parameters, and thus support large-scale camera movement effects. Through the temporal attention unit, the video generation model can learn temporal information, enabling smooth transitions between time sequences in the generated video, making it more realistic and natural.” and page 6, lines 26-44, “Virtual camera parameters can also be called virtual camera trajectory parameters. Virtual camera parameters are a series of data used to describe the motion state of a virtual camera when it is moving or shooting a dynamic scene. The motion state of the virtual camera itself can be controlled by the virtual camera parameters, thereby controlling the camera movement trajectory of the generated video. Specifically, the virtual camera parameters include the camera extrinsic parameters at each time t throughout the entire time period T, such as the parameters of the virtual camera's position and orientation in the world coordinate system, namely the rotation matrix and translation vector.”)
for each initial training example:
processing a first input comprising (i) the target video in the initial training example or (ii) features of the target video in the initial training example using a camera motion classifier neural network to generate a camera motion output that characterizes a motion of a camera that captured the target video during the corresponding time window, [[wherein the camera motion output comprises a respective score for each of a set of a plurality of camera motion labels that each represent a different camera motion;]] and (Gao, page 16, lines 1-5, “In another optional embodiment of this disclosure, in addition to selecting a video generation model from the model library, a pre-trained generation model in the model library can be trained based on the sample video and the corresponding sample object image, sample action sequence, and sample camera parameters in the request information to obtain a video generation model. That is, the request information includes the sample video of the target video task and the corresponding sample object image, sample action sequence, and sample camera parameters.” and page 10, lines 37-39, “Input the sample object image, sample action sequence and sample camera parameters into the pre-trained generative model to obtain the first predicted feature output by the generative unit in the pre-trained generative model; Input the sample video into the image encoding unit in the pre-trained generative model to obtain the first sample features …”)
generating a final training example that comprises (i) a conditioning input that comprises the initial conditioning input in the initial training example and a camera motion input generated from the camera motion output and (ii) the target video in the initial training example; and (Gao, page 16, lines 1-7, “In another optional embodiment of this disclosure, in addition to selecting a video generation model from the model library, a pre-trained generation model in the model library can be trained based on the sample video and the corresponding sample object image, sample action sequence, and sample camera parameters in the request information to obtain a video generation model. That is, the request information includes the sample video of the target video task and the corresponding sample object image, sample action sequence, and sample camera parameters. The above-mentioned acquisition of the video generation model based on the request information may include the following steps: Based on the sample video and the corresponding sample object image, sample action sequence and sample camera parameters, the pre-trained generation model corresponding to the target video task is trained to obtain the trained video generation model.” and page 10, lines 37-39, “Input the sample object image, sample action sequence and sample camera parameters into the pre-trained generative model to obtain the first predicted feature output by the generative unit in the pre-trained generative model; Input the sample video into the image encoding unit in the pre-trained generative model to obtain the first sample features …” and page 6, lines 26-44, “Virtual camera parameters can also be called virtual camera trajectory parameters. Virtual camera parameters are a series of data used to describe the motion state of a virtual camera when it is moving or shooting a dynamic scene. The motion state of the virtual camera itself can be controlled by the virtual camera parameters, thereby controlling the camera movement trajectory of the generated video. Specifically, the virtual camera parameters include the camera extrinsic parameters at each time t throughout the entire time period T, such as the parameters of the virtual camera's position and orientation in the world coordinate system, namely the rotation matrix and translation vector.” and page 11, line 45-page 12, line 4, “This training phase can be called the video generation training phase. At this stage, the goal of model training is to generate temporally stable video clips while being able to control camera movement. ... By adding parameters across attention units, the model learns information about the camera trajectory and generates videos that match the input camera parameters.” and page 7, lines 10-14, “Specifically, the action extraction model is used to extract action sequences from the input video. Action extraction models can be deep learning models trained on the basis of CNN models. The parameter extraction model is used to extract camera parameters from the input video.”; Examiner’s note: Based on sample video from the video generation model library, a “target video task” is created and corresponds to a “final training example.” Also, as disclosed by Gao in page 11, line 45-page 12, line 4, the training process generates video clips that match the input camera parameters and thus would be considered a camera motion input.)
training the video generation neural network on the final training examples. (Gao, page 16, lines 6-11, “Based on the sample video and the corresponding sample object image, sample action sequence and sample camera parameters, the pre-trained generation model corresponding to the target video task is trained to obtain the trained video generation model. It should be noted that the implementation method of "training the pre-trained generation model corresponding to the target video task based on the sample video and the sample object image, sample action sequence and sample camera parameters to obtain the trained video generation model" is the same as the training method of the above-mentioned "video generation model" ...”; Examiner’s note: Geo discloses training a pre-trained model to create a final trained task specific to a target video.)
14. Gao doesn't explicitly disclose but Han discloses: [[processing a first input comprising (i) the target video in the initial training example or (ii) features of the target video in the initial training example using a camera motion classifier neural network to generate a camera motion output that characterizes a motion of a camera that captured the target video during the corresponding time window,]] wherein the camera motion output comprises a respective score for each of a set of a plurality of camera motion labels that each represent a different camera motion; [[and]] (Han, [0054], “In an implementation, a content side model may be adopted for determining the content score of the candidate video as discussed above. ... The content side model 230 may be established based on various techniques, e.g., machine learning, deep learning, etc. Features adopted by the content side model 230 may comprise at least one of: shot transition, camera motion, scene, human, human motion, object, object motion, text information, audio attribute, and video metadata, as discussed above. In terms of function, the content side model 230 may be, e.g., a regression model, a classification model, etc. ... Training data for the content side model 230 may be obtained through: obtaining a group of videos to be used for training; for each video in the group of videos, labeling respective values corresponding to the features of the content side model, and labeling a content score for the video; and forming training data from the group of videos with respective labels.” and [0053], “It should be appreciated that any two or more of the above discussed shot transition, camera motion, scene, human, human motion, object, object motion, text information, audio attribute, and video metadata may be combined together so as to determine the content score of the candidate video. For example, for a video recording a cute dog's activities, this video may contain a large amount of camera motions and object motions but does not include any speech or music, and thus a content score indicating a high importance of visual information may be determined for this video.”)
15. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of Gao to include the disclosure of the camera motion output comprises a respective score for each of a set of a plurality of camera motion labels that each represent a different camera motion, of Han. The motivation for this modification could have been to assist with evaluating the motion of a camera. In particular, this allows for comparison of various motions to determine how similar they are and also for classification purposes.
16. As per claim 19, Gao in view of Han discloses: The method of claim 18, wherein:
each target video depicts a respective subject, the method further comprises, for each initial training example: (Gao, page 10, lines 33-36, “In one optional embodiment of this disclosure, the training method of the video generation model is described. That is, before inputting the video generation data into the video generation model to obtain the target video of the target object, the following steps may be included: Obtain sample videos and corresponding sample object images, sample action sequences, and sample camera parameters …”)
processing a second input comprising (i) the target video in the initial training example or (ii) the features of the target video in the initial training example using a subject motion classifier neural network to generate a subject motion output that characterizes a motion of the respective subject depicted in the target video during the corresponding time window; (Gao, page 10, lines 37-39, “Input the sample object image, sample action sequence and sample camera parameters into the pre-trained generative model to obtain the first predicted feature output by the generative unit in the pre-trained generative model; Input the sample video into the image encoding unit in the pre-trained generative model to obtain the first sample features …”)
wherein the conditioning input in the final training example for the initial training example further comprises the subject motion output. (Gao, page 16, lines 6-11, “Based on the sample video and the corresponding sample object image, sample action sequence and sample camera parameters, the pre-trained generation model corresponding to the target video task is trained to obtain the trained video generation model. It should be noted that the implementation method of "training the pre-trained generation model corresponding to the target video task based on the sample video and the sample object image, sample action sequence and sample camera parameters to obtain the trained video generation model" is the same as the training method of the above-mentioned "video generation model" ...”)
17. Claims 1-4, 16, 22, and 24-25 are rejected under 35 U.S.C. 103 as being unpatentable over Gao et al (WO-2026/001197-A1, hereinafter "Gao") in view of Han et al. (US-2021/0144418-A1, hereinafter "Han"), and further in view of Osotsi et al. (US-2024/0202604-A1, hereinafter "Osotsi").
18. As per claim 1, Gao discloses: A method performed by one or more computers, the method comprising: (Gao, page 5, lines 32-37, “This disclosure provides a video generation method, and also relates to a method for generating motion videos of virtual objects, a video editing method, a video generation model training method, ... a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.”)
obtaining a conditioning input for generating an output video that spans a particular time window, wherein the conditioning input comprises a camera motion input that specifies a target motion of a camera that captures the output video during the particular time window, [[wherein the camera motion input is one of a set of camera motion labels that each represent a different camera motion;]] and (Gao, page 1, lines 33-39, “According to a second aspect of the present disclosure, a method for generating motion videos of virtual objects is provided, comprising: acquiring motion video generation data of virtual objects, wherein the motion video generation data includes virtual camera parameters, a reference action sequence of the virtual object, and an image of the virtual object; inputting the motion video generation data into a video generation model to obtain a target motion video of the virtual object, wherein the video generation model is obtained by adjusting the parameters of the parameter encoding unit, parameter cross-attention unit, and temporal attention unit in a pre-trained generation model based on the sample video and the sample object image, sample action sequence, and sample camera parameters of the sample video, and the pre-trained generation model is obtained by adjusting the parameters of the object encoding unit, action encoding unit, and generation unit in an initial generation model based on the sample video.” and page 3, lines 31-40, “By inputting virtual camera parameters into the video generation model, camera movement control of the target video is achieved.” and page 5, lines 27-31, “Through the parameter cross-attention unit, the model can associate virtual camera parameters and video frames in the feature dimension, learn the control and generation of camera parameters, and thus support large-scale camera movement effects. Through the temporal attention unit, the video generation model can learn temporal information, enabling smooth transitions between time sequences in the generated video, making it more realistic and natural.” and page 6, lines 26-44, “Virtual camera parameters can also be called virtual camera trajectory parameters. Virtual camera parameters are a series of data used to describe the motion state of a virtual camera when it is moving or shooting a dynamic scene. The motion state of the virtual camera itself can be controlled by the virtual camera parameters, thereby controlling the camera movement trajectory of the generated video. Specifically, the virtual camera parameters include the camera extrinsic parameters at each time t throughout the entire time period T, such as the parameters of the virtual camera's position and orientation in the world coordinate system, namely the rotation matrix and translation vector.”)
processing the conditioning input using a video generation neural network to generate the output video, comprising: (Gao, page 1, lines 33-39, “According to a second aspect of the present disclosure, a method for generating motion videos of virtual objects is provided, comprising: acquiring motion video generation data of virtual objects, wherein the motion video generation data includes virtual camera parameters, a reference action sequence of the virtual object, and an image of the virtual object; inputting the motion video generation data into a video generation model to obtain a target motion video of the virtual object, wherein the video generation model is obtained by adjusting the parameters of the parameter encoding unit, parameter cross-attention unit, and temporal attention unit …” and page 7, lines 10-14, “Specifically, the action extraction model is used to extract action sequences from the input video. Action extraction models can be deep learning models trained on the basis of CNN models. The parameter extraction model is used to extract camera parameters from the input video.”)
[[tokenizing the camera motion label into a set of one or more tokens and conditioning the video generation neural network on the set of one or more tokens]] when generating the output video. (Gao, page 1, lines 33-39, “According to a second aspect of the present disclosure, a method for generating motion videos of virtual objects is provided, comprising: acquiring motion video generation data of virtual objects, wherein the motion video generation data includes virtual camera parameters, a reference action sequence of the virtual object, and an image of the virtual object; inputting the motion video generation data into a video generation model to obtain a target motion video of the virtual object, wherein the video generation model is obtained by adjusting the parameters of the parameter encoding unit, parameter cross-attention unit, and temporal attention unit …” and page 7, lines 10-14, “Specifically, the action extraction model is used to extract action sequences from the input video. Action extraction models can be deep learning models trained on the basis of CNN models. The parameter extraction model is used to extract camera parameters from the input video.”)
19. Gao doesn't explicitly disclose but Han discloses: [[obtaining a conditioning input for generating an output video that spans a particular time window, wherein the conditioning input comprises a camera motion input that specifies a target motion of a camera that captures the output video during the particular time window,]] wherein the camera motion input is one of a set of camera motion labels that each represent a different camera motion; [[and]] (Han, [0054] ,”In an implementation, a content side model may be adopted for determining the content score of the candidate video as discussed above. ... The content side model 230 may be established based on various techniques, e.g., machine learning, deep learning, etc. Features adopted by the content side model 230 may comprise at least one of: shot transition, camera motion, scene, human, human motion, object, object motion, text information, audio attribute, and video metadata, as discussed above. In terms of function, the content side model 230 may be, e.g., a regression model, a classification model, etc. ... Training data for the content side model 230 may be obtained through: obtaining a group of videos to be used for training; for each video in the group of videos, labeling respective values corresponding to the features of the content side model, and labeling a content score for the video; and forming training data from the group of videos with respective labels.”)
[[tokenizing the]] camera motion label [[into a set of one or more tokens and conditioning the video generation neural network on the set of one or more tokens when generating the output video.]] (See Han, [0054] above.)
20. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of Gao to include the disclosure of a set of camera motion labels that each represent a different camera motion, of Han. The motivation for this modification could have been to help classify various different movements of a camera so each can be uniquely defined within the machine learning model. Doing so helps to categorize similar movements with each other and keep track of the properties of the camera motion.
21. Gao in view of Han doesn't explicitly disclose but Osotsi discloses: tokenizing the [[camera motion label]] into a set of one or more tokens and conditioning the video generation neural network on the set of one or more tokens [[when generating the output video.]] (Osotsi, [0095], “The process 800 begins at step/operation 802 when the predictive data analysis computing entity 106 generates, via an embedding machine learning model, for each token of the input sequence of tokens, a token embedding. The input sequence of tokens may comprise a tokenization of a plurality of classification labels associated with predictive output data from a predictive system.” and [0027], “The term “classification label” may refer to a data construct that describes a label that associates target features, properties, or characteristics to data. For example, a classification label may be generated as output from a predictive system or a predictive machine learning model for describing a prediction on given prediction input data. Classification labels may be applied to various kinds of data types such as images, text, audio, video files, and application files that may be used to, for example, train a machine learning model. In some embodiments, classification labels may comprise descriptions, tags, or identifiers that classify or emphasize features present in training data entries which may be analyzed by machine learning models for relationships or patterns to perform a predictive inference.” and [0026], “A predictive machine learning model may be trained to generate predictive output data by learning from a training dataset. A training dataset may (e.g., supervised learning via labeled data) or may not (e.g., unsupervised learning via unlabeled data) include classification labels that characterize data in the training dataset. As such, a predictive machine learning model may learn known (e.g., supervised) or unknown (e.g., unsupervised) relationships or patterns from a training dataset.” and [0064], “The predictive data analysis computing entity 106 may also include, or be in communication with, one or more output elements (not shown), such as audio output, video output ...”)
22. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of Gao in view of Han to include the disclosure of tokenizing the camera motion label into a set of one or more tokens and conditioning the video generation neural network on the set of one or more tokens, of Osotsi. The motivation for this modification could have been to represent motion of a camera (and its as various properties) into smaller tokenized data that can be used for queries and analysis. Tokenizing this data further helps with classification of the specific movement and allows for quick query and comparison with other data.
23. As per claim 2, Gao in view of Han, and further in view of Osotsi discloses: The method of claim 1, wherein the output video depicts a subject, and wherein the conditioning input comprises a subject motion input that specifies a target motion of the subject during the particular time window. (Gao, page 1, lines 33-39, “According to a second aspect of the present disclosure, a method for generating motion videos of virtual objects is provided, comprising: acquiring motion video generation data of virtual objects, wherein the motion video generation data includes virtual camera parameters, a reference action sequence of the virtual object, and an image of the virtual object; inputting the motion video generation data into a video generation model to obtain a target motion video of the virtual object …” and page 6, lines 26-44, “The target object refers to the object appearing in the target video. The target object can be a person, an animal, a virtual character, etc.” and page 6, lines 26-44, “Specifically, the virtual camera parameters include the camera extrinsic parameters at each time t throughout the entire time period T, such as the parameters of the virtual camera's position and orientation in the world coordinate system, namely the rotation matrix and translation vector.”)
24. As per claim 3, Gao in view of Han, and further in view of Osotsi discloses: The method of claim 2, wherein the subject motion input specifies the target motion of the subject relative to the target motion of the camera during the particular time window. (Gao, page 7, lines 30-35, “Specifically, a target video refers to a video in which the actions of the target object conform to the reference action sequence and the video camera movement effect conforms to the virtual camera parameters.” and page 6, lines 26-44, “Specifically, the virtual camera parameters include the camera extrinsic parameters at each time t throughout the entire time period T, such as the parameters of the virtual camera's position and orientation in the world coordinate system, namely the rotation matrix and translation vector.”)
25. As per claim 4, Gao in view of Han, and further in view of Osotsi discloses: A method performed by one or more computers, the method comprising: (Gao, page 5, lines 32-37, “This disclosure provides a video generation method, and also relates to a method for generating motion videos of virtual objects, a video editing method, a video generation model training method, ... a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.”)
obtaining a conditioning input for generating an output video that spans a particular time window, wherein the conditioning input comprises a camera motion input that specifies a target motion of a camera that captures the output video during the particular time window, (Gao, page 1, lines 33-39, “According to a second aspect of the present disclosure, a method for generating motion videos of virtual objects is provided, comprising: acquiring motion video generation data of virtual objects, wherein the motion video generation data includes virtual camera parameters, a reference action sequence of the virtual object, and an image of the virtual object; inputting the motion video generation data into a video generation model to obtain a target motion video of the virtual object, wherein the video generation model is obtained by adjusting the parameters of the parameter encoding unit, parameter cross-attention unit, and temporal attention unit in a pre-trained generation model based on the sample video and the sample object image, sample action sequence, and sample camera parameters of the sample video, and the pre-trained generation model is obtained by adjusting the parameters of the object encoding unit, action encoding unit, and generation unit in an initial generation model based on the sample video.” and Gao, page 3, lines 31-40, “By inputting virtual camera parameters into the video generation model, camera movement control of the target video is achieved.” and Gao, page 5, lines 27-31, “Through the parameter cross-attention unit, the model can associate virtual camera parameters and video frames in the feature dimension, learn the control and generation of camera parameters, and thus support large-scale camera movement effects. Through the temporal attention unit, the video generation model can learn temporal information, enabling smooth transitions between time sequences in the generated video, making it more realistic and natural.” and Gao, page 6, lines 26-44, “The target object refers to the object appearing in the target video. The target object can be a person, an animal, a virtual character, etc.” and Gao, page 6, lines 26-44, “Virtual camera parameters can also be called virtual camera trajectory parameters. Virtual camera parameters are a series of data used to describe the motion state of a virtual camera when it is moving or shooting a dynamic scene. The motion state of the virtual camera itself can be controlled by the virtual camera parameters, thereby controlling the camera movement trajectory of the generated video. Specifically, the virtual camera parameters include the camera extrinsic parameters at each time t throughout the entire time period T, such as the parameters of the virtual camera's position and orientation in the world coordinate system, namely the rotation matrix and translation vector.”) wherein the camera motion input comprises a respective score for each of a set of a plurality of camera motion labels that each represent a different camera motion; and (Han, [0054], “In an implementation, a content side model may be adopted for determining the content score of the candidate video as discussed above. ... The content side model 230 may be established based on various techniques, e.g., machine learning, deep learning, etc. Features adopted by the content side model 230 may comprise at least one of: shot transition, camera motion, scene, human, human motion, object, object motion, text information, audio attribute, and video metadata, as discussed above. In terms of function, the content side model 230 may be, e.g., a regression model, a classification model, etc. ... Training data for the content side model 230 may be obtained through: obtaining a group of videos to be used for training; for each video in the group of videos, labeling respective values corresponding to the features of the content side model, and labeling a content score for the video; and forming training data from the group of videos with respective labels.” and Han, [0053], “It should be appreciated that any two or more of the above discussed shot transition, camera motion, scene, human, human motion, object, object motion, text information, audio attribute, and video metadata may be combined together so as to determine the content score of the candidate video. For example, for a video recording a cute dog's activities, this video may contain a large amount of camera motions and object motions but does not include any speech or music, and thus a content score indicating a high importance of visual information may be determined for this video.”)
processing the conditioning input using a video generation neural network to generate the output video , comprising: (Gao, page 1, lines 33-39, “According to a second aspect of the present disclosure, a method for generating motion videos of virtual objects is provided, comprising: acquiring motion video generation data of virtual objects, wherein the motion video generation data includes virtual camera parameters, a reference action sequence of the virtual object, and an image of the virtual object; inputting the motion video generation data into a video generation model to obtain a target motion video of the virtual object, wherein the video generation model is obtained by adjusting the parameters of the parameter encoding unit, parameter cross-attention unit, and temporal attention unit …” and Gao, page 7, lines 10-14, “Specifically, the action extraction model is used to extract action sequences from the input video. Action extraction models can be deep learning models trained on the basis of CNN models. The parameter extraction model is used to extract camera parameters from the input video.”)
tokenizing the respective scores for the plurality of camera motion labels into a set of one or more tokens and conditioning the video generation neural network on the set of one or more tokens. (Osotsi, [0095], “The process 800 begins at step/operation 802 when the predictive data analysis computing entity 106 generates, via an embedding machine learning model, for each token of the input sequence of tokens, a token embedding. The input sequence of tokens may comprise a tokenization of a plurality of classification labels associated with predictive output data from a predictive system.” and Osotsi, [0027], “The term “classification label” may refer to a data construct that describes a label that associates target features, properties, or characteristics to data. For example, a classification label may be generated as output from a predictive system or a predictive machine learning model for describing a prediction on given prediction input data. Classification labels may be applied to various kinds of data types such as images, text, audio, video files, and application files that may be used to, for example, train a machine learning model. In some embodiments, classification labels may comprise descriptions, tags, or identifiers that classify or emphasize features present in training data entries which may be analyzed by machine learning models for relationships or patterns to perform a predictive inference.” and Osotsi, [0026], “A predictive machine learning model may be trained to generate predictive output data by learning from a training dataset. A training dataset may (e.g., supervised learning via labeled data) or may not (e.g., unsupervised learning via unlabeled data) include classification labels that characterize data in the training dataset. As such, a predictive machine learning model may learn known (e.g., supervised) or unknown (e.g., unsupervised) relationships or patterns from a training dataset.” and Osotsi, [0064], “The predictive data analysis computing entity 106 may also include, or be in communication with, one or more output elements (not shown), such as audio output, video output ...”; Examiner’s note: As disclosed in Osotsi, [0026], the predictive machine can generate predictive output data. Predictive output data includes scoring data for classification purposes.)
26. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of Gao to include the disclosure of the camera motion input comprises a respective score for each of a set of a plurality of camera motion labels that each represent a different camera motion, of Han. The motivation for this modification could have been to assist with evaluating the motion of a camera. In particular, this allows for comparison of various motions to determine how similar they are and also for classification purposes.
In addition, before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of Gao in view of Han to include the disclosure of tokenizing the respective scores for the plurality of camera motion labels into a set of one or more tokens and conditioning the video generation neural network on the set of one or more tokens, of Osotsi. The motivation for this modification could have been to represent scores for various motions of a camera into smaller tokenized data that can be used for queries and analysis. Tokenizing this data further helps with classification of the specific movement and allows for quick query and comparison with other data.
27. As per claim 16, Gao in view of Han, and further in view of Osotsi discloses: The method of claim 1, wherein the conditioning input further comprises one or more context images for the output video. (Gao, page 1, lines 33-39, “According to a second aspect of the present disclosure, a method for generating motion videos of virtual objects is provided, comprising: acquiring motion video generation data of virtual objects, wherein the motion video generation data includes virtual camera parameters, a reference action sequence of the virtual object, and an image of the virtual object …”)
28. Claim 22, which is similar in scope to dependent claim 16 and independent claim 4, is thus rejected under the same rationale as described above.
29. Claim 24, which is similar in scope to dependent claim 2 and independent claim 4, is thus rejected under the same rationale as described above.
30. Claim 25, which is similar in scope to dependent claims 2-3 and independent claim 4, is thus rejected under the same rationale as described above.
31. Claims 7-8 and 11-12 are rejected under 35 U.S.C. 103 as being unpatentable over Gao et al (WO-2026/001197-A1, hereinafter "Gao") in view of Han et al. (US-2021/0144418-A1, hereinafter "Han"), further in view of Osotsi et al. (US-2024/0202604-A1, hereinafter "Osotsi"), and further in view of Meier et al. (US-2024/0314406-A1, hereinafter "Meier").
32. As per claim 7, Gao in view of Han, and further in view of Osotsi discloses: The method of claim 2, wherein the subject motion input [[is one of a set of subject motion labels that each represent a different subject motion.]] (See rejection for claim 2.)
33. Gao in view of Han, and further in view of Osotsi doesn't explicitly disclose but Meier discloses: [[The method of claim 4, wherein the subject motion input]] is one of a set of subject motion labels that each represent a different subject motion. (Meier, [0047], “Similarly, in some instances, the motion attribute classifier 710 can be trained and fine-tuned with video segments depicting objects or subjects exhibiting different motion and including one or more labels characterizing the respective motion exhibited by the objects or subjects.”)
34. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 2 of Gao in view of Han, and further in view of Osotsi to include the disclosure of a set of subject motion labels that each represent a different subject motion, of Meier. The motivation for this modification could have been to help classify various different movements of a subject so each can be uniquely defined within the machine learning model. Doing so helps to categorize similar movements with each other and keep track of the properties of the subject motion.
35. As per claim 8, Gao in view of Han, further in view of Osotsi, and further in view of Meier discloses: The method of claim 7, wherein processing the conditioning input using a video generation neural network to generate the output video comprises tokenizing the subject motion label into a set of one or more tokens and conditioning the video generation neural network on the set of one or more tokens. (Meier, [0047], “Similarly, in some instances, the motion attribute classifier 710 can be trained and fine-tuned with video segments depicting objects or subjects exhibiting different motion and including one or more labels characterizing the respective motion exhibited by the objects or subjects.” and Osotsi, [0095], “The process 800 begins at step/operation 802 when the predictive data analysis computing entity 106 generates, via an embedding machine learning model, for each token of the input sequence of tokens, a token embedding. The input sequence of tokens may comprise a tokenization of a plurality of classification labels associated with predictive output data from a predictive system.” and Osotsi, [0027], “The term “classification label” may refer to a data construct that describes a label that associates target features, properties, or characteristics to data. For example, a classification label may be generated as output from a predictive system or a predictive machine learning model for describing a prediction on given prediction input data. Classification labels may be applied to various kinds of data types such as images, text, audio, video files, and application files that may be used to, for example, train a machine learning model. In some embodiments, classification labels may comprise descriptions, tags, or identifiers that classify or emphasize features present in training data entries which may be analyzed by machine learning models for relationships or patterns to perform a predictive inference.” and Osotsi, [0026], “A predictive machine learning model may be trained to generate predictive output data by learning from a training dataset. A training dataset may (e.g., supervised learning via labeled data) or may not (e.g., unsupervised learning via unlabeled data) include classification labels that characterize data in the training dataset. As such, a predictive machine learning model may learn known (e.g., supervised) or unknown (e.g., unsupervised) relationships or patterns from a training dataset.” and Osotsi, [0064], “The predictive data analysis computing entity 106 may also include, or be in communication with, one or more output elements (not shown), such as audio output, video output ...”)
36. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 7 of Gao in view of Han, and further in view of Meier to include the disclosure of tokenizing the subject motion label into a set of one or more tokens and conditioning the video generation neural network on the set of one or more tokens, of Osotsi. The motivation for this modification could have been to represent motion of a subject (and its as various properties) into smaller tokenized data that can be used for queries and analysis. Tokenizing this data further helps with classification of the specific movement and allows for quick query and comparison with other data.
37. As per claim 11, Gao in view of Han, further in view of Osotsi, and further in view of Meier discloses: The method of claim 2, wherein the subject motion input comprises a respective score for each of a set of a plurality of subject motion labels that each represent a different subject motion. (Meier, [0047], “Similarly, in some instances, the motion attribute classifier 710 can be trained and fine-tuned with video segments depicting objects or subjects exhibiting different motion and including one or more labels characterizing the respective motion exhibited by the objects or subjects.” and [0056], “The match component 1030 includes a feature extractor that is configured to extract features from the video segments 402 and the candidate videos 1070 and a match scorer that is configured to determine match scores and trimming positions for the candidate videos 1070. ... In some instances, the feature extractor is configured to extract features from each video segment of the video segments 402 and each candidate video of the candidate videos 1070 by using one or more image processing algorithms and/or natural language processing algorithms. In some instances, the image processing algorithms include object/subject detection algorithms, pose detection algorithms, motion detection/optical flow algorithms, and the like.” and [0057], “The match scorer is configured to determine a match score and trimming positions for each candidate video of the candidate videos 1070. In some instances, the match scorer is configured to determine a match score for each candidate video of the candidate videos 1070 by determining the similarity between the features extracted from the respective candidate video and the features extracted from a video segment corresponding to the respective candidate video.”)
38. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 2 of Gao in view of Han, and further in view of Osotsi to include the disclosure of the subject motion input comprising a respective score for each of a set of a plurality of subject motion labels that each represent a different subject motion, of Meier. The motivation for this modification could have been to assist with evaluating the motion of a subject. In particular, this allows for comparison of various motions to determine how similar they are and also for classification purposes.
39. As per claim 12, Gao in view of Han, further in view of Osotsi, and further in view of Meier discloses: The method of claim 11, wherein processing the conditioning input using a video generation neural network to generate the output video comprises tokenizing the respective scores for the subject motion labels into a set of one or more tokens and conditioning the video generation neural network on the set of one or more tokens. (Osotsi, [0095], “The process 800 begins at step/operation 802 when the predictive data analysis computing entity 106 generates, via an embedding machine learning model, for each token of the input sequence of tokens, a token embedding. The input sequence of tokens may comprise a tokenization of a plurality of classification labels associated with predictive output data from a predictive system.” and [0027], “The term “classification label” may refer to a data construct that describes a label that associates target features, properties, or characteristics to data. For example, a classification label may be generated as output from a predictive system or a predictive machine learning model for describing a prediction on given prediction input data. Classification labels may be applied to various kinds of data types such as images, text, audio, video files, and application files that may be used to, for example, train a machine learning model. In some embodiments, classification labels may comprise descriptions, tags, or identifiers that classify or emphasize features present in training data entries which may be analyzed by machine learning models for relationships or patterns to perform a predictive inference.” and [0026], “A predictive machine learning model may be trained to generate predictive output data by learning from a training dataset. A training dataset may (e.g., supervised learning via labeled data) or may not (e.g., unsupervised learning via unlabeled data) include classification labels that characterize data in the training dataset. As such, a predictive machine learning model may learn known (e.g., supervised) or unknown (e.g., unsupervised) relationships or patterns from a training dataset.” and [0064], “The predictive data analysis computing entity 106 may also include, or be in communication with, one or more output elements (not shown), such as audio output, video output ...”; Examiner’s note: As disclosed in Osotsi, [0026], the predictive machine can generate predictive output data. Predictive output data includes scoring data for classification purposes.)
40. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 11 of Gao in view of Han, and further in view of Meier to include the disclosure of tokenizing the respective scores for the subject motion labels into a set of one or more tokens and conditioning the video generation neural network on the set of one or more tokens, of Osotsi. The motivation for this modification could have been to represent scores for various motions of a subject into smaller tokenized data that can be used for queries and analysis. Tokenizing this data further helps with classification of the specific movement and allows for quick query and comparison with other data.
41. Claims 15 and 21 is rejected under 35 U.S.C. 103 as being unpatentable over Gao et al (WO-2026/001197-A1, hereinafter "Gao") in view of Han et al. (US-2021/0144418-A1, hereinafter "Han"), further in view of Osotsi et al. (US-2024/0202604-A1, hereinafter "Osotsi"), and further in view of Affleck-Boldt (US-12511837-B1).
42. As per claim 15, Gao in view of Han, and further in view of Osotsi discloses: The method claim 1, wherein the conditioning input [[further comprises text or audio charactering one or more properties of the output video.]] (See rejection for claim 1.)
43. Gao in view of Han, and further in view of Osotsi doesn't explicitly disclose but Affleck-Boldt discloses: [[The method of claim 4, wherein the conditioning input]] further comprises text or audio charactering one or more properties of the output video. (Affleck-Boldt, col. 4, lines 51-61, “Training with video data, video metadata and/or Lidar data not only enhances the AI model's ability to simulate professional camera movements and adjustments but also significantly improves the spatial awareness of the AI model, contributing to the realism and quality of video content generated using the AI model. Furthermore, the present techniques may include a user interface (UI) that allows users to specify video characteristics using inputs (e.g., text, images, audio, etc.). In some aspects, the inputs may be the same metadata language the AI was trained on.”)
44. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 1 of Gao in view of Han, and further in view of Osotsi to include the disclosure of the conditioning input further comprises text or audio charactering one or more properties of the output video, of Affleck-Boldt. The motivation for this modification could have been to allow a user to provide further details or parameters in text or audio form regarding the video output. By doing so, this could provide a more unique and specific generated video that includes the parameters specified by a user.
45. Claim 21, which is similar in scope to dependent claim 15 and independent claim 4, is thus rejected under the same rationale as described above. The motivation for this modification is the same as claim 15.
46. Claim 17 and 23 is rejected under 35 U.S.C. 103 as being unpatentable over Gao et al (WO-2026/001197-A1, hereinafter "Gao") in view of Han et al. (US-2021/0144418-A1, hereinafter "Han"), further in view of Osotsi et al. (US-2024/0202604-A1, hereinafter "Osotsi"), and further in view of Meingast (US-2023/0016900-A1).
47. As per claim 17, Gao in view of Han, and further in view of Osotsi discloses: The method of claim 1, wherein the video generation neural network [[has been trained on training examples that are generated by mapping camera track data for an input video segment to a camera motion label for the input video segment using a generative neural network.]] (See rejection for claim 1.)
48. Gao in view of Han, and further in view of Osotsi doesn't explicitly disclose but Meingast discloses: [[The method of claim 5, wherein the video generation neural network]] has been trained on training examples that are generated by mapping camera track data for an input video segment to a camera motion label for the input video segment using a generative neural network. (Meingast, [0067], “For example, the inter-camera training data set 193 may be input to the machine learning system 130 in order to train the machine learning system 130. The inter-camera training data set 193 may be input to the machine learning system 130 in any suitable manner. The inter-camera training data set 193 may include labels that correlate motion paths across different motion path data and indicate whether the correlated motion paths are normal or anomalous motion paths. The labels may be generated in any suitable manner, including, for example, through human labeling or through cluster analysis of the motion paths in the inter-camera training data set 193.” and [0035], “A machine learning system may be trained using the data generated by a camera to generate a model for the camera that may be used to identify normal and anomalous paths of motion in the field of view of the camera. For example, motion paths and contextual data from a camera in an environment may be input to a machine learning system. The machine learning system may be, for example, a recurrent neural network, including a deep learning neural network.”)
49. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 1 of Gao in view of Han, and further in view of Osotsi to include the disclosure that the video generation neural network has been trained on training examples that are generated by mapping camera track data for an input video segment to a camera motion label for the input video segment using a generative neural network, of Meingast. The motivation for this modification could have been to train unique track movements of a camera so that it can be reproduced in other target videos. This simplifies the use of unique camera tracks for other generated videos, especially if the camera track might otherwise be difficult execute.
50. Claim 23, which is similar in scope to dependent claim 17 and independent claim 4, is thus rejected under the same rationale as described above. The motivation for this modification is the same as claim 17.
51. Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Gao et al (WO-2026/001197-A1, hereinafter "Gao") in view of Han et al. (US-2021/0144418-A1, hereinafter "Han"), and further in view of Ji et al. (CN-108171222-A, hereinafter "Ji"). (Examiner’s note: Citations to Ji use the original CN-108171222-A document locations.)
52. As per claim 20, Gao in view of Han discloses: The method of claim 18, wherein the first input comprises the features and wherein the features comprise features [[of an optical flow prediction for the video frames in the target video.]] (See rejection for claim 18.)
53. Gao in view of Han doesn't explicitly disclose but Ji discloses: [[The method of claim 18, wherein the first input comprises the features and wherein the features]] comprise features of an optical flow prediction for the video frames in the target video. (Ji, page 4, [0006]-[0007], “According to one aspect of this disclosure, a video classification method is provided, comprising: extracting video frames and motion vectors from a video to be classified; extracting optical flow from the video to be classified using an optical flow neural network; adjusting the motion vectors using the optical flow; inputting the video frames, the extracted optical flow, and the adjusted motion vectors into a multi-flow neural network, and determining the category of the video to be classified based on the output of the multi-flow neural network. In one possible implementation, the method further includes: training the optical flow neural network with adjacent video frames and motion vectors corresponding to the adjacent video frames as inputs and optical flow corresponding to the adjacent video frames as ground truth.”)
54. Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of claim 18 of Gao in view of Han to include the disclosure of the first input comprising of features of an optical flow prediction for video frames in the target video, of Ji. The motivation for this modification could have been to help keep track of the movement of the camera and/or objects in a video. This may help provide further information about the motion of a video and can add to its properties. In addition, this may serve as an alternative way to determine motion and/or orientation of a camera or object.
Conclusion
55. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Ruiz Gutierrez et al. (US-2026/0120720-A1) discloses methods and systems for generating output videos that have new camera motion using a generative neural network.
56. THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
57. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW CLOTHIER whose telephone number is (571)272-4667. The examiner can normally be reached Mon-Fri 8:00am-4:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached at (571)272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MATTHEW CLOTHIER/Examiner, Art Unit 2614
/KENT W CHANG/Supervisory Patent Examiner, Art Unit 2614