DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “obtainment unit that obtains” and “generation unit that generates” in claim 1.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 14 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter. Claim 14 recites “A program that causes a computer to implement” in line 1. This is pure software and thus directed to non-statutory subject matter.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-7 10, 13 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Yang et al. (“TransMoMo: Invariance-Driven Unsupervised Video Motion Retargeting”) and Gallo et al. (US 20230137403).
Regarding claim 1, Yang et al. discloses an information processing device comprising:
an obtainment unit that obtains a motion feature indicating a feature generated from motion data (“The Motion Retargeting Network decomposes 2D joint input sequences as a motion code that represents the movements of the actor” at section 3, paragraph 2, line 1); and
a generation unit that generates an image corresponding to the motion data used to generate the motion feature obtained by the obtainment unit (“For transferring motion from a source video to a target video, we first use an off-the-shelf 2D keypoints detector to extract joint sequences from videos. By combining the motion code encoded from the source sequence and the structure code encoded from the target sequence, our model then yields a transferred 3D joint sequence. The transferred sequence is then projected back to 2D with any desired view-angle. Finally, we convert the 2D joint sequence frame-by-frame to a pixel-level representation, i.e., label maps. These label maps are fed into a pre-trained image-to-image generator to render the transferred video” at section 3, paragraph 3).
Yang et al. does not explicitly disclose that the motion feature is delivered from a communication terminal.
However, as demonstrated by Gallo et al., motion retargeting is commonly conducted in a distributed computing environment (see paragraph 0393) where data is transferred from a server to a client device (see paragraph 0060). As such, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to utilize a distributed computing as taught by Gallo et al. to provide a user of the rendered content while leveraging the computing power at the server end.
Regarding claim 2, Yang et al. discloses a device wherein the motion feature of the motion data is estimated using an estimation encoder obtained by learning a relationship between the motion data and the motion feature (“The motion encoder uses several layers of one dimen-sional temporal convolution to extract motion information: Em(x) = m ∈ RM×Cm, where M is the sequence length after encoding and Cm is the number of channels. Note that the motion code m is variable in length so as to preserve temporal information.” at section 3.1, paragraph 2).
Regarding claim 3, Yang et al. discloses a device wherein the generation unit generates an image corresponding to the motion data based on a relationship between the motion feature and the image, the relationship having been learned in advance (“The decoder takes the motion, body and view codes as input and reconstructs a 3D joint sequence G(m, ¯s, ¯v) = ˆX ∈ RT×3N through convolution layers, in symmetry with the encoders. Our discriminator D is a temporal convolutional network that is similar to our motion encoder: D(x) ∈ RM.” at section 3.1, last paragraph); the encoder and decoder is trained using the loss functions in sections 3.2.1-3.2.4).
Regarding claim 4, Yang et al. discloses a device wherein the generation unit generates an image corresponding to the motion data based on a relationship between: the motion feature and a camera parameter pertaining to viewpoint information indicating a viewpoint from which the image is viewed; and the image, the relationship having been learned in advance (“The Motion Retargeting Network decomposes 2D joint input sequences as a motion code that represents the movements of the actor, a structure code that represents the body shape of the actor and a view-angle code that represents the camera angle. The decoder takes any combination of the latent codes and produces a reconstructed 3D joint sequence, which automatically isolates view from motion and structure.” at section 3, paragraph 2; the encoder and decoder is trained using the loss functions in sections 3.2.1-3.2.4).
Regarding claim 5, Yang et al. discloses a device wherein the motion data includes designated motion data indicating motion data designated by a user, and the obtainment unit obtains the motion feature of the designated motion data (as demonstrated in Figure 2, the user selects the Source Video which constitutes the motion data that will be used to determine the motion feature).
Regarding claim 6, Yang et al. discloses a device wherein the generation unit generates an image corresponding to the designated motion data based on a relationship between: the motion feature, a camera parameter pertaining to viewpoint information indicating a viewpoint from which the image is viewed (“The Motion Retargeting Network decomposes 2D joint input sequences as a motion code that represents the movements of the actor, a structure code that represents the body shape of the actor and a view-angle code that represents the camera angle. The decoder takes any combination of the latent codes and produces a reconstructed 3D joint sequence, which automatically isolates view from motion and structure.” at section 3, paragraph 2), and character information indicating a character to perform the motion data (e.g. the target video character); and the image, the relationship having been learned in advance (the encoder and decoder is trained using the loss functions in sections 3.2.1-3.2.4).
Regarding claim 7, Yang et al. discloses a device wherein the generation unit generates an image corresponding to the designated motion data at a time of input (“For transferring motion from a source video to a target video, we first use an off-the-shelf 2D keypoints detector to extract joint sequences from videos. By combining the motion code encoded from the source sequence and the structure code encoded from the target sequence, our model then yields a transferred 3D joint sequence. The transferred sequence is then projected back to 2D with any desired view-angle. Finally, we convert the 2D joint sequence frame-by-frame to a pixel-level representation, i.e., label maps. These label maps are fed into a pre-trained image-to-image generator to render the transferred video” at section 3, paragraph 3).
Regarding claim 10, Yang et al. discloses a device wherein the motion feature of the designated motion data is estimated using an encoder corresponding to the designated motion data (as seen in Figure 2, there is a designated encoder for the Source Video).
Regarding claim 13, Yang et al. discloses an information processing method comprising:
obtaining a motion feature indicating a feature generated from motion data (“The Motion Retargeting Network decomposes 2D joint input sequences as a motion code that represents the movements of the actor” at section 3, paragraph 2, line 1); and
generating an image corresponding to the motion data used to generate the motion feature obtained by the obtainment unit (“For transferring motion from a source video to a target video, we first use an off-the-shelf 2D keypoints detector to extract joint sequences from videos. By combining the motion code encoded from the source sequence and the structure code encoded from the target sequence, our model then yields a transferred 3D joint sequence. The transferred sequence is then projected back to 2D with any desired view-angle. Finally, we convert the 2D joint sequence frame-by-frame to a pixel-level representation, i.e., label maps. These label maps are fed into a pre-trained image-to-image generator to render the transferred video” at section 3, paragraph 3).
Yang et al. does not explicitly disclose that the motion feature is delivered from a communication terminal.
However, as demonstrated by Gallo et al., motion retargeting is commonly conducted in a distributed computing environment (see paragraph 0393) where data is transferred from a server to a client device (see paragraph 0060). As such, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to utilize a distributed computing as taught by Gallo et al. to provide a user of the rendered content while leveraging the computing power at the server end.
Regarding claim 14, Yang et al. discloses a program (implied that the system is a programmed computer) that causes a computer to implement:
an obtainment function that obtains a motion feature indicating a feature generated from motion data (“The Motion Retargeting Network decomposes 2D joint input sequences as a motion code that represents the movements of the actor” at section 3, paragraph 2, line 1); and
a generation function that generates an image corresponding to the motion data used to generate the motion feature obtained by the obtainment unit (“For transferring motion from a source video to a target video, we first use an off-the-shelf 2D keypoints detector to extract joint sequences from videos. By combining the motion code encoded from the source sequence and the structure code encoded from the target sequence, our model then yields a transferred 3D joint sequence. The transferred sequence is then projected back to 2D with any desired view-angle. Finally, we convert the 2D joint sequence frame-by-frame to a pixel-level representation, i.e., label maps. These label maps are fed into a pre-trained image-to-image generator to render the transferred video” at section 3, paragraph 3).
Yang et al. does not explicitly disclose that the motion feature is delivered from a communication terminal.
However, as demonstrated by Gallo et al., motion retargeting is commonly conducted in a distributed computing environment (see paragraph 0393) where data is transferred from a server to a client device (see paragraph 0060). As such, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to utilize a distributed computing as taught by Gallo et al. to provide a user of the rendered content while leveraging the computing power at the server end.
Claim(s) 8 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of Yang et al. and Gallo et al. as applied to claim 7 above, and further in view of Alattar et al. (“A System for Mitigating the Problem of Deepfake News Videos Using Watermarking”).
The Yang et al. and Gallo et al. combination discloses the elements of claim 7 as described above.
The Yang et al. and Gallo et al. combination does not explicitly disclose that the motion feature includes noise information that adds predetermined modification processing to the image generated by the generation unit.
Alattar et al. teaches a device in the same field of endeavor of generating hallucinated videos, wherein the motion feature includes noise information that adds predetermined modification processing to the image generated by the generation unit (“The proposed system requires that trusted news entities embed unique digital watermarks in their video during video capturing or before video distribution. They can automatically serialize and embed the watermark at the time of streaming or downloading of the video by a recipient” at section 3, paragraph 2, line 1).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to utilize a watermarking as taught by Alattar et al. on the generated retargeted image of the Yang et al. and Gallo et al. combination as it has the “extra benefit of enabling tracking distribution of copies or derivative works of content” (Alattar et al. at section 3, paragraph 2, line 5) and prevent the misuse of the generated content (see Alattar et al. in the remaining paragraphs of section 3).
Claim(s) 9 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of Yang et al. and Gallo et al. as applied to claim 7 above, and further in view of Wu et al. (US 20220237879).
The Yang et al. and Gallo et al. combination discloses the elements of claim 7 as described above.
The Yang et al. and Gallo et al. combination does not explicitly disclose that the generation unit generates an image in which predetermined modification processing is added to the motion data used to generate the motion feature obtained by the obtainment unit.
Wu et al. teaches a device in the same field of endeavor of motion transfer to avatar, wherein the generation unit generates an image in which predetermined modification processing is added to the motion data used to generate the motion feature obtained by the obtainment unit (“Model 1000 animates single-layer avatars 1021A-1, 1021A-2, 1021B-1, and 1021B-2 (hereinafter, collectively referred to as “single-layer avatars 1021-1 and 1021-2”) via random sampling of a unit Gaussian distribution (e.g., clothing inputs 604B), and use the resulting noise values for imputation of the latent code, where available. For the sampled latent code in avatars 1021A-2 and 1021-B-2, model 1000 feeds the skeletal pose and facial keypoints together, into the decoder networks (e.g., networks 600)” at paragraph 0079, line 9).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to utilize a latent noise vector as taught by Wu et al. in the encoder of the Yang et al. and Gallo et al. combination to generate a realistic animation.
Claim(s) 1 and 11-14 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Liu et al. (“Aggregated Multi-GANs for Controlled 3D Human Motion Prediction”) and Gallo et al.
Regarding claim 1, Liu et al. discloses an information processing device comprising:
an obtainment unit that obtains a motion feature indicating a feature generated from motion data (“The overall framework of our method is illustrated in Fig. 2. There are five sub-GANs (local GANs), each taking care of a kinematic chain. The sub-GANs will now model motion generation on a lower dimensional manifold which is advantageous from a computational perspective and may be favorable in avoiding mode collapse. After a local discrimination process, the outputs from all sub-GANs are then combined via an aggregation layer. The output of the aggregation layer is then passed to the global critic for assessing the quality of the generated motion sequence in its entirety” at page 2227, Architecture Overview, paragraph 1); and
a generation unit that generates an image corresponding to the motion data used to generate the motion feature obtained by the obtainment unit (“The critic module consists of Local Critics for each sub-GAN Dj and a Global Critic DG after the aggregation layer. A three-layers fully connected feedforward network is used in both types of critics. The local critics judge the generation of the five branches of the human body to ensure the accuracy of local prediction. The rationale is to first focus on the local context and reduce interference between local components or body parts at the base level. However, we cannot simply ignore the correlation and synchronization between different body parts in executing a motion. We also include a global critic to judge the feasibility of the motion and to ensure the naturalness and accuracy of the entire pose.” at page 2228, Critic).
Liu et al. does not explicitly disclose that the motion feature is delivered from a communication terminal.
However, as demonstrated by Gallo et al., motion retargeting is commonly conducted in a distributed computing environment (see paragraph 0393) where data is transferred from a server to a client device (see paragraph 0060). As such, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to utilize a distributed computing as taught by Gallo et al. to provide a user of the rendered content while leveraging the computing power at the server end.
Regarding claim 11, Liu et al. discloses a device wherein the motion feature of the designated motion data is estimated using a plurality of encoders corresponding to the designated motion data (“The overall framework of our method is illustrated in Fig. 2. There are five sub-GANs (local GANs), each taking care of a kinematic chain.” at page 2227, Architecture Overview, paragraph 1, line 1).
Regarding claim 12, the Liu et al. and Gallo et al. combination discloses a device further comprising: a display unit that displays the image generated by the generation unit (“The observed sequence has 25 frames and the predicted output sequence is also set to 25 frames” Liu et al., page 2230, Dataset and Preprocessing, last sentence; implied that the output sequence is displayed; “In at least one embodiment, client device 502 receiving this content or data can provide this content or data to a corresponding application 504, which may also or alternatively include a data manager 510, content processor 512, or rendering engine 514 for generating or processing at least some of this content or data, such as for usage or presentation via client device 502, such as video content through a display 506” Gallo et al. at paragraph 0060, fifth to last sentence).
Regarding claim 13, Liu et al. discloses an information processing method executed by a computer, the method comprising:
obtaining a motion feature indicating a feature generated from motion data (“The overall framework of our method is illustrated in Fig. 2. There are five sub-GANs (local GANs), each taking care of a kinematic chain. The sub-GANs will now model motion generation on a lower dimensional manifold which is advantageous from a computational perspective and may be favorable in avoiding mode collapse. After a local discrimination process, the outputs from all sub-GANs are then combined via an aggregation layer. The output of the aggregation layer is then passed to the global critic for assessing the quality of the generated motion sequence in its entirety” at page 2227, Architecture Overview, paragraph 1); and
generating an image corresponding to the motion data used to generate the motion feature obtained by the obtainment unit (“The critic module consists of Local Critics for each sub-GAN Dj and a Global Critic DG after the aggregation layer. A three-layers fully connected feedforward network is used in both types of critics. The local critics judge the generation of the five branches of the human body to ensure the accuracy of local prediction. The rationale is to first focus on the local context and reduce interference between local components or body parts at the base level. However, we cannot simply ignore the correlation and synchronization between different body parts in executing a motion. We also include a global critic to judge the feasibility of the motion and to ensure the naturalness and accuracy of the entire pose.” at page 2228, Critic).
Liu et al. does not explicitly disclose that the motion feature is delivered from a communication terminal.
However, as demonstrated by Gallo et al., motion retargeting is commonly conducted in a distributed computing environment (see paragraph 0393) where data is transferred from a server to a client device (see paragraph 0060). As such, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to utilize a distributed computing as taught by Gallo et al. to provide a user of the rendered content while leveraging the computing power at the server end.
Regarding claim 14, Liu et al. discloses a program that causes a computer to implement:
an obtainment function that obtains a motion feature indicating a feature generated from motion data (“The overall framework of our method is illustrated in Fig. 2. There are five sub-GANs (local GANs), each taking care of a kinematic chain. The sub-GANs will now model motion generation on a lower dimensional manifold which is advantageous from a computational perspective and may be favorable in avoiding mode collapse. After a local discrimination process, the outputs from all sub-GANs are then combined via an aggregation layer. The output of the aggregation layer is then passed to the global critic for assessing the quality of the generated motion sequence in its entirety” at page 2227, Architecture Overview, paragraph 1); and
a generation function that generates an image corresponding to the motion data used to generate the motion feature obtained by the obtainment unit (“The critic module consists of Local Critics for each sub-GAN Dj and a Global Critic DG after the aggregation layer. A three-layers fully connected feedforward network is used in both types of critics. The local critics judge the generation of the five branches of the human body to ensure the accuracy of local prediction. The rationale is to first focus on the local context and reduce interference between local components or body parts at the base level. However, we cannot simply ignore the correlation and synchronization between different body parts in executing a motion. We also include a global critic to judge the feasibility of the motion and to ensure the naturalness and accuracy of the entire pose.” at page 2228, Critic).
Liu et al. does not explicitly disclose that the motion feature is delivered from a communication terminal.
However, as demonstrated by Gallo et al., motion retargeting is commonly conducted in a distributed computing environment (see paragraph 0393) where data is transferred from a server to a client device (see paragraph 0060). As such, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to utilize a distributed computing as taught by Gallo et al. to provide a user of the rendered content while leveraging the computing power at the server end.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KATRINA R FUJITA whose telephone number is (571)270-1574. The examiner can normally be reached Monday - Friday 9:30-5:30 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sumati Lefkowitz can be reached at 5712723638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KATRINA R FUJITA/ Primary Examiner, Art Unit 2672