DETAILED ACTION
This Office action is in response to the Application filed on November 15, 2024, claims priority to Chinese Patent Application No. 202311594465.X, filed on November 27, 2023. An action on the merits follows. Claims 1-20 are pending on the application.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recites the limitation “determining n noised codes corresponding to n video frames of an original video… determining a text code corresponding to a description text guiding video editing… obtain n denoised codes” in lines 2-5 of the claim. However, the claimed “noised codes” , “text code”, and “denoised codes” terms are not defined by the claims. Additionally, the claimed “n” term recited in claims 1, 3, 5-6, 8, 10-14, 16, and 19-20, respectively, is also not defined by the claims. Specifically, the scope of the claims is unclear without specifying what “n” means because “n” can be interpreted as “a predetermined number,” “any number,” or “a plurality”, for example, and the specification does not define the set of items or conditions that “n” refers to. Therefore, the metes and bounds of the claim are not clearly set forth and the examiner cannot clearly determine which elements are encompassed by the claim language, which renders the claim indefinite.
Claim 1 recites the limitation “performing, using n Unet models obtained by using the text code and copying a Unet model denoising processing on the n noised codes, to obtain n denoised codes, wherein a pre-trained text-to-image model comprises the Unet model, wherein each Unet model comprises a self-attention layer” in lines 4-7 of the claim. However, it is not clear how the claimed “n Unet models” are “obtained by using the text code and copying a Unet model denoising processing on the n noised codes” because it merely recites a use (e.g. “obtained by using“) without reciting any active, positive steps delimiting how this use is actually practiced, and the specification does not describe the actual steps or process for “using n Unet models obtained by using” refers to, which renders the claim indefinite. See MPEP § 2173.05(q).
Claim 1 recites the limitation “a self-attention layer of any ith Unet model, attention calculation based on an output of a target network layer of the ith Unet model and an output of a target network layer in a predetermined target Unet model” in lines 8-10 of the claim. However, the claimed “ith” term recited in claims 1, 4-5 and 19-20, respectively, is also not defined by the claims. Specifically, the scope of the claims is unclear without specifying what “ith” means because “ith” can be interpreted as “a predetermined number,” “any number,” or “a plurality”, for example. Additionally, the claimed “any ith Unet model” limitation is ambiguous or fails to define the scope of the claim with sufficient clarity for a person of ordinary skill in the art to understand and practice the invention and the specification does not define the set of items or conditions that “any” refers to. Therefore, the metes and bounds of the claim are not clearly set forth and the examiner cannot clearly determine which elements are encompassed by the claim language, which renders the claim indefinite.
Claim 1 recites the limitation “performing, using an image decoder, decoding processing on the n denoised codes to obtain n target images” in lines 11-12 of the claim. However, it is not clear how the claimed “decoding processing” is performed by “using an image decoder” because the claimed “decoding processing” term is a functional description and the specification does not describe the actual steps or process for performing the claimed “decoding processing”, which renders the claim indefinite.
Claims 2-18 are rejected by virtue of being dependent upon rejected base claim 1.
Claim 19 recites the limitation “determining n noised codes corresponding to n video frames of an original video… determining a text code corresponding to a description text guiding video editing… obtain n denoised codes” in lines 3-6 of the claim. However, the claimed “noised codes” , “text code”, and “denoised codes” terms are not defined by the claims. Additionally, the claimed “n” term recited in claims 1, 3, 5-6, 8, 10-14, 16, and 19-20, respectively, is also not defined by the claims, as indicated above. Specifically, the scope of the claims is unclear without specifying what “n” means because “n” can be interpreted as “a predetermined number,” “any number,” or “a plurality”, for example, and the specification does not define the set of items or conditions that “n” refers to. Therefore, the metes and bounds of the claim are not clearly set forth and the examiner cannot clearly determine which elements are encompassed by the claim language, which renders the claim indefinite.
Claim 19 recites the limitation “performing, using n Unet models obtained by using the text code and copying a Unet model denoising processing on the n noised codes, to obtain n denoised codes, wherein a pre-trained text-to-image model comprises the Unet model, wherein each Unet model comprises a self-attention layer” in lines 5-8 of the claim. However, it is not clear how the claimed “n Unet models” are “obtained by using the text code and copying a Unet model denoising processing on the n noised codes” because it merely recites a use (e.g. “obtained by using“) without reciting any active, positive steps delimiting how this use is actually practiced, and the specification does not describe the actual steps or process for “using n Unet models obtained by using” refers to, which renders the claim indefinite. See MPEP § 2173.05(q).
Claim 19 recites the limitation “a self-attention layer of any ith Unet model, attention calculation based on an output of a target network layer of the ith Unet model and an output of a target network layer in a predetermined target Unet model” in lines 9-11 of the claim. However, the claimed “ith” term recited in claims 1, 4-5 and 19-20, respectively, is also not defined by the claims, as indicated above. Specifically, the scope of the claims is unclear without specifying what “ith” means because “ith” can be interpreted as “a predetermined number,” “any number,” or “a plurality”, for example. Additionally, the claimed “any ith Unet model” limitation is ambiguous or fails to define the scope of the claim with sufficient clarity for a person of ordinary skill in the art to understand and practice the invention and the specification does not define the set of items or conditions that “any” refers to. Therefore, the metes and bounds of the claim are not clearly set forth and the examiner cannot clearly determine which elements are encompassed by the claim language, which renders the claim indefinite.
Claim 19 recites the limitation “performing, using an image decoder, decoding processing on the n denoised codes to obtain n target images” in lines 12-13 of the claim. However, it is not clear how the claimed “decoding processing” is performed by “using an image decoder” because the claimed “decoding processing” term is a functional description and the specification does not describe the actual steps or process for performing the claimed “decoding processing”, which renders the claim indefinite.
Claim 20 recites the limitation “determining n noised codes corresponding to n video frames of an original video… determining a text code corresponding to a description text guiding video editing… obtain n denoised codes” in lines 7-10 of the claim. However, the claimed “noised codes” , “text code”, and “denoised codes” terms are not defined by the claims. Additionally, the claimed “n” term recited in claims 1, 3, 5-6, 8, 10-14, 16, and 19-20, respectively, is also not defined by the claims, as indicated above. Specifically, the scope of the claims is unclear without specifying what “n” means because “n” can be interpreted as “a predetermined number,” “any number,” or “a plurality”, for example, and the specification does not define the set of items or conditions that “n” refers to. Therefore, the metes and bounds of the claim are not clearly set forth and the examiner cannot clearly determine which elements are encompassed by the claim language, which renders the claim indefinite.
Claim 20 recites the limitation “performing, using n Unet models obtained by using the text code and copying a Unet model denoising processing on the n noised codes, to obtain n denoised codes, wherein a pre-trained text-to-image model comprises the Unet model, wherein each Unet model comprises a self-attention layer” in lines 9-12 of the claim. However, it is not clear how the claimed “n Unet models” are “obtained by using the text code and copying a Unet model denoising processing on the n noised codes” because it merely recites a use (e.g. “obtained by using“) without reciting any active, positive steps delimiting how this use is actually practiced, and the specification does not describe the actual steps or process for “using n Unet models obtained by using” refers to, which renders the claim indefinite. See MPEP § 2173.05(q).
Claim 20 recites the limitation “a self-attention layer of any ith Unet model, attention calculation based on an output of a target network layer of the ith Unet model and an output of a target network layer in a predetermined target Unet model” in lines 14-16 of the claim. However, the claimed “ith” term recited in claims 1, 4-5 and 19-20, respectively, is also not defined by the claims, as indicated above. Specifically, the scope of the claims is unclear without specifying what “ith” means because “ith” can be interpreted as “a predetermined number,” “any number,” or “a plurality”, for example. Additionally, the claimed “any ith Unet model” limitation is ambiguous or fails to define the scope of the claim with sufficient clarity for a person of ordinary skill in the art to understand and practice the invention and the specification does not define the set of items or conditions that “any” refers to. Therefore, the metes and bounds of the claim are not clearly set forth and the examiner cannot clearly determine which elements are encompassed by the claim language, which renders the claim indefinite.
Claim 20 recites the limitation “performing, using an image decoder, decoding processing on the n denoised codes to obtain n target images” in lines 17-18 of the claim. However, it is not clear how the claimed “decoding processing” is performed by “using an image decoder” because the claimed “decoding processing” term is a functional description and the specification does not describe the actual steps or process for performing the claimed “decoding processing”, which renders the claim indefinite.
Conclusion
The prior art made of record cited in PTO-892 and not relied upon is considered pertinent to applicant’s disclosure. In particular, CN 116681630 A, Applicant cited prior art, appears to disclose an inventive concept similar to applicant’s claimed invention. However, due to the inability to determine a reasonable interpretation of the claims, as indicated above, no prior art rejection or determination of allowability over the prior art is possible during this examination.
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GUILLERMO M RIVERA-MARTINEZ whose telephone number is (571) 272-4979. The examiner can normally be reached on 9 am to 5 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Bee can be reached on 571-270-5183. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/GUILLERMO M RIVERA-MARTINEZ/ Primary Examiner, Art Unit 2677