Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of Applicant’s claim of the present application claiming priority and benefit under 35 U.S.C. 119(e) to US Provisional Application No. 63/679,214 filed 08/05/2024.
Information Disclosure Statement
The information disclosure statement (“IDS”) filed 12/29/2025 has been reviewed and the listed references were noted.
Drawings
The 5-page drawings have been considered and placed in the file.
Status of Claims
Claims 1-20 are pending.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-4, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Liang et al. (“PIE: Simulating Disease Progression Via Progressive Image Editing”), in view of Chen et al. (“Two-Stage Video-Based Convolutional Neural Networks for Adult Spinal Deformity Classification”).
Regarding claim 1, Liang teaches, “A method of enhancing dataset for use in a medical diagnostic system, comprising the step of: receiving a static medical image capturing a diagnostic target;” (Liang, Pg. 4, Section 4 discloses; “The inputs to PIE are a discrete medical image x(0) depicting any start or middle stage of a disease”) “and generating, based on the received static medical image, a series of video frames arranged to combine to a dynamic video representing a clinical motion of the diagnostic target over a predetermined period of time;” (Liang, Pg. 4, Section 4 discloses; “The output generated is a sequence of images, {x(0) ,x(1) ,...,x(N) },illustrating the progression of the disease as per the input report.”) “(Chen, Pg. 5 discloses; “In this study, we trained a ResNet-style 3D CNN using videos to recognize a patient's gait posture.” Examiner submits that it is obvious to train medical diagnostic systems with medical video datasets. Thus, it would be obvious to use the generated video of Liang in a medical training dataset.)
Liang and Chen are considered to be analogous to the claimed invention because they are in the same field of image analysis of medical images/videos. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Liang to incorporate the teachings of Chen in order to train a medical diagnostic system on a synthetic medical video dataset. One of ordinary skill in the art would have been motivated to combine the previously described method of Liang with the teachings of Chen to have a quick way to generate medical videos for training of a system. Accordingly, it would have been obvious to combine Liang and Chen to obtain claim 1.
Regarding claim 2, the combination of Liang and Chen teaches, “The method of Claim 1, wherein the series of video frames are generated by an AI-based video generator.” (Liang, Pg. 5, “Algorithm 1” discloses; “Original input image x(0) 0 at the start point, input image x(n−1) 0 at stage n, number of diffusion steps T, text conditional vector y, noise strength γ, stable diffusion parameterized denoiser” Liang uses a diffusion model, which is interpreted as AI-based.)
Regarding claim 3, the combination of Liang and Chen teaches, “The method of Claim 2, wherein the series of video frames are generated based on augmentation of the static medical image.” (Liang, Abstract discloses; “To address this issue, we develop a novel framework termed Progressive Image Editing (PIE) that enables controlled manipulation of disease-related image features, facilitating precise and realistic disease progression simulation in imaging space.”)
Regarding claim 4, the combination of Liang and Chen teaches, “The method of Claim 3, wherein the step of generating the series of video frames comprises the step of generating N video frames using a Stable Video Diffusion process.” (Liang, Pg. 4, Section 4 discloses; “The output generated is a sequence of images, {x(0) ,x(1) ,...,x(N) },illustrating the progression of the disease as per the input report.” and Liang, Pg. 6, Section 5.1 discloses; “Specifically, we evaluate the pretrained domain-specific stable diffusion model on three different types of disease data sets in classification tasks”).
Regarding claim 19, the combination of Liang and Chen teaches, “A method of synthesizing video for medical diagnosis, comprising the step of: providing a static medical image capturing a diagnostic target;” (Liang, Pg. 4, Section 4 discloses; “The inputs to PIE are a discrete medical image x(0) depicting any start or middle stage of a disease”) “and generating a series of video frames using the method in accordance with claim 1.” (Liang, Pg. 4, Section 4 discloses; “The output generated is a sequence of images, {x(0) ,x(1) ,...,x(N) },illustrating the progression of the disease as per the input report.”)
Claims 5 and 6 are rejected under 35 U.S.C. 103 as being unpatentable over Liang et al., in view of Chen et al., in further view of Shi et al. (“Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling”).
Regarding claim 5, the combination of Liang and Chen does not explicitly teach, “The method of Claim 4, wherein stable video diffusion process is formulated with a Markov chain, arranged to generate video data from noise in the static medical image via a T-step denoising process.” Since the combination of Liang and Chen does not explicitly disclose these limitations, Examiner relies on the teachings of Shi in an analogous field of endeavor. Specifically, Shi teaches, “The method of Claim 4, wherein stable video diffusion process is formulated with a Markov chain, arranged to generate video data from noise in the static medical image via a T-step denoising process.” (Shi, Pg. 4 discloses; “
PNG
media_image1.png
258
344
media_image1.png
Greyscale
” Stable diffusion Markov chains and T-step denoising processes are known in the art. Thus, it would be obvious to use this method with the static images of Liang to generate video data.)
Liang, Chen, and Shi are considered to be analogous to the claimed invention because they are in the same field of analysis of images/videos. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Liang and Chen to incorporate the teachings of Shi in order to use a T-step denoising process to generate video data. One of ordinary skill in the art would have been motivated to combine the previously described method of Liang and Chen with the teachings of Shi to have the system learn latent space distribution. Accordingly, it would have been obvious to combine Liang, Chen, and Shi to obtain claim 5.
Regarding claim 6, the combination of Liang, Chen, and Shi teaches, “The method of Claim 5, wherein a plurality of static medical images are provided as sample images each captures the respective diagnostic target” (Liang, Pg. 4, Section 4 discloses; “The inputs to PIE are a discrete medical image x(0) depicting any start or middle stage of a disease”) “and wherein the sample images are processed by the stable video diffusion process” (Liang, Pg. 6, Section 5.1 discloses; “Specifically, we evaluate the pretrained domain-specific stable diffusion model on three different types of disease data sets in classification tasks”) “to obtained a set of synthesized videos” (Liang, Pg. 4, Section 4 discloses; “The output generated is a sequence of images, {x(0) ,x(1) ,...,x(N) },illustrating the progression of the disease as per the input report.”) “wherein each of the synthesized video comprises the N video frames generated by the each of the sample images being augmented.” (Liang, Pg. 4, Section 4 discloses; “The output generated is a sequence of images, {x(0) ,x(1) ,...,x(N) },illustrating the progression of the disease as per the input report.”).
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Liang et al., in view of Chen et al., in further view of Shi et al., and still in view of Madani et al. (US 20190197368 A1).
Regarding claim 7, the combination of Liang, Chen, and Shi does not explicitly teach, “The method of Claim 6, wherein the sample images including labelled and unlabeled medical images capturing a respective diagnostic target.” Since the combination of Liang, Chen, and Shi does not explicitly disclose these limitations, Examiner relies on the teachings of Madani in an analogous field of endeavor. Specifically, Madani teaches, “The method of Claim 6, wherein the sample images including labelled and unlabeled medical images capturing a respective diagnostic target.” (Madani, Para. [0005] discloses; “The method further comprises generating, by a generator of the GAN, one or more generated medical images and inputting, to the discriminator of the GAN, a training medical image set comprising a first subset of labeled medical images, a second subset of unlabeled medical images, and a third subset comprising the one or more generated medical images.”)
Liang, Chen, Shi, and Madani are considered to be analogous to the claimed invention because they are in the same field of analysis of images/videos. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Liang, Chen, and Shi to incorporate the teachings of Madani in order to use labeled and unlabeled training data. One of ordinary skill in the art would have been motivated to combine the previously described method of Liang, Chen, and Shi with the teachings of Madani to increase system accuracy. Accordingly, it would have been obvious to combine Liang, Chen, Shi, and Madani to obtain claim 7.
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Liang et al., in view of Chen et al., in further view of Ren et al. (“ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation”).
Regarding claim 8, the combination of Liang and Chen does not explicitly teach, “The method of Claim 1, wherein the clinical motion includes at least one of spatial translation, liquid flow and shake blur.” Since the combination of Liang and Chen does not explicitly disclose these limitations, Examiner relies on the teachings of Ren in an analogous field of endeavor. Specifically, Ren teaches, “The method of Claim 1, wherein the clinical motion includes at least one of spatial translation, liquid flow and shake blur.” (Ren, Figure 5 shows an input frame, and a video being generated representing spatial translation.
PNG
media_image2.png
399
498
media_image2.png
Greyscale
)
Liang, Chen, and Ren are considered to be analogous to the claimed invention because they are in the same field of analysis of images/videos. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Liang and Chen to incorporate the teachings of Ren in order to ensure the generated video includes spatial translation. One of ordinary skill in the art would have been motivated to combine the previously described method of Liang and Chen with the teachings of Ren to ensure the imaged object translates and progresses in the generated video. Accordingly, it would have been obvious to combine Liang, Chen, and Ren to obtain claim 8.
Claims 9, 10, and 14-18 are rejected under 35 U.S.C. 103 as being unpatentable over Liang et al., in view of Chen et al., in further view of Wang et al. (“Dancing with Still Images: Video Distillation via Static-Dynamic Disentanglement”).
Regarding claim 9, the combination of Liang and Chen teaches, “The method of Claim 1, further comprising the step of generating, (Liang, Pg. 18 discloses; “To fine-tune the Stable Diffusion model, we center-crop and resize the input images to 512 × 512 resolution”). The combination of Liang and Chen does not explicitly teach, “based on the dynamic video being generated, a series of reversed-generated images embedding inherent motion information associated with the diagnostic target over the predetermined period of time”. Since the combination of Liang and Chen does not explicitly disclose these limitations, Examiner relies on the teachings of Wang in an analogous field of endeavor. Specifically, Wang teaches, “based on the dynamic video being generated, a series of reversed-generated images embedding inherent motion information associated with the diagnostic target over the predetermined period of time” (Figure 1(c) of Wang shows their proposal generating a static image from a video for motion embedding. Wang, Abstract discloses; “It first distills the videos into still images as static memory and then compensates the dynamic and motion information with a learnable dynamic memory block.” It would be obvious to use the generated image of Liang with the generating of an image from a video of Wang.)
Liang, Chen, and Wang are considered to be analogous to the claimed invention because they are in the same field of analysis of images/videos. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Liang and Chen to incorporate the teachings of Wang in order to embed motion information from images generated from the generated video. One of ordinary skill in the art would have been motivated to combine the previously described method of Liang and Chen with the teachings of Wang to ensure the system accounts for motion during the video. Accordingly, it would have been obvious to combine Liang, Chen, and Wang to obtain claim 9.
Regarding claim 10, the combination of Liang, Chen, and Wang teaches, “The method of Claim 9, further comprising the step of processing the series of reversed- generated images and the dynamic video using a video-to-image distillation process to distill motion-aware cue information from the dynamic video.” (Wang, Pg. 5, Section 4 discloses; “Thus, based on the analysis in Sec.3.3 and considering the tradeoff between efficiency and efficacy, we propose a video dataset distillation paradigm by disentangling the static and dynamic information in videos.”) The proposed combination as well as the motivation for combining the Liang, Chen, and Wang references presented in the rejection of claim 9, apply to claim 10 and are incorporated herein by reference. Thus, the method recited in claim 10 is met by Liang, Chen, and Wang.
Regarding claim 14, the combination of Liang, Chen, and Wang teaches, “A method for training a medical diagnostic system in accordance with claim 9, comprising the step of training a classifier with the medical dataset comprising the dynamic video and/or the series of reversed-generated images.” (Chen, Pg. 5 discloses; “In this study, we trained a ResNet-style 3D CNN using videos to recognize a patient's gait posture.” Examiner submits that it is obvious to train medical diagnostic systems with medical video datasets. Thus, it would be obvious to use the generated video of Liang in a medical training dataset.) The proposed combination as well as the motivation for combining the Liang, Chen, and Wang references presented in the rejection of claim 9, apply to claim 14 and are incorporated herein by reference. Thus, the method recited in claim 14 is met by Liang, Chen, and Wang.
Regarding claim 15, the combination of Liang, Chen, and Wang teaches, “The method of Claim 14, wherein the dynamic videos are labelled.” (Chen, Pg. 7, Section 4.1 discloses; “The gait posture’s label attached to each video was based on the spine surgeon’s diagnosis from diagnostic radiographical assessment using standing whole-spine X-ray images and the clinical symptoms of the patients.”) The proposed combination as well as the motivation for combining the Liang, Chen, and Wang references presented in the rejection of claim 9, apply to claim 15 and are incorporated herein by reference. Thus, the method recited in claim 15 is met by Liang, Chen, and Wang.
Regarding claim 16, the combination of Liang, Chen, and Wang teaches, “The method of Claim 14, further comprising the step of training an image encoder arranged to generated the series of reversed-generated image embeddings based on the dynamic video.” (Figure 1(c) of Wang shows their proposal generating a static image from a video for motion embedding. Wang, Abstract discloses; “It first distills the videos into still images as static memory and then compensates the dynamic and motion information with a learnable dynamic memory block.” It is implied that the image encoder of Wang is trained to generate these images, since the system can perform this task.) The proposed combination as well as the motivation for combining the Liang, Chen, and Wang references presented in the rejection of claim 9, apply to claim 16 and are incorporated herein by reference. Thus, the method recited in claim 16 is met by Liang, Chen, and Wang.
Regarding claim 17, the combination of Liang, Chen, and Wang teaches, “The method of Claim 16, wherein the classifier and/or the image encoder is a machine learning network.” (Wang, Pg. 6, Section 5.3 discloses; “We use DC [31] to distill static memory with random real frame initialization. We utilize a 4-layer 2D convolutional neural network for distillation (ConvNetD4) and perform an early stop in the distillation training when the loss converges.”) The proposed combination as well as the motivation for combining the Liang, Chen, and Wang references presented in the rejection of claim 9, apply to claim 17 and are incorporated herein by reference. Thus, the method recited in claim 17 is met by Liang, Chen, and Wang.
Regarding claim 18, Examiner does not need to analyze the claim because it depends on claims 16 and 14. Claim 14 states “dataset comprising the dynamic video and/or the series of reverse-generated images”. Claim 18 is directed towards the unelected limitation of claim 14. Thus, claim 18 is still rejected.
Claims 12 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Liang et al., in view of Chen et al., in further view of Wang et al., and still in view of Zhai et al. (“Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation”).
Regarding claim 12, the combination of Liang, Chen, and Wang does not explicitly teach, “The method of claim 10, further comprising the step of enhancing cross-image consistency within imaging modality of the series of reversed-generated images.” Since the combination of Liang, Chen, and Wang does not explicitly disclose these limitations, Examiner relies on the teachings of Zhai in an analogous field of endeavor. Specifically, Zhai teaches, “The method of claim 10, further comprising the step of enhancing cross-image consistency within imaging modality of the series of reversed-generated images.” (Zhai, Abstract discloses; “Specifically, MCM includes a video consistency model that distills motion from the video teacher model, and an image discriminator that enhances frame appearance to match high-quality image data.”)
Liang, Chen, Wang, and Zhai are considered to be analogous to the claimed invention because they are in the same field of analysis of images/videos. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Liang, Chen, and Wang to incorporate the teachings of Zhai in order to enhance cross-image consistency. One of ordinary skill in the art would have been motivated to combine the previously described method of Liang, Chen, and Wang with the teachings of Zhai to ensure the distillation is accurate. Accordingly, it would have been obvious to combine Liang, Chen, Wang, and Zhai to obtain claim 12.
Regarding claim 13, the combination of Liang, Chen, Wang, and Zhai teaches, “The method of Claim 12, wherein a plurality pairs of reversed-generated images in the series of reversed-generated images associated with each video frame pair in the dynamic video are enhanced via consistency loss.” (Zhai, Pg. 9, Section 4.3 discloses; “Starting from LCM [38], we progressively add adversarial learning, motion consistency distillation loss, and mixed trajectory distillation. Each addition shows an improvement in video quality, demonstrating the effectiveness of these components.”) The proposed combination as well as the motivation for combining the Liang, Chen, Wang, and Zhai references presented in the rejection of claim 12, apply to claim 13 and are incorporated herein by reference. Thus, the method recited in claim 13 is met by Liang, Chen, Wang, and Zhai.
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Liang et al., in view of Chen et al., in further view of Tian et al. (US 20250301154 A1 w/ EFD of 5/10/2023)
Regarding claim 20, the combination of Liang and Chen does not explicitly teach, “The method of Claim 19, further comprising the step of generating a dynamic video with the series of video frame using a frozen video encoder.” Since the combination of Liang and Chen does not explicitly disclose this limitation, Examiner relies on the teachings of Tian in an analogous field of endeavor. Specifically, Tian teaches, “The method of Claim 19, further comprising the step of generating a dynamic video with the series of video frame using a frozen video encoder.” (Tian, Abstract discloses; “A method includes: extracting a video frame sequence from a sample video, the video frame sequence including a key frame and an estimated frame; performing encoding and decoding processing on the key frame via a pre-trained key frame network of a video encoding and decoding model” Frozen is interpreted to mean a pre-trained model.)
Liang, Chen, and Tian are considered to be analogous to the claimed invention because they are in the same field of analysis of images/videos. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Liang and Chen to incorporate the teachings of Tian in order to freeze the video encoder while generating a video. One of ordinary skill in the art would have been motivated to combine the previously described method of Liang and Chen with the teachings of Tian to stop the further training of the video encoder. Accordingly, it would have been obvious to combine Liang, Chen, and Tian to obtain claim 20.
Allowable Subject Matter
Claim 11 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JUSTIN M. OAKES whose telephone number is (571)272-9379. The examiner can normally be reached 7:30am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached at (571) 272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JUSTIN M OAKES/Examiner, Art Unit 2662
/Siamak Harandi/Primary Examiner, Art Unit 2662