Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 6-9, and 13-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kopp (US 20240185518 A1), and further in view of Dutta (US 20230296516 A1).
Regarding claim 1, Kopp teaches an image editing apparatus comprising:
an input/output interface configured to obtain a drag input instruction (par. 0462: “In one embodiment, processing logic provides a user interface for altering a dental site. For example, processing logic load the received or generated 3D models of the upper and/or lower dental arches and present the 3D models in the user interface. A user may then select individual teeth or groups of teeth and may move the one or more selected teeth (e.g., by dragging a mouse), may rotate the one or more selected teeth, may change one or more properties of the one or more selected teeth (e.g., changing a size, shape, color, presence of dental conditions such as caries, cracks, wear, stains, etc.), or perform other alterations to the selected one or more teeth.”) and an image (par. 0496: “Accordingly, a user interface may enable a user to update image/frame selection based on manipulating 3D models of dental arches, and may additionally enable a user to manipulate 3D models of dental arches based on selection of images/frames.”); and
a controller configured to obtain an optical flow based on the drag input instruction and the image by using a first artificial intelligence model that is trained to receive a drag input instruction and an image as input and output an optical flow (par. 0323: “In one embodiment a machine learning model is trained to receive current and previously generated labels (for current and previous frames) as well as a previously generated frame and to compute an optical flow between the current post-treatment contours and the previous generated frame. The optical flow may be computed in the feature space in embodiments.”), to input the optical flow and the image to a second artificial intelligence model that is different from the first artificial intelligence model (par. 0240: “Additionally, or alternatively, different machine learning (ML) models may be trained to perform different combinations of the tasks.”), thereby obtaining an edited image as an output of the second artificial intelligence model (par. 0251: “In one embodiment, the generative model has three main stages, including a shared feature extraction stage, a scale-agnostic motion estimation stage, and a fusion stage that outputs a resulting color image.”), and to provide the edited image (par. 0467: “Once the altered image or video is generated, it may be stored, transmitted to a client device (e.g., if method 1900 is performed by a service executing on a server), output to a display, and so on.”),
wherein the controller trains the second artificial intelligence model based on a diffusion model (par. 0249: “In one embodiment, a generative model is used for one or more machine learning models. The generative model may be a generative adversarial network (GAN), encoder/decoder model, diffusion model, variational autoencoder (VAE), neural radiance field (NeRF), or other type of generative model. The generative model may be used, for example, in modified frame generator 336.”), with the second artificial intelligence model being trained to gradually remove noise required to restore an image from random noise based on the image and the optical flow (par. 0252: “In one embodiment, one or more machine learning model is a conditional generative adversarial (cGAN) network, such as pix2pix or vid2vid. These networks not only learn the mapping from input image to output image, but also learn a loss function to train this mapping. GANs are generative models that learn a mapping from random noise vector z to output image y, G:z.fwdarw.y. In contrast, conditional GANs learn a mapping from observed image x and random noise vector z, to y, G:{x, z}.fwdarw.y. The generator G is trained to produce outputs that cannot be distinguished from “real” images by an adversarially trained discriminator, D, which is trained to do as well as possible at detecting the generator's “fakes”.”), and
wherein the controller generates an image by filling a background of a first image, which is one of the two images, with a second image, which is a remaining image (par. 0227: “In embodiments, video processing logic 208 performs a sequence of operations to identify an area of interest in frames of the video, determine replacement content to insert into the area of interest, and generate modified frames that integrate the original frames and the replacement content.”), and trains the second artificial intelligence model by using a training dataset that includes a plurality of samples each constructed by replacing the first image with the generated image (par. 0237: “In embodiments, one or more trained machine learning models of the video processing workflow 305 are trained at a server, and the trained models are provided to a video processing logic 208 on another computing device (e.g., computing device 205 of FIG. 2), which may perform the video processing workflow 305.”; par. 0338: “The training dataset may additionally or alternatively be augmented. Training of large-scale neural networks generally uses tens of thousands of images, which are not easy to acquire in many real-world applications. Data augmentation can be used to artificially increase the effective sample size. Common techniques include random rotation, shifts, shear, flips and so on to existing images to increase the sample size.”).
Kopp fails to teach wherein the controller trains the first artificial intelligence model including a generator and a discriminator of a generative adversarial network (GAN), with the first artificial intelligence model being trained by allowing the generator to generate a fake synthetic optical flow based on the image and a conditional drag input and by allowing the discriminator to perform a process of distinguishing between the fake synthetic optical flow and a genuine optical flow, and
wherein the controller obtains a training dataset including a plurality of samples each including two images, two masks, and optical flows based on random video data, thereby preprocessing the random video data to generate preprocessed video data, and trains the first artificial intelligence model by using the preprocessed video data.
Dutta teaches wherein the controller trains the first artificial intelligence model including a generator and a discriminator of a generative adversarial network (GAN) (par. 0098: “The training context NN section 130 comprises a GAN having Generator (G) 132, Discriminator (D) 134, and Pixel-wise Loss function elements 136.”), with the first artificial intelligence model being trained by allowing the generator to generate a fake synthetic optical flow based on the image and a conditional drag input and by allowing the discriminator to perform a process of distinguishing between the fake synthetic optical flow and a genuine optical flow (par. 0101: “Returning to the training, the training NN Generator (G) 132 learns to generate so-called fake unreduced power images that closely resemble the (collected) unreduced power images 124 of the flow cells 110 and provides them to the training NN Discriminator (D) (conceptually indicated by the arrow labeled Fake 138). The Discriminator (D) 134 learns to distinguish between the fake unreduced power images 138 and the collected unreduced power images 124.”), and
wherein the controller obtains a training dataset including a plurality of samples each including two images, two masks, and optical flows based on random video data (par. 0097: “Some implementations of AI-driven signal enhancement of sequencing images are based on supervised training (e.g., directed to an autoencoder or some variations of a GAN), using so-called paired images. An example of a paired image is a pair of images of a same sample area (such as a same tile of a same flow cell).”), thereby preprocessing the random video data to generate preprocessed video data, and trains the first artificial intelligence model by using the preprocessed video data (par. 0117: “The top row illustrates raw images, the middle row illustrates normalized images, and the bottom row illustrates intensity histograms of the images. In some implementations, normalized images are used and/or produced by one or more preprocessing operations for use by subsequent neural network processing.”).
It would have been obvious to one familiar in the art prior to the effective filing date of the claimed invention to incorporate the training methods of Dutta into the video generation method of Kopp, as both are in the same field of endeavor of generative AI training. Such techniques are well-known in the art and necessary to properly implementing the video generation method of Kopp.
Regarding claim 6, Kopp and Dutta teach the image editing apparatus of claim 1. Kopp further teaches wherein the controller performs fixed-size normalization on the optical flow from the first artificial intelligence model to generate a normalized optical flow and inputs the normalized optical flow to the second artificial intelligence model (par. 0268: “The segmentation model may return for each pixel a score distribution of multiple classes that can be normalized and interpreted as a probability distribution. In one embodiment, an operation that finds the argument that gives the maximum value from a target function (e.g., argmax) is performed on the class distribution to assign a single class to each pixel. If two classes have a similar score at a certain pixel, small image changes can lead to changes in pixel assignment.”).
Regarding claim 7, Kopp and Dutta teach the image editing apparatus of claim 1. Kopp further teaches wherein the controller generates a random drag input instruction based on a sparse flow fs∈R2×h×w initialized with random values sampled from U(0, 1), and trains the first artificial intelligence model by using the generated random drag input instruction (par. 0252: “In one embodiment, one or more machine learning model is a conditional generative adversarial (cGAN) network, such as pix2pix or vid2vid. These networks not only learn the mapping from input image to output image, but also learn a loss function to train this mapping. GANs are generative models that learn a mapping from random noise vector z to output image y, G:z.fwdarw.y. In contrast, conditional GANs learn a mapping from observed image x and random noise vector z, to y, G:{x, z}.fwdarw.y.”).
Regarding claim 8, Kopp and Dutta teach the image editing apparatus of claim 1. Kopp further teaches wherein the controller performs sample-wise normalization on the optical flow when training the first artificial intelligence model, and performs fixed-size normalization on the optical flow when training the second artificial intelligence model (par. 0268: “The segmentation model may return for each pixel a score distribution of multiple classes that can be normalized and interpreted as a probability distribution.”; par. 0269: “Also input into segmenter 318 are one or more optical flows, including a first optical flow 608 between the cropped mouth area of previous frame 602 and the cropped mouth area of current frame 606 and/or a second optical flow 610 between the cropped mouth area of previous frame 604 and the cropped mouth area of current frame 606.”).
Claim 9 is substantially similar to claim 1, except that it teaches a method as opposed to an apparatus. As such, it is rejected on a similar basis to claim 1.
Claim 13 is substantially similar to claim 6, except that it depends from claim 9 as opposed to claim 1. As such, it is rejected on a similar basis to claim 6.
Claim 14 is substantially similar to claim 7, except that it depends from claim 9 as opposed to claim 1. As such, it is rejected on a similar basis to claim 7.
Claim 15 is substantially similar to claim 8, except that it depends from claim 9 as opposed to claim 1. As such, it is rejected on a similar basis to claim 8.
Response to Arguments
Applicant’s arguments with respect to claim(s) 1, 6-9, and 13-15 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RYAN A BARHAM whose telephone number is (571)272-4338. The examiner can normally be reached Mon-Fri, 8:30am-5pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu, can be reached at (571) 272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RYAN ALLEN BARHAM/Examiner, Art Unit 2613
/XIAO M WU/Supervisory Patent Examiner, Art Unit 2613