DETAILED ACTION
Response to Amendment
Applicant’s response to the last Office Action, filed on 7/21/2026 has been entered and made of record.
Examiner maintains the current prior art of record; accordingly, this action is made Final.
New rejections under 35 USC 112(a) and 112(b) appear below.
Response to Arguments
Applicant's arguments filed on 7/21/2026 have been fully considered but they are not persuasive.
Regarding the rejection under 35 USC 112(a) directed to the ‘ground truth’ class as output from the recognition network, Applicant noted “the Applicant has amended claim 1 to remove the "ground truth" phrase” (it appears Applicant intended to refer to the first independent claim 15 rather than claim 1). However, Examiner notes that claim 15 has not been amended as such. The claim continues to recite, “an image recognition network which takes an image containing an object as input and predicts one or more classes of the object, each predicted class being a ground truth class.”
Regarding the rejection under 35 USC 112(a) directed to enablement and/or written description Applicant remarked, “A person of ordinary skill in the art of machine learning understands how to implement standard DNN or GAN architectures once the specific functional arrangement and feedback loops are disclosed. Requiring a layer-by-layer architectural breakdown for known generic models is contrary to the Federal Circuit's recent guidance in Recentive. The Recentive decision reinforces the idea that standard machine learning architectures, such as the DNNs and GAN of claim 1 are well-understood "generic" components in the eyes of one or skill in the art . . . The case supports the argument that requiring a granular, layer-by-layer description of a standard neural network is equivalent to requiring a patent applicant to describe the internal circuitry of a standard microprocessor, as argued in the previous response. Because models like GANs and classifiers (recognition networks) are now considered "off-the-shelf" algorithmic tools, the Recentive case suggests that their internal mathematical structure does not need to be reinvented or exhaustively detailed in every filing.”
Examiner disagrees that the analogy is appropriate here and finds that the situation in the instant 112(a) rejection is not at all akin to requiring Applicant to “describe the internal circuitry of a standard microprocessor” in the context of a computer software-based invention or the like. A more appropriate analogy would be a case where an uncommon microprocessor architecture is directly claimed and Examiner requires Applicant to disclose how that particular uncommon microprocessor architecture would actually operate, even at a high level.
It is important to note that Applicant is claiming a particular architecture of integrating a label-conditional GAN variation (not whatsoever a generic GAN block) with multiple other neural network blocks without so much as explaining how these multiple non-generic blocks might interact with one another. Applicant fails to mention, for example, that the discriminator block is not an off-the-shelf GAN block (which would not contain any label input at all here) but instead appears to be some type of label-conditional variation. Unfortunately, a basic explanation of what this is or how it may function, even at a high level, is not present in the disclosure.
Thus, the question at issue is not whether GANs exist and can be considered generic modules but whether the description requirement is met for the claimed combination of networks. For examples, it appears the discriminator conditionally “punishes” its multiple individual subnetworks in different circumstances but how this might be accomplished (beyond vague statements that gradients can be backpropagated to layers that are themselves unmentioned) is simply not described with enough detail to be legible. All the reader knows is that “"fake" output of discriminator 202 will result in image generation 104 and/or recognition network 102 being punished.” This “and/or” disclosure does not even reveal if one or both of the subnetworks is punished, and if only one, which one would be punished in what circumstance. Again, these details of how this label-conditional GAN variation might work are not found in generic off-the-shelf components well-known in the art.
Applicant remarked that Gauthier and Sharma are both cited for separate teachings of adversarial networks including a generative network and a discriminator network. The Applicant submits that neither Gauthier nor Sharma teaches a discriminator network capable of determining if a generated image is "photorealistic". Instead, Applicant argues, the discriminators of both cited references discriminate between generated images and images that were part of the training dataset. Applicant argues that Gauthier thus teaches a discriminator that determines if an input image is generated or is part of the training dataset. Applicant argues that Sharma teaches a discriminator that determines if an input image is a captured image or a generated image (or if the input image even contains a face). Applicant argues that nether discriminator taught by the cited art determines if a generated image is photorealistic.”
First Examiner notes that the prior Office Action maps the limitations in question to the primary reference Gauthier, which should make some arguments related to Sharma moot. See Gauthier pg. 2, right column, second paragraph from bottom teaches the function of the discriminator network “which accepts an image x and condition y [predicted class] and predicts the probability under condition y that x came from the empirical data distribution rather than from the generative model” (i.e., that is found to be photorealistic/indistinguishable from the original photo in the empirical dataset). Also see another explicit explanation of the generator producing a photorealistic image at pg. 3, right column, ¶ 2, “This process [improving the generator by punishing it when the discriminator can tell it is not a photo] continues ad infinitum, until the discriminator is maximally confused. Since the discriminator outputs the probability that an input image was sampled from the training data, we would expect a “maximally confused” discriminator to consistently output a probability of 0.5 for inputs both from the training data and from the generator.”
Additionally, Examiner notes that contrary to Applicant’s remarks Sharma also provides a clear teaching for the limitation. The discriminator also punishes the GAN generator/image generation network as shown in Fig. 4 and ¶ 0032 when the GAN discriminator, via the modified noise parameter, punishes the GAN generator when the determining that the generated image is a generated image and not a photo image from the database of training images, directly testing whether the generated image is photorealistic. Examiner notes that the term in the claim need not appear word-for-word in the cited reference.
Applicant argues that the dependent claims are allowable by virtue of their dependence on allowable claims, but directs no independent arguments to these claims. Please see detailed response to arguments above.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 15-24 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Claim 15 currently requires “an image recognition network which takes an image containing an object as input and predicts one or more classes of the object, each predicted class being a ground truth class representing an attribute of the object depicted in the image”, however the only support in the original disclosure is for the ground truth class being input to the image generation network. See ¶ 0010 of the Specification, as filed, as well as the originally presented claim 15. The image generation network takes real/ground truth classes as input and not the class predicted from the image recognition network. Examiner also notes that this is consistent with the term ‘ground truth’ in the art which refers to the real class in contrast with the predicted class. Regardless, there is no disclosure for obtaining a ground truth class from the image recognition network. Appropriate correction is required.
Claims 15-24 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Examiner notes that ¶ 10 and 15 of the specification disclose that the image generation network takes real/ground truth classes as input and not a ‘predicated’ class from the image recognition network. Claim 15 recites, “an image generation network which takes the one or more predicted classes as input and generates an image”.
Claims 15-24 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
Claims 15 requires an image recognition network to predict one or more classes of an image, an image generation network that generates an image fitted to certain classes, and a discriminator network which takes as input the generated image and the one or more predicted classes and returns a result indicating whether the classification of the generated image conforms to the one or more predicted classes and is photorealistic, wherein the discriminator network punishes the image generation network if the generated image is not photorealistic and punishes the image recognition network if the classification of the generated image does not conform to the one or more predicted classes.
However, Examiner notes that the disclosure as filed contains essentially no detail on how any of these individual steps might be accomplished, beyond the very high-level description of what their intended function is. For example, no diagram or description of the network architectures is provided whatsoever relating to the image recognition network, image generation network and the discriminator network. As an example, the Specification as filed contains only eight paragraphs in total devoted to the detailed description of the entire system and key details are missing, such as any basic outline for how these networks might actually work. Many systems in the art use convolutional neural nets with many layers to accomplish similar functions but the disclosure contains no such mention of convolutional networks and only a passing reference to “Deep neural network” at ¶ 0010 and vague explanation of training at ¶ 0014, “the punishment may be in the form of gradient to be backpropagated to the various layers of the respective networks” though no description of any network’s layers has been provided. The only drawings are flow diagrams of using the networks and provide no detail on their actual individual function.
Without any clear explanation of the details of the system, it would not be at all clear to ordinary skill in the art whether the inventor has been able to use these network components to accomplish the stated function of generating a photorealistic image according to certain input classes by punishing the various networks described, or how this would be done. As such, Examiner finds that the disclosure is not adequate to reasonably convey to one skilled in the relevant art that the inventor had possession of the claimed invention. The disclosure has not clearly conveyed that the Applicant has invented the required claimed subject matter. In re Barker, 559 F.2d 588, 592 n.4, 194 USPQ 470, 473 n.4 (CCPA 1977).
To satisfy the written description requirement, a patent specification must describe the claimed invention in sufficient detail that one skilled in the art can reasonably conclude that the inventor had possession of the claimed invention. See, e.g., Moba, B.V. v. Diamond Automation, Inc., 325 F.3d 1306, 1319, 66 USPQ2d 1429, 1438 (Fed. Cir. 2003); MPEP 2163 (I)(A) states that “issues of adequate written description may arise even for original claims, for example, when an aspect of the claimed invention has not been described with sufficient particularity such that one skilled in the art would recognize that the inventor had possession of the claimed invention at the time of filing. The claimed invention as a whole may not be adequately described if the claims require an essential or critical feature which is not adequately described in the specification and which is not conventional or known in the art.”
Claims 15-24 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the enablement requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to enable one skilled in the art to which it pertains, or with which it is most nearly connected, to make and/or use the invention.
Examiner notes that the language of claim 15, was not described in the specification in such a way as to enable one skilled in the art to make and/or use the invention. The enablement requirement of 35 U.S.C. 112(a) or pre-AIA 35 U.S.C. 112, first paragraph, is separate and distinct from the written description requirement. Vas-Cath,Inc. v. Mahurkar, 935 F.2d 1555, 1563, 19 USPQ2d 1111, 1116-17 (Fed. Cir. 1991)
As above, claim 15 requires an image recognition network to predict one or more classes of an image, an image generation network that generates an image fitted to certain classes, and a discriminator network which takes as input the generated image and the one or more predicted classes and returns a result indicating whether the classification of the generated image conforms to the one or more predicted classes and is photorealistic, wherein the discriminator network punishes the image generation network if the generated image is not photorealistic and punishes the image recognition network if the classification of the generated image does not conform to the one or more predicted classes.
Without repeating the entire analysis and claim construction above, Examiner notes that the disclosure as filed contains essentially no detail on how any of these individual steps might be accomplished, beyond the very high-level description of what their intended function is. For example, no diagram or description of the network architectures is provided whatsoever relating to the image recognition network, image generation network and the discriminator network. As such, the only conclusion is that language was not described in the specification in such a way as to enable one skilled in the art to make and/or use the invention. Nowhere does the disclosure explain or suggest how or what the process would be to perform the claimed function.
Examiner notes that as per In re Wands, 858 F.2d at 737, 8 USPQ2d at 1404. the determination that “undue experimentation” would have been needed to make and use the claimed invention is not a single, simple factual determination. Rather, it is a conclusion reached by weighing the following noted factual considerations. (A) The breadth of the claims; (B) The nature of the invention; (C) The state of the prior art; (D) The level of one of ordinary skill; (E) The level of predictability in the art; (F) The amount of direction provided by the inventor; (G) The existence of working examples; (H) The quantity of experimentation needed to make or use the invention based on the content of the disclosure. As per MPEP 2164.04 it is not necessary to discuss each factor in the enablement rejection. Here in particular, amount of direction provided by the inventor (factor F) is found to be lacking. The level of one of ordinary skill and predictability of the art was noted as someone familiar with generative adversarial learning. The quantity of experimentation (factor H) was assessed and noted that experimentation would be very high as the claim requirement describes something which the description contains no genuine explanation for.
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 15-24 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as failing to set forth the subject matter which the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the applicant regards as the invention. Claim 15 recites, “an image generation network which takes the one or more predicted classes as input and generates an image exhibiting features specified by the one or more classes, wherein the one or more classes are ground truth classes of the object depicted in the image”. It is not clear what the antecedent basis support is for the underlined language above, ground truth classes or predicted classes. Examiner notes that it also appears Applicant may be attempting to assert that the predicted classes and ground truth classes are one and the same, or possibly that the set of predicted classes and the set of ground truth classes are the same set of classes. Regardless this has not been claimed in such a way as to define a definite claim scope.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: an image recognition network, image generation network, discriminator network in claims 15-23.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 15-24 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gauthier ("Conditional generative adversarial nets for convolutional face generation"; cited by Examiner in the parent application and provided by Applicant here) in view of Sharma (US PGPub 2019/0332850).
Regarding claim 15, Gauthier discloses a method comprising:
an image generation network which takes one or more classes as input and generates an image exhibiting features specified exhibiting features specified by the one or more classes, wherein the one or more classes are ground truth classes of the object depicted in the image; (Gautier teaches a generative adversarial network which includes an image generation network and a discriminator network. The image generator network takes a class as input and generates an image fitted to the one or more classes. Pg. 2, right column, ¶ 2, “The contribution of this paper is to add a conditioning ability to this framework. We can establish some arbitrary condition y for generation, which restricts both the generator in its output and the discriminator in its expected input.” The condition is a label such as a class. Also see Fig. 1 inputs y and z and pg. 6, right column, “4.3. Conditional GAN” which teaches using face classes as input to generate face images fitted to the class.)
and a discriminator network which takes as input the generated image and the one or more predicted classes and returns a result indicating whether the classification of the generated image conforms to the one or more predicted classes, and whether the generated image is photorealistic. (Pg. 2, right column, second paragraph from bottom teaches the function of the discriminator network “which accepts an image x and condition y [predicted class] and predicts the probability under condition y that x came from the empirical data distribution rather than from the generative model” (i.e., that is found to be photorealistic/indistinguishable from the original photo in the empirical dataset). As above, Pg. 2, right column, ¶ 2, “The contribution of this paper is to add a conditioning ability to this framework. We can establish some arbitrary condition y for generation, which restricts both the generator in its output and the discriminator in its expected input.” The condition is a label such as a class. Also see Fig. 1 inputs y which conditions the discriminator's other input image I based on the class of label y. Also see, pg. 6, right column, “4.3. Conditional GAN” which teaches using face classes as input to discriminate face images of the class.)
wherein the discriminator network punishes the image generation network if the generated image is not photorealistic (As above, See Gauthier pg. 2, right column, second paragraph from bottom teaches the function of the discriminator network “which accepts an image x and condition y [predicted class] and predicts the probability under condition y that x came from the empirical data distribution rather than from the generative model.” See pg. 3, left column, ¶ 1 which teaches Equation (2) in which G (generator) and D (discriminator) play a minimax game to minimize each loss. In this way the discriminator is used to punish/provide loss in training the generator based whether the discriminator can distinguish real facial features from “trick” images from the generator.)
In the field of image recognition Sharma teaches a system with an image recognition network which takes an image containing an object as input and predicts one or more classes of the object, each predicted class being a ground truth class representing an attribute of the object depicted in the image (Sharma teaches a system for training a generative adversarial network (GAN) for use in facial recognition and for training a facial recognition system via a GAN, see Abstract. Facial recognition to predict a face class is taught at ¶ 0020 and 0031. The predicted face classes are attributes of the face object. ¶ 0017 teaches a processor system.)
wherein the discriminator network punishes the image generation network if the generated image is not photorealistic and punishes the image recognition network if the classification of the generated image does not conform to the one or more predicted classes. (Sharma ¶ 0025-0028 teaches the GAN discriminator and generator training the facial image recognition network on the basis of generated and discriminated fake images in which generated negative face images are used to train the recognition network on what is not the particular face input to the GAN. The recognizer learns to discriminate the negative images from the images of the particular face. This way the system punishes the facial image recognition network by training until it recognizes the accurate classification of the generated image as the particular face. The discriminator also punishes the GAN generator/image generation network as shown in Fig. 4 and ¶ 0032 when the GAN discriminator, via the modified noise parameter, punishes the GAN generator when the determining that the generated image is a generated image and not a photo image from the database of training images.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined Gauthier's neural network-based image generation with Sharma’s neural network-based image recognition. Gauthier teaches a method of using a conditional generative adversarial network in order to accomplish conditional image generation according to some class label (e.g., generating a face according to a face type). Gauthier relies on a dataset of labeled images as input in order to input predicted class images into the generator and discriminator for training. Instead of using a pre-existing dataset, any real image can be used as long as any widely-available image recognition algorithm is used to predict the image class prior to inputting into the generator and discriminator steps. Sharma teaches such a recognition algorithm and teaches using it in the context of a facial GAN. Sharma also teaches training the facial recognizer via the system. The combination constitutes the repeatable and predictable result of simply applying image recognition here for the purpose of allowing Gauthier to function without a pre-existing dataset of labelled images. This cannot be considered a non-obvious improvement in view of the relevant prior art here. Using known engineering design, no “fundamental” operating principle of the teachings are changed; they continue to perform the same functions as originally taught prior to being combined.
Regarding claim 16, the above combination discloses the system of claim 15 wherein the discriminator network returns a real result if the generated image input is real and the predicted class input is real. (Gauthier pg. 2, right column, second paragraph from bottom teaches the function of the discriminator network “which accepts an image x and condition y [predicted class] and predicts the probability under condition y that x came from the empirical data distribution rather than from the generative model.”)
Regarding claim 17, the above combination discloses the system of claim 15 wherein the discriminator network returns a fake result if the generated input image is real and the predicted class input is fake. (As above, See Gauthier pg. 2, right column, second paragraph from bottom teaches the function of the discriminator network “which accepts an image x and condition y [predicted class] and predicts the probability under condition y that x came from the empirical data distribution rather than from the generative model.” Also see pg. right column, ¶ 1, “The discriminator becomes more attuned to real facial features in order to distinguish between the simple “trick” images from the generator and real face images. Furthermore, the discriminator learns to use signals in the conditional data y to look for particular triggers in the image”)
Regarding claim 18, the above combination discloses the system of claim 15 wherein the discriminator network returns a fake result if the generated input image is fake and the predicted class input is real. (As above, See Gauthier pg. 2, right column, second paragraph from bottom teaches the function of the discriminator network “which accepts an image x and condition y [predicted class] and predicts the probability under condition y that x came from the empirical data distribution rather than from the generative model.” That is, the prediction is based on the probability under condition y [the class] that the image x is real. Also see pg. right column, ¶ 1, “The discriminator becomes more attuned to real facial features in order to distinguish between the simple “trick” images from the generator and real face images. Furthermore, the discriminator learns to use signals in the conditional data y to look for particular triggers in the image”)
Regarding claim 19, the above combination discloses the system of claim 15, wherein discriminator network generates a gradient to be backpropagated to the image recognition network and/or the image generation network as the punishment. (As above, see Gauthier, pg. 3, left column, ¶ 1 which teaches Equation (2) in which G (generator) and D (discriminator) play a minimax game to minimize each loss. In this way the discriminator is used to punish/provide loss in training the generator based whether the discriminator can distinguish real facial features from “trick” images from the generator. Pg. 3, right column, ¶ 4, “If both G [generator] and D [discriminator] are MLPs, we can train the framework by alternating between performing gradient-based updates on G and D.”)
Regarding claim 20, the above combination discloses the system of claim 15 wherein the image generation network takes as additional input random noise to introduce class independent semantic variations into the generated image. (Gauthier, pg. 2, right column, “Z is a noise space used to seed the generative model”)
Regarding claim 21, the above combination discloses the system of claim 15 wherein the images generated by the image generation network are used to train the image recognition network. (Sharma ¶ 0027-0028 teaches the GAN discriminator and generator training the facial image recognition network on the basis of generated and discriminated fake images in which generated negative face images are used to train the recognition network on what is not a face.)
Regarding claim 22, the above combination discloses the system of claim 15 wherein the image input to the image recognition network is a facial image. (See rejection of claim 15, Chauhan teaches CNN models for image recognition which image as input and predicts one or more classes of the image, see Pg. 279, right column, ¶ 2.)
Regarding claim 23, the above combination discloses the system of claim 22 wherein the facial images are generated by the image generation network and preserve a class identity of the face depicted in the facial image. (Gauthier, pg. 6, right column, “4.3. Conditional GAN” which teaches using face classes as input to generate face images fitted to that same class.)
Regarding claim 24, the above combination discloses the system of claim 15 further comprising: a processor; memory, storing software that, when executed by the processor, implement the image recognition network, the image generation network, and the discriminator network. (See rejection of claim 15 for the image recognition network, the image generation network, and the discriminator network. Sharma ¶ 0017 teaches a processor and memory.)
Conclusion
Based on these facts, THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Raphael Schwartz whose telephone number is (571)270-3822. The examiner can normally be reached Monday to Friday 9am-5pm CT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent Rudolph can be reached on (571) 272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RAPHAEL SCHWARTZ/ Examiner, Art Unit 2671