Prosecution Insights
Last updated: October 02, 2026
Application No. 18/439,748

TRAINING AND USING A NEURAL NETWORK FOR DEFECT DETECTION

Non-Final OA §101§102§103§112
Filed
Feb 12, 2024
Examiner
GERMICK, JOHNATHAN R
Art Unit
Tech Center
Assignee
Applied Materials Israel Ltd.
OA Round
1 (Non-Final)
46%
Grant Probability
Moderate
1-2
OA Rounds
1y 11m
Est. Remaining
77%
With Interview

Examiner Intelligence

Grants 46% of resolved cases
46%
Career Allowance Rate
48 granted / 104 resolved
-13.8% vs TC avg
Strong +31% interview lift
Without
With
+30.6%
Interview Lift
resolved cases with interview
Typical timeline
4y 7m
Avg Prosecution
27 currently pending
Career history
125
Total Applications
across all art units

Statute-Specific Performance

§101
28.8%
-11.2% vs TC avg
§103
39.1%
-0.9% vs TC avg
§102
16.9%
-23.1% vs TC avg
§112
14.4%
-25.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 104 resolved cases

Office Action

§101 §102 §103 §112
DETAILED ACTION This action is responsive to the Claims filed on 02/12/2024. Claims 1-22 are pending in the case. Claims 1, 13, 19-22 are independent claims. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claim 2 and 6 are objected to because of the following informalities: The claim does not end with a period. As noted in MPEP 608.01(m) “Each claim begins with a capital letter and ends with a period” Appropriate correction is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 1-22 rejected under 35 U.S.C. 101 because the claim(s) is directed to an abstract idea Regarding Claim 1/19/21 Under step 1, claim 1 is directed to a system for training a model which is directed to a machine, one of the statutory categories. Under step 1, claim 19 is directed to a method for training a model which is directed to a process, one of the statutory categories Under step 1, claim 21 is directed to a non-transitory computer readable storage medium which is directed to an article of manufacture, one of the statutory categories. Under Step 2A Prong 1, the claim recites the following limitations which are considered mental evaluations and/or mathematical calculations: using the desired probability function to repeatedly transform, until a specified criterion is met, elements in the input space to equal the plurality of elements in the latent space in compliance with an actual probability function that is indicative of an actual allocation of the elements to the one or more respective clusters determining a training loss value L associated with the elements in the latent space and testing if said training loss value L meets the specified criterion; said training loss value L being determined based on at least: a. a first term LRec indicative of a distance between the elements in the input space and elements in an output space reconstructed from the elements in the latent space; and b. a second term, LProb indicative of statistical distance between the desired probability function and the actual probability function. Each of these limitations describe using a transformation function on abstract data and making a determination of training value bases on certain features. Each of these steps reflect abstract evaluations which can be performed in the mind. Therefore, the claim recites an abstract idea. The claim recites the following additional element(s): representing a plurality of elements in the input space, each having N dimensions and being associated with at least one image of a semiconductor specimen, to a latent space representing an equal plurality of elements, each having M (M N) dimensions (is generally linking the use of the judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h)) the system comprising a processing and memory circuitry (PMC) configured to: (amounts to mere instructions to apply a computer technology to an abstract idea, see MPEP 2106.05(f)) obtain a desired probability function for transformation of the elements in the input space into one or more respective clusters of elements in the latent space; (that amounts to adding insignificant extra-solution activity to the judicial exception, because the limitation describe mere data gathering. See MPEP 2106.05(g)) The claim therefore is directed to an abstract idea. Under step 2B, The additional elements of obtain a desired probability function for transformation of the elements in the input space into one or more respective clusters of elements in the latent space] is well understood, routine, and conventional activity because it amounts to “transmitting or receiving data over a network" (see MPEP 2106.05(d)(II)(i). The recited additional elements when considered alone or in combination neither integrates the abstract idea into a practical application nor provides significantly more than the abstract idea itself. Regarding Claim 2 The claim(s) depends from claim 1 Each of the limitations described in the claim, under Step 2A Prong 1, only serve to describe the abstract ideas addressed in the parent claim, in particular the limitations describe mental evaluations. Furthermore, under step 2A Prong 2 and 2B, the claim does not recite additional elements to consider. Regarding Claim 3 The claim(s) depends from claim 1 The claim recites the following additional element(s), in addition to those already identified in the parent claim: for facilitating more efficient analysis of elements associated with the actual probability function that is sufficiently similar to the desired probability function in the latent space, rather than hypothetical analysis of the elements in the input space which inherently do not comply with the specified desired probability function. (is generally linking the use of the judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h)) The recited additional elements when considered alone or in combination neither integrates the abstract idea into a practical application nor provides significantly more than the abstract idea itself. Regarding Claim 4 The claim(s) depends from claim 1 The claim recites the following additional element(s), in addition to those already identified in the parent claim: wherein said training is a semi or fully supervised learning such that at least some of said plurality of elements in the input space are labeled with respective class of at least two classes for transformation of the elements in the input space into elements in one or more respective clusters of elements in the latent space. (is generally linking the use of the judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h)) The recited additional elements when considered alone or in combination neither integrates the abstract idea into a practical application nor provides significantly more than the abstract idea itself. Regarding Claim 5 The claim(s) depends from claim 4 Under Step 2A Prong 1, the claim recites the following limitations which are considered mental evaluations and/or mathematical calculations: elements in the input space are allocated to a number (I 2) of mutually discernible clusters in the latent space label each element of at least a subset of the input space with a selected class of said K classes; determine the training loss L value based also on: a third term, Lsup indicative of a degree of allocation of the transformed elements, labeled with classes each associated with a respective group of said I group of classes, to a corresponding cluster of said I clusters, such that the fewer the transformed elements that are allocated to other than said corresponding cluster of said I clusters, the lower said Lsup value. Each of these limitations describe manipulations of abstract data. Arranging data into clusters and labeling and making determinations about abstract values associated to the data are all steps which reflect abstract mental evaluation of data capable of being performed in the mind.. Therefore, the claim recites an abstract idea. The claim recites the following additional element(s): the (PMC) is further configured to: (amounts to mere instructions to apply a computer technology to an abstract idea, see MPEP 2106.05(f)) obtain data indicative of a number (K>_I) of classes each associated with a respective group of classes for each cluster, giving rise to I groups of classes; (that amounts to adding insignificant extra-solution activity to the judicial exception, because the limitation describe mere data gathering. See MPEP 2106.05(g)) The claim therefore is directed to an abstract idea. Under step 2B, The additional elements of obtain data indicative of a number (K>_I) of classes each associated with a respective group of classes for each cluster, giving rise to I groups of classes; is well understood, routine, and conventional activity because it amounts to “transmitting or receiving data over a network" (see MPEP 2106.05(d)(II)(i). The recited additional elements when considered alone or in combination neither integrates the abstract idea into a practical application nor provides significantly more than the abstract idea itself. Regarding Claim 6 The claim(s) depends from claim 4 Each of the limitations described in the claim, under Step 2A Prong 1, only serve to describe the abstract ideas addressed in the parent claim, in particular the limitations describe mental evaluations. Furthermore, under step 2A Prong 2 and 2B, the claim does not recite additional elements to consider. Regarding Claim 7 The claim(s) depends from claim 3 Under Step 2A Prong 1, the claim recites the following limitations which are considered mental evaluations and/or mathematical calculations: elements in the input space are allocated to a number (I 2) of mutually discernible clusters in the latent space b) label each element of at least a subset of the input space with a selected class of said K classes; c) determine the training loss L value based also on: a third term, Lsup indicative of a degree of allocation of the transformed elements, labeled with k classes to a corresponding cluster of said I clusters, such that the fewer the transformed elements that are allocated to other than said corresponding cluster of said I clusters, the lower said Ls~, value. Each of these limitations describe manipulations of abstract data. Arranging data into clusters and labeling and making determinations about abstract values associated to the data are all steps which reflect abstract mental evaluation of data capable of being performed in the mind.. Therefore, the claim recites an abstract idea. The claim recites the following additional element(s): the (PMC) is further configured to: (amounts to mere instructions to apply a computer technology to an abstract idea, see MPEP 2106.05(f)) obtain data indicative of a number (k< I) of classes;; (that amounts to adding insignificant extra-solution activity to the judicial exception, because the limitation describe mere data gathering. See MPEP 2106.05(g)) The claim therefore is directed to an abstract idea. Under step 2B, The additional elements of obtain data indicative of a number (k< I) of classes; is well understood, routine, and conventional activity because it amounts to “transmitting or receiving data over a network" (see MPEP 2106.05(d)(II)(i). The recited additional elements when considered alone or in combination neither integrates the abstract idea into a practical application nor provides significantly more than the abstract idea itself. Regarding Claim 8 The claim(s) depends from claim 1 The claim recites the following additional element(s), in addition to those already identified in the parent claim: wherein said model being a neural network that includes an encoder and decoder. (is generally linking the use of the judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h)) The recited additional elements when considered alone or in combination neither integrates the abstract idea into a practical application nor provides significantly more than the abstract idea itself. Regarding Claim 9 The claim(s) depends from claim 1 The claim recites the following additional element(s), in addition to those already identified in the parent claim: wherein the statistical distance between the desired probability function and the actual probability function is calculated utilizing Jenson Shannon or Kuliback-Leibler divergences. (is generally linking the use of the judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h)) The recited additional elements when considered alone or in combination neither integrates the abstract idea into a practical application nor provides significantly more than the abstract idea itself. Regarding Claim 10 The claim(s) depends from claim 1 The claim recites the following additional element(s), in addition to those already identified in the parent claim: wherein at least one of said clusters is characterized by a known statistical distribution (is generally linking the use of the judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h)) The recited additional elements when considered alone or in combination neither integrates the abstract idea into a practical application nor provides significantly more than the abstract idea itself. Regarding Claim 11 The claim(s) depends from claim 1 The claim recites the following additional element(s), in addition to those already identified in the parent claim: wherein at least two of said clusters are characterized by Gaussian Mixture Modeling (GMM) or Gaussian. (is generally linking the use of the judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h)) The recited additional elements when considered alone or in combination neither integrates the abstract idea into a practical application nor provides significantly more than the abstract idea itself. Regarding Claim 12 The claim(s) depends from claim 1 The claim recites the following additional element(s), in addition to those already identified in the parent claim: wherein the N dimensions input space associated with at least one image are informative of pixel values and/or at least two of the following: average intensity level, deviation from average pixel value, defect size, SNR (signal to noise ratio), correlation with predefined template, image moments. (is generally linking the use of the judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h)) The recited additional elements when considered alone or in combination neither integrates the abstract idea into a practical application nor provides significantly more than the abstract idea itself. Regarding Claim 13/20/22 Under step 1, claim 13 is directed to A system for utilizing a trained model for analyzing elements in a latent space which is directed to a machine, one of the statutory categories. Under step 1, claim 20 is directed to A method for utilizing a trained model for analyzing elements in a latent space which is directed to a process, one of the statutory categories. Under step 1, claim 22 is directed to A non-transitory computer readable storage medium which is directed to an article of manufacture, one of the statutory categories. Under Step 2A Prong 1, the claim recites the following limitations which are considered mental evaluations and/or mathematical calculations: transforming the at least one element to an equal number of elements in the latent space; c) for each transformed element, determine the distance between the element and reference to the probability function, wherein examination of the element is based on the determined distance. Each of these limitations describe using a transformation on abstract data and making a determination of distance based on certain features. Each of these steps reflect abstract evaluations which can be performed in the mind. Therefore, the claim recites an abstract idea. The claim recites the following additional element(s): the latent space representing a plurality of elements each having M dimensions that were transformed from elements in an input space having N (MAN) dimensions and being associated with at least one image of a semiconductor specimen; the transformed elements comply with a probability function (is generally linking the use of the judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h)) comprising a processing and memory circuitry (PMC) configured to… utilizing the trained model (amounts to mere instructions to apply a computer technology to an abstract idea, see MPEP 2106.05(f)) obtain at least one element in the input space that is associated with an image of a semiconductor specimen; (that amounts to adding insignificant extra-solution activity to the judicial exception, because the limitation describe mere data gathering. See MPEP 2106.05(g)) The claim therefore is directed to an abstract idea. Under step 2B, The additional elements of obtain at least one element in the input space that is associated with an image of a semiconductor specimen; is well understood, routine, and conventional activity because it amounts to “transmitting or receiving data over a network" (see MPEP 2106.05(d)(II)(i). The recited additional elements when considered alone or in combination neither integrates the abstract idea into a practical application nor provides significantly more than the abstract idea itself. Regarding Claim 14 The claim(s) depends from claim 13 Each of the limitations described in the claim, under Step 2A Prong 1, only serve to describe the abstract ideas addressed in the parent claim, in particular the limitations describe mental evaluations. Furthermore, under step 2A Prong 2 and 2B, the claim does not recite additional elements to consider. Regarding Claim 15 The claim(s) depends from claim 13 Under Step 2A Prong 1, the claim recites the following limitations which are considered mental evaluations and/or mathematical calculations: using the desired probability function to repeatedly transform, until a specified criterion is met, elements in the input space to equal the plurality of elements in the latent space in compliance with an actual probability function that is indicative of an actual allocation of the elements to the one or more respective clusters; c) determining a training loss value L associated with the elements in the latent space and testing if said training loss value L meets the specified criterion; said training loss value L being determined based on at least: a. a first term LReC indicative of a distance between the elements in the input space and elements in an output space reconstructed from the elements in the latent space; and b. a second term, Lep indicative of statistical distance between the desired probability function and the actual probability function. Each of these limitations describe manipulations of abstract data. Transforming abstract data and making determinations about abstract values associated to the data are all steps which reflect abstract mental evaluation of data capable of being performed in the mind.. Therefore, the claim recites an abstract idea. The claim recites the following additional element(s): wherein said model was trained by a PMC (amounts to mere instructions to apply a computer technology to an abstract idea, see MPEP 2106.05(f)) obtaining a desired probability function for transformation of the elements in the input space into one or more respective clusters of elements in the latent space;; (that amounts to adding insignificant extra-solution activity to the judicial exception, because the limitation describe mere data gathering. See MPEP 2106.05(g)) The claim therefore is directed to an abstract idea. Under step 2B, The additional elements of obtaining a desired probability function for transformation of the elements in the input space into one or more respective clusters of elements in the latent space; is well understood, routine, and conventional activity because it amounts to “transmitting or receiving data over a network" (see MPEP 2106.05(d)(II)(i). The recited additional elements when considered alone or in combination neither integrates the abstract idea into a practical application nor provides significantly more than the abstract idea itself. Regarding Claim 16-18 The claim(s) depends from claim 13 Each of the limitations described in the claim, under Step 2A Prong 1, only serve to describe the abstract ideas addressed in the parent claim, in particular the limitations describe mental evaluations. Furthermore, under step 2A Prong 2 and 2B, the claim does not recite additional elements to consider. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 2 and 6 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 2 and 6 use undefined variables α and β. These variables should be defined in the text of the claim. For the purposes of examination these are understood to be numerical values which may hold any value as they are weighting parameters. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1-4, 8-10, 12-22 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Yitian Wang “A class imbalanced wafer defect classification framework based on variational autoencoder generative adversarial network” Claim 1/19/21 Wang teaches, A system for training a model representing a plurality of elements in the input space [from claim 19] A method for training a model representing a plurality of elements in the input space [from claim 21 ] A non-transitory computer readable storage medium tangibly embodying a program of instructions that, when executed by a computer, cause the computer to perform a method for training a model representing a plurality of elements in the input space …the system comprising a processing and memory circuitry (PMC) configured to: (pg 8 “This training is implemented under the acceleration of GTX2070s through Python3.8.5, Tensorflow2.2, and Keras2.3.1.”, Figure 2 caption pg 5 “The overall flowcharts of VAE-GAN based data augmentation. The process is mainly divided into data processing; generative model training; sample generation and final CNN classification model training.”) each having N dimensions and being associated with at least one image of a semiconductor specimen (pg 4 Section 2 “we use the nearest interpolation method to resize the wafer maps of different sizes to 64 ×64 while ensuring that the information loss of the wafer maps is minimized. After that, to classify the amorphous part, the intact part and the defective part more intuitively, we use One-hot encoding to convert the original grayscale images into RGB images” here wafers are semiconductor specimen) to a latent space representing an equal plurality of elements, each having M (M<=N) dimensions, ( pg 8 Figure 3 “The overall model structure of VAE-GAN based CI-WDC framework” PNG media_image1.png 613 1036 media_image1.png Greyscale as shown in the figures the internal layers of the model amounts to mapping the input to a latent space which represents the vary same input having M dimensions as claimed) a) obtain a desired probability function for transformation of the elements in the input space into one or more respective clusters of elements in the latent space; (pg 5 Section 3.2 In this study, the basis for new wafer map generation is the VAE network. Define the wafer map samples as x; to generate a similar probability distribution of defected wafer map, the probability distribution p(x) of x needed to be obtained… p(x|z) describes a model that generates x from z, and p(z) obeys the standard normal distribution N(0,I)… We can generate the desired wafer maps by using this loss function to optimize the VAE model.” The generative VAE model defines a probability function for transforming input space into elements in latent space representing a cluster defined by the Normal distribution) b) using the desired probability function to repeatedly transform, until a specified criterion is met, elements in the input space to equal the plurality of elements in the latent space in compliance with an actual probability function that is indicative of an actual allocation of the elements to the one or more respective clusters; (pg 5 “By optimizing the generative model with such loss, we can achieve the constrain that ˜x(i) is infinitely close to x(i)” pg 6 “We can generate the desired wafer maps by using this loss function to optimize the VAE model” Algorithm 1 pg 6 PNG media_image2.png 67 266 media_image2.png Greyscale as shown in the algorithm the input is continuously transformed until the Epoch criterion is mean such that the loss function results in a VAE which is in compliance with an actual desired probability function as the model approaches the optimal.) determining a training loss value L associated with the elements in the latent space and testing if said training loss value L meets the specified criterion; said training loss value L being determined based on at least: (pg 5 algorithm 1 PNG media_image3.png 279 523 media_image3.png Greyscale as shown in the algorithm the training loss L is calculated. The value is considered to have met the criteria upon reaching the number of epochs E) a first term LRec indicative of a distance between the elements in the input space and elements in an output space reconstructed from the elements in the latent space; and (pg 5 Section 3.2 “The decoding process can realize the construction of the original defect wafer maps through transposed convo lution and restoring the original size of the wafer maps through upsampling. After ˜ x(i) is obtained, by minimizing the reconstruction loss between ˜ x(i) and x(i), the convolution ker nel re-updates the weights and make the probability distribu tion of the reconstruction on z(i) more close to the original wafer maps. The reconstruction loss are defined as follow: PNG media_image4.png 50 282 media_image4.png Greyscale ”) a second term, LProb indicative of statistical distance between the desired probability function and the actual probability function. ( pg 6 Section 3.2 “Therefore we need to introduce Kullback–Leibler divergence (KLD) to solve this problem. KLD measures the asymmetry of the difference between two probability distributions. When the KLD of two probability distribution drops, p(z|x(i)) gets closer to p(z) PNG media_image5.png 123 395 media_image5.png Greyscale “therefore, the loss function of the entire variational autoencoder is defined as” PNG media_image6.png 40 175 media_image6.png Greyscale ) Claim 2 Wang teaches claim 1 Wang teaches, wherein said training loss value L complies with the following equation: PNG media_image7.png 32 195 media_image7.png Greyscale (pg 6 Section 3.2 “therefore, the loss function of the entire variational autoencoder is defined as” PNG media_image6.png 40 175 media_image6.png Greyscale ” as noted in the 112 rejection alpha is undefined by the claims. When alpha is chosen to be ½ the total loss complies with the claimed formula.) Claim 3 Wang teaches claim 1 Wang teaches, for facilitating more efficient analysis of elements associated with the actual probability function that is sufficiently similar to the desired probability function in the latent space, rather than hypothetical analysis of the elements in the input space which inherently do not comply with the specified desired probability function. ( pg 6 Section 3.2 “When we continue to reduce KLD by neural networks, the noise intensity can gradually increase. This way, by decoding the noisy samples z(i), we get a new wafer map that differs in several pixels while keeping the general defect characteristics the same.” Examiner notes the claim describes intended use and is not given patentable weight because the limitations do not limit the systems function and instead describe the purpose or intended use. Nevertheless, the model is for analysis which model wafers having general defect characteristics similar to the desired probability function as claimed, rather than hypothetical analysis) Claim 4 Wang teaches claim 1 Wang teaches, wherein said training is a semi or fully supervised learning such that at least some of said plurality of elements in the input space are labeled with respective class of at least two classes for transformation of the elements in the input space into elements in one or more respective clusters of elements in the latent space ( pg 7 “a CNN model is adopted as a classifier for wafer defect maps….In order to better calculate the error between the predicted value and the true value, an appropriate loss function needs to be selected. We take categorical-cross entropy loss as the loss function for the wafer defect map classifier. PNG media_image8.png 78 265 media_image8.png Greyscale ” the described cross entropy loss is a supervised learning method as it is based on a loss which incorporates underlying true values. Thus the composite system is trained at least in part on fully supervised learning which elements labeled by respective classes via the classifier. The classification is based on the transformation from latent space, which as shown in figure 3 is od multiple classes.) Claim 8 Wang teaches claim 1 Wang teaches, wherein said model being a neural network that includes an encoder and decoder. (pg 6 Section 3.3 “In VAE-GAN network, there are three major models, encoder, decoder and discriminator.”) Claim 9 Wang teaches claim 1 Wang teaches, wherein the statistical distance between the desired probability function and the actual probability function is calculated utilizing Jenson Shannon or Kuliback-Leibler divergences. ( pg 6 Section 3.2 “Therefore we need to introduce Kullback–Leibler divergence (KLD) to solve this problem. KLD measures the asymmetry of the difference between two probability distributions.”) Claim 10 Wang teaches claim 1 Wang teaches, wherein at least one of said clusters is characterized by a known statistical distribution. (pg 5 Section 3.2 “In this study, the basis for new wafer map generation is the VAE network… Define the probability distribution p(x) as follow… where z here is the latent variable from normal distribution… describes a model that generates x from z, and p(z) obeys the standard normal distribution” a normal distribution is a known statistical distribution) Claim 12 Wang teaches claim 1 Wang teaches, wherein the N dimensions input space associated with at least one image are informative of pixel values and/or at least two of the following: average intensity level, deviation from average pixel value, defect size, SNR (signal to noise ratio), correlation with predefined template, image moments. (pg 4 Section 2 “we use One-hot encoding to convert the original grayscale images into RGB images. In the wafer maps, each pixel has three different possible values cor responding to three cases: when the pixel area is outside the wafer (Case 1); when the area is inside the wafer and normal (Case 2); when the area is inside the wafer and is defective (Case 3)” the image of N dimensions in input space is associated with informative pixel values. Further is informative of deviation from normal or average pixel value, and defect size) Claim 13/20/22 Wang teaches, A system for utilizing a trained model for analyzing elements in a latent space; [from claim 20] A method for utilizing a trained model for analyzing elements in a latent space [from claim 22] A non-transitory computer readable storage medium tangibly embodying a program of instructions that, when executed by a computer, cause the computer to perform a method for utilizing a trained model for analyzing elements in a latent space…the system comprising a processing and memory circuitry (PMC) configured to: (Figure 2 caption pg 5 “The overall flowcharts of VAE-GAN based data augmentation. The process is mainly divided into data processing; generative model training; sample generation and final CNN classification model training.” pg 8 “This training is implemented under the acceleration of GTX2070s through Python3.8.5, Tensorflow2.2, and Keras2.3.1.”) the latent space representing a plurality of elements each having M dimensions that were transformed from elements in an input space having N (M<=N) dimensions and being associated with at least one image of a semiconductor specimen… a) obtain at least one element in the input space that is associated with an image of a semiconductor specimen; (pg 4 Section 2 “we use the nearest interpolation method to resize the wafer maps of different sizes to 64 ×64 while ensuring that the information loss of the wafer maps is minimized. After that, to classify the amorphous part, the intact part and the defective part more intuitively, we use One-hot encoding to convert the original grayscale images into RGB images” here wafers are semiconductor specimen. pg 8 Figure 3 “The overall model structure of VAE-GAN based CI-WDC framework” PNG media_image1.png 613 1036 media_image1.png Greyscale as shown in the figures the internal layers of the model amounts to mapping the input to a latent space which represents the very same input having M dimensions as claimed)) the transformed elements comply with a probability function (pg 5 Section 3.2 “In this study, the basis for new wafer map generation is the VAE network. Define the wafer map samples as x; to generate a similar probability distribution of defected wafer map, the probability distribution p(x) of x needed to be obtained”) utilizing the trained model for transforming the at least one element to an equal number of elements in the latent space; (pg 6 Section 3.4 “Firstly, we performed data processing to get ideal sized wafer maps and used them for training our VAE-GAN models. After training the processed with multiple batches in VAE, we can get a well-trained encoder, decoder, and discriminator of these three models for the data generation process… The encoder can generate high dimensional vectors of variance and mean from the original data and add noise to these two parameters to sample the latent variables of the wafer map features. The decoder can decode these latent variables into a generated image. The resulting image is usually similar to the original in overall defect distribution” the encoder and decoder transform an elements to an equal number of elements in the latent space. The transformation is a representation of the same input elements thus an equal number as claimed) for each transformed element, determine the distance between the element and reference to the probability function, wherein examination of the element is based on the determined distance. (pg 6 Section 3.2 “Therefore we need to introduce Kullback–Leibler divergence (KLD) to solve this problem. KLD measures the asymmetry of the difference between two probability distributions... PNG media_image9.png 44 261 media_image9.png Greyscale ” the KL divergence measures the distance between the transformed element and the reference distribution. Here difference and distance are synonomous) Claim 14 Wang teaches claim 1 Wang teaches, wherein said reference to the probability function being the center of said probability function. (pg 6 Section 3.2 “Therefore we need to introduce Kullback–Leibler divergence (KLD) to solve this problem. KLD measures the asymmetry of the difference between two probability distributions... PNG media_image9.png 44 261 media_image9.png Greyscale ” the KL divergence measures the distance between two distributions which includes the center of the probability function) Claim 15 Wang teaches claim 13 Wang teaches, obtaining a desired probability function for transformation of the elements in the input space into one or more respective clusters of elements in the latent space; (pg 5 Section 3.2 In this study, the basis for new wafer map generation is the VAE network. Define the wafer map samples as x; to generate a similar probability distribution of defected wafer map, the probability distribution p(x) of x needed to be obtained… p(x|z) describes a model that generates x from z, and p(z) obeys the standard normal distribution N(0,I)… We can generate the desired wafer maps by using this loss function to optimize the VAE model.” The generative VAE model defines a probability function for transforming input space into elements in latent space representing a cluster defined by the Normal distribution) b) using the desired probability function to repeatedly transform, until a specified criterion is met, elements in the input space to equal the plurality of elements in the latent space in compliance with an actual probability function that is indicative of an actual allocation of the elements to the one or more respective clusters; (pg 5 “By optimizing the generative model with such loss, we can achieve the constrain that ˜x(i) is infinitely close to x(i)” pg 6 “We can generate the desired wafer maps by using this loss function to optimize the VAE model” Algorithm 1 pg 6 PNG media_image2.png 67 266 media_image2.png Greyscale as shown in the algorithm the input is continuously transformed until the Epoch criterion is mean such that the loss function results in a VAE which is in compliance with an actual desired probability function as the model approaches the optimal.) c) determining a training loss value L associated with the elements in the latent space and testing if said training loss value L meets the specified criterion; said training loss value L being determined based on at least: (pg 5 algorithm 1 PNG media_image3.png 279 523 media_image3.png Greyscale as shown in the algorithm the training loss L is calculated. The value is considered to have met the criteria upon reaching the number of epochs E) a. a first term LReC indicative of a distance between the elements in the input space and elements in an output space reconstructed from the elements in the latent space; and (pg 5 Section 3.2 “The decoding process can realize the construction of the original defect wafer maps through transposed convo lution and restoring the original size of the wafer maps through upsampling. After ˜ x(i) is obtained, by minimizing the reconstruction loss between ˜ x(i) and x(i), the convolution ker nel re-updates the weights and make the probability distribu tion of the reconstruction on z(i) more close to the original wafer maps. The reconstruction loss are defined as follow: PNG media_image4.png 50 282 media_image4.png Greyscale ”) b. a second term, Lep indicative of statistical distance between the desired probability function and the actual probability function ( pg 6 Section 3.2 “Therefore we need to introduce Kullback–Leibler divergence (KLD) to solve this problem. KLD measures the asymmetry of the difference between two probability distributions. When the KLD of two probability distribution drops, p(z|x(i)) gets closer to p(z) PNG media_image5.png 123 395 media_image5.png Greyscale “therefore, the loss function of the entire variational autoencoder is defined as” PNG media_image6.png 40 175 media_image6.png Greyscale ) Claim 16 Wang teaches claim 13 Wang teaches, wherein said analysis includes determining anomality of the transformed elements. (pg 7 Section 3.6 “In the wafer defect classifier, we use the softmax activation function to achieve classification of multiple defect classes” the classifier determines defect classes with is a determining of a type of anomality as compared to a normal non defect case) Claim 17 Wang teaches claim 13 Wang teaches, wherein said analysis includes determining association of each transformed element to a cluster of said clusters. ( pg 13 Figure 9 caption “t-SNE comparison before and after VAE-GAN processing….The colors correspond to different types of defect. The relationships are follows: red: Center, tomato: Donut, orange:… PNG media_image10.png 303 484 media_image10.png Greyscale the figure depicts the association each element which was transformed after VAE-GAN processing to a cluster of a plurality of clusters.) Claim 18 Wang teaches claim 13 Wang teaches, wherein said analysis includes generating at least one new element in the latent space being re-constructible to a corresponding at least one output element in the output space (Section 3.2 pg 5 “In this study, the basis for new wafer map generation is the VAE network…. z(i) is then constructed by the decoder into ˜ x(i). ˜ x(i) is the new wafer map we generate through the entire VAE network... The decoding process can realize the construction of the original defect wafer maps through transposed convolution and restoring the original size of the wafer maps through upsampling” the purpose of an encoder/decoder network is to generate new samples from latent space, i.e the latent space element being re-constructible) each of the at least one output element constitutes a new synthetic input element for training a model. (pg 6 “It generates judgments on the confidence of wafer defect maps by directly understanding the difference between the original image and the generated image. Discriminator learns the feature of new wafer maps that are continuously updated by the VAE so that in each training batch, the ability to identify unqualified wafer maps can be improved” the Discriminator is trained based on the new synthetic input element. Pg 13 Section 6 “In particular, the VAE network is responsible for generating new samples. The introduced discriminator can identify unqualified samples by training the generated and original samples. When we perform data augmentation, the images generated by VAE will go through the discriminator’s validation test, the samples that pass the test are responsible for CNN training.”) Claim Rejections - 35 U.S.C. § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. §§ 102 and 103 (or as subject to pre-AIA 35 U.S.C. §§ 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. § 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 5-7 are rejected under 35 U.S.C. § 103 as being unpatentable over Wang further in view of Kim et al “Dynamic Clustering for Wafer Map Patterns Using Self-Supervised Learning on Convolutional Autoencoders”, further still in view of Bai et al. “Gaussian Mixture Variational Autoencoder with Contrastive Learning for Multi-Label Classification” Claim 5 Wang teaches claim 4 Wang teaches, label each element of at least a subset of the input space with a selected class of said K classes ( pg 7 Section 3.6 “In this framework, a CNN model is adopted as a classifier for wafer defect maps… The vectorized categories correspond to the multiple probability values trained by the neural network. Thus, the vector corresponding to the output with the largest probability value is the obtained label.” The system selects a class of the set of k classes based on the input space.) Wang does not explicitly teach, elements in the input space are allocated to a number (I 2) of mutually discernible clusters in the latent space and the (PMC) is further configured to: a) obtain data indicative of a number (K>_I) of classes each associated with a respective group of classes for each cluster, giving rise to I groups of classes; c) determine the training loss L value based also on:a third term, Lsup indicative of a degree of allocation of the transformed elements, labeled with classes each associated with a respective group of said I group of classes, to a corresponding cluster of said I clusters, such that the fewer the transformed elements that are allocated to other than said corresponding cluster of said I clusters, the lower said Lsup value. Kim however when addressing clustering algorithms for classification of semiconductor wafers teaches, elements in the input space are allocated to a number (I 2) of mutually discernible clusters in the latent space and the (PMC) is further configured to: a) obtain data indicative of a number (K>_I) of classes each associated with a respective group of classes for each cluster, giving rise to I groups of classes; (pg 4 Section 4 “Training the proposed model consists of three phases, as depicted in Fig. 3:… extracting semantic visual features… determining the number of clusters and assigning cluster membership … to each WBM based on the DPMM” the model is trained via extracting elements in input space and assigning corresponding classes to the latent space depicted in Figure 3) determine the training loss L value based also on ( pg 5 “Algorithm1 describes the whole procedure of the TDP…The convergence criterion of the EM algorithm is determined by the evidence lower bound (ELBO), which is maximized by updating the mixture parameters for the responsibilities (line 19 in Algorithm1)” algorithm 1 depicts the training loss.) Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify neural network wafer classification system of Wang to comprise latent space clustering as described by Kim. One would have been motivated to make such a combination because Wang and Kim are directed to classification of semiconductor wafer images. Further, Kim notes “However, conventional methods not only require a large amount of labeled data for pattern classification but also need to fix the number of patterns. To overcome these limitations, we propose self-supervised learning-based dynamic clustering to identify failure patterns in an uninterrupted and continuous semiconductor manufacturing process” (Kim pg 10) Wang/Kim does not explicitly teach, a third term, Lsup indicative of a degree of allocation of the transformed elements, labeled with classes each associated with a respective group of said I group of classes, to a corresponding cluster of said I clusters, such that the fewer the transformed elements that are allocated to other than said corresponding cluster of said I clusters, the lower said Lsup value. Bai however when addressing gaussian clustering for neural network classification teaches, a third term, Lsup indicative of a degree of allocation of the transformed elements, labeled with classes each associated with a respective group of said I group of classes, to a corresponding cluster of said I clusters, such that the fewer the transformed elements that are allocated to other than said corresponding cluster of said I clusters, the lower said Lsup value. ( pg 5 Section 2.2.2 “Our objective function also includes a supervised cross entropy loss term to further facilitate the training. With the label embeddings wl i and the feature embedding wf x, the cross entropy loss for each (x,y) is given by: PNG media_image11.png 56 389 media_image11.png Greyscale … PNG media_image12.png 45 318 media_image12.png Greyscale The final objective function to minimize is simply the summation of different losses,… This is partly because we also learn a latent space that is closely connected to label embeddings.” The cross entropy loss learns to minimize the degree of distance or allocation between the learned embeddings via the latent space and the label embeddings, thus resulting in a lower value when the alignment between the transformed elements and the cluster labels is low as claimed.) Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify neural network wafer classification system of Wang/Kim to comprise a joint loss function for an improved neural network described by Bai. One would have been motivated to make such a combination because Wang/Kim and Bai are directed to classification based on the latent space of input data. Further, Bai notes “A joint training strategy reconciles different modules. We show in the experiments that the learnt embed dings are semantically meaningful and can reveal the label correlations.” (Bai pg 5) Claim 6 Wang/Kim/Bai teaches claim 5 Bai teaches, wherein said training loss value L complies with the following equation: PNG media_image13.png 27 318 media_image13.png Greyscale ( pg 5 Section 2.2.3 “The final objective function to minimize is simply the sum mation of different losses, PNG media_image14.png 33 272 media_image14.png Greyscale ” the KL loss corresponds to the Lprob. Neither the specification nor the claims provide values for alpha and beta. A value of 1 for alpha and 1 for Beta complies with the equation in the art) Claim 7 Wang teaches claim 3 Wang teaches, a) obtain data indicative of a number (k< I) of classes; b) label each element of at least a subset of the input space with a selected class of said K classes; ( pg 7 Section 3.6 “In this framework, a CNN model is adopted as a classifier for wafer defect maps… The vectorized categories correspond to the multiple probability values trained by the neural network. Thus, the vector corresponding to the output with the largest probability value is the obtained label.” The system selects a class of the set of k classes based on the input space.) Wang does not explicitly teach, elements in the input space are allocated to a number (I 2) of mutually discernible clusters in the latent space and the (PMC) is further configured to: c) determine the training loss L value based also on:a third term, Lsup indicative of a degree of allocation of the transformed elements, labeled with k classes to a corresponding cluster of said I clusters, such that the fewer the transformed elements that are allocated to other than said corresponding cluster of said I clusters, the lower said Lsup, value. Kim however when addressing clustering algorithms for classification of semiconductor wafers teaches, elements in the input space are allocated to a number (I 2) of mutually discernible clusters in the latent space and the (PMC) is further configured to: (pg 4 Section 4 “Training the proposed model consists of three phases, as depicted in Fig. 3:… extracting semantic visual features… determining the number of clusters and assigning cluster membership … to each WBM based on the DPMM” the model is trained via extracting elements in input space and assigning corresponding classes to the latent space depicted in Figure 3) c) determine the training loss L value ( pg 5 “Algorithm1 describes the whole procedure of the TDP…The convergence criterion of the EM algorithm is determined by the evidence lower bound (ELBO), which is maximized by updating the mixture parameters for the responsibilities (line 19 in Algorithm1)” algorithm 1 depicts the training loss.) Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify neural network wafer classification system of Wang to comprise latent space clustering as described by Kim. One would have been motivated to make such a combination because Wang and Kim are directed to classification of semiconductor wafer images. Further, Kim notes “However, conventional methods not only require a large amount of labeled data for pattern classification but also need to fix the number of patterns. To overcome these limitations, we propose self-supervised learning-based dynamic clustering to identify failure patterns in an uninterrupted and continuous semiconductor manufacturing process” (Kim pg 10) Wang/Kim does not explicitly teach, a third term, Lsup indicative of a degree of allocation of the transformed elements, labeled with k classes to a corresponding cluster of said I clusters, such that the fewer the transformed elements that are allocated to other than said corresponding cluster of said I clusters, the lower said Lsup, value. Bai however when addressing gaussian clustering for neural network classification teaches, a third term, Lsup indicative of a degree of allocation of the transformed elements, labeled with k classes to a corresponding cluster of said I clusters, such that the fewer the transformed elements that are allocated to other than said corresponding cluster of said I clusters, the lower said Lsup, value. ( pg 5 Section 2.2.2 “Our objective function also includes a supervised cross entropy loss term to further facilitate the training. With the label embeddings wl i and the feature embedding wf x, the cross entropy loss for each (x,y) is given by: PNG media_image11.png 56 389 media_image11.png Greyscale … PNG media_image12.png 45 318 media_image12.png Greyscale The final objective function to minimize is simply the summation of different losses,… This is partly because we also learn a latent space that is closely connected to label embeddings.” The cross entropy loss learns to minimize the degree of distance or allocation between the learned embeddings via the latent space and the label embeddings, thus resulting in a lower value when the alignment between the transformed elements and the cluster labels is low as claimed.) Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify neural network wafer classification system of Wang/Kim to comprise a joint loss function for an improved neural network described by Bai. One would have been motivated to make such a combination because Wang/Kim and Bai are directed to classification based on the latent space of input data. Further, Bai notes “A joint training strategy reconciles different modules. We show in the experiments that the learnt embed dings are semantically meaningful and can reveal the label correlations.” (Bai pg 5) Claim(s) 11 are rejected under 35 U.S.C. § 103 as being unpatentable over Wang further in view of Bai et al. “Gaussian Mixture Variational Autoencoder with Contrastive Learning for Multi-Label Classification” Claim 11 Wang teaches claim 4 Wang does not explicitly teach, wherein at least two of said clusters are characterized by Gaussian Mixture Modeling (GMM) or Gaussian. Bai however when addressing gaussian clustering for neural network classification teaches, wherein at least two of said clusters are characterized by Gaussian Mixture Modeling (GMM) or Gaussian. ( pg 2 Section 2.1.1 “In our work, we adopt the Gausian mixture prior… where i is the cluster index of k Gaussian clusters with mean µi and covariance σ2… Our intuition is that each label embedding could correlate to a Gaussian subspace. Given a label set, the mixture of the positive Gaussian subspaces forms a unique multimodal prior distribution” a number of Gaussian clusters characterize the model) Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify neural network wafer classification system of Wang to comprise a Gaussian mixtures model described by Bai. One would have been motivated to make such a combination because Wang/Kim and Bai are directed to classification based on the latent space of input data. Further, Bai describes a limitation of standard VAE models suggesting a Gaussian prior as an improvement “A standard VAE… One weakness of this formulation is the unimodality of its latent space, inhibiting the learning of more complex representations. Another concern is over-regularization… Numerous works extend the prior to be more complex… In our work, we adopt the Gausian mixture prior” (Bai pg 2) Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOHNATHAN R GERMICK whose telephone number is (571)272-8363. The examiner can normally be reached M-F 9:30-4:30. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached on 571-272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /J.R.G./ Examiner, Art Unit 2122 /KAKALI CHAKI/Supervisory Patent Examiner, Art Unit 2122
Read full office action

Prosecution Timeline

Feb 12, 2024
Application Filed
Sep 16, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749029
ONLINE OPERATING MODE TRAJECTORY OPTIMIZATION FOR PRODUCTION PROCESSES
7y 2m to grant Granted Sep 29, 2026
Patent 12725012
LOW-LATENCY TIME-ENCODED SPIKING NEURAL NETWORK
3y 4m to grant Granted Sep 01, 2026
Patent 12711384
Neural Network Initialization
5y 3m to grant Granted Aug 18, 2026
Patent 12699893
SELF-SUPERVISED REPRESENTATION LEARNING USING BOOTSTRAPPED LATENT REPRESENTATIONS
5y 2m to grant Granted Aug 04, 2026
Patent 12694266
CORRELATION RECURRENT UNIT FOR IMPROVING PREDICTION PERFORMANCE OF TIME-SERIES DATA AND CORRELATION RECURRENT NEURAL NETWORK
3y 7m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
46%
Grant Probability
77%
With Interview (+30.6%)
4y 7m (~1y 11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 104 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month