DETAILED ACTION
Status of Claims
Claim(s) 16 and 18-30 are pending and are examined herein.
Claim(s) 16, 18, 21, 27, and 29-30 have been Amended. Claim(s) 17 has been Canceled.
Claim(s) 16 and 18-30 remain rejected under 35 U.S.C. § 103.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of the applicant’s claim for foreign priority under 35 U.S.C. 119 to the foreign applications identified in the Application.
Response to Amendment
The amendment filed on May 26, 2026 has been entered. Claims 16 and 18-30 are pending in the application. Applicant’s amendments to the claims have overcome the rejections under 35 USC § 112(b) and nonstatutory double patenting, as well as the objections set forth in the Non-Final Office Action mailed on February 27, 2026. However, the present amendments introduce new issues, which are addressed below. The present amendments have been fully considered and are addressed in the rejections below.
Response to Arguments
Applicant's arguments with respect to the rejection under 35 U.S.C. § 101, filed on 05/26/2026, have been fully considered and are persuasive. (See Remarks pp. 13-25).
Specifically, the following remarks are considered persuasive:
Applicant argues that the pending claim as a whole provides a specific architectural improvement: and operational limitations-a modular autoencoder hat (i) processes heterogenous inputs.
Applicant further points out that the improvement disclosed in the specification is reflected in the claim language, and thus integrate any judicial exception into a practical application. In particular, Applicant argues that the claim recites an architecture that “solves a concrete technical problem in semiconductor metrology and manufacturing-namely opacity, inflexibility, and bias in monolithic autoencoders-by enabling interpretable, extensible, physics-aligned parameter estimation, which conventional ML systems do not provide. In other words, this is an improvement to the architecture of the autoencoder model.”
For at least the above reasons, Applicant’s arguments are persuasive and the rejection under 35 U.S.C. § 101 is withdrawn.
Applicant's arguments, with respect to the rejection under 35 U.S.C. § 103, filed on 05/26/2026 (see Remarks pp. 25-34), have been fully considered but are not persuasive and are moot in view of the new grounds of rejection necessitated by amendments.
Applicant argues that the cited references fail to disclose or suggest all limitations in currently amended claims 16, 27, 29, and 30.
Specifically, Applicant contends that "Shazeer and Rothberg are silent as to processing one or more inputs to a first level of dimensionality using individual input models and using one or more output models to generate one or more outputs from expanded data, where individual input and output models comprise two or more sub-models associated with different portions of a sensing operation in a semiconductor manufacturing process and/or a semiconductor manufacturing process." Applicant further asserts that “Schlake does not disclose or suggest "wherein individual input models comprise two or more sub-models, the two or more sub-models associated with different portions of a sensing operation in a semiconductor manufacturing process and/or a semiconductor manufacturing process.”
The Examiner respectfully disagrees with Applicant’s arguments for the following reasons.
It is important to note that one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986). The claim limitations are rejected using the combination of references, and the proper test is what the combined teachings would have suggested to a person of ordinary skill in the art. See MPEP § 2143.
Shazeer teaches the core architectural limitations. As noted in the Non-Final Office Action, Shazeer teaches “processing one or more inputs to a first level of dimensionality using individual input models and using one or more output models to generate one or more outputs from expanded data.” Specifically, Shazeer teaches a plurality of multiple input modality neural networks 102a-102c that are architecturally separate preprocessors mapping raw modality inputs into a common dimensionality (i.e., first level of dimensionality), and a plurality of multiple output modality neural networks 108-a108c that are architecturally distinct networks using the decoder’s output (expanded representation) to generate one or more outputs. See Shazeer e.g., [0011], [0014], [0043], and [0047].
Shazeer also teaches that “individual input and output models comprise two or more sub-models.” Specifically, Shazeer teaches that each of the input and output modality neural networks consist of multiple neural network layers and/or stack of ConvRes blocks. See Shazeer e.g., [0056] and [0062].
With respect to the recitations of “the two or more sub-models associated with different portions of a sensing operation in a semiconductor manufacturing process and/or a semiconductor manufacturing process” (individual input side) and “the two or more sub-models associated with different portions of a sensing operation in a semiconductor manufacturing process and/or a semiconductor manufacturing process” (individual output side). Under the broadest reasonable interpretation (BRI), the recitation of “associated with different portions of a sensing operation in a semiconductor manufacturing process and/or a semiconductor manufacturing process” merely define the domain or field with which the sub-models are associated. They do not describe any structural feature of the sub-models themselves, active method step, or any operational requirement on the sub-models. Nor do these recitations give "meaning and purpose to the manipulative steps" of the claimed method, including the processing, combining, expanding, using, and estimating steps of claim 16. Accordingly, these recitations are non-limiting intended-use or field-of-use language. For examination purposes, the Examiner interprets the recitations as reciting “the two or more sub-models associated with different portions of a sensing operation or process.
As explained above, while Shazeer is silent on whether the input and output sub-models are associated with different portions of a sensing operation or a process. Schlake teaches this limitation. Specifically, Schlake teaches a plurality of input and output sub-models processing input and outputting predicted future values, and further describes the models and/or sub-models as associated with different portions of measurement operations and physical processes. See Schlake e.g., [0041], [0063], [0069] and [0075]-[0077].
With respect to applicant’s assertion that cited references for dependent claims fails to disclose or suggest this subject matter in claim 16. The examiner respectfully disagrees for similar reasons described above.
Accordingly, for at least the above reasons, Applicant’s arguments are not persuasive, and the rejection under 35 U.S.C. § 103 is maintained. The examiner refers to the updated rejection below for more details.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION. —The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim(s) 16-30 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, for pre-AIA the applicant regards as the invention.
Regarding Currently Amended Claim 16, the claim recites limitations that renders the scope of the claimed invention indefinite for the following reasons:
The claim recites the limitation “the two or more sub-models associated with different portions of a sensing operation in a semiconductor manufacturing process and/or a semiconductor manufacturing process” lines 4-6. This limitation recites two elements using the inclusive disjunctive “and/or”:
First element: “a sensing operation in a semiconductor manufacturing process”.
Second element: “a semiconductor manufacturing process”.
The use of inclusive disjunctive is broadly construed to cover the first element alone, the second element, or both elements together. However, the second element appear to define a separate distinct element “a semiconductor manufacturing process” and not referring to the “semiconductor manufacturing process” recited earlier. It is therefore unclear whether the two or more sub-models associated with different portions of a sensing operation in a semiconductor manufacturing process, the two or more sub-models associated with different portions of a semiconductor manufacturing process, or the two or more sub-models associated with different portions of both: (i) a sensing operation in a semiconductor manufacturing process and (ii) a semiconductor manufacturing process. Additionally, it is unclear whether the second element refers to the same semiconductor manufacturing process as that embedded in the first element, or introduces an additional, distinct semiconductor manufacturing process. This lack of clarity renders the claim scope indefinite.
The claim recites the limitation “the two or more sub-models associated with different portions of a sensing operation in a semiconductor manufacturing process and/or a semiconductor manufacturing process” lines 21-22. Similarly, this limitation introduces issues similar to those identified in the previous limitation and renders the scope of the claim indefinite.
For at least the above reasons, claim 16 does not particularly point out and distinctly claim the invention.
Regarding Currently Amended Claim 18, the claim recites limitation “wherein two or more sub-models of an individual output model comprise a sensor model and a stack model for a semiconductor sensor operation” without a proper antecedent basis for the recitation of “two or more sub-models of an individual output model.”
Claim 16, from which claim 18 depends, already recite that “individual output models comprise two or more sub-models.” Thus, it is unclear whether the “two or more sub-models” and the “individual model” recited in claim 18 refer to the same “two or more sub-models” and the same “individual output models” already introduced in claim 16, or whether claim 18 is introducing new elements. This lack of clarity renders the scope of the claim unclear. For examination, these elements are broadly interpreted as referring to the same elements already introduced in the parent claim 16.
Regarding Currently Amended Claims 27, 29, and 30, the claims recite substantially similar limitations as those of claim 16 and are rejected for similar reasons and rationale.
Regarding dependent claims 18-26 and 28, these claims depend from a rejected claim (16 and 27) and therefore inherit the deficiencies of the respective parent claim.
In view of the above, Examiner respectfully requests that Applicant thoroughly review the claims for compliance with the requirements set forth under 35 U.S.C. § 112.
Claim Interpretation
The recitations of “two or more sub-models associated with different portions of a sensing operation in a semiconductor manufacturing process and/or a semiconductor manufacturing process” for both individual input models (lines 4-6) and individual output models (lines 21-22) are interpreted as non-limiting field-of-use language. Under the broadest reasonable interpretation, these recitations merely define the domain or field with which the sub-models are associated. They don’t describe any active steps, any structural features of the sub-models themselves, or any operational requirement on the sub-models that impose a structural or functional limitation on the claimed invention. Nor do these recitations give "meaning and purpose to the manipulative steps" of the claimed method. See, e.g., Griffin v. Bertina, 285 F.3d 1029, 1034, 62 USPQ2d 1431 (Fed. Cir. 2002). Accordingly, these recitations are non-limiting intended-use or field-of-use statement. See MPEP § 2103.I.C. For examination purposes, the Examiner interprets these limitations as reciting “the two or more sub-models associated with different portions of a sensing operation and/or a manufacturing process.”
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 16, 20-23, 27, and 29-30 are rejected under 35 U.S.C. 103 as being unpatentable over Shazeer et al., (Pub. No.: US 20200364405 A1) in view of Rothberg et al., (Pub. No.: US 20190347523 A1), and further in view of Schlake et al., (Pub. No.: US 20210349453 A1), hereinafter, the combination of Shazeer, Rothberg, and Schlake teaches the following.
Regarding Currently Amended Claim 16,
Shazeer discloses the following:
A method for parameter estimation, the method comprising: (Shazeer, [0038] “This specification describes a multi model neural network architecture including a single deep learning model that can simultaneously learn different machine learning tasks from different machine learning domains.” [0063] “The multi task multi modal machine learning model 100 can be trained to perform different machine learning tasks from different machine learning domains or modalities using training data.”)
processing, with one or more input models of a modular autoencoder model, one or more inputs to a first level of dimensionality suitable for combination with other inputs; (Shazeer, [0043]-[0045] “The multi task multi modal machine learning model 100 includes multiple input modality neural networks 102 a-102 c, ... Data inputs, e.g., data input 110, received by the multi task multi modal machine learning model 100 are provided to the multiple input modality neural networks 102 a-102 c and processed by an input modality neural network corresponding to the modality (domain) of the data input. ... The input modality neural networks 102 a-102 c are configured to process received data inputs and to generate as output mapped data inputs from a unified representation space, e.g., mapped data 112. ... Each input modality neural network of the multiple input modality networks 102 a-c is configured to map received machine learning model data inputs of one of multiple machine learning domains or modalities to mapped data inputs of a unified representation space. That is, each input modality neural network is specific to a respective modality (and not necessarily a respective machine learning task) and defines transformations between the modality and the unified representation. For example, input modality neural network 102 a may be configured to map received machine learning model data inputs of a first modality, e.g., data inputs 110, to mapped data inputs of the unified representation space. Mapped data inputs of the unified representation space can vary in size.” [Examiner’s Note: The input modality neural network 102a-c would read on the claimed input models.]) wherein individual input models of the one or more input models comprise two or more sub-models, (Shazeer, [0056] “The input output mixer neural network may include one or more attention neural network layers configured to perform respective attention mechanisms, and one or more convolutional neural network layers. An example input output mixer neural network is illustrated and described in more detail below with reference to FIG. 4.” [0067] “In some implementations the multiple input modality neural networks 102 a-c and output modality neural networks 108 a-c may include audio modality networks. For example, the modality neural networks 102 a-c or 108 a-c may include neural networks that receive audio inputs in the form of a 1-dimensional waveform over time (or as a 2-dimensional spectrogram) and include a stack of the ConvRes blocks described with reference to the image input modality neural network, e.g., where an ith block has the form li=ConvRes(li−1, 2i). In this example, the spectral modality does not perform any striding along the frequency bin dimension, preserving full resolution in the spectral domain.”)
combining, with a common model of the modular autoencoder model, the processed inputs and reducing a dimensionality of the combined processed inputs to generate low dimensional data in a latent space, the low dimensional data in the latent space having a second level of resulting reduced dimensionality that is less than the first level; (Shazeer, [0049]-[0051] “The encoder neural network 104 is a neural network that is configured to process mapped data inputs from the unified representation space, e.g., mapped data input 112, to generate respective encoder data outputs in the unified representation space, e.g., encoder data output 114. Encoder data outputs are in the unified representation space. ... the encoder neural network 104 and decoder neural network 106 may include (i) one or more convolutional neural network layers, e.g., a stack of multiple convolutional layers with various types of connections between the layers, (ii) one or more attention neural network layers configured to perform respective attention mechanisms, (iii) one or more sparsely gated neural network layers.”) [Examiner’s Note: a shared encoder neural network receives the mapped outputs form all input modality networks to generate encoded unified representation. The autoencoder model consist of the encoder-decoder architecture reads on the common model.]
expanding, with the common model, the low dimensional data in the latent space into one or more expanded versions of the one or more inputs, the one or more expanded versions of the one or more inputs having increased dimensionality compared to the low dimensional data in the latent space, the one or more expanded versions of the one or more inputs suitable for generating one or more different outputs; (Shazeer, [0005] “a decoder neural network that is configured to process encoder data outputs to generate respective decoder data outputs from the unified representation space;” [0049]-[0051] “The decoder neural network 106 is a neural network, e.g., an autoregressive neural network, that is configured to process encoder data outputs from the unified representation space, e.g., encoder data output 114, to generate respective decoder data outputs from an output space, e.g., decoder data output 116. ... the encoder neural network 104 and decoder neural network 106 may include (i) one or more convolutional neural network layers, e.g., a stack of multiple convolutional layers with various types of connections between the layers, (ii) one or more attention neural network layers configured to perform respective attention mechanisms, (iii) one or more sparsely gated neural network layers.” Further see [0080].) [Examiner’s Note: The autoencoder model consist of the encoder-decoder architecture reads on the common model. The decoder network takes the encoded data to generate/reconstructed outputs. The encoded representation reads on the low dimensional latent space in a latent space and the decoder data outputs reads on the expanded version of the encoded data.]
using, with one or more output models of the modular autoencoder model, the one or more expanded versions of the one or more inputs to generate the one or more different outputs, ..., the one or more different outputs having the same or increased dimensionality compared to the expanded versions of the one or more inputs, (Shazeer, [0005] “a plurality of multiple output modality neural networks, wherein each output modality neural network corresponds to a different modality and is configured to map decoder data outputs from the unified representation space that correspond to received data inputs of the corresponding modality to data outputs of the corresponding modality.” [0047]-[0048] “each output modality neural network of the multiple output modality networks 108 a-c is configured to map data outputs of the unified representation space received from the decoder neural network, e.g., decoder data output 116, to mapped data outputs of one of the multiple modalities. That is, each output modality neural network is specific to a respective modality and defines transformations between the unified representation and the modality. For example, output modality neural network 108 c may be configured to map decoder data output 116 to mapped data outputs of a second modality, e.g., data output 118.” Further See [0081].) [Examiner’s Note: The output modality networks receive decoder outputs and amp them to approximated outputs. These outputs correspond to the “one or more different outputs” generated by the “output models.” The output modality network process the expanded/decoded signal from the decoder and produce outputs at the modality level.] wherein individual output models comprise two or more sub-models, (Shazeer, [0056] “The input output mixer neural network may include one or more attention neural network layers configured to perform respective attention mechanisms, and one or more convolutional neural network layers. An example input output mixer neural network is illustrated and described in more detail below with reference to FIG. 4.” [0062] “In some implementations the multiple input modality neural networks 102 a-c and output modality neural networks 108 a-c may include audio modality networks. For example, the modality neural networks 102 a-c or 108 a-c may include neural networks that receive audio inputs in the form of a 1-dimensional waveform over time (or as a 2-dimensional spectrogram) and include a stack of the ConvRes blocks described with reference to the image input modality neural network, e.g., where an ith block has the form li=ConvRes(li−1, 2i). In this example, the spectral modality does not perform any striding along the frequency bin dimension, preserving full resolution in the spectral domain.”) and
estimating, ..., one or more parameters based on the low dimensional data in the latent space and/or the one or more outputs. (Shazeer, [0003] “Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. ... Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters. Neural networks may be trained on machine learning tasks using training data to determine trained values of the layer parameters and may be used to perform machine learning tasks on neural network inputs.”)
While Shazeer describes the Multi-task multi-modal machine learning system that includes multiple input models, shared autoencoder, and output models. Additionally, Shazeer also teaches that each of the input and output modality neural network includes multiple layers and/or a stack of blocks (i.e., two or more sub-models). However, Shazeer is salient on whether the two or more input and output sub-models are associated with different portions of a sensing operation or a process. Furthermore, while Shazeer does not clearly define the dimensionality relationships of the processed data, these dimensionality relationship properties are obvious in the context of encoder/decoder processing architecture.
Shazeer does not appear to explicitly teach:
estimating, with a prediction model of the modular autoencoder model, one or more parameters based on the low dimensional data in the latent space and/or the one or more outputs.
the two or more sub-models of the input and output models are associated with different portions of a sensing operations or process.
However, Shazeer in view of Rothberg teaches the following:
combining, with a common model of the modular autoencoder model, the processed inputs and reducing a dimensionality of the combined processed inputs to generate low dimensional data in a latent space, the low dimensional data in the latent space having a second level of resulting reduced dimensionality that is less than the first level; (Rothberg, [0067)] “encoder 104 may be configured to receive input and output a latent representation (which may have a lower dimensionality than the dimensionality of the input data) ...” [0105] “The output of the first encoder (e.g., feature representation 206), the joint modality representation (e.g., knowledge base 230), and the first modality embedding (e.g., one of modality embeddings 232), may be used to generate input (e.g., feature representation 208) to a first decoder ...” Further see [0092].)
expanding, with the common model, the low dimensional data in the latent space into one or more expanded versions of the one or more inputs, the one or more expanded versions of the one or more inputs having increased dimensionality compared to the low dimensional data in the latent space, the one or more expanded versions of the one or more inputs suitable for generating one or more different outputs; (Rothberg, [0067] “encoder 104 may be configured to receive input and output a latent representation (which may have a lower dimensionality than the dimensionality of the input data) and the first decoder may be configured to reconstruct the input data from the latent representation. In some embodiments, the encoder and decoder may be part of an auto-encoder.” [0105] “The output of the first encoder (e.g., feature representation 206), the joint modality representation (e.g., knowledge base 230), and the first modality embedding (e.g., one of modality embeddings 232), may be used to generate input (e.g., feature representation 208) to a first decoder for the first modality (e.g., decoder 210). In turn, the output of the decoder 210 may be compared with the input provided to the first encoder ...”)
using, with one or more output models of the modular autoencoder model, the one or more expanded versions of the one or more inputs to generate the one or more different outputs, the one or more different outputs being approximations of the one or more inputs, the one or more different outputs having the same or increased dimensionality compared to the expanded versions of the one or more inputs; (Rothberg, [0068] “Accordingly, in some embodiments, during training, output of the statistical model 100 is compared to the input and the parameter values of the memory 105 are updated iteratively, based on a measure of distance between the input and the output, using stochastic gradient descent (with gradients calculated using backpropagation when the encoder and decoder are neural networks) or any other suitable training algorithm.” [0104]-[0105] “generating output using the respective decoders, comparing the input with the generated output, and updating the parameters values of the joint modality representation and/or the modality embeddings based on the difference between the input and output. ... In turn, the output of the decoder 210 may be compared with the input provided to the first encoder ...”) [Examiner’s Note: the decoder outputs are reconstructions/approximations of the original inputs, during training stage is used to minimize difference between input and output (reconstruction loss). This explicit reconstruction combination with the network outputs disclosed by Shazeer reads on the claimed “outputs being approximations of the one or more inputs.”] and
estimating, with a prediction model of the modular autoencoder model, one or more parameters based on the low dimensional data in the latent space and/or the one or more outputs. (Rothberg, [0082]-[0085] “As shown in FIG. 2B, the multi-modal statistical model 250 includes predictor 252 for prediction task 256 and task embeddings 254. ... These weighted feature representations may then be aggregated (e.g., as a weighted sum or product) via operation 260 to generate input for the predictor 252.” [0090] “the multi-modal statistical model 250 of FIG. 2B comprises encoder 204, encoder 214, knowledge base 230, modality embeddings 232, predictor 252, and task embeddings 254 and parameters of the components 230, 232, 252, and 254 may be estimated as part of process 300.” [0108] “act 308 may comprise estimating parameter values for one or more components of the multi-modal statistical model using a supervised learning technique. ... the parameters of a predictor (e.g., predictor 252 in the example of FIG. 2B) may be estimated at act 308. Additionally, in some embodiments, parameters of one or more task embeddings (e.g., one or more of task embeddings 254) may be estimated at act 308. ... the parameter values estimated as part of act 306 may be estimated using supervised learning based on the labeled training data accessed at act 304.”) [Examiner’s Note: Rothberg teaches a predictor component that is separate from the encoder-decoder, it takes as input latent/feature representations form the encoded inputs to estimate parameters for prediction tasks.]
Accordingly, at the effective filing date, it would have been prima facie obvious to one ordinarily skilled in the art of machine learning to modify the combination of Shazeer and Rothberg to incorporate the techniques for performing a prediction task using a multi-modal statistical model as taught by Rothberg. One would have been motivated to make such a combination in order to enable generating multi-modal statistical models for any suitable number of data modalities, but also improves the computer technology used to train and deploy such machine learning systems (Rothberg [0051]).
As outlined above, while Shazeer and Rothberg are salient on whether the input and output sub-models are associated with different portions of a sensing operation or process. However, Schlake, in combination with Shazeer and Rothberg, teaches the limitation:
wherein the two or more sub-models of individual input models associated with different portions of a sensing operation ... or ... process, and wherein the two or more sub-models of individual output models associated with different portions of a sensing operation ... or ... process. (Schlake, [0041] “a dynamic predictive model for an industrial plant that comprises a plurality of sub-models, wherein the inputs of the model are distributed across the inputs of the sub-models, the outputs of the model are compiled from the outputs of the sub-models, and at least one output of one sub-model is processed into at least one input of one other sub-model.” [0075] “FIG. 3 illustrates the model 10 generated in this manner. The model 10 consists of two sub-modules 13 a and 13 b that correspond to sub-units 1 a and 1 b. The first sub-module 13 a gets the current values of process variables 41-44 and 48, as well as the set-points 51-54, that pertain to vessel V1 as inputs. Based on these inputs, the sub-model 13 a predicts how the process variables 41-44 and 48 will evolve to future values 41′-44′ and 48′.” [0015] “The plant also has a plurality of sensors. Each such sensor measures at least one process variable of one of the physical process, and/or of the plant as a whole. For example, there may be a sensor measuring the temperature in a vessel, and/or a sensor for the turbidity of a mixture inside the vessel that is a measure of how homogeneous the mixture is.” [0077] “Second sub-model 13 b predicts future values 45′-47′ and 49′ for process variables 45-47 and 49 based on their current values, the set-points 55-57 that directly act upon them, and the future values 41′-43′ of relevant process variables 41-43, as obtained from first sub-model 13 a.”) [Examiner’s Note: Schlake teaches a plurality of input and output sub-models processing input data and outputting predicted feature values, and further describes the models/sub-models as associated with different portions of measurement operations and/or physical processes. See Schlake e.g., [0041], [0063], [0069] and [0075]-[0077].)
Accordingly, at the effective filing date, it would have been prima facie obvious to one ordinarily skilled in the art to modify the combination of Shazeer, Rothberg, and Schlake to incorporate the method for generating a dynamic model as taught by Schlake. One would have been motivated to make such a combination in order to generate a dynamic model needed for the model predictive control of the overall plant makes the creation of the model much more efficient and much more transparent (Schlake [0026]).
Regarding Previously Presented Claim 20, the combination of Shazeer, Rothberg, and Schlake teaches the elements of claim 16 as outlined above, and further teaches:
determining a quantity of the one or more input models, and/or a quantity of the one or more output models, based on process physics differences in different parts of a manufacturing process and/or a sensing operation. (Schlake, [0018]-[0019] “a division of a representation of the plant into sub-units is obtained. ... If the plant is a modular industrial plant, then the already existing division into the physical process modules that make up the plant may be used.” [0021] “If no division into physical process modules exists, the layout of the plant may be actively searched for suitable sub-units. In other words, the representation of the plant may be actively divided into sub-units. Several strategies for doing this, which may be used individually or in arbitrary combinations, are detailed below.” [0028] “Moreover, the model may be much more easily adapted to any changes in the plant. For example, the trend is going from monolithic plants to modular plants where modules may be added and removed, or brought on-line and off-line, as needed. Whenever a change of this type happens, the overall model for the plant may be adapted in a straight-forward manner by simply adding and removing identical copies of one and the same sub-model in the right place.” Further See [0038].)
Regarding Currently Amended Claim 21, the combination of Shazeer, Rothberg, and Schlake teaches the elements of claim 16 as outlined above, and further teaches:
Shazeer further teaches: wherein a quantity of input models is different than the quantity of output models. (Shazeer, [0044] “For convenience, the example multi task multi modal machine learning model 100 is shown as including three input modality networks and three output modality neural networks. However, in some implementations the number of input or output modality neural networks may be less or more, in addition the number of input modality neural networks may not equal the number of output modality neural networks.”)
Regarding Previously Presented Claim 22, the combination of Shazeer, Rothberg, and Schlake teaches the elements of claim 16 as outlined above, and further teaches:
Wherein:
the common model comprises encoder-decoder architecture and/or variational encoder-decoder architecture; (Shazeer, [0043] “The multi task multi modal machine learning model 100 includes multiple input modality neural networks 102 a-102 c, an encoder neural network 104, a decoder neural network 106, and multiple output modality neural networks 108 a-108 c.”)
processing the one or more inputs to the first level of dimensionality, and reducing the dimensionality of the combined processed inputs comprises encoding; (Shazeer, [0045] “Each input modality neural network of the multiple input modality networks 102 a-c is configured to map received machine learning model data inputs of one of multiple machine learning domains or modalities to mapped data inputs of a unified representation space.” [0049] “The encoder neural network 104 is a neural network that is configured to process mapped data inputs from the unified representation space, e.g., mapped data input 112, to generate respective encoder data outputs in the unified representation space, e.g., encoder data output 114. Encoder data outputs are in the unified representation space.”) and reducing the dimensionality of the combined processed inputs comprises encoding; (Rothberg, [0127]-[0128] “In some embodiments, the input data may be converted or otherwise pre-processed into a representation suitable for providing to the encoder for the first modality. For example, categorical data may be one-hot encoded prior to being provided to the encoder for the first modality. As another example, image data may be resized prior to being provided to the encoder for the first modality. ... Next, process 400 proceeds to act 406, where the input data is provided as input to the first encoder, which generates a first feature vector as output. FIG. 2B, input 202 for modality “A”, is provided as input to the encoder 204 for modality “A”, and the encoder 204 produces a first feature vector (e.g., feature representation 206 as output).” [0092] “The first encoder may be configured to receive, as input, data having the first modality and output a latent representation (which may have a lower dimensionality than the dimensionality of the input data) and the first decoder may be configured to reconstruct the input data from the latent representation.”) and
expanding the low dimensional data in the latent space into the one or more expanded versions of the one or more inputs comprises decoding. (Shazeer, [0050] “The decoder neural network 106 is a neural network, e.g., an autoregressive neural network, that is configured to process encoder data outputs from the unified representation space, e.g., encoder data output 114, to generate respective decoder data outputs from an output space, e.g., decoder data output 116.”)
Regarding Previously Presented Claim 23, the combination of Shazeer, Rothberg, and Schlake teaches the elements of claim 16 as outlined above, and further teaches:
training the modular autoencoder model by comparing the one or more different outputs to corresponding inputs, and adjusting a parameterization of the one or more input models, the common model, and/or the one or more output models to reduce or minimize a difference between an output and a corresponding input. (Shazeer, [0063] “The multi task multi modal machine learning model 100 can be trained to perform different machine learning tasks from different machine learning domains or modalities using training data. ... The training data may be used to adjust the input modality neural networks 102 a-c, encoder neural network 104, decoder neural network 106, and output modality neural networks 108 a-c weights from initial values to trained values, e.g., by processing the training examples and adjusting the neural network weights to minimize a corresponding loss function.” Further see Rothberg [0068] and [0104]-[0105].)
Regarding Currently Amended Claim 27,
The claim recites substantially similar limitations as corresponding claim 16 and is rejected for similar reasons as claim 16 using similar teachings and rationale. Claim 16 is directed to a method, and claim 27 is directed to a non-transitory computer readable medium.
The combination of Shazeer, Rothberg, and Schlake also discloses methods, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods. A system of one or more computers can be configured to perform particular operations or actions by virtue of software, firmware, hardware, or any combination thereof installed on the system that in operation may cause the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by virtue of including instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.
Regarding Currently Amended Claim 29,
The claim recites substantially similar limitations as corresponding claim 16 and is rejected for similar reasons as claim 16 using similar teachings and rationale. Claim 16 is directed to a method, and claim 29 is directed to a system.
The combination of Shazeer, Rothberg, and Schlake also discloses methods, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods. A system of one or more computers can be configured to perform particular operations or actions by virtue of software, firmware, hardware, or any combination thereof installed on the system that in operation may cause the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by virtue of including instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.
Regarding Currently Amended Claim 30,
The claim recites substantially similar limitations as corresponding claim 16 and is rejected for similar reasons as claim 16 using similar teachings and rationale. Claim 16 is directed to a method and claim 30 is directed to a non-transitory computer readable medium. The primary difference is that claim 30 broadly recites a “machine-learning model” rather than a “modular autoencoder model,” and generically labels the components as first, second, third, and fourth models instead of input, common, output, and prediction models. Accordingly, the same teaching and rationale applied to claim 16 are equally applicable to the broader scope of claim 30.
Claim(s) 18 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of Shazeer, Rothberg, and Schlake as described above, and further in view of Kim et al., (Pub. No.: US 20140291678 A1).
Regarding Currently Amended Claim 18, the combination of Shazeer, Rothberg, and Schlake teaches the elements of claim 16 as outlined above.
While Shazeer in view of Rothberg teaches that an individual output model comprises the two or more sub-models, Shazeer in view of Rothberg does not appear to explicitly teach:
wherein the two or more sub-models of the individual output model comprise a sensor model and a stack model for a semiconductor sensor operation.
However, Kim, in combination with Shazeer, Rothberg, and Schlake, teaches the limitation:
wherein the two or more sub-models of the individual output model comprise a sensor model and a stack model for a semiconductor sensor operation. (Kim, [0041]-[0042] “embodiments of the present invention provide a semiconductor sensor reliability system and method. Specifically, the present invention provides in-situ positioning of a reliability sensor (hereinafter sensors) within each functional block, as well as at critical locations, of a semiconductor system. ... In general, the sensor models a behavior (e.g., aging process) of the location (e.g., functional block) in which it is positioned and comprises a plurality of stages connected as a network and a self-digitizer. ... semiconductor 10 can include one or more wafer/layer 12A-N, each of which has various functional blocks 14.” [0045]-[0046] “sensors 42A-N are configured to model the behavior of the functional block in which they are positioned. ... Sensors 42A-N are positioned at locations within circuit block 40 subject to semiconductor system manufacturing process variation. Sensors 42A-N collect, process, and store sensed data about both predicted and unpredicted semiconductor degradation.”) [Examiner’s Note: Accordingly to MPEP § 2103 (I), the recitation of “for a semiconductor sensor operation” is non-limiting field-of-use language.]
Therefore, it would have been prima facie obvious to one of ordinary skill in the art, before the effective date of the claimed invention, having the combination of Shazeer, Rothberg, Schlake, and Kim to incorporate the method for sensing semiconductor reliability operation as taught by Kim. One would have been motivated to make such a combination in order to provide higher-level reliability across the system hierarchy, as well as a need to identify and isolate the most likely runtime component failure to avoid catastrophic system failure. Doing so would reduce semiconductor system degradation in real-time (Kim [0041]).
Claim(s) 19, 25, 26, and 28 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Shazeer, Rothberg, and Schlake as described above, and further in view of et al., Sha et al., (Pub. No.: US 20200151538 A1).
Regarding Previously Presented Claim 19, the combination of Shazeer, Rothberg, and Schlake teaches the elements of claim 16 as outlined above, and further teaches:
wherein the one or more input models, the common model, and the one or more output models are separate from each other and correspond to process physics differences in different parts of ... operation such that each of the one or more input models, such that each of the one or more input models, the common model, and/or the one or more output models are trained together and/or separately, but individually configured based on the process physics for a corresponding part of the ... operation, apart from other models in the modular autoencoder model. (Shazeer, [0043] “The multi task multi modal machine learning model 100 includes multiple input modality neural networks 102 a-102 c, an encoder neural network 104, a decoder neural network 106, and multiple output modality neural networks 108 a-108 c.” [0027] “ The model can be trained to perform the multiple machine learning tasks jointly, thus simplifying and improving the efficiency of the training process. In addition, by training the model jointly, in some cases less training data may be required to train the model (to achieve the same performance) compared to when separate training processes are performed for separate machine learning tasks.” [0063] “The multi task multi modal machine learning model 100 can be trained to perform different machine learning tasks from different machine learning domains or modalities using training data. The multi task multi modal machine learning model 100 can be trained jointly to perform different machine learning tasks from different machine learning domains, so that the multi task multi modal machine learning model 100 simultaneously learns multiple machine learning tasks from different machine learning domains.”)
While Shazeer in view of Rothberg describes the separate components of the machine learning architecture performing different machine learning tasks from different domains while being jointly trained. Shazeer in view of Rothberg does not define the domains tasks correspond to process physics for a corresponding part of the manufacturing process and/or sensing operation.
However, it would have been obvious in view of Sha. Hereinafter, Sha, in combination with Shazeer, Rothberg, and Schlake, teaches:
individually configured based on the process physics for a corresponding part of the manufacturing process and/or sensing operation, apart from other models in the modular autoencoder model. (Sha, [0026] “As such the technical solutions are rooted in and/or tied to computer technology in order to overcome a problem specifically arising in the realm of computers, specifically manufacturing semiconductors, such as integrated chips.” [00311] “In one or more examples, the synthesis controller 116 uses lithography, such as photolithography, for manufacturing the physical implementation, that is chip or semiconductor 120. The photo-lithography process 220 in semiconductor fabrication consists in duplicating desired mask shapes 210 as best as possible onto a semiconductor wafer 120.” [0047] “Referring to the flowchart of FIG. 5, training the DNN 610 can include accessing multiple chip layouts (511) and generating aerial images 330 for each of the chip layouts (512).” [0061] “FIG. 8 illustrates a flowchart of another method of automatic feature extraction from aerial images for test pattern sampling and pattern coverage inspection according to one or more embodiments of the present invention. In this case, training the neural network (510) includes using the simulated aerial images 330 of existing chip layouts (511, 512) as input for unsupervised training. The method includes generating a set of codings based on the aerial images 330 that are obtained from aerial image generation 410 using an artificial neural network in an unsupervised manner, such as using an autoencoder (613). An autoencoder learns to compress data (“encode”) from the input layer into a short code, and then decompress that code (“decode”) in the output layer so that the output closely matches the original input data.”)
Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Shazeer, Rothberg, Schlake, and Sha, to incorporate the aerial image generation system for a physical synthesis system used to synthesize a physical design such as a semiconductor chip as taught by sha. One would have been motivated to make such a combination in order to improve the chip layout early in the fabrication process so that the chip can be reworked to eliminate the defects from the end product (sha [0065]).
Regarding Previously Presented Claim 25, the combination of Shazeer, Rothberg, and Schlake teaches the elements of claim 16 as outlined above.
the one or more input models and/or the one or more output models comprise dense feed-forward layers, convolutional layers, and/or residual network architecture of the modular autoencoder model; the common model comprises feed forward and/or residual layers; (Shazeer, [0067]-[0068] “The stack of depth wise separable convolutional neural network layers includes two skip- connections 220, 222 between the stack input 210 and the outputs of (i) the second convolutional step 204 and (ii) the fourth convolutional step 208. The stack of depth wise separable convolutional neural network layers also includes two residual connections 214 and 216. ... The example encoder neural network 104 includes a residual connection 306 between the data input 302 and a timing signal 304. After the timing signal 304 has been added to the input 302, the combined input is provided to the convolutional module 308 for processing. The convolutional module 308 includes multiple convolutional neural network layers, e.g., depth wise separable convolutional neural network layers, as described above with reference to FIGS. 1 and 2. The convolutional module 308 generates as output a convolutional output, e.g., convolutional output 322.”) and
the prediction model comprises feed forward and/or residual layers. (Rothberg, [0083]-[0084] “In the embodiment illustrated in FIG. 2B, it is assumed that the encoders 204 and 214, the decoders 210 and 220, the knowledge base 230, and the modality embeddings 232 have been previously trained as shown by the fill pattern having diagonal lines extending downward from left to right, and that the predictor 252 and task embeddings 254, ... the predictor 252 may comprise a linear model (e.g., a linear regression model), a generalized linear model (e.g., logistic regression, probit regression), a neural network or other non-linear regression model, ...”)
the combination of Shazeer, Rothberg, and Schlake does not appear to explicitly teach:
the one or more parameters are semiconductor manufacturing process parameters;
However, Sha, in combination with Shazeer, Rothberg, and Schlake, teaches:
the one or more parameters are semiconductor manufacturing process parameters (Sha, [0024] “Alternatively, or in addition, aerial image parameters of the electronic circuit can be categorized using predetermined features to be inspected from the aerial images. In one or more examples, manual selection of image parameters for the aerial images are provided for analyzing the pattern coverage. In one or more examples, such aerial image parameters are limited to one-dimensional (1D) cutlines. As can be understood by a person skilled in the art, such existing solutions operate using engineering judgment, where a user provides particular features to inspect in aerial images.” [0026] “Described herein are technical solutions for test pattern sampling and pattern coverage inspection, which are used for semiconductor manufacturing, particularly using photolithography.”)
The same motivation that was utilized for combining Shazeer, Rothberg, Schlake, and Sha as set forth in claim 19 is equally applicable to claim 25.
Regarding Previously Presented Claim 26,
the combination of Shazeer, Rothberg, and Schlake teaches the elements of claim 16 as outlined above:
While the combination of Shazeer, Rothberg, and Schlake teaches the modular autoencoder and parameter estimation/prediction, Shazeer in view of Rothberg does not appear to explicitly recite:
generating, with one or more auxiliary models of the modular autoencoder model, labels for at least some of the low dimensional data in the latent space, the labels configured to be used by the prediction model for estimations.
However, Sha, in combination with Shazeer, Rothberg, and Schlake, teaches the limitation:
generating, with one or more auxiliary models of the modular autoencoder model, labels for at least some of the low dimensional data in the latent space, the labels configured to be used by the prediction model for estimations. (Sha, [0047]- [0050] “Referring to the flowchart of FIG. 5, training the DNN 610 can include accessing multiple chip layouts (511) and generating aerial images 330 for each of the chip layouts (512). [0048] In one or more examples, the aerial images 330 are generated using the aerial image generation 410. The method further includes creating training data by identifying classification labels or types of the aerial images 330 (513). In one or more examples, the aerial images 330 that are obtained by the simulation are labeled with the corresponding types of aerial image. The DNN 610 is trained using the created training data (514). ... In summary, the DNN 610 processes the records that include aerial images 330 and corresponding labels in the training data one at a time, using the weights and functions in the hidden layers 614, then compares the resulting outputs against the desired outputs.)
The same motivation that was utilized for combining Shazeer, Rothberg, Schlake, and Sha as set forth in claim 19 is equally applicable to claim 26.
Regarding Previously Presented Claim 28,
The claim recites substantially similar limitations as corresponding claim 26 and is rejected for similar reasons as claim 26 using similar teachings and rationale.
Claim(s) 24 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of Shazeer, Rothberg, and Schlake as described above, and further in view of et al., Sjolund et al., (Pub. No.: US 20190332900 A1).
Regarding Previously Presented Claim 24,
the combination of Shazeer, Rothberg, and Schlake teaches the elements of claim 16 as outlined above, and further teaches:
wherein the common model comprises an encoder and a decoder, the method further comprising training the modular autoencoder model by: ... providing the decoder signal to the encoder to generate new low dimensional data; comparing the new low dimensional data to the low dimensional data; and adjusting one or more components of the modular autoencoder model based on the comparison to reduce or minimize a difference between the new low dimensional data and the low dimensional data. (Rothberg, [0092] “the first trained statistical model may include an auto-encoder and may comprise a first encoder and a first decoder, each having a respective set of parameters, which may be accessed at act 302. The first encoder may be configured to receive, as input, data having the first modality and output a latent representation (which may have a lower dimensionality than the dimensionality of the input data) and the first decoder may be configured to reconstruct the input data from the latent representation.” [0104]-[0105]) “The iterative learning algorithm may involve providing at least some of the unlabeled training data as input to the encoders of the multi-modal statistical model, generating output using the respective decoders, comparing the input with the generated output, and updating the parameters values of the joint modality representation and/or the modality embeddings based on the difference between the input and output. For example, in some embodiments, training data of a first modality may be provided as input to a first encoder for the first modality (e.g., encoder 204). The output of the first encoder (e.g., feature representation 206), the joint modality representation (e.g., knowledge base 230), and the first modality embedding (e.g., one of modality embeddings 232), may be used to generate input (e.g., feature representation 208) to a first decoder for the first modality (e.g., decoder 210). In turn, the output of the decoder 210 may be compared with the input provided to the first encoder and at least some of the parameter values of the joint modality representation and/or the first modality embedding may be updated based on the difference between the input to the first encoder and the output of the first decoder.”)
As outlined above, while the combination of Shazeer, Rothberg, and Schlake describes the training process of the components of the autoencoder modular including iteratively updating the models’ parameters during training based on the loss function and using backpropagation. The combination of Shazeer, Rothberg, and Schlake does not appear to explicitly teach:
applying variation to the low dimensional data in the latent space such that the common model decodes a relatively more continuous latent space to generate a decoder signal;
recursively providing the decoder signal to the encoder to generate new low dimensional data;
However, Sjolund, in combination with Shazeer, Rothberg, and Schlake, teaches the following:
applying variation to the low dimensional data in the latent space such that the common model decodes a relatively more continuous latent space to generate a decoder signal; recursively providing the decoder signal to the encoder to generate new low dimensional data; comparing the new low dimensional data to the low dimensional data; and adjusting one or more components of the modular autoencoder model based on the comparison to reduce or minimize a difference between the new low dimensional data and the low dimensional data. (Sjolund, [0081] “As a result, it may be necessary to enforce some type of regularity (e.g. smoothness) of the latent representation during encoding, which can be done using, for instance, a variational autoencoder. It may also be preferable to have a latent representation which has the same spatial coordinates as the images so that the variables can be interpreted locally. This technique is introduced using Gaussian random fields in various examples below.” [0106] “As an overview of auto-encoders and variational auto-encoders, consider that auto-encoders, work by training two networks simultaneously. The encoder E takes data as input and maps it to a latent variable z, while the decoder D takes the latent variable z and reconstructs the input from it. The networks are trained by minimizing the reconstruction loss, typically taken to be the expected mean squared error...” [0137]-[0138] “This first approach for variational autoencoding is depicted in FIGS. 7 and 8, whereas this second approach for encoding and decoding is depicted in FIGS. 9 and 10. Each of these approaches in the following examples may be usable to learn the latent representation, but with different latent space dimensions and network architecture. FIG. 7 illustrates a data flow diagram of a variational autoencoder, employed in an exemplary encoding process for generating a latent representation of an imaging input. The variational autoencoder is composed of an encoder and a decoder. As shown, the variational autoencoder includes three blocks of convolutional layers in both the encoder (710) and decoder (740) portions of the network. Input data is encoded to a latent vector in the latent space. It is then decoded to reconstruct the input data. As an example, a latent vector is sampled with the mean and standard deviation by reparameterization. As shown, the variational autoencoder involves use of convolution operations (710), and fully connected operations (720A, 720B) that involve sampling (730). The decoding portion of the variational autoencoder further involves transpose convolution operations (740) that can recreate the original data input. FIG. 8 more specifically illustrates these and more detailed data processing operations performed by a neural network of a variational autoencoder, in connection with input convolution processing operations 810, result generation operations 820, and output transpose convolution processing operations 830 (e.g., producing a value useful for reconstruction).”)
Accordingly, it would have been prima facie obvious to one having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Shazeer, Rothberg, Schlake, and Sjolund, to incorporate the method for generating a modality-agnostic image processing model using variational autoencoder models as taught by Sjolund. One would have been motivated to make such a combination in order to improve the performance of the neural network by leveraging information from the unified representation that has been learned on other tasks (Sjolund [0037]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
(Pub. No.: US 20170200265 A1) – “Kris Bhaskar” relates to “Generating simulated output for a specimen.”
[0078] “The component(s), e.g., component(s) 100 shown in FIG. 1, executed by the computer subsystem(s), e.g., computer subsystem 36 and/or computer subsystem(s) 102, include learning based model 104. The learning based model is configured for mapping a triangular relationship between optical images, electron beam images, and design data. For example, as shown in FIG. 2, the learning based model may be configured for mapping a triangular relationship between three different spaces: design 200, electron beam 202, and optical 204. The embodiments described herein, therefore, provide a generalized representation mapping system between optical, electron beam, and design (e.g., CAD) for semiconductor wafers and masks for inspection, metrology, and other use cases. The embodiments described herein, therefore, also provide a systematic approach to perform transformations between different observable representations of specimens such as wafers and reticles in semiconductor process control applications. The representations include design data, EDA design, (with 0.1 nm placement accuracy), scanning electron microscopy (SEM) review/inspection/etc. (with 1 nm placement accuracy), and patch images from optical tools such as inspection (with 10 nm to 100 nm placement accuracy).”
(Pub. No.: US 20200294224 A1) – “Ohad SHAUBI” relates to “Method of deep learning-based examination of a semiconductor specimen and system thereof.”
[0011]-[0012] “By way of non-limiting example, the one or more first modalities can differ from the one or more second modalities by at least one of: examination tool, channel of the same examination tool, operational parameters of the same examination tool and/or channel, layers of the semiconductor specimen corresponding to respective FP images, nature of obtaining the FP images and deriving techniques applied to the captured images. In accordance with other aspects of the presently disclosed subject matter, there is provided a system usable for examination of a semiconductor specimen in accordance with the method above.”
[0060]-[0061] “DNN module 114 can comprise a plurality of input subnetworks (denoted 302-1-302-3), each given input subnetwork configured to process a certain type of FP input data (denoted 301-1-301-3) specified for the given subnetwork. The architecture of a given input subnetwork can correspond to respectively specified type(s) of input data or, alternatively, can be independent of the type of input data. The input subnetworks can be connected to an aggregation subnetwork 305 that is further connected to an output subnetwork 306 configured to output application-specific examination-related data. Optionally, at least part of the input subnetworks can be directly connected to the aggregation subnetwork 305 or output subnetwork 306. Optionally, aggregation and output subnetworks can be organized in a single subnetwork.”
NPL: Yu, Jianbo. "Enhanced stacked denoising autoencoder-based feature learning for recognition of wafer map defects." (2019).
[Abstract] “In semiconductor manufacturing systems, defects on wafer maps tend to cluster and then these spatial patterns provide important process information for helping operators in finding out root-causes of abnormal processes. Promptly recognizing wafer map defects is an effective way to increase manufacturing process stability and then to improve yields. Deep learning has been widely applied and obtained many successes in image and visual analysis. This paper proposes an effective deep learning method, enhanced stacked denoising autoencoder (ESDAE) with manifold regularization for wafer map pattern recognition (WMPR) in manufacturing processes. This study will concentrate on developing a deep learning model to learn effective discriminative features from wafer maps through a deep network architecture for WMPR improvement.”
Fig. 7. The application procedure of the proposed wafer map defect detection and recognition system.
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SADIK ALSHAHARI whose telephone number is (703)756-4749. The examiner can normally be reached Monday Friday, 9 A.M - 6 P.M. ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li Zhen can be reached on (571) 272-3768. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/S.A.A./Examiner, Art Unit 2121
/Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121