DETAILED ACTION
Claims 1-20 are presented for examination.
Claims 1, 11, and 15-20 have been amended.
This office action is in response to the amendment submitted on 06-July-2026.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments – 35 USC 101
In the Applicant/Arguments Remarks, Applicant argues the amended claims have overcome the rejection under 35 USC 101. The argument have been fully considered and are persuasive. The rejection has been withdrawn.
Response to Arguments – 35 USC 103
Applicant’s arguments with respect to the 103 rejections have been considered but are moot in view of the new ground(s) of rejection provided below.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Ypma et al. (US20160246185A1) in view of Smola et al. (A tutorial on support vector regression).
Regarding Claim 1, Ypma teaches obtaining measurement data relating to a performance parameter for at least a portion of the substrate physically processed in a lithographic process performed using a lithographic apparatus; and fitting, by a hardware computer system the measurement data to a model (Ypma, [0100-0106] “Standard alignment models have six parameters (effectively three per direction X & Y) and in addition there are more advanced alignment models. On the other hand, for the most demanding processes currently in use and under development, to achieve the desired overlay performance requires more detailed corrections of the wafer grid. While standard models might use fewer than ten parameters, advanced alignment models typically use more than 15 parameters, or more than 30 parameters. Examples of advanced models are higher order wafer alignment (HOWA) models, zone-alignment (ZA) and radial basis function (RBF) based alignment models. HOWA is a published technique based on second, third and higher order polynomial functions. Zone alignment is described for example in Huang et al, “Overlay improvement by zone alignment strategy”, Proc. SPIE 6922, 69221G (2008). RBF modeling is described in published patent application US 2012/0218533. Different versions and extensions of these advanced models can be devised. The advanced models generate a complex description of the wafer grid that is corrected for, during the exposure of the target layer. RBF and latest versions of HOWA provide particularly complex descriptions based on tens of parameters. This implies a great many measurements are required to obtain a wafer grid with sufficient detail. FIGS. 3 & 4 illustrate the form of alignment information that can be used to correct for wafer grid distortion as measured by the alignment sensor AL on alignment marks (targets) 400 in a previous layer on wafer (substrate) W. Each target has a nominal position, defined usually in relation to a regular, rectangular grid 402 with axes X and Y. Measurements of the real position 404 of each target reveal deviations from the nominal grid. The alignment marks may be provided within device areas of the substrate, and/or they may be provided in so-called “scribe lane” areas between device areas….FIG. 6 illustrates the collection of object data in the embodiment of FIG. 2, during performance of a patterning operation by litho tool 200. As already described, measurement station 202 of the litho tool 200 uses alignment sensors AS to measure positional deviations 404 of individual marks, spatially distributed across the substrate W. As mentioned above with reference to FIGS. 4 and 5, alignment models used in lithography can be of low order or high order (advanced) type. In the present example, a higher order correction module 602 calculates an alignment model 406 according to the HOWA method, mentioned above. This alignment model, is used at the exposure station 204 to apply a pattern to substrate. For the purposes of the PCA apparatus 250, we propose to use residual data, rather than the deviations 404 as measured by the alignment sensors. This is because, in a modern high-performance lithography apparatus, most of the measured deviation will be compensated by the alignment model. Therefore performance improvements and diagnostic methods concentrate on detecting and eliminating the small deviations that remain uncorrected by the model. One option would therefore be to use as the object data, residual deviations that are not corrected by the HOWA model. In the present example, however, the designers have made a different choice. In the present embodiment it is chosen to use residuals after subtraction of only a low order correction, so that high order deviations, even though some of them may be compensated by the HOWA model in operation of the litho tool, are nevertheless revealed in the object data. Leaving high order deviations in the residuals may facilitate diagnostic interpretation of the resulting component vectors. The HOWA model corrects low order and high order deviations simultaneously. To make a low order correction accessible for calculation of residuals, in the present embodiment, a traditional 6-parameter (6 PAR) model 402′ is separately calculated by a unit 604. The 6 PAR calculating unit calculating unit 604 may be provided already as part of the litho tool management software, or it may be provided specially as part of the diagnostic apparatus. The low order model 402′ is subtracted from the measured deviations 404 to obtain residual deviations 404′. These residual variations 404′ are collected as the object data for use in the PCA apparatus 250. In embodiments using a different higher order model, or no higher order model at all, the 6PAR calculation unit 604 may be provided already, and the residuals 404′ may be calculated already. For example, the RBF model described in the prior art mentioned above, is generally applied to correct only the higher order deviations, after low order deviations have been corrected by a low order model such as the 6PAR model.” And [0123] “FIG. 10 illustrates how projections onto various ones of the component vector axes can be used to identify product units of interest. FIG. 10(a) illustrates the projection onto an axis represented by the first component vector PC1. We see how the vector AL(i) of each product unit in the multidimensional space is reduced a single-dimensional value, namely the coefficient c(PC1). Comparing roughly with the distributions seen in three dimensions in FIG. 9, the clusters 900 to 904 are recognizable, as well as the outlier 906. Applying a statistical threshold to this distribution allows outliers such as point 906 to be identified. For example, 910 in the drawing indicates a Gaussian distribution curve that has been fitted to the data, with its mean centered on the mean value of the coefficient c(PC1). Statistical significance thresholds can be established, as indicated at 912, 914. Point 906 and point 916 lie outside these thresholds, and are identified as being of interest.”)
the fitting including determining fitting parameters of the model (Ypma, [0100-0106])
controlling or configuring the lithographic process based on the fitted model, wherein the controlling or configuring comprises: generating, based on the fitted model, a set-point correction for a controllable parameter of the lithographic apparatus ([0032-0034] “The method may further comprise the step of generating one or more sets of correction data for use in controlling the industrial process when performed on further product units. The correction data may be applied for example as alignment corrections in a future lithographic step to correct distortions of the products introduced by a chemical and physical processing steps. The corrections may be applied selectively based on context criteria. The corrections may be applied so as to correct some of the identified component vectors and not others. Where the industrial process comprises a mixture of lithographic pattering operations and physical and/or chemical operations, the diagnostic apparatus may be programmed to generate said correction data for applying corrections in a lithographic pattering operation. The apparatus may further comprise a controller arranged to control a lithographic apparatus by applying corrections based on the extracted diagnostic information.”)
outputting the set-point correction to the lithographic apparatus for adjusting the controllable parameter and thereby controlling the lithographic process ([0224] “The steps of the methods described above can be automated within any general purpose data processing hardware (computer), so long as it has access to the object data and, if desired performance data and context data. The apparatus may be integrated with existing processors such as the lithography apparatus control unit LACU shown in FIG. 1 or an overall process control system. The hardware can be remote from the processing apparatus, even being located in a different country. Components of a suitable data processing apparatus (DPA) are shown in FIG. 22. The apparatus may be arranged for loading a computer program product comprising computer executable code. This may enable the computer assembly, when the computer program product is downloaded, to implement the functions of the PCA apparatus and/or RCA apparatus as described above.” )
However, Ypma is not relied on for:
performing support vector machines regression,
by minimizing a complexity metric applied to the fitting parameters
while not allowing a deviation between the measurement data and the fitted model to exceed a threshold value
Smola teaches performing support vector machines regression, (Pg. 1, “The purpose of this paper is twofold. It should serve as a selfcontained introduction to Support Vector regression for readers new to this rapidly developing field of research.1 On the other hand, it attempts to give an overview of recent developments in the field. To this end, we decided to organize the essay as follows. We start by giving a brief overview of the basic techniques in Sections 1, 2 and 3, plus a short summary with a number of figures and diagrams in Section 4. Section 5 reviews current algorithmic techniques used for actually implementing SV machines. This may be of most interest for practitioners. The following section covers more advanced topics such as extensions of the basic SV algorithm, connections between SV machines and regularization and briefly mentions methods for carrying out model selection. We conclude with a discussion of open questions and problems and current directions of SV research. Most of the results presented in this review paper already have been published elsewhere, but the comprehensive presentations and some details are new. “ and Pg. 56, “This is the so-called Support Vector expansion, i.e. w can be completely described as a linear combination of the training patterns xi. In a sense, the complexity of a function’s representation by SVs is independent of the dimensionality of the input space X, and depends only on the number of SVs.”
PNG
media_image1.png
100
775
media_image1.png
Greyscale
62, “Lagrange function: The Lagrange function is given by the primal objective function minus the sum of all products between constraints and corresponding Lagrange multipliers (cf. e.g. Fletcher 1989, Bertsekas 1995). Optimization can be seen as minimzation of the Lagrangian wrt. the primal variables and simultaneous maximization wrt. the Lagrange multipliers, i.e. dual variables. It has a saddle point at the solution.” EN: Eq (11), w is the model parameters, which is a linear combination of the design matrix x(i) and the optimized values using the LaGrange multipliers.).
by minimizing a complexity metric applied to the fitting parameters (Pg. 55, Eqs 2 and 3 show the minimization of the complexity metric w applied to the parameters being fitted.)
while not allowing a deviation between the measurement data and the fitted model to exceed a threshold value (Pg. 2, Eq(3) and “Analogously to the “soft margin” loss function (Bennett and Mangasarian 1992) which was used in SV machines by Cortes and Vapnik (1995), one can introduce slack variables ξi, ξ∗i to cope with otherwise infeasible constraints of the optimization problem (2). Hence we arrive at the formulation stated in Vapnik (1995)… The constant C > 0 determines the trade-off between the flatness of f and the amount up to which deviations larger than
ε are tolerated. This corresponds to dealing with a so calledε-insensitive loss function |ξ |ε described by" EN: C is the coefficient for weighting the slack variables).
PNG
media_image2.png
231
786
media_image2.png
Greyscale
Ypma, and Smola are analogous art because they are from the same field of endeavor in modeling and error minimization. Smola provides the known technique of using support vectors for fitting data while maintaining a threshold for possible errors. Ypma provides the environment for correcting errors in a lithographic equipment while maintaining an error threshold using various fitting techniques such as MSE, LS and RBF. Before the effective filing date of the invention, it would have been obvious to a person of ordinary skill in the art, to combine Ypma and Smola to utilize SVR in place of LS for modeling and error minimization with expected results. A PHOSITA would be motivated to do so as SVR provides a more rigorous method for dealing with outliers.
Regarding Claim 2, Ypma in view of Smola teaches the method of claim 1. Smola further teaches the complexity metric is 1-norm or 2-norm of the fitting parameters, or is 1-norm or 2-norm of weighted model parameters (Pg. 55, Eq 2 minimizes the 2 norm form. “Flatness in the case of (1) means that one seeks a small w. One way to ensure this is to minimize the norm,3 i.e. _w_2 = _w,w_. We can write this problem as a convex optimization problem: Eq (2)”).
Regarding Claim 3, Ypma in view of Smola teaches the method of claim 1. Smola further teaches the complexity metric further comprises :one or more slack variables to accommodate any one or more outliers comprised within the measurement data, the deviation between the measurement data and the fitted model being allowed to exceed the threshold value for the one or more outliers, and one or more coefficients for weighting the slack variables (Pg. 2, Eq(3) and “Analogously to the “soft margin” loss function (Bennett and Mangasarian 1992) which was used in SV machines by Cortes and Vapnik (1995), one can introduce slack variables ξi, ξ∗i to cope with otherwise infeasible constraints of the optimization problem (2). Hence we arrive at the formulation stated in Vapnik (1995)… The constant C > 0 determines the trade-off between the flatness of f and the amount up to which deviations larger than
ε are tolerated. This corresponds to dealing with a so calledε-insensitive loss function |ξ |ε described by" EN: C is the coefficient for weighting the slack variables).
PNG
media_image2.png
231
786
media_image2.png
Greyscale
Regarding Claim 4, Ypma in view of Smola teaches the method of claim 3. Smola further teaches the one or more coefficients is a complexity coefficient which can be selected and/or optimized to determine the degree to which the one or more outliers are penalized against the complexity of the fitting (Pg 2, “The constant C > 0 determines the trade-off between the flatness of f and the amount up to which deviations larger thanε are tolerated. This corresponds to dealing with a so calledε-insensitive loss function |ξ |ε described by").
Regarding Claim 5, Ypma in view of Smola teaches the method of claim 1. Ypma further teaches the measurement data comprises at least two-dimensional measurement data ([0100] uses the x,y 2-dimensional measurement as a vector for fitting and [0108] “FIGS. 7-9 illustrate steps in the analysis performed by the first diagnostic apparatus in the example embodiment. In FIG. 7(a), we show the representation of the residual deviations on a first substrate W(1) as a vector AL(1). Each measured deviation has x and y components. It is assumed that each wafer has n alignment marks to be measured (or at least, for the purposes of this analysis, residual deviations for n marks are collected in the object data). The x deviation for the first mark on wafer number 1 is labeled x1,1, while the x deviation for the n—the mark on the first substrate is labeled x1,n. The vector AL(1) comprises all the x and y values for the marks on the first substrate. Similarly, as shown in FIG. 7(b), the residual deviations for a second wafer W(2) are stored as a vector AL(2). The components of this vector are the residual deviations for the n marks as measured on the second wafer, with labels x2,1 to x2,n and y2,1 to y2,n. In an alternative implementation, the data can be organized into a vector per mark position, that is to say a vector X(1) would comprise the first x value for all the wafers, a wafer X(2) would comprise the second x value for all the wafers and so forth. The alternative implementation will be explained in a separate section, further below.”)
Regarding Claim 6, Ypma in view of Smola teaches the method of claim 5. Ypma further teaches the fitting comprises determining a two-dimensional fingerprint describing a spatial distribution of the performance parameter ([0100] and [0108]).
Regarding Claim 7, Ypma in view of Smola teaches the method of claim 1. Smola further teaches defining Lagrange multipliers for the complexity metric, converting the complexity metric into a Lagrangian function using the Lagrange multipliers and converting the Lagrangian function into a quadratic programming optimization (Pg. 2, " 1.3. Dual problem and quadratic programs, The key idea is to construct a Lagrange function from the objective function (it will be called the primal objective function in the rest of this article) and the corresponding constraints, by introducing a dual set of variables. It can be shown that this function has a saddle point with respect to the primal and dual variables at the solution… Here L is the Lagrangian and ηi, η∗i , αi, α∗i are Lagrange multipliers."
PNG
media_image3.png
311
769
media_image3.png
Greyscale
PNG
media_image4.png
386
802
media_image4.png
Greyscale
Eq 5 shows the complexity metric as a Lagrange function using the Lagrange multipliers. Eq 10 is the quadratic programing optimization.).
Regarding Claim 8, Ypma in view of Smola teaches the method of claim 7. Smola further teaches the fitting comprises determining model parameters as a linear combination of a design matrix and optimized values for the Lagrange multipliers (Pg. 56, “This is the so-called Support Vector expansion, i.e. w can be completely described as a linear combination of the training patterns xi. In a sense, the complexity of a function’s representation by SVs is independent of the dimensionality of the input space X, and depends only on the number of SVs.”
PNG
media_image1.png
100
775
media_image1.png
Greyscale
62, “Lagrange function: The Lagrange function is given by the primal objective function minus the sum of all products between constraints and corresponding Lagrange multipliers (cf. e.g. Fletcher 1989, Bertsekas 1995). Optimization can be seen as minimzation of the Lagrangian wrt. the primal variables and simultaneous maximization wrt. the Lagrange multipliers, i.e. dual variables. It has a saddle point at the solution.” EN: Eq (11), w is the model parameters, which is a linear combination of the design matrix x(i) and the optimized values using the LaGrange multipliers.).
Regarding Claim 9, Ypma in view of Smola teaches the method of claim 1. Ypma further teaches the measurement data describes one or more selected from: a characteristic of the substrate; a characteristic of a patterning device which defines a pattern which is to be applied to the substrate; a position of one or both of a substrate stage for holding the substrate and a reticle stage for holding the patterning device; or a characteristic of a pattern transfer system which transfers the pattern on the patterning device to the substrate([0088] “Also shown in FIG. 2 is a metrology apparatus 240 which is provided for making measurements of parameters of the products at desired stages in the manufacturing process. A common example of a metrology station in a modern lithographic production facility is a scatterometer, for example an angle-resolved scatterometer or a spectroscopic scatterometer, and it may be applied to measure properties of the developed substrates at 220 prior to etching in the apparatus 222. Using metrology apparatus 240, it may be determined, for example, that important performance parameters such as overlay or critical dimension (CD) do not meet specified accuracy requirements in the developed resist. Prior to the etching step, the opportunity exists to strip the developed resist and reprocess the substrates 220 through the litho cluster. As is also well known, the metrology results from the apparatus 240 can be used for quality control. They can also be used as inputs for a process monitoring system. This system can be to maintain accurate performance of the patterning operations in the litho cluster, by making small adjustments over time, thereby minimizing the risk of products being made out-of-specification, and requiring re-work. Of course, metrology apparatus 240 and/or other metrology apparatuses (not shown) can be applied to measure properties of the processed substrates 232, 234, and incoming substrates 230.” Also [0102] for the reticle stage.).
Regarding Claim 10, Ypma in view of Smola teaches the method of claim 1. Ypma further teaches the measurement data comprises one or more selected from: overlay data, critical dimension data, alignment data, focus data, or and levelling data ([0087-0088] “The previous and/or subsequent processes may be performed in other lithography apparatuses, as just mentioned, and may even be performed in different types of lithography apparatus. For example, some layers in the device manufacturing process which are very demanding in parameters such as resolution and overlay may be performed in a more advanced lithography tool than other layers that are less demanding. Therefore some layers may be exposed in an immersion type lithography tool, while others are exposed in a ‘dry’ tool. Some layers may be exposed in a tool working at DUV wavelengths, while others are exposed using EUV wavelength radiation. ).
Regarding Claim 11, Ypma in view of Smola teaches the method of claim 1. Ypma further teaches the controllable parameter comprises at least one of: exposure trajectory control in directions parallel to a substrate plane; exposure trajectory control in a direction perpendicular to the substrate plane; lens aberration correction; dose control; or and laser bandwidth control for a source laser of a lithographic apparatus ([0032-0034] and [0077-0079] “1. In step mode, the mask table MT and the substrate table WTa/WTb are kept essentially stationary, while an entire pattern imparted to the radiation beam is projected onto a target portion C at one time (i.e. a single static exposure). The substrate table WTa/WTb is then shifted in the X and/or Y direction so that a different target portion C can be exposed. In step mode, the maximum size of the exposure field limits the size of the target portion C imaged in a single static exposure. The velocity and direction of the substrate table WTa/WTb relative to the mask table MT may be determined by the (de-)magnification and image reversal characteristics of the projection system PS. In scan mode, the maximum size of the exposure field limits the width (in the non-scanning direction) of the target portion in a single dynamic exposure, whereas the length of the scanning motion determines the height (in the scanning direction) of the target portion.” And [0073-0075] for the radiation beam dose control).
Regarding Claim 12, Ypma in view of Smola teaches the method of claim 11. Ypma further teaches controlling the lithographic process according to the optimized control ([0032-0034).
Regarding Claim 13, Ypma in view of Smola teaches the method of claim 11. Ypma further teaches the lithographic process comprises exposure of a layer on a substrate, and the forming part of a manufacturing process is for manufacturing an integrated circuit ([0082-0083] “The apparatus further includes a lithographic apparatus control unit LACU which controls all the movements and measurements of the various actuators and sensors described. LACU also includes signal processing and data processing capacity to implement desired calculations relevant to the operation of the apparatus. In practice, control unit LACU will be realized as a system of many sub-units, each handling the real-time data acquisition, processing and control of a subsystem or component within the apparatus. For example, one processing subsystem may be dedicated to servo control of the substrate positioner PW. Separate units may even handle coarse and fine actuators, or different axes. Another unit might be dedicated to the readout of the position sensor IF. Overall control of the apparatus may be controlled by a central processing unit, communicating with these sub-systems processing units, with operators and with other apparatuses involved in the lithographic manufacturing process. FIG. 2 at 200 shows the lithographic apparatus LA in the context of an industrial production facility for semiconductor products. Within the lithographic apparatus (or “litho tool” 200 for short), the measurement station MEA is shown at 202 and the exposure station EXP is shown at 204. The control unit LACU is shown at 206. Within the production facility, apparatus 200 forms part of a “litho cell” or “litho cluster” that contains also a coating apparatus 208 for applying photosensitive resist and other coatings to substrate W for patterning by the apparatus 200. At the output side of apparatus 200, a baking apparatus 210 and developing apparatus 212 are provided for developing the exposed pattern into a physical resist pattern.”).
Regarding Claim 14, Ypma in view of Smola teaches the method of claim 1. Ypma further teaches the complexity metric is operable to minimize one or more selected from: overlay error, edge placement error, critical dimension error, focus error, alignment error or levelling error ([0092] “In this way, the object data can include measurements directly or indirectly of parameters such as overlay, CD, side wall angle, mark asymmetry, leveling and focus. Further below, an embodiment will be described in which such object data can be used and analyzed to implement an improved process monitoring system in the manufacturing facility of FIG. 2. It is also possible that these parameters can be measured by apparatus within the litho tool 200 itself. Various prior publications describe special marks and/or or measurement techniques for this. For example information on mark asymmetry can be obtained using signals obtained at different wavelengths by the alignment sensors. “ [0101] “As illustrated in FIG. 4 the measured positions 404 of all the targets can be processed numerically to set up a model of a distorted wafer grid 406 for this particular wafer. This alignment model is used in the patterning operation to control the position of the patterns applied to the substrate. In the example illustrated, the straight lines of the nominal grid have become curves, indicating use of a higher order (advanced) alignment model. It goes without saying that the distortions illustrated are exaggerated compared to the real situation. Alignment is a unique part of the lithographic process, because it is the correction mechanism able to correct for deviations (distortions) in each exposed wafer. The alignment measures positions of alignment targets formed in a previous layer. The inventors have recognized that alignment data (and related data such as level sensor data) is always collected and always available. By finding a way to exploit this data as a resource for use in root cause analysis, the methods and apparatuses described herein greatly increase the practicality of such analysis.” ).
Regarding Claim 15, Ypma teaches a non-transitory computer readable medium comprising program instructions ([0223-0227])
The remaining limitations are similar to claim 1 and are rejected under the same rationale.
Claims 16-20 are medium claims reciting limitations similar to claims 2-4 and 11-12, and are rejected under the same rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Chuang et al. (Robust Least Squares-Support Vector Machines for Regression with Outliers): Discloses the foundations for least square SVM with the detailed definitions for all the equation terms used by You et al.
Onose et al. (US20190129313A1): Discloses model fitting/estimation using MSE and threshold in the context of manufacturing semiconductors.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to AMIR DARWISH whose telephone number is (571)272-4779. The examiner can normally be reached 7:30-5:30 M-Thurs.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Lewis Bullock can be reached on 571-272-3759. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/A.E.D./Examiner, Art Unit 2199
/LEWIS A BULLOCK JR/Supervisory Patent Examiner, Art Unit 2199