DETAILED ACTION
Response to Amendment
This Action is responsive to Applicant’s response filed on 08/16/2026. All claims are still pending
in the present application. This Action is made FINAL.
Amendment
Applicant submitted amendments on 08/16/2026. The Examiner acknowledges the amendment and has reviewed the claims accordingly.
Response to Arguments
During prosecution, claim scope not solely on the basis of claim language, but also on giving claims their broadest reasonable construction in light of the specification as it would be interpreted by one of ordinary skill in the art. In re Am. Acad. of Sci. Tech. Ctr., 367 F.3d 1359, 1364 (Fed. Cir. 2004). See also Superguide Corp. v. DirecTV Enterprises, Inc., 358 F.3d 870, 875 (Fed. Cir. 2004) (“Though understanding the claim language may be aided by explanations contained in the written description, it is important not to import into a claim limitations that are not part of the claim.”). Additionally, “[t]hough understanding the claim language may be aided by the explanations contained in the written description, it is important not to import into a claim limitations that are not a part of the claim. For example, a particular embodiment appearing in the written description may not be read into a claim when the claim language is broader than the embodiment.” Superguide Corp. v. DirecTV Enterprises, Inc., 358 F.3d 870, 875 (Fed. Cir. 2004).
Applicant argues that Alfassi does not teach or suggest the recited pre-trained vision foundation model (VFM). Applicant’s arguments have been fully considered but are not persuasive. However, this argument addresses Alfassi individually rather than the combined teachings of Alfassi and Yuan upon which the rejection is based. The Office Action expressly acknowledged that Alfassi fails to teach the claimed pre-trained VFM and relied upon Yuan to cure that deficiency. Alfassi was relied upon for teaching a computerized defect-detection system that acquires multiple runtime images of a semiconductor specimen, including an inspection image and a corresponding reference image, and feeds the images to a trained machine-learning model “to generate a feature vector representative of the given TOI candidate” for determining whether the candidate is a target of interest or non-target of interest (Alfassi, Paragraph [0010], Paragraph [0027], and Paragraph [0035]). Alfassi further teaches that the specimen images may be generated using optical or electron-based inspection machines, including SEM, AFM, and TEM systems (Alfassi, Paragraph [0004] and Paragraph [0056]). The proper obviousness inquiry is not whether Alfassi individually discloses every claimed limitation, but “what the combined teachings of those references would have suggested to those of ordinary skill in the art.” In re Keller, 642 F.2d 413, 425, 208 USPQ 871, 881 (CCPA 1981); See also MPEP 2145(IV); In re Mouttet, 686 F.3d 1322, 1333, 103 USPQ2d 1219, 1226 (Fed. Cir. 2012).
As a result, the argued features are written such that they read upon the cited references. Therefore, the previous rejection still applies.
Applicant further argues that Yuan’s VFM is configured to accept non-image inputs because Yuan’s pre-training model 212 may include an image encoder 220 and a language encoder 222. This argument improperly treats Yuan’s separately disclosed image-processing and language-processing branches as though each branch accepts both types of inputs. Yuan expressly discloses “a two-tower architecture including an image encoder 220 and a language encoder 222” (Yuan, Paragraph [0022]). Language data are supplied to language encoder 222 to produce encoded language data, whereas images are supplied to image encoder 220, which includes a hierarchical vision transformer with shifted windows and convolutional embedding, to produce visual feature representations (Yuan, Paragraphs [0028] – Paragraph [0031]). Thus, the text identified by Applicant is input to a separately identified language encoder, not to image encoder 220 corresponding to the claimed pre-trained VFM configured for projecting multiple images to high-dimensional embeddings. The presence of a separate language encoder in Yuan’s broader pre-training architecture does not establish that the relied-upon image encoder accepts text or other non-image inputs. Furthermore, the obviousness analysis does not require the bodily incorporation of Yuan’s entire pre-training architecture into Alfassi. “A determination of obviousness based on teachings from multiple references does not require an actual, physical substitution of elements.” In re Mouttet, 686 F.3d at 1332; See also MPEP 2145(III); In re Keller, 642 F.2d at 425. The issue is what the image-encoding teachings of Yuan would have suggested when applied to Alfassi’s image-based defect-detection system, not whether Yuan’s complete two-tower training architecture must be physically incorporated into Alfassi.
As a result, the argued features are written such that they read upon the cited references. Therefore, the previous rejection still applies.
The claims do not require that every component involved in training the VFM, or the computer system as a whole, be incapable of receiving non-image information. During examination, claims are given their broadest reasonable interpretation consistent with the specification. MPEP 2111; In re Morris, 127 F.3d 1048, 1054-55, 44 USPQ2d 1023, 1027-28 (Fed. Cir. 1997). The claims require the recited pre-trained VFM that projects the multiple images to high-dimensional embeddings to be configured for accepting only inputs in image formats. Under the broadest reasonable interpretation, Yuan’s pre-trained image encoder satisfies this limitation because the inputs supplied to that vision encoder are images, while language data are supplied to a separate language encoder. Moreover, the open-ended term “comprising” does not exclude the presence of an additional language encoder elsewhere in the system. MPEP 2111.03; Genentech, Inc. v. Chiron Corp., 112 F.3d 495, 501, 42 USPQ2d 1608, 1613 (Fed. Cir. 1997). Therefore, the presence of an additional language encoder elsewhere in Yuan’s system does not remove the relied-upon image encoder from the scope of the claims. Moreover, Yuan expressly discloses support for “single modalities (image only),” RGB-only modalities, and downstream tasks including image retrieval, image classification, object detection, action recognition, and object tracking (Yuan, Paragraphs [0015], Paragraph [0018], and Paragraph [0023]). Accordingly, Yuan expressly contemplates image-only operation and does not require non-image inputs for every use of its pre-trained vision model.
As a result, the argued features are written such that they read upon the cited references. Therefore, the previous rejection still applies.
Applicant’s teaching-away argument is also unpersuasive. The proposed combination does not require removing Yuan’s language encoder, preventing cross-modal pretraining, or rendering Yuan’s broader architecture incapable of processing language information. Rather, the combination applies Yuan’s pre-trained image-processing model to the multiple semiconductor-specimen images acquired and evaluated by Alfassi. In re Fulton, 391 F.3d 1195, 1201, 73 USPQ2d 1141, 1146 (Fed. Cir. 2004); See also MPEP 2145(D). Yuan does not criticize, discredit, or discourage image-only operation; to the contrary, Yuan expressly identifies image-only modalities and multiple image-based downstream tasks. A person of ordinary skill in the art would therefore have been motivated to incorporate Yuan’s pre-trained image encoder into Alfassi’s semiconductor-specimen inspection system to provide transferable visual representations learned from large and diverse data sets, thereby predictably improving the adaptability and efficiency of Alfassi’s image-based defect analysis. Accordingly, Applicant’s arguments do not overcome the rejection of independent claims 1, 29, and 30, or the claims dependent therefrom, under 35 U.S.C. § 103 over Alfassi in view of Yuan. As a result, the argued features are written such that they read upon the cited references. Therefore, the previous rejection still applies.
Office Action Summary
Claim(s) 2 and 10 is/are cancelled.
Claim(s) 1, 3, 11, 13-15, 17, 22-24, and 26-30 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alfassi et al (US 2025/0191177 A1) in view of Yuan et al (US 2023/0162481 A1).
Claim(s) 4-5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alfassi et al (US 2025/0191177 A1) in view of Yuan et al (US 2023/0162481 A1), further in view of Cohen et al (US 2005/0030315 A1).
Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alfassi et al (US 2025/0191177 A1) in view of Yuan et al (US 2023/0162481 A1) and Cohen et al (US 2005/0030315 A1), further in view of Yati (US 2016/0371831 A1).
Claim(s) 7-9 and 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alfassi et al (US 2025/0191177 A1) in view of Yuan et al (US 2023/0162481 A1), further in view of Yati (US 2016/0371831 A1).
Claim(s) 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alfassi et al (US 2025/0191177 A1) in view of Yuan et al (US 2023/0162481 A1), further in view of Wang et al (WO 2022/221045 A1).
Claim(s) 18-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alfassi et al (US 2025/0191177 A1) in view of Yuan et al (US 2023/0162481 A1), further in view of Eser et al (US 2019/0370616 A1).
Claim(s) 20-21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alfassi et al (US 2025/0191177 A1) in view of Yuan et al (US 2023/0162481 A1) and Eser et al (US 2019/0370616 A1), further in view of Wang et al (WO 2022/221045 A1).
Claim(s) 25 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alfassi et al (US 2025/0191177 A1) in view of Yuan et al (US 2023/0162481 A1), further in view of Roham et al (US 2024/0378347 A1).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1, 3, 11, 13-15, 17, 22-24, and 26-30 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alfassi et al (US 2025/0191177 A1) in view of Yuan et al (US 2023/0162481 A1).
Regarding claim(s) 1, 29, and 30, Alfassi teaches a non-transitory computer-readable medium, storing program instructions executable on a computer system (Figure 1; and Paragraph [0061]) for performing a computer-implemented method for determining information for a specimen (Paragraph [0010]: “[…] there is provided a computerized system of defect detection on a semiconductor specimen, the system comprising a processing circuitry […]”), wherein the computer-implemented method comprises:
inputting multiple images for a specimen into a pre-trained vision foundation model (VFM) (Paragraph [0027]: “acquire, in runtime, a set of runtime images of a specimen to be examined, comprising an inspection image comprising one or more TOI candidates, and a reference image corresponding to the inspection image […] feed the inspection patch and the reference patch together to a trained machine learning (ML) model, to generate a feature vector representative of the given TOI candidate”), wherein the multiple images comprise an image generated for the specimen with one or more modes of an imaging system (Paragraph [0004]: “generating certain output (e.g., images, signals, etc.) for a specimen by directing light or electrons to the wafer, and detecting the light or electrons from the wafer”; and Paragraph [0056]: “as inspection machines of various types, such as optical inspection machines, electron beam inspection machines (e.g., a Scanning Electron Microscope (SEM), an Atomic Force Microscopy (AFM), or a Transmission Electron Microscope (TEM), etc.), and so on”); and
determining information for the specimen from the high dimensional embeddings, wherein the pre-trained VFM is included in one or more components executed by the computer system (Figure 1; Paragraph [0027]: “to generate a feature vector representative of the given TOI candidate, wherein the ML model is previously trained to map targets of interest (TOIs) and non-TOIs to corresponding feature vectors in an attribute space such that feature vectors of the TOIs are relatively close to each other with respect to feature vectors of the non-TOIs; and evaluate the feature vector of the given TOI candidate to provide a likelihood of the given TOI candidate being a TOI or non-TOI”; and Paragraph [0035]: “The processing circuitry can be configured to evaluate the feature vector of the given candidate using a classification model operatively connected to the ML model. The classification model is configured to classify the given TOI candidate based on the feature vector thereof to provide a probability score”).
Alfassi fails to teach a pre-trained vision foundation model (VFM) configured for projecting the multiple images to high dimensional embeddings via continuous pretraining, and wherein the pre-trained VFM is further configured for accepting only inputs in image formats.
However, Yuan teach a pre-trained vision foundation model (VFM) configured for projecting the multiple images to high dimensional embeddings via continuous pretraining (Paragraph [00016]: “a computer vision foundation model may be pre-trained based on […] encoded using a hierarchical vision transformer with shifted windows and convolutional embedding. The encoded images and encoded language may then be used to pre-train the computer vision foundation model via unified image-text contrastive learning”; Paragraph [0013]: “may be applied to any model that is trained from broad data sets at a scale that is capable of being adapted (e.g., fine-tuned) to a wide range of downstream tasks […]”; Paragraph [0023]: “Computer vision foundation model 210 may be pre-trained on increasingly large data sets via scalable training infrastructure 240. Once pre-trained and adapted, computer vision foundation model 210 may be trained on any number of tasks 250 […]”; Paragraph [0029]: “Image encoder 220 may include a hierarchical vision transformer with shifted windows and convolutional embedding […]”; Paragraph [0033]: “[…] For example, two or more linear projection layers may be added on top of the image encoder and language encoder to match the dimensions of image and language features […]”; and Paragraph [0037]: “The image encoder 220 and language encoder 222 may be denoted fθ and fϕ, respectively. u and v are the normalized visual feature vector and language feature vector, respectively […]”), and wherein the pre-trained VFM is further configured for accepting only inputs in image formats (Paragraph [0015]: “While existing vision foundation models focus mainly on mapping images and textual representations to a cross-modal shared representation […] a computer vision foundation model that builds expansive representations able to support space-based tasks (e.g., coarse (scene) to fine (object)), time-based tasks (e.g., static (images) to dynamic (videos)), and modality-based tasks (e.g., single modalities (image only) to multiple modalities (caption, depth))”; Paragraph [0019]: “[…] Modality-based tasks may range from monochromatic or red-green-blue (RGB) only to multiple senses (e.g. captioning and depth) […]”; Paragraph [0022]: “[…] pre-training model 212 may comprise a two-tower architecture including an image encoder 220 and a language encoder 222 […]”; Paragraph [0028]: “Language encoder 222 may be any suitable language encoder that produces encoded language data usable by unified image-text contrastive learning module 225 […]”; and Paragraph [0049]: “[…] Computer vision foundation model 300 may include at least a pre-training model 302 and a set of adaptation models 305. Pre-training model 302 may be an example of pre-training model 212, and may include a hierarchical vision transformer with shifted windows and convolutional embedding, such as a modified swin transformer”).
determining information for the specimen from the high dimensional embeddings, wherein the pre-trained VFM is included in one or more components executed by the computer system (Paragraph [0022]: “pre-training model 212 may comprise a two-tower architecture including an image encoder 220 and a language encoder 222. The two-tower architecture allows for end-to-end training of pre-training model 212 via unified image-text contrastive learning module 225”; and Paragraph [0023]: “Once pre-trained and adapted, computer vision foundation model 210 may be trained on any number of tasks 250, such as image retrieval, image classification, object detection, VQA, action recognition, object tracking, etc.”).
Therefore, it would have been obvious to a person of ordinary skill in the art at the time of the invention to modify the system of Alfassi to incorporate the pre-trained computer vision foundation model as taught by Yuan, in order to process multiple images of a specimen and generate feature representations (embeddings) that can be used for determining information such as defect likelihood, because Yuan teaches that such pre-trained foundation models can be adapted to a wide range of downstream computer vision tasks, including classification and analysis of images. The motivation for this combination of references would have been to provide a computer vision foundation model that is trained on large-scale diverse data sets and can be adapted to a wide range of downstream tasks, thereby enabling efficient image analysis such as classification and retrieval. This motivation for the combination of Alfassi and Yuan is supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Regarding claim(s) 3, Alfassi as modified by Yuan teaches the system of claim 1, where Alfassi teaches wherein the multiple images further comprise multi-mode images generated for the specimen with multiple modes of the imaging system (Paragraph [0027]: “acquire, in runtime, a set of runtime images of a specimen to be examined, comprising an inspection image comprising one or more TOI candidates, and a reference image corresponding to the inspection image”; Paragraph [0031]: “The set of runtime images can further comprise a difference image representative of a difference between the inspection image and the reference image”; and Paragraph [0056]: “as inspection machines of various types, such as optical inspection machines, electron beam inspection machines (e.g., a Scanning Electron Microscope (SEM), an Atomic Force Microscopy (AFM), or a Transmission Electron Microscope (TEM), etc.), and so on”).
Regarding claim(s) 11, Alfassi as modified by Yuan teaches the system of claim 1, wherein the one or more components further comprise a multi-image encoder configured for projecting the multiple images into a latent space embedding (where Alfassi teaches in Paragraph [0027]: “acquire, in runtime, a set of runtime images of a specimen to be examined, comprising an inspection image comprising one or more TOI candidates, and a reference image corresponding to the inspection image […] feed the inspection patch and the reference patch together to a trained machine learning (ML) model, to generate a feature vector representative of the given TOI candidate”; and where Yuan teaches in Paragraph [00016]: “a computer vision foundation model may be pre-trained based on […] encoded using a hierarchical vision transformer with shifted windows and convolutional embedding. The encoded images and encoded language may then be used to pre-train the computer vision foundation model via unified image-text contrastive learning”; Paragraph [0013]: “may be applied to any model that is trained from broad data sets at a scale that is capable of being adapted (e.g., fine-tuned) to a wide range of downstream tasks […]”; and Paragraph [0023]: “Computer vision foundation model 210 may be pre-trained on increasingly large data sets via scalable training infrastructure 240. Once pre-trained and adapted, computer vision foundation model 210 may be trained on any number of tasks 250 […]”).
Regarding claim(s) 13, Alfassi as modified by Yuan teaches the system of claim 1, where Yuan teaches wherein the computer system is configured for pre-training an initial VFM from scratch with unlabeled training images through self-supervised learning thereby generating the pre-trained VFM (Paragraph [00016]: “a computer vision foundation model may be pre-trained based on […] encoded using a hierarchical vision transformer with shifted windows and convolutional embedding. The encoded images and encoded language may then be used to pre-train the computer vision foundation model via unified image-text contrastive learning”; Paragraph [0022]: “The pre-trained model generated by pre-training model 212 may be provided to a plurality of adaptation models 215 […]”; Paragraph [0013]: “may be applied to any model that is trained from broad data sets at a scale that is capable of being adapted (e.g., fine-tuned) to a wide range of downstream tasks […]”; and Paragraph [0023]: “Computer vision foundation model 210 may be pre-trained on increasingly large data sets via scalable training infrastructure 240. Once pre-trained and adapted, computer vision foundation model 210 may be trained on any number of tasks 250 […]”).
Regarding claim(s) 14, Alfassi as modified by Yuan teaches the system of claim 1, where Yuan teaches wherein the computer system is configured for pre-training an initial VFM from pre-trained parameters with unlabeled training images through self-supervised learning thereby generating the pre-trained VFM (Paragraph [0013]: “may be applied to any model that is trained from broad data sets at a scale that is capable of being adapted (e.g., fine-tuned) to a wide range of downstream tasks […]”; Paragraph [0023]: “Computer vision foundation model 210 may be pre-trained on increasingly large data sets via scalable training infrastructure 240. Once pre-trained and adapted, computer vision foundation model 210 may be trained on any number of tasks 250 […]”; and Paragraph [0022]: “The pre-trained model generated by pre-training model 212 may be provided to a plurality of adaptation models 215 […]”).
Regarding claim(s) 15, Alfassi as modified by Yuan teaches the system of claim 1, where Yuan teaches wherein the computer system is configured for pre-training an initial VFM thereby generating the pre-trained VFM and fine-tuning the one or more components with labeled training data (Paragraph [00016]: “a computer vision foundation model may be pre-trained based on […] encoded using a hierarchical vision transformer with shifted windows and convolutional embedding. The encoded images and encoded language may then be used to pre-train the computer vision foundation model via unified image-text contrastive learning”; and Paragraph [0013]: “may be applied to any model that is trained from broad data sets at a scale that is capable of being adapted (e.g., fine-tuned) to a wide range of downstream tasks […]”).
Regarding claim(s) 17, Alfassi as modified by Yuan teaches the system of claim 15, where Yuan teaches wherein the fine-tuning comprises modifying one or more pre-trained parameters of the pre-trained VFM and one or more parameters of said determining information (Paragraph [00016]: “a computer vision foundation model may be pre-trained based on […] encoded using a hierarchical vision transformer with shifted windows and convolutional embedding. The encoded images and encoded language may then be used to pre-train the computer vision foundation model via unified image-text contrastive learning”; Paragraph [0013]: “may be applied to any model that is trained from broad data sets at a scale that is capable of being adapted (e.g., fine-tuned) to a wide range of downstream tasks […]”; and Paragraph [0023]: “Computer vision foundation model 210 may be pre-trained on increasingly large data sets via scalable training infrastructure 240. Once pre-trained and adapted, computer vision foundation model 210 may be trained on any number of tasks 250 […]”).
Regarding claim(s) 22, Alfassi as modified by Yuan teaches the system of claim 1, where Yuan teaches wherein the one or more additional components are further configured for learning by supervised fine-tuning (Paragraph [00016]: “a computer vision foundation model may be pre-trained based on […] encoded using a hierarchical vision transformer with shifted windows and convolutional embedding […]”; Paragraph [0013]: “may be applied to any model that is trained from broad data sets at a scale that is capable of being adapted (e.g., fine-tuned) to a wide range of downstream tasks […]”; and Paragraph [0023]: “Computer vision foundation model 210 may be pre-trained on increasingly large data sets via scalable training infrastructure 240. Once pre-trained and adapted, computer vision foundation model 210 may be trained on any number of tasks 250 […]”).
Regarding claim(s) 23, Alfassi as modified by Yuan teaches the system of claim 1, where Yuan teaches wherein the one or more additional components are further configured for learning by reinforcement learning (Paragraph [00016]: “a computer vision foundation model may be pre-trained based on […] encoded using a hierarchical vision transformer with shifted windows and convolutional embedding. The encoded images and encoded language may then be used to pre-train the computer vision foundation model via unified image-text contrastive learning”; Paragraph [0013]: “may be applied to any model that is trained from broad data sets at a scale that is capable of being adapted (e.g., fine-tuned) to a wide range of downstream tasks […]”; and Paragraph [0023]: “Computer vision foundation model 210 may be pre-trained on increasingly large data sets via scalable training infrastructure 240. Once pre-trained and adapted, computer vision foundation model 210 may be trained on any number of tasks 250 […]”).
Regarding claim(s) 24, Alfassi as modified by Yuan teaches the system of claim 1, where Alfassi teaches wherein determining the information comprises detecting defects on the specimen (Paragraph [0004]: “inspection of a specimen followed by review of sampled locations of potential defects”) and where Yuan teaches based on the high dimensional embeddings (Paragraph [00016]: “a computer vision foundation model may be pre-trained based on […] encoded using a hierarchical vision transformer with shifted windows and convolutional embedding. The encoded images and encoded language may then be used to pre-train the computer vision foundation model via unified image-text contrastive learning”; and Paragraph [0013]: “may be applied to any model that is trained from broad data sets at a scale that is capable of being adapted (e.g., fine-tuned) to a wide range of downstream tasks […]”).
Regarding claim(s) 26, Alfassi as modified by Yuan teaches the system of claim 1, where Alfassi teaches wherein determining the information comprises classifying defects detected on the specimen (Paragraph [0007]: “detect and classify defects on specimens, as well as perform metrology related operations. Effectiveness of examination can be improved by automatization of process(es) such as, for example, defect detection, Automatic Defect Classification (ADC), Automatic Defect Review (ADR), image segmentation, automated metrology-related operations, etc.”) and where Yuan teaches based on the high dimensional embeddings (Paragraph [00016]: “a computer vision foundation model may be pre-trained based on […] encoded using a hierarchical vision transformer with shifted windows and convolutional embedding. The encoded images and encoded language may then be used to pre-train the computer vision foundation model via unified image-text contrastive learning”; and Paragraph [0013]: “may be applied to any model that is trained from broad data sets at a scale that is capable of being adapted (e.g., fine-tuned) to a wide range of downstream tasks […]”).
Regarding claim(s) 27, Alfassi as modified by Yuan teaches the system of claim 1, where Yuan teaches wherein determining the information comprises segmenting one or more of the multiple images based on the high dimensional embeddings (Paragraph [00016]: “a computer vision foundation model may be pre-trained based on […] encoded using a hierarchical vision transformer with shifted windows and convolutional embedding. The encoded images and encoded language may then be used to pre-train the computer vision foundation model via unified image-text contrastive learning”; Paragraph [0013]: “may be applied to any model that is trained from broad data sets at a scale that is capable of being adapted (e.g., fine-tuned) to a wide range of downstream tasks […]”; and Paragraph [0019]: “[…] space-based 110, time-based 120, and modality-based 130. Space-based tasks may range from coarse recognition (e.g. scene-level classification) to fine-grained recognition (e.g. object detection, segmentation)”).
Regarding claim(s) 28, Alfassi as modified by Yuan teaches the system of claim 1, where Alfassi teaches wherein determining the information comprises selecting one or more modes of the imaging system for a process performed on the specimen or another specimen based on the high dimensional embeddings (Paragraph [0027]: “acquire, in runtime, a set of runtime images of a specimen to be examined, comprising an inspection image comprising one or more TOI candidates, and a reference image corresponding to the inspection image”; Paragraph [0004]: “generating certain output (e.g., images, signals, etc.) for a specimen by directing light or electrons to the wafer, and detecting the light or electrons from the wafer”; Paragraph [0031]: “The set of runtime images can further comprise a difference image representative of a difference between the inspection image and the reference image”; and Paragraph [0056]: “as inspection machines of various types, such as optical inspection machines, electron beam inspection machines (e.g., a Scanning Electron Microscope (SEM), an Atomic Force Microscopy (AFM), or a Transmission Electron Microscope (TEM), etc.), and so on”).
Claim(s) 4-5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alfassi et al (US 2025/0191177 A1) in view of Yuan et al (US 2023/0162481 A1), further in view of Cohen et al (US 2005/0030315 A1).
Regarding claim(s) 4, Alfassi as modified by Yuan teaches the system of claim 1, but do not specifically teach wherein the one or more components further comprise an image packing component configured for generating a single image that contains information from the multiple images.
However, Cohen teaches wherein the one or more components further comprise an image packing component configured for generating a single image that contains information from the multiple images (Paragraph [0009]: “An image stack is a set of identically sized registered images […]”; Paragraph [0011]: “Filters may be applied to the 3D image stack, or a portion thereof, to create one or more new 2D intermediate images. A filter is a function that operates on the 3D image stack to create a 2D image”; and Abstract: “[…] image stack is employed in creating an enhanced image from a stack of registered images. This paradigm combines pixels using multi-image operations on the image stack […]”).
Therefore, it would have been obvious to a person having ordinary skill in the art at the time of the invention to modify the system of Alfassi to include the image combination technique of Cohen, in order to pack multiple images into a single packed image. Cohen teaches combining multiple registered images in an image stack and applying multi-image operations to create a single enhanced composite image, including “creating an enhanced image from a stack of registered images” and “combining pixels using multi-image operations on the image stack.” The motivation for this combination of references would have been to combine multiple registered images using multi-image operations on an image stack to generate a single enhanced image. This motivation for the combination of Alfassi, Yuan, and Cohen is supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Regarding claim(s) 5, Alfassi as modified by Yuan and Cohen teaches the system of claim 4, where Alfassi teaches wherein the multiple images further comprise multi-mode images generated for the specimen with multiple modes of the imaging system (Paragraph [0027]: “acquire, in runtime, a set of runtime images of a specimen to be examined, comprising an inspection image comprising one or more TOI candidates, and a reference image corresponding to the inspection image”; Paragraph [0031]: “The set of runtime images can further comprise a difference image representative of a difference between the inspection image and the reference image”; and Paragraph [0056]: “as inspection machines of various types, such as optical inspection machines, electron beam inspection machines (e.g., a Scanning Electron Microscope (SEM), an Atomic Force Microscopy (AFM), or a Transmission Electron Microscope (TEM), etc.), and so on”).
Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alfassi et al (US 2025/0191177 A1) in view of Yuan et al (US 2023/0162481 A1) and Cohen et al (US 2005/0030315 A1), further in view of Yati (US 2016/0371831 A1).
Regarding claim(s) 6, Alfassi as modified by Yuan and Cohen teaches the system of claim 4, where Alfassi teaches wherein the multiple images (Paragraph [0027]: “acquire, in runtime, a set of runtime images of a specimen to be examined, comprising an inspection image comprising one or more TOI candidates, and a reference image corresponding to the inspection image”; Paragraph [0031]: “The set of runtime images can further comprise a difference image representative of a difference between the inspection image and the reference image”; and Paragraph [0056]: “as inspection machines of various types, such as optical inspection machines, electron beam inspection machines (e.g., a Scanning Electron Microscope (SEM), an Atomic Force Microscopy (AFM), or a Transmission Electron Microscope (TEM), etc.), and so on”).
Alfassi, Yuan, and Cohen fails to teach wherein
However, Yati teaches wherein (Figure 6; and Paragraph [0016]: “aligning using a controller a design file for a current layer of the wafer to an image of the current layer; aligning using the controller a design file for a previous layer of the wafer to the design file for the current layer; and identifying a region of the image of the current layer based on a coordinate of a defect in the previous layer using the controller”).
Therefore, it would have been obvious to a person having ordinary skill in the art at the time of the invention to modify the system of Alfassi in view of Cohen and Yati to include multiple images corresponding to different layers of a specimen, including a design layer image and a prior layer image. Cohen teaches combining multiple images into a single composite image using multi-image operations on an image stack, while Yati teaches aligning a design file for a current layer to an image of the current layer and utilizing an image of a previous layer formed prior to the current layer. The motivation for this combination of references would have been to utilize images corresponding to multiple layers of a specimen, including aligning a design file for a current layer to an image of the current layer and using an image of a previous layer formed prior to the current layer, as taught by Yati. This motivation for the combination of Alfassi, Yuan, Cohen, and Yati is supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Claim(s) 7-9 and 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alfassi et al (US 2025/0191177 A1) in view of Yuan et al (US 2023/0162481 A1), further in view of Yati (US 2016/0371831 A1).
Regarding claim(s) 7, Alfassi as modified by Yuan teaches the system of claim 1, where Alfassi teaches wherein the multiple images further comprise the image generated for the specimen with the one or more modes of the imaging system (Paragraph [0027]: “acquire, in runtime, a set of runtime images of a specimen to be examined, comprising an inspection image comprising one or more TOI candidates, and a reference image corresponding to the inspection image”; Paragraph [0004]: “generating certain output (e.g., images, signals, etc.) for a specimen by directing light or electrons to the wafer, and detecting the light or electrons from the wafer”; and Paragraph [0056]: “as inspection machines of various types, such as optical inspection machines, electron beam inspection machines (e.g., a Scanning Electron Microscope (SEM), an Atomic Force Microscopy (AFM), or a Transmission Electron Microscope (TEM), etc.), and so on”)
Alfassi and Yuan fails to teach at least one image generated from design data for the specimen. However, Yati teaches at least one image generated from design data for the specimen (Figure 6; and Paragraph [0016]: “aligning using a controller a design file for a current layer of the wafer to an image of the current layer; aligning using the controller a design file for a previous layer of the wafer to the design file for the current layer; and identifying a region of the image of the current layer based on a coordinate of a defect in the previous layer using the controller”).
Therefore, it would have been obvious to a person having ordinary skill in the art at the time of the invention to modify the system of Alfassi to further include the use of design data as taught by Yati, in order to utilize both images generated by an imaging system and images generated from design data for a specimen. Alfassi teaches generating images of a specimen using one or more imaging system modes, such as inspection images generated by directing light or electrons to the specimen. Yati teaches using design data for a specimen, including aligning a design file for a layer to an image of the specimen, thereby incorporating design-based information in image form. The motivation for this combination of references would have been to utilize design-based information together with inspection images by aligning a design file for a current layer to an image of the current layer, as taught by Yati. This motivation for the combination of Alfassi, Yuan, and Yati is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Regarding claim(s) 8, Alfassi as modified by Yuan and Yati teaches the system of claim 7, where Alfassi teaches wherein the image generated for the specimen with the one or more modes of the imaging system (Paragraph [0027]: “acquire, in runtime, a set of runtime images of a specimen to be examined, comprising an inspection image comprising one or more TOI candidates, and a reference image corresponding to the inspection image”; Paragraph [0031]: “The set of runtime images can further comprise a difference image representative of a difference between the inspection image and the reference image”; and Paragraph [0056]: “as inspection machines of various types, such as optical inspection machines, electron beam inspection machines (e.g., a Scanning Electron Microscope (SEM), an Atomic Force Microscopy (AFM), or a Transmission Electron Microscope (TEM), etc.), and so on”) and where Yati teaches the at least one image generated from the design data for the specimen are generated for the same layer on the specimen (Figure 6; and Paragraph [0016]: “aligning using a controller a design file for a current layer of the wafer to an image of the current layer; aligning using the controller a design file for a previous layer of the wafer to the design file for the current layer; and identifying a region of the image of the current layer based on a coordinate of a defect in the previous layer using the controller”).
Regarding claim(s) 9, Alfassi as modified by Yuan and Yati teaches the system of claim 7, where Alfassi teaches wherein the image generated for the specimen with the one or more modes of the imaging system (Paragraph [0027]: “acquire, in runtime, a set of runtime images of a specimen to be examined, comprising an inspection image comprising one or more TOI candidates, and a reference image corresponding to the inspection image”; Paragraph [0031]: “The set of runtime images can further comprise a difference image representative of a difference between the inspection image and the reference image”; and Paragraph [0056]: “as inspection machines of various types, such as optical inspection machines, electron beam inspection machines (e.g., a Scanning Electron Microscope (SEM), an Atomic Force Microscopy (AFM), or a Transmission Electron Microscope (TEM), etc.), and so on”) and where Yati teaches the at least one image generated from the design data for the specimen are generated for different layers on the specimen (Figure 6; and Paragraph [0016]: “aligning using a controller a design file for a current layer of the wafer to an image of the current layer; aligning using the controller a design file for a previous layer of the wafer to the design file for the current layer; and identifying a region of the image of the current layer based on a coordinate of a defect in the previous layer using the controller”).
Regarding claim(s) 12, Alfassi as modified by Yuan teaches the system of claim 11, where Alfassi teaches wherein the multiple images (Paragraph [0027]: “acquire, in runtime, a set of runtime images of a specimen to be examined, comprising an inspection image comprising one or more TOI candidates, and a reference image corresponding to the inspection image”; Paragraph [0004]: “generating certain output (e.g., images, signals, etc.) for a specimen by directing light or electrons to the wafer, and detecting the light or electrons from the wafer”; and Paragraph [0056]: “as inspection machines of various types, such as optical inspection machines, electron beam inspection machines (e.g., a Scanning Electron Microscope (SEM), an Atomic Force Microscopy (AFM), or a Transmission Electron Microscope (TEM), etc.), and so on”).
Alfassi and Yuan fails to teach wherein Yati teaches wherein (Figure 6; and Paragraph [0016]: “aligning using a controller a design file for a current layer of the wafer to an image of the current layer; aligning using the controller a design file for a previous layer of the wafer to the design file for the current layer; and identifying a region of the image of the current layer based on a coordinate of a defect in the previous layer using the controller. The previous layer is formed prior to the current layer. The region corresponds to the coordinate of the defect in the previous layer. The image of the current layer can be a scanning electron microscope image”).
Therefore, it would have been obvious to one of ordinary skill in the art to combine Alfassi, Yuan, and Yati before the effective filing date of the claimed invention. The motivation for this combination of references would have been to improve image-based analysis of a specimen by incorporating layer-specific information, including aligning a design file for a current layer to an image of the current layer and utilizing information for a previous layer formed prior to the current layer, as taught by Yati, in combination with the embedding-based processing of Yuan. This motivation for the combination of Alfassi, Yuan, and Yati is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Claim(s) 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alfassi et al (US 2025/0191177 A1) in view of Yuan et al (US 2023/0162481 A1), further in view of Wang et al (WO 2022/221045 A1).
Regarding claim(s) 16, Alfassi as modified by Yuan teaches the system of claim 15, but do not specifically teach wherein the fine-tuning comprises fixing the pre-trained VFM to extract the high dimensional embeddings of the labeled training data and only fine-tuning parameters of said determining information.
However, Wang teaches wherein the fine-tuning comprises fixing the pre-trained VFM to extract the high dimensional embeddings of the labeled training data and only fine-tuning parameters of said determining information (Paragraph [0038]: “The multi-task model may be trained through pre training the shared encoder, and in the case of fixing parameters of the pre-trained shared encoder, optimizing multiple task-specific encoder parameter sets of the multiple task- specific encoders and/or multiple linear layer parameter sets of the multiple task-specific linear layers”; Paragraph [0058]: “Parameters of the shared encoder may be fixed during performing of the multiple tasks”; and Paragraph [0075]: “[…] obtain a text input, generate a set of shared representations of the text input in multiple layers”).
Therefore, it would have been obvious to a person having ordinary skill in the art at the time of the invention to modify the system of Yuan in view of Wang to perform fine-tuning by fixing a pre-trained vision foundation model while only updating task-specific parameters. Yuan teaches a pre-trained vision foundation model that encodes input data into embeddings and is adapted for downstream tasks. Wang teaches a multi-task model including a shared encoder that generates shared representations and multiple task-specific encoders and linear layers, wherein parameters of the shared encoder may be fixed during operation, while task-specific encoder parameters and/or linear layer parameters are optimized. The motivation for this combination of references would have been to improve computational efficiency and preserve performance of a pre-trained model across multiple tasks by fixing parameters of a shared encoder while optimizing only task-specific encoder and linear layer parameter sets, as taught by Wang et al., thereby enabling efficient adaptation without affecting previously learned representations. This motivation for the combination of Alfassi, Yuan, and Wang is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Claim(s) 18-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alfassi et al (US 2025/0191177 A1) in view of Yuan et al (US 2023/0162481 A1), further in view of Eser et al (US 2019/0370616 A1).
Regarding claim(s) 18, Alfassi as modified by Yuan teaches the system of claim 1, where Yuan teaches wherein the one or more components further comprise a pre-trained multi-image encoder configured for projecting the multiple images into a latent space embedding, wherein the pre-trained VFM is further configured as a pre-trained latent VFM (LVFM) (Paragraph [00016]: “a computer vision foundation model may be pre-trained based on […] encoded using a hierarchical vision transformer with shifted windows and convolutional embedding. The encoded images and encoded language may then be used to pre-train the computer vision foundation model via unified image-text contrastive learning”; Paragraph [0022]: “pre-training model 212 may comprise a two-tower architecture including an image encoder 220 and a language encoder 222. The two-tower architecture allows for end-to-end training of pre-training model 212 via unified image-text contrastive learning module 225”; Paragraph [0013]: “may be applied to any model that is trained from broad data sets at a scale that is capable of being adapted (e.g., fine-tuned) to a wide range of downstream tasks […]”; and Paragraph [0023]: “Computer vision foundation model 210 may be pre-trained on increasingly large data sets via scalable training infrastructure 240. Once pre-trained and adapted, computer vision foundation model 210 may be trained on any number of tasks 250 […]”).
Alfassi and Yuan fails to teach wherein the computer system is configured for simultaneously training an initial multi-image encoder and an initial LVFM together through self-supervised learning thereby generating the pre-trained multi-image encoder and the pre-trained LVFM.
However, Eser teaches wherein the computer system is configured for simultaneously training an initial multi-image encoder and an initial LVFM together through self-supervised learning thereby generating the pre-trained multi-image encoder and the pre-trained LVFM (Paragraph [0008]: “accessing training data including training data for the first modality and training data for the second modality, training the statistical model, the statistical model comprising first and second encoders, first and second decoders, and a joint-modality representation coupling the first and second encoders to the first and second decoders. The training comprises estimating values for parameters of the first and second encoders and the first and second decoders using a self-supervised learning technique, at least some of the training data, and information describing at least one link between data pairs in the training data, and storing information specifying the statistical model at least in part by storing the estimated values for parameters of the first and second encoders and the first and second decoders of the statistical model”).
Therefore, it would have been obvious to a person having ordinary skill in the art at the time of the invention to modify the system of Alfassi in view of Yuan and further in view of Eser to include jointly trained encoding components for processing multiple inputs. Alfassi teaches processing multiple images of a specimen using a machine learning model to generate feature vectors. Yuan teaches a pre-trained vision foundation model including an image encoder configured to encode images into latent embeddings. Eser teaches a statistical model comprising a plurality of encoders and a joint-modality representation, wherein parameters of the encoders are estimated together using a self-supervised learning technique. The training process updates the parameters of multiple encoders within a shared training loop, thereby inherently training the encoders simultaneously and in conjunction with the shared latent representation. The motivation for this combination of references would have been to improve representation learning across multiple data inputs by jointly estimating parameters of multiple encoders within a shared self-supervised learning framework to produce a common latent representation, as taught by Eser et al., thereby enabling consistent and efficient multi-input embedding generation. This motivation for the combination of Alfassi, Yuan, and Eser is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Regarding claim(s) 19, Alfassi as modified by Yuan and Eser teaches the system of claim 18, where Yuan teaches wherein the computer system is further configured for fine-tuning the one or more components (Paragraph [0013]: “may be applied to any model that is trained from broad data sets at a scale that is capable of being adapted (e.g., fine-tuned) to a wide range of downstream tasks […]”) by where Eser teaches modifying one or more parameters of the pre-trained multi-image encoder, the pre-trained LVFM (Paragraph [0008]: “[…] The training comprises estimating values for parameters of the first and second encoders and the first and second decoders using a self-supervised learning technique, at least some of the training data, and information describing at least one link between data pairs in the training data […]”), and where Alfassi teaches said determining information (Paragraph [0027]: “acquire, in runtime, a set of runtime images of a specimen to be examined, comprising an inspection image comprising one or more TOI candidates, and a reference image corresponding to the inspection image […] feed the inspection patch and the reference patch together to a trained machine learning (ML) model, to generate a feature vector representative of the given TOI candidate”).
Claim(s) 20-21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alfassi et al (US 2025/0191177 A1) in view of Yuan et al (US 2023/0162481 A1) and Eser et al (US 2019/0370616 A1), further in view of Wang et al (WO 2022/221045 A1).
Regarding claim(s) 20, Alfassi as modified by Yuan and Eser teaches the system of claim 18, but do not specifically teach wherein the computer system is further configured for fine-tuning the one or more components by fixing the pre-trained LVFM and only fine-tuning parameters of the pre-trained multi-image encoder and said determining information.
However, Wang teaches teach wherein the computer system is further configured for fine-tuning the one or more components by fixing the pre-trained LVFM and only fine-tuning parameters of the pre-trained multi-image encoder and said determining information (Paragraph [0038]: “The multi-task model may be trained through pre training the shared encoder, and in the case of fixing parameters of the pre-trained shared encoder, optimizing multiple task-specific encoder parameter sets of the multiple task- specific encoders and/or multiple linear layer parameter sets of the multiple task-specific linear layers”; Paragraph [0058]: “Parameters of the shared encoder may be fixed during performing of the multiple tasks”; and Paragraph [0075]: “[…] obtain a text input, generate a set of shared representations of the text input in multiple layers”).
Therefore, it would have been obvious to one of ordinary skill in the art to combine Alfassi, Yuan, Eser and Wang is/are before the effective filing date of the claimed invention. The motivation for this combination of references would have been to improve computational efficiency and preserve performance of a pre-trained model across multiple tasks by fixing parameters of a shared encoder while optimizing only task-specific encoder and linear layer parameter sets, as taught by Wang et al., thereby enabling efficient adaptation without affecting previously learned representations. This motivation for the combination of Alfassi, Yuan, Eser and Wang is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Regarding claim(s) 21, Alfassi as modified by Yuan and Eser teaches the system of claim 18, but do not specifically teach wherein the computer system is further configured for fine-tuning the one or more components by fixing the pre-trained multi-image encoder and the pre-trained LVFM and only fine-tuning parameters of said determining information.
However, Wang teaches wherein the computer system is further configured for fine-tuning the one or more components by fixing the pre-trained multi-image encoder and the pre-trained LVFM and only fine-tuning parameters of said determining information (Paragraph [0038]: “The multi-task model may be trained through pre training the shared encoder, and in the case of fixing parameters of the pre-trained shared encoder, optimizing multiple task-specific encoder parameter sets of the multiple task- specific encoders and/or multiple linear layer parameter sets of the multiple task-specific linear layers”; Paragraph [0058]: “Parameters of the shared encoder may be fixed during performing of the multiple tasks”; and Paragraph [0075]: “[…] obtain a text input, generate a set of shared representations of the text input in multiple layers”).
Therefore, it would have been obvious to one of ordinary skill in the art to combine Alfassi, Yuan, Eser and Wang is/are before the effective filing date of the claimed invention. The motivation for this combination of references would have been to improve computational efficiency and preserve performance of a pre-trained model across multiple tasks by fixing parameters of a shared encoder while optimizing only task-specific encoder and linear layer parameter sets, as taught by Wang et al., thereby enabling efficient adaptation without affecting previously learned representations. This motivation for the combination of Alfassi, Yuan, Eser and Wang is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Claim(s) 25 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alfassi et al (US 2025/0191177 A1) in view of Yuan et al (US 2023/0162481 A1), further in view of Roham et al (US 2024/0378347 A1).
Regarding claim(s) 25, Alfassi as modified by Yuan teaches the system of claim 1, where Yuan teaches wherein determining the information comprises (Paragraph [00016]: “a computer vision foundation model may be pre-trained based on […] encoded using a hierarchical vision transformer with shifted windows and convolutional embedding. The encoded images and encoded language may then be used to pre-train the computer vision foundation model via unified image-text contrastive learning”; and Paragraph [0013]: “may be applied to any model that is trained from broad data sets at a scale that is capable of being adapted (e.g., fine-tuned) to a wide range of downstream tasks […]”).
Alfassi and Yuan fails to teach wherein determining the information comprises generating a digital twin of a manufacturing process performed on the specimen prior to generation of the image generated
However, Roham teaches wherein determining the information comprises generating a digital twin of a manufacturing process performed on the specimen prior to generation of the image generated (Paragraph [0011]: “a computer program product for generating digital twins of process chambers […] generating a digital twin by […]”; Paragraph [0005]: “a digital twin of a process chamber of semiconductor manufacturing equipment, comprising one or more […]”; and Paragraph [0064]: “A digital twin can include multiple models that can be coupled together to form the digital twin […] a surface of a simulated wafer being fabricated by the process chamber, a particular pipe, etc. Additionally, each model can represent a particular class of physical phenomena […]”).
Therefore, it would have been obvious to one of ordinary skill in the art to combine Alfassi, Yuan, and Roham before the effective filing date of the claimed invention. The motivation for this combination of references would have been to improve modeling and prediction of manufacturing processes by leveraging data-driven representations within AI/ML-based digital twin frameworks, thereby enabling more accurate simulation and analysis of process conditions and outcomes. This motivation for the combination of Alfassi, Yuan, and Roham is/are supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III).
Relevant Prior Art Directed to State of Art
Cho et al (US 2024/0355456 A1) are relevant prior art not applied in the rejection(s) above. Cho discloses a method performed by one or more processors, the method comprising: receiving medical image data, wherein the medical image data is associated with a plurality of body parts and configured for training a generative model for medical images; acquiring label data associated with the medical image data, wherein the label data comprises a score associated with an anatomical location of at least one body part of the plurality of body parts; and training, based on the received medical image data and the acquired label data, the generative model for medical images.
Zhang et al (US 2017/0345140 A1) are relevant prior art not applied in the rejection(s) above. Zhang discloses a system configured to generate a simulated image from an input image, comprising: one or more computer subsystems; and one or more components executed by the one or more computer subsystems, wherein the one or more components comprise: a neural network, wherein the neural network comprises: two or more encoder layers configured for determining features of an image for a specimen; and two or more decoder layers configured for generating one or more simulated images from the determined features, wherein the neural network does not comprise a fully connected layer thereby eliminating constraints on size of the image input to the two or more encoders layers.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JONGBONG NAH whose telephone number is (571)272-1361. The examiner can normally be reached M - F: 7:30am - 4:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Edward Urban can be reached on 571-272-7899. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JONGBONG NAH/Examiner, Art Unit 2674
/ONEAL R MISTRY/Supervisory Patent Examiner, Art Unit 2674