Prosecution Insights
Last updated: October 01, 2026
Application No. 18/749,144

NEUROMORPHIC OPTICAL COMPUTING ARCHITECTURE SYSTEM AND APPARATUS

Non-Final OA §103§112
Filed
Jun 20, 2024
Priority
Jun 20, 2023 — CN 202310735709.5
Examiner
WU, NICHOLAS S
Art Unit
Tech Center
Assignee
Tsinghua University
OA Round
1 (Non-Final)
52%
Grant Probability
Moderate
1-2
OA Rounds
1y 9m
Est. Remaining
84%
With Interview

Examiner Intelligence

Grants 52% of resolved cases
52%
Career Allowance Rate
33 granted / 63 resolved
-7.6% vs TC avg
Strong +31% interview lift
Without
With
+31.4%
Interview Lift
resolved cases with interview
Typical timeline
4y 0m
Avg Prosecution
17 currently pending
Career history
89
Total Applications
across all art units

Statute-Specific Performance

§101
25.0%
-15.0% vs TC avg
§103
53.5%
+13.5% vs TC avg
§102
3.8%
-36.2% vs TC avg
§112
16.7%
-23.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 63 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: the multi-channel representation module is configured to encode in claim 1. No sufficient structure provided in the specification. the TD optical attention module performs, based on the trained attention-aware optical neural network, spectral and spatial transmittance modulation in claim 1. Specification ⁋30 describes TD optical attention module as a type of convolutional unit. multi-dimensional sparse features extracted by the BU optical attention module in claim 1. Specification ⁋30 describes BU optical attention module as a type of convolutional unit. the output module is configured to detect and identify in claim 1. No sufficient structure provided in the specification. the TD optical attention module…processes the input and feeds back to the second BU optical attention module to adjust the second BU optical attention module in claim 10. Specification ⁋30 describes TD optical attention module as a convolutional unit. a multidimensional sparse feature output by the first BU optical attention module in claim 10. No sufficient structure provided in the specification. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-15 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claim 1, the claim limitations multi-channel representation module and output module invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. The Specification is devoid on any structure that performs the functions of the modules listed. Therefore, the claim is indefinite and is rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph. Applicant may: (a) Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph; (b) Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (c) Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)). If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either: (a) Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (b) Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181. Regarding claims 2-9, the claims are rejected for at least their dependence on claim 1. Regarding claim 10, the claim limitation first BU optical attention module invokes 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. The Specification is devoid on any structure that performs the functions of the module listed. Therefore, the claim is indefinite and is rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph. Applicant may: (a) Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph; (b) Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (c) Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)). If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either: (a) Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (b) Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181. Regarding claims 11-15, the claims are rejected for at least their dependence on claim 10. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim 1 is rejected under 35 U.S.C. 103 as being unpatentable over Wu, et al., Foreign Patent Publication CN117671454A (“Wu”) (please use the provided copy for mapping purposes) in view of Duan, et al., Non-Patent Literature “Optical multi-task learning using multi-wavelength diffractive deep neural networks” (“Duan”) and further in view of Cao, et al., Non-Patent Literature “Look and Think Twice: Capturing Top-Down Visual Attention with Feedback Convolutional Neural Networks” (“Cao”). Regarding claim 1, Wu discloses: A neuromorphic optical computing architecture system,…an attention-aware optical neural network module and an output module,…; (Wu, pg. 3, “The application provides an optoelectronic computing system and method integrating attention mechanisms, which can reduce electronic computing load and avoid preparing complex optical systems. In a first aspect, embodiments of the present application provide an attention-mechanism-fused optoelectronic computing system, including: the device comprises an electronic neural network, [A neuromorphic optical computing architecture system,…an attention-aware optical neural network module] a laser, a spatial light modulator and a photoelectric detection unit [and an output module,…;]”). the attention-aware optical neural network module comprises: a…first optical attention module and a…second optical attention module, wherein the coherent light having…wavelengths is input to the…first optical attention module (Wu, pg. 3, “In a second aspect, an embodiment of the present application provides an attention-fused photoelectric computing method of an attention fused photoelectric computing system using the foregoing attention-fused photoelectric computing system, [the attention-aware optical neural network module comprises:] where the attention-fused photoelectric computing method includes: the electronic neural network performs feature extraction and conversion on the received picture to be processed [wherein the coherent light having…wavelengths is input to the…first optical attention module] based on an attention mechanism to obtain phase distribution data; [a…first optical attention module] the spatial light modulator modulates single-wavelength laser generated by the laser based on the phase distribution data, [and a…second optical attention module,] and transmits the modulated laser to the photoelectric detection unit through a preset light path”). and network training is performed on an attention-aware optical neural network, (Wu, pg. 6, “It should be appreciated that the fused attention mechanism optoelectronic computing system in this embodiment may be built by training [and network training is performed on an attention-aware optical neural network,]”). and the…second optical attention module performs, based on the trained attention-aware optical neural network,…and spatial transmittance modulation of multi-dimensional sparse features extracted by the…first optical attention module to obtain a final spatial light output; (Wu, pg. 3, “In a second aspect, an embodiment of the present application provides an attention-fused photoelectric computing method of an attention fused photoelectric computing system using the foregoing attention-fused photoelectric computing system, where the attention-fused photoelectric computing method includes: the electronic neural network performs feature extraction and conversion on the received picture to be processed based on an attention mechanism to obtain phase distribution data; [of multi-dimensional sparse features extracted by the…first optical attention module] the spatial light modulator modulates single-wavelength laser generated by the laser based on the phase distribution data, [and the…second optical attention module performs, based on the trained attention-aware optical neural network,…and spatial transmittance modulation] and transmits the modulated laser to the photoelectric detection unit through a preset light path [to obtain a final spatial light output;]”). and the output module is configured to detect and identify the final spatial light output on an output plane to obtain a location of an object in a light field and an identification result. (Wu, pg. 6, “Referring to fig. 4, assuming that the target picture output based on the attention mechanism neural network A1 is a handwritten digital picture (i.e. a picture containing the number 3), the handwritten digital picture is taken as an input to the electronic deep neural network A2, and then the output of the electronic deep neural network A2 is a phase diagram, i.e. phase distribution data, which is input to the spatial light modulator 7; the spatial light modulator 7 performs wavefront (phase) modulation on the collimated laser light, so that the modulated laser light forms a light intensity distribution at the photodetection unit 8 to obtain a photodetection result, [and the output module is configured to detect and identify the final spatial light output on an output plane] for example, the photodetection unit 8 is divided into 10 regions, and the light intensity of the region corresponding to the numeral 3 on the handwritten digital picture is strongest [to obtain a location of an object in a light field and an identification result.].”). While Wu discloses an optical neural network that incorporates an attention mechanism, Wu does not explicitly teach: …comprising a multi-channel representation module,… …wherein: the multi-channel representation module is configured to encode, via a multi-spectral laser, an originally inputted target light field signal into coherent light having different wavelengths… bottom-up (BU) module top-down (TD) module different wavelengths spectral features Duan teaches: …comprising a multi-channel representation module,… (Duan, pg. 894 col. 1, “Here, we propose an optical multitask learning mono lithic system design that can simultaneously perform multiple classification tasks on different databases without mechanical movement by developing multi-wavelength D2NNs. Different from previous broadband D2NNs [22, 23], the wavelength dimension is exploited in this work to improve the computing throughput, which encodes different inputs into different wavelength channels and per forms photonic computing in both spatial and spectral dimensions […comprising a multi-channel representation module,…].”). …wherein: the multi-channel representation module is configured to encode, via a multi-spectral laser, an originally inputted target light field signal into coherent light having different wavelengths… (Duan, pg. 894 col. 1 and see Figure 1, “Here, we propose an optical multitask learning mono lithic system design that can simultaneously perform multiple classification tasks on different databases without mechanical movement by developing multi-wavelength D2NNs. Different from previous broadband D2NNs [22, 23], the wavelength dimension is exploited in this work to improve the computing throughput, which encodes different inputs into different wavelength channels and per forms photonic computing in both spatial and spectral dimensions […wherein: the multi-channel representation module is configured to encode, via a multi-spectral laser, an originally inputted target light field signal into coherent light having different wavelengths…].”). different wavelengths (Duan, pg. 894 col. 1 and see Figure 1, “Here, we propose an optical multitask learning mono lithic system design that can simultaneously perform multiple classification tasks on different databases without mechanical movement by developing multi-wavelength D2NNs [different wavelengths].”). spectral features (Duan, pg. 894 col. 1 and see Figure 1, “Different from previous broadband D2NNs [22, 23], the wavelength dimension is exploited in this work to improve the computing throughput, which encodes different inputs into different wavelength channels and per forms photonic computing in both spatial and spectral dimensions [spectral features].”). Wu and Duan are both in the same field of endeavor (i.e. optical neural networks). It would have been obvious for a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Wu and Duan to teach the above limitation(s). The motivation for doing so is that using multiple channels for different wavelengths allows for improved processing (cf. Duan, abstract, “By encoding multitask inputs into multiwavelength channels, the system can increase the computing throughput and significantly alleviate the competition to perform multiple tasks in parallel with high accuracy.”). While Wu in view of Duan teaches a multiwavelength optical neural network with an attention mechanism, the combination does not explicitly teach: bottom-up (BU) module top-down (TD) module Cao teaches: bottom-up (BU) module (Cao, pg. 2956 col. 2, “As shown in Figure 1, during the feedforward stage, the proposed networks perform inference from input images in a bottom-up manner as traditional Convolutional Networks [bottom-up (BU) module]; while in feedback loops, it sets up high-level semantic labels, (e.g., outputs of class nodes) as the “goal” in visual search to infer the activation status of hidden layer neurons.”). top-down (TD) module (Cao, pg. 2956 col. 2, “As shown in Figure 1, during the feedforward stage, the proposed networks perform inference from input images in a bottom-up manner as traditional Convolutional Networks; while in feedback loops, it sets up high-level semantic labels, (e.g., outputs of class nodes) as the “goal” in visual search to infer the activation status of hidden layer neurons [top-down (TD) module].”). Wu, in view of Duan, and Cao are both in the same field of endeavor (i.e. attention in neural networks). It would have been obvious for a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Wu, in view of Duan, and Cao to teach the above limitation(s). The motivation for doing so is that combining bottom-up and top-down attention improves a model’s decision making (cf. Cao, pg. 2956 col. 2, “From a machine learning perspective, the proposed feedback networks add extra flexibility to Convolutional Networks, to help in capturing visual attention and improving feature detection.”). Allowable Subject Matter Claims 2-9 would be allowable if rewritten or amended to overcome the rejection(s) under 35 U.S.C. 112(b) set forth in this Office action and to include all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for indication of allowable subject matter: Regarding claim 2, Below are the closest cited references, each of which disclose various aspects of the claimed invention: Katsuki, et al., “Bottom-Up and Top-Down Attention: Different Processes and Overlapping Neural Systems” discloses a background of bottom-up and top-down attention in the field of cognitive understanding and how that differ and work together. Katsuki discloses that bottom-up attention is the part of the neural system that is automatic and reactive to the stimuli in the environment like a reaction or impulse. Katsuki discloses that top-down attention is the part of the neural system that is goal-driven. While Katsuki teaches the definition of the terms bottom-up and top-down attention, Katsuki’s teachings are directed to the physical brain and is not directed to optical neural networks. Therefore, Katsuki does not explicitly teach wherein the BU optical attention module comprises a first BU optical attention module and a second BU optical attention module, the TD optical attention module takes output features of the first BU optical attention module as input, processes the input and feeds back to the second BU optical attention module to adjust the second BU optical attention module. Anderson, et al., “Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering” discloses a neural network model that uses top-down and bottom-up attention for visual question answering. The bottom-up attention determines the image regions and the top-down attention determines the feature weightings. While Anderson teaches using bottom-up and top-down attention for a visual neural network task, Anderson does not explicitly teach wherein the BU optical attention module comprises a first BU optical attention module and a second BU optical attention module, the TD optical attention module takes output features of the first BU optical attention module as input, processes the input and feeds back to the second BU optical attention module to adjust the second BU optical attention module. Rasolzadeh, et al., “An Attentional System Combining Top-Down and Bottom-Up Influences” discloses a neural network system that combines bottom-up and top-down attention for computer vision tasks. The paper equates bottom-up attention as being an exploration mode where a model scans a scene with no influence and purely based on instinct. The paper then equates top-down attention as a driving force that dictates where the model scans with intent. The two attention mechanisms are combined to aid in creating biased saliency maps for scene understanding. While Rasolzadeh teaches combining bottom-up and top-down attention, Rasolzadeh does not explicitly teach wherein the BU optical attention module comprises a first BU optical attention module and a second BU optical attention module, the TD optical attention module takes output features of the first BU optical attention module as input, processes the input and feeds back to the second BU optical attention module to adjust the second BU optical attention module. Mittal, et al., “Learning to Combine Top-Down and Bottom-Up Signals in Recurrent Neural Networks with Attention over Modules” discloses using top-down and bottom-up attention signals in a recurrent neural network to improve an recurrent neural network’s reliability over sequence data. While Mittal teaches combining top-down and bottom-up attention in a recurrent neural network, Mittal does not explicitly teach wherein the BU optical attention module comprises a first BU optical attention module and a second BU optical attention module, the TD optical attention module takes output features of the first BU optical attention module as input, processes the input and feeds back to the second BU optical attention module to adjust the second BU optical attention module. While the above prior arts disclose the aforementioned concepts, however, none of the prior arts, individually or in reasonable combination, discloses all the limitations in the manner recited in claim 2. Specifically, the claim requires wherein the BU optical attention module comprises a first BU optical attention module and a second BU optical attention module, the TD optical attention module takes output features of the first BU optical attention module as input, processes the input and feeds back to the second BU optical attention module to adjust the second BU optical attention module. While the references cited above mention aspects of a bottom-up attention module or a top-down attention module, they do not recite wherein the BU optical attention module comprises a first BU optical attention module and a second BU optical attention module, the TD optical attention module takes output features of the first BU optical attention module as input, processes the input and feeds back to the second BU optical attention module to adjust the second BU optical attention module. Therefore, claim 2 is allowable over the prior art. Regarding claims 3, 5-6, and 8-9, the claims are also considered allowable over the prior art at least by virtue of their dependence on claim 2. Regarding claim 4, Below are the closest cited references, each of which disclose various aspects of the claimed invention: Xu, et al., “A multichannel optical computing architecture for advanced machine vision” discloses an optical neural network named Monet: a multichannel optical neural network architecture for a universal multiple input multiple channel optical computing based on a novel projection interference prediction framework where the inter and intra channel connections are mapped to optical interference and diffraction. While Xu teaches the use of an optical neural network for diffraction with optical components, Xu does not explicitly teach wherein each unit of a metasurface-based optical filter consists of two layers of film, a first layer is a GeSbTe (GST) unit and a second layer is an intensity mask unit, the GST unit comprises an amorphous state and a crystalline state corresponding to different transmittance spectrums, and the amorphous state and the crystalline state are switched instantaneously by converting light. Yu, et al., “Optical Diffractive Convolutional Neural Networks Implemented in an All-Optical Way” discloses a system that combines a diffractive optical neural network with a convolutional neural network to perform image processing tasks. For the convolution, the system uses a 4f system because it is well-known that a lens can perform the two-dimensional Fourier transform at the speed of light. Therefore, an optical 4f system is used to achieve complicated convolutional operations, thereby reducing computational costs. While Yu teaches an optical neural network using optical components, Yu does not explicitly teach wherein each unit of a metasurface-based optical filter consists of two layers of film, a first layer is a GeSbTe (GST) unit and a second layer is an intensity mask unit, the GST unit comprises an amorphous state and a crystalline state corresponding to different transmittance spectrums, and the amorphous state and the crystalline state are switched instantaneously by converting light. Gao, et al., US20230359879A1 discloses a multipath deep diffractive neural network that comprises a first optical path for performing a first task and a second optical path for performing a second task. The system improves the efficiency of a diffractive neural network by incorporating multi-task learning into the diffractive model’s training. While Gao teaches the use of a optical neural network with optical components, Gao does not explicitly teach wherein each unit of a metasurface-based optical filter consists of two layers of film, a first layer is a GeSbTe (GST) unit and a second layer is an intensity mask unit, the GST unit comprises an amorphous state and a crystalline state corresponding to different transmittance spectrums, and the amorphous state and the crystalline state are switched instantaneously by converting light. While the above prior arts disclose the aforementioned concepts, however, none of the prior arts, individually or in reasonable combination, discloses all the limitations in the manner recited in claim 4. Specifically, the claim requires wherein each unit of a metasurface-based optical filter consists of two layers of film, a first layer is a GeSbTe (GST) unit and a second layer is an intensity mask unit, the GST unit comprises an amorphous state and a crystalline state corresponding to different transmittance spectrums, and the amorphous state and the crystalline state are switched instantaneously by converting light. While the references cited above mention aspects of using an optical filter, they do not recite wherein each unit of a metasurface-based optical filter consists of two layers of film, a first layer is a GeSbTe (GST) unit and a second layer is an intensity mask unit, the GST unit comprises an amorphous state and a crystalline state corresponding to different transmittance spectrums, and the amorphous state and the crystalline state are switched instantaneously by converting light. Therefore, claim 4 is allowable over the prior art. Regarding claim 7, Below are the closest cited references, each of which disclose various aspects of the claimed invention: Xu, et al., “A multichannel optical computing architecture for advanced machine vision” discloses an optical neural network named Monet: a multichannel optical neural network architecture for a universal multiple input multiple channel optical computing based on a novel projection interference prediction framework where the inter and intra channel connections are mapped to optical interference and diffraction. Monet is trained with a combined loss function of a mixed mean squared error and structural similarity index metric. While Xu teaches the use of an optical neural network for diffraction trained with a loss function, Xu does not explicitly teach using a loss function for training an attention optical neural network that is based on an optical Fourier transform performed twice and a resulting loss is propagated backward to optimize spectral and spatial coefficients of BU and TD branches. Yu, et al., “Optical Diffractive Convolutional Neural Networks Implemented in an All-Optical Way” discloses a system that combines a diffractive optical neural network with a convolutional neural network to perform image processing tasks. For the convolution, the system uses a 4f system because it is well-known that a lens can perform the two-dimensional Fourier transform at the speed of light. The system is trained with a loss function based on a mean square error for classification training. While Yu teaches training an optical neural network using a loss function, Yu does not explicitly teach using a loss function for training an attention optical neural network that is based on an optical Fourier transform performed twice and a resulting loss is propagated backward to optimize spectral and spatial coefficients of BU and TD branches. Gao, et al., US20230359879A1 discloses a multipath deep diffractive neural network that comprises a first optical path for performing a first task and a second optical path for performing a second task. The system improves the efficiency of a diffractive neural network by incorporating multi-task learning into the diffractive model’s training. The multipath diffractive neural network is trained with a multi-task loss function. While Gao teaches the use of a loss function for training an optical neural network, Gao does not explicitly teach using a loss function for training an attention optical neural network that is based on an optical Fourier transform performed twice and a resulting loss is propagated backward to optimize spectral and spatial coefficients of BU and TD branches. While the above prior arts disclose the aforementioned concepts, however, none of the prior arts, individually or in reasonable combination, discloses all the limitations in the manner recited in claim 7. Specifically, the claim requires a specific loss function used for network training of an attention optical neural network that is based on an optical Fourier transform performed twice and a resulting loss is propagated backward to optimize spectral and spatial coefficients of BU and TD branches. While the references cited above mention aspects of using an loss function to train optical neural networks, they do not recite the specific loss function and resulting loss as claimed in claim 7. Therefore, claim 7 is allowable over the prior art. Claims 10-15 are allowable over the prior art would be allowable if rewritten or amended to overcome the rejections under 35 U.S.C. 112(b) set forth in this Office Action. The following is an examiner’s statement of reasons for allowance: Claims 10-15 are considered allowable over the art since when reading the claims in light of the specification, per MPEP 2111.01, the references of record alone or in combination do not disclose or suggest the limitations found within claim 10. The claim as a whole with regards to technical features recited by the claim limitations are directed to: (exemplar claim 10 limitations): “a first bottom-up (BU) optical attention module, a second BU optical attention module, an optical filter, a top-down (TD) optical attention module…after propagation, the TD optical attention module takes a multidimensional sparse feature output by the first BU optical attention module as input, processes the input and feeds back to the second BU optical attention module to adjust the second BU optical attention module; the optical filter is used to control connections between optical neurons in the second BU optical attention module and perform spectral and spatial modulation on the optical neurons;” Below are the closest prior arts, each of which disclose various aspects of the claimed invention: Katsuki, et al., “Bottom-Up and Top-Down Attention: Different Processes and Overlapping Neural Systems” discloses a background of bottom-up and top-down attention in the field of cognitive understanding and how that differ and work together. Katsuki discloses that bottom-up attention is the part of the neural system that is automatic and reactive to the stimuli in the environment like a reaction or impulse. Katsuki discloses that top-down attention is the part of the neural system that is goal-driven. While Katsuki teaches the definition of the terms bottom-up and top-down attention, Katsuki’s teachings are directed to the physical brain and is not directed to optical neural networks. Therefore, Katsuki does not explicitly teach a TD optical attention module that takes a multidimensional sparse feature output by a first BU optical attention module as input, processes the input and feeds back to a second BU optical attention module to adjust the second BU optical attention module. Anderson, et al., “Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering” discloses a neural network model that uses top-down and bottom-up attention for visual question answering. The bottom-up attention determines the image regions and the top-down attention determines the feature weightings. While Anderson teaches using bottom-up and top-down attention for a visual neural network task, Anderson does not explicitly teach a TD optical attention module that takes a multidimensional sparse feature output by a first BU optical attention module as input, processes the input and feeds back to a second BU optical attention module to adjust the second BU optical attention module. Rasolzadeh, et al., “An Attentional System Combining Top-Down and Bottom-Up Influences” discloses a neural network system that combines bottom-up and top-down attention for computer vision tasks. The paper equates bottom-up attention as being an exploration mode where a model scans a scene with no influence and purely based on instinct. The paper then equates top-down attention as a driving force that dictates where the model scans with intent. The two attention mechanisms are combined to aid in creating biased saliency maps for scene understanding. While Rasolzadeh teaches combining bottom-up and top-down attention, Rasolzadeh does not explicitly teach a TD optical attention module that takes a multidimensional sparse feature output by a first BU optical attention module as input, processes the input and feeds back to a second BU optical attention module to adjust the second BU optical attention module. Mittal, et al., “Learning to Combine Top-Down and Bottom-Up Signals in Recurrent Neural Networks with Attention over Modules” discloses using top-down and bottom-up attention signals in a recurrent neural network to improve an recurrent neural network’s reliability over sequence data. While Mittal teaches combining top-down and bottom-up attention in a recurrent neural network, Mittal does not explicitly teach a TD optical attention module that takes a multidimensional sparse feature output by a first BU optical attention module as input, processes the input and feeds back to a second BU optical attention module to adjust the second BU optical attention module. In summary, the references made of record, fail to disclose the required claimed technical features recited by the noted claim limitations as a whole and the related dependent claims are found allowable over the art. Furthermore, the references of record alone or in combination fail to disclose or suggest the combination of limitations found within the independent claim as a whole without hindsight reasoning. The dependent claims, being further limiting to the independent claim, definite, and enable by the Specification are also allowed over the prior art. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Ozcan, et al., US20230401447A1 discloses the creation of a special type of optical neural network called a diffractive optical neural network. Any inquiry concerning this communication or earlier communications from the examiner should be directed to NICHOLAS S WU whose telephone number is (571)270-0939. The examiner can normally be reached Monday - Friday 8:00 am - 4:00 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle Bechtold can be reached at 571-431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /N.S.W./Examiner, Art Unit 2148 /MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148
Read full office action

Prosecution Timeline

Jun 20, 2024
Application Filed
Sep 11, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737619
OPTIMIZING ALGORITHMS FOR HARDWARE DEVICES
3y 11m to grant Granted Sep 15, 2026
Patent 12725027
PROACTIVE ANOMALY DETECTION
5y 9m to grant Granted Sep 01, 2026
Patent 12725405
LEARNING APPARATUS, ESTIMATION APPARATUS, DATA GENERATION APPARATUS, LEARNING METHOD, AND COMPUTER-READABLE STORAGE MEDIUM STORING A LEARNING PROGRAM
5y 0m to grant Granted Sep 01, 2026
Patent 12645939
SPIKING NEURAL NETWORK
3y 5m to grant Granted Jun 02, 2026
Patent 12619880
METHODS, DEVICES AND MEDIA FOR RE-WEIGHTING TO IMPROVE KNOWLEDGE DISTILLATION
5y 0m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
52%
Grant Probability
84%
With Interview (+31.4%)
4y 0m (~1y 9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 63 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month