Prosecution Insights
Last updated: October 04, 2026
Application No. 18/777,007

reconstruction method and system of aerosol chemical components based on CNN-BiLSTM-BO

Non-Final OA §103§112
Filed
Jul 18, 2024
Priority
Apr 30, 2024 — CN 2024105442971
Examiner
DASGUPTA, SHOURJO
Art Unit
Tech Center
Assignee
Institute Of Atmospheric Physics Chinese Academy Of Sciences
OA Round
1 (Non-Final)
65%
Grant Probability
Favorable
1-2
OA Rounds
1y 2m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 65% — above average
65%
Career Allowance Rate
303 granted / 465 resolved
+5.2% vs TC avg
Strong +39% interview lift
Without
With
+39.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
26 currently pending
Career history
491
Total Applications
across all art units

Statute-Specific Performance

§101
12.6%
-27.4% vs TC avg
§103
57.9%
+17.9% vs TC avg
§102
11.0%
-29.0% vs TC avg
§112
16.3%
-23.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 465 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f), because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “the data preprocessing module”, “the feature analysis and evaluation module”, and “the model training optimization module” in claim 8 (the latter is also in claim 10). “the characteristic analysis and evaluation module” in claim 9. Because these claim limitations are being interpreted under 35 U.S.C. 112(f), they are being interpreted to cover a corresponding structure described in the specification as performing the claimed function, and equivalents thereof. However, a review of Applicants’ specification finds a lack of any teaching that would provide the corresponding structure to enable the recitation of these module elements recited in the present claims. See the Examiner’s rejection under 35 U.S.C. 112(a). If Applicants do not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f), Applicants may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f). Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. Claims 8-10 are rejected under 35 U.S.C. 112(a) as failing to comply with the enablement requirement. The claims contain subject matter which was not described in the specification in such a way as to enable one skilled in the art to which it pertains, or with which it is most nearly connected, to make and/or use the invention. Specifically, the claims, as noted above in the section titled Claim Interpretation, invoke means plus function claiming, under 35 U.S.C. 112(f), but are not otherwise supported in the specification or the claims themselves with the appropriate hardware/processing structures to successfully enable that sort of claiming. The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. Claims 2-10 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention. Regarding claim 2, the claim features improper antecedent basis for the recited directionality, see e.g., “… Each time step of the multi-source environment observation data is independently convolved along the horizontal and vertical directions …”, and further it is not clear what the directionality is in respect to. Similarly, the claim features improper antecedent basis for recited layer and dimension instances, see e.g., “… the time series structure of the multi-source environment observation data is reconstructed using the sequence expansion layer, and the spatial dimension is folded into the variable dimension by the flattening layer.” The effect is that the present claim 2 is vague and indefinite. Accordingly, the aforementioned claim, and the claims which depend from it – such as claims 8-10 – are rejected on this basis. Claim 3 is similarly rejected, as it features improper antecedent basis for recitations of “the bidirectional long short-term memory layer”, “the fully connected layer”, and “the regression output layer.” The effect is that the present claim 3 is vague and indefinite. Accordingly, the aforementioned claim, and the claims which depend from it – such as claims 8-10 – are rejected on this basis. Claim 4 is rejected as featuring improper antecedent basis for the following recitations: “the objective function of the CNN-BiLSTM model”, “the initial decision vector”, “the value of the objective function is calculated and the prior probability distribution is obtained, and the maximum number of iterations and calculation time are controlled by the control parameters”, “the likelihood distribution of the Gaussian process regression model of the objective function is updated by the actual observed values”, and “based on the posterior probability distribution, the acquisition function value is maximized to determine the next optimal sampling point.” The effect is that the present claim 4 is vague and indefinite. Accordingly, the aforementioned claim, and the claims which depend from it – such as claims 8-10 – are rejected on this basis. Regarding claims 5-6, both claims recite a step 211, which has no prior mention in the claim and bears no obvious logical relation to the aforementioned prior steps S1-S3 recited in claim 1, from which these claims indirectly depend. Hence, the step, as recited in this manner, is vague and indefinite. On this basis, the present claims 5-6, and the claims which depend from it – such as claims 7-10 – are rejected on this basis. Regarding claim 8, the claim is rejected for improper antecedent basis for recitations of “the data preprocessing module”, “the feature analysis and evaluation module”, and “the model training optimization module.” The effect is that the present claim 8 is vague and indefinite. Accordingly, the aforementioned claim, and the claims which depend from it – such as claims 9-10 – are rejected on this basis. Regarding claim 9, the claim is rejected for improper antecedent basis for recitations of “the characteristic analysis and evaluation module.” The effect is that the present claim 9 is vague and indefinite. Accordingly, the aforementioned claim, and claim 10 which depends from it, are rejected on this basis. Claim Objections Claims 8-10 are objected to under 37 CFR 1.75(c) as being in improper form because a multiple dependent claim should refer to other claims in the alternative only, and/or, cannot depend from any other multiple dependent claim. See MPEP § 608.01(n), referring explicitly to 35 U.S.C. 112(e). For Applicants’ benefit, the Examiner notes that Applicants’ language in claim 8, and via dependency inherited by claims 9-10, resembles the example provided in MPEP 608.01(n) of an “Unacceptable Multiple Dependent Claim Wording” because the 1. Claim Does Not Refer Back in the Alternative Only”: PNG media_image1.png 94 392 media_image1.png Greyscale The Examiner recommends a claim amendment to claim 8 that falls within the compliant examples provided in that same section MPEP 608.01(n) under it’s heading A.Acceptable Multiple Dependent Claim Wording, which provide examples of multiple dependent claiming that properly refer back in the alternative to more than one preceding claim. Accordingly, in view of the guidance provided per MPEP 608.01(n) (and specifically, form paragraph 7.45), the claims 8-10 have not been further treated on the merits. However, the claims appear upon review feature explicit limitations/language that mirror those found in claims 1-3; hence, it would reasonable for Applicants to expect prior art rejections on the same grounds as those provided by the Examiner with respect to claims 1-3 for these claims 8-10. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office Action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim 1 is rejected under 35 U.S.C. 103 as being unpatentable over Non-Patent Literature “Design a regional and multistep air quality forecast model based on deep learning and domain knowledge” (“Mo”) in view of Non-Patent Literature “An Intelligent Bayesian Optimization with Stacked BiLSTM Model for Air Quality Index Prediction” (“Manikandan”). Regarding claim 1, MO teaches A reconstruction method of aerosol chemical components based on CNN-BILSTM (Abstract @ p.1: “… Thirdly, the features of pollution source and meteorological condition are learned and predicted by CNN-BiLSTM-Attention model, the integrated model of convolutional neural network and Bidirectional long short-term memory network based on Sequence to Sequence framework with Attention mechanism, and then Convolutional Long Short-Term Memory Neural Network (ConvLSTM) integrates the two determinant features to obtain predicted pollutant concentration …”; and section 2.3.1 teaching a CNN specifically), comprising the following steps: Step S1, collecting multi-source environmental observation data through observation equipment (section 2.1 @ p.3: “Beijing-Tianjin-Hebei Air Pollution Transmission Channel (“2+26” cities) as Figure 1 shows is chosen as study area for its important strategic position and space correlation of air pollution.”; and section 2.2 @ p.3: “As to meteorological elements, we collect data of wind speed, temperature, relative humidity and precipitation based on the domain knowledge of Atmospheric Science. Hourly monitoring data of 6 air pollutants and 4 meteorological elements during 2017.12–2019.2 of 28 cities are collected”; and sections 3.1.1 and 3.1.2 @ p.6, for pollution source and meteorological condition, respectively), and conduct pre-processing (section 3.2 @ p.7: “min-max normalization is applied to input data”, and see also the “data preprocessing” stage in Figure 6 found on p. 9) and extract key feature variables (section 3.2 @ p.7: “output data are processed by reverse normalization for the evaluation of model’s performance” and see also the “feature selection” stage in Figure 6 found on p. 9, where the features are applied in a supervised learning process to arrive at a forecast strategy in the stage rendered just below in the same Figure (and see similar feature-based processing in Figure 7’s workflow, e.g. as being fed into the CNN-BiLSTM models)); Step S2, inputting the pre-processed multi-source environment observation data into the CNN-BILSTM model for feature analysis (Abstract @ p.1: “… Thirdly, the features of pollution source and meteorological condition are learned and predicted by CNN-BiLSTM-Attention model, the integrated model of convolutional neural network and Bidirectional long short-term memory network based on Sequence to Sequence framework with Attention mechanism, and then Convolutional Long Short-Term Memory Neural Network (ConvLSTM) integrates the two determinant features to obtain predicted pollutant concentration …”, which the Examiner reasons includes feature analysis to perform the prediction as taught (and see also the feeding of feature information into the CNN-BiLSTM models as shown in p.9’s Figure 7)), wherein, convolutional neural network extracts spatial features (section. 3.1.1 @ p.6: “The time series of the above variables are mined by deep learning algorithm to learn their complex interactions and spatial-temporal variations.” (and see also “space feature” in p.9’s Figure 7 workflow model, as being fed as input into the top CNN-BiLSTM model representation)) and uses bidirectional long short-term memory neural network to process time series information (section 2.3.3 @ p.4 discussing the use of BiLSTM in relation to time-series data, and see also p.9’s Figure 7 workflow model), and …, a reconstructed model of aerosol chemical components is generated (“pollutant concentration” processing result in p. 9’s Figure 6, and also “predicted concentration” processing result in Figure 7); Step S3, after verifying the performance and stability of the reconstructed model (section 2.2 specifically @ p.4: “Data from 2017.12 to 2018.12 are used as training set, and the data of 2019.1 are used as validation set while the data of 2019.2 are selected as test set. The hyperparameters are tuned manually and determined based on the performance of model on validation set. In addition, in order to test the generalization ability of the model, the data of 2019.6 are supplemented as an additional test set.”), outputting the predicted results of the chemical components of the aerosol (“pollutant concentration” processing result in p. 9’s Figure 6, and also “predicted concentration” processing result in Figure 7). Mo alone does not teach the entirety of the limitation for the method that further adjusts hyperparameters of the CNN-BILSTM through Bayesian optimization algorithm. At best, Mo teaches manual hyperparameter tuning per its section 2.2. Rather, to more fully address the limitation, the Examiner relies upon MANIKANDAN to teach what Mo otherwise lacks, see e.g., Manikandan’s comparable BilLSTM-driven architecture for air quality analysis/prediction (Abstract, Fig. 1 on p. 1700)), where the process flow per Fig. 1 explicitly shows use of a Bayesian Optimization Algorithm, and where the last paragraph before section 2 (on the same p. 1700) teaches “In this study a Bayesian Optimization with Stacked deep learning based Air Quality Index Prediction (BOSDL-AQIP) approach. The goal of the BOSDL-AQIP approach lies in the effectual identification and classification of Air Quality (AQ) into multiple class labels. To attain this, the presented BOSDL AQIP technique employs min-max normalization for data scaling purposes. Next, the BOSDL-AQIP system utilizes Stacked Bidirectional Long Short-Term Memory (SBiLSTM) technique for prediction process. Moreover, BO technique was utilized for adjusting the hyperparameter values of the SBiLSTM technique and thereby improve the predictive outputs.” Both Mo and Manikandan teach the design and implementation of similar BiLSTM-type architectures to accomplish a similar data processing/prediction task relating to air quality. Hence, they are similarly directed and therefore analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the use of Manikandan’s Bayesian Optimization algorithm approach to improve the more manual and hence labor/time intensive hyperparameter tuning that Mo teaches, with a reasonable expectation of success, for purposes of saving time, labor, cost, etc. considerations by more fully automating the hyperparameter tuning and hence the model preparation aspects of Mo’s framework via Manikandan’s approach. Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Mo in view of Manikandan and further in view of Non-Patent Literature “Air-pollution prediction in smart city, deep learning approach” (“Bekkar”) and Non-Patent Literature “A Deep CNN-LSTM Model for Particulate Matter (PM2.5) Forecasting in Smart Cities” (“Huang”). Regarding claim 3, Mo in view of Manikandan teach the reconstruction method of aerosol chemical components described in claim 1, as discussed above, and further wherein in step S2, the bidirectional long short-term memory neural network is used to process time series information (Mo: section 2.3.3 @ p.4 discussing the use of BiLSTM in relation to time-series data, and see also p.9’s Figure 7 workflow model), including that the bidirectional long short-term memory layer inputs the time characteristics of the data from the convolutional neural network in a forward and backward way (Mo, same section 2.3.3 as just referenced above: “The input at each time will be provided to the forward and backward LSTM at the same time.”), and simultaneously captures the past and future information (Mo section 2.3.3, still: “However, in some cases, the output of the current moment is closely related to the history and future state, so considering the context information at the same time is conducive to comprehensive judgment. Bidirectional long short-term memory network (BiLSTM) solves this problem. It consists of two unidirectional LSTM … The input at each time will be provided to the forward and backward LSTM at the same time. The two hidden layers calculate the state and output independently.”) … the output result of the bidirectional long short term memory layer is dimensionally transformed to match the target output size using the fully connected layer and finally the regression output layer outputs the regression estimate of the aerosol chemical components (Mo: Figure 2, rightmost two columns in the Fig. showing an adaptation of more neurons to less as it determines its final layer output, and where this evidences a sort of regression to convert the varied inputs detailed above per Mo to amount to a concentration/composition prediction as taught). Neither Mo nor Manikandan explicitly teach the further limitation to randomly, remove a certain percentage of neurons by dropping layers. Rather, the Examiner relies upon Bekkar and Huang to teach what they otherwise lack, see e.g., the following: In a comparable air quality prediction framework, Bekkar teaches 10% neuron dropout. See p. 14, in the bullet point just above Fig. 11: “In LSTM, GRU, BI-LSTM, and BI-GRU each network, four hidden layers are used, 200 units in the first layer, then 100 in the second layer, and 50 units in the last two layers. The last layer output of the network is linked to a dense layer with a single output neuron. Between the layers, a dropout equal to 10% is used.” In a comparable air quality prediction framework, Huang teaches random neuron dropout. See p. 8’s description of dropout as a way to avoid overfitting, specifically by randomly stopping the operation of a neuron. Bekkar and Huang are similarly directed to Mo and Manikandan, both in terms of the types of machine learning architecture but also type of data and problem to be solved. Hence, the references are similarly directed and therefore analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate a dropout aspect, as is known in the state of the art – for example, Bekkar as clarified by Huang, with a reasonable expectation of success, into the modified Mo in view of Manikandan framework, for purposes of mitigating overfitting of the model relative to the data. Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Mo in view of Manikandan and further in view of Non-Patent Literature “Convolutional Neural Networks for Multi-Step Time Series Forecasting” (“Brownlee/CNN”), Non-Patent Literature “Long-Short-Term-Memory-Based Deep Stacked Sequence-to-Sequence Autoencoder for Health Prediction of Industrial Workers in Closed Environments Based on Wearable Devices” (“Xu”), Non-Patent Literature “RECURRENT BATCH NORMALIZATION” (“Cooijmans”), and Non-Patent Literature “Flatten Layer — Implementation, Advantage and Disadvantages” (“Srivatsavaya”). Regarding claim 2, Mo in view of Manikandan teach The reconstruction method of aerosol chemical components described in claim 1, as discussed above, wherein in step S2, the convolutional neural network extracts spatial features (Mo’s section. 3.1.1 @ p.6: “The time series of the above variables are mined by deep learning algorithm to learn their complex interactions and spatial-temporal variations.” (and see also “space feature” in p.9’s Figure 7 workflow model, as being fed as input into the top CNN-BiLSTM model representation)), including folding the multivariable time series of the multi-source environmental observation data … to complete independent convolution operations at each time step (Mo’s feature input is explicitly multi-variate (section 3.3, and also FIG. 7’s left-most two boxes that are essentially inputs and diverse) and time-series (2.3.3, and also FIG. 6’s supervised learning via a sliding time window (right-most sub-stage in the feature selection stage))), and local features are extracted by convolution kernel (which the Examiner reasons is inherent to any sort of convolutional neural network’s fundamental operation, and certainly is taught by Mo when Figure 7’s pipeline of [feature selection and construction]-[CNN-BiLSTM-Attention]-[Pollution source / Meteorological condition] is understood, as that sequence clearly teaches the convolution of certain features from the data to arrive at a CNN processing result) and normalized (Mo’s section 3.5.1 @ p.8 discussing the normalization via NMB of the processed data). Applicants’ claim more fully recites including “folding the multivariable time series of the multi-source environmental observation data into a multivariable array to complete independent convolution operations at each time step.” While Mo in view of Manikandan, teach the above more generally, it is not clear whether an array as specifically recited is used, as obvious as it may be to use such a data structure to encompass/manage different numerical values. Rather, the Examiner relies upon BROWNLEE to more explicitly provide that level of detail which Mo etc. may otherwise lack, see e.g., Brownlee’s page 13, under the heading Multi-step Time Series Forecasting with a Multichannel CNN, in which diverse data is each a different channel of input provided to the CNN, such that together they are considered in an arrayed manner which the Examiner likens to the present recitation. The claim further recites, in part, limitations wherein each time step of the multi-source environment observation data is independently convolved, which the Examiner reasons is taught by Mo: section 4.1 @ p.7 discussing multi-step results in relation to its convolutional/iterative operation of the time-series data. Regarding the further limitation of convolving along the horizontal and vertical directions, the Examiner relies upon Brownlee again to teach what Mo etc. otherwise lack, see e.g., Brownlee’s p.7 discussing “Convolutional layers read an input, such as a 2D image or a 1D signal using a kernel that reads in small segments at a time and steps across the entire input field. Each read results in an interpretation of the input that is projected onto a filter map and represents an interpretation of the input.”, where the Examiner reasons that 2D horizontal and vertical direction operation of a CNN is known, e.g. as this reference teaches, and could be applied in the consideration and processing of the data that Mo contemplates, e.g., as generally discussed per claim 1. The Examiner reasons that Brownlee’s discussion, as relied-upon, provides an analogous input mechanism for a CNN, and is apt to the similar mechanism of Mo’s Figure 7’s two left-most boxes (e.g., indicating diverse multi-source inputs of pollution concentration and meteorological element variables), which are also subject to processing by a CNN per that reference’s framework. Hence, Brownlee is a broader and more generalized problem-agnostic teaching of an input and processing mechanism that the Examiner reasons Mo specifically contemplates for its problem space. Accordingly, the references are analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the management of diverse inputs per an array or similar structure, as Brownlee contemplates, to the input processing and management aspect of Mo’s framework, with a reasonable expectation of success, so that a more detailed implementation that generally works for the same type of similar input aspect can be more specially applied to Mo’s more focused framework. The claim further recites additional limitations that the Examiner does not believe Mo or Manikandan or Brownlee teach, and rather the Examiner relies upon additional prior art references: nonlinear processing is performed by batch normalization and correction of linear elements (Cooijmans: “We propose a reparameterization of LSTM that brings the benefits of batch normalization to recurrent neural networks. Whereas previous works only apply batch normalization to the input-to-hidden transformation of RNNs, we demonstrate that it is both possible and beneficial to batch-normalize the hidden-to-hidden transition, thereby reducing internal covariate shift between time steps.” (Abstract), which is further clarified in section 3 @ p.3 which details the incorporation of batch normalization into a LSTM network that better addresses covariate shift between time step intervals). Cooijmans teaches a similar architecture to that contemplated by Mo, and hence a similar time-series data processing aspect. Hence, it similarly directed and therefore analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Cooijmans’s batch normalization application for LSTMs to Mo’s modified framework, with a reasonable expectation of success, so as to more readily account for covariate shift and produce a more accurate model as a function of time and use. the time series structure of the multi-source environment observation data is reconstructed using the sequence expansion layer (Xu’s Abstract and sections 2.2, and 2.4, which teach the reconstruction of sequences in time series data processed by a LSTM to more meaningfully extract features from the data) Xu teaches a similar architecture to that contemplated by Mo, and hence a similar time-series data processing aspect. Hence, it similarly directed and therefore analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Xu’s sequence reconstruction aspect specific to LSTM networks processing time series data into Mo’s modified framework, with a reasonable expectation of success, so as to more enable more meaningful feature extraction as Xu contemplates. the spatial dimension is folded into the variable dimension by the flattening layer (Srivatsavaya’s pages 1-2: “The Flatten layer is a crucial component in neural network architectures, especially when transitioning from convolutional layers (Conv2D) or recurrent layers (LSTM, GRU) to fully connected layers (Dense) in deep learning models. Its primary use is to reshape the input data or tensor into a one-dimensional (1D) vector so that it can be fed into subsequent fully connected layers …” and more specifically “Transition from Convolutional Layers to Fully Connected Layers: In convolutional neural networks (CNNs), convolutional and pooling layers typically operate on multi-dimensional tensors (e.g., 2D for images). However, before passing data to fully connected layers, which require 1D inputs, we need to flatten the tensor. The Flatten layer serves this purpose by converting a 2D or 3D tensor into a 1D vector.”, which the Examiner reasons is a necessary step/layer in a framework such as Mo’s (as detailed in Mo’s Figures 6-7, featuring the similar transition from convolution layers to output layers that make use of the convolution result)). Srivatsavaya teaches a similar architecture to that contemplated by Mo, and hence a similar time-series data processing aspect. Hence, it similarly directed and therefore analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Srivatsavaya’s flattening aspect specific to LSTM networks processing time series data into Mo’s modified framework featuring a CNN-LSTM network, with a reasonable expectation of success, so as to transition from convolutional layer operation to final processing steps/stages that can readily make use of the convolutional layer results. Claims 4-7 are rejected under 35 U.S.C. 103 as being unpatentable over Mo in view of Manikandan and then Bekkar and Huang and further in view of Non-Patent Literature “Tune Experiment Hyperparameters by Using Bayesian Optimization” (“Matlab”), Non-Patent Literature “Hyperparameter Optimization With Random Search and Grid Search” (“Brownlee/HPO”), and Non-Patent Literature “Exploring Bayesian Optimization” (“Agnihotri”). Regarding claim 4, Mo in view of Manikandan and then Bekkar and Huang teach The reconstruction method of aerosol chemical components described in claim 3, as discussed above, and further wherein in step S2, the hyperparameters are adjusted by the Bayesian optimization algorithm (Manikandan modifying Mo, as discussed per claim 1), including, the objective function of the CNN-BiLSTM model is constructed (Mo for example teaches a framework, per Abstract @ p.1, that “… Thirdly, the features of pollution source and meteorological condition are learned and predicted by CNN-BiLSTM-Attention model, the integrated model of convolutional neural network and Bidirectional long short-term memory network based on Sequence to Sequence framework with Attention mechanism, and then Convolutional Long Short-Term Memory Neural Network (ConvLSTM) integrates the two determinant features to obtain predicted pollutant concentration …”; and section 2.3.1 teaching a CNN specifically, and the Examiner reasons that the design, training/implementation, etc. of these machine learning elements, both individually and together as arranged in the taught configuration, would comprise at least one if not many error/loss functions so that they can be appropriately trained to an acceptable performance level prior to deployment, reliance, etc.), and a set of initial hyperparameters are … selected in the decision space as the initial decision vector (as discussed per claim 1, Mo teaches manual hyperparameter tuning per its section 2.2, and Manikandan teaches use of a Bayesian Optimization Algorithm, such that (the last paragraph before section 2, on p. 1700) “… BO technique was utilized for adjusting the hyperparameter values of the SBiLSTM technique and thereby improve the predictive outputs.”), the value of the objective function is calculated and the prior probability distribution is obtained (the Examiner reasons that Mo in view of Manikandan, as discussed per claim 1, would necessarily involve the use and calculation of an objective function, as the Examiner has mentioned above, and also that as part of the Bayesian Optimization pr Manikandan, a prior probability distribution is involved), and the likelihood distribution of the Gaussian process regression model of the objective function is updated by the actual observed values, and the posterior probability distribution is derived to reflect the update of the objective function (the Examiner reasons this is inherent to Bayesian Optimization generally, and hence would be involved in accordance with Manikandan’s employment of the technique), … and the corresponding prior probability distribution is continued to be updated, iterating continuously until the combination of hyperparameters corresponding to the minimum value of the objective function is determined, and the hyperparameter combination is the corresponding optimal solution (again, the Examiner reasons this is inherent to Bayesian Optimization generally, and hence would be involved in accordance with Manikandan’s employment of the technique). Regarding the selection of hyperparameters, the aforementioned references do not teach a set of initial hyperparameters are randomly selected in the decision space as the initial decision vector. Rather, the Examiner relies upon BROWNLEE to teach what they lack, see e.g., Brownlee’s page 1 generally discussing hyperparameter selection process that contemplates inclusion of random search, with page 2 clarifying further that “A range of different optimization algorithms may be used, although two of the simplest and most common methods are random search and grid search.” And that “Random Search” is performed to “Define a search space as a bounded domain of hyperparameter values and randomly sample points in that domain.” Brownlee teaches a framework/technique for hyperparameter selection using approaches known and common in the state of the art. Similarly, both Mo and Manikandan contemplate similar hyperparameter selection and tuning steps, so that the relevant models used perform sufficiently, with Manikandan selected by the Examiner to improve upon Mo’s more manual approach. Hence, the references are similarly directed in the aspect of the machine learning modeling they address and seek to improve, and are therefore analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Manikandan’s HPO as improved by Brownlee’s random search technique, which is well known, to improve upon Mo’s manual aspect of the same training and model building task, with a reasonable expectation of success, so as to automate hyperparameter selection and model building more fully to offload work/costs from a human as involved. Mo etc. do not teach the further limitation of the maximum number of iterations and calculation time are controlled by the control parameters, and rather the Examine relies upon Matlab to teach what they lack, see e.g., Matlab’s framework for hyperparameter tuning using Bayesian Optimization, and specifically p. 2 showing a configuration screen where maximum number of trials and maximum time in seconds are provided as configuration parameters to perform the HPO using BO. Matlab teaches a framework/technique for hyperparameter selection/tuning as related to Bayesian Optimization, which is the approach Manikandan uses and was selected by the Examiner to improve upon Mo’s more manual approach. Hence, the references are similarly directed in the aspect of the machine learning modeling they address and seek to improve, and are therefore analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Matlab’s configurable settings to provide more user configurable and flexible inputs to drive the Bayesian Optimization process that Manikandan contemplates, for improving Mo’s framework with a reasonable expectation of success, so as to better automate hyperparameter selection and model building more fully to offload work/costs from a human as involved. Mo etc. do not teach the further limitation that, based on the posterior probability distribution, the acquisition function value is maximized to determine the next optimal sampling point, and rather the Examine relies upon AGNIHOTRI to teach what they lack, see e.g., Agnihotri’s p. 3 teaching active learning to determine the “next query point”, as part of the iterative process to train the model until convergence or budget met. Agnihotri teaches a framework/technique for hyperparameter selection/tuning as related to Bayesian Optimization, which is the approach Manikandan uses and was selected by the Examiner to improve upon Mo’s more manual approach. Hence, the references are similarly directed in the aspect of the machine learning modeling they address and seek to improve, and are therefore analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Agnihotri’s active learning aspect to make efficient the Bayesian Optimization process that Manikandan contemplates, for improving Mo’s framework with a reasonable expectation of success, so as to better automate hyperparameter selection and model building more fully to offload work/costs from a human as involved. Regarding claim 5, Mo in view of Manikandan and then Bekkar and Huang and then further in view of Brownlee, Matlab, and Agnihotri teach The reconstruction method of aerosol chemical components mentioned in claim 4, as discussed above, further wherein in step S211, the objective function of the CNN-BILSTM model is constructed, and a set of initial hyperparameters are randomly selected in the decision space as the initial decision vector, the value of the objective function is calculated and the prior probability distribution is obtained, including, p (f (x ❘ input) ❘ Y) = p (Y ❘ f (x ❘ input)) p (f (x ❘ input)) / p (Y) : x* = argmin f (x ❘ input), x ∈ X ⊆ Rd: wherein, x is the decision vector composed of d hyperparameters, X is the decision space, input is the input data, x* is the optimal hyperparameter combination vector, f (x | input) is the objective function, Y is the actual observation value, p (Y | f (x | input)) is the likelihood distribution, p (f (x | input)) is the prior probability distribution, that is, the estimate of f (x | input), p(Y) is the edge likelihood distribution, p ( f (x | input) | Y) is the posterior probability distribution, that is, the confidence coefficient of f (x | input). The Examiner reasons that these are merely the discrete steps involved in any Bayesian Optimization process, such as the one Manikandan is cited for by the Examiner for example in relation to claim 1. Regarding claim 6, Mo in view of Manikandan and then Bekkar and Huang and then further in view of Brownlee, Matlab, and Agnihotri teach The reconstruction method of aerosol chemical components described in claim 5, as discussed above, and further wherein in step S211, the likelihood distribution of the Gaussian process regression model of the objective function is updated by actual observed values, and a posterior probability distribution is derived from it to reflect the update of the objective function, including, f (x ❘ input) ~GP (mean (x), k (x, x′; θ)); k (x,x′;θ) = θ1 (1 + {[squareroot (5) ❘ x - x′ ❘] / θ2} + {(5 (x-x′)2) / 3 θ2} exp {squareroot (5) ❘ x - x′ ❘] / θ2}; θ1 = σf2 , σf > 0; θ2 = σ1, σ1 > 0; wherein, GP(mean(x),k(x,x′;θ)) is the Gaussian process regression model, which is composed of mean value function(mean(x)) and covariance kernel function(k(x,x′;θ)), k(x,x′;θ) is The Matern 5/2 covariance kernel function, where x and x′ is expressed as any two coordinate points, θ is the kernel parameter vector, σf is the signal standard deviation, and σ1 is the feature length scale. The Examiner reasons that these are merely the discrete steps involved in any Bayesian Optimization process, such as the one Manikandan is cited for by the Examiner for example in relation to claim 1. Regarding claim 7, Mo in view of Manikandan and then Bekkar and Huang and then further in view of Brownlee, Matlab, and Agnihotri teach The reconstruction method of aerosol chemical components described in claim 6, as discussed above and further wherein based on the posterior probability distribution, the acquisition function value is maximized to determine the next optimal sampling point, and the corresponding prior probability distribution is continuously updated, and the continuous iteration is continued until the hyperparameter combination corresponding to the minimum value of the objective function is determined, the hyperparameter combination is the corresponding optimal solution, including, EI (x, P) = EP [max (0, μP (xbest) – f (x ❘ input))]; EIpS (x) = EIP (x) / μs (x); σP2(x) = σF2 (x) + σ2, σF (x) < tσ σ; where, El (x,P) is the acquisition function, x is the same as x in the objective function(f(x|input)), P is the objective function(p(f(x|input))), μP(xbest) is the minimum posterior mean, xbest is the currently known optimal solution, ElpS(x) represents the expected improvement function per second and μs(x) is the posterior mean of the time Gaussian process model, σP2(x) represents the audition function, σF(x) is the standard deviation of f(x|input), σ is the posteriori standard deviation of the additional noise, and tσ is the value of the mining proportion in the acquisition function. The Examiner reasons that these are merely the discrete steps involved in any Bayesian Optimization process, such as the one Manikandan is cited for by the Examiner for example in relation to claim 1. Conclusion The prior art made of record and not relied upon is considered pertinent to Applicants’ disclosure: Non-Patent Literature “Application of LSTM models in predicting particulate matter (PM2.5) levels for urban area” (Balaraman) Non-Patent Literature “Air-quality prediction based on the ARIMA-CNN-LSTM combination model optimized by dung beetle optimizer” (Duan) Non-Patent Literature “A hybrid CLSTM-GPR model for forecasting particulate matter (PM2.5)” (He) Non-Patent Literature “1D & 3D Convolutions explained with… MS Excel!” (Lane) Non-Patent Literature “Air Pollution Prediction by Deep Learning Model” (Jeya) Non-Patent Literature “Attention-LSTM architecture combined with Bayesian hyperparameter optimization for indoor temperature prediction” (Jiang) Non-Patent Literature “A Hybrid CNN-LSTM Model for Forecasting Particulate Matter (PM2.5)” (Li) Non-Patent Literature “Predicting of air pollutant concentrations based on spatio-temporal attention convolutional LSTM networks” (Jiang) Non-Patent Literature “Deep Learning for Air Quality Forecasts: a Review” (Liao) Non-Patent Literature “Multi-directional temporal convolutional artificial neural network for PM2.5 forecasting with missing values: A deep learning approach” (Samal) Non-Patent Literature “Flatten Layer — Implementation, Advantage and Disadvantages” (Srivatsavaya) Non-Patent Literature “Application of CNN-LSTM Algorithm for PM2.5 Concentration Forecasting in the Beijing-Tianjin-Hebei Metropolitan Area” (Su) Non-Patent Literature “PM2.5 Prediction with a Novel Multi-Step-Ahead Forecasting Model Based on Dynamic Wind Field Distance” (Yang) Non-Patent Literature “Deep-learning architecture for PM2.5 concentration prediction: A review” (Zhou) Non-Patent Literature “24-Hour prediction of PM₂.₅ concentrations by combining empirical mode decomposition and bidirectional long short-term memory neural network” (Teng) Non-Patent Literature “Deep learning PM2.5 concentrations with bidirectional LSTM RNN” (Tong) Non-Patent Literature “Updated Prediction of Air Quality Based on Kalman-Attention-LSTM Network” (Zhou) Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHOURJO DASGUPTA whose telephone number is (571)272-7207. The examiner can normally be reached M-F 8am-5pm CST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tamara Kyle can be reached at 571 272 4241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SHOURJO DASGUPTA/Primary Examiner, Art Unit 2144
Read full office action

Prosecution Timeline

Jul 18, 2024
Application Filed
Sep 23, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731146
System, Method, and Computer Program Product for Determining a Reason for a Deep Learning Model Output
4y 6m to grant Granted Sep 08, 2026
Patent 12731061
QUANTUM CIRCUIT FOR ESTIMATING MATRIX SPECTRAL SUMS
4y 1m to grant Granted Sep 08, 2026
Patent 12730554
DYNAMIC RESIZABLE MEDIA ITEM PLAYER
2y 2m to grant Granted Sep 08, 2026
Patent 12699932
System, Method, and Computer Program Product for Generating Error Rate Predictions Based on Machine Learning Using Incremental Backpropagation
4y 3m to grant Granted Aug 04, 2026
Patent 12694285
Method and Apparatus for Training a Quantized Classifier
4y 11m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
65%
Grant Probability
99%
With Interview (+39.3%)
3y 5m (~1y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 465 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month