Prosecution Insights
Last updated: August 16, 2026
Application No. 18/564,687

METHODS AND APPARATUSES FOR TRAINING A MODEL BASED REINFORCEMENT LEARNING MODEL

Non-Final OA §103§112
Filed
Nov 28, 2023
Priority
May 28, 2021 — nonprovisional of PCTEP2021064416
Examiner
CHEN, KUANG FU
Art Unit
Tech Center
Assignee
Telefonaktiebolaget LM Ericsson
OA Round
1 (Non-Final)
80%
Grant Probability
Favorable
1-2
OA Rounds
2m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
216 granted / 270 resolved
+20.0% vs TC avg
Strong +68% interview lift
Without
With
+68.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
26 currently pending
Career history
295
Total Applications
across all art units

Statute-Specific Performance

§101
16.9%
-23.1% vs TC avg
§103
50.3%
+10.3% vs TC avg
§102
11.1%
-28.9% vs TC avg
§112
15.2%
-24.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 270 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is responsive to the claims filed 11/28/2023. Claims 1-18 are presented for examination. Drawings The drawings are objected to under 37 CFR 1.83 and MPEP 608.02(g). Figure 1 depicts only the conventional process of manually tuning a cavity filter by a human expert (the expert 100 observing the S-parameter measurements 101 on the Vector Network Analyser 102 and turning the screws 103), which the specification describes as background and which is admitted prior art; Figure 1 is presently designated only "(background)." Figure 1 should be designated by a legend such as --Prior Art-- because only that which is old is illustrated. See MPEP 608.02(g). Corrected drawings in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. The replacement sheet(s) should be labeled "Replacement Sheet" in the page header (as per 37 CFR 1.84(c)) so as not to obstruct any portion of the drawing figures. If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they do not include the following reference sign mentioned in the description: reference sign 1200 (the apparatus). The apparatus 1200 is introduced in the specification in the passage describing Figure 12 (the paragraph beginning "Figure 12 illustrates an apparatus 1200 comprising processing circuitry (or logic) 1201"), but the reference number 1200 does not appear in Figure 12, in which the outer enclosure representing the apparatus is unlabeled. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either "Replacement Sheet" or "New Sheet" pursuant to 37 CFR 1.121(d). Alternatively, the applicant may amend the specification to delete the reference sign 1200 rather than adding it to Figure 12. The objection to the drawings will not be held in abeyance. Specification The disclosure is objected to because of the following informalities: (a) On page 3, in the Summary, both the method embodiment and the apparatus embodiment recite "determining means and standard deviations for based on the latent states st." The word "for" is extraneous; the phrase should read "determining means and standard deviations based on the latent states st," consistent with claim 1 and with the description of the observation model in the detailed description. (b) On page 13, the description of how step 203 of Figure 2 is performed, the sentence "Step 203 of Figure 2 may comprise minimizing a second loss function to update network parameters of the critic model 601 and the actor model 602" designates the actor model by reference number 602, whereas the actor model is designated by reference number 600 elsewhere in the same discussion ("the actor model 600 and the critic model 601 may be updated") and is designated 600 in Figure 6. The reference number used for the actor model should be made consistent as 600; reference number 602 does not appear in the drawings. (c) On page 8, in the description of the critic model neural network, the phrase "The critic model neural network may comprise a sequence of, for fully connected layers" is grammatically incomplete and appears to contain a typographical error. (d) On page 10, in the description of the training procedure of Figure 2, the sentence "The method may then continue until the network parameters of the world model and the actor-critic model converge, or until the performs at a desired level" appears to omit a word (for example, "until the model performs at a desired level"). (e) On page 14, in the description of using the trained model in a cavity filter environment, the phrase "Using the a trained MBRL model in the environment" contains the successive articles "the a" and should read "Using the trained MBRL model" or "Using a trained MBRL model." (f) On page 15, in the description of the wireless-device environment, the sentence ending "to obtain a desired value of the performance parameter.," contains a stray comma following the terminal period. Appropriate correction is required. The disclosure is objected to because it contains an embedded hyperlink and/or other form of browser-executable code. For example, the Background includes the embedded uniform resource locators "http://urn.kb.se/resolve?urn=urn:nbn:se:kth:diva-254422" and "https://arxiv.org/abs/2010.02193." Applicant is required to delete the embedded hyperlink and/or other form of browser-executable code; references to websites should be limited to the top-level domain name without any prefix such as http:// or other browser-executable code. See MPEP 608.01 (VII). Appropriate correction is required. Claim Objections Claims 3 and 16 are objected to because of the following informalities: Claim 3 recites "wherein the step minimizing the first loss function is further used to update network parameters of the reward model"; the phrase "the step minimizing" should read "the step of minimizing." Claim 3 further recites "a component relating to the how well the reward rt represents a real reward"; the article "the" before "how well" is extraneous, and the phrase should read "a component relating to how well the reward rt represents a real reward." Claim 16 recites "and wherein the using the trained model in the environment comprises adjusting"; the article "the" before "using" should be deleted so that the phrase reads "and wherein using the trained model in the environment comprises adjusting," consistent with claim 14 and claim 15. Appropriate correction is required. Claim Rejections - 35 U.S.C. 112(b) The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 5, 6, 15, and 16 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 5 depends from claim 3 and recites, in the wherein clause, that the critic model determines state values based on the transitional latent states strans,t and the actor model determines actions at based on the transitional latent states strans,t. The limitation the transitional latent states strans,t lacks proper antecedent basis. A transitional latent state strans,t is first introduced only in claim 4 (estimating a transitional latent state strans,t, using a transition model). Claim 5, however, depends from claim 3 rather than from claim 4, and claim 5 does not itself introduce any transitional latent state; the transitional latent state strans,t recited in claim 5 therefore has no antecedent in claim 5 or in any claim in its dependency chain (claims 1, 3, and 5). Claim 5 recites the limitation the transitional latent states strans,t in the clauses reciting the critic model and the actor model. There is insufficient antecedent basis for this limitation in the claim. See MPEP 2173.05(e). Under the broadest reasonable interpretation, the transitional latent states strans,t is interpreted, consistent with the specification (pages 13-14), as the latent state at time t produced by a transition model from the previous transitional latent state strans,t-1 and the previous action at-1 without using the current observation, and from which the critic model determines state values and the actor model determines actions; the claim is examined on that reading for prior art. Claim 6 depends from claim 5. It incorporates and does not cure the lack of antecedent basis discussed above, and it independently recites transitional latent states, strans,t associated with high state values without antecedent basis for the same reason. Claim 6 is therefore rejected under 35 U.S.C. 112(b) for the same reason as claim 5. Claim 15 depends from claim 14, which depends from claim 1. Claim 15 recites that the observations, ot, each comprise S-parameters of the cavity filter and that using the trained model comprises tuning the characteristics of the cavity filter to produce desired S-parameters. The limitation the cavity filter lacks proper antecedent basis. A cavity filter is first introduced in claim 7 (the environment comprises a cavity filter being controlled by a control unit), which is not an ancestor of claim 15; neither claim 14 nor claim 1 introduces a cavity filter. Likewise, the characteristics of the cavity filter has no antecedent, tuning characteristics of the cavity filter first being introduced in claim 9, which is also not an ancestor of claim 15. Claim 15 recites the limitations the cavity filter and the characteristics of the cavity filter. There is insufficient antecedent basis for these limitations in the claim. See MPEP 2173.05(e). Under the broadest reasonable interpretation, the cavity filter is interpreted, consistent with the specification (pages 15-16), as the cavity filter of the disclosed environment whose scattering (S-) parameters serve as the observations, and the characteristics of the cavity filter as the tunable characteristics adjusted (for example, by turning screws to move the poles and zeros) to bring the S-parameters to target values; the claim is examined on that reading for prior art. Claim 16 depends from claim 14, which depends from claim 1. Claim 16 introduces a wireless device, a cell, and a radio transmission beam pattern, and recites adjusting one of: the transmission power of the wireless device; the modulation and coding scheme used by the wireless device; and a radio transmission beam pattern, to obtain a desired value of the performance parameter. The limitations the transmission power of the wireless device, the modulation and coding scheme used by the wireless device, and the performance parameter lack proper antecedent basis. A transmission power of the wireless device and a modulation and coding scheme used by the wireless device are first introduced in claim 13, and a performance parameter is first introduced in claim 11 (which depends from claim 10); none of claims 10, 11, and 13 is an ancestor of claim 16, and claim 16 does not itself introduce a transmission power, a modulation and coding scheme, or a performance parameter with an indefinite article. Claim 16 recites the limitations the transmission power of the wireless device, the modulation and coding scheme used by the wireless device, and the performance parameter. There is insufficient antecedent basis for these limitations in the claim. See MPEP 2173.05(e). This rejection is not based on the alternative one of ... and recitation itself, which is a definite alternative limitation, but solely on the absence of antecedent basis for the three definite-article limitations identified above. Under the broadest reasonable interpretation, the performance parameter is interpreted, consistent with the specification (page 16), as a performance parameter experienced by the wireless device (for example, a signal-to-interference-and-noise ratio, traffic in the cell, or a transmission budget), and the transmission power and the modulation and coding scheme as the correspondingly controllable transmission attributes of the wireless device; the claim is examined on that reading for prior art. Claim Rejections - 35 U.S.C. 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-6, 14 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Hafner et al. (hereinafter Hafner) "Dream to Control: Learning Behaviors by Latent Imagination" (2020) in view of Kingma et al. (hereinafter Kingma) "Auto-Encoding Variational Bayes" (2014). Hafner was disclosed in an IDS dated 11/28/2023. Regarding independent claim 1, Hafner teaches a method for training a model based reinforcement learning, MBRL, model for use in an environment (Hafner: page 1, Abstract, "We present Dreamer, a reinforcement learning agent that solves long-horizon tasks from images purely by latent imagination"; Hafner presents a reinforcement learning agent (a model based reinforcement learning, MBRL, model) that learns a world model of its environment, so the training of that agent is a method for training a model based reinforcement learning model for use in an environment), the method comprising: obtaining a sequence of observations, ot, representative of the environment at a time t (Hafner: page 3, Algorithm 1, "Draw B data sequences {(at, ot, rt)}"; Hafner draws sequences of past experience whose elements include the observations ot (a sequence of observations, ot) collected from the environment at successive times); estimating latent states st at time t using a representation model, wherein the representation model estimates the latent states st based on the previous latent states st-1, previous actions at-1 and the observations ot (Hafner: page 2, Latent dynamics, "The representation model encodes observations and actions to create continuous vector-valued model states st with Markovian transitions"; page 2, equation (1), defining the representation model as p(st | st-1, at-1, ot); Hafner's representation model (a representation model) encodes the observation ot and the action into the model state st (the latent states st), and equation (1) expressly conditions the model state st on the previous model state st-1, the previous action at-1 and the current observation ot); generating modelled observations, om,t, using an observation model, wherein the observation model generates the modelled observations based on the respective latent states st (Hafner: page 6, Reconstruction, "the observation model is only used to provide a learning signal"; page 6, Reconstruction, "the observation model as a transposed CNN"; Hafner's observation model (an observation model) is a transposed convolutional decoder that reconstructs the observation from the model state st, and that reconstruction is a modelled observation om,t generated from the respective model state st); and minimizing a first loss function to update network parameters of the representation model and the observation model, wherein the first loss function comprises a component comparing the modelled observations, om,t to the respective observations ot (Hafner: page 6, Reconstruction, "The components are optimized jointly to increase the variational lower bound"; page 6, Reconstruction, "the bound includes reconstruction terms for observations and rewards and a KL regularizer"; the reconstruction term for observations scores the reconstructed observation against the true observation ot and, being optimized jointly, updates the network parameters of both the representation model and the observation model, so it is a first loss function whose component compares the modelled observation om,t to the respective observation ot). Hafner does not expressly teach wherein the step of generating comprises determining means and standard deviations based on the latent states st. However, Kingma teaches wherein the step of generating comprises determining means and standard deviations based on the latent states st (Kingma: page 3, Section 2.1, "a probabilistic decoder, since given a code z it produces a distribution over the possible corresponding values of x"; page 5, Section 3, "whose distribution parameters are computed from z with a MLP…a multivariate Gaussian with a diagonal covariance structure …where the mean and s.d. of the approximate posterior are outputs"; page 11, Appendix C.2, "let encoder or decoder be a multivariate Gaussian with a diagonal covariance structure"; Kingma's probabilistic decoder maps the latent code z to a diagonal-covariance Gaussian over the observation whose mean and whose diagonal standard deviation are each produced as a separate output of the decoder network from the latent code, the distribution parameters being expressly computed from the latent code z, so both a mean and a standard deviation are determined from the latent state). Because Hafner and Kingma are analogous art within the same field of endeavor, specifically the machine learning of generative models that reconstruct observations from a learned latent representation, and each is reasonably pertinent to the problem of accurately modelling high dimensional observations from a compact latent state, accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Kingma's Gaussian decoder, which determines both a mean and a standard deviation from the latent code, with Hafner's observation model, with a reasonable expectation of success, such that Hafner's observation model determines means and standard deviations based on the latent states st, to teach wherein the step of generating comprises determining means and standard deviations based on the latent states st. This modification would have been motivated by the desire to perform efficient inference and learning in directed probabilistic models, in the presence of continuous latent variables with intractable posterior distributions, an large datasets (Kingma: page 1, Abstract). Regarding dependent claim 2, Hafner, in view of Kingma, teach the method of claim 1, wherein the step of generating further comprises sampling distributions generated from the means and standard deviations to generate respective modelled observations, om,t (Kingma: page 3, Section 2.1, "a probabilistic decoder, since given a code z it produces a distribution over the possible corresponding values of x"; page 2, Section 2.1, "a value x(i) is generated from some conditional distribution pθ∗(x|z)"; page 5, Section 3, "whose distribution parameters are computed from z with a MLP"; Kingma's generative step expressly generates, that is samples, the data value from the conditional distribution whose parameters are computed from the latent code, and sampling the diagonal Gaussian observation distribution that Kingma's decoder generates from the determined mean and standard deviation yields the reconstructed observation om,t, so the modelled observation is generated by sampling a distribution generated from the means and standard deviations). Regarding dependent claim 3, Hafner, in view of Kingma, teach the method of claim 1, determining a reward rt based on a reward model, wherein the reward model determines the reward rt based on the latent state st (Hafner: page 2, Latent dynamics, "The reward model predicts the rewards given the model states"; Hafner's reward model (a reward model) predicts the reward rt from the model state st), wherein the step minimizing the first loss function is further used to update network parameters of the reward model, and wherein the first loss function further comprises a component relating to how well the reward rt represents a real reward for the observation ot (Hafner: page 6, Reconstruction, "the bound includes reconstruction terms for observations and rewards and a KL regularizer"; the reward reconstruction term is optimized jointly with the observation term, so minimizing the first loss updates the reward-model parameters and scores how well the predicted reward matches the true reward). Regarding dependent claim 4, Hafner, in view of Kingma, teach the method of claim 1, estimating a transitional latent state strans,t, using a transition model, wherein the transition model estimates the transitional latent state strans,t based on the previous transitional latent state strans,t-1 and a previous action at-1 (Hafner: page 2, Latent dynamics, "The transition model predicts future model states without seeing the corresponding observations that will later cause them"; Hafner's transition model (a transition model) predicts the next model state (the transitional latent state strans,t) from the previous model state and previous action at-1 without using the current observation), wherein the step of minimizing the first loss function is further used to update network parameters of the transition model, and wherein the first loss function further comprises a component relating to how similar the transitional latent state strans,t is to the latent state st (Hafner: page 6, Reconstruction, "the bound includes reconstruction terms for observations and rewards and a KL regularizer", page 15, Appendix B Derivations; Hafner's Kullback-Leibler regularizer penalizes the divergence between the representation-model state and the transition-model prediction and, optimized jointly, updates the transition-model parameters, so the first loss includes a component relating to how similar the transitional latent state is to the latent state). Regarding dependent claim 5, Hafner, in view of Kingma, teach the method of claim 4, after minimizing the first loss function, minimizing a second loss function to update network parameters of a critic model and an actor model, wherein the critic model determines state values based on the transitional latent states strans,t and the actor model determines actions at based on the transitional latent states strans,t (Hafner: page 4, Action and value models, "We learn an action model and a value model in the latent space of the world model…The action model implements the policy and aims to predict actions that solve the imagination environment. The value model estimates the expected imagined rewards that the action model achieves from each state…trained cooperatively as typical in policy iteration"; page 3, Algorithm 1 ("Dynamics learning" before "Behavior learning" in each update step); after the world model is fit, Hafner trains, in a separate behavior-learning objective, an action model (an actor model) and a value model (a critic model) that respectively determine actions and state values from the imagined latent states produced by the transition model, that is, from the transitional latent states; the action and value models are trained cooperatively in a single behavior-learning objective, a second loss function that updates the network parameters of both the critic model and the actor model; and Algorithm 1 places the dynamics-learning update that minimizes the first loss function before the behavior-learning update of the action and value models within each update step, so the second loss function is minimized after minimizing the first loss function). Regarding dependent claim 6, Hafner, in view of Kingma, teach the method of claim 5, wherein the second loss function comprises a component relating to ensuring the state values are accurate, and a component relating to ensuring the actor model leads to transitional latent states, strans,t associated with high state values (Hafner: page 4, Action and value models, "the action model aims to maximize an estimate of the value, while the value model aims to match an estimate of the value"; Hafner's value-model term drives the critic to match, and hence be accurate about, the value estimate, while the action-model term drives the actor toward states of high value, so the second loss has a component for accurate state values and a component for an actor that leads to high-value transitional latent states). Regarding dependent claim 14, Hafner, in view of Kingma, teach the method of claim 1, using the trained model in the environment (Hafner: page 2, Section 2, "Executing the learned action model in the world to collect new experience for growing the dataset"; page 3, Figure 3(c), "Act in the environment"; Hafner executes the trained agent in the environment to collect new experience, as depicted in Figure 3(c), so the trained model is used in the environment). Regarding dependent claim 17, Hafner, in view of Kingma, teach an apparatus for training a model based reinforcement learning, MBRL, model for use in an environment, the apparatus comprising processing circuitry configured to cause the apparatus to perform the method as claimed in claim 1 (Hafner: page 3, Algorithm 1, "Initialize neural network parameters"; page 8, Implementation, "We use a single Nvidia V100 GPU and 10 CPU cores for each training run"; the Dreamer agent is realized as neural networks whose parameters are initialized and updated by a computer, and Hafner expressly trains on a GPU and CPU cores, that is, on processing circuitry that causes the apparatus to carry out the training method of claim 1). Claims 7-9, 15 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Hafner, in view of Kingma, and further in view of Lindstahl "Reinforcement Learning with Imitation for Cavity Filter Tuning" (2019). Lindstahl was disclosed in an IDS dated 11/28/2023. Regarding dependent claim 7, Hafner in view of Kingma, teach all the elements of claim 1. Hafner and Kingma do not expressly teach wherein the environment comprises a cavity filter being controlled by a control unit. However, Lindstahl teaches wherein the environment comprises a cavity filter being controlled by a control unit (Lindstahl: page 20, Section 3.1.2, "Treating the S-parameters as the state, and the relative turning of each screw as a collective action, cavity filter tuning may be regarded as an MDP"; page 1, Chapter 1, "it is necessary to tune the filters, which is done by turning certain tuning screws on the surface of the component"; page 21, Section 3.2, "the S-parameters as state and tuning screw deviations as actions"; Lindstahl expressly casts cavity filter tuning as a Markov decision process whose environment is the cavity filter (a cavity filter), and an agent (a control unit) controls the filter by turning its tuning screws, so the environment comprises a cavity filter being controlled by a control unit). Because Hafner, in view of Kingma, and Lindstahl are analogous art, Hafner and Lindstahl being within the same field of endeavor of reinforcement learning for control and Lindstahl being reasonably pertinent to the problem of tuning a physical cavity filter with few interactions, accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to apply the model based reinforcement learning method of Hafner in view of Kingma to Lindstahl's cavity filter tuning environment, with a reasonable expectation of success, such that the environment comprises a cavity filter being controlled by a control unit, to teach wherein the environment comprises a cavity filter being controlled by a control unit. This modification would have been motivated by the desire to overcome the high number of training interactions that model-free reinforcement learning requires for cavity filter tuning (Lindstahl: page 1, Chapter 1). Regarding dependent claim 8, Hafner, in view of Kingma, and Lindstahl, teach the method of claim 7, wherein the observations, ot, each comprise S-parameters of the cavity filter (Lindstahl: page 21, Section 3.2, "the S-parameters as state"; Lindstahl uses the cavity filter's S-parameters (S-parameters of the cavity filter) as the observed state, so each observation ot comprises S-parameters of the cavity filter). Regarding dependent claim 9, Hafner, in view of Kingma, and Lindstahl, teach the method of claim 7, wherein the previous actions at-1 relate to tuning characteristics of the cavity filter (Lindstahl: page 21, Chapter 3, "tuning screw deviations as actions"; Lindstahl's actions are deviations of the cavity filter's tuning screws, which set the filter's tuning characteristics, so the previous actions at-1 relate to tuning characteristics of the cavity filter). Regarding dependent claim 15, Hafner, in view of Kingma, teach all the elements of claim 14. Hafner and Kingma do not expressly teach wherein the observations, ot, each comprise S-parameters of a cavity filter and wherein using the trained model in the environment comprises tuning characteristics of the cavity filter to produce desired S-parameters. However, Lindstahl teaches wherein the observations, ot, each comprise S-parameters of a cavity filter and wherein using the trained model in the environment comprises tuning characteristics of the cavity filter to produce desired S-parameters (Lindstahl: page 21, Section 3.2, "the S-parameters as state and tuning screw deviations as actions", page 22, "e11, e21 are the distance vectors between the S-parameters and specifications"; Lindstahl observes the filter's S-parameters (S-parameters of a cavity filter) and applies the trained agent to turn the tuning screws so the S-parameters approach the specification, so using the trained model tunes characteristics of the cavity filter to produce desired S-parameters). Because Hafner, in view of Kingma, and Lindstahl are analogous art within the same field of endeavor of reinforcement learning for control, and Lindstahl is reasonably pertinent to the problem of driving a cavity filter to a target response, accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to deploy the trained model based reinforcement learning agent of Hafner and Kingma in Lindstahl's cavity filter environment, with a reasonable expectation of success, such that the observations comprise S-parameters of a cavity filter and the trained model tunes the filter to produce desired S-parameters, to teach wherein the observations, ot, each comprise S-parameters of a cavity filter and wherein using the trained model in the environment comprises tuning characteristics of the cavity filter to produce desired S-parameters. This modification would have been motivated by the desire to overcome the high number of training interactions that model-free reinforcement learning requires for cavity filter tuning (Lindstahl: page 1, Chapter 1). Regarding dependent claim 18, Hafner, in view of Kingma, teach all the elements of claim 17. Hafner and Kingma do not expressly teach wherein the apparatus comprises a control unit for a cavity filter. However, Lindstahl teaches wherein the apparatus comprises a control unit for a cavity filter (Lindstahl: page 1, Chapter 1, "it is necessary to tune the filters, which is done by turning certain tuning screws on the surface of the component"; page 21, Section 3.2, "tuning screw deviations as actions"; Lindstahl's agent actuates the tuning screws of a cavity filter, so the apparatus that carries out the agent comprises a control unit for a cavity filter). Because Hafner, in view of Kingma, and Lindstahl are analogous art within the same field of endeavor of reinforcement learning for control, and Lindstahl is reasonably pertinent to the problem of automating cavity filter tuning, accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to embody the training apparatus of Hafner in view of Kingma as the control unit that tunes Lindstahl's cavity filter, with a reasonable expectation of success, such that the apparatus comprises a control unit for a cavity filter, to teach wherein the apparatus comprises a control unit for a cavity filter. This modification would have been motivated by the desire to replace slow manual tuning with an automated agent acting directly on the filter screws (Lindstahl: page 1, Chapter 1). Claims 10-13 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Hafner, in view of Kingma, and further in view of Mota et al. (hereinafter Mota) "Adaptive Modulation and Coding Based on Reinforcement Learning for 5G Networks" (2019). Regarding dependent claim 10, Hafner in view of Kingma, teach all the elements of claim 1. Hafner and Kingma do not expressly teach wherein the environment comprises a wireless device performing transmissions in a cell. However, Mota teaches wherein the environment comprises a wireless device performing transmissions in a cell (Mota: page 1, Abstract, "the BS chooses the MCS based on the channel quality indicator (CQI) reported by the user equipment (UE). A transmission is made with the chosen MCS"; page 2, Section III, "Consider a single cell system whose BS is equipped with M antennas serving one UE"; page 4, Algorithm 1, "The UE observes the state s : CQI and feeds it back to the BS"; Mota's environment is a single cell system in which a user equipment (a wireless device) transmits in a cell served by a base station, so the environment comprises a wireless device performing transmissions in a cell). Because Hafner, in view of Kingma, and Mota are analogous art, Hafner and Mota being within the same field of endeavor of reinforcement learning for control and Mota being reasonably pertinent to the problem of adapting a wireless transmission to a changing channel, accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to apply the model based reinforcement learning method of Hafner in view of Kingma to Mota's cellular environment, with a reasonable expectation of success, such that the environment comprises a wireless device performing transmissions in a cell, to teach wherein the environment comprises a wireless device performing transmissions in a cell. This modification would have been motivated by the desire to learn a transmission-adaptation policy that matches the time-varying wireless channel from few interactions by leveraging a learned world model to choose a suitable modulation and coding scheme (MCS) that maximizes the spectral efficiency (Mota: page 1). Regarding dependent claim 11, Hafner, in view of Kingma, and Mota, teach the method of claim 10, wherein the observations, ot, each comprise a performance parameter experienced by a wireless device (Mota: page 4, Algorithm 1, "The UE observes the state s : CQI"; the channel quality indicator (a performance parameter) that Mota's user equipment observes is experienced by that wireless device, so each observation ot comprises a performance parameter experienced by a wireless device). Regarding dependent claim 12, Hafner, in view of Kingma, and Mota, teach the method of claim 11, wherein the performance parameter comprises one or more of: a signal to interference and noise ratio; traffic in the cell and a transmission budget (Mota: page 1, Section I, "the selection of the MCS is based on the received signal-to-interference-plus-noise ratio (SINR)"; Mota's performance parameter is the received signal-to-interference-plus-noise ratio (a signal to interference and noise ratio), which satisfies the recited alternative group). Regarding dependent claim 13, Hafner, in view of Kingma and Mota, teach the method of claim 10, wherein the previous actions at-1 relate to controlling one or more of: a transmission power of the wireless device; a modulation and coding scheme used by the wireless device; and a radio transmission beam pattern (Mota: page 4, Algorithm 1, "The BS takes an action a : MCS"; Mota's action sets the modulation and coding scheme (a modulation and coding scheme used by the wireless device) used for the transmission, which satisfies the recited alternative group). Regarding dependent claim 16, Hafner in view of Kingma, teach all the elements of claim 14. Hafner and Kingma do not expressly teach wherein the environment comprises a wireless device performing transmissions in a cell and wherein using the trained model in the environment comprises adjusting one of: a transmission power of the wireless device; a modulation and coding scheme used by the wireless device; and a radio transmission beam pattern, to obtain a desired value of a performance parameter. However, Mota teaches wherein the environment comprises a wireless device performing transmissions in a cell and wherein using the trained model in the environment comprises adjusting one of: a transmission power of the wireless device; a modulation and coding scheme used by the wireless device; and a radio transmission beam pattern, to obtain a desired value of a performance parameter (Mota: page 2, Section III, "Consider a single cell system whose BS is equipped with M antennas serving one UE"; page 4, Algorithm 1, "The BS takes an action a : MCS using the policy driven by Q"; page 4, Equation (10), "the agent will try to maximize the spectral efficiency"; Mota applies the trained agent to a user equipment transmitting in a single cell system, that is in a cell, and selects, that is adjusts, the modulation and coding scheme so as to maximize the spectral efficiency, so the trained model adjusts a modulation and coding scheme used by the wireless device to obtain a desired value of a performance parameter; only one member of the recited alternative group need be taught). Because Hafner, in view of Kingma, and Mota are analogous art within the same field of endeavor of reinforcement learning for control, and Mota is reasonably pertinent to the problem of adapting a wireless transmission toward a target performance, accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to deploy the trained model based reinforcement learning agent of Hafner in view of Kingma in Mota's cellular environment, with a reasonable expectation of success, such that the trained model adjusts the modulation and coding scheme to obtain a desired value of a performance parameter, to teach wherein the environment comprises a wireless device performing transmissions in a cell and wherein using the trained model in the environment comprises adjusting one of: a transmission power of the wireless device; a modulation and coding scheme used by the wireless device; and a radio transmission beam pattern, to obtain a desired value of a performance parameter. This modification would have been motivated by the desire to learn a transmission-adaptation policy that matches the time-varying wireless channel from few interactions by leveraging a learned world model to choose a suitable modulation and coding scheme (MCS) that maximizes the spectral efficiency (Mota: page 1). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Karaletsos et al., US 2020/0372410 A1 (Nov. 26, 2020) (A machine learning model for reinforcement learning uses parameterized families of Markov decision processes (MDP) with latent variables. The system uses latent variables to improve ability of models to transfer knowledge and generalize to new tasks. Accordingly, trained machine learning based models are able to work in unseen environments or combinations of conditions/factors that the machine learning model was never trained on. For example, robots or self-driving vehicles based on the machine learning based models are robust to changing goals and are able to adapt to novel reward functions or tasks flexibly while being able to transfer knowledge about environments and agents to new tasks). Any inquiry concerning this communication or earlier communications from the examiner should be directed to KUANG FU CHEN whose telephone number is (571)272-1393. The examiner can normally be reached M-F 9:00-5:30pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached on (571) 272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KC CHEN/Primary Patent Examiner, Art Unit 2143
Read full office action

Prosecution Timeline

Nov 28, 2023
Application Filed
Jul 15, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12682047
Machine Learning Time Series Anomaly Detection
4y 1m to grant Granted Jul 14, 2026
Patent 12675771
System for Online Interaction with Content
5y 8m to grant Granted Jul 07, 2026
Patent 12664448
AVERAGE TREATMENT EFFECT FOR PAIRED DATA
4y 11m to grant Granted Jun 23, 2026
Patent 12657260
SIMULATING TRAINING DATA TO MITIGATE BIASES IN MACHINE LEARNING MODELS
4y 1m to grant Granted Jun 16, 2026
Patent 12657494
LEARNING SYSTEM, LEARNING METHOD, AND STORAGE MEDIUM
2y 12m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
80%
Grant Probability
99%
With Interview (+68.4%)
2y 11m (~2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 270 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month