DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This is a Final Action on the Merits. Claims 1-20 are currently pending and are addressed below.
Response to Amendments
The amendment filed on June 9th, 2026 has been considered and entered. Accordingly, claims 1, 10, and 165 have been amended.
Response to Arguments
The previous rejection of claims 1-20 under 35 USC 101 has been overcome due to the applicant’s amendments.
The applicant’s arguments with respect to claims 1-20 have been considered but are moot in view of the newly formulated grounds of rejections necessitated by the applicant’s amendments.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-3, 5-6, 9-17, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Guzman (Heteroscedastic Bayesian Optimisation for Stochastic Model Predictive Control) (“Guzman”) (Attached) in view of Jiang (Learning-Based Vehicle Dynamics Residual Correction Model for Autonomous Driving Simulation) (“Jiang”) (Attached) in view of Slutskyy (US 20200132488 A1) (“Slutskyy”).
With respect to claim 1, Guzman teaches an apparatus for determining improved hyperparameters for use in an autonomous vehicle motion planner, the apparatus comprising: one or more processors; and a memory storing data in non-transient form defining a program code executable by the one or more processors (See at least Guzman Page 6 “To assess the effects of real heteroscedastic noise in a physical system, we performed experiments on tuning an MPPI controller for a physical robot.”), wherein the one or more processors execute the program code to cause the apparatus to:
receive data comprising at least one data pair, each data pair comprising a set of hyperparameters and a utility score defining a corresponding utility of a motion planner outcome resulting from the set of hyperparameters (See at least Guzman Pages 4-5 “We optimise the controller with BO, a global optimisation method, by maximising the episodic or cumulative reward g dependent on controller hyper-parameters x to solve x∗ = argmaxg(x). Considering g is stochastic, we maximise the expected cumulative reward ˆgt = E[g(xt)]. The controller hyper-parameters are optimised following Algorithm 2. At each BO iteration, we fit the GP model Mwith observations collected up to the current iteration t … an optimal action a∗ i is returned by the MPC controller configured with x and sent to the system actuators. This returns a reward ri that is accumulated in gj(xt), where j is the current repetition. Finally, the optimal controller hyper-parameters x* correspond to those with maximum expected cumulative after nBO optimisation iterations”);
provide a model, based on the at least one data pair, wherein the model defines a relationship between the set of hyperparameters and the corresponding utility score (See at least Guzman Page 4 “The controller hyper-parameters are optimised following Algorithm 2. At each BO iteration, we fit the GP model Mwith observations collected up to the current iteration t.”);
generate at least one trial set of hyperparameters using a guidance objective configured to evaluate a quality of trial sets of hyperparameters in dependence on the model (See at least Guzman Page 4 “Next, we select controller hyper-parameters xt by maximising the acquisition function h with a global optimisation method.”);
determine a trial outcome of the motion planner in dependence on the trial set of hyperparameters (See at least Guzamn Page 4 “We then compute the expected cumulative reward ˆgt empirically by averaging the cumulative rewards obtained after nr episodes of ne time steps each. At each time-step i, an optimal action a∗ i is returned by the MPC controller configured with x and sent to the system actuators.”);
determine a new utility score of the trial set of hyperparameters; generate a new data pair comprising the trial set of hyperparameters and the new utility score; determined the improved hyperparameters from the new data pair (See at least Guzman Pages 4-5 “This returns a reward ri that is accumulated in gj(xt), where j is the current repetition. Finally, the optimal controller hyper-parameters x* correspond to those with maximum expected cumulative after nBO optimisation iterations”).
Guzman fails to explicitly disclose that a trial outcome of the motion planner in dependence on the trial set of hyperparameters and predetermined journey data and that the new utility score is determined in dependence on a comparison of the trial outcome with truth outcome data associated with the predetermined journey data; and operate the motion planner of the autonomous vehicle using the improved hyperparameters to select a trajectory from multiple potential trajectories to control motion of the autonomous vehicle, wherein the improved hyperparameters indicate a relative importance of environmental factors obtained by the autonomous vehicle, wherein the relative importance indicates a weight that influences how the motion planner processes each environmental factor to select the trajectory that minimizes a trajectory cost.
Jiang teaches that that a trial outcome of the motion planner in dependence on the trial set of hyperparameters and predetermined journey data (See at least Jiang Page 785 “Evaluation and Tuning: Evaluation dataset is different from training and validation. Four types of scenarios (left-turn, right-turn, U-turn and zig-zag) are designed and collected from open-loop driving data. Each lasts around 20 seconds. The model performance is evaluated by the similarity between residual corrected trajectories and the ground truth trajectories under the same control commands, i.e., m-ATE. The pipeline utilizes these evaluation metrics to construct a loss function for tuning. We select Bayesian-based optimization to tune eight hyperparameters, such as kernel size, loss weights. Based on past evaluations, Bayesian optimization constructs a posterior distribution of functions that best describes the objective function, hence, can efficiently find the optimal set of hyperparameters. All these tuned models are tracked and versioned so that appropriate ones can be deployed on the simulation platform to provide a vehicle dynamics model that truthfully reflects real-world dynamic characteristics.”) and that the new utility score is determined in dependence on a comparison of the trial outcome with truth outcome data associated with the predetermined journey data (See at least Jiang Page 786 “The overall model performance is defined as its trajectory accuracy improvements compared to its plugged-in dynamics base model in all evaluation scenarios. The overall accuracy improvement compared to LB is denoted as IMPLB, and IMPRB when compared to RB … where Js is the total number of scenarios and δj traj,. is the trajectory residual of model predicted trajectory compared to the ground truth trajectory in scenario j. We choose the mean trajectory error, i.e., m-ATE, to represent the trajectory residual, calculated as … where pm,i is a model predicted two dimensional trajectory point at t = i, while pgt,i is the ground truth trajectory point at the same time t = i. The distance between this pair of points is calculated by function dist(.).”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the apparatus of Guzman to include that a trial outcome of the motion planner in dependence on the trial set of hyperparameters and predetermined journey data and that the new utility score is determined in dependence on a comparison of the trial outcome with truth outcome data associated with the predetermined journey data, as taught by Jiang as disclosed above, in order to ensure an accurate updated hyperparameters (Jiang Page 782 “In this paper, we present a learning-based dynamics residual correction mechanism, which corrects the prediction residual of a dynamics base model to increase the overall model prediction accuracy.”).
Guzman in view of Jiang, however, fail to explicitly disclose to operate the motion planner of the autonomous vehicle using the improved hyperparameters to select a trajectory from multiple potential trajectories to control motion of the autonomous vehicle, wherein the improved hyperparameters indicate a relative importance of environmental factors obtained by the autonomous vehicle, wherein the relative importance indicates a weight that influences how the motion planner processes each environmental factor to select the trajectory that minimizes a trajectory cost.
Slutskyy teaches to operate the motion planner of the autonomous vehicle using the improved hyperparameters to select a trajectory from multiple potential trajectories to control motion of the autonomous vehicle, wherein the improved hyperparameters indicate a relative importance of environmental factors obtained by the autonomous vehicle, wherein the relative importance indicates a weight that influences how the motion planner processes each environmental factor to select the trajectory that minimizes a trajectory cost (See at least Slutskyy Paragraphs 128-129 “The planning module 1324 uses a directed graph representation of the drivable regions in the environment 1304 to generate a trajectory 414 including a plurality of travel segments. Each travel segment (e.g., edge 1010 a in FIG. 10) represents a portion of the trajectory 414. In one embodiment, a travel segment includes a section of a road, a bridge, a change of lanes, an elevation, etc. Each travel segment of the plurality of travel segments begins at a first spatiotemporal location (an intermediate point in the trajectory) of a plurality of spatiotemporal locations and terminates at a second spatiotemporal location (another intermediate point in the trajectory) of the plurality of spatiotemporal locations. The plurality of spatiotemporal locations includes the initial spatiotemporal location and the destination spatiotemporal location. Each travel segment in the trajectory is associated with a plurality of operational metrics. The operational metrics represent an N-tuple of costs associated with navigating the AV 1316 along the travel segment from the first spatiotemporal location to the second spatiotemporal location. In one embodiment, if eight different operational metrics are used, the N-tuple is represented as (m1, m2, m3, m4, m5, m6, m7, m8). In this example, m1 represents the length of a travel segment, m2 represents the number of traffic lights on the travel segment, m3 represents the number of predicted collisions with other vehicles on the travel segment, etc. A cost function of the plurality of operational metrics is used to determine the cost of a candidate trajectory. Each operational metric in the N-tuple is maximized or summed across the travel segments in the candidate trajectory. In one embodiment, the cost function is represented as (+, +, max, +, max, max, +, +), indicating that the first and second elements in the N-tuple are summed, the third element is maximized, and so on. The operational metrics m1, m2, m4, m7, and m8 are summed, while the operational metrics m3, m5 and m6 are maximized. Therefore, the length of the travel segments and the number of traffic lights are summed, while the number of predicted collisions is maximized. If the maximum number of predicted collisions equals one or more, the candidate trajectory is discarded.” | Paragraph 134 “Each travel segment (e.g., 1436) is associated with a plurality of operational metrics. For example, the operational metrics (0, 1) are associated with 1436 as shown in FIG. 14. The operational metrics are associated with navigating the AV 1316 from the first spatiotemporal location 1404 to the second spatiotemporal location 1416 connected by the travel segment 1436. Each operational metric refers to a parametric cost that the AV 1316 will incur by traveling along the travel segment 1436. Each operational metric (e.g., parametric cost) of the plurality of operational metrics is optimized across the plurality of travel segments to generate the trajectory.” | Paragraph 141 “In one embodiment, different cost functions of the plurality of operational metrics are used to determine the cost of a trajectory. A cost function may include a Boolean indicator indicating whether a candidate trajectory satisfies a strategic guideline or all the strategic guidelines of a priority group of operational metrics. In one embodiment, the optimizing of each operational metric of the plurality of operational metrics across the plurality of travel segments to generate the trajectory includes ranking each operational metric of the plurality of operational metrics that is associated with navigational safety higher than an operational metric of the plurality of operational metrics that is not associated with navigational safety. The operational metrics of the plurality of operational metrics that are associated with navigational safety are optimized before the operational metrics not associated with navigational safety. The cost evaluation is performed based on the ranked priority of the plurality of operational metrics. The cost evaluation iterates through higher-ranked operational metrics to lower-ranked operational metrics. For example, an operational metric (e.g., a predicted number of collisions of the AV 1316 with other vehicles) related to human safety is ranked higher, and the cost evaluation begins with this higher-ranked operational metric. When two candidate travel segments or two candidate trajectories have the same value for a higher-ranked operational metric, the cost evaluation proceeds to the next lower-ranked operational metric, e.g., reducing driving time. In this manner, the cost evaluation iterates down the lower-ranked operational metrics until an optimal trajectory is generated. The ranking and priority of rules for navigation is explained in additional detail above with reference to FIG. 9.” | Paragraph 154 “In one embodiment, the optimizing of each operational metric across the plurality of travel segments (e.g., 1536 and 1540) to generate the trajectory includes ranking each operational metric that is associated with navigational safety (e.g., lateral clearance of the AV 1316 from object 1508) higher than an operational metric that is not associated with navigational safety (e.g., maximum change in steering angle). The operational metrics that are associated with navigational safety are optimized ahead in priority of the operational metrics not associated with navigational safety. For example, the lateral clearance operational metric is related to safety and is ranked higher than the change in steering angle operational metric. Therefore, the AV 1316 will optimize (reduce) the lateral clearance operational metric and select travel segment 1536 instead of the travel segment 1540.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the apparatus of Guzman in view of Jiang to operate the motion planner of the autonomous vehicle using the improved hyperparameters to select a trajectory from multiple potential trajectories to control motion of the autonomous vehicle, wherein the improved hyperparameters indicate a relative importance of environmental factors obtained by the autonomous vehicle, wherein the relative importance indicates a weight that influences how the motion planner processes each environmental factor to select the trajectory that minimizes a trajectory cost, as taught by Slutskyy as disclosed above, in order to ensure optimal trajectory selection (Slutskyy Paragraph 2 “This description relates generally to operation of vehicles and specifically to generation of optimal trajectories for navigation of vehicles.”).
With respect to claim 2, and similarly claim 16, Guzman in view of Jiang in view of Slutskyy teach that the model is a probabilistic surrogate function, and wherein the guidance objective is configured to generate the at least one trial set of hyperparameters by: fitting the probabilistic surrogate function to one or more of the at least one data pair; and searching a domain space of hyperparameter inputs in dependence on sampling the probabilistic surrogate function (See at least Guzamn Page 3 “C. Gaussian Processes A Gaussian process [20] represents a probability distribution over a space of functions. A GP prior over a function g : X →R is completely specified by a mean m : X → R and a positive-definite covariance function, k : X × X → R. Under the GP prior, the values of g at a finite collection of points {xi}n i=1 ⊂ X follow a multivariate normal distribution g(X) ∼ N(m,K), where g(X) = [g(x1),...,g(xn)]T, m:=m(X), and K is the n-by-n covariance matrix given by [K]i,j = k(xi,xj). 1) Inference: Now suppose we observe y ∈ Rn, where each yi = g(xi) + νi represents a function evaluation corrupted by jointly Gaussian noise ν ∼ N(0,Σν). The joint distribution of the observations and the function value at a point x ∈ X is then given by … where k(x) := [k(x,x1),...,k(x,xn)]T, Conditioning g(x) on the observations yields a Gaussian predictive distribution g(x)|y ∼ N(µ(x),σ2(x)), where … allowing us to infer function values at unobserved locations. Noise model: In general, observation noise ν is assumed to be homoscedastic, which means its distribution is not dependent on the inputs x. However, many applications present noise with a heteroscedastic behaviour, i.e. the noise distribution varies across the domain X. Under the Gaussian assumption, observation noise is simply another (zero-mean) Gaussian process with covariance function kν : X × X → R, so that [Σν]i,j := kν(xi,xj). In the homoscedastic case, the noise covariance function is simply kν(x,x) := σ2 ν, where σν ∈ R is constant, and kν(x,x) = 0 for x= x, yielding the classic Σν = σ2 νI. More generally, however, kν can be an arbitrary positive-definite covariance function. D. Bayesian Optimisation Consider the problem of searching for the global optimum of a function g : X → R over a given compact search space S ⊂X such as determining x∗ ∈ argmaxx∈S g(x). Assume that g is possibly non-convex and only partially observable via noisy estimates yt = g(xt) + νt with νt ∼ N(0,σ2 νt ). In addition, we can only observe the function up to N times. Bayesian optimisation [6] assumes that g is a random variable itself and models it as a stochastic process, which is usually a GP, indexed by X. To select points at which to observe g, BO uses an acquisition function h(x) as a guide that incorporates prior information provided by the GP model and the observations. Each query point xt ∈ S is then selected by maximising h. After collecting an observation yt, BO updates the GP model with the pair (xt,yt) and starts the next iteration with an improved belief about f. The BO loop repeats until we reach the given budget of N evaluations of the objective function. See Algorithm 1 for a summary. The acquisition function h determines which values to sample next. A common and simple acquisition function is the upper confidence bound (UCB) [26] … where κ ∈ R+ is a balance factor. UCB allows balancing exploration and exploitation by valuing points where there is high uncertainty (exploration) or where the GP predictive mean is high (exploitation). Keeping the balance factor κ biased towards exploration avoids local minima”).
With respect to claim 3, and similarly claim 17, Guzman in view of Jiang in view of Slutskyy teach that the probabilistic surrogate function is a gaussian process model (See at least Guzamn Page 3 “C. Gaussian Processes A Gaussian process [20] represents a probability distribution over a space of functions. A GP prior over a function g : X →R is completely specified by a mean m : X → R and a positive-definite covariance function, k : X × X → R. Under the GP prior, the values of g at a finite collection of points {xi}n i=1 ⊂ X follow a multivariate normal distribution g(X) ∼ N(m,K), where g(X) = [g(x1),...,g(xn)]T, m:=m(X), and K is the n-by-n covariance matrix given by [K]i,j = k(xi,xj). 1) Inference: Now suppose we observe y ∈ Rn, where each yi = g(xi) + νi represents a function evaluation corrupted by jointly Gaussian noise ν ∼ N(0,Σν). The joint distribution of the observations and the function value at a point x ∈ X is then given by … where k(x) := [k(x,x1),...,k(x,xn)]T, Conditioning g(x) on the observations yields a Gaussian predictive distribution g(x)|y ∼ N(µ(x),σ2(x)), where … allowing us to infer function values at unobserved locations. Noise model: In general, observation noise ν is assumed to be homoscedastic, which means its distribution is not dependent on the inputs x. However, many applications present noise with a heteroscedastic behaviour, i.e. the noise distribution varies across the domain X. Under the Gaussian assumption, observation noise is simply another (zero-mean) Gaussian process with covariance function kν : X × X → R, so that [Σν]i,j := kν(xi,xj). In the homoscedastic case, the noise covariance function is simply kν(x,x) := σ2 ν, where σν ∈ R is constant, and kν(x,x) = 0 for x= x, yielding the classic Σν = σ2 νI. More generally, however, kν can be an arbitrary positive-definite covariance function. D. Bayesian Optimisation Consider the problem of searching for the global optimum of a function g : X → R over a given compact search space S ⊂X such as determining x∗ ∈ argmaxx∈S g(x). Assume that g is possibly non-convex and only partially observable via noisy estimates yt = g(xt) + νt with νt ∼ N(0,σ2 νt ). In addition, we can only observe the function up to N times. Bayesian optimisation [6] assumes that g is a random variable itself and models it as a stochastic process, which is usually a GP, indexed by X. To select points at which to observe g, BO uses an acquisition function h(x) as a guide that incorporates prior information provided by the GP model and the observations. Each query point xt ∈ S is then selected by maximising h. After collecting an observation yt, BO updates the GP model with the pair (xt,yt) and starts the next iteration with an improved belief about f. The BO loop repeats until we reach the given budget of N evaluations of the objective function. See Algorithm 1 for a summary. The acquisition function h determines which values to sample next. A common and simple acquisition function is the upper confidence bound (UCB) [26] … where κ ∈ R+ is a balance factor. UCB allows balancing exploration and exploitation by valuing points where there is high uncertainty (exploration) or where the GP predictive mean is high (exploitation). Keeping the balance factor κ biased towards exploration avoids local minima”).
With respect to claim 5, and similarly claim 19, Guzman in view of Jiang in view of Slutskyy teach that the search of the domain space of hyperparameter inputs is guided by an acquisition function configured to determine the quality of trial sets of hyperparameters based at least in part on a predicted uncertainty of a value of the surrogate function resulting from a trial set of hyperparameters (See at least Guzman Page 3 “D. Bayesian Optimisation Consider the problem of searching for the global optimum of a function g : X → R over a given compact search space S ⊂X such as determining x∗ ∈ argmaxx∈S g(x). Assume that g is possibly non-convex and only partially observable via noisy estimates yt = g(xt) + νt with νt ∼ N(0,σ2 νt ). In addition, we can only observe the function up to N times. Bayesian optimisation [6] assumes that g is a random variable itself and models it as a stochastic process, which is usually a GP, indexed by X. To select points at which to observe g, BO uses an acquisition function h(x) as a guide that incorporates prior information provided by the GP model and the observations. Each query point xt ∈ S is then selected by maximising h. After collecting an observation yt, BO updates the GP model with the pair (xt,yt) and starts the next iteration with an improved belief about f. The BO loop repeats until we reach the given budget of N evaluations of the objective function. See Algorithm 1 for a summary. The acquisition function h determines which values to sample next. A common and simple acquisition function is the upper confidence bound (UCB) [26] … where κ ∈ R+ is a balance factor. UCB allows balancing exploration and exploitation by valuing points where there is high uncertainty (exploration) or where the GP predictive mean is high (exploitation). Keeping the balance factor κ biased towards exploration avoids local minima”).
With respect to claim 6, and similarly claim 20, Guzman in view of Jiang in view of Slutskyy teach that the acquisition function comprises one or more of the following functions: an expected utility function, a probability of improvement function, and an upper confidence bound function (See at least Guzman Page 3 “The acquisition function h determines which values to sample next. A common and simple acquisition function is the upper confidence bound (UCB) [26] … where κ ∈ R+ is a balance factor. UCB allows balancing exploration and exploitation by valuing points where there is high uncertainty (exploration) or where the GP predictive mean is high (exploitation). Keeping the balance factor κ biased towards exploration avoids local minima”).
With respect to claim 9, Guzman in view of Jiang in view of Slutskyy teach that the trial outcome is a vehicle trajectory comprising a plurality of vehicle motion decisions corresponding to at least one type of vehicle motion action (See at least Guzman Pages 6-7 “E. Experiments with a physical robot To assess the effects of real heteroscedastic noise in a physical system, we performed experiments on tuning an MPPI controller for a physical robot. The four-wheel-drive skid-steer robot (Fig. 10a) was tasked with following a circular path at a set speed. The cost function was formulated as c(st) = d2 t +(vr −vt)2, where dt represents the robot’s distance to the edge of the circle, vr = 0.2 m/s is a reference linear speed, and vt is the current speed. The robot was localised using a particle filter on a prebuilt map. Internally, MPPI employed a kinematic model of the robot [34] for trajectory rollouts which is challenging for MPC as the model does not simulate the dynamics of skid-steering platforms accurately. The controller was configured with M = 50 rollouts and a time horizon T = 400. Episodes lasted 20 seconds with the robot starting from a fixed initial position. The search space S for BO was set as the box defined by the intervals σ ∈ [0.3,0.5] and λ ∈ [0.01,0.21] … Performance results are in Fig. 10c. We compared BOhetero against BOhomo. Both algorithms are eventually able to find high reward regions. However, due to its uniform noise model, BOhomo is led to a more exploratory behaviour, instead of concentrating on promising regions, as evidenced by the query locations in Fig. 11. As a consequence, we observe a significant drop in performance during the optimisation, as shown in Fig. 10c. In contrast, BOhetero maintains a steady high performance, which means lower tracking error with respect to the circular path specified by the cost function.”) (See at least Jiang Page 785 “) Evaluation and Tuning: Evaluation dataset is different from training and validation. Four types of scenarios (left-turn, right-turn, U-turn and zig-zag) are designed and collected from open-loop driving data. Each lasts around 20 seconds. The model performance is evaluated by the similarity between residual corrected trajectories and the ground truth trajectories under the same control commands, i.e., m-ATE. The pipeline utilizes these evaluation metrics to construct a loss function for tuning. We select Bayesian-based optimization to tune eight hyperparameters, such as kernel size, loss weights. Based on past evaluations, Bayesian optimization constructs a posterior distribution of functions that best describes the objective function, hence, can efficiently find the optimal set of hyperparameters. All these tuned models are tracked and versioned so that appropriate ones can be deployed on the simulation platform to provide a vehicle dynamics model that truthfully reflects real-world dynamic characteristics.”).
With respect to claim 10, Guzman in view of Jiang in view of Slutskyy teach that the corresponding utility score represents accuracy of a trial outcome compared to the truth outcome data, the truth outcome data comprising human-labelled vehicle motion decisions (See at least Jiang Pages 782-783 “In this paper, we present a learning-based dynamics residual correction mechanism, which corrects the prediction residual of a dynamics base model to increase the overall model prediction accuracy. It achieves a mean average trajectory error (m-ATE) [14] of 3.266 m in different 20s-long scenarios, which is 80.93% trajectory accuracy improvement compared with a rule-based (RB) commercial model widely used in industry, and 52.12% trajectory accuracy improvement compared with our previous learning based (LB) model [15]. This is done by building a residual correction mechanism of vehicle dynamics such as heading angle and speed throughout model structure, training loss design and online inference … In model structure, we extend the output of the residual predictor to include speed difference, heading angle difference and heading angle change rate difference prediction. • In loss function, we include the speed mean squared error (MSE) loss as well. • In online inference, we apply the residual corrections to the additional vehicle dynamics mentioned above instead of applying corrections to vehicle position.”).
With respect to claim 11, Guzman in view of Jiang in view of Slutskyy teach that the corresponding utility score is determined in dependence on an objective function which rewards correct decisions in the trial outcome (See at least Jiang Pages 782-783 “In this paper, we present a learning-based dynamics residual correction mechanism, which corrects the prediction residual of a dynamics base model to increase the overall model prediction accuracy. It achieves a mean average trajectory error (m-ATE) [14] of 3.266 m in different 20s-long scenarios, which is 80.93% trajectory accuracy improvement compared with a rule-based (RB) commercial model widely used in industry, and 52.12% trajectory accuracy improvement compared with our previous learning based (LB) model [15]. This is done by building a residual correction mechanism of vehicle dynamics such as heading angle and speed throughout model structure, training loss design and online inference … In model structure, we extend the output of the residual predictor to include speed difference, heading angle difference and heading angle change rate difference prediction. • In loss function, we include the speed mean squared error (MSE) loss as well. • In online inference, we apply the residual corrections to the additional vehicle dynamics mentioned above instead of applying corrections to vehicle position.”) and/or penalises decisions in the trial outcome which are incorrect and which have previously been determined as correct based on an initial set of hyperparameter inputs in the received data
With respect to claim 12, Guzman in view of Jiang in view of Slutskyy teach that predetermined journey data comprises a plurality of journeys and the one or more processors further execute the program code to cause the apparatus to determine a plurality of trial outcomes, one per journey, using the trial set of hyperparameters; and determine the corresponding utility score based on the plurality of trial outcomes (See at least Jiang Pages 782-783 “In this paper, we present a learning-based dynamics residual correction mechanism, which corrects the prediction residual of a dynamics base model to increase the overall model prediction accuracy. It achieves a mean average trajectory error (m-ATE) [14] of 3.266 m in different 20s-long scenarios, which is 80.93% trajectory accuracy improvement compared with a rule-based (RB) commercial model widely used in industry, and 52.12% trajectory accuracy improvement compared with our previous learning based (LB) model [15]. This is done by building a residual correction mechanism of vehicle dynamics such as heading angle and speed throughout model structure, training loss design and online inference … In model structure, we extend the output of the residual predictor to include speed difference, heading angle difference and heading angle change rate difference prediction. • In loss function, we include the speed mean squared error (MSE) loss as well. • In online inference, we apply the residual corrections to the additional vehicle dynamics mentioned above instead of applying corrections to vehicle position.”).
With respect to claim 13, Guzman in view of Jiang in view of Slutskyy teach that the one or more processors further execute the program code to cause the apparatus to: determine the trial outcome of the motion planner using a static simulator (See at least Guzman Page 5 “A. Control Problem Simulations We conducted experiments on benchmark control problems from OpenAI Gym1 [27] and Mujoco [28]: Acrobot, Cartpole, Half-Cheetah, Pendulum, and Reacher. Each control problem has a particular state reward function r(s, a) shown in Table I. We made slight modifications in Reacher and Half-Cheetah. We reduced the effect of actions and gave more priority to the distance to the target in the case of the Reacher problem. For Half-Cheetah, we added more priority to the inclination, since Half-Cheetah would tend to turn upside down as its speed increases. The actuation is then set to finish when such inclination is greater than π/2 or lower than −π/2. These modifications make the rewards more informative for MPPI, enabling it to solve these two tasks. We can then focus the analysis on tuning the controller. The expected cumulative reward represents the expected time the pendulum stays in an upright position in the Acrobot, Cartpole, and Pendulum. It represents the distance traversed in Half-Cheetah, and the speed to reach the target in Reacher. High expected cumulative rewards are the result of motions that increased the reward accordingly, e.g. Half-Cheetah would be expected to reach farther distances. Now, to evaluate the expected cumulative rewards for each problem, we determined fixed values for time horizon T, number of trajectory rollouts M, and MPPI hyper-parameter intervals are shown in Table II. These were found by narrowing down large-enough intervals from near-zero values to 500, taking into account usual values for these hyper-parameters that tend to be close to 0. These are typical in several applications [17], [18]. The table also shows optimal values found within these narrowed intervals via grid search.”) (See at least Jiang Pages 785-786 “2) Data Balancing and Labeling: To generate and balance data, features are first aggregated based on the configurable history segment length (N) and over lapping size. The history segment length indicates the number of consecutive data points from the past in a training input, whereas overlapping size specifies the amount of the same data points to use in the next training input. In each training input, there are N historical datapoints, but only inputs from control commands and the initial vehicle states are from real-world driving data (aka. The ground truth data). The rest of the vehicle states are from the vehicle dynamics model that predicts over ground truth control commands sequences, as shown in Figure 5. A base vehicle dynamics model, the yellow block, updates vehicle dynamics under control command sequences at each step. The vehicle dynamics sequence and the control commands sequence form a N×7matrix as model inputs, blue block in Figure 5. Training output is the residual error between ground truth values and the base model predicted outputs … Labeled data are then classified into different categories based on 50 percentile values of those features in speed, throttle, steering, and brake. Each area is broken down into four to seven sub-categories to reflect the characteristics of different scenarios. For instance, different values of speed can map to very low, low, medium, and high speed scenarios. And steering angles are broken down into seven kinds of turning: slight left/right-turn, left/right-turn, sharp left/right turn or go straight. The combination of the three control commands plus speed suggests the category of a sample. In order to evenly distribute the samples, 8000 is the maximum number of samples in each category, extra ones are discarded. We obtained the training dataset and validation dataset by randomly splitting these samples in each category with a 9:1 ratio.”).
With respect to claim 14, Guzman in view of Jiang in view of Slutskyy teach that the one or more processors further execute the program code to cause the apparatus to: repeatedly perform the following processes: generating at least one trial set of hyperparameters; determining a new utility score; and wherein the received data comprises a previously generated new data pair (See at least Guzman Pages 4-5 “We optimise the controller with BO, a global optimisation method, by maximising the episodic or cumulative reward g dependent on controller hyper-parameters x to solve x∗ = argmaxg(x). Considering g is stochastic, we maximise the expected cumulative reward ˆgt = E[g(xt)]. The controller hyper-parameters are optimised following Algorithm 2. At each BO iteration, we fit the GP model M with observations collected up to the current iteration t … an optimal action a∗ i is returned by the MPC controller configured with x and sent to the system actuators. This returns a reward ri that is accumulated in gj(xt), where j is the current repetition. Finally, the optimal controller hyper-parameters x* correspond to those with maximum expected cumulative after nBO optimisation iterations”).
With respect to claim 15, Guzman teaches an method for determining improved hyperparameters for use in an autonomous vehicle motion planner, the method which applied to an electronic apparatus comprising:
receive data comprising at least one data pair, each data pair comprising a set of hyperparameters and a utility score defining a corresponding utility of a motion planner outcome resulting from the set of hyperparameters (See at least Guzman Pages 4-5 “We optimise the controller with BO, a global optimisation method, by maximising the episodic or cumulative reward g dependent on controller hyper-parameters x to solve x∗ = argmaxg(x). Considering g is stochastic, we maximise the expected cumulative reward ˆgt = E[g(xt)]. The controller hyper-parameters are optimised following Algorithm 2. At each BO iteration, we fit the GP model Mwith observations collected up to the current iteration t … an optimal action a∗ i is returned by the MPC controller configured with x and sent to the system actuators. This returns a reward ri that is accumulated in gj(xt), where j is the current repetition. Finally, the optimal controller hyper-parameters x* correspond to those with maximum expected cumulative after nBO optimisation iterations”);
provide a model, based on the at least one data pair, wherein the model defines a relationship between the set of hyperparameters and the corresponding utility score (See at least Guzman Page 4 “The controller hyper-parameters are optimised following Algorithm 2. At each BO iteration, we fit the GP model Mwith observations collected up to the current iteration t.”);
generate at least one trial set of hyperparameters using a guidance objective configured to evaluate a quality of trial sets of hyperparameters in dependence on the model (See at least Guzman Page 4 “Next, we select controller hyper-parameters xt by maximising the acquisition function h with a global optimisation method.”);
determine a trial outcome of the motion planner in dependence on the trial set of hyperparameters (See at least Guzamn Page 4 “We then compute the expected cumulative reward ˆgt empirically by averaging the cumulative rewards obtained after nr episodes of ne time steps each. At each time-step i, an optimal action a∗ i is returned by the MPC controller configured with x and sent to the system actuators.”);
determine a new utility score of the trial set of hyperparameters; generate a new data pair comprising the trial set of hyperparameters and the new utility score; determined the improved hyperparameters from the new data pair (See at least Guzman Pages 4-5 “This returns a reward ri that is accumulated in gj(xt), where j is the current repetition. Finally, the optimal controller hyper-parameters x* correspond to those with maximum expected cumulative after nBO optimisation iterations”).
Guzman fails to explicitly disclose that a trial outcome of the motion planner in dependence on the trial set of hyperparameters and predetermined journey data and that the new utility score is determined in dependence on a comparison of the trial outcome with truth outcome data associated with the predetermined journey data; and operate the motion planner of the autonomous vehicle using the improved hyperparameters to select a trajectory from multiple potential trajectories to control motion of the autonomous vehicle, wherein the improved hyperparameters indicate a relative importance of environmental factors obtained by the autonomous vehicle, wherein the relative importance indicates a weight that influences how the motion planner processes each environmental factor to select the trajectory that minimizes a trajectory cost.
Jiang teaches that that a trial outcome of the motion planner in dependence on the trial set of hyperparameters and predetermined journey data (See at least Jiang Page 785 “Evaluation and Tuning: Evaluation dataset is different from training and validation. Four types of scenarios (left-turn, right-turn, U-turn and zig-zag) are designed and collected from open-loop driving data. Each lasts around 20 seconds. The model performance is evaluated by the similarity between residual corrected trajectories and the ground truth trajectories under the same control commands, i.e., m-ATE. The pipeline utilizes these evaluation metrics to construct a loss function for tuning. We select Bayesian-based optimization to tune eight hyperparameters, such as kernel size, loss weights. Based on past evaluations, Bayesian optimization constructs a posterior distribution of functions that best describes the objective function, hence, can efficiently find the optimal set of hyperparameters. All these tuned models are tracked and versioned so that appropriate ones can be deployed on the simulation platform to provide a vehicle dynamics model that truthfully reflects real-world dynamic characteristics.”) and that the new utility score is determined in dependence on a comparison of the trial outcome with truth outcome data associated with the predetermined journey data (See at least Jiang Page 786 “The overall model performance is defined as its trajectory accuracy improvements compared to its plugged-in dynamics base model in all evaluation scenarios. The overall accuracy improvement compared to LB is denoted as IMPLB, and IMPRB when compared to RB … where Js is the total number of scenarios and δj traj,. is the trajectory residual of model predicted trajectory compared to the ground truth trajectory in scenario j. We choose the mean trajectory error, i.e., m-ATE, to represent the trajectory residual, calculated as … where pm,i is a model predicted two dimensional trajectory point at t = i, while pgt,i is the ground truth trajectory point at the same time t = i. The distance between this pair of points is calculated by function dist(.).”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the apparatus of Guzman to include that a trial outcome of the motion planner in dependence on the trial set of hyperparameters and predetermined journey data and that the new utility score is determined in dependence on a comparison of the trial outcome with truth outcome data associated with the predetermined journey data, as taught by Jiang as disclosed above, in order to ensure an accurate updated hyperparameters (Jiang Page 782 “In this paper, we present a learning-based dynamics residual correction mechanism, which corrects the prediction residual of a dynamics base model to increase the overall model prediction accuracy.”).
Guzman in view of Jiang, however, fail to explicitly disclose to operate the motion planner of the autonomous vehicle using the improved hyperparameters to select a trajectory from multiple potential trajectories to control motion of the autonomous vehicle, wherein the improved hyperparameters indicate a relative importance of environmental factors obtained by the autonomous vehicle, wherein the relative importance indicates a weight that influences how the motion planner processes each environmental factor to select the trajectory that minimizes a trajectory cost.
Slutskyy teaches to operate the motion planner of the autonomous vehicle using the improved hyperparameters to select a trajectory from multiple potential trajectories to control motion of the autonomous vehicle, wherein the improved hyperparameters indicate a relative importance of environmental factors obtained by the autonomous vehicle, wherein the relative importance indicates a weight that influences how the motion planner processes each environmental factor to select the trajectory that minimizes a trajectory cost (See at least Slutskyy Paragraphs 128-129 “The planning module 1324 uses a directed graph representation of the drivable regions in the environment 1304 to generate a trajectory 414 including a plurality of travel segments. Each travel segment (e.g., edge 1010 a in FIG. 10) represents a portion of the trajectory 414. In one embodiment, a travel segment includes a section of a road, a bridge, a change of lanes, an elevation, etc. Each travel segment of the plurality of travel segments begins at a first spatiotemporal location (an intermediate point in the trajectory) of a plurality of spatiotemporal locations and terminates at a second spatiotemporal location (another intermediate point in the trajectory) of the plurality of spatiotemporal locations. The plurality of spatiotemporal locations includes the initial spatiotemporal location and the destination spatiotemporal location. Each travel segment in the trajectory is associated with a plurality of operational metrics. The operational metrics represent an N-tuple of costs associated with navigating the AV 1316 along the travel segment from the first spatiotemporal location to the second spatiotemporal location. In one embodiment, if eight different operational metrics are used, the N-tuple is represented as (m1, m2, m3, m4, m5, m6, m7, m8). In this example, m1 represents the length of a travel segment, m2 represents the number of traffic lights on the travel segment, m3 represents the number of predicted collisions with other vehicles on the travel segment, etc. A cost function of the plurality of operational metrics is used to determine the cost of a candidate trajectory. Each operational metric in the N-tuple is maximized or summed across the travel segments in the candidate trajectory. In one embodiment, the cost function is represented as (+, +, max, +, max, max, +, +), indicating that the first and second elements in the N-tuple are summed, the third element is maximized, and so on. The operational metrics m1, m2, m4, m7, and m8 are summed, while the operational metrics m3, m5 and m6 are maximized. Therefore, the length of the travel segments and the number of traffic lights are summed, while the number of predicted collisions is maximized. If the maximum number of predicted collisions equals one or more, the candidate trajectory is discarded.” | Paragraph 134 “Each travel segment (e.g., 1436) is associated with a plurality of operational metrics. For example, the operational metrics (0, 1) are associated with 1436 as shown in FIG. 14. The operational metrics are associated with navigating the AV 1316 from the first spatiotemporal location 1404 to the second spatiotemporal location 1416 connected by the travel segment 1436. Each operational metric refers to a parametric cost that the AV 1316 will incur by traveling along the travel segment 1436. Each operational metric (e.g., parametric cost) of the plurality of operational metrics is optimized across the plurality of travel segments to generate the trajectory.” | Paragraph 141 “In one embodiment, different cost functions of the plurality of operational metrics are used to determine the cost of a trajectory. A cost function may include a Boolean indicator indicating whether a candidate trajectory satisfies a strategic guideline or all the strategic guidelines of a priority group of operational metrics. In one embodiment, the optimizing of each operational metric of the plurality of operational metrics across the plurality of travel segments to generate the trajectory includes ranking each operational metric of the plurality of operational metrics that is associated with navigational safety higher than an operational metric of the plurality of operational metrics that is not associated with navigational safety. The operational metrics of the plurality of operational metrics that are associated with navigational safety are optimized before the operational metrics not associated with navigational safety. The cost evaluation is performed based on the ranked priority of the plurality of operational metrics. The cost evaluation iterates through higher-ranked operational metrics to lower-ranked operational metrics. For example, an operational metric (e.g., a predicted number of collisions of the AV 1316 with other vehicles) related to human safety is ranked higher, and the cost evaluation begins with this higher-ranked operational metric. When two candidate travel segments or two candidate trajectories have the same value for a higher-ranked operational metric, the cost evaluation proceeds to the next lower-ranked operational metric, e.g., reducing driving time. In this manner, the cost evaluation iterates down the lower-ranked operational metrics until an optimal trajectory is generated. The ranking and priority of rules for navigation is explained in additional detail above with reference to FIG. 9.” | Paragraph 154 “In one embodiment, the optimizing of each operational metric across the plurality of travel segments (e.g., 1536 and 1540) to generate the trajectory includes ranking each operational metric that is associated with navigational safety (e.g., lateral clearance of the AV 1316 from object 1508) higher than an operational metric that is not associated with navigational safety (e.g., maximum change in steering angle). The operational metrics that are associated with navigational safety are optimized ahead in priority of the operational metrics not associated with navigational safety. For example, the lateral clearance operational metric is related to safety and is ranked higher than the change in steering angle operational metric. Therefore, the AV 1316 will optimize (reduce) the lateral clearance operational metric and select travel segment 1536 instead of the travel segment 1540.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the apparatus of Guzman in view of Jiang to operate the motion planner of the autonomous vehicle using the improved hyperparameters to select a trajectory from multiple potential trajectories to control motion of the autonomous vehicle, wherein the improved hyperparameters indicate a relative importance of environmental factors obtained by the autonomous vehicle, wherein the relative importance indicates a weight that influences how the motion planner processes each environmental factor to select the trajectory that minimizes a trajectory cost, as taught by Slutskyy as disclosed above, in order to ensure optimal trajectory selection (Slutskyy Paragraph 2 “This description relates generally to operation of vehicles and specifically to generation of optimal trajectories for navigation of vehicles.”).
Claims 4, 7-8, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Guzman (Heteroscedastic Bayesian Optimisation for Stochastic Model Predictive Control) (“Guzman”) (Attached) in view of Jiang (Learning-Based Vehicle Dynamics Residual Correction Model for Autonomous Driving Simulation) (“Jiang”) (Attached) in view of Slutskyy (US 20200132488 A1) (“Slutskyy”) further in view of Zhang (On the Importance of Hyperparameter Optimization for Model-based Reinforcement Learning) (“Zhang”) (Attached).
With respect to claim 4, and similarly claim 18, Guzman in view of Jiang in view of Slutskyy fail to explicitly disclose that the probabilistic surrogate function is a gaussian mixture model formed from the combination of a plurality of neural networks.
Zhang, however, teaches that the probabilistic surrogate function is a gaussian mixture model formed from the combination of a plurality of neural networks (See at least Zhang Pages 4-5 “Optimizee To demonstrate the importance of HPO for MBRL we use the current state-of-the-art Probabilistic Ensembles With Trajectory Sampling (PETS) (Chua et al., 2018) algorithm as optimizee. PETS uses an ensemble of neural networks to learn a model of the environment which provides aleatoric and epistemic uncertainty estimates. In PETS, the dynamics model is chosen to be an ensemble of neural networks whose outputs parameterize anisotropic Gaussians. Model Predictive Control (MPC) is then used to get the pol icy, by directly optimizing the expected sum of rewards over a fixed planning horizon. PETS, in particular, performs model predictive control with a cross-entropy method (CEM) optimizer for action selection. To eval uate an action sequence, PETS first samples a model from the ensemble, and rolls out the action sequence using the selected model, and computes the sum of re wards. Action sequences are then evaluated by perform ing this process multiple times iteratively and averaging the simulated returns over the ensemble members. Environments For the experiments, we consider four test environments: Pusher, Reacher, Hopper, Halfcheetah from MuJoCo (Todorov et al., 2012) and a simulation environment of Daisy, a robot hexapod to accomplish locomotion tasks. The reward signal for Daisy is similar to Hopper and HalfCheetah, where we use the forward speed as the reward signal. The number of trials is fixed to 80 for pusher and 300 for the rest, with rollout lengths of 150 and 1000 steps, respectively. The actions in the initial trial are sampled randomly to collect data to train the model before it is used by PETS in future trials. We use Hopper and HalfCheetah for illustrative plots in the main paper; quantitative results for the other three tasks can be found in Appendix A.5. Configuration Space We split the hyperparame ters of PETS into two groups: (i) Model Training and (ii) CEM Optimizer; to clearly differentiate the influence these parameter spaces have in the MBRL setting. This allows optimizers to learn interaction effects within each group. When optimizing one group of hyperparameters, the others are set to the default value of the best manually tuned PETS hyperparameters as reported by Chua et al. (2018). We also optimized all hyperparameters together in Appendix A.4. The full configuration spaces are given in Appendix A.3. HPOObjective We use the average returns of the 3 most recent trials as the objective for all HPO methods. This gives us a better estimate of the noisy reward in the MBRL tasks. As discussed in Section 4, we consider two scenarios: 1) the transferability of the hyperparameter schedule (static or dynamic) learned by the HPO methods across environments; and 2) where we are interested in the final learned model and policy reward across multiple runs. In the first scenario, we consider the mean performance of the top 5 members during the search. For the latter scenario, we use the best found schedules of each HPO method and report the performance of PETS over 5 seeds”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the apparatus of Guzman in view of Jiang in view of Slutskyy to include that the probabilistic surrogate function is a gaussian mixture model formed from the combination of a plurality of neural networks, as taught by Zhang as disclosed above, in order to ensure an accurate generation of the trial set of hyperparameters (Zhang Page 1 “Finally, our experiments provide valuable insights into the effects of several hyperparameters, such as plan horizon or learning rate and their influence on the stability of training and resulting rewards.”).
With respect to claim 7, Guzman in view of Jiang in view of Slutskyy fail to explicitly disclose that the search of the domain space of hyperparameter inputs is guided by an evolutionary algorithm.
Zhang, however, teaches that the search of the domain space of hyperparameter inputs is guided by an evolutionary algorithm (See at least Zhang Page 3 “3.1 Population Based Training Population based training (PBT) is an evolutionary approach for dynamic HPO and allows to optimize hyperparameters during the training of the members of its population. PBT starts out with a randomly initialized population. As a result, all members of its population start from different regions in the hyperparameter configuration space. The members are ranked according to their current performance at regular intervals. The worst performing members in the population are then replaced by the best ones, by copying over their parameters (network weights, in case of neural networks NNs)), as well as their hyperparameters in an exploitation step. It is unclear from the original PBT paper, however, whether the data on which the members are trained are also copied over. We discuss ablations of this setting in the experiments section. To allow searching for potentially better performing hyperparameter configurations, the copied hyperparameters are perturbed to allow for small steps in their immediate neighborhoods in an exploration step. Over time, different configurations are evaluated during the training process and potentially kept if they improve performance. PBT comes with its own hyperparameters. In the original paper, members from the population are selected using truncation during the exploitation step. Thereby, agents from the bottom 20% of the population are replaced by agents from the top 20%. Additionally, continuous hyperparameter values are multiplied at random by 0.8 or 1.2 in the exploration step”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the apparatus of Guzman in view of Jiang in view of Slutskyy to include that the search of the domain space of hyperparameter inputs is guided by an evolutionary algorithm, as taught by Zhang as disclosed above, in order to ensure an accurate generation of the trial set of hyperparameters (Zhang Page 1 “Finally, our experiments provide valuable insights into the effects of several hyperparameters, such as plan horizon or learning rate and their influence on the stability of training and resulting rewards.”).
With respect to claim 8, Guzman in view of Jiang in view of Slutskyy fail to explicitly disclose that the model is the motion planner, and the guidance objective is configured to generate the at least one trial set of hyperparameters using an evolutionary algorithm to stochastically determine new data pairs in dependence on evaluating a utility score of one or more new sets of hyperparameter.
Zhang teaches that the model is the motion planner, and the guidance objective is configured to generate the at least one trial set of hyperparameters using an evolutionary algorithm to stochastically determine new data pairs in dependence on evaluating a utility score of one or more new sets of hyperparameters (See at least Zhang Page 3 “3.1 Population Based Training Population based training (PBT) is an evolutionary approach for dynamic HPO and allows to optimize hyperparameters during the training of the members of its population. PBT starts out with a randomly initialized population. As a result, all members of its population start from different regions in the hyperparameter configuration space. The members are ranked according to their current performance at regular intervals. The worst performing members in the population are then replaced by the best ones, by copying over their parameters (network weights, in case of neural networks NNs)), as well as their hyperparameters in an exploitation step. It is unclear from the original PBT paper, however, whether the data on which the members are trained are also copied over. We discuss ablations of this setting in the experiments section. To allow searching for potentially better performing hyperparameter configurations, the copied hyperparameters are perturbed to allow for small steps in their immediate neighborhoods in an exploration step. Over time, different configurations are evaluated during the training process and potentially kept if they improve performance. PBT comes with its own hyperparameters. In the original paper, members from the population are selected using truncation during the exploitation step. Thereby, agents from the bottom 20% of the population are replaced by agents from the top 20%. Additionally, continuous hyperparameter values are multiplied at random by 0.8 or 1.2 in the exploration step”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the apparatus of Guzman in view of Jiang in view of Slutskyy to include that the model is the motion planner, and the guidance objective is configured to generate the at least one trial set of hyperparameters using an evolutionary algorithm to stochastically determine new data pairs in dependence on evaluating a utility score of one or more new sets of hyperparameters, as taught by Zhang as disclosed above, in order to ensure an accurate generation of the trial set of hyperparameters (Zhang Page 1 “Finally, our experiments provide valuable insights into the effects of several hyperparameters, such as plan horizon or learning rate and their influence on the stability of training and resulting rewards.”).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to IBRAHIM ABDOALATIF ALSOMAIRY whose telephone number is (571)272-5653. The examiner can normally be reached M-F 7:30-5:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Faris Almatrahi can be reached at 313-446-4821. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/IBRAHIM ABDOALATIF ALSOMAIRY/ Examiner, Art Unit 3667 /KENNETH J MALKOWSKI/Primary Examiner, Art Unit 3667