DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 02/16/2026 has been entered.
Status of Claims
Applicant filed an RCE on 02/16/2026. Claims 1, 9, 19, and 20 have been amended and claim 21 was newly added. Claims 1-21 are currently pending examination.
Response to Arguments
Regarding the claim rejections under 35 USC 103: Applicant's arguments filed 02/16/2026 with respect to Yadmellat et al . (US20220219726A1) in view of Wolff et al. (US 20230415772 A1), have been fully considered but they are moot because the new ground of rejection does not rely on all reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument in view of Zadeh et al. (US20220121884A1).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-3, 5, 8-12, 14-15 and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over Yadmellat et al . (US20220219726A1) in view of Wolff et al. (US 20230415772 A1) and further in view of Zadeh et al. (US20220121884A1), herein after referred to as Yadmellat, Wolff and Zadeh.
Regarding claims 1 and 19, Yadmellat discloses
collecting a set of world scene parameters from an ego autonomous vehicle (AV) (“The sensor system 110 includes various sensing units, such as a radar unit 112, a LIDAR unit 114, and a camera 116, for collecting information about an environment surrounding the vehicle 100 as the vehicle 100 operates in the environment. ” [0045] and ¶ [0020]–[0021]);
selecting a rule hierarchy from a set of traveling rules for the ego AV (“Criteria are ranked by priority rank, which determines the sorting order.” [0009] see also FIG. 4 and ¶ [0076] (Venn diagram of prioritized criteria hierarchy);
calculating a robustness vector for one or more rules in the rule hierarchy utilizing a trajectory derived from each of the respective one or more rules and the set of world scene parameters (“Existing rule-based motion planning techniques typically require evaluation (e.g., comparison and sorting) of the generated trajectories according to explicitly defined cost functions that may take into account various objectives to calculate a cost associated with each trajectory being compared.” [0005],
“An overall cost function is used to calculate an overall cost 644 for each trajectory 522, 524, 526, 528 based on the objective costs 604, 606, 608, 610, 612 combined with a corresponding set of weights (wPhysical Limit, wNo Collision, wSafety, wComfort, wMobility).” [0105],
“the trajectory evaluator 334 in the present disclosure implements a rules-based trajectory evaluation operation for sorting candidate trajectories using a combination of hard constraints, called criteria, and soft constraints, called objectives, wherein each criterion is associated with an objective. The criteria are ranked relative to each other, with each criterion having a priority rank indicating the relative importance or necessity of satisfying the criterion. In some embodiments, each objective may also include other information, such as an objective cost function and/or a weight associated with the objective. Objectives, criteria, priority ranks, objective cost functions and weights are discussed in greater detail below.” [0067]).
assigning a reward parameter for the one or more rules in proportion to the respective robustness vector and the rank of the one or more rules (“The trajectory evaluator 334 receives as input the vehicle state, the planned path (or intermediate target positions), the vehicle kinodynamic parameters, and the candidate trajectories generated by the trajectory generator 332, and assigns an evaluation value to each candidate trajectory. The assigned evaluation value may be reflective of whether the candidate trajectory successfully achieves the goal of relatively safe, comfortable and speedy driving (and also satisfies the behavior decision), in accordance with various predetermined objectives. The trajectory selector 336 selects the candidate trajectory with the highest evaluation value (as assigned by the trajectory evaluator 334) among the candidate trajectories generated by the trajectory generator 332. In some embodiments, the candidate trajectories are sorted by the trajectory evaluator in order of evaluation value, such that the candidate trajectory with the highest evaluation value is first in sorted order.” [0064]);
generating a set of reward parameters from the one or more rules and respective of the assigned reward parameters (“The trajectory evaluator 334 receives as input the vehicle state, the planned path (or intermediate target positions), the vehicle kinodynamic parameters, and the candidate trajectories generated by the trajectory generator 332, and assigns an evaluation value to each candidate trajectory. The assigned evaluation value may be reflective of whether the candidate trajectory successfully achieves the goal of relatively safe, comfortable and speedy driving (and also satisfies the behavior decision), in accordance with various predetermined objectives.” [0064]);
selecting a preferred trajectory derived from the rule hierarchy using the set of reward parameters, wherein the preferred trajectory evaluates to a highest reward in the set of reward parameters (“The trajectory selector 336 selects the candidate trajectory with the highest evaluation value (as assigned by the trajectory evaluator 334) among the candidate trajectories generated by the trajectory generator 332. In some embodiments, the candidate trajectories are sorted by the trajectory evaluator in order of evaluation value, such that the candidate trajectory with the highest evaluation value is first in sorted order.” [0064]).
and controlling a movement of the ego AV utilizing the preferred trajectory (“software system for control of the vehicle (e.g. a vehicle control system) receives the trajectory from the planning system and generates control commands to control operation of the vehicle to follow the trajectory.” [0041]).
Yadmellat does not explicitly disclose
wherein the robustness vector represents an N-dimensional vector-valued function whose elements comprise a robustness scalar for each of the one or more rules,
where the robustness scalar is the degree of satisfaction for each of the one or more rules,
wherein each rule in the one or more rules is ranked so that the rank of each rule in the one or more rules is in proportion to the importance of each respective rule and higher importance rules in the one or more rules have a higher rank than lower importance rules in the one or more rules;
revising the rank of each rule in the one or more rules using a rank-preserving reward function, wherein the rank-preserving reward function applies a multiplier to the rank of each rule, determined in the calculating of the robustness vector, where the multiplier grows exponentially in relation to the increase in priority of the rule;
However, Wolff does teach wherein each rule in the one or more rules is ranked so that the rank of each rule in the one or more rules is in proportion to the importance of each respective rule and higher importance rules in the one or more rules have a higher rank than lower importance rules in the one or more rules; (“the trajectory evaluation policy can score the trajectories higher if one or more features of the trajectories satisfy corresponding feature thresholds and lower if they do not.” [0135],
“the trajectory evaluation policy can indicate that trajectories are to be evaluated based on safety and/or comfort thresholds. Accordingly, trajectories with a higher safety and/or comfort score can receive a higher score.” [0139]).
wherein the robustness vector represents an N-dimensional vector-valued function whose elements comprise a robustness scalar for each of the one or more rules, (“the feature embeddings can include one or more n-dimensional feature vectors. In some such cases, an individual feature vector may not correspond to an object attribute, but a combination of multiple n-dimensional feature vectors can contain information about an object's attributes, such as, but not limited to, its classification, width, length, height, etc.” [0105]).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to modify the prioritized-criteria trajectory evaluator of Yadmellat to also include ranking each rule so that the rank of each rule is in proportion to the importance of that rule and higher-importance rules have a higher rank than lower-importance rules and representing the robustness vector as an N-dimensional vector-valued function whose elements comprise a robustness scalar for each of the one or more rules, as taught by Wolff, with a reasonable expectation of success.. Using Wolff’s importance-proportional scoring as the rank of Yadmellat’s criteria keeps the existing hierarchy while making rank track relative necessity of each rule. (With regard to this reasoning, see at least Wolff, [0105],[0135], [0139].)
However, Zadeh does teach where the robustness scalar is the degree of satisfaction for each of the one or more rules, (“We can change the assumptions one at a time, and see the results again, until “satisfied”, which is also a fuzzy concept (for the degree of “satisfaction”). ….. based on some fuzzy rules.” [2052],
“First, we fuzzify the input space. Then, using data, we produce fuzzy rules. Then, for each rule, we assign a degree, followed by the creation of the combined rule library.” [2244],
“we use a vector matching representation (for possible partial matching), using non-binary weights to index terms in documents or queries (for degree of similarity). Thus, the cosine of angle between 2 given vectors is an indication of similarity of the 2 vectors, which can be obtained by a probability ranking principle, or ranking based on relevant and non-relevant information.” [2425]).
revising the rank of each rule in the one or more rules using a rank-preserving reward function, wherein the rank-preserving reward function applies a multiplier to the rank of each rule, determined in the calculating of the robustness vector, where the multiplier grows exponentially in relation to the increase in priority of the rule; (“we use Crank to adjust the list or ranking, as a feedback to the system (which, in one embodiment, generally is not a linear function of or proportional to Crank at all).” [2390],
“the ranking order is higher for more relevant match, more reliable match, more similar match, in-stock status, higher merchant's score, and/or higher ……..a ranking function is calculated based on for example, weighted dependency on such ranking factors with linear, quadratic, polynomial, or exponential, or reciprocal (e.g., inverse of powers of inverse) dependencies.” [2964]).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to further modify the combined Yadmellat/Wolff robustness vector so that the robustness scalar is the degree of satisfaction for each of the one or more rules and revising the rank of each rule using a rank-preserving reward function that applies a multiplier to the rank of each rule, determined in the calculating of the robustness vector, where the multiplier grows exponentially in relation to the increase in priority of the rule, as taught by Zadeh, with a reasonable expectation of success. Using Zadeh’s per-rule degree of satisfaction as the scalar inside Wolff’s n-dimensional vector converts Yadmellat’s binary-or-cost criterion result into a continuous satisfaction measure without changing the evaluator architecture. (With regard to this reasoning, see at least Zadeh, [2052], [2244], [2425] , [2390], [2964]).
Regarding claims 2 and 10, Yadmellat discloses The method as recited in Claim 1, further comprising: communicating the preferred trajectory to an AV controller to select an appropriate motion primitive control action for the ego AV (“The method 700 generally begins with trajectory generation, followed by trajectory evaluation, then initially sorting the trajectories according to their criterion satisfaction data, then sorting within the categories assigned by the initial trajectory sorting operation based on a cost function, then trajectory selection based on the sorted order of the candidate trajectories, then commanding the autonomous vehicle to execute the selected trajectory.” [0112]).
Regarding claims 3 and 14, Yadmellat discloses The method as recited in Claim 1, wherein the calculating the robustness vector further comprises:
selecting a set of motion primitives that the ego AV can utilize (“The vehicle control system 140 serves to control operation of the vehicle 100 based on the trajectory output by the planning system 130. The vehicle control system 140 may be used to generate control signals for the electromechanical components of the vehicle 100 to control the motion of the vehicle 100. The electromechanical system 150 receives control signals from the vehicle control system 140 to operate the electromechanical components of the vehicle 100 such as an engine, transmission, steering system and braking system.” [0051]);
computing a set of trajectory locations for the ego AV over a determined number of time-steps by simulating a propagation of the ego AV using the set of motion primitives (“In the present disclosure, a trajectory is a sequence, over multiple time steps, of a spatial position for the autonomous vehicle (in a geometrical coordinate system) and other parameters. Other parameters may include vehicle orientation, vehicle velocity, vehicle acceleration, vehicle jerk or any combination thereof.” [0003]
(“The objectives used by a trajectory evaluation operation may include objectives related to safety, comfort, and mobility (i.e. moving the vehicle toward its destination). Existing rule-based motion planning techniques typically require evaluation (e.g., comparison and sorting) of the generated trajectories according to explicitly defined cost functions that may take into account various objectives to calculate a cost associated with each trajectory being compared. The trajectories being compared may then be sorted by their associated estimated costs. A typical cost function is computed by combining weighted costs associated with the various objectives into an overall cost function. For example, a cost function in a conventional trajectory evaluator may consider the objectives (comfort, safety, mobility) in determining a cost function.” [0005]);
and modifying the robustness vector using an evaluation of the one or more rules at one or more trajectory locations in the set of trajectory locations (“Weights may be any linear or non-linear modifier applied to an objective cost function to enable its combination with other objective cost functions to generate an overall cost function used to sort candidate trajectories, as described above with reference to FIG. 3. In some embodiments, weights are constant scalar values used as multipliers to generate an overall cost function consisting of a weighted sum of individual objective costs, such that the weights are used as coefficients multiplying the respective objective costs to generate the weighted sum. In other embodiments, weights may apply a non-linear modifier to an objective cost function, such as an exponential modifier. The overall cost function may combine the weighted objective cost functions for a plurality of objectives through addition or any other means of combination.” [0084]).
Regarding claim 5, Yadmellat discloses The method as recited in Claim 3, wherein the set of world scene parameters are updated at one or more time-steps in the number of time-steps (“Selecting a route to travel through a set of roads is an example of mission planning. Generally, the final destination point, once set (e.g., by user input) is unchanging through the duration of the journey. Although the final destination point may be unchanging, the path planned by mission planning may change through the duration of the journey. For example, changing traffic conditions may require mission planning to dynamically update the planned path to avoid a congested road.” [0059]).
Regarding claim 8, Yadmellat discloses The method as recited in Claim 1, wherein the assigning the reward parameter further comprises: modifying the respective robustness vector by applying an average robustness (“Weights may be any linear or non-linear modifier applied to an objective cost function to enable its combination with other objective cost functions to generate an overall cost function used to sort candidate trajectories, as described above with reference to FIG. 3.” [0084]).
Regarding claims 9 and 20, Yadmellat discloses a system, comprising:
a receiver, operational to receive a set of world scene parameters (“Information collected by each sensing unit of the sensor system 110 is provided as sensor data to the perception system 120. The perception system 120 processes the sensor data received from each sensing unit to generate data about the vehicle and data about the surrounding environment.” [0046]),
a set of hyperparameters (“which may be in absolute geographical longitude/latitudinal values and/or values that reference other frames of reference), data representing kinodynamic parameters, and data representing the physical parameters of the vehicle, such as width and length, mass, inertia, wheelbase, slip angle, cornering forces, and data about the motion of the vehicle, such as linear speed and acceleration, travel direction, angular acceleration, pose (e.g., pitch, yaw, roll), and vibration, and mechanical system operating parameters such as engine RPM, throttle position, brake position, and transmission gear ratio, etc.)” [0046]),
a set of ego autonomous vehicle (AV) states for an ego AV (“The data about the environment and the data about the vehicle 100 output by the perception system 120 is received by the state generator 125. The state generator 125 processes about the environment and the data about the vehicle 100 to generate a state for the vehicle 100 (hereinafter vehicle state). Although the state generator 125 is shown in FIG. 3 as a separate software system, in some embodiments, the state generator 125 may be included in the perception system 120 or in the planning system 130.” [0050]),
and a set of traveling rules for the ego AV, wherein the ego AV is intending to move (“e behavior planner 320 may generate a behavior decision that is in accordance with certain rules or driving preferences.” [0062]);
and one or more processors (“The perception system 120, the planning system 130, and the vehicle control system 140 in this example are distinct software systems that include machine readable instructions that may be executed by one or more processors in a processing system of the vehicle 100.” [0043]), operational to calculate a robustness vector for one or more trajectories derived from one or more rules in a rule hierarchy and select a preferred trajectory from the one or more trajectories using a reward parameter derived from the robustness vector, (“The objectives used by a trajectory evaluation operation may include objectives related to safety, comfort, and mobility (i.e. moving the vehicle toward its destination). Existing rule-based motion planning techniques typically require evaluation (e.g., comparison and sorting) of the generated trajectories according to explicitly defined cost functions that may take into account various objectives to calculate a cost associated with each trajectory being compared. The trajectories being compared may then be sorted by their associated estimated costs. A typical cost function is computed by combining weighted costs associated with the various objectives into an overall cost function. For example, a cost function in a conventional trajectory evaluator may consider the objectives (comfort, safety, mobility) in determining a cost function.” [0005]),
where the rule hierarchy is selected from the set of traveling rules, a calculation for the robustness vector utilizes the set of world scene parameters, the set of hyperparameters, and the set of ego AV states and where a movement of the ego AV is controlled by the preferred trajectory. (“The trajectory evaluator 334 receives as input the vehicle state, the planned path (or intermediate target positions), the vehicle kinodynamic parameters, and the candidate trajectories generated by the trajectory generator 332, and assigns an evaluation value to each candidate trajectory. The assigned evaluation value may be reflective of whether the candidate trajectory successfully achieves the goal of relatively safe, comfortable and speedy driving (and also satisfies the behavior decision), in accordance with various predetermined objectives. The trajectory selector 336 selects the candidate trajectory with the highest evaluation value (as assigned by the trajectory evaluator 334) among the candidate trajectories generated by the trajectory generator 332. In some embodiments, the candidate trajectories are sorted by the trajectory evaluator in order of evaluation value, such that the candidate trajectory with the highest evaluation value is first in sorted order.” [0064]).
Yadmellat does not explicitly disclose
wherein each rule in the one or more rules is ranked so that the rank of each rule in the one or more rules is in proportion to the importance of each respective rule and higher importance rules in the one or more rules have a higher rank than lower importance rules in the one or more rules,
and the robustness vector represents an N- dimensional vector-valued function whose elements comprise a robustness scalar for each of the one or more rules,
where the robustness scalar is the degree of satisfaction for each of the one or more rules;
revising the rank of each rule in the one or more rules using a rank-preserving reward function, wherein the rank-preserving reward function applies a multiplier to the rank of each rule, determined in the calculating of the robustness vector, where the multiplier grows exponentially in relation to the increase in priority of the rule,
However, Wolff does teach wherein each rule in the one or more rules is ranked so that the rank of each rule in the one or more rules is in proportion to the importance of each respective rule and higher importance rules in the one or more rules have a higher rank than lower importance rules in the one or more rules; (“the trajectory evaluation policy can score the trajectories higher if one or more features of the trajectories satisfy corresponding feature thresholds and lower if they do not.” [0135],
“the trajectory evaluation policy can indicate that trajectories are to be evaluated based on safety and/or comfort thresholds. Accordingly, trajectories with a higher safety and/or comfort score can receive a higher score.” [0139]).
wherein the robustness vector represents an N-dimensional vector-valued function whose elements comprise a robustness scalar for each of the one or more rules, (“the feature embeddings can include one or more n-dimensional feature vectors. In some such cases, an individual feature vector may not correspond to an object attribute, but a combination of multiple n-dimensional feature vectors can contain information about an object's attributes, such as, but not limited to, its classification, width, length, height, etc.” [0105]).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to modify the prioritized-criteria trajectory evaluator of Yadmellat to also include ranking each rule so that the rank of each rule is in proportion to the importance of that rule and higher-importance rules have a higher rank than lower-importance rules and representing the robustness vector as an N-dimensional vector-valued function whose elements comprise a robustness scalar for each of the one or more rules, as taught by Wolff, with a reasonable expectation of success.. Using Wolff’s importance-proportional scoring as the rank of Yadmellat’s criteria keeps the existing hierarchy while making rank track relative necessity of each rule. (With regard to this reasoning, see at least Wolff, [0105],[0135], [0139]).
However, Zadeh does teach where the robustness scalar is the degree of satisfaction for each of the one or more rules, (“We can change the assumptions one at a time, and see the results again, until “satisfied”, which is also a fuzzy concept (for the degree of “satisfaction”). ….. based on some fuzzy rules.” [2052],
“First, we fuzzify the input space. Then, using data, we produce fuzzy rules. Then, for each rule, we assign a degree, followed by the creation of the combined rule library.” [2244],
“we use a vector matching representation (for possible partial matching), using non-binary weights to index terms in documents or queries (for degree of similarity). Thus, the cosine of angle between 2 given vectors is an indication of similarity of the 2 vectors, which can be obtained by a probability ranking principle, or ranking based on relevant and non-relevant information.” [2425]).
revising the rank of each rule in the one or more rules using a rank-preserving reward function, wherein the rank-preserving reward function applies a multiplier to the rank of each rule, determined in the calculating of the robustness vector, where the multiplier grows exponentially in relation to the increase in priority of the rule; (“we use Crank to adjust the list or ranking, as a feedback to the system (which, in one embodiment, generally is not a linear function of or proportional to Crank at all).” [2390],
“the ranking order is higher for more relevant match, more reliable match, more similar match, in-stock status, higher merchant's score, and/or higher ……..a ranking function is calculated based on for example, weighted dependency on such ranking factors with linear, quadratic, polynomial, or exponential, or reciprocal (e.g., inverse of powers of inverse) dependencies.” [2964]).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to further modify the combined Yadmellat/Wolff robustness vector so that the robustness scalar is the degree of satisfaction for each of the one or more rules and revising the rank of each rule using a rank-preserving reward function that applies a multiplier to the rank of each rule, determined in the calculating of the robustness vector, where the multiplier grows exponentially in relation to the increase in priority of the rule, as taught by Zadeh, with a reasonable expectation of success. Using Zadeh’s per-rule degree of satisfaction as the scalar inside Wolff’s n-dimensional vector converts Yadmellat’s binary-or-cost criterion result into a continuous satisfaction measure without changing the evaluator architecture. (With regard to this reasoning, see at least Zadeh, [2052], [2244], [2425] , [2390], [2964]).
Regarding claim 11, Yadmellat discloses The system as recited in Claim 9, where the one or more processors are a reward rule hierarchy analyzer (“The trajectory selector 336 selects the candidate trajectory with the highest evaluation value (as assigned by the trajectory evaluator 334) among the candidate trajectories generated by the trajectory generator 332.” [0064]).
Regarding claim 12, Yadmellat discloses The system as recited in Claim 9, where the one or more processors are one or more of a central processing unit or one or more of a graphics processing unit (“The processing system 200 includes one or more processors 210. The one or more processors 210 may include a central processing unit (CPU), a graphical processing unit (GPU), a tensor processing unit (TPU), a neural processing unit (NPU), a digital signal processor, and/or another computational element.”[0052]).
Regarding claim 15, Yadmellat discloses The system as recited in Claim 9, wherein the one or more processors is further operational to utilize one or more optimizers when deriving the reward parameter from the robustness vector (“The objectives used by a trajectory evaluation operation may include objectives related to safety, comfort, and mobility (i.e. moving the vehicle toward its destination). Existing rule-based motion planning techniques typically require evaluation (e.g., comparison and sorting) of the generated trajectories according to explicitly defined cost functions that may take into account various objectives to calculate a cost associated with each trajectory being compared. The trajectories being compared may then be sorted by their associated estimated costs. A typical cost function is computed by combining weighted costs associated with the various objectives into an overall cost function. For example, a cost function in a conventional trajectory evaluator may consider the objectives (comfort, safety, mobility) in determining a cost function.” [0005]), Examiner Notes: the calculation of the cost functions based on the rules of the planning techniques increases the robustness of best trajectory selection.
Regarding claim 17, Yadmellat discloses The system as recited in Claim 9, further comprising: a driver assisted vehicle including a vehicle control system that directs an operation of the driver assisted vehicle, wherein the one or more processors provide input to the vehicle control system (“or example, any vehicle that includes advanced driver-assistance system for a vehicle that includes a planning system may benefit from a motion planner that performs the trajectory generation, trajectory evaluation, trajectory selection operations of the present disclosure.” [0130]).
Regarding claim 18, Yadmellat discloses The system as recited in Claim 9, wherein the receiver and the one or more processors are part of an integrated circuit (“Alternatively, the perception system 120, the planning system 130, and the vehicle control system 140 may be distinct systems on one or more chips (e.g., application-specific integrated circuit (ASIC), field-programmable gate array (FGPA), and/or other type of chip).” [0043]).
Claims 4, 6-7, 13 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Yadmellat in view of Wolff and further in view of Zadeh as applied to claim 1 above, and in further view of He et al. (US20200363813A1), hereinafter referred to as He.
Regarding claims 4 and 13,
Yadmellat in view of Wolff does not explicitly disclose wherein the computing the set of trajectory locations and the modifying the robustness vector is performed in parallel for the one or more rules
However, He does teach wherein the computing the set of trajectory locations and the modifying the robustness vector is performed in parallel for the one or more rules (“In one embodiment, RL agent 803 includes an actor critic framework. The actor can include a policy function generator to generate a list of controls (or actions) from a current trajectory state and the critic includes a value function to determine value predictions for the controls generated by the actor. In one embodiment, the actor-critic framework includes an actor neural network 805 coupled to a critic neural network 807. In one embodiment, actor neural network 805 and/or critic neural network 807 are deep neural networks, isolated from each other. In another embodiment, actor neural network 805 and critic neural network 807 runs in parallel.” [0065]).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to modify the trajectory-location computation and robustness-vector modification of Yadmellat in view of Wolff and Zadeh so that those two operations are performed in parallel for the one or more rules, as taught by He, with a reasonable expectation of success. Executing Yadmellat’s per-rule location simulation and robustness update on He’s parallel actor-critic pattern evaluates every rule on the same time-step set without serializing the hierarchy. (With regard to this reasoning, see at least He, [0065].)
Regarding claim 6,
Yadmellat in view of Wolff does not explicitly disclose wherein a sigmoid function is applied to a step function for calculating the set of reward parameters.
However, He does teach wherein a sigmoid function is applied to a step function for calculating the set of reward parameters (“In one embodiment, critic neural network can include a MLP. The critic neural network includes one or more hidden layers that include a second set of weights that must be optimized separately from the first set of weights of the actor neural network. Note, the hidden layers and/or the output layers can have different activation function, such as a linear, sigmoid, tanh, RELU, softmax, etc.” [0066]).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to modify the reward-parameter calculation of Yadmellat in view of Wolff and Zadeh to also include applying a sigmoid function to a step function for calculating the set of reward parameters, as taught by He, with a reasonable expectation of success. Passing Yadmellat’s step-like criterion-satisfaction result through He’s sigmoid yields a smooth, bounded reward that remains differentiable for later ranking. (With regard to this reasoning, see at least He, [0066].)
Regarding claim 7,
Yadmellat in view of Wolff does not explicitly disclose wherein the assigning the reward parameter further comprises: applying a stochastic optimization to the reward parameter over a determined set of time-steps, wherein a motion primitive is applied to the ego AV at one or more time-steps in the set of time-steps
However, He does teach wherein the assigning the reward parameter further comprises: applying a stochastic optimization to the reward parameter over a determined set of time-steps, wherein a motion primitive is applied to the ego AV at one or more time-steps in the set of time-steps (“Based on driving statistics 123, machine learning engine 122 generates or trains a set of rules, algorithms, and/or predictive models 124 for a variety of purposes. In one embodiment, algorithms/models 124 may include a bicycle model to model the vehicle dynamics for the ADV, an open space optimization model or an RL agent/environment model to plan a trajectory for the ADV in an open space. Algorithms/models 124 can then be uploaded on ADVs (e.g., models 313 of FIG. 3A) to be utilized by the ADVs in real-time.” [0038]).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to modify the reward-parameter assignment of Yadmellat in view of Wolff and Zadeh to also include applying a stochastic optimization to the reward parameter over a determined set of time-steps, wherein a motion primitive is applied to the ego AV at one or more time-steps in that set, as taught by He, with a reasonable expectation of success. Running Yadmellat’s reward through He’s stochastic optimizer, while applying a motion primitive at each time-step as He does, searches the same time-indexed action space instead of scoring only a finished open-loop path. (With regard to this reasoning, see at least He, [0038].)
Regarding claim 16,
Yadmellat in view of Wolff does not explicitly disclose wherein the one or more processors is further operational to utilize a machine learning system utilizing a trajectory model and the rule hierarchy
However, He does teach wherein the one or more processors is further operational to utilize a machine learning system utilizing a trajectory model and the rule hierarchy (“Based on driving statistics 123, machine learning engine 122 generates or trains a set of rules, algorithms, and/or predictive models 124 for a variety of purposes.” [0038]).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to modify the processors of Yadmellat in view of Wolff and Zadeh to also utilize a machine-learning system that uses a trajectory model and the rule hierarchy, as taught by He, with a reasonable expectation of success. Substituting He’s learned trajectory-plus-rules model for Yadmellat’s hand-weighted cost functions lets the same hierarchy drive a trained policy rather than a static weighted sum. (With regard to this reasoning, see at least He, [0038].)
Claim 21 is rejected under 35 U.S.C. 103 as being unpatentable over Yadmellat in view of Wolff and further in view of Zadeh as applied to claim 1 above, and further in view of Gray et al. (US20090158223A1), herein after referred to as Gary.
Regarding claim 21,
The combination of Yadmellat, Wolff and Zadeh does not explicitly teach wherein the multiplier is equal to a base number raised to a power, where the base number is greater than 2 and the power is a total number of rules in the rule hierarchy minus a current rank position of each rule plus one, where the one or more rules are ordered from highest rank to lowest rank.
However, Gary does teach wherein the multiplier is equal to a base number raised to a power, where the base number is greater than 2 and the power is a total number of rules in the rule hierarchy minus a current rank position of each rule plus one, where the one or more rules are ordered from highest rank to lowest rank. (“The embodiments herein use a tradeoff multiplier to differentiate priorities:
W(e i,j)=m i% W(e i−1,j′)
where mi is the multiplier between the adjacent priorities, ei,jχpi, and ei−1,j′χpi−1. The weights of location perturbation (p0), W(e0j), are determined based on the input layout. Then the other weights, W(eij), i>0, are determined using the formula.” [0045],
“to differentiate different priorities of objectives (e.g., to differentiate spacePert from locPert), the embodiments herein assign the weight of a higher priority (pi) to be a multiple of the weight of a lower priority (pi−1) where w(pi)=miw(pi−1). …….. the multipliers can be scaled up to further differentiate different priorities if desired as long as the total cost is within the trustable range.” [0014]).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to modify the rank-preserving multiplier of Yadmellat in view of Wolff and Zadeh so that the multiplier is equal to a base number raised to a power, where the base number is greater than 2 and the power is the total number of rules in the rule hierarchy minus the current rank position of each rule plus one, with the rules ordered from highest rank to lowest rank, as taught by Gray, with a reasonable expectation of success. Zadeh already teaches an exponential dependence of ranking weight on priority; Gray differentiates adjacent priorities by setting the weight of a higher priority equal to a multiplier m times the weight of the next-lower priority, and allows that multiplier to be scaled up so that higher-priority objectives dominate. Instantiating Zadeh’s exponential multiplier as Gray’s (m^{N-r+1}) form (base (m>2), exponent equal to remaining ranks plus one) produces a concrete rank-preserving scale that still cannot be overturned by any combination of lower-ranked rules. (With regard to this reasoning, see at least Gray, [0014], [0045].)
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to AHMED ALKIRSH whose telephone number is (703) 756-4503. The examiner can normally be reached M-F 9:00 am-5:00 pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, FADEY JABR can be reached on (571) 272-1516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/A.A./Examiner, Art Unit 3668
/Fadey S. Jabr/Supervisory Patent Examiner, Art Unit 3668