Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
This action is in response to the original filing on 02/19/2024. Claims 1-11 are pending for examination.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “acquisition unit configured to”, “output unit configured to”, “determination unit configured to”, “learning unit configured to” as recited in Claim 1, “determination unit configured to” as recited in Claim 2, “determination unit configured to” and “learning unit is configured to” as recited in Claim 3, “determination unit is configured to” as recited in Claim 4, “determination unit is configured to” as recited in Claim 5, “goal setting unit configured to ” and “determination unit is configured to” as recited in Claim 6, and “goal setting unit configured to ” as recited in Claims 7-9,
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-10 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Claims 1 and 10
Step 1: Claims 1 and 10 recite a machine, a method, and a computer programming product. As such, they are directed to the statutory categories of a machine, method, and product of manufacture.
Step 2A Prong 1: The claims recite, inter alia: “acquire observation information including information on a speed of a control target point at a control target time;” under its broadest reasonable interpretation in light of the specification, this amounts to no more than a mental process comprising simple observation.
The claims further recite: “output control information including information on speed control of the control target point, the control information being determined in accordance with the observation information and a control policy;” under its broadest reasonable interpretation in light of the specification, this limitation amounts to no more than the mental process of making a decision when presented with observational data.
The claims further recite: “determine a corrected reward obtained by correcting a reward in accordance with a speed of the control target point derived from the observation information, the reward being higher as an error between a value of an evaluation parameter and a goal is smaller, the evaluation parameter being a parameter other than a speed derived from the observation information; and” under its broadest reasonable interpretation in light of the specification, this limitation amounts to no more than the mental process of a determination and/or judgement, as what is being done amounts to adjustment of a reward based on performance data between speed and another observed parameter. This is thus easily accomplished by a human being with aid of a pen and paper.
Step 2A Prong 2: The claims recites the additional elements of “an acquisition unit configured to”, “an output unit configured to”, and “a corrected reward determination unit configured to” amounts to no more than mere instructions to apply an exception, wherein a generic computer is linked to an abstract idea (see MPEP 2106.05(f)), and is known to be well-understood, routine, and conventional within the art (see MPEP 2106.05(d) Subsection II).
The claims further recite: “a learning unit configured to perform reinforcement learning of the control policy based on the observation information and the corrected reward.”, amounts to no more than a mere high level of generality recitation of the words “apply it”, as no details that reflect the inventive concept i.e. no recitation as to how the reinforcement learning is performed other than that it is based on observed information, corrected reward, and corrected discount rate (See MPEP 2106.05(f)).
Step 2B: The claims do not recite significantly more than the judicial exception. The claims recites the additional elements of “an acquisition unit configured to”, “an output unit configured to”, and “a corrected reward determination unit configured to” amounts to no more than mere instructions to apply an exception, wherein a generic computer is linked to an abstract idea (see MPEP 2106.05(f)).
The claims further recite: “a learning unit configured to perform reinforcement learning of the control policy based on the observation information and the corrected reward.”, amounts to no more than a mere high level of generality recitation of the words “apply it”, as no details that reflect the inventive concept i.e. no recitation as to how the reinforcement learning is performed other than that it is based on observed information, corrected reward, and corrected discount rate (See MPEP 2106.05(f)).
Claim 2
Step 1: Claim 2 recites a machine. As such, the claim is directed to the statutory category of a machine.
Step 2A Prong 1: The claim recites: “determine the corrected reward that is corrected so as to be lower as a speed of the control target point is higher.” Under its broadest reasonable interpretation in light of the specification, this amounts to no more than making a determination and/or judgement such that one value remains lower than another.
Step 2A Prong 2: The claim recites the additional element of: “the corrected reward determination unit configured to”, amounts to no more than mere instructions to apply an exception, wherein a generic computer is linked to an abstract idea (see MPEP 2106.05(f)).
Step 2B: The claims do not recite significantly more than the judicial exception. he claim recites the additional element of: “the corrected reward determination unit configured to”, amounts to no more than mere instructions to apply an exception, wherein a generic computer is linked to an abstract idea (see MPEP 2106.05(f)), and is known to be well-understood, routine, and conventional within the art (see MPEP 2106.05(d) Subsection II).
Claim 3
Step 1: Claim 3 recites a machine. As such, the claim is directed to the statutory category of a machine.
Step 2A Prong 1: The claim recites: “determine a corrected discount rate obtained by correcting a discount rate of the corrected reward in accordance with a speed of the control target point derived from the observation information,” Under its broadest reasonable interpretation in light of the specification, this amounts to no more than making a determination and/or judgement of a value when taking into account two other values derived from an observation.
Step 2A Prong 2: The claim recites the additional element of: “a corrected discount rate determination unit configured to”, amounts to no more than mere instructions to apply an exception, wherein a generic computer is linked to an abstract idea (see MPEP 2106.05(f)).
The claims also recite the additional element of: “wherein the learning unit is configured to perform reinforcement learning of the control policy based on the observation information, the corrected reward, and the corrected discount rate.” Which amounts to no more merely linking an abstract idea to a particular technological environment, the environment being reinforcement learning (see MPEP 2106.05(h))
Step 2B: The claims do not recite significantly more than the judicial exception. he claim recites the additional element of: “a corrected discount rate determination unit configured to”, amounts to no more than mere instructions to apply an exception, wherein a generic computer is linked to an abstract idea (see MPEP 2106.05(f)), and is known to be well-understood, routine, and conventional within the art (see MPEP 2106.05(d) Subsection II).
The claims also recite the additional element of: “wherein the learning unit is configured to perform reinforcement learning of the control policy based on the observation information, the corrected reward, and the corrected discount rate.” Which amounts to no more merely linking an abstract idea to a particular technological environment, the environment being reinforcement learning (see MPEP 2106.05(h))
Claim 4
Step 1: Claim 4 recites a machine. As such, the claim is directed to the statutory category of a machine.
Step 2A Prong 1: The claim recites: “determine the corrected discount rate that is corrected such that a value of the discount rate is smaller as the speed of the control target point is higher.” Under its broadest reasonable interpretation in light of the specification, this amounts to no more than making a determination and/or judgement of a value and determining that the specified value is lower than another.
Step 2A Prong 2: The claim recites the additional element of: “wherein the corrected discount rate determination unit is configured to determine the corrected discount rate that is corrected such that a value of the discount rate is smaller as the speed of the control target point is higher.”, amounts to no more than a mere high level of generality recitation of the words “apply it”, as no details that reflect the inventive concept i.e. no recitation as to how the reinforcement
Step 2B: The claims do not recite significantly more than the judicial exception. he claim recites the additional element of: “wherein the corrected discount rate determination unit is configured to”, amounts to no more than mere instructions to apply an exception, wherein a generic computer is linked to an abstract idea (see MPEP 2106.05(f)), and is known to be well-understood, routine, and conventional within the art (see MPEP 2106.05(d) Subsection II).
Claim 5
Step 1: Claim 2 recites a machine. As such, the claim is directed to the statutory category of a machine.
Step 2A Prong 1: The claim recites: “determine the corrected discount rate obtained by correcting, in accordance with the speed of the control target point, the discount rate in accordance with an input discount rate for an input speed that has been input.” Under its broadest reasonable interpretation in light of the specification, this amounts to no more than making a determination and/or judgement such that one value remains lower than another.
Step 2A Prong 2: The claim recites the additional element of: “wherein the corrected discount rate determination unit is configured to”, amounts to no more than mere instructions to apply an exception, wherein a generic computer is linked to an abstract idea (see MPEP 2106.05(f)).
Step 2B: The claims do not recite significantly more than the judicial exception. he claim recites the additional element of: “wherein the corrected discount rate determination unit is configured to”, amounts to no more than mere instructions to apply an exception, wherein a generic computer is linked to an abstract idea (see MPEP 2106.05(f)), and is known to be well-understood, routine, and conventional within the art (see MPEP 2106.05(d) Subsection II).
Claim 6
Step 1: Claim 6 recites a machine. As such, the claim is directed to the statutory category of a machine.
Step 2A Prong 1: The claim recites: “as goals, a plurality of evaluation parameter values, based on a group consisting of: Setting a goal first experience data including a value of an evaluation parameter derived from the acquired observation information; and” Under its broadest reasonable interpretation in light of the specification, this amounts to no more than the mental process of obtaining data from an evaluated parameter, a process which is easily performed by the human mind.
The claim also recites: “one or more pieces of second experience data including values of an evaluation parameter derived from one or more pieces of other observation information different in control target time from the acquired observation information, each of the plurality of evaluation parameter values selected from a plurality of evaluation parameter values included in the group,” Under its broadest reasonable interpretation in light of the specification, this amounts to no more than the mental process of obtaining data from an evaluated parameter, a process which is easily performed by the human mind.
“determine, for the set goals, corrected rewards obtained by correcting rewards in accordance with the speed of the control target point included in the observation information,” Under its broadest reasonable interpretation in light of the specification, this amounts to no more than making a determination and/or judgement comprising determining a given reward for goals that have been set in order to reach those goals.
“the rewards being higher as errors between the values of the evaluation parameter derived from the acquired observation information and the goals are smaller.” Under its broadest reasonable interpretation in light of the specification, this amounts to no more than making a determination and/or judgement comprising determining a given reward for goals, the reward being higher should the error rate between two observed values remain low.
Step 2A Prong 2: The claim recites the additional element of: “further comprising a goal setting unit configured to set,”, and “wherein the corrected reward determination unit is configured to” these limitations amount to no more than mere instructions to apply an exception, wherein a generic computer is linked to an abstract idea (see MPEP 2106.05(f)).
Step 2B: The claims do not recite significantly more than the judicial exception. he claim recites the additional element of: “further comprising a goal setting unit configured to set,”, and “wherein the corrected reward determination unit is configured to” these limitations amount to no more than mere instructions to apply an exception, wherein a generic computer is linked to an abstract idea (see MPEP 2106.05(f)), and is known to be well-understood, routine, and conventional within the art (see MPEP 2106.05(d) Subsection II).
Claim 7
Step 1: Claim 2 recites a machine. As such, the claim is directed to the statutory category of a machine.
Step 2A Prong 1: The claim recites: “as the goals, the value of the evaluation parameter included in the first experience data, and”, Under its broadest reasonable interpretation in light of the specification, this amounts to no more than making a determination and/or judgement made when setting the value of a goal.
The claims also recite: “a noise-added value of an evaluation parameter in which noise is added to a value of an evaluation parameter included in the second experience data.” Under its broadest reasonable interpretation in light of the specification, this amounts to no more than the mental process of adding randomness to data or a set of data, something easily performable my a human being with the aid of pen and paper.
Step 2A Prong 2: The claim recites the additional element of: “wherein the goal setting unit is configured to set,”, amounts to no more than mere instructions to apply an exception, wherein a generic computer is linked to an abstract idea (see MPEP 2106.05(f)).
Step 2B: The claims do not recite significantly more than the judicial exception. he claim recites the additional element of: “wherein the goal setting unit is configured to set,”, amounts to no more than mere instructions to apply an exception, wherein a generic computer is linked to an abstract idea (see MPEP 2106.05(f)), and is known to be well-understood, routine, and conventional within the art (see MPEP 2106.05(d) Subsection II).
Claim 8
Step 1: Claim 2 recites a machine. As such, the claim is directed to the statutory category of a machine.
Step 2A Prong 1: The claim recites: “set the goals in accordance with a goal selecting method selected by a user.” Under its broadest reasonable interpretation in light of the specification, this amounts to no more than setting a goal such that the goal aligns with a previous selection made by a user, a mental process easily performed in the mind.
Step 2A Prong 2: The claim recites the additional element of: “wherein the goal setting unit is configured to”, amounts to no more than mere instructions to apply an exception, wherein a generic computer is linked to an abstract idea (see MPEP 2106.05(f)).
Step 2B: The claims do not recite significantly more than the judicial exception. he claim recites the additional element of: “wherein the goal setting unit is configured to”, amounts to no more than mere instructions to apply an exception, wherein a generic computer is linked to an abstract idea (see MPEP 2106.05(f)), and is known to be well-understood, routine, and conventional within the art (see MPEP 2106.05(d) Subsection II).
Claim 9
Step 1: Claim 2 recites a machine. As such, the claim is directed to the statutory category of a machine.
Step 2A Prong 1: Claim 9 merely narrows the previously recited abstract limitations. For the reasons described above with respect to Claim 6, this judicial exception is not meaningfully integrated into a practical application, or significantly more than the abstract idea. The claims disclose similar limitations described for the independent claim above and does not provide anything more than the mental processes that are practically capable of being performed in the human mind with the assistance of pen and paper and mathematical concepts that are achievable through mathematical computation.
Step 2A Prong 2: The claim recites the additional element of: “wherein the goal setting unit is configured to set the goals, a number of which is selected by a user”, amounts to no more than insignificant extra solution activity, such as selection of a particular data source or data type to be manipulated, most similar to limiting a database index to only be made of XML tags (see MPEP 2106.05(g)).
Step 2B: The claims do not recite significantly more than the judicial exception. he claim recites the additional element of: “wherein the goal setting unit is configured to set the goals, a number of which is selected by a user”, amounts to no more than insignificant extra solution activity, such as selection of a particular data source or data type to be manipulated, most similar to limiting a database index to only be made of XML tags (see MPEP 2106.05(g)).
Claim 11 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim does not fall within at least one of the four categories of patent eligible subject matter.
Claim 11 recites “a computer readable medium”. The specification did not define what is considered as a computer-readable medium. When the specification is silent, the BRI of a CRM in view of the state of the art covers a signal per se. carrier wave signal as non-statutory subject matter. Therefore, claim 11 is directed to a non-statutory subject matter.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1 and 10-11 are rejected under 35 U.S.C. 103 as being unpatentable over Nishitani et al. (US 20210001857 A1, hereinafter Nishitani) in view of KANEMARU (US 20180356793 A1, hereinafter Kanemaru).
Regarding Claim 1, Nishitani teaches a machine learning device (Paragraph [0032] Aspects of the present disclosure improve the learning efficiency of deep reinforcement learning by using an embedded network for estimating a controlled vehicle speed.) comprising:
an acquisition unit configured to acquire observation information (Paragraph [0026] This image or traffic data can be acquired by using a road side unit (RSU), in-vehicle sensors, road map information, and/or vehicle-to-vehicle (V2V) communication (e.g., in connected vehicle environments). (Traffic data (observation information) is acquired through sensors))
including information on a speed of a control target point at a control target time; (Paragraph [0028] For example V2V communications use wireless signals to send information back and forth between other connected vehicles (e.g., location, speed, and/or direction). Conversely, V2I communications involve vehicle to infrastructure (e.g., road signs or traffic signals) communications, generally involving vehicle safety issues. Paragraph [0067] In this configuration, the traffic state estimation network 530 disentangles an estimated traffic state (e.g., auxiliary task output 532) from the features extracted from the input 506 of the vehicle behavior controller 310. For example, a current and/or future position, a speed and/or a density of vehicles can be applied to the estimated traffic state (e.g., 532). (The sensors acquire information of a targeted vehicle, including its current speed at a point in time))
an output unit configured to output control information including information on speed control of the control target point, the control information being determined in accordance with the observation information and a control policy; (Paragraph [0031] In aspects of the present disclosure, current and past input traffic data images (e.g., of a merging section) are used as input when selecting a target speed of a controlled vehicle. An impact of a control vehicle merging behavior on traffic flow is realized by setting an average speed of all surrounding vehicles after merging as a reward for deep reinforcement learning (RL) of a vehicle merging controller. Unfortunately, it is difficult to make a machine learning network perform appropriate feature extraction for output speed selection from the input traffic data images. Paragraph [0032] Aspects of the present disclosure improve the learning efficiency of deep reinforcement learning by using an embedded network for estimating a controlled vehicle speed. In this aspect of the present disclosure, the embedded network is configured to estimate dynamic traffic conditions (e.g., controlled vehicle speed) to a vehicle merging controller. The embedded network enables appropriate feature extraction for vehicle behavior control, which boosts a learning efficiency of the deep reinforcement learning. This system enables controlled vehicles to effectively merge onto the highway main-lane, while reducing the traffic impact on the highway main-lane and on-ramp. [0074] In the present reinforcement learning problem of FIG. 6, an agent 610 (e.g., speed controller) and the highway environment 400 interact at discrete time steps, in which the highway environment 400 may be formulated as a Markov Decision Process (MDP). The agent 610 observes state s ∈S and selects an action a ∈A according to its policy π. (The traffic controller uses observational data and reinforcement learning in order to select target speeds of the vehicle, the reinforcement learning system acting as a control policy by providing higher or lower rewards based on selected actions.))
a corrected reward determination unit configured to determine a corrected reward obtained by correcting a reward in accordance with a speed of the control target point derived from the observation information (Paragraph [0067] For example, a Q-value of each action (e.g., a weighted sum of the expected rewards of all future steps starting from the current state) can be applied to the estimated impact 524 by the impact calculator 520. Paragraph [0071] The pre-trained vehicle behavior controller 510 implements the impact calculator 520 of FIG. 5A using a Q-value calculation 560. In this example, an estimated Q-value 564 is provided by the Q-value estimation network 516. In addition, a Q-value of each action (e.g., a weighted sum of the expected rewards of all future steps starting from the current state) can be applied to a temporal answer Q-value 562 by the Q-value calculation 560. In this configuration, an error between the estimated Q-value 564 from the Q-value estimation network 516 and the temporal answer Q-value 562 is provided as feedback to the Q-value estimation network 516. Paragraph [0072] As further illustrated in FIG. 5B, the block for selecting the action with the highest Q-value 518 of the pre-trained vehicle behavior controller 510 selects the optimal action at the output 508. In this configuration, the controlled vehicle speed estimation network 570 includes the auxiliary task output 572 to provide the estimated controlled vehicle speed. In addition, the controlled vehicle speed calculation 580 provides the actual controlled vehicle speed 582. The actual controlled vehicle speed 582 and the temporal answer Q-value 562 are calculated from the images 502 and the traffic data 505 received from the state manager 504 at the input 506 of the pre-trained vehicle behavior controller 510. The error between the estimated Q-value 564 and the temporal answer Q-value 562, as well as the actual controlled vehicle speed 582 (e.g., actual speed of the ego vehicle) and the estimated controlled vehicle speed (e.g, 572) are fed back to update the Q-value estimation network 516 of the pre-trained vehicle behavior controller 510 to adjust the Q-values of each action to enable selection of the optimal action at output 508. (An error is determined between expected and actual Q-values (rewards). The error rate is then fed back into the system to determine an updated (corrected) Q-value based on an input of (observed) traffic data and a goal. The Q-value is analogous to a sum of rewards (paragraph [0067] and [0071]).))
the reward being higher as an error between a value of an evaluation parameter and a goal is smaller (Paragraph [0072] The actual controlled vehicle speed 582 and the temporal answer Q-value 562 are calculated from the images 502 and the traffic data 505 received from the state manager 504 at the input 506 of the pre-trained vehicle behavior controller 510. The error between the estimated Q-value 564 and the temporal answer Q-value 562, as well as the actual controlled vehicle speed 582 (e.g., actual speed of the ego vehicle) and the estimated controlled vehicle speed (e.g, 572) are fed back to update the Q-value estimation network 516 of the pre-trained vehicle behavior controller 510 to adjust the Q-values of each action to enable selection of the optimal action at output 508. (The reward (Q-value) is higher or lower based on how well the estimated speed of the vehicle corresponds to the actual vehicle speed, and then used to select an optimal action.))
the evaluation parameter being a parameter other than a speed derived from the observation information; and (Paragraph [0067] In this configuration, the traffic state estimation network 530 disentangles an estimated traffic state (e.g., auxiliary task output 532) from the features extracted from the input 506 of the vehicle behavior controller 310. For example, a current and/or future position, a speed and/or a density of vehicles can be applied to the estimated traffic state (e.g., 532). This enables the feature extraction network 314 to explicitly understand the state of the surrounding environment (including controlled vehicles, surrounding vehicles, and road information). Subsequently, the impact estimation network 316 disentangles the estimated impact 524 from the extracted features. For example, a Q-value of each action (e.g., a weighted sum of the expected rewards of all future steps starting from the current state) can be applied to the estimated impact 524 by the impact calculator 520. (Speed is not necessarily the only observed information. Other possible evaluation parameters are position or density of vehicles in traffic)).
Nishitani does not teach a learning unit configured to perform reinforcement learning of the control policy based on the observation information and the corrected reward
In the same field of endeavor, Kanemaru teaches: a learning unit configured to perform reinforcement learning of the control policy based on the observation information and the corrected reward (Paragraph [0133] In step S20, the learning unit 120 determines whether or not a condition for finishing the reinforcement learning has been satisfied. The reinforcement learning is to be finished on the condition that the foregoing processes have been repeated a predetermined number of times or repeated for a predetermined period of time, for example. (Paragraphs [0131]-[0133] teach that the learning unit is configured to perform reinforcement learning based on the information in the state information and the reward. Paragraphs [0083]-[0085] teach that the reward also includes the discount rate. Therefore, the reinforcement learning determines condition satisfaction based on the condition))
It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to have combined the concept of a learning unit configured to performing reinforcement learning of the control policy based on the observation information and the corrected reward as taught by Kanemaru into Nishitani as both references are in the same field of observing states of machines and performing actions in order to acquire rewards, and doing so would be desirable by allowing for increased capacity of machining tools or other devices to perform their functions due to the increased scale and higher speeds of new, modern requirements (Kanemaru Paragraph [0004])
Regarding Claims 10-11, Claims 10-11 are method and computer programming product claims which correspond to the machine claim of Claim 1. As such, they are rejected for the same reasons.
Claims 2-9 are rejected under 35 U.S.C. 103 as being unpatentable over Nishitani in view of Kanemaru in further view of SATOU (US 20250174009 A1, hereinafter Satou).
Regarding Claim 2, the combination of Nishitani and Kanemaru teaches all the limitations as claimed in Claim 1. However, the combination does not teach: the corrected reward determination unit configured to determine the corrected reward that is corrected so as to be lower as a speed of the control target point is higher.
In the same field of endeavor, Satou teaches: the corrected reward determination unit configured to determine the corrected reward that is corrected so as to be lower as a speed of the control target point is higher (Paragraph [0131] For example, if at least one of the position and posture of the workpiece W can be detected, the reward R is 100 points, and if neither the position nor posture of the workpiece W can be detected, reward R is 0 points. (The reward is lowered or raised based on failure or completion of a stated goal parameter, such as speed (paragraph [0049])))
It would have been obvious to have incorporated the concept of a corrected reward determination that corrects a reward to be lower as the speed of a target is higher as taught by Satou into the combination of Nishitani and Kanemaru as all three references are in the same field of controlling machines through observations of their current states and undertaking actions based on those states, and this combination would be desirable in order for the increased capacity of detecting and observing the actions of machining arms (Satou Paragraph [0003]-[0004]).
Regarding Claim 3, the combination of Nishitani, Kanemaru and Satou teaches all the limitations as claimed in Claim 1, including: a corrected discount rate determination unit configured to determine a corrected discount rate obtained by correcting a discount rate of the corrected reward in accordance with a speed of the control target point derived from the observation information, (Kanemaru Paragraph [0069] A controller for a machine tool is responsible for a control process to be performed in a fixed cycle or a fixed period of time for achieving real-time control over the position or speed of an axis. Paragraph [0083] To satisfy a desire to maximize a total reward to be given in the future, a goal is to eventually satisfy Q(s, a)=E[Σ(γ.sup.t)r.sub.t]. Here, E[ ] is an expected value, t is time, γ is a parameter called a discount rate described later Paragraph [0085] Conversely, if the best behavior value max.sub.a Q(s.sub.t+1, a) is smaller than the value Q(s.sub.t, a.sub.t), Q(s.sub.t, a.sub.t) is reduced. In other words, the value of a certain behavior in a certain state is approximated to a best behavior value determined by the same behavior in a subsequent state. A difference between these values is changed by a way of determining the discount rate γ and the reward r.sub.t+1. (Determination of a discount rate y in order to maximize a reward associated with a control target of speed of a machine axis))
wherein the learning unit is configured to perform reinforcement learning of the control policy based on the observation information, the corrected reward, and the corrected discount rate (Kanemaru Paragraph [0131] If the reward takes a positive value, a determination “positive value” is made in step S15. Then, the processing proceeds to step S16. In step S16, the positive value is output as the reward to the value function update unit 122. If the reward takes zero, a determination “zero” is made in step S15. Then, the processing proceeds to step S17. In step S17, zero is output as the reward to the value function update unit 122. If the reward takes a negative value, a determination “negative value” is made in step S15. Then, the processing proceeds to step S18. In step S18, the negative value is output as the reward to the value function update unit 122. If any of step S16, step S17, and step S18 is finished, the processing proceeds to step S19. Paragraph [0133] In step S20, the learning unit 120 determines whether or not a condition for finishing the reinforcement learning has been satisfied. (Performing reinforcement learning of the controlling policy based on the observed information, and the reward and discount rate, the reward containing the discount rate as seen in paragraph[0083]))
Regarding Claim 4, the combination of Nishitani, Kanemaru and Satou teaches all the limitations as claimed in Claim 3, including: wherein the corrected discount rate determination unit is configured to determine the corrected discount rate that is corrected such that a value of the discount rate is smaller as the speed of the control target point is higher. (Kanemaru Paragraph [0085] The foregoing [Formula 1] shows a method of updating a value Q(s.sub.t, a.sub.t) of the behavior a.sub.t in the state s.sub.t based on the reward r.sub.t+1 given in response to doing the behavior a.sub.t tentatively. This update formula shows that, if a best behavior value max.sub.a Q(s.sub.t+1, a) determined by the behavior a.sub.t in the subsequent state s.sub.t+1 becomes larger than a value Q(s.sub.t, a.sub.t) determined by the behavior at in the state s.sub.t, Q(s.sub.t, a.sub.t) is increased. Conversely, if the best behavior value max.sub.a Q(s.sub.t+1, a) is smaller than the value Q(s.sub.t, a.sub.t), Q(s.sub.t, a.sub.t) is reduced. In other words, the value of a certain behavior in a certain state is approximated to a best behavior value determined by the same behavior in a subsequent state. A difference between these values is changed by a way of determining the discount rate γ and the reward r.sub.t+1. Meanwhile, the basic mechanism is such that a best behavior value in a certain state is propagated to a behavior value in a state previous to the state of the best behavior value. (The discount rate is altered in response to a change in behavior of a certain state, such as speed))
Regarding Claim 5, the combination of Nishitani, Kanemaru and Satou teaches all the limitations as claimed in Claim 3, including: wherein the corrected discount rate determination unit is configured to determine the corrected discount rate obtained by correcting, in accordance with the speed of the control target point, the discount rate in accordance with an input discount rate for an input speed that has been input. (Kanemaru Paragraph [0058] The machining program is information indicating a machining program used as a program for controlling a machine tool for machining on a workpiece. Paragraph [0085] The foregoing [Formula 1] shows a method of updating a value Q(s.sub.t, a.sub.t) of the behavior a.sub.t in the state s.sub.t based on the reward r.sub.t+1 given in response to doing the behavior a.sub.t tentatively. This update formula shows that, if a best behavior value max.sub.a Q(s.sub.t+1, a) determined by the behavior a.sub.t in the subsequent state s.sub.t+1 becomes larger than a value Q(s.sub.t, a.sub.t) determined by the behavior at in the state s.sub.t, Q(s.sub.t, a.sub.t) is increased. Conversely, if the best behavior value max.sub.a Q(s.sub.t+1, a) is smaller than the value Q(s.sub.t, a.sub.t), Q(s.sub.t, a.sub.t) is reduced. In other words, the value of a certain behavior in a certain state is approximated to a best behavior value determined by the same behavior in a subsequent state. A difference between these values is changed by a way of determining the discount rate γ and the reward r.sub.t+1. Meanwhile, the basic mechanism is such that a best behavior value in a certain state is propagated to a behavior value in a state previous to the state of the best behavior value. (The discount rate is computed, determinant on a change in state such as a speed of a machining process and the difference of the reward, meaning the discount rate is determined in accordance with an input speed))
Regarding Claim 6, the combination of Nishitani and Kanemaru teaches all the limitations as claimed in Claim 1. However, the combination does not teach: further comprising a goal setting unit configured to set, as goals, a plurality of evaluation parameter values, based on a group consisting of: first experience data including a value of an evaluation parameter derived from the acquired observation information; and one or more pieces of second experience data including values of an evaluation parameter derived from one or more pieces of other observation information different in control target time from the acquired observation information, each of the plurality of evaluation parameter values selected from a plurality of evaluation parameter values included in the group, wherein the corrected reward determination unit is configured to determine, for the set goals, corrected rewards obtained by correcting rewards in accordance with the speed of the control target point included in the observation information, the rewards being higher as errors between the values of the evaluation parameter derived from the acquired observation information and the goals are smaller.
In the same field of endeavor, Satou teaches: further comprising a goal setting unit configured to set, as goals, a plurality of evaluation parameter values, based on a group consisting of: (Paragraph [0041] The teaching device 4 has an online teaching function such as a playback method or direct teaching method, which teaches the position and posture of the control target part by actually moving the machine 2, or an offline teaching function which teaches the position and posture of the control target part by moving a virtual model of the machine 2 in a computer-generated virtual space The teaching device 4 generates an operation program for the machine 2 by associating the taught position, posture, operation speed, etc., of the control target part with various operation commands. The operation commands include various commands such as linear movement, circular arcuate movement, and movement of each axis. (A target position (goal) is set, consisting of multiple parameters of evaluation of success such as speed, correct posture, etc.))
first experience data including a value of an evaluation parameter derived from the acquired observation information; and (Paragraph [0041] The teaching device 4 receives the state of the machine 2 from the controller 3 and displays the state of the machine 2 on a display or the like. Paragraph [0042] The controller 3 detects at least one of the position and posture of the workpiece W by obtaining an image in which the workpiece W is captured using the visual sensor 5, extracting a feature of the workpiece W from the image in which workpiece W is captured, and comparing the extracted feature of the workpiece W with a model feature of the workpiece W extracted from a model image in which the workpiece W, for which at least one of the position and posture thereof is known, is captured. (First data of an observation is acquired, being the state the machine is in, such as position and posture))
one or more pieces of second experience data including values of an evaluation parameter derived from one or more pieces of other observation information different in control target time from the acquired observation information, each of the plurality of evaluation parameter values selected from a plurality of evaluation parameter values included in the group, (Paragraph [0049] The control part 32 corrects at least one of the position and posture of a control target part of the machine 2 based on at least one of the detected position and posture of the workpiece W. For example, the control part 32 may correct the position and posture data of the control target part used in the operation program of the machine 2, or may provide visual feedback by calculating the position deviation, speed deviation, acceleration deviation, etc., of one or more electric motors based on inverse kinematics from the position and posture correction amounts of the control target part during operation of the machine 2. (After making a change in the position and posture of the target, the control part then observes the change in state and re-evaluates the effectiveness of the motion as a second point in time from the initial observation))
wherein the corrected reward determination unit is configured to determine, for the set goals, corrected rewards obtained by correcting rewards in accordance with the speed of the control target point included in the observation information, (Paragraph [0130] The learning part 52 searches for the optimal action A through trial and error so as to maximize the total future reward R, rather than the immediate reward R. Paragraph [0131] Furthermore, the reward R is a score obtained as a result of detecting at least one of the position and posture of the workpiece W by comparing the feature extraction image in a certain state S with the model feature extraction image. For example, if at least one of the position and posture of the workpiece W can be detected, the reward R is 100 points, and if neither the position nor posture of the workpiece W can be detected, reward R is 0 points. Paragraph [0132] When the learning part 52 executes a certain action A (setting the composite ratio for each predetermined section), the state S (state of the feature extraction image) in the object detection device 33 changes, and the learning data acquisition part 51 acquires the changed state S and its result as the reward R, and feeds the reward R back to the learning part 52. The learning part 52 searches for the optimal action A (optimum composite ratio setting for each predetermined section) through trial and error so as to maximize the total future reward R, rather than the immediate reward R. (The reward is determined, then fed back in order to maximize future values of a position, including deviations in speed (paragraph [0049])))
the rewards being higher as errors between the values of the evaluation parameter derived from the acquired observation information and the goals are smaller. (Paragraph [0131] Furthermore, the reward R is a score obtained as a result of detecting at least one of the position and posture of the workpiece W by comparing the feature extraction image in a certain state S with the model feature extraction image. For example, if at least one of the position and posture of the workpiece W can be detected, the reward R is 100 points, and if neither the position nor posture of the workpiece W can be detected, reward R is 0 points. (The reward is higher or lower as a result of accuracy (error rate) when reaching a position))
It would have been obvious to have incorporated the concepts of a goal unit configured to set, as goals, a plurality of evaluation parameter values based on a group of first experience data including a value of an evaluation parameter derived from the acquired observation information; and one or more pieces of second experience data including values of an evaluation parameter derived from one or more pieces of other observation information different in control target time from the acquired observation information, each of the plurality of evaluation parameter values selected from a plurality of evaluation parameter values included in the group, wherein the corrected reward determination unit is configured to determine, for the set goals, corrected rewards obtained by correcting rewards in accordance with the speed of the control target point included in the observation information, with the rewards being higher as errors between the values of the evaluation parameter derived from the acquired observation information and the goals are smaller as taught by Satou into the combination of Nishitani and Kanemaru as all three references are in the same field of controlling machines through observations of their current states and undertaking actions based on those states, and this combination would be desirable in order for the increased capacity of detecting and observing the actions of machining arms (Satou Paragraph [0003]-[0004]).
Regarding Claim 7, the combination of Nishitani, Kanemaru and Satou teaches all the limitations as claimed in Claim 6, including: wherein the goal setting unit is configured to set, as the goals, the value of the evaluation parameter included in the first experience data, and (Satou Paragraph [0050] As described above, the machining system 1 detects at least one of the position and posture of the workpiece W from the image in which the workpiece is captured W using the visual sensor 5, and controls the operations of the machine 2 based on at least one of the position and posture of the workpiece W. (A goal from an image (such as position) is detected and used as learning data in order to achieve the goal in the simulated or real machine piece))
a noise-added value of an evaluation parameter in which noise is added to a value of an evaluation parameter included in the second experience data. (Satou Paragraph [0050] Depending on the type, size, etc., of the filter F used by the feature extraction part 34, there may be locations where the reaction of the filter F is weak. By setting a low threshold in threshold-processing after filter processing, it is possible to extract contours from areas where reaction is weak, but unnecessary noise will also be extracted, increasing the time required for feature matching. Furthermore, the feature of workpiece W may not be extracted due to a slight change in imaging conditions. Paragraph [0059] The image reception part 36 adds one or more changes to the received model image, such as brightness, enlargement or reduction, shearing, translation, rotation, etc., and may receive one or more model images having changes added thereto. (One implementation allows the optional extraction and analysis of images with noise, such as positioning and posture. Another possibility remains where noise is added at the later step of the image reception part, as noise is a randomized variation of brightness and color information in an image, which the image reception part is capable of modifying.))
Regarding Claim 8, the combination of Nishitani, Kanemaru and Satou teaches all the limitations as claimed in Claim 6, including: wherein the goal setting unit is configured to set the goals in accordance with a goal selecting method selected by a user. (Kanemaru Paragraph [0046] In this embodiment, a machining program is executed based on a machining condition set by a user. In response to this, multiple control processes corresponding to a particular process or function in the machining program are selected. (In response to a user selecting a goal (machining condition), processes for achieving the outcome are selected))
Regarding Claim 9, the combination Nishitani, Kanemaru and Satou teaches all the limitations as claimed in Claim 6, including: wherein the goal setting unit is configured to set the goals, a number of which is selected by a user. (Kanemaru Paragraph [0046] In this embodiment, a machining program is executed based on a machining condition set by a user. (A user defines a goal for the program))
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
SINGH et al. (US 20240037193 A1) discusses reinforcement-based management of infrastructure control-state spaces.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JUSTIN A CARDOSO whose telephone number is (571)272-8512. The examiner can normally be reached M-F 7:30 - 5:00, alternate Friday's off.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached at (571) 272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JUSTIN CARDOSO/
Patent Examiner, Art Unit 2143
/JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143