DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Responsive to the communication dated 8/26/2026.
Claims 1 – 20 are presented for examination.
Priority
ADS dated 5/2/2023 claims domestic priority to provisional 63338313 dated 5/4/2022.
Information Disclosure Statement
IDS dated 9/25/2023, 1/29/2026, 4/15/2026 have been reviewed. Only items provided in English are considered. See the attached IDS.
Drawings
The drawings dated 5/2/2023 have been reviewed. They are accepted.
Specification
The abstract dated 5/2/2023 has 156 words, 11 lines and no legal phraseology. The abstract is objected to because it has more than the maximum allowed 150 words.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1 – 20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception without significantly more.
Claim 1.
STEP 1: YES. The claim recites “a method”.
STEP 2A PRONG ONE: YES. The claim recites:
“A method the method comprising: receiving a natural language input data that reflects a user request to automatically generate actions for performing a task in dependence on a corresponding state of a domain [i.e., hearing a conditional verbal request].
Generating a request embedding based on processing the natural language input data [i.e., thinking about where information about similar actions for performing a task in dependence on a similarly corresponding state of a domain might be found: thinking about analogous scenario while answering a question];
Performing a simulation [i.e., imagining: mental imagery as mental emulation], of the task, by implementing, in a simulated [i.e., imagined] environment that reflects the corresponding state of the domain, a predicted action set generated based on processing [i.e., thinking about] the request embedding [i.e., analogous scenario] using one or more trained action models [e.g., probability distributions, reinforcement policy];
Determining, based on the simulation [i.e., imagining: mental imagery as mental emulation], that the predicted action set is not suitable for performing the task [i.e., make a determination/judgement];
In response to determining that the predicted action set is not suitable for performing the task [i.e., make a determination/judgement]:
Generating [i.e., imagining/thinking or selecting] an alternative predicted action set for performing the task, wherein generating the alternate predicted set comprises: utilizing an alternate request embedding [i.e., analogous scenario] in generating the alternate predicted action set, and/or
Utilizing predicted action set, and/or utilizing at least one alternate trained action model [e.g., probability distributions, reinforcement policy] in generating the alternate predicted action set;
Determining that the alternate predicted action set is suitable for performing the task;
In response to determining that the alternate predicted action set is suitable for performing the task [i.e., making a determination/judgement];
which is a mental process of receiving a verbal request to perform an action corresponding to a scenario, thinking about actions of analogous scenarios, imagining performing the action of the analogous scenario based on probability or experience (spec par 15), making a judgment, based on thinking about the imagined performance, that an action is not suitable and thinking about an action of an alternative analogous scenario based on probability or experience and deciding that the alternative action is suitable.
This is the very mental process used by humans when deciding which course of action they will choose given a request to perform some activity.
Regarding the trained action model, the instant application discloses that an action model can be a probability distributions or reinforcement learning policy. See Par 15 of the instant application which states: “… for example, the action mode can be a model that is used to generate a probability distribution… for example, the action model can generate a probability distribution… as another example, the alternative action model can be an RL policy…”.
An example of a reinforcement learning (RL) policy in human learning is a child's strategy for deciding whether to touch a hot stove/pot based on past painful experiences and current visual cues. In reinforcement learning, a policy is the strategy or rule that a person uses to map a specific situation or environment (i.e., state of a domain) to the behavior they choose (i.e., action).
Hot Stove Example:
A child hears a request to get the pot from the stove.
child sees a glowing red burner under the pot on the stove.
The child imagines picking up the hot pot.
The behavioral rule (i.e., policy) the child follows: "If the stove burner is visible (i.e., red = hot), do not touch it or the pot with your hand (i.e., touching the stove or pot with your hand is an unsuitable action); if the stove burner is visible (i.e., red = hot) tell a parent"
touching a hot stove results in severe pain (a strong negative reward or punishment), while avoiding it keeps the child safe (neutral/positive continuation) and over time, the child updates their mental policy to assign a very high probability to a "keep hands away" action whenever they encounter the state of a kitchen stove.
Alternatively, the child imagines telling their parent that the stove is hot.
The child has learned that there is a high probability that if they tell the parent the stove is hot that the parent will praise them for not touching the hot stove and being safe. Accordingly, the child determines that a suitable alternative action is not to touch the stove and communicate to the parents the decision not to touch the stove.
Regarding the recited embedding, when a human recalls an analogous scenario to answer a question, their brain is performing a process that is functionally identical to the concept of embeddings. Mathematically, the concept of embeddings is a vector that points in a direction and similar things point is a similar direction. When a human is mentally thinking of an analogous scenario they are mentally evaluating the similarity of two things by ignoring the literal surface details and focusing on the underlying relational meaning (i.e., directionality) by compressing complex, real-world information into an abstract conceptual "space" where items with similar meanings are stored close together. Mentally, humans group concepts in their mind as demonstrated when humans use taxonomy/hierarchical classifications. For example, humans group car, truck, boat, airplane as vehicles (i.e., storing these items in the same conceptual space). Further, the human mind more closely associates cars and trucks because they are more similar vehicles when compared to boats and airplane because, for example, cars and trucks both operate on the land while boats and airplanes do not. The human mind takes aspects (i.e., dimensions) of the items (e.g., no. of tires, size, types of loads, operating domain (land, air, sea), fuel type (gas, diesel, jet fuel), etc.) and reduces those dimensions into a similarity/analogy. Therefore, a human’s mental ability to generate/identify similarities/analogies is the human mind generating an embedding. Indeed, mathematical embedding is the attempt to recreate a human’s abstract mental embedding capability by use of a further abstract mathematical calculation. Accordingly, embeddings are abstract and the human mind is capable of performing an embedding.
Take the further example of a verbal request to take a vehicle from Atlanta Georgia to Bermingham Alabama. A human, using the vehicle embedding and imagining driving a boat from Atlanta to Bermingham, would decide that that action of driving a boat is not suitable because there is a very low probability of driving a boat where there is no waterway. A human would then, using the vehicle embedding and imagining driving a car west along I-20, decide there is a high probability of success and conclude that the action of driving a car or truck is a suitable action.
Therefore, it is found that the claim recites the mental process of decision making.
STEP 2A PRONG TWO: NO.
While the claim recites the abstract idea is “implemented by one or more processors” merely executing an abstract idea on a generally recited “processor” is not a practical application as this merely invokes the processor as a tool for the performance of the abstract idea. See MPEP 2105.05(f) which states: “… claims that amount to nothing more than an instruction to apply the abstract idea using a generic computer do not render an abstract idea eligible…”.
While the claim recites: “… transmitting data to cause the alternate predicted action set to be implemented, in a real-world environment, to perform the task” is insufficient application. MPEP 2105.05(f) states: “… A claim having broad applicability across many fields of endeavor may not provide meaningful limitations that integrate a judicial exception into a practical application or amount to significantly more. For instance, a claim that generically recites an effect of the judicial exception or claims every mode of accomplishing that effect, amounts to a claim that is merely adding the words "apply it" to the judicial exception…” These elements claim any and all actions applied to any and all real-world environments to perform any and all tasks and accordingly are merely a recitation to “apply” the mental decision. MPEP 2105.05(f) further states: ”…in order to make a claim directed to a judicial exception patent-eligible, the additional element or combination of elements must do "‘more than simply stat[e] the [judicial exception] while adding the words ‘apply it’…’".
Further, MPEP 2106.05(g) indicate that insignificant extra-solution activity is not indicative of a practical application and provide examples of insignificant application that include cutting hair after first determining the hair style. This is similar to the instant claims because the instant claims recite to make a decision (i.e., determining a suitable action) and then recites to perform that action. In the
Therefore, it is found that the claim does not recite a practical application as there are no elements in the claim that rely upon or use the judicial exception in a meaningful way.
STEP 2B: NO
MPEP 2106.05, regarding STEP 2B, states: “… limitations that the courts have found not to be enough to qualify as “significantly more” when recited in a claim with a judicial exception include… adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer… simply appending well-understood, routine, and conventional activities previously known in the industry… adding insignificant extra solution activity to the judicial exception…”
In the instant claims, the claim recites to perform the abstract idea “by one or more processors.” This is simply an instruction to implement an abstract idea on a generically recited “processor.”
Further, the recitation of “transmitting data” is a generically recites step that is common and routine.
Additionally, reciting “to cause the alternate predicted action set to be implemented, in a real-world environment, to perform the task” is equivalent to adding the words “apply it.”
Therefore, the claim is not significantly more than the abstract idea itself.
Claim 20. The limitations of claim 20 are substantially the same as those of claim 1 and are rejected due to the same reasons as outlined above for claim 1.
Claim 2 recites: “Further comprising: generating the predicted action set based on processing the request embedding using the one or more trained action models” which is further part of the abstract idea itself. These elements simply evaluate (i.e., process) a similarity by one or more probabilities or policies. This is simply decision making which is a mental process. These elements do not recite additional elements that rely upon or use the decision and do not recite any elements that are significantly more that the abstract idea itself.
Claim 3 recites “Further comprising: generating the predicted action set based on processing the request embedding using the one or more trained action models, wherein generating the predicted action set occurs during performing the simulation and is further based on processing, using the one or more trained action models, simulated state data generated during performing the simulation” which is further part of the abstract idea itself. These elements simply recite to imagine actions and outcome based on probabilities or policies. An example might be imagining turning left across traffic which has a 53% probability of an accident while a right turn has about 6% probability of an accident and imagining taking three right hand turns to achieve a left direction results in 16.9% probability of an accident. Which is 6 accidents on first turn, 5.6 on second turn, and 5.3 on third turn = 16.9% leading to the mental decision to take 3 right hand turn rather than 1 left hand turn because it is below 53% and accordingly, much safer. These elements are purely abstract decision making.
Claim 4 recites ”Wherein determining that the alternate predicted actions are suitable for performing the task comprises: performing an additional simulation, of the task, by implementing the alternate predicted actions in the simulated environment; and determining, based on the additional simulation, that the alternate predicted actions are suitable for performing the task” which is further part of the abstract idea itself. These elements simply recite to imagine performing alternative actions and to make a judgement about the suitability of those alternative actions. Humans routinely imagine performing alternative actions and image the outcome of those actions and make decisions based on those imagining.
Claim 5 recites ”herein determining, based on the simulation, that the predicted actions are not suitable for performing the task, comprises: processing simulation data, form the simulation using the predicted action, to generate natural language output that describes the processed simulation data; generating a metric based on comparing the natural language output to the natural language input data; and determining, based on the metric failing to satisfy a threshold, that the predicted actions are not suitable for performing the task” however, this is further part of the abstract idea itself. Humans routinely use natural language to communicate their opinions and judgements after imagining and considering alternative actions. Take for example the situation of imagining right-hand turns vs. left-hand turns and mentally concluding that a left hand turn, being 53% likely to cause an accident while three left hand turns being 16.9% likely to cause an accident is an unsuitable action because it’s accident likelihood (53%) is above the other alternative (16.9%) and then speaking/explaining this decision to someone else. These elements doe not recite additional elements that rely upon or use the abstract idea nor are they significantly more than the abstract idea itself. They merely recite to perform judgements and to communicate those judgements.
Claim 6 recites “Wherein the simulation data comprises a final state, of the simulated environment, form the simulation using the predicted actions” which is further part of the abstract idea itself. These elements simply indicate to image an outcome of an action. These elements do not rely upon on use the abstract idea and are not significantly more than the abstract idea as they are part of the abstract idea itself.
Claim 7 recites “Wherein determining, based on the simulation, that the predicted actions are not suitable for performing the task, comprises: processing simulation data, from the simulation using the predicted action, to determine whether one or more domain or task specific rules are violated; in response to determining at least one of the one or more domain or task specific rules are violated: determining that the predicted actions are not suitable for performing the task” which are further part of the abstract idea itself. The elements simply recite a mental process of determining a violation of a condition. This is a judgement. These elements do not rely upon on use the abstract idea and are not significantly more than the abstract idea as they are part of the abstract idea itself.
Claim 8 Recites: “Wherein determining, based on the simulation, that the predicted actions are not suitable for performing the task, comprises: causing simulation data, form the simulation using the predicted action to be rendered at a client device via which the user request was received; receiving user interface input, provided at the client device, response to causing the simulation data to be rendered at the client device; determining, based on the user interface input, that the predicted actions are not suitable for performing the task.” While this claim includes “rendered at a client device via which the user request was received; receiving user interface input, provided at the client device, response to causing the simulation data to be rendered at the client device” this is merely a general high-level recitation of a standard computer interface. While these may be elements other than the abstract idea these are not significantly more than the abstract idea nor are they a practical application because computers commonly render output to a user interface by sending information back and fourth between, for example, clients/terminals and servers. MPEP 2106.05(g) indicates that data gathering and outputting is not indicative of a practical application nor significantly more.
Claim 9 Recites “Wherein the simulation data comprises a final state, of the simulated environment, form the simulation using the predicted actions” which merely indicates that the imagined scenario includes an imagined outcome. This is part of the mental abstract idea. These elements do not rely upon on use the abstract idea and are not significantly more than the abstract idea as they are part of the abstract idea itself.
Claim 10 Recites “Further comprising: determining, based on the user interface input being directed to a particular feature of the final state, a particular action, of the predicted actions, whose implementation in simulation resulted in the particular feature; wherein determining that the alternate predicted actions are suitable for performing the task comprises determining that the alternate predicted actions lack the particular action” merely recites to receive input via a user interface. MPEP 2106.05(g) indicates that data gathering and outputting is not indicative of a practical application nor significantly more. While these limitations recite that the input is “particular” does not result in a practical application nor does it make the limitations significantly more than the abstract idea. While a claim may be particular if it actually recite elements that limit the claim, however, merely characterizing a feature or a action as “particular” when the action and feature are unlimited is a characterization without distinction.
Claim 11 Recites “Wherein generating the alternate predicted action set comprises using the alternate request embedding in generating the alternate predicted action set” which is the abstract idea of thinking about similar actions. As discussed above, an embedding is a mental process. These elements do not rely upon on use the abstract idea and are not significantly more than the abstract idea as they are part of the abstract idea itself.
Claim 12 recites “Wherein generating the request embedding comprises: generating a natural language embedding based on processing the natural language input data using a language model; and generating the request embedding based on the natural language embedding; wherein generating the alternative request embedding comprises: generating alternate natural language input data by modifying and/or supplementing the natural language input data using one or more supplemental terms from a domain specific knowledge based for the task; generating an alterative natural language embedding based on processing the alternative natural language input data using a language model; and generating the alternate request embedding based on the alternate natural language embedding; wherein the one or more supplemental terms are not utilized in generating the request embedding” which merely recites the mental processes of using language including domain specific terms. A human mind is fully capable of using natural language and domain specific terms. Using language to generate a similarity/analogy (i.e., embedding) is not significantly more than the abstract idea and is not a practical application because humans think with language. This is known as an internal voice.
Claim 13 recites “wherein generating the request embedding comprises: generating a natural language embedding based on processing the natural language input data using a language model; and generating the request embedding based on the natural language embedding; wherein generating the alternate request embedding comprises: generating alternate natural language input data by modifying and/or supplementing the natural language input data using one or more supplemental terms from an external knowledge source that is not specific to the domain or to the task; generating an alternate natural language embedding based on processing the alternate natural language input data using a language model; and generating the alternate request embedding based on the alternate natural language embedding; and wherein the one or more supplemental terms are not utilized in generating the request embedding” which merely recites the mental processes of using language including domain specific terms. A human mind is fully capable of using natural language and domain specific terms. Using language to generate a similarity/analogy (i.e., embedding) is not significantly more than the abstract idea and is not a practical application because humans think with language. This is known as an internal voice.
Claim 14 recites “wherein the generating the alternate request embedding comprises: causing a clarification prompt to be rendered at a client device via which the user request was received; receiving user feedback that is provided in response to the clarification prompt and via one or more user interface inputs at the client device; and generating the alternate request embedding based on processing the user feedback; and wherein the user feedback is not utilized in generating the request embedding” which is merely data gathering. MPEP 2106.05(g) states that data gathering and outputting are insignificant extra solution activity that does not amount to a practical application or significantly more. MPEP 2106.05(g) provides example of insignificant data gather that includes:
i. Performing clinical tests on individuals to obtain input for an equation, In re Grams, 888 F.2d 835, 839-40; 12 USPQ2d 1824, 1827-28 (Fed. Cir. 1989);
ii. Testing a system for a response, the response being used to determine system malfunction, In re Meyers, 688 F.2d 789, 794; 215 USPQ 193, 196-97 (CCPA 1982);
iii. Presenting offers to potential customers and gathering statistics generated based on the testing about how potential customers responded to the offers; the statistics are then used to calculate an optimized price, OIP Technologies, 788 F.3d at 1363, 115 USPQ2d at 1092-93;
iv. Obtaining information about transactions using the Internet to verify credit card transactions, CyberSource v. Retail Decisions, Inc., 654 F.3d 1366, 1375, 99 USPQ2d 1690, 1694 (Fed. Cir. 2011);
v. Consulting and updating an activity log, Ultramercial, 772 F.3d at 715, 112 USPQ2d at 1754; and
vi. Determining the level of a biomarker in blood, Mayo, 566 U.S. at 79, 101 USPQ2d at 1968. See also PerkinElmer, Inc. v. Intema Ltd., 496 Fed. App'x 65, 73, 105 USPQ2d 1960, 1966 (Fed. Cir. 2012) (assessing or measuring data derived from an ultrasound scan, to be used in a diagnosis).
Claim 15 recites “wherein generating the request embedding comprises: generating the request embedding based on processing the natural language input data and processing first context data, without processing second context data; and wherein generating the alternate request embedding comprises: generating the request embedding based on processing the natural language input data and processing the second context data, and without processing the first context data” which is merely the notion of mentally ignoring context while thinking and making a judgement about an action. The human mind is fully capable of ignoring context information. Further the human mind is fully capable of selecting which contextual information is relevant while making decisions.
Claim 16 recites “wherein generating the request embedding comprises: generating the request embedding based on processing the natural language input data and processing first context data, without processing second context data; and wherein generating the alternate request embedding comprises: generating the request embedding based on processing the natural language input data and processing the second context data, and without processing the first context data” which is merely the notion of mentally ignoring context while thinking and making a judgement about an action. The human mind is fully capable of ignoring context information. Further the human mind is fully capable of selecting which contextual information is relevant while making decisions.
Claim 17 recites “Wherein he first context data represents a current state of the domain at a first level of abstraction and the second context data represents the current state of the domain at a second level of abstract” which Is merely describing types of context. This is not a practical application nor significantly more.
Claim 18 recites “determining, based on one or more probabilities, whether to utilize the alternate request embedding or to instead utilize the at least one alternate trained action model, in generating the alternate predicted action set, wherein the one or more probabilities are for the predicted action set and are generated based on processing the request embedding using the one or more trained action models” and this merely further recite elements of the abstract idea itself. A human mind is fully capable of processing natural language and making determinations based on an understanding of probable outcomes of an action. This is not a practical application nor is it significantly more.
Claim 19 recites “wherein transmitting the data to cause the alternate predicted action set to be implemented, in the real-world environment, to perform the task, comprises: transmitting the data to cause the alternate action automatically implemented in response to the user request, and automatically implemented without requiring any further user input after providing the user request.”
2106.05(a) i states that courts have indicated that mere automation of manual processes, such as using a generic computer to process [data]… by avoiding physical interaction is not, for example, an improvement to a computer.
MPEP 2106.05(e) states that meaningful limitations may be relevant and states Diamond v. Diehr provides an example of a claim that recited meaningful limitations beyond generally linking the use of the judicial exception to a particular technological environment. 450 U.S. 175, 209 USPQ 1 (1981). In Diehr, the claim was directed to the use of the Arrhenius equation (an abstract idea or law of nature) in an automated process for operating a rubber-molding press. 450 U.S. at 177-78, 209 USPQ at 4. The Court evaluated additional elements such as the steps of installing rubber in a press, closing the mold, constantly measuring the temperature in the mold, and automatically opening the press at the proper time, and found them to be meaningful because they sufficiently limited the use of the mathematical equation to the practical application of molding rubber products. 450 U.S. at 184, 187, 209 USPQ at 7, 8. In contrast, the claims in Alice Corp. v. CLS Bank International did not meaningfully limit the abstract idea of mitigating settlement risk. 573 U.S. 208, 110 USPQ2d 1976 (2014). In particular, the Court concluded that the additional elements such as the data processing system and communications controllers recited in the system claims did not meaningfully limit the abstract idea because they merely linked the use of the abstract idea to a particular technological environment (i.e., "implementation via computers") or were well-understood, routine, conventional activity recited at a high level of generality. 573 U.S. at 225-26, 110 USPQ2d at 1984-85
This instant claim, however, does not have additional elements that sufficiently limit the use of the abstract idea to any particular application. Rather, the claims simply recite to apply any action automatically. Therefore, the claim is simply saying and “apply it” automatically. Because the claim recites that abstract mental idea of choosing a suitable action, the claim is merely claiming to make a decision and then automatically apply the decision.
MPEP 2106.05(f) states “a claim having broad applicability across many fields of endeavor may not provide meaningful limitations that integrate a judicial exception into a practical application or amount to significantly more.” In the instant application, the claim broadly recites to automatically apply the mental process of choosing a course of action across any and all fields of endeavor. There is no particularity with regard to any field of applicability. The Applicant is therefore attempting to claim the using any decision to do anything.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1 – 6, 8 – 13, 15 - 20 are rejected under 35 U.S.C. 103 as being unpatentable over Ahn_2022 (Do as I can, Not as I say: grounding Language in Robotic Affordances, aug 16, 2022) in view of Chang_2010 (US 7,681,186 B2).
Claim 1. Ahn_2022 makes obvious “A method , the method comprising: (a) receiving natural language input data that reflects a user request to automatically generate actions for performing a task in dependence on a corresponding state of a domain (section 4, algorithm 1 : "A high level instruction i, state s0"); (b) generating a request embedding based on processing the natural language input data (section 4.1 "text embeddings are used as the input to the policy and value function that specify which skill should be performed"); (c) performing a simulation, of the task, by implementing, in a simulated environment that reflects the corresponding state of the domain, a predicted action set generated based on processing the request embedding using one or more trained action models (section 2 : "The goal of TD methods is to learn state or state-action value functions (Q-function) OTT (s, a), which represents the discounted sum of rewards when starting from state sand action a, followed by the actions produced by the policy TT"); (d) determining, based on the simulation, that the predicted action set is not suitable for performing the task (section 2 "to determine whether a given command is feasible from the given state. It is worth noting that in the sparse reward case, where the agent receives the reward of 1.0 at the end of the episode if it was successful and 0.0 otherwise, the value function trained via RL ... specifies whether a skill is possible in a given state"); in response to determining that the predicted action set is not suitable for performing the task: (e) generating an alternate predicted action set for performing the task (Algorithm 1 is carried out for all more than one possible skills (i.e actions), see line 4 to 9) wherein generating the alternate predicted action set comprises: (f) utilizing an alternate request embedding in generating the alternate predicted action set, and/or utilizing at least one alternate trained action model in generating the alternate predicted action set (Algorithm 1, line 4-9, each skill in the skill set is scored); (g) determining that the alternate predicted action set is suitable for performing the task (Algorithm 1, line 10, the highest scored skill is selected); in response to determining that the alternate predicted action set is suitable for performing the task: (h) transmitting data to cause the alternate predicted action set to be implemented, in a real-world environment, to perform the task (Algorithm 1, line 11 "Execute TTn(sn) in the environment, updating state sn+ 1 ").
Chang_2010 makes obvious “implemented by one or more processors” (FIG. 1 ; COL 3: “… well-known computing systems, environments, and/or configuration that may be suitable for use with the invention include, but are not limited to, personal computers, server computers, hand-held or laptop devices, multiprocessor systems, microprocessor-based systems… network PCs… mainframe computers… distributed computing environments…”).
Ahn_2022 and Chang_2010 are analogous art because they are from the same field of endeavor called natural language. Before the effective filing date, it would have been obvious to a person of ordinary skill in the art to combine Ahn_2022 and Chang_2010. The rationale for doing so would have been that Ahn_2022 teaches to use natural language models to encode semantic knowledge but doesn’t explicitly teach a computer processor the encodes semantic knowledge and Chang_2010 teaches a system and method for modeling semantics of natural language and explicitly teaches a computer processor capable of performing encoding of semantic knowledge from natural language. Therefore, it would have been obvious to combine Ahn_2022 and Chang_2010 for the benefit of having equipment to process and encode semantic knowlege to obtain the invention as specified in the claims.
Claim 20. The limitations of claim 20 are substantially the same as those of claim 1 and are rejected due to the same reasons as outlined above for claim 1.
Claim 2. Ahn_2022 makes obvious “Further comprising: generating the predicted action set based on processing the request embedding using the one or more trained action models” (page 5 “we utilize both BC and RP policy…”).
Claim 3. Ahn_2022 makes obvious “Further comprising: generating the predicted action set based on processing the request embedding using the one or more trained action models, wherein generating the predicted action set occurs during performing the simulation and is further based on processing, using the one or more trained action models, simulated state data generated during performing the simulation” (page 5: “… use the embeddings generated by passing in text descriptions… embeddings are used as the input to the policy … the specify which skill should be performed…”).
Claim 4. Ahn_2022 makes obvious “Wherein determining that the alternate predicted actions are suitable for performing the task comprises: performing an additional simulation, of the task, by implementing the alternate predicted actions in the simulated environment; and determining, based on the additional simulation, that the alternate predicted actions are suitable for performing the task” (FIG. 3/13 illustrates “find an apple” is suitable. Page 3: “… specify the task and the response structure which the model will emulate…”).
Claim 5. Ahn_2022 makes obvious “Wherein determining, based on the simulation, that the predicted actions are not suitable for performing the task, comprises: processing simulation data, form the simulation using the predicted action, to generate natural language output that describes the processed simulation data; generating a metric based on comparing the natural language output to the natural language input data; and determining, based on the metric failing to satisfy a threshold, that the predicted actions are not suitable for performing the task” ( appendix D affordance function, Appendix E).
Claim 6. Ahn_2022 makes obvious “Wherein the simulation data comprises a final state, of the simulated environment, form the simulation using the predicted actions” (page 31 Figure 14 illustrates step 4 = “Done” step 4 = “Done”; page 4 “… run again until a termination token (e.g., “done”)).
Claim 8. Ahn_2022 makes obvious “Wherein determining, based on the simulation, that the predicted actions are not suitable for performing the task, comprises: causing simulation data, form the simulation using the predicted action to be rendered at a client device via which the user request was received; receiving user interface input, provided at the client device, response to causing the simulation data to be rendered at the client device; determining, based on the user interface input, that the predicted actions are not suitable for performing the task” (page 5: “… robot simulator using RetinaGAN sim-to-real transfer… the performance of simulation policies by utilizing simulation demonstrations…” page 21: “… the raters to mark the episodes as unsafe (i.e., if the robot collided with the environment), undesirable (i.e., if the robot perturbed objects that were not relevant to the skill) or infeasible (i.e., if the skill cannot be done or is already accomplished). If any of these conditions are met, the episode is excluded from training.).
Claim 9. Ahn_2022 makes obvious “wherein the simulation data comprises a final state, of the simulated environment, from the simulation using the predicted actions” (page 31 Figure 14 illustrates step 4 = “Done” step 4 = “Done” page 4 “… run again until a termination token (e.g., “done”)).
Claim 10. Ahn_2022 makes obvious “Further comprising: determining, based on the user interface input being directed to a particular feature of the final state, a particular action, of the predicted actions, whose implementation in simulation resulted in the particular feature; wherein determining that the alternate predicted actions are suitable for performing the task comprises determining that the alternate predicted actions lack the particular action” (page 21: “… raters to mark the episodes as unsafe (i.e., if the robot collided with the environment), undesirable (i.e., if the robot perturbed objects that were not relevant to the skill) or infeasible (i.e., if the skill cannot be done or is already accomplished)…” EXAMINER NOTE: a collision is, for example, a particular features and if that features is is lacking (i.e., no collision) then that skill is suitable because it doesn’t cause a collision).
Claim 11. Ahn_2022 makes obvious “wherein generating the alternate predicted action set comprises using the alternate request embedding in generating the alternate predicted action set” (Fig. 3, 12, 13: the alternative predicted action set is the set of next actions).
Claim 12. Ahn_2022 makes obvious “The method of wherein generating the request embedding comprises: generating a natural language embedding based on processing the natural language input data using a language model; and generating the request embedding based on the natural language embedding; wherein generating the alternate request embedding comprises: generating alternate natural language input data by modifying and/or supplementing the natural language input data using one or more supplemental terms from a domain specific knowledge base for the task; generating an alternate natural language embedding based on processing the alternate natural language input data using a language model; and generating the alternate request embedding based on the alternate natural language embedding; wherein the one or more supplemental terms are not utilized in generating the request embedding” (Algorithm 1, Figure 3, 13)
Claim 13. Ahn_2022 makes obvious “The method of wherein generating the request embedding comprises: generating a natural language embedding based on processing the natural language input data using a language model; and generating the request embedding based on the natural language embedding; wherein generating the alternate request embedding comprises: generating alternate natural language input data by modifying and/or supplementing the natural language input data using one or more supplemental terms from an external knowledge source that is not specific to the domain or to the task; generating an alternate natural language embedding based on processing the alternate natural language input data using a language model; and generating the alternate request embedding based on the alternate natural language embedding; and wherein the one or more supplemental terms are not utilized in generating the request embedding” (Algorithm 1, Figure 3, 13)
Claim 15. Ahn_2022 makes obvious “wherein generating the alternate request embedding comprises: generating a context embedding based on processing context data; and generating the alternate request embedding further based on the context embedding; wherein the context data is not utilized in generating the request embedding” (Algorithm 1: affordance probability given language description of a skill in state Sn. NOTE: the first context is the current state Sn and he second context is the next state Sn+1 )
Claim 16. Ahn_2022 makes obvious “The method of wherein generating the request embedding comprises: generating the request embedding based on processing the natural language input data and processing first context data, without processing second context data; and wherein generating the alternate request embedding comprises: generating the request embedding based on processing the natural language input data and processing the second context data, and without processing the first context data” (Algorithm 1: affordance probability given language description of a skill in state Sn. NOTE: the first context is the current state Sn and he second context is the next state Sn+1 )
Claim 17. Ahn_2022 makes obvious “wherein the first context data represents a current state of the domain at a first level of abstraction and the second context data represents the current state of the domain at a second level of abstraction” (Algorithm 1: affordance probability given language description of a skill in state Sn. NOTE: the first context is the current state Sn and he second context is the next state Sn+1 )
Claim 18. Ahn_2022 makes obvious “further comprising: determining, based on one or more probabilities, whether to utilize the alternate request embedding or to instead utilize the at least one alternate trained action model, in generating the alternate predicted action set, wherein the one or more probabilities are for the predicted action set and are generated based on processing the request embedding using the one or more trained action models” (page 3 an affordance function p (C|sl) which indicates the probability of completing the skill… successfully from state s. Intuitively means “if I ask the robot to do… will it do it?... the probability that skill… with textual label… successfully completes if executed from state s…”).
Claim 19. Ahn_2022 makes obvious “wherein transmitting the data to cause the alternate predicted action set to be implemented, in the real-world environment, to perform the task, comprises: transmitting the data to cause the alternate action automatically implemented in response to the user request, and automatically implemented without requiring any further user input after providing the user request” (abstract: “… this approach is capable of completing long-horizon, abstract, natural language instructions on a mobile manipulator…”; Figure 1: “… LLMs via value functions of pretrained skills, allowing them to execute real-world, abstract, long-horizon commands on robots…”; page 6: “… with a mobile manipulator and a set of object manipulation and navigation skills in two office kitchen environments… a real office kitchen…”; Figure 5; NOTE: after receiving the natural language request the LLM automatically sends instructions to the manipulator that autonomously executes the series of actions without human intervention).
Claims 7 are rejected under 35 U.S.C. 103 as being unpatentable over Ahn_2022 (Do as I can, Not as I say: grounding Language in Robotic Affordances, aug 16, 2022) in view of Chang_2010 in view of Bai_2022 (Constitutional AI: Harmlessness from AI Feedback, Dec 15, 2022).
Claim 7. Bai_2022 makes obvious “Wherein determining, based on the simulation, that the predicted actions are not suitable for performing the task, comprises: processing simulation data, from the simulation using the predicted action, to determine whether one or more domain or task specific rules are violated; in response to determining at least one of the one or more domain or task specific rules are violated: determining that the predicted actions are not suitable for performing the task” (page 7 section 3: “… we refer to these instructions as ‘principles’ forming a ‘constitution’, i.e., as set of rules with which to steer the model…” EXAMINER NOTE: Bai_2022 outlined a two tier multi-stage model framework using natural language instructions with a “constitution” consisting of specific operational rules. First Tier (The Supervisor/Critic): A model acts as an evaluator/supervisor that takes human requests, analyzes the responses, and explicitly pulls out structural critiques. Second Tier (Task/Domain Rules): A second generation tier is forced to adhere strictly to targeted, task-specific, and domain-specific rules (the constitutional principles) to ground its responses safely and accurately).
Ahn_2022 and Bai_2022 are analogous art because they are from the same field of endeavor called learning models. Before the effective filing date, it would have been obvious to a person of ordinary skill in the art to combine Ahn_2022 and Bai_2022. The rationale for doing so would have been Ahn_2022 teaches an AI that processes natural language to perform tasks. Bai_2022 teaches that as AI systems become more capable it is important to ensure that these autonomous systems are helpful and don’t cause harm. To accomplish oversight without direct human oversight, Bai_2022 teaches to provide a through list of rules or principles. The purpose is to make it possible to control AI behavior more precisely and with far fewer human labels. Therefore, it would have been obvious to combine Ahn_2022 and Bai_2022 for the benefit of controlling AI more precisely to train the AI to follow rules that ensures it is harmless to obtain the invention as specified in the claims.
Claims 14 are rejected under 35 U.S.C. 103 as being unpatentable over Ahn_2022 (Do as I can, Not as I say: grounding Language in Robotic Affordances, aug 16, 2022) in view of Chang_2010 in view of Boukaram_2021 (Mitigating the Effects of Delayed Virtual Agent Response Time Using Conversational Fillers, HAI ’21, November 9 – 11, 2021, Nagoua, Japan).
Claim 14. Boukaram_2021 makes obvious “wherein generating the alternate request embedding comprises: causing a clarification prompt to be rendered at a client device via which the user request was received; receiving user feedback that is provided in response to the clarification prompt and via one or more user interface inputs at the client device; and generating the alternate request embedding based on processing the user feedback; and wherein the user feedback is not utilized in generating the request embedding” (abstract: “… fillers are used by the agent to keep the user engaged until the response is ready… conversational fillers… contextualized fillers that assume some semantic knowledge of the input and contain some of its elements… we ran task-based experiments… at a restaurant… contextualized fillers positively affected participants’ rating of the agent…” EXAMINER NOTE: “The paper evaluates a conversational pipeline that generates specific "holding expressions" or small talk designed to acknowledge the user's intent and fill time. Instead of showing an idle loading animation, the system uses natural dialogue structures to keep the user engaged and distract them from background processing latency. The paper details how conversational agents handle request fulfillment (such as executing long-running API tasks like retrieving a recipe or searching restaurant databases) asynchronously while the user interface immediately engages the user to pass the time. Natural language dialog “fillers” that seek to acknowledge the users intent while asynchronous tasks are being performed in the background do not utilize the feedback from the filler dialog prompts.)
Ahn_2022 and Boukaram_2021 are analogous art because they are from the same field of endeavor called interactions with natural language models. Before the effective filing date, it would have been obvious to a person of ordinary skill in the art to combine Ahn_2022 and Boukaram_2021.
The rationale for doing so would have been that Ahn_2022 teaches to use an LLM that interacts with humans at a restaurant. Boukaram_2021 teaches to use contextual conversational fillers to discuss requests an LLM receives from humans while the LLM is fulfilling the request to pass the time and that this positively affects users rating of the LLM. Therefore, it would have been obvious to combine Ahn_2022 and Boukaram_2021 for the benefit of passing the time and giving a better rating of the LLM to obtain the invention as specified in the claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRIAN S COOK whose telephone number is (571)272-4276. The examiner can normally be reached 8:00 AM - 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emerson Puente can be reached at 571-272-3652. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BRIAN S COOK/Primary Examiner, Art Unit 2187