DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claims 5, 9, 10, 11, 14, 15, 17, 19 are missing an “and” prior to the last limitation.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-4, 6-8, 10-13 and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chokjaroenwathanakul et a. (Pub 20260120029) (hereafter Chok) in view of Bavishi et al. (Pub 20250298495) (hereafter Bavishi).
As per claim 1, Chok teaches:
A computer-implemented method of generating workflow definitions in workflow definition language, the method employing a large language model, LLM, the method comprising:
receiving a natural language description of an automated workflow;
generating a plan generation prompt including the natural language description and plan generation instructions; ([Paragraph 3], More specifically, implementations of the present disclosure relate to computerized systems and techniques that improve user interactions with large language models (“LLMs”) through analysis, updating, supplementing, summarizing, and/or the like natural language prompts from users, as well as responses from the LLMs. [Paragraph 44], Prompt (or “LLM Prompt” or “Natural Language Prompt” or “Model Input”): a term, phrase, question, and/or statement written in a human language (e.g., English, Chinese, Spanish, and/or the like) that serves as a starting point for a language model and/or other language processing… [Paragraph 68], The prompt can include a natural language prompt. For example, the prompt can include instructions for instructing the LLM to generate workflow documentation that describes some or all of the workflow of the user, such as in using software application 180. The prompt can include various different workflow information, such a screen captures, event metadata, transcribed audio, intermediate documentation documents, and/or the like. For example, workflow data can include workflow events, such as events that trigger a screen capture of the app 180 by the documentation generation module 160. Event triggers can be associated with an important aspect of the workflow of the user and associated with relevant metadata.)
providing the plan generation prompt as input to the LLM, and receiving in response a structured plan comprising a first action and a second action; ([Paragraph 68], The prompt can include a natural language prompt. For example, the prompt can include instructions for instructing the LLM to generate workflow documentation that describes some or all of the workflow of the user, such as in using software application 180. The prompt can include various different workflow information, such a screen captures, event metadata, transcribed audio, intermediate documentation documents, and/or the like. For example, workflow data can include workflow events, such as events that trigger a screen capture of the app 180 by the documentation generation module 160. Event triggers can be associated with an important aspect of the workflow of the user and associated with relevant metadata. [Paragraph 50], An ontology may also include respective link types/definitions associated with data object types, which may include indications of how data object types may be related to one another. An ontology may also include respective actions associated with data object types or data object instances. The actions may include defined changes to values of properties based on various inputs. An ontology may also include respective functions, or indications of associated functions, associated with data object types, which functions may be executed when a data object of the associated type is accessed. An ontology may constitute a way to represent things in the world. An ontology may be used by an organization to model a view on what objects exist in the world, what their properties are, and how they are related to each other. An ontology may be user-defined, computer-defined, or some combination of the two. An ontology may include hierarchical relationships among data object types. [Paragraph 76], Tool information may indicate, for example, how data that may be accessed by the LLM (via tool calls) is structured, such as in an ontology or other format. [Paragraph 90], The documentation generation module 160 may be configured to access workflow data associated with interactions of the user with app 180 and/or actions performed by app 180 via an API call.)
Although Chok teaches workflow having actions.
Chok does not explicitly disclose generating a first segment of workflow definition language corresponding to the first action;
generating a second segment of workflow definition language corresponding to the second action; and
generating a workflow definition corresponding to the natural language description by combining the first segment of workflow definition language with the second segment of workflow definition language.
Bavishi teaches generating a first segment of workflow definition language corresponding to the first action;
generating a second segment of workflow definition language corresponding to the second action; and
generating a workflow definition corresponding to the natural language description by combining the first segment of workflow definition language with the second segment of workflow definition language. ([Paragraph 355], As indicated by block 302-1, in some examples, the task is segmented in a plurality of sub-tasks that form a workflow (e.g., interface workflow). As indicated by block 302-2, in some examples, one (e.g., a current) sub-task in the plurality of sub-tasks is a result of executing one or more preceding sub-tasks in the plurality of sub-tasks. [Paragraph 388], FIG. 57 is a pictorial illustration showing example inputs and outputs of the disclosed systems and methods. As illustrated (top), some current approaches utilize natural language inputs (prompts) and output commands. As further illustrated (bottom), the disclosed systems and methods are trained on and receive, as input, a state of the interface (e.g., screenshot), a workflow (e.g., functions, parameters, key-value pair, descriptions) and outputs actuation commands. [Paragraph 393],indicated by block 402-1, in some examples, the first workflow definition is a natural language description of the first workflow. As indicated by block 402-2, in some examples, the first workflow definition is a first tuple that translates the first workflow into a first set of functions and a first set of parameters. As indicated by block 402-3, in some examples, the first set of parameters are key-value pairs or include descriptions, or both. As indicated by block 402-4, in some examples, the first set of parameters include descriptions of the first set of functions. [Paragraph 389], FIG. 58 is a pictorial illustration showing one example execution loop of the disclosed systems and methods. As illustrated, a function planner (e.g., model(s)) receive, as input a state of the interface (e.g., screenshot), a workflow (e.g., functions, parameters, key-value pair, descriptions), and actions already taken (e.g., preceding tasks executed). The function planner outputs actuation commands (e.g., “click(‘login field’)”). An actuation model (e.g., actuator, actuation logic, etc.) executes machine-actuated actions based on the actuation commands and sends another screenshot of the interface to the function planner. The function planner outputs another actuation command (e.g., “type(username)”). This loop is repeated until, at some point, the function planner outputs a token (e.g., EOS token). )
It would have been obvious to a person with ordinary skill in the art, before the effective filing date of the invention, to combine the teachings of Chok wherein workflow definitions are generated by employing a large language model (LLM) based on receiving natural language description of an automated workflow, plan generation prompt is generated to prompt users which include natural language prompt, structured prompt comprising action(s) is/are received, into teachings of Bavishi wherein action(s) is/are task(s) which is/are segmented into segment(s) (i.e. sub-tasks) and the automated workflow comprises tasks/sub-tasks, because this would enhance the teachings of Chok wherein by segmenting the action(s)/task(s) into segments, it allows action(s) which has dependencies which include multiple sub-tasks to be completed/executed to be tracked via state of preceding action(s)/segment(s) and sequence of the workflow to be maintained. [Bavishi paragraph 289, 413, 425]
As per claim 2, rejection of claim 1 is incorporated:
Chok teaches deploying the workflow definition to a workflow executor;
displaying the automated workflow defined by the workflow definition in an editing user interface associated with the workflow executor.([Paragraph 62], In the example of FIG. 1, a user 150 (which generally refers to human user and/or a computing device of any type that may be operated by a human user) provides user input that causes the documentation generation module 160 to begin or end capture of a workflow session (e.g., capture of data and/or context associated with the interactions of the user with software application 180 and/or actions performed by software application 180). In some examples, the user 150 may provide user input indicative of instructions intended for an LLM to generate workflow documentation that describes the workflow of the user, such as in using software application 180. In some embodiments, the user 150 may provide user input to modify instructions generated by the documentation generation module 160 for the LLM to generate workflow documentation that describes the workflow of the user, such as in using software application 180. The instructions may be intended to instruct the LLM to generate workflow documentation that provides a step-by-step guide of the user's use of app 180, such as in performing the particular workflow. In some embodiments, the user 150 may speak while in a workflow session, which may be recorded by the documentation generation module 160. The documentation generation module 160 may receive audio input, such as audio signals indicative of words spoken by the user. In some embodiments, the user 150 may provide user input that causes the documentation generation module 160 to select and/or modify one or more of the workflow information (e.g., the workflow data and/or context) to be added to an LLM prompt.)
Bavishi also teaches ([Paragraph 378], As shown, the agent is operable to plan and execute a workflow. [Paragraph 269], Further, the extension interface provides a user interactable interface element (“Edit” button) that allows a user to edit the end-to-end workflow. [Paragraph 500], FIG. 119 is a pictorial illustration showing an example tool corresponding to the disclosed systems and methods. The illustrated example shows an example workflow editor that provides a user capability to author workflows.)
As per claim 3, rejection of claim 1 is incorporated:
Chok teaches wherein the natural language description is a second natural language description and the method further comprises:
receiving a first natural language description from a user;
generating a rephrasing prompt including the first natural language description and rephrasing instructions;
providing the rephrasing prompt as input to the LLM and receiving in response a second natural language description, the second natural language description providing a more detailed description of the automated workflow than the first natural language description. ([Paragraph 7], An improved artificial intelligence system (or simply “system”) achieves automated data monitoring and identification, gathering and collation of relevant data in combination with an LLM. In particular, the system facilitates generating LLM prompts that can increase the usefulness (e.g., accuracy, relevance, effectiveness, and/or the like) of LLM responses. In particular, the system discussed herein generates an LLM prompt (or prompts) that provide instructions to an LLM to respond with some or all of a workflow's documentation. For example, an LLM response can include a documentation document that accurately, completely, and clearly describes, in detail, a user's workflow. [Paragraph 9], The system can generate meaningful prompts (e.g., prompts that can induce useful LLM responses) based on at least some of the workflow information. For example, the system may transform at least some of the screen captures, associated metadata, audio transcript, and/or other context information into a meaningful prompt that is provided to the LLM such that the LLM can generate a useful response. In some embodiments, the system may augment user input with at least some of the workflow information, which can reduce the burden of prompt engineering on the user and increase the effectiveness of the prompt in inducing the LLM to generate a useful response such as accurate, complete, clear, and/or detailed workflow documentation. [Paragraph 68], The prompt can include a natural language prompt. For example, the prompt can include instructions for instructing the LLM to generate workflow documentation that describes some or all of the workflow of the user, such as in using software application 180. The prompt can include various different workflow information, such a screen captures, event metadata, transcribed audio, intermediate documentation documents, and/or the like. For example, workflow data can include workflow events, such as events that trigger a screen capture of the app 180 by the documentation generation module 160. Event triggers can be associated with an important aspect of the workflow of the user and associated with relevant metadata.)
Bavishi also teaches ([Paragraph 34], A system for constructing prompts that cause an agent to automate multimodal interface workflows includes agent specification logic and agent calling logic. The agent specification logic is configured to construct agent specifications using prompts and agent functions, wherein the agent specifications are configured to automate a multimodal interface workflow. The agent calling logic is in communication with the agent specification logic and is configured to translate the agent specifications into agent calls that cause an agent to implement the agent functions to produce outputs that are responsive to the prompts. [Paragraph 393], As indicated by block 402-1, in some examples, the first workflow definition is a natural language description of the first workflow. As indicated by block 402-2, in some examples, the first workflow definition is a first tuple that translates the first workflow into a first set of functions and a first set of parameters. As indicated by block 402-3, in some examples, the first set of parameters are key-value pairs or include descriptions, or both. As indicated by block 402-4, in some examples, the first set of parameters include descriptions of the first set of functions.)
As per claim 4, rejection of claim 3 is incorporated:
Bavishi teaches retrieving a rephrasing shot from a shot data store comprising an example first natural language description and a corresponding example second natural language description, based on a similarity of the example first natural language description to the received first natural language description; and including the rephrasing shot in the rephrasing prompt. ([Paragraph 258], FIG. 15 is a pictorial illustration showing one example of translating user intent into actions. As shown in FIG. 15, the systems and methods disclosed herein disclosed herein provide for generating and executing end-to-end workflows. FIG. 15 shows an example of planning (or generation an action plan, such as an end-to-end workflow). As shown in FIG. 15, a user has provided the prompt “Help me contact the venues listed for Happy Hour”. Though not shown in FIG. 15, the user also provides context information (the UI includes information) listing various venues. The systems and methods disclosed herein provide for translating the user intent (represented by the prompt “Help me contact the venues listed for Happy Hour”) into one or more actions. The systems and methods disclosed herein infer the user intent (model(s) infer that “The user needs help with Happy Hour planning.”). Further, the systems and methods disclosed translate the user intent into an action of planning or generating an action plan (end-to-end workflow) including one or more actions. As shown, a first action, of the one or more actions, of the action plan, comprises a search action. As shown, the systems and methods disclosed herein have generated an action plan that include searching for the venues provided in the context information (i.e., “I'll start by searching for venues listed on the screen.”). [Paragraph 265], As shown in FIG. 21, the browser screen is open to a document (or page) that includes various information corresponding to a Happy Hour event. The browser extension includes a prompt (“What can I help you with?”) that prompts a user to provide a prompt into the prompt input bar. As shown in FIG. 23, a user has begin interacting with the prompt input bar to provide a prompt. As shown in FIGS. 22 and 24, the user has provided the prompt “Help me contact these venues for Happy Hour.” In response, the disclosed systems and methods begin planning and executing an end-to-end work flow. As illustrated, the disclosed systems and methods infer the user's intent (“The user needs helps with Happy Hour planning”) corresponding to the user's prompt. The disclosed systems and methods then plan the work flow (“I'll start by searching for the venues listed on the screen” which results in the action of extracting information from the UI (“The venues I see on the screen are Casements, Left Door, Bar Iris, Arcana, Part Time.”). The disclosed systems and methods then plan the work flow (“I will now find the websites of these venues.”) which results in the actions of navigating to the website for each of the identified venues (i.e., multiple iterations of “Run ‘Happy Hour Search’ at stage ‘PRODUCTION’ with kwargs Object expression in the background”). [Paragraph 213], The reasoning behind this concept is that words with similar meanings occur in similar contexts. Different methods take the context of words into account. Some methods, like GloVe, base their context embedding on co-occurrence statistics from corpora (large texts) such as Wikipedia. Words with similar co-occurrence statistics have similar word embeddings. Other methods use neural networks to train the embeddings. For example, they train their embeddings to predict the word based on the context (Common Bag of Words), and/or to predict the context based on the word (Skip-Gram). Training these contextual embeddings is time intensive. For this reason, pre-trained libraries exist. Other deep learning methods can be used to create embeddings. For example, the latent space of a variational autoencoder (VAE) can be used as the embedding of the input. Another method is to use 1D convolutions to create embeddings. This causes a sparse, high-dimensional input space to be converted to a denser, low-dimensional feature space. [Paragraph 32], A system for generating training data to train agents to automate tasks otherwise done by users includes an intermediary disposed between an interface and a user. The intermediary is configured to: intercept one or more user-actuated actions directed towards the interface by the user, the user-actuated actions, if received by the interface, execute a task on the interface; preserve a state of the interface prior to the execution of the task; translate the user-actuated actions into one or more actuation commands, the actuation commands configured to trigger one or more machine-actuated actions that replicate the user-actuated actions on the interface to cause automation of the task; and generate a training dataset to train an agent to automate the task, wherein the training dataset requires the agent to process, as input, the state of the interface prior to the execution of the task, and to generate, as output, the actuation commands.)
Chok also teaches ([Paragraph 7], An improved artificial intelligence system (or simply “system”) achieves automated data monitoring and identification, gathering and collation of relevant data in combination with an LLM. In particular, the system facilitates generating LLM prompts that can increase the usefulness (e.g., accuracy, relevance, effectiveness, and/or the like) of LLM responses. In particular, the system discussed herein generates an LLM prompt (or prompts) that provide instructions to an LLM to respond with some or all of a workflow's documentation. For example, an LLM response can include a documentation document that accurately, completely, and clearly describes, in detail, a user's workflow. [Paragraph 9], The system can generate meaningful prompts (e.g., prompts that can induce useful LLM responses) based on at least some of the workflow information. For example, the system may transform at least some of the screen captures, associated metadata, audio transcript, and/or other context information into a meaningful prompt that is provided to the LLM such that the LLM can generate a useful response. In some embodiments, the system may augment user input with at least some of the workflow information, which can reduce the burden of prompt engineering on the user and increase the effectiveness of the prompt in inducing the LLM to generate a useful response such as accurate, complete, clear, and/or detailed workflow documentation. [Paragraph 68], The prompt can include a natural language prompt. For example, the prompt can include instructions for instructing the LLM to generate workflow documentation that describes some or all of the workflow of the user, such as in using software application 180. The prompt can include various different workflow information, such a screen captures, event metadata, transcribed audio, intermediate documentation documents, and/or the like. For example, workflow data can include workflow events, such as events that trigger a screen capture of the app 180 by the documentation generation module 160. Event triggers can be associated with an important aspect of the workflow of the user and associated with relevant metadata. [Paragraph 38], A language model may calculate the probability of different word combinations based on the patterns learned during training (based on a set of text data from books, articles, websites, audio files, and/or the like). [Paragraph 39], A Large Language Model (“LLM”) is any type of language model that has been trained on a larger data set and has a larger number of training parameters compared to a regular language model. An LLM can understand more intricate patterns and generate text that is more coherent and contextually relevant due to its extensive training. Thus, an LLM may perform well on a wide range of topics and tasks. LLMs may work by taking an input text and repeatedly predicting the next word or token (e.g., a portion of a word, a combination of one or more words or portions of words, punctuation, and/or any combination of the foregoing and/or the like).)
As per claim 6, rejection of claim 1 is incorporated:
Chok teaches retrieving a plan generation shot from a shot data store comprising an example natural language description and a corresponding example structured plan, based on a similarity of the example natural language description to the natural language description; and including the plan generation shot in the plan generation prompt. ([Paragraph 7], An improved artificial intelligence system (or simply “system”) achieves automated data monitoring and identification, gathering and collation of relevant data in combination with an LLM. In particular, the system facilitates generating LLM prompts that can increase the usefulness (e.g., accuracy, relevance, effectiveness, and/or the like) of LLM responses. In particular, the system discussed herein generates an LLM prompt (or prompts) that provide instructions to an LLM to respond with some or all of a workflow's documentation. For example, an LLM response can include a documentation document that accurately, completely, and clearly describes, in detail, a user's workflow. [Paragraph 9], The system can generate meaningful prompts (e.g., prompts that can induce useful LLM responses) based on at least some of the workflow information. For example, the system may transform at least some of the screen captures, associated metadata, audio transcript, and/or other context information into a meaningful prompt that is provided to the LLM such that the LLM can generate a useful response. In some embodiments, the system may augment user input with at least some of the workflow information, which can reduce the burden of prompt engineering on the user and increase the effectiveness of the prompt in inducing the LLM to generate a useful response such as accurate, complete, clear, and/or detailed workflow documentation. [Paragraph 68], The prompt can include a natural language prompt. For example, the prompt can include instructions for instructing the LLM to generate workflow documentation that describes some or all of the workflow of the user, such as in using software application 180. The prompt can include various different workflow information, such a screen captures, event metadata, transcribed audio, intermediate documentation documents, and/or the like. For example, workflow data can include workflow events, such as events that trigger a screen capture of the app 180 by the documentation generation module 160. Event triggers can be associated with an important aspect of the workflow of the user and associated with relevant metadata. [Paragraph 38], A language model may calculate the probability of different word combinations based on the patterns learned during training (based on a set of text data from books, articles, websites, audio files, and/or the like). [Paragraph 39], A Large Language Model (“LLM”) is any type of language model that has been trained on a larger data set and has a larger number of training parameters compared to a regular language model. An LLM can understand more intricate patterns and generate text that is more coherent and contextually relevant due to its extensive training. Thus, an LLM may perform well on a wide range of topics and tasks. LLMs may work by taking an input text and repeatedly predicting the next word or token (e.g., a portion of a word, a combination of one or more words or portions of words, punctuation, and/or any combination of the foregoing and/or the like).)
As per claim 7, rejection of claim 1 is incorporated:
Chok teaches displaying the structured plan on a user interface; receiving user input comprising comments in relation to the structured plan; generating a corrective plan generation prompt comprising the comments, the structured plan and corrective plan instructions; and providing the corrective plan generation prompt as input to the LLM and receiving in response a corrected structured plan based on user the comments. ([Paragraph 7], An improved artificial intelligence system (or simply “system”) achieves automated data monitoring and identification, gathering and collation of relevant data in combination with an LLM. In particular, the system facilitates generating LLM prompts that can increase the usefulness (e.g., accuracy, relevance, effectiveness, and/or the like) of LLM responses. In particular, the system discussed herein generates an LLM prompt (or prompts) that provide instructions to an LLM to respond with some or all of a workflow's documentation. For example, an LLM response can include a documentation document that accurately, completely, and clearly describes, in detail, a user's workflow. [Paragraph 9], The system can generate meaningful prompts (e.g., prompts that can induce useful LLM responses) based on at least some of the workflow information. For example, the system may transform at least some of the screen captures, associated metadata, audio transcript, and/or other context information into a meaningful prompt that is provided to the LLM such that the LLM can generate a useful response. In some embodiments, the system may augment user input with at least some of the workflow information, which can reduce the burden of prompt engineering on the user and increase the effectiveness of the prompt in inducing the LLM to generate a useful response such as accurate, complete, clear, and/or detailed workflow documentation. [Paragraph 68], The prompt can include a natural language prompt. For example, the prompt can include instructions for instructing the LLM to generate workflow documentation that describes some or all of the workflow of the user, such as in using software application 180. The prompt can include various different workflow information, such a screen captures, event metadata, transcribed audio, intermediate documentation documents, and/or the like. For example, workflow data can include workflow events, such as events that trigger a screen capture of the app 180 by the documentation generation module 160. Event triggers can be associated with an important aspect of the workflow of the user and associated with relevant metadata. [Paragraph 38], A language model may calculate the probability of different word combinations based on the patterns learned during training (based on a set of text data from books, articles, websites, audio files, and/or the like). [Paragraph 39], A Large Language Model (“LLM”) is any type of language model that has been trained on a larger data set and has a larger number of training parameters compared to a regular language model. An LLM can understand more intricate patterns and generate text that is more coherent and contextually relevant due to its extensive training. Thus, an LLM may perform well on a wide range of topics and tasks. LLMs may work by taking an input text and repeatedly predicting the next word or token (e.g., a portion of a word, a combination of one or more words or portions of words, punctuation, and/or any combination of the foregoing and/or the like). [Paragraph 66], As noted above, the user input can include starting or stopping a workflow capture session, modification of workflow information (e.g., editing and/or updating workflow information intended for an LLM), modification of instructions instructing the LLM to generate workflow documentation, and/or other information.)
Bavishi also teaches ([Paragraph 269], As shown in FIG. 30, the disclosed systems and methods further provide for surfacing a verification (or feedback) interface with elements prompting user response (“Does this look good to you?”) and elements providing for user response (“Yes” and “No” buttons). Further, as can be seen in at least FIG. 30, the extension interface provides a user interactable interface element (“Stop” button) that allows a user to stop the end-to-end workflow generation and execution. Further, the extension interface provides a user interactable interface element (“Edit” button) that allows a user to edit the end-to-end workflow.)
As per claim 8, rejection of claim 1 is incorporated:
Chock teaches wherein generating each segment of workflow definition language comprises: generating a workflow definition prompt including the respective action and workflow definition generation instructions; and providing the workflow definition prompt as input to the LLM and receiving in response the respective segment of workflow definition language. ([Paragraph 7], An improved artificial intelligence system (or simply “system”) achieves automated data monitoring and identification, gathering and collation of relevant data in combination with an LLM. In particular, the system facilitates generating LLM prompts that can increase the usefulness (e.g., accuracy, relevance, effectiveness, and/or the like) of LLM responses. In particular, the system discussed herein generates an LLM prompt (or prompts) that provide instructions to an LLM to respond with some or all of a workflow's documentation. For example, an LLM response can include a documentation document that accurately, completely, and clearly describes, in detail, a user's workflow. [Paragraph 9], The system can generate meaningful prompts (e.g., prompts that can induce useful LLM responses) based on at least some of the workflow information. For example, the system may transform at least some of the screen captures, associated metadata, audio transcript, and/or other context information into a meaningful prompt that is provided to the LLM such that the LLM can generate a useful response. In some embodiments, the system may augment user input with at least some of the workflow information, which can reduce the burden of prompt engineering on the user and increase the effectiveness of the prompt in inducing the LLM to generate a useful response such as accurate, complete, clear, and/or detailed workflow documentation. [Paragraph 68], The prompt can include a natural language prompt. For example, the prompt can include instructions for instructing the LLM to generate workflow documentation that describes some or all of the workflow of the user, such as in using software application 180. The prompt can include various different workflow information, such a screen captures, event metadata, transcribed audio, intermediate documentation documents, and/or the like. For example, workflow data can include workflow events, such as events that trigger a screen capture of the app 180 by the documentation generation module 160. Event triggers can be associated with an important aspect of the workflow of the user and associated with relevant metadata. [Paragraph 38], A language model may calculate the probability of different word combinations based on the patterns learned during training (based on a set of text data from books, articles, websites, audio files, and/or the like). [Paragraph 39], A Large Language Model (“LLM”) is any type of language model that has been trained on a larger data set and has a larger number of training parameters compared to a regular language model. An LLM can understand more intricate patterns and generate text that is more coherent and contextually relevant due to its extensive training. Thus, an LLM may perform well on a wide range of topics and tasks. LLMs may work by taking an input text and repeatedly predicting the next word or token (e.g., a portion of a word, a combination of one or more words or portions of words, punctuation, and/or any combination of the foregoing and/or the like)
As per claim 10, rejection of claim 8 is incorporated:
Bavishi teaches validating the respective segment of workflow definition language; in response to the respective segment of workflow definition language failing validation, receiving an error message indicative of reasons for the respective segment failing the validation; generating a segment correction prompt including the error message, the respective segment and validation correction instructions; providing the segment correction prompt to the LLM and receiving in response a corrected segment of workflow definition language. ([Paragraph 324], At block 1004, output evaluation logic makes the output available to the annotator for review and receives approval or disapproval from the annotator on the output. [Paragraph 327], At block 1010, output revision logic causes the agent to generate a revised output in response to determining that the annotator has disapproved the output and receiving corrective instructions from the annotator, makes the revised output available to the annotator for review, and receives approval or disapproval from the annotator on the revised output. [Paragraph 348], FIG. 162 is a block diagram showing an example system 2800 corresponding to the disclosed systems and methods. The system 2800, in one example, can be used to perform the method described in FIG. 146. The system 2800 is operable to effectively collect on-policy feedback for ongoing agent fine-tuning. As shown, system 2800 includes prompt processing logic 2802, prompt 2804, annotator 2806, agent 2808, output 2810, output evaluation logic 2812, training data construction logic 2814, training data 2816, run continuation logic 2818, subsequent output 2820, output revision logic 2822, revised output 2824, corrective instructions 2826, and can include various other items and functionality 2899. [Paragraph 353], Output revision logic 2822 is configured to cause the agent 2808 to generate a revised output 2824 in response to determining that the annotator 2806 has disapproved the output 2804 and receiving corrective instructions 2826 from the annotator. Output revision logic 2822 is configured to make the revised output 2824 available to the annotator 2806 for review and to receive approval or disapproval from the annotator 2806 on the revised output 2824. [Paragraph 362], FIG. 45 is a pictorial illustration illustrating one example of the operation of the recorder of the disclosed systems and methods. As illustrated in FIG. 45, the recorder provides for a user to instruct (e.g., prompt), oversee, and intervene on the planning and execution of the workflow and to provide feedback when the systems and methods provide incorrect workflow. In some examples, the recorder provides interface elements (e.g., text) describing each step of the workflow generated by the disclosed systems and methods and interface elements allowing a user to approve or deny the step or to provide input (e.g., a hint) to do something differently.)
As per claim 11, rejection of claim 1 is incorporated:
Chok teaches detecting a trigger action of the workflow from the natural language description; including, in the plan generation prompt, the trigger action. ([Paragraph 68], The prompt can include a natural language prompt. For example, the prompt can include instructions for instructing the LLM to generate workflow documentation that describes some or all of the workflow of the user, such as in using software application 180. The prompt can include various different workflow information, such a screen captures, event metadata, transcribed audio, intermediate documentation documents, and/or the like. For example, workflow data can include workflow events, such as events that trigger a screen capture of the app 180 by the documentation generation module 160. Event triggers can be associated with an important aspect of the workflow of the user and associated with relevant metadata. [Paragraph 25], An improved artificial intelligence system (“AIS” or simply “system”) facilitates automated generation of workflow documentation using an LLM. In response to a user initiating monitoring of a software application during performance of a workflow and the system providing workflow information obtained while monitoring the software application during performance of a workflow, an LLM can provide a response including some or all of a workflow's documentation. A workflow can include a user's use of a software application, such as in performing a task. The system, based on user input indicating initiation of monitoring, may monitor a user's use of the software application (e.g., user interactions with the software application and/or actions performed by the software application) and capture workflow information, such as workflow information that may be associated with an important aspect of the workflow. In this context, “workflow information” generally refers to data and/or context associated with interactions of a user with a software application and/or actions performed by the software application (e.g., workflow data and/or context). The system can capture selective and meaningful portions of the user's workflow, such as in response to detection of event triggers as the user performs the workflow in the software application. For example, the system can proctor/filter event capture during monitoring as the user can select which types of events (e.g., event triggers) and/or metadata is to be captured by the system.)
Bavishi also teaches ([Paragraph 32], A system for generating training data to train agents to automate tasks otherwise done by users includes an intermediary disposed between an interface and a user. The intermediary is configured to: intercept one or more user-actuated actions directed towards the interface by the user, the user-actuated actions, if received by the interface, execute a task on the interface; preserve a state of the interface prior to the execution of the task; translate the user-actuated actions into one or more actuation commands, the actuation commands configured to trigger one or more machine-actuated actions that replicate the user-actuated actions on the interface to cause automation of the task; and generate a training dataset to train an agent to automate the task, wherein the training dataset requires the agent to process, as input, the state of the interface prior to the execution of the task, and to generate, as output, the actuation commands.)
As per claim 12, rejection of claim 11 is incorporated:
Chok teaches wherein detecting the trigger action comprises: generating a trigger detection prompt comprising the natural language description and trigger detection instructions; and providing the trigger detection prompt as input to the LLM and receiving in response the trigger action. ([Paragraph 68], The prompt can include a natural language prompt. For example, the prompt can include instructions for instructing the LLM to generate workflow documentation that describes some or all of the workflow of the user, such as in using software application 180. The prompt can include various different workflow information, such a screen captures, event metadata, transcribed audio, intermediate documentation documents, and/or the like. For example, workflow data can include workflow events, such as events that trigger a screen capture of the app 180 by the documentation generation module 160. Event triggers can be associated with an important aspect of the workflow of the user and associated with relevant metadata. [Paragraph 67], Next, at interaction 2 the documentation generation module 160 generates a prompt based on at least the user input. The prompt can include the user input and/or may be generated based on other context, such as may be accessed by the context module 110 and/or that originating with app 180. In some embodiments, as further described herein, responsive to user input indicating initiation of monitoring, the documentation generation module 160 may capture workflow information (e.g., workflow data and/or context) associated with a user's workflow, such as based on selective and meaningful event-based monitoring of the workflow of the user, such as in using software application 180. The AIS may capture workflow information associated with detected event triggers (e.g., particular user interactions with software application 180). The documentation generation module 160 may capture workflow data and/or context to potentially add to an LLM prompt. The documentation generation module 160 may add at least some of the workflow information to the LLM prompt. For example, the documentation generation module 160 may add relevant workflow data and/or context to the LLM prompt. In some embodiments, the prompt can include instructions intended for an LLM, such as instructions intended to instruct the LLM to generate workflow documentation that describes the workflow of the user, such as in using app 180.)
Bavishi also teaches ([Paragraph 32], A system for generating training data to train agents to automate tasks otherwise done by users includes an intermediary disposed between an interface and a user. The intermediary is configured to: intercept one or more user-actuated actions directed towards the interface by the user, the user-actuated actions, if received by the interface, execute a task on the interface; preserve a state of the interface prior to the execution of the task; translate the user-actuated actions into one or more actuation commands, the actuation commands configured to trigger one or more machine-actuated actions that replicate the user-actuated actions on the interface to cause automation of the task; and generate a training dataset to train an agent to automate the task, wherein the training dataset requires the agent to process, as input, the state of the interface prior to the execution of the task, and to generate, as output, the actuation commands.)
As per claim 13, rejection of claim 1 is incorporated:
Chok teaches wherein the automated workflow is a security workflow forming part of a security playbook. ([Paragraph 24], Furthermore, automated workflow capture can raise privacy and/or security concerns (especially in sensitive environments such as healthcare, finance, the legal field, and/or the like) where content that is not meant to be captured may indeed be captured by the system, which can discourage use of such systems altogether. [Paragraph 35], In some embodiments, as further described herein, the artificial intelligence system may be configured to blur private, sensitive, and/or confidential information that may otherwise appear in a generated screen capture. This can permit the automated workflow capture to comply with various data privacy and security requirements.)
As per claims 18-20, these are device claims corresponding to the method claims 1, 3 and 11. Therefore, rejected based on similar rationale.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 14-17 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Chok.
As per claim 14, Chok teaches:
A computer-implemented method of generating a shot data store for use in generation of automated workflows in workflow definition language from natural language, comprising: ([Paragraph 3], More specifically, implementations of the present disclosure relate to computerized systems and techniques that improve user interactions with large language models (“LLMs”) through analysis, updating, supplementing, summarizing, and/or the like natural language prompts from users, as well as responses from the LLMs.)
receiving a workflow definition in a workflow definition language and a user description of the workflow definition; ([Paragraph 50], Ontology: stored information that provides a data model for storage of data in one or more databases and/or other data stores. For example, the stored data may include definitions for data object types and respective associated property types. An ontology may also include respective link types/definitions associated with data object types, which may include indications of how data object types may be related to one another. An ontology may also include respective actions associated with data object types or data object instances. The actions may include defined changes to values of properties based on various inputs. An ontology may also include respective functions, or indications of associated functions, associated with data object types, which functions may be executed when a data object of the associated type is accessed. An ontology may constitute a way to represent things in the world. An ontology may be used by an organization to model a view on what objects exist in the world, what their properties are, and how they are related to each other. An ontology may be user-defined, computer-defined, or some combination of the two. An ontology may include hierarchical relationships among data object types. [Paragraph 58], A documentation generation module 160 is configured to generate a prompt to a language model, such as LLM 130. As described in further detail below, the documentation generation module 160 may generate a prompt based on data provided by the user interface module 104 (e.g., a user input, tool information, and/or the like), the context module 110 (e.g., conversation history and/or other contextual information), and/or a software application such as app 180 (e.g., user workflow and corresponding workflow information).
generating, from the workflow definition, a structured plan comprising a plurality of actions corresponding to the workflow definition;
generating, from the structured plan, a natural language description of the workflow definition; ([Paragraph 3], More specifically, implementations of the present disclosure relate to computerized systems and techniques that improve user interactions with large language models (“LLMs”) through analysis, updating, supplementing, summarizing, and/or the like natural language prompts from users, as well as responses from the LLMs. [Paragraph 44], Prompt (or “LLM Prompt” or “Natural Language Prompt” or “Model Input”): a term, phrase, question, and/or statement written in a human language (e.g., English, Chinese, Spanish, and/or the like) that serves as a starting point for a language model and/or other language processing… [Paragraph 68], The prompt can include a natural language prompt. For example, the prompt can include instructions for instructing the LLM to generate workflow documentation that describes some or all of the workflow of the user, such as in using software application 180. The prompt can include various different workflow information, such a screen captures, event metadata, transcribed audio, intermediate documentation documents, and/or the like. For example, workflow data can include workflow events, such as events that trigger a screen capture of the app 180 by the documentation generation module 160. Event triggers can be associated with an important aspect of the workflow of the user and associated with relevant metadata. [Paragraph 76], Tool information may indicate, for example, how data that may be accessed by the LLM (via tool calls) is structured, such as in an ontology or other format. [Paragraph 90], The documentation generation module 160 may be configured to access workflow data associated with interactions of the user with app 180 and/or actions performed by app 180 via an API call. )
storing the workflow definition, the structured plan, the natural language description and the user description in a shot data store. ([Paragraph 50], Ontology: stored information that provides a data model for storage of data in one or more databases and/or other data stores. For example, the stored data may include definitions for data object types and respective associated property types. An ontology may also include respective link types/definitions associated with data object types, which may include indications of how data object types may be related to one another. An ontology may also include respective actions associated with data object types or data object instances. The actions may include defined changes to values of properties based on various inputs. An ontology may also include respective functions, or indications of associated functions, associated with data object types, which functions may be executed when a data object of the associated type is accessed. An ontology may constitute a way to represent things in the world. An ontology may be used by an organization to model a view on what objects exist in the world, what their properties are, and how they are related to each other. An ontology may be user-defined, computer-defined, or some combination of the two. An ontology may include hierarchical relationships among data object types. [Paragraph 58], A documentation generation module 160 is configured to generate a prompt to a language model, such as LLM 130. As described in further detail below, the documentation generation module 160 may generate a prompt based on data provided by the user interface module 104 (e.g., a user input, tool information, and/or the like), the context module 110 (e.g., conversation history and/or other contextual information), and/or a software application such as app 180 (e.g., user workflow and corresponding workflow information).
As per claim 15, rejection of claim 14 is incorporated:
Chok teaches generating a plan shot generation prompt including the workflow definition and plan generation instructions; providing the plan shot generation prompt to the LLM and receiving the structured plan in response. ([Paragraph 50], Ontology: stored information that provides a data model for storage of data in one or more databases and/or other data stores. For example, the stored data may include definitions for data object types and respective associated property types. An ontology may also include respective link types/definitions associated with data object types, which may include indications of how data object types may be related to one another. An ontology may also include respective actions associated with data object types or data object instances. The actions may include defined changes to values of properties based on various inputs. An ontology may also include respective functions, or indications of associated functions, associated with data object types, which functions may be executed when a data object of the associated type is accessed. An ontology may constitute a way to represent things in the world. An ontology may be used by an organization to model a view on what objects exist in the world, what their properties are, and how they are related to each other. An ontology may be user-defined, computer-defined, or some combination of the two. An ontology may include hierarchical relationships among data object types. [Paragraph 58], A documentation generation module 160 is configured to generate a prompt to a language model, such as LLM 130. As described in further detail below, the documentation generation module 160 may generate a prompt based on data provided by the user interface module 104 (e.g., a user input, tool information, and/or the like), the context module 110 (e.g., conversation history and/or other contextual information), and/or a software application such as app 180 (e.g., user workflow and corresponding workflow information). [Paragraph 76], The prompt can include information associated with one or more tools selected by the user, such as in the form of tool information, which enables the LLM 130 to generate a tool call that can be used by the AIS to communicate with a data processing service. Tool information may indicate, for example, how data that may be accessed by the LLM (via tool calls) is structured, such as in an ontology or other format. Tool information can indicate properties associated with a particular object type, such as an object type selected by the user in the user input at interaction 1. Tool information can include instructions for implementing a tool, instructions for generating a tool call, including instructions for formatting a tool call, tool implementation examples for executing one or more tool operations, and/or other information that may allow the LLM to provide more meaningful responses to the AIS. Tool implementation examples included in an LLM prompt can include pre-defined examples (e.g., the same for each use of the tool), user-selected or user-generated examples, and/or examples that are dynamically configured by the AIS 102 based on context.)
As per claim 16, rejection of claim 15 is incorporated:
Chok teaches generating respective embeddings of the workflow definition, the structured plan, the natural language description and the user description, and storing the generated embeddings in the shot data store. ([Paragraph 50], Ontology: stored information that provides a data model for storage of data in one or more databases and/or other data stores. For example, the stored data may include definitions for data object types and respective associated property types. An ontology may also include respective link types/definitions associated with data object types, which may include indications of how data object types may be related to one another. An ontology may also include respective actions associated with data object types or data object instances. The actions may include defined changes to values of properties based on various inputs. An ontology may also include respective functions, or indications of associated functions, associated with data object types, which functions may be executed when a data object of the associated type is accessed. An ontology may constitute a way to represent things in the world. An ontology may be used by an organization to model a view on what objects exist in the world, what their properties are, and how they are related to each other. An ontology may be user-defined, computer-defined, or some combination of the two. An ontology may include hierarchical relationships among data object types. [Paragraph 58], A documentation generation module 160 is configured to generate a prompt to a language model, such as LLM 130. As described in further detail below, the documentation generation module 160 may generate a prompt based on data provided by the user interface module 104 (e.g., a user input, tool information, and/or the like), the context module 110 (e.g., conversation history and/or other contextual information), and/or a software application such as app 180 (e.g., user workflow and corresponding workflow information). [Paragraph 76], The prompt can include information associated with one or more tools selected by the user, such as in the form of tool information, which enables the LLM 130 to generate a tool call that can be used by the AIS to communicate with a data processing service. Tool information may indicate, for example, how data that may be accessed by the LLM (via tool calls) is structured, such as in an ontology or other format. Tool information can indicate properties associated with a particular object type, such as an object type selected by the user in the user input at interaction 1. Tool information can include instructions for implementing a tool, instructions for generating a tool call, including instructions for formatting a tool call, tool implementation examples for executing one or more tool operations, and/or other information that may allow the LLM to provide more meaningful responses to the AIS. Tool implementation examples included in an LLM prompt can include pre-defined examples (e.g., the same for each use of the tool), user-selected or user-generated examples, and/or examples that are dynamically configured by the AIS 102 based on context.)
As per claim 17, rejection of claim 14 is incorporated:
Chok teaches generating a natural language shot generation prompt including the structured plan and natural language shot instructions; providing the natural language shot generation prompt to the LLM and receiving the natural language description in response. ([Paragraph 38], A language model can be useful for natural language processing, including receiving natural language prompts and providing natural language responses based on the text on which the model is trained. A language model may include an n-gram, exponential, positional, neural network, and/or other type of model. [Paragraph 50], Ontology: stored information that provides a data model for storage of data in one or more databases and/or other data stores. For example, the stored data may include definitions for data object types and respective associated property types. An ontology may also include respective link types/definitions associated with data object types, which may include indications of how data object types may be related to one another. An ontology may also include respective actions associated with data object types or data object instances. The actions may include defined changes to values of properties based on various inputs. An ontology may also include respective functions, or indications of associated functions, associated with data object types, which functions may be executed when a data object of the associated type is accessed. An ontology may constitute a way to represent things in the world. An ontology may be used by an organization to model a view on what objects exist in the world, what their properties are, and how they are related to each other. An ontology may be user-defined, computer-defined, or some combination of the two. An ontology may include hierarchical relationships among data object types. [Paragraph 58], A documentation generation module 160 is configured to generate a prompt to a language model, such as LLM 130. As described in further detail below, the documentation generation module 160 may generate a prompt based on data provided by the user interface module 104 (e.g., a user input, tool information, and/or the like), the context module 110 (e.g., conversation history and/or other contextual information), and/or a software application such as app 180 (e.g., user workflow and corresponding workflow information). [Paragraph 76], The prompt can include information associated with one or more tools selected by the user, such as in the form of tool information, which enables the LLM 130 to generate a tool call that can be used by the AIS to communicate with a data processing service. Tool information may indicate, for example, how data that may be accessed by the LLM (via tool calls) is structured, such as in an ontology or other format. Tool information can indicate properties associated with a particular object type, such as an object type selected by the user in the user input at interaction 1. Tool information can include instructions for implementing a tool, instructions for generating a tool call, including instructions for formatting a tool call, tool implementation examples for executing one or more tool operations, and/or other information that may allow the LLM to provide more meaningful responses to the AIS. Tool implementation examples included in an LLM prompt can include pre-defined examples (e.g., the same for each use of the tool), user-selected or user-generated examples, and/or examples that are dynamically configured by the AIS 102 based on context.)
Allowable Subject Matter
Claim(s) 5, 9 is/are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DONG U KIM whose telephone number is (571)270-1313. The examiner can normally be reached 9:00am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bradley Teets can be reached at 5712723338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DONG U KIM/Primary Examiner, Art Unit 2197