Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment/Arguments
1. Applicant’s arguments filed on July 21, 2026 regarding the rejection under 35 U.S.C. 101 have been fully considered but are not persuasive.
Applicant’s arguments on pages 10-12 regarding step 2A, Prong One have been fully considered but are not persuasive. Claims 2 and 14 are no longer rejected under 35 U.S.C. 101. It is acknowledged that claims 1, 12, and 13 fall within the statutory categories under Step 1 of the Alice/Mayo analysis.
Applicant argues that amended claim 1 does not recite the mental processes identified in the previous Office Action in view of the newly recited programmatic structural representation, the determination of an action based on the task and the representation, the interaction with an identified component, the receipt of the assistive-technology output, the updating of the representation, and the repetition of the navigation-testing operation. This argument is not persuasive. Paragraph [0048] of the Specification states that the visual representation may be a “a DOM/accessibility tree representation, focus order graph, state transition graph, timeline, diagram, image, or the like.” Thus, the Specification broadly describes the representation as a diagrammatic organization of webpage components, relationships, and states. A person could visit and explore a webpage, observe and evaluate its structure and functionality, identify the user-interface components and their relationships, and draw a corresponding DOM-like tree, accessibility-tree diagram, graph, timeline, or other structural diagram using pen and paper. The person could then review the task and diagram, identify the component relevant to the task, and determine an action to perform with respect to that component. These activities involve observation, evaluation, and judgement that can practically be performed in the human mind with the aid of pen and paper.
Applicant next argues that the subsequent action is not selected merely by observing information and making a judgement since it is determined using a computer-maintained representation of the current interface state that was updated based on an output generated by the assistive technology. This argument is not persuasive. Characterizing the representation as “computer-maintained” does not alter the nature of the information represented or the determination made from the information. A person performing accessibility testing could visit the webpage, observe its visible structure and functionality, click links, activate buttons, enter text, or otherwise interact with user-interface components to determine how the webpage functions. A sighted tester could listen to the outputs generated by the assistive technology and compare those outputs with the visually observed webpage structure, functionality, and resulting interface state to evaluate whether the assistive technology correctly interprets and communicates the structure and functionality of the webpage. The test could then modify a hand-drawn DOM-like or accessibility tree diagram to reflect the resulting interface state and determine the subsequent navigation action based on the preceding assistive-technology output and the updated diagram. Reviewing the output, comparing the output with the observed webpage state, updating the diagram, and deciding which action should be performed next constitute observation, comparison, evaluation, and judgement.
Applicant further argues that the accessibility result is not determined merely by observing outputs and judging whether an accessibility issue exists since the determination depends on operations performed on a rendered webpage, observing how the webpage responses. The person can listen to the corresponding assistive-technology outputs and compare those outputs with the components, functions, and interface states actually presented by the webpage. The person can then determine whether the assistive-technology correctly interprets the webpage and whether the sequence of navigation actions reaches an interface state associated with completion of the task. The webpage interactions and assistive-technology outputs provide the information being evaluated, while determining whether the sequence completes the task remains an observation, comparison, evaluation, and judgement.
Paragraph [0009] of the Specification itself acknowledges that webpage accessibility testing may be performed manually, while explaining that such testing may require substantial time and resources and may produce subjective outcomes. The fact that manual testing may be inefficient or time-consuming does not establish that the recited testing cannot be practically performed by a person. Moreover, claim 1 merely requires repeating the navigation-testing operation “at least once” and does not require testing every possible combination of webpage components or processing a quantity of information that the human mind is not equipped to evaluate.
Applicant next argues next that the claimed limitations cannot practically be performed in the human mind since a person cannot mentally cause the first AI agent to interact with an identified webpage component, receive an output generated by assistive technology, update the programmatic structural representation, and cause a language model to determine a subsequent action. This argument improperly conflates the identified mental processes with the computer elements used to perform or implement those processes. The Office does not maintain that a person cannot mentally operate an AI agent, language model, browser, or assistive-technology system. Rather, the identified mental processes include exploring the webpage, observing its structure and functionality, organizing the observed components and interface states into a diagram, determining an action based on the task and diagram, listening to and evaluating the assistive-technology output, comparing the output with observed webpage structure and functionality, updating the diagram to reflect the resulting interface state, determining a subsequent action, and judging whether the sequence reaches completion of the task.
The MPEP expressly provides that observations, evaluations, judgements, and opinions fall within the mental process grouping and that a claim may recite a mental process even when the process is performed on a computer, within a computer environment, or using a computer as a tool. The use of pen and paper to construct or update a diagram likewise does not negate the mental nature of the recited observations, comparisons, evaluations, and judgements. The first AI agent, language model, assistive-technology, and corresponding computer operations are additional elements considered separately under Step 2A, Prong Two.
Finally, Applicant concludes that amended claim 1 is directed to a specific computer-implemented process for accessibility testing and is not directed merely to generate instructions, selecting an action, or identifying an accessibility issue. However, Step 2A, Prong One determines whether the claim recites a judicial exception. Claim 1 recites observations, comparisons, evaluations, and judgements that can be practically performed by a person manually navigating the webpage, listening to the assistive-technology outputs, comparing those outputs with the observed webpage structure and functionality, and recording the webpage structure and successive interface states using diagrams or other written representations. The additional recitation of computer implementation does not remove those limitations from the mental process grouping. Whether the additional computer elements integrate the recited mental processes into a practical application is considered separately under Step 2A, Prong Two. Accordingly, Applicant’s arguments do not establish that amended claim 1 fails to recite a mental process under Step 2A, Prong One.
Applicant’s arguments under Step 2A, Prong Two have been fully considered but are not persuasive. The additional elements have been considered both individually and in combination with the recited mental processes. All additional elements are given weight in the Prong Two analysis without considering as a whole, the additional elements do not integrate the recited mental processes into a practical application.
Applicant first argues that the claimed arrangement integrates any alleged abstract idea into a practical application since the programmatic structural representation identifies webpage components, the language model determines an action concerning an identified component, the first AI agent performs the action, the assistive technology produces an output, the structural representation is updated based on the output, and the output and the updated representation are used to determine a subsequent action. Applicant further asserts that this arrangement improves accessibility testing by using task-oriented interaction with a rendered webpage rather than an isolated static or syntactic analysis. This argument is not persuasive. The claim uses the webpage, programmatic structural representation, language model, AI agent, and assistive technology as tools for performing the recited navigation analysis and determining whether a task can be completed. The claim does not recite an improvement to the functioning of the webpage, browser, assistive technology, language model, AI agent, DOM, or accessibility tree. Nor does the claim recite a particular technical technique for generating the structural representation, processing the assistive-technology output, updating the representation, or determining the navigation actions. Instead, the claim identifies the information to be considered and the results to be produced. The asserted improvement is therefore an improvement in carrying out the recited accessibility evaluation, rather than an improvement to the operation of a computer or another technology.
The recitation of task-oriented interaction with a rendered webpage also does not, by itself, provide a practical application. A person could perform the same type of accessibility testing by navigating the webpage, observing its structure and functionality, listening to the assistive-technology outputs, comparing those outputs with the observed webpage state, and determining whether the task can be completed. Claim 1 assigns those observations, evaluations, and judgements to computer components and applies them in the technological environment of a webpage and assistive technology. Generally linking the use of an abstract idea to a particular technological environment does not impose a meaningful limitation on the judicial exception. See MPEP 2106.04(d) and 2106.05(h).
Applicant next argues that the amended limitations address the prior characterization of the prompt engine, language model, and first AI agent as being recited at a high level. Applicant points to the recitation of the information included in the prompt, the basis for determining the action, the relationship between the action and identified component, the action performed by the AI agent, the output generated by the assistive technology, and the information used to determine the subsequent action. Although the amended claim provides additional detail regarding the information supplied to the components and the results expected from them, it still recites those components functionally and at a high level. The claim does not recite how the prompt engine technically constructs the prompt, how the language model processes the task and structural representation to determine the action, how the first AI agent technically performs the interaction, or how the assistive-technology output causes a particular technical modification to the structural representation. Rather, the claim instructs the components to perform the recited mental analysis and implement the resulting action. Specifying the information used by a generic computer component and desired results of its processing does not, without a particular technological implementation, transform the mental process into a practical application. See MPEP 2106.05(f).
Applicant further argues that receiving the output from the assistive technology is not insignificant data gathering since the output is used to update the structural representation and determine the subsequent action. The fact that the received output is subsequently used does not establish that its receipt meaningfully limits the judicial exception. The assistive-technology output supplies information for the recited evaluation of the current webpage state and determination of the next navigation action. Claim 1 does not recite an improvement to how the assistive technology generates the output, how the output is received, or how the output is technically converted into an updated structural representation. The output therefore serves as input data for the recited observation, comparison, evaluation, and decision-making process. Collecting information required to perform an abstract analysis does not integrate the abstract analysis into a practical application merely because the collected information is used during the analysis.
Applicant next argues that performing the determined action is not insignificant extra-solution activity since the action is performed on identified component, changes the interface state, and causes the assistive technology to generate the output used in the subsequent operation. The Office acknowledges that the action occurs within the claimed sequence rather than solely before or after the sequence. Nevertheless, its placement within the sequence does not establish integration into a practical application. The limitation merely instructs the first AI agent to carry out an action selected through the recited mental determination by interacting with a webpage component. The claim does not recite a particular technical mechanism for performing that interaction or an improvement to webpage control, browser operation, or the operation of the assistive technology. The resulting change in interface state is the ordinary result of performing an action on a webpage and provides additional information for the next abstract determination. Thus, even when considered as an operative part of the sequence, performing the action merely applies the abstract determination using a computer as a tool in the environment of a webpage.
Applicant further argues that the accessibility result is not merely a presentation or labeling of information and that determining whether the sequence reaches a task-completion state constitutes the technical application of the preceding operations. This argument is not persuasive. Determining the accessibility result is itself part of the identified mental process. The limitation requires reviewing the sequence of actions, the corresponding assistive technology outputs, and the represented interface states and judging whether the sequence reached a state associated with completion of the task. The accessibility result does not control or modify the webpage, correct an accessibility issue, alter the operation of the assistive technology, or otherwise produce a technological improvement. It merely indicates the conclusion reached from the recited observation and evaluation. The judicial exception cannot provide its own practical application merely by producing the result of the abstract analysis.
When considered as an ordered combination, claim 1 uses a prompt engine to organize information into a prompt, a language model to determine navigation actions, an AI agent to performed the determined actions, assistive technology to provide outputs, and a structural representation to record or reflect webpage states. These elements interact to collect and organize information, perform the selected webpage actions, and evaluate whether the task can be completed. However, the ordered arrangement does not recite a particular technological implementation that improves the operation of any recited computer component or another technology. It instead uses the additional computer elements as tools to perform and repeat the recited mental processes within the field of webpage accessibility testing. Accordingly, the additional elements, individually or in combination, do not impose a meaningful limitation on the judicial exception, and claim 1 does not integrate the recited mental processes into a practical application under Step 2A, Prong Two. The MPEP requires consideration of the claim as a whole and of the interaction among all limitations, but the ultimate inquiry remains whether those limitations meaningfully apply the exception rather than merely implement it with computer tools or in a technological environment.
Applicant’s arguments under Step 2B have been fully considered but are not persuasive. Step 2B evaluates whether the claim, considered as a whole, includes elements that amount to significantly more than the judicial exception. The additional elements have therefore been considered both individually and in combination.
Applicant argues that claim 1 recites significantly more than the identified abstract idea due to the particular arrangement of the programmatic structural representation, language model, first AI agent, assistive-technology output, and updated representation used to determine subsequent actions. This argument is not persuasive. The determination of an action based on a preceding output and updated representation, and the determination of whether the sequence reaches task completion are part of the abstract observations, evaluations, and judgements identified under Step 2A, Prong One. Those abstract determinations cannot themselves supply the inventive concept.
The additional elements implement those determinations through computer components recited according to their ordinary functions. The structural representation organizes webpage-component and interface-state information, the language model determines an action from supplied information, the fist AI agent performs the determined action, and the assistive technology produces output used in the next determination. Claim 1 does not recite a particular technique for generating or updating the structural representation, a particular technical manner in which the language model determines the action, or a special mechanism through which the first AI agent interacts with the webpage. Rather, the claim recites the information by using the components, the functional relationships among the components, and the results to be produced.
Applicant asserts that the characterization of the individual operations as generic or conventional does not established the ordered combination as well-understood, routine, and conventional. It is acknowledged that the conventionality of the individual elements does not, standing alone, establish that the ordered combination lacks an inventive concept. However, when considered in the claimed order, the additional elements merely implement and repeat the abstract decision-making and evaluation process. The output from one generically recited operation supplies information for the next determination, but the claim does not recite a nonconventional technical implementation or a technological improvement arising from that arrangement.
Applicant further argues that the accessibility result is determined from the actions, output, and interface states produced through the claimed arrangement. This argument is not persuasive. Determining whether the sequence reaches an interface state associated with completion of the task is itself part of the abstract evaluation identified under Step 2A, Prong One. The accessibility issue, modify webpage behavior, or alter the operation of the assistive technology, correct an accessibility issue, modify webpage behavior, or alter the operation of the assistive technology. It reflects the conclusion reached from the recited observation, evaluation, and judgement.
Applicant also argue that independent claims 12 and 13 incorporate the limitations of amended claim 1 and are eligible for the same reasons. Claims 12 and 13 recite substantially the same abstract determinations and additional elements in computer-readable medium and system form. Accordingly, Applicant’s arguments are unpersuasive for the same reasons discussed above. Applicant presents no separate eligibility argument for the dependent claims beyond their dependency from the independent claims. Therefore, Applicant’s arguments do not overcome the rejection under 35 U.S.C. 101.
2. Applicant’s arguments filed on July 21, 2026 regarding the rejection under 35 U.S.C. 103 have been fully considered but are not persuasive.
Applicant’s arguments regarding Kumar on pages 16-18 have been fully considered but are not persuasive. Applicant’s arguments focus on limitations allegedly absent from Kumar. The present rejection, however, does not rely on Kumar alone, but the combined teachings of Kumar, Deshmukh, and Azose. Kumar supplies the instruction following language model and virtual agent framework, Deshmukh supplies the assistive-technology based webpage navigation and testing operations, and Azose supplies the DOM or accessibility-tree representation identifying webpage components and the use of the structural webpage information by a generative language model. Accordingly, the amended limitations are evaluated based on the combined teachings of Kumar, Deshmukh, and Azose rather than whether Kumar individually discloses the entire claimed process.
Applicant argues that Kumar paragraphs [0031] and [0051] do not disclose a prompt including a programmatic structural representation identifying webpage components or determining an action based on the task and that representation. The current rejection does not rely on Kumar alone for these limitations. Kumar teaches an instruction following language model that supports a virtual agent and receives a prompt and request. Azose paragraph [0028] teaches providing a DOM or accessibility tree to a generative language model, inspecting nodes identifying webpage components, and determining context or instructions based on that structural information. Thus, Azose supplies the structural webpage context use by a language model and directly connects with Kumar’s language model and virtual agent framework. Deshmukh paragraph [0052] further teaches a navigation-testing task identifying an operation that should be performed on a webpage, such as accessing an identified link, and generating an input command to perform that operation. The combined teachings therefore determine, based on the task and the programmatic structural representation, a navigation action for the AI agent to perform with respect to a component identified in the representation.
Applicant next argues that the actions described in Kumar paragraph [0064] merely correspond to conversational intents and are not actions performed by the virtual agent on an identified webpage component. This argument does not address the combined discloses relied upon in the current rejection. Kumar teaches an instruction following language model supporting a virtual agent and executing actions corresponding to a determine intent. When Kumar’s virtual agent framework is applied to the accessibility-testing process of Deshmukh, the determined intent corresponds to the webpage-navigation objective or goal, such as accessing a particular link, navigating to a target page, or otherwise performing an interaction needed to complete the testing task. Azose teaches identifying the webpage component associated with the navigation objective through a node in a DOM or accessibility tree provided to a generative language model. Deshmukh paragraphs [0029], [0052], and [0053] teach generating and executing an input command to interact with an identified webpage link and simulate the behavior of a low-vision user interacting with the webpage. Thus, Kumar supplies the language-model and virtual agent framework for determining and executing an action corresponding to the navigation intent, Azose supplies the structural identification of the webpage component on which the action is to be performed, and Deshmukh supplies the navigation task and execution of the interaction on that identified component. The rejection does not rely on Kumar’s conversational routing alone as disclosing the entire navigation-testing operation.
Applicant’s argument concerning Kumar paragraph [0096] is likewise not persuasive with respect to the present rejection. Applicant argues that paragraph [0096] merely describes a browser through which a user interacts with Kumar’s implementation and does not disclose an AI agent interacting with a webpage. The current rejection does not rely on paragraph [0096] for the amended limitation requiring the first AI agent to perform the determined action on the identified webpage component. The limitation is mapped to the combined disclosures of Kumar paragraph [0064], Azose paragraph [0028], and Deshmukh paragraphs [0029], [0052], and [0053]. Applicant’s distinction between a user accessing Kumar’s system through a browser and an AI agent interacting with a webpage component therefore does not address the mapping presently relied upon.
Applicant further argues that Kumar’s prompt state and dialog history are not the claimed programmatic structural representation and do not reflect the current interface state of the tested webpage. The current rejection does not rely on Kumar’s prompt state or dialog history as the programmatic structural representation. Deshmukh paragraphs [0054] and [0060] teach receiving screen-reader output, converting the output to a navigation command, executing the command, monitoring the resulting webpage behavior, determining how the webpage changes, and analyzing the HTML displayed after execution of the command. Azose paragraphs [0028] and [0057] teach obtaining webpage context from a DOM or accessibility tree. Under the broadest reasonable interpretation, obtaining the DOM or accessibility tree context after Deshmukh’s output-driven webpage interaction corresponds to updating the programmatic structural representation to reflect the current interface state. Thus, the screen-reader output informs the interaction that changes the webpage, and the resulting current webpage state is reflected in the refreshed DOM or accessibility-tree context.
Applicant further argues that Kumar’s determination of a “next prompt” does not disclose repeating the navigation-testing operation and determining a subsequent action based on both a preceding assistive-technology output and an updated structural representation. The present rejection again does not rely on Kumar alone. Kumar paragraph [0076] and [0078] teach repeated language model processing and determining the next operation using a prior state and history. Azose paragraphs [0057] and [0063] teach obtaining current webpage context from a DOM or accessibility tree and repeating generative language model processing using prior information as context. Deshmukh paragraphs [0059] and [0060] teach using a preceding screen-reader output associated with a webpage element to generate a navigation command and determining how the webpage changes after execution of that command. Accordingly, the combined teachings repeat the navigation-testing operation and determine a subsequent action using a preceding assistive-technology output and the updated DOM or accessibility tree representation of the current webpage state.
Applicant’s arguments regarding Deshmukh on pages 18-19 have been fully considered but are not persuasive in view of the presently applied combination of Kumar, Deshmukh, and Azose. Applicant principally argues that Deshmukh does not disclose the claimed language-model and programmatic structural representation limitations. However, the current rejection relies on the references collectively. Kumar teaches an instruction following language model supporting a virtual agent, Azose teaches providing DOM or accessibility tree webpage context to a generative language model, and Deshmukh teaches using machine learning based natural language processing to convert screen-reader information into simulated webpage interactions that are executed in a browser. Thus, all three references involve language-based processing and that directly supports the combined accessibility-testing process, while each reference supplies a different portion of the claimed arrangement.
Applicant first argues that Deshmukh paragraphs [0029]-[0033] merely convert screen-reader voice signals into simulated user interactions and do not teach a language model determining an action based on a task and programmatic structural representation. This argument does not account for Deshmukh’s complete disclosure or the references in combination. Deshmukh paragraph [0029] expressly teaches converting the screen-reader voice signal into text and then using “a machine learning algorithm trained to convert textual instructions into simulated user interactions,” including “a natural language processing algorithm.” Paragraph [0030] similarly teaches using a machine learning algorithm to convert the text into a simulated user interaction and executing the interaction in the browser. Thus, Deshmukh does not merely perform a fixed conversion of a screen-reader signal into a command; it uses machine learning based natural language processing to determine a webpage interaction form the screen-reader information.
Deshmukh’s NLP-based interaction determination directly complements Kumar’s instruction following language model and Azose’s generative language model. Kumar paragraph [0031] and [0064] teach a language mode supporting a virtual agent and executing actions corresponding to a determined intent. Azose paragraph [0028] teaches providing a DOM or accessibility tree identifying webpage components to a generative model and determining context or instructions based on that structural webpage information. Deshmukh paragraph [0052] further teaches a navigation-testing task specifying an operation that should be performed on a webpage, such as accessing an identified link. Accordingly, Kumar supplies the language model and virtual agent framework, Azose supplies the programmatic structural representation used by the generative language model, and Deshmukh supplies the NLP-based determination and execution of simulated webpage interactions for assistive-technology testing.
Applicant next argues Deshmukh does not update a programmatic structural representation based on assistive-technology output and that monitoring browser behavior is not the claimed update. The current rejection does not rely on Deshmukh alone for the programmatic structural representation. Deshmukh paragraphs [0054] and [0060] teach receiving a screen-reader output, converting that output into a navigation command, executing the command, determining how the webpage changes, and analyzing the HTML displayed after execution of the command. Azose paragraphs [0028] and [0057] teach obtaining webpage context from a DOM or accessibility tree. Under the broadest reasonable interpretation, obtaining the DOM or accessibility tree context after Deshmukh’s output-driven webpage interaction corresponds to updating the programmatic structural representation to reflect the current interface state. Deshmukh supplies the output driven interaction and resulting webpage change, while Azose supplies the structural representation reflecting that changed state.
Applicant’s argument regarding Deshmukh’s static HTML analysis likewise does not address the present mapping. The current rejection does not rely on paragraph [0033] alone as teaching the claimed structural representation or the update. Rather, Deshmukh’s monitoring and analysis of the webpage after execution of the navigation command is combined with Azose’s teaching of obtaining current DOM or accessibility-tree context. The combined teachings therefore provide more than a static review of HTML code.
Applicant further argues that Deshmukh paragraphs [0059] and [0060] do not determine a subsequent action based on additionally on updated programmatic structural representation. Again, the current rejection relies on all three references. Deshmukh teaches using a preceding screen-reader output associated with a webpage element to determine and execute a navigation command and then determining how the webpage changes. Kumar paragraphs [0076] and [0078] teach repeated language model processing using prior state and history to determine a next operation. Azose paragraphs [0057] and [0063] teach obtaining current DOM or accessibility tree context and repeating generative language model processing using prior information as context. Accordingly, the combined teachings determine a subsequent action using the preceding assistive-technology output and the updated DOM or accessibility tree representation of the current webpage state.
Applicant also argues that Deshmukh does not teach the claimed accessibility result based on the sequence of actions, assistive-technology outputs, and represented interface states, The rejection again relies on the references collectively. Deshmukh paragraph [0039] teaches a test script specifying ordered webpage actions and a task objective, such as reaching a login in page. Paragraph [0055] teaches determining accessibility compliance based on screen-reader outputs, simulated interactions, resulting browser behavior, and the test script, and paragraph [0057] teaches generating an accessibility result. Kumar paragraph [0021] teaches determining completion when a specified goal has been achieved, while Azose paragraph [0028] supplies the DOM or accessibility tree representation of the interface states produced during the sequence. The combined teachings therefore determine whether the assistive-technology interaction sequence reaches an interface state associated with completion of the task.
Applicant’s concluding arguments on page 20 have been fully considered but are not persuasive. Applicant’s conclusion address only the combination of Kumar and Deshmukh and repeats the previously addressed assertions that the references do not teach the programmatic structural representation, interaction with identified webpage components, updating of that representation, determination of subsequent actions, and the claimed accessibility result. As explained above, the present rejection relies on the combined teachings of Kumar, Deshmukh, and Azose, not on Kumar and Deshmukh alone.
Azose supplies the DOM or accessibility tree representation identifying webpage components and the use of that structural context by a generative language model. Kumar supplies the prompt orchestration, instruction following language model, and virtual agent framework. Deshmukh supplies the assistive-technology based navigation, NLP-based determination and execution of simulated webpage interactions, monitoring of resulting webpage changes, and accessibility testing. Applicant’s arguments directed to what Kumar and Deshmukh individually or collectively fail to disclose do not address the presently applied three reference combination and therefore do not overcome the rejection.
To the extent Applicant generally asserts that there is a lack of motivation to combine the references, this argument is not persuasive. As set forth in the rejection, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, having a combination of Kumar, Deshmukh, and Azose before them, to incorporate Azose’s DOM or accessibility tree webpage context into Kumar’s prompt orchestration virtual agent when performing the screen-reader based accessibility testing of Deshmukh. One would have been motivated to make such a combination in order to determine whether a blind or visually impaired user could reliably navigate a website and complete intended tasks using assistive technology, and to identify points at which inaccessible webpage elements prevent or hinder further navigation. This would allow website accessibility issues to be detected and corrected, thereby enabling users who rely on assistive technology to navigate websites with greater confidence.
Applicant presents no separate substantive arguments for independent claims 12 and 13 or for the dependent claims, but instead repeats the arguments presented for claim 1. As the rejection of claim 1 is maintained, the rejections of claims 3-13, and 15-23 are likewise maintained.
Accordingly, the rejection under 35 U.S.C. 103 is maintained.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1, 3-13, and 15-23 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more.
101 Subject Matter Eligibility Analysis
Step 1: Claims 1, 3-13, and 15-23 are within the four statutory categories (a process, machine, manufacture or composition of matter).
Step 2A Prong One, Step 2A Prong Two, and Step 2B Analysis:
Step 2A Prong One asks if the claim recites a judicial exception (abstract idea, law of nature, or natural phenomenon). If the claim recites a judicial exception, analysis proceeds to Step 2A Prong Two, which asks if the claim recites additional elements that integrate the abstract idea into a practical application. If the claim does not integrate the judicial exception, analysis proceeds to Step 2B, which asks if the claim amounts to significantly more than the judicial exception. If the claim does not amount to significantly more than the judicial exception, the claim is not eligible subject matter under 35 U.S.C. 101.
None of the claims represent an improvement to technology.
Claims 1 and 3-11 are directed to a method consisting of a series of steps, meaning that it is directed to the statutory category of process. Claims 12-13, and 15-23 are directed to storage mediums and processors which are machines.
Regarding claim 1, the following claim elements are abstract ideas:
generating a prompt…the prompt including a task for the first Al agent to complete on a web page using actions configured to mimic interactions expected of a user of an assistive-technology (This is an abstract idea of a mental process. This involves observation of a task, judgement in adopting how the user of assistive technology would approach the task, evaluation of how the user would interact with the web page, and decision-making in determining actions to perform. A person could mentally consider how a user relying on assistive technology would navigate a webpage and determine corresponding steps to complete the task. Such activities can be performed in the human mind and therefore falls within the mental process grouping of abstract ideas. See MPEP 2106.04(a)(2)(III).);
a programmatic structural representation of the web page identifying components in a user interface of the web page (This is an abstract idea of a mental process. The limitation involves observing a webpage, identifying the user-interface components present on the webpage, and organizing these components into a structural mapping. A person could inspect the webpage, identify buttons, links, headings, text fields, and other components, and record the components and their relationships in a list, hierarchy, tree, or diagram using observation and judgement. This identification and structural mapping can be practically performed in the human mind with the aid of pen and paper or basic computational tools. The recitation of the representation is “programmatic” merely implements the underlying mental process using a computer and does not require a particular technique for generating the representation. Therefore, the limitation falls within the mental process grouping of abstract ideas.);
performing a navigation-testing operation, including: determining…and based on the task and the programmatic structural representation, an action for the first Al agent to perform on the web page with respect to one of the components identified in the programmatic structural representation (This is an abstract idea of a mental process. The limitation involves observing a programmatic structural representation, such as a displayed DOM or accessibility tree, identifying the component relevant to the task, and selecting an action to perform on that component. For example, a person could review a DOM tree, locate a login button, and determine that the button should be selected to advance the task. This type of observation, evaluation, and judgement can be practically performed in the human mind with the aid of pen and paper or basic computational tools and therefore falls within the mental process grouping of abstract ideas.);
repeating the performance of the navigation-testing operation at least once, wherein, in each repetition… determines a subsequent action based on at least one preceding output from the accessibility-technology and the updated programmatic structural representation (This is an abstract idea of a mental process. The limitation involves reviewing a preceding assistive-technology output and an updated structural mapping of webpage components and selecting a subsequent action based on that information. For example, a person could review a screen-reader output and an updated DOM-like tree, determine which webpage component should be addressed next, and repeat the analysis for another action. This type of observation, evaluation, and decision-making can be practically performed in the human mind with the aid of pen and paper or basic computational tools and therefore falls within the mental process grouping of abstract ideas.);
performing an accessibility test of the web page to determine, based on a sequence including the action and the subsequent action, corresponding outputs from the assistive-technology, and interface states reflected in the updated programmatic structural representation, an accessibility result indicating whether a sequence of the actions performed by the first Al agent reaches an interface state associated with completion of the task using the assistive-technology (This is an abstract idea of a mental process. The limitation involves reviewing a sequence of webpage actions, corresponding assistive-technology outputs, and represented interface states, and determining whether the sequence reached a state associated with a completed task. For example, a person could review a written list of actions, screen-reader outputs, and drawings of the webpage states and determine whether the intended webpage task was successfully completed. This type of observation, evaluation, and judgement can be practically performed in the human mind with the aid of pen and paper or basic computational tools and therefore falls within the mental process grouping of abstract ideas.).
The following claim elements are additional elements which, taken alone or in combination with the other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
by a prompt engine, for a language model integrated with a first artificial intelligence (Al) agent (This limitation merely provides instructions for performing the abstract idea using generic computer components. The “prompt engine,” “language model,” and “AI agent” are recited at a high level and are used as tools to carry out the abstract idea, amounting to no more than instructions to apply the judicial exception on a computer. Further, specifying that the instructions are provided to a language model integrated with an AI agent constitutes insignificant extra-solution activity, as it merely identifies the environment or tool used to perform the abstract idea and does not impose a meaningful limit on the claim.),
providing the prompt to the language model (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).);
by the first Al agent (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).)
receiving, by the first Al agent, at least one output from the assistive-technology in response to performing the determined action on the web page (This limitation recites receiving a first output, which is a well-understood, routine, and conventional activity. This step merely involves obtaining information resulting from a performed action and does not impose a meaningful limit on the abstract idea. Such data gathering is a basic computer function and amounts to insignificant extra-solution activity.);
performing, by the first Al agent, the determined action on the web page by interacting with the component in the component identified in the programmatic structural representation (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).)
updating, based on the at least one output, the programmatic structural representation to reflect a current interface state of the web page (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).);
Regarding claim 3, the rejection of claim 1 is incorporated herein. Further, claim 3 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
when performing any action does not further the first Al agent's completion of the task to a predetermined degree, invoking an action assistance from a second Al agent to achieve the task, and wherein the second Al agent is configured to assist the first Al agent to complete the task, wherein any action includes the first action and the at least one subsequent action (This limitation recites merely providing instructions for applying the abstract idea based on a condition. The step directs that, upon determining insufficient progress, assistance is invoked from another agent, which amounts to insignificant extra-solution activity, as it merely identifies an additional generic component used to carry out the abstract idea and does not impose any meaningful limit on the claim.).
Regarding claim 4, the rejection of claim 3 is incorporated herein. Further, claim 4 recites the following abstract ideas:
determining…that an action does not further the first Al agent's completion of the task to a predetermined degree when a number of subsequent actions are continuations of a preceding action without furthering completion of the task exceeds a pre-determined threshold (This is an abstract idea of a mental process. The limitation recites determining whether progress toward a task has been made based on prior actions and a threshold. A person could observe a series of actions, recognize that the actions are not advancing completion of a task, and determine that progress has not been made once the number of such actions exceeds at threshold. This type of evaluation and judgement can be performed in the human mind, or with the aid of computational tools, and therefore falls within the mental process grouping of abstract ideas.).
The following claim elements are additional elements which, taken alone or in combination with the other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
by the first Al agent (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).),
Regarding claim 5, the rejection of claim 3 is incorporated herein. Further, claim 5 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
allocating, by the first Al agent, control of the task to the second Al agent to invoke the action assistance (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).).
Regarding claim 6, the rejection of claim 1 is incorporated herein. Further, claim 6 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
updating at least one visual representation of the web page with information based on the at least one output received during the navigation-testing operation and at least one output received during a repetition of the navigation-testing operation, wherein the visual representation of the web page includes representation of components of the user interface of the web page (This limitation merely instructs that information obtained during navigation-testing operations be used to update a visual representation of the webpage, without reciting any particular technical manner of performing the update. The limitation merely uses a generic computer as a tool to apply the abstract analysis and does not impose a meaningful limitation on the judicial exception.).
Regarding claim 7, the rejection of claim 1 is incorporated herein. Further, claim 7 recites the following abstract ideas:
generating…inputs to be included in the generated prompt (This is an abstract idea of a mental process. The limitation recites selecting and organizing information to be used in formulating instructions. This type of evaluation and judgement can be performed in the human mind and therefore falls within the mental process grouping of abstract ideas.).
Regarding claim 8, the rejection of claim 7 is incorporated herein. Further, claim 8 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
wherein the inputs further include at least one of: a role description, configuration information for a particular screen reader software, information about a particular user device, a starting point, a log, a list of specific actions, instructions to use a particular navigation tool, and various execution strategies based on the task and the programmatic structural representation (This limitation recites additional information to be included in the inputs and merely describes the content of information used in performing the abstract idea. The recited items identify type of data or instructions that may be provided and do not impose any meaningful limit on the claim. Such recitation constitutes insignificant extra-solution activity because it merely specifies information used or presenting in carrying out the abstract idea.).
Regarding claim 9, the rejection of claim 1 is incorporated herein. Further, claim 9 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
iteratively testing a plurality of patterns of interaction with various components of the user interface of the web page (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).).
Regarding claim 10, the rejection of claim 1 is incorporated herein. Further, claim 10 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
wherein the accessibility result indicates whether the web page has an accessibility issue relating to the perceivability, operability, understandability, or robustness of a web page for mobile-device assistive technology use (This limitation recites merely specifying the type of issue being identified and describes categories of information associated with the abstract idea. The recitation does not impose any meaningful limit on the claim and constitutes insignificant extra-solution activity, as it merely characterizes the information being evaluated and presented.).
Regarding claim 11, the rejection of claim 1 is incorporated herein. Further, claim 11 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
wherein the assistive-technology is a screen-reader technology for visually-impaired users to operate web pages (This limitation amounts to insignificant extra-solution activity.).
Regarding claim 12, the following claim elements are abstract ideas:
generate a prompt…the prompt including a task for the first Al agent to complete on a web page using actions configured to mimic interactions expected of a user of an assistive-technology (This is an abstract idea of a mental process. This involves observation of a task, judgement in adopting how the user of assistive technology would approach the task, evaluation of how the user would interact with the web page, and decision-making in determining actions to perform. A person could mentally consider how a user relying on assistive technology would navigate a webpage and determine corresponding steps to complete the task. Such activities can be performed in the human mind and therefore falls within the mental process grouping of abstract ideas. See MPEP 2106.04(a)(2)(III).);
a programmatic structural representation of the web page identifying components in a user interface of the web page (This is an abstract idea of a mental process. The limitation involves observing a webpage, identifying the user-interface components present on the webpage, and organizing these components into a structural mapping. A person could inspect the webpage, identify buttons, links, headings, text fields, and other components, and record the components and their relationships in a list, hierarchy, tree, or diagram using observation and judgement. This identification and structural mapping can be practically performed in the human mind with the aid of pen and paper or basic computational tools. The recitation of the representation is “programmatic” merely implements the underlying mental process using a computer and does not require a particular technique for generating the representation. Therefore, the limitation falls within the mental process grouping of abstract ideas.);
perform a navigation-testing operation, including: determining…and based on the task and the programmatic structural representation, an action for the first Al agent to perform on the web page with respect to one of the components identified in the programmatic structural representation (This is an abstract idea of a mental process. The limitation involves observing a programmatic structural representation, such as a displayed DOM or accessibility tree, identifying the component relevant to the task, and selecting an action to perform on that component. For example, a person could review a DOM tree, locate a login button, and determine that the button should be selected to advance the task. This type of observation, evaluation, and judgement can be practically performed in the human mind with the aid of pen and paper or basic computational tools and therefore falls within the mental process grouping of abstract ideas.);
repeat the performance of the navigation-testing operation at least once, wherein, in each repetition… determines a subsequent action based on at least one preceding output from the accessibility-technology and the updated programmatic structural representation (This is an abstract idea of a mental process. The limitation involves reviewing a preceding assistive-technology output and an updated structural mapping of webpage components and selecting a subsequent action based on that information. For example, a person could review a screen-reader output and an updated DOM-like tree, determine which webpage component should be addressed next, and repeat the analysis for another action. This type of observation, evaluation, and decision-making can be practically performed in the human mind with the aid of pen and paper or basic computational tools and therefore falls within the mental process grouping of abstract ideas.);
perform an accessibility test of the web page to determine, based on a sequence including the action and the subsequent action, corresponding outputs from the assistive-technology, and interface states reflected in the updated programmatic structural representation, an accessibility result indicating whether a sequence of the actions performed by the first Al agent reaches an interface state associated with completion of the task using the assistive-technology (This is an abstract idea of a mental process. The limitation involves reviewing a sequence of webpage actions, corresponding assistive-technology outputs, and represented interface states, and determining whether the sequence reached a state associated with a completed task. For example, a person could review a written list of actions, screen-reader outputs, and drawings of the webpage states and determine whether the intended webpage task was successfully completed. This type of observation, evaluation, and judgement can be practically performed in the human mind with the aid of pen and paper or basic computational tools and therefore falls within the mental process grouping of abstract ideas.).
The following claim elements are additional elements which, taken alone or in combination with the other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
A non-transitory computer-readable medium (This a high-level recitation of generic computer components for performing the abstract idea. See MPEP 2106.05(f).)
one or more processing circuitries (This a high-level recitation of generic computer components for performing the abstract idea. See MPEP 2106.05(f).)
by a prompt engine, for a language model integrated with a first artificial intelligence (Al) agent (This limitation merely provides instructions for performing the abstract idea using generic computer components. The “prompt engine,” “language model,” and “AI agent” are recited at a high level and are used as tools to carry out the abstract idea, amounting to no more than instructions to apply the judicial exception on a computer. Further, specifying that the instructions are provided to a language model integrated with an AI agent constitutes insignificant extra-solution activity, as it merely identifies the environment or tool used to perform the abstract idea and does not impose a meaningful limit on the claim.),
providing the prompt to the language model (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).);
by the first Al agent (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).)
receiving, by the first Al agent, at least one output from the assistive-technology in response to performing the determined action on the web page (This limitation recites receiving a first output, which is a well-understood, routine, and conventional activity. This step merely involves obtaining information resulting from a performed action and does not impose a meaningful limit on the abstract idea. Such data gathering is a basic computer function and amounts to insignificant extra-solution activity.);
performing, by the first Al agent, the determined action on the web page by interacting with the component in the component identified in the programmatic structural representation (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).)
updating, based on the at least one output, the programmatic structural representation to reflect a current interface state of the web page (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).);
Regarding claim 13, the following claim elements are abstract ideas:
generate a prompt…the prompt including a task for the first Al agent to complete on a web page using actions configured to mimic interactions expected of a user of an assistive-technology (This is an abstract idea of a mental process. This involves observation of a task, judgement in adopting how the user of assistive technology would approach the task, evaluation of how the user would interact with the web page, and decision-making in determining actions to perform. A person could mentally consider how a user relying on assistive technology would navigate a webpage and determine corresponding steps to complete the task. Such activities can be performed in the human mind and therefore falls within the mental process grouping of abstract ideas. See MPEP 2106.04(a)(2)(III).);
a programmatic structural representation of the web page identifying components in a user interface of the web page (This is an abstract idea of a mental process. The limitation involves observing a webpage, identifying the user-interface components present on the webpage, and organizing these components into a structural mapping. A person could inspect the webpage, identify buttons, links, headings, text fields, and other components, and record the components and their relationships in a list, hierarchy, tree, or diagram using observation and judgement. This identification and structural mapping can be practically performed in the human mind with the aid of pen and paper or basic computational tools. The recitation of the representation is “programmatic” merely implements the underlying mental process using a computer and does not require a particular technique for generating the representation. Therefore, the limitation falls within the mental process grouping of abstract ideas.);
perform a navigation-testing operation, including: determining…and based on the task and the programmatic structural representation, an action for the first Al agent to perform on the web page with respect to one of the components identified in the programmatic structural representation (This is an abstract idea of a mental process. The limitation involves observing a programmatic structural representation, such as a displayed DOM or accessibility tree, identifying the component relevant to the task, and selecting an action to perform on that component. For example, a person could review a DOM tree, locate a login button, and determine that the button should be selected to advance the task. This type of observation, evaluation, and judgement can be practically performed in the human mind with the aid of pen and paper or basic computational tools and therefore falls within the mental process grouping of abstract ideas.);
repeat the performance of the navigation-testing operation at least once, wherein, in each repetition… determines a subsequent action based on at least one preceding output from the accessibility-technology and the updated programmatic structural representation (This is an abstract idea of a mental process. The limitation involves reviewing a preceding assistive-technology output and an updated structural mapping of webpage components and selecting a subsequent action based on that information. For example, a person could review a screen-reader output and an updated DOM-like tree, determine which webpage component should be addressed next, and repeat the analysis for another action. This type of observation, evaluation, and decision-making can be practically performed in the human mind with the aid of pen and paper or basic computational tools and therefore falls within the mental process grouping of abstract ideas.);
perform an accessibility test of the web page to determine, based on a sequence including the action and the subsequent action, corresponding outputs from the assistive-technology, and interface states reflected in the updated programmatic structural representation, an accessibility result indicating whether a sequence of the actions performed by the first Al agent reaches an interface state associated with completion of the task using the assistive-technology (This is an abstract idea of a mental process. The limitation involves reviewing a sequence of webpage actions, corresponding assistive-technology outputs, and represented interface states, and determining whether the sequence reached a state associated with a completed task. For example, a person could review a written list of actions, screen-reader outputs, and drawings of the webpage states and determine whether the intended webpage task was successfully completed. This type of observation, evaluation, and judgement can be practically performed in the human mind with the aid of pen and paper or basic computational tools and therefore falls within the mental process grouping of abstract ideas.).
The following claim elements are additional elements which, taken alone or in combination with the other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
a processing circuitry (This a high-level recitation of generic computer components for performing the abstract idea. See MPEP 2106.05(f).);
a memory (This a high-level recitation of generic computer components for performing the abstract idea. See MPEP 2106.05(f).),
by a prompt engine, for a language model integrated with a first artificial intelligence (Al) agent (This limitation merely provides instructions for performing the abstract idea using generic computer components. The “prompt engine,” “language model,” and “AI agent” are recited at a high level and are used as tools to carry out the abstract idea, amounting to no more than instructions to apply the judicial exception on a computer. Further, specifying that the instructions are provided to a language model integrated with an AI agent constitutes insignificant extra-solution activity, as it merely identifies the environment or tool used to perform the abstract idea and does not impose a meaningful limit on the claim.),
providing the prompt to the language model (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).);
by the first Al agent (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).)
receiving, by the first Al agent, at least one output from the assistive-technology in response to performing the determined action on the web page (This limitation recites receiving a first output, which is a well-understood, routine, and conventional activity. This step merely involves obtaining information resulting from a performed action and does not impose a meaningful limit on the abstract idea. Such data gathering is a basic computer function and amounts to insignificant extra-solution activity.);
performing, by the first Al agent, the determined action on the web page by interacting with the component in the component identified in the programmatic structural representation (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).)
updating, based on the at least one output, the programmatic structural representation to reflect a current interface state of the web page (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).);
Regarding claim 15, the rejection of claim 13 is incorporated herein. Further, claim 15 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
when performance of any action does not further the first Al agent's completion of the task to a predetermined degree, invoking an action assistance from a second Al agent to achieve the task, and wherein the second Al agent is configured to assist the first Al agent to complete the task, wherein any action includes the first action and the at least one subsequent action (This limitation recites merely providing instructions for applying the abstract idea based on a condition. The step directs that, upon determining insufficient progress, assistance is invoked from another agent, which amounts to insignificant extra-solution activity, as it merely identifies an additional generic component used to carry out the abstract idea and does not impose any meaningful limit on the claim.).
Regarding claim 16, the rejection of claim 15 is incorporated herein. Further, claim 16 recites the following abstract ideas:
determine…that an action does not further the first Al agent's completion of the task to a predetermined degree when a number of subsequent actions are continuations of a preceding action without furthering completion of the task exceeds a pre-determined threshold (This is an abstract idea of a mental process. The limitation recites determining whether progress toward a task has been made based on prior actions and a threshold. A person could observe a series of actions, recognize that the actions are not advancing completion of a task, and determine that progress has not been made once the number of such actions exceeds at threshold. This type of evaluation and judgement can be performed in the human mind, or with the aid of computational tools, and therefore falls within the mental process grouping of abstract ideas.).
The following claim elements are additional elements which, taken alone or in combination with the other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
by the first Al agent (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).),
Regarding claim 17, the rejection of claim 15 is incorporated herein. Further, claim 17 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
allocate, by the first Al agent, control of the task to the second Al agent to invoke the action assistance (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).).
Regarding claim 18, the rejection of claim 13 is incorporated herein. Further, claim 18 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
update at least one visual representation of the web page with information based on the at least one output received during the navigation-testing operation and at least one output received during a repetition of the navigation-testing operation, wherein the visual representation of the web page includes representation of components of the user interface of the web page (This limitation merely instructs that information obtained during navigation-testing operations be used to update a visual representation of the webpage, without reciting any particular technical manner of performing the update. The limitation merely uses a generic computer as a tool to apply the abstract analysis and does not impose a meaningful limitation on the judicial exception.).
Regarding claim 19, the rejection of claim 13 is incorporated herein. Further, claim 19 recites the following abstract ideas:
generate…inputs to be included in the generated prompt (This is an abstract idea of a mental process. The limitation recites selecting and organizing information to be used in formulating instructions. This type of evaluation and judgement can be performed in the human mind and therefore falls within the mental process grouping of abstract ideas.).
Regarding claim 20, the rejection of claim 19 is incorporated herein. Further, claim 20 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
wherein the inputs further include at least one of: a role description, configuration information for a particular screen reader software, information about a particular user device, a starting point, a log, a list of specific actions, instructions to use a particular navigation tool, and various execution strategies based on the task and the programmatic structural representation (This limitation recites additional information to be included in the inputs and merely describes the content of information used in performing the abstract idea. The recited items identify type of data or instructions that may be provided and do not impose any meaningful limit on the claim. Such recitation constitutes insignificant extra-solution activity because it merely specifies information used or presenting in carrying out the abstract idea.).
Regarding claim 21, the rejection of claim 13 is incorporated herein. Further, claim 21 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
iteratively testing a plurality of patterns of interaction with various components of the user interface of the web page (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).).
Regarding claim 22, the rejection of claim 13 is incorporated herein. Further, claim 22 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
wherein the accessibility result indicates whether the web page has an accessibility issue relating to the perceivability, operability, understandability, or robustness of a web page for mobile-device assistive technology use (This limitation recites merely specifying the type of issue being identified and describes categories of information associated with the abstract idea. The recitation does not impose any meaningful limit on the claim and constitutes insignificant extra-solution activity, as it merely characterizes the information being evaluated and presented.).
Regarding claim 23, the rejection of claim 13 is incorporated herein. Further, claim 23 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception:
wherein the assistive-technology is a screen-reader technology for visually-impaired users to operate web pages (This limitation amounts to insignificant extra-solution activity.).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3-13, and 15-23 are rejected under the 35 U.S.C. 103 as being unpatentable over Kumar et al., (Pub. No.: US 20250307563 A1 (Filed: 2024)) in view of Deshmukh et al., (Pub. No.: US 20210081165 A1 (Filed: 2020)) further in view of Azose et al., (Pub. No.: US 20250117573 A1 (Filed: 2023)).
Regarding claim 1, Kumar teaches the following limitation:
providing the prompt to the language model (Kumar, paragraph [0051] “The topic prompt and the request may then be provided to the language model (206).”);
However, Kumar does not teach but Kumar in view of Deshmukh teaches the following limitations:
generating a prompt, by a prompt engine, for a language model integrated with a first artificial intelligence (Al) agent, the prompt including a task for the first Al agent to complete on a web page using actions configured to mimic interactions expected of a user of an assistive-technology (Kumar, paragraph [0005] “process the request in response to a prompt to determine, from a language model, a topic prompt of a plurality of topic prompts associated with the request…to provide the topic prompt and the request to the language model, receive, from the language model and in response to the topic prompt and the request, a response to the request, and provide the response using the virtual agent.” Deshmukh, paragraph [0004] “ This disclosure contemplates a machine learning website accessibility testing tool…The tool is designed to operate in conjunction with a screen reader…The tool converts the audio output of the screen reader to simulated user interactions (for example, keystrokes)” [0029] “Human interaction simulator 140 is configured to (1) receive voice signals generated by screen reader 114, while screen reader 114 reads webpage 112, (2) convert the voice signals into simulated user interactions, and (3) execute the simulated user interactions with webpage 112.” [0031] “ This disclosure contemplates that the simulated user interactions may include any type of interactions that simulate the interactions that a user 104 may perform with webpages 112. For example, the simulated user interactions may include keystrokes that simulate a user 104 using a keyboard connected to device 106 to navigate webpage 112. As another example, the simulated user interactions may include information associated with cursor movements and/or clicks that simulate a user 104 using a mouse connected to device 106 to navigate webpage 112. As a further example, the simulated user interactions may include information that simulates one or more gestures that a user 104 may perform on display 108 to navigate webpage 112.” – Kumar teaches generating a prompt for a language model used by a virtual agent, where the prompt is used to process a request. Deshmukh teaches performing actions on a web page by converting outputs from a screen reader into simulated user interactions and executing those interactions on a web page, where the interactions simulate those performed by the user. Since a screen reader is an assistive technology, the simulated interactions correspond to actions configured to mimic interactions expected of a user of an assistive technology.)
receiving, by the first Al agent, at least one output from the assistive-technology in response to performing the determined action on the web page (Kumar, paragraph [0029] “an instance of a virtual agent 109 may be generated in response to each request for assistance received from the user 105 “ Deshmukh, paragraph [0059] “tool 102 launches screen reader 114. Screen reader 114 is configured to read and/or process the content of the webpage displayed in browser 110, in order to generate voice signals 144 that may be used by low-vision users 104 to navigate around the website…In step 406, tool 102 receives a voice signal 144 from screen reader 114…Accordingly, tool 102 may listen for a voice signal 144 associated with the login page. For example, tool 102 may listen for a voice signal 144 describing a link that leads to the login page.” – Kumar teaches a virtual agent. Deshmukh teaches that a screen reader, which is an assistive technology, generates voice signals associated with elements of a web page, and that the system receives the voice signals from the screen reader. The voice signals correspond to outputs from the assistive technology generated in response to navigation of a web page.);
However, Kumar in view of Deshmukh does not teach but Kumar in view of Deshmukh further in view of Azose teaches the following limitations:
a programmatic structural representation of the web page identifying components in a user interface of the web page (Azose, paragraph [0028] “the context is identified by analyzing the document object model (DOM) for text or metadata relevant to the text box TB1. In some implementations, the context is identified by analyzing an accessibility tree for text or metadata relevant to the text box TB1… the drafting assistant may inspect the nodes (in the DOM, the accessibility tree, or both) related to the text box TB1 for attributes describing the text box, such as a character size limit, text describing the purpose of the text box TB1, etc.” – Azose’s DOM or accessibility tree corresponds to the programmatic structural representation of the webpage, and the nodes associated with the text box TB1 identify a component in the webpage’s user interface.);
performing a navigation-testing operation, including: determining, by the language model and based on the task and the programmatic structural representation, an action for the first AI agent to perform on the webpage with respect to one of the components identified in the programmatic structural representation (Kumar, paragraph [0031] “a language model 112 represents one or more language models leveraged by the orchestration engine 102 to support operations of virtual agents, such as the virtual agent 109… the language model 112 may be implemented as an instruction-following large language model (LLM), which is trained on, and designed to reproduce, interactive and instruction-following behavior.” [0064] “The prompts 304, 306 provide examples of a set of prompts (e.g., in the prompt store 118 of FIG. 1) that each define a topic of conversation and execute actions corresponding to each determined, corresponding intent determined by the router prompt 302.” Azose, paragraph [0028] “ In some implementations, the context is identified by analyzing the document object model (DOM) for text or metadata relevant to the text box TB1. In some implementations, the context is identified by analyzing an accessibility tree for text or metadata relevant to the text box TB1… the drafting assistant may inspect the nodes (in the DOM, the accessibility tree, or both) related to the text box TB1 for attributes describing the text box… a DOM and/or an accessibility tree may be provided to the model and the model may determine the context and/or the instructions generated based on the context.” Deshmukh, paragraph [0052] “ test script 128a may include instructions indicating aspects of a website that webpage accessibility tool 102 should test for ADA compliance and/or a description of operations that user 104 should be able to perform with the website. As an example, test script 128a may indicate that the first action user 104 should be able to perform with webpage 112, after webpage 112 is loaded in browser 110, is to access a link on webpage 112 through which the user may navigate to a login page… human interaction simulator 140 may wait until it receives a voice signal 144 identifying a link on webpage 112 through which a user may access the login page, and may then generate input commands configured to simulate an attempt by a user to access the link to the login page.” – Kumar teaches an instruction-following language model that supports a virtual agent and executes actions corresponding to a determined intent. Azose teaches providing a DOM or accessibility tree to a model and inspecting nodes identifying webpage components, which corresponds to the claimed programmatic structural representation. Deshmukh teaches a navigation-testing task and generating an input command to access an identified webpage link. Accordingly, the combined teachings determine, based on the task and structural representation, a navigation action for the AI agent to perform on a component identified in the representation.);
performing, by the first Al agent, the determined action on the web page by interacting with the component identified in the programmatic structural representation (Kumar, paragraph [0064] “The prompts 304, 306 provide examples of a set of prompts (e.g., in the prompt store 118 of FIG. 1) that each define a topic of conversation and execute actions corresponding to each determined, corresponding intent determined by the router prompt 302.” Azose, paragraph [0028] “the drafting assistant may inspect the nodes (in the DOM, the accessibility tree, or both) related to the text box TB1 for attributes describing the text box, such as a character size limit, text describing the purpose of the text box TB1, etc.” Deshmukh, paragraph [0029] “Human interaction simulator 140 is configured to (1) receive voice signals generated by screen reader 114, while screen reader 114 reads webpage 112, (2) convert the voice signals into simulated user interactions, and (3) execute the simulated user interactions with webpage 112.” [0052] “ human interaction simulator 140 may wait until it receives a voice signal 144 identifying a link on webpage 112 through which a user may access the login page, and may then generate input commands configured to simulate an attempt by a user to access the link to the login page.” [0053] “In response to human interaction simulator 140 generating input commands configured to simulate the behavior of a user interacting with webpage 112, human interaction simulator 140 is configured to execute the input commands, thereby simulating the behavior of a low-vision user interacting with the webpage.” – Kumar teaches executing the action determined for the virtual agent. Azose teaches identifying a webpage component through a node in a DOM or accessibility tree. Deshmukh teaches executing the determined input command to interact with an identified webpage link. Accordingly, the combined teachings perform, through the first AI agent, the determined action on the component identified in the programmatic structural representation.);
updating, based on the at least one output, the programmatic structural representation to reflect a current interface state of the web page (Deshmukh, paragraph [0060] “ In step 408, tool 102 converts the voice signal 144 into an input command… In step 410, tool 102 executes the input command. In step 412 tool 102 monitors a behavior of browser 110 in response to tool 102 executing the input command… tool 102 determines the manner by which webpage 112 and/or any information displayed by webpage 112 changes, in response to tool 102 executing the input command.” [0054] “For example, in certain embodiments, monitoring the behavior of browser 110 includes monitoring the voice signals 144 received from screen reader 114 after human interaction simulator 140 has executed the input commands. As another example, in certain embodiments, monitoring the behavior of browser 110 includes analyzing the html code of the webpage displayed in browser 110 after human interaction simulator 140 has executed the input commands.” Azose, paragraph [0028] “In some implementations, the context is identified by analyzing the document object model (DOM) for text or metadata relevant to the text box TB1. In some implementations, the context is identified by analyzing an accessibility tree for text or metadata relevant to the text box TB1.” [0057] “At step 604, the system may obtain context identified using the web page… The context can be identified from a DOM tree. The context may be identified from an accessibility tree.” – Deshmukh teaches that the screen-reader output is converted to an input command and that, after executing the command, the system determines how the webpage changes and analyzes the resulting HTML. Azose teaches obtaining webpage context from a DOM or accessibility tree. Under BRI, obtaining the DOM/accessibility tree context after Deshmukh’s output-driven interaction corresponds to updating the programmatic structural representation to reflect the webpage’s current interface state.);
repeating the performance of the navigation-testing operation at least once, wherein, in each repetition, the language model determines a subsequent action based on at least one preceding output from the accessibility-technology and the updated programmatic structural representation (Kumar, paragraph [0076] “ the language model 112 thus receives inputs including, e.g., the router prompt 506, the user query 504, relevant state information, and relevant conversation history, and determines a “next prompt.”” [0078] “ the new prompt text, the user query 504, state information, and history information may be sent to the language model 112 for inference… the user query 504 may be directed back to the orchestration engine and routing may be repeated or a new input received until an appropriate prompt is found” Azose, paragraph [0057] “ At step 604, the system may obtain context identified using the web page… The context can be identified from a DOM tree. The context may be identified from an accessibility tree.” [0063] “ In response to the selection of the retry control the system may restart all or part of method 600... selection of retry control may generate a new prompt that tells the generative language model to “draft a different response” using the prior modified prompt and prior generated response may as context for the new prompt.” Deshmukh, paragraph [0059] “each voice signal 144 may be associated with one or more elements of the website and may provide information that may be used by users 104 to navigate to the one or more elements.” [0060] “ tool 102 converts the voice signal 144 into an input command… In step 410, tool 102 executes the input command… tool 102 determines the manner by which webpage 112 and/or any information displayed by webpage 112 changes, in response to tool 102 executing the input command.” – Kumar teaches repeating language-model processing and determining a next operation using prior state and history. Azose teaches restarting the model operation using prior response as context and obtaining a current webpage context from a DOM or accessibility tree. Deshmukh teaches using a preceding screen-reader output to determine a navigation command and monitoring the resulting webpage changes. Accordingly, the combined teachings repeat the navigation testing operation and determine a subsequent action using a preceding assistive-technology output and the refreshed DOM or accessibility tree representation of the current webpage state.); and
performing an accessibility test of the web page to determine, based on a sequence including the action and the subsequent action, corresponding outputs from the assistive-technology, and interface states reflected in the updated programmatic structural representation, an accessibility result indicating whether a sequence of the actions performed by the first Al agent reaches an interface state associated with completion of the task using the assistive-technology (Deshmukh, paragraph [0039] “ test scripts 128 may indicate the actions that user 104 should be able to perform with webpage 112 and/or the order by which elements of webpage 112 should be presented to user 104… test script 128 may indicate that when first landing on webpage 112, user 104 should be presented with the option of accessing a login page.” [0055] “Machine learning algorithm 136 may include any algorithm configured to identify ADA compliance issues associated with webpage 112 based on: (1) ADA rules 126, (2) voice signals 144 generated by screen reader 114 while reading webpage 112, (3) the behavior of browser 110 in response to human interaction simulator 140 executing an input command configured to simulate an interaction between a user and webpage 112, based on voice signals 144, (4) any information that may be provided in test scripts 128” [0057] “ compliance detector 142 may generate test results 302 indicating whether or not each such aspect is ADA compliant.” Kumar, paragraph [0021] “ one or more goal(s) of each unit of conversation may be specified for each prompt, so that each prompt may be considered complete (and a next or subsequent prompt predicted) once its corresponding goal or goals have been achieved.” Azose, paragraph [0028] “the context is identified by analyzing the document object model (DOM) for text or metadata relevant to the text box TB1. In some implementations, the context is identified by analyzing an accessibility tree for text or metadata relevant to the text box TB1.” – Deshmukh teaches testing a sequence of ordered webpage actions and determining an accessibility result based on screen-reader outputs, the simulated interactions, the resulting browser behavior, and the test-script objective, such as accessing a login page. Kumar teaches determining completion when the specified goal is achieved. Azose’s DOM or accessibility tree represents the interface states produced during the sequence. Accordingly, the combined teachings determine whether the assistive technology interaction sequence reaches the interface state associated with a completed task.).
Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, having a combination of Kumar, Deshmukh, and Azose before them, to incorporate Azose’s DOM or accessibility-tree webpage context into Kumar’s prompt orchestration virtual agent when performing the screen-reader based accessibility testing of Deshmukh. One would have been motivated to make such a combination in order to determine whether a blind or visually impaired user could reliably navigate a website and complete intended tasks using assistive technology, and to identify points at which inaccessible webpage elements prevent or hinder further navigation. This would allow website accessibility issues to be detected and corrected, thereby enabling uses who rely on assistive technology to navigate websites with greater confidence.
Regarding claim 3, Kumar in view of Deshmukh further in view of Azose teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. Kumar in view of Deshmukh further in view of Azose further teaches:
when performing any action does not further the first Al agent's completion of the task to a predetermined degree, invoking an action assistance from a second Al agent to achieve the task, wherein the second Al agent is configured to assist the first Al agent to complete the task, and wherein any action includes the first action and the subsequent action (Kumar, paragraph [0031] “Further in FIG. 1, a language model 112 represents one or more language models leveraged by the orchestration engine 102 to support operations of virtual agents, such as the virtual agent 109, generated by the virtual agent generator 108.” [0083] “ processing of the digression request 602 includes routing (using the routing prompt 506) the digression request 602 to the question answering prompt 508, which (together with the language model 112) responds with response 604 of “there is one outage.” The user 502b responds with a related query 606 of “Summarize it for me,” which is sent to the summarization prompt 510 to obtain a result 608 “ [0084] “ Upon determining that the result 608 is a final answer to the digression or otherwise determining that the digression has ended, the pre-digression state is retrieved and the user is provided with a reiteration of, and return to, the query 520” – under the broadest reasonable interpretation, a “second AI agent” includes another agent instance or process utilized by an orchestration engine to assist in completing a task. Kumar teaches that an orchestration engine supports operations of virtual agents and routes requests, including digression requests, to different prompts to obtain responses. When a digression occurs, the current task is not being furthered, and the request is routed to another prompt to obtain assistance. The different prompts correspond to different agent processes that assist in completing the task. Kumar further teaches that after digression is resolved, processing returns to the original task, thereby assisting completion of the task. Thus, when performance of an action does not further completion of the task, assistance is invoked from another agent to achieve the task.).
Regarding claim 4, Kumar in view of Deshmukh further in view of Azose teaches all the elements of claim 3, therefore is rejected for the same reasons as those presented for claim 3. Kumar in view of Deshmukh further in view of Azose further teaches:
determining, by the first Al agent, that an action does not further the first Al agent's completion of the task to a predetermined degree when a number of subsequent actions are continuations of a preceding action without furthering completion of the task exceeds a pre-determined threshold (Kumar, paragraph [0077] “the determined “next prompt” may be associated with a confidence score. If the confidence score is less than a threshold, or there is any ambiguity, the language model 112 may determine two or more most-plausible or highest-confidence prompts, and request clarification from the user 502.” – under BRI, determining an action does not further completion of a task to a predetermined degree includes determining that continued actions are not sufficiently progressing toward resolving the task based on a threshold. Kumar teaches that when a confidence score is less than a threshold, the language model determines alternative prompts indicating that the current interaction is not sufficiently progressing toward completing the task. This corresponds to determining that continued actions are not furthering completion of the task to a predetermined degree when a threshold condition is met.).
Regarding claim 5, Kumar in view of Deshmukh further in view of Azose teaches all the elements of claim 3, therefore is rejected for the same reasons as those presented for claim 3. Kumar in view of Deshmukh further in view of Azose further teaches:
allocating, by the first Al agent, control of the task to the second Al agent to invoke the action assistance (Kumar, paragraph [0031] “a language model 112 represents one or more language models leveraged by the orchestration engine 102 to support operations of virtual agents, such as the virtual agent 109, generated by the virtual agent generator 108.” [0083] “processing of the digression request 602 includes routing (using the routing prompt 506) the digression request 602 to the question answering prompt 508, which (together with the language model 112) responds with response 604 of “there is one outage.” The user 502b responds with a related query 606 of “Summarize it for me,” which is sent to the summarization prompt 510 to obtain a result” – under BRI, allocating control of a task to another agent includes routing the task to another agent or process to perform actions associated with a task. Kumar teaches routing a request to different prompts, including routing a question answering prompt and a summarization prompt, which then handle the request and generate responses. The routing of the request to another prompt corresponds to allocating control of the task to another agent to invoke assistance.).
Regarding claim 6, Kumar in view of Deshmukh further in view of Azose teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. Kumar in view of Deshmukh further in view of Azose further teaches:
updating at least one visual representation of the web page with information based on the at least one output received during the navigation-testing operation and at least one output received during a repetition of the navigation-testing operation, wherein the visual representation of the web page includes representation of components of the user interface of the web page (Deshmukh, paragraph [0058] “ Loading the website into browser 110 may include loading a webpage 112 belonging to the website into the browser. This disclosure contemplates that the website includes one or more elements, where each element corresponds to content that may be displayed by the website” [0059] “Screen reader 114 is configured to read and/or process the content of the webpage displayed in browser 110, in order to generate voice signals 144…each voice signal 144 may be associated with one or more elements of the website and may provide information that may be used by users 104 to navigate to the one or more elements.” [0060] “ For example, tool 102 determines the manner by which webpage 112 and/or any information displayed by webpage 112 changes, in response to tool 102 executing the input command.” – under BRI, the webpage displayed in the browser constitutes a visual representation including the webpage’s user interface component. Deshmukh teaches that screen-reader outputs are associated with those components and are used to produce navigation commands after which the system determines how the displayed webpage changes. As mapped in claim 1, the navigation operation is repeated using successive screen-reader outputs. Thus, the displayed webpage is updated based on the output received during the initial navigation-testing operation and the output received during the repeated operation.).
Regarding claim 7, Kumar in view of Deshmukh further in view of Azose teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. Kumar in view of Deshmukh further in view of Azose teaches:
generating, by the prompt engine, inputs to be included in the generated prompt (Kumar, paragraph [0076] “the language model 112 thus receives inputs including, e.g., the router prompt 506, the user query 504, relevant state information, and relevant conversation history, and determines a “next prompt.”” – teaches that inputs, including a router prompt, user query, state information, and conversation history, are provided to the language model. These inputs correspond to inputs generated by a prompt engine to be included in a generated prompt.).
Regarding claim 8, Kumar in view of Deshmukh teaches all the elements of claim 7, therefore is rejected for the same reasons as those presented for claim 7. Kumar in view of Deshmukh further teaches:
wherein the inputs further include at least one of: a role description, configuration information for a particular screen reader software, information about a particular user device, a starting point, a log, a list of specific actions, instructions to use a particular navigation tool, and various execution strategies based on the task and the programmatic structural representation (Kumar, paragraph [0020] “a subsequent prompt is predicted by a current prompt and corresponding LLM response.” [0021] “one or more goal(s) of each unit of conversation may be specified for each prompt, so that each prompt may be considered complete (and a next or subsequent prompt predicted) once its corresponding goal or goals have been achieved.” [0076] “the language model 112 thus receives inputs including, e.g., the router prompt 506, the user query 504, relevant state information, and relevant conversation history” Deshmukh, paragraph [0037] “Error log 124 may include a list of ADA rule violations and/or compliance issues associated with webpage” [0039] “Test scripts 128 may include information that may be used by webpage accessibility tool 102 to determine whether webpage 112 is ADA compliant.” [0059] “For example, each voice signal 144 may be associated with one or more elements of the website and may provide information that may be used by users 104 to navigate to the one or more elements…test script 128 may indicate that the first action that a user should be able to perform when visiting the website is to access a login page.” Azose, paragraph [0028] “ the context is identified by analyzing the document object model (DOM) for text or metadata relevant to the text box TB1. In some implementations, the context is identified by analyzing an accessibility tree for text or metadata relevant to the text box TB1…a DOM and/or an accessibility tree may be provided to the model and the model may determine the context and/or the instructions generated based on the context.” – Kumar teaches language-model inputs including a user query, state information, and conversation history, and teaches determining subsequent processing based on those inputs. Deshmukh teaches test scripts identifying a starting point and specific actions or instructions for navigating a webpage. Azose teaches providing a DOM or accessibility tree to a generative language model and determining instructions based on that structural webpage context. Under BRI, the DOM or accessibility tree corresponds to the claimed programmatic structural representation, and the resulting instructions, together with Kumar’s task-based processing, correspond to execution strategies based on the task and the programmatic structural representation.),
Regarding claim 9, Kumar in view of Deshmukh further in view of Azose teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. Kumar in view of Deshmukh further teaches:
iteratively testing a plurality of patterns of interaction with various components of the user interface of the web page (Deshmukh, paragraph [0026] “webpage accessibility tool 102 implements machine learning trainer 138 to train machine learning algorithm 136 to detect ADA compliance issues with webpages 112” [0029] “ Human interaction simulator 140 is configured to (1) receive voice signals generated by screen reader 114, while screen reader 114 reads webpage 112, (2) convert the voice signals into simulated user interactions, and (3) execute the simulated user interactions with webpage 112.” [0031] “the simulated user interactions may include keystrokes that simulate a user 104 using a keyboard connected to device 106 to navigate webpage 112. As another example, the simulated user interactions may include information associated with cursor movements and/or clicks that simulate a user 104 using a mouse connected to device 106 to navigate webpage 112. As a further example, the simulated user interactions may include information that simulates one or more gestures that a user 104 may perform on display 108 to navigate webpage 112.” [0060] “ tool 102 determines the manner by which webpage 112 and/or any information displayed by webpage 112 changes, in response to tool 102 executing the input command.” – Deshmukh teaches executing simulated user interactions with a web page, including keystrokes, cursor movements, and gestures, which corresponds to different patterns of interactions with components of the user interface and web page. Deshmukh further teaches monitoring behavior of the web page in response to the interactions and detecting accessibility issues for web pages. The execution of multiple interactions and evaluation of resulting behavior of the web page corresponds to iteratively testing a plurality of patterns of interactions with components of the user interface of the web page.).
Regarding claim 10, Kumar in view of Deshmukh teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. Kumar in view of Deshmukh further teaches:
wherein the accessibility result indicates whether the webpage has an accessibility issue relating to perceivability, operability, understandability, or robustness of a web page for mobile-device assistive technology use (Deshmukh, paragraph [0020] “device 106 may be a smart phone, a tablet, a laptop, an automated assistant, or any other suitable device that includes a display as an integral part of the device.” [0021] “Screen reader 114 is designed for low-vision users 104, and may be used to assist such users in navigating the information displayed on browser 110.” [0038] “ADA rules 126 may include Web Content Accessibility Guidelines (WCAG). For example, ADA rules 126 may include rules associated with WCAG 2.0 Level A, WCAG 2.0 Level AA, and/or WCAG 2.0 Level AAA…ADA rules 126 may include rules specifying that the html code associated with webpage 112 include alternative text for any images displayed on webpage 112; webpage 112 present content in a meaningful order; webpage 112 does not rely on color alone to convey information; all content and functions of webpage 112 are accessible by keyboard” [0057] “compliance detector 142 may generate test results 302 indicating whether or not each such aspect is ADA compliant.” – Deshmukh teaches generating test results indicating whether aspects of a webpage comply with ADA rules incorporating the Web Content Accessibility Guidelines. The test results correspond to the claimed accessibility result and indicate whether the webpage has an accessibility issue relating to perceivability, operability, understandability, or robustness. Deshmukh further teaches accessing the webpage using devices such as smart phones and tablets and using a screen reader to assist low-vision users, thereby meeting the mobile-device assistive-technology limitation.).
Regarding claim 11, Kumar in view of Deshmukh further in view of Azose teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. Kumar in view of Deshmukh further in view of Azose further teaches:
wherein the assistive-technology is a screen-reader technology for visually-impaired users to operate web pages (Deshmukh, paragraph [0021] Screen reader 114 is designed for low-vision users 104, and may be used to assist such users in navigating the information displayed on browser 110...In addition to reading the text of webpage 112, as displayed on display 108, screen reader 120 may also read alternative text, descriptive headings, and/or descriptive link text that is included in the html code of webpage 112.”).
Regarding claim 12, Kumar teaches the following limitation:
A non-transitory computer-readable medium storing a set of instructions for performing accessibility testing in web pages for assistive-technology users, the set of instructions comprising: one or more instructions that, when executed by one or more processing circuitries of a device, cause the device to (Kumar, paragraph [0046] “the orchestration engine 102 is illustrated as being implemented using at least one computing device 128, including at least one processor 130, and a non-transitory computer-readable storage medium 132. That is, the non-transitory computer-readable storage medium 132 may store instructions that, when executed by the at least one processor 130, cause the at least one computing device 128 to provide the functionalities of the orchestration engine 102 and related functionalities.”):
providing the prompt to the language model (Kumar, paragraph [0051] “The topic prompt and the request may then be provided to the language model (206).”);
However, Kumar does not teach but Kumar in view of Deshmukh teaches the following limitations:
generate a prompt, by a prompt engine, for a language model integrated with a first artificial intelligence (Al) agent, the prompt including a task for the first Al agent to complete on a web page using actions configured to mimic interactions expected of a user of an assistive-technology (Kumar, paragraph [0005] “process the request in response to a prompt to determine, from a language model, a topic prompt of a plurality of topic prompts associated with the request…to provide the topic prompt and the request to the language model, receive, from the language model and in response to the topic prompt and the request, a response to the request, and provide the response using the virtual agent.” Deshmukh, paragraph [0004] “ This disclosure contemplates a machine learning website accessibility testing tool…The tool is designed to operate in conjunction with a screen reader…The tool converts the audio output of the screen reader to simulated user interactions (for example, keystrokes)” [0029] “Human interaction simulator 140 is configured to (1) receive voice signals generated by screen reader 114, while screen reader 114 reads webpage 112, (2) convert the voice signals into simulated user interactions, and (3) execute the simulated user interactions with webpage 112.” [0031] “ This disclosure contemplates that the simulated user interactions may include any type of interactions that simulate the interactions that a user 104 may perform with webpages 112. For example, the simulated user interactions may include keystrokes that simulate a user 104 using a keyboard connected to device 106 to navigate webpage 112. As another example, the simulated user interactions may include information associated with cursor movements and/or clicks that simulate a user 104 using a mouse connected to device 106 to navigate webpage 112. As a further example, the simulated user interactions may include information that simulates one or more gestures that a user 104 may perform on display 108 to navigate webpage 112.” – Kumar teaches generating a prompt for a language model used by a virtual agent, where the prompt is used to process a request. Deshmukh teaches performing actions on a web page by converting outputs from a screen reader into simulated user interactions and executing those interactions on a web page, where the interactions simulate those performed by the user. Since a screen reader is an assistive technology, the simulated interactions correspond to actions configured to mimic interactions expected of a user of an assistive technology.)
receiving, by the first Al agent, at least one output from the assistive-technology in response to performing the determined action on the web page (Kumar, paragraph [0029] “an instance of a virtual agent 109 may be generated in response to each request for assistance received from the user 105 “ Deshmukh, paragraph [0059] “tool 102 launches screen reader 114. Screen reader 114 is configured to read and/or process the content of the webpage displayed in browser 110, in order to generate voice signals 144 that may be used by low-vision users 104 to navigate around the website…In step 406, tool 102 receives a voice signal 144 from screen reader 114…Accordingly, tool 102 may listen for a voice signal 144 associated with the login page. For example, tool 102 may listen for a voice signal 144 describing a link that leads to the login page.” – Kumar teaches a virtual agent. Deshmukh teaches that a screen reader, which is an assistive technology, generates voice signals associated with elements of a web page, and that the system receives the voice signals from the screen reader. The voice signals correspond to outputs from the assistive technology generated in response to navigation of a web page.);
However, Kumar in view of Deshmukh does not teach but Kumar in view of Deshmukh further in view of Azose teaches the following limitations:
a programmatic structural representation of the web page identifying components in a user interface of the web page (Azose, paragraph [0028] “the context is identified by analyzing the document object model (DOM) for text or metadata relevant to the text box TB1. In some implementations, the context is identified by analyzing an accessibility tree for text or metadata relevant to the text box TB1… the drafting assistant may inspect the nodes (in the DOM, the accessibility tree, or both) related to the text box TB1 for attributes describing the text box, such as a character size limit, text describing the purpose of the text box TB1, etc.” – Azose’s DOM or accessibility tree corresponds to the programmatic structural representation of the webpage, and the nodes associated with the text box TB1 identify a component in the webpage’s user interface.);
perform a navigation-testing operation, including: determining, by the language model and based on the task and the programmatic structural representation, an action for the first AI agent to perform on the webpage with respect to one of the components identified in the programmatic structural representation (Kumar, paragraph [0031] “a language model 112 represents one or more language models leveraged by the orchestration engine 102 to support operations of virtual agents, such as the virtual agent 109… the language model 112 may be implemented as an instruction-following large language model (LLM), which is trained on, and designed to reproduce, interactive and instruction-following behavior.” [0064] “The prompts 304, 306 provide examples of a set of prompts (e.g., in the prompt store 118 of FIG. 1) that each define a topic of conversation and execute actions corresponding to each determined, corresponding intent determined by the router prompt 302.” Azose, paragraph [0028] “ In some implementations, the context is identified by analyzing the document object model (DOM) for text or metadata relevant to the text box TB1. In some implementations, the context is identified by analyzing an accessibility tree for text or metadata relevant to the text box TB1… the drafting assistant may inspect the nodes (in the DOM, the accessibility tree, or both) related to the text box TB1 for attributes describing the text box… a DOM and/or an accessibility tree may be provided to the model and the model may determine the context and/or the instructions generated based on the context.” Deshmukh, paragraph [0052] “ test script 128a may include instructions indicating aspects of a website that webpage accessibility tool 102 should test for ADA compliance and/or a description of operations that user 104 should be able to perform with the website. As an example, test script 128a may indicate that the first action user 104 should be able to perform with webpage 112, after webpage 112 is loaded in browser 110, is to access a link on webpage 112 through which the user may navigate to a login page… human interaction simulator 140 may wait until it receives a voice signal 144 identifying a link on webpage 112 through which a user may access the login page, and may then generate input commands configured to simulate an attempt by a user to access the link to the login page.” – Kumar teaches an instruction-following language model that supports a virtual agent and executes actions corresponding to a determined intent. Azose teaches providing a DOM or accessibility tree to a model and inspecting nodes identifying webpage components, which corresponds to the claimed programmatic structural representation. Deshmukh teaches a navigation-testing task and generating an input command to access an identified webpage link. Accordingly, the combined teachings determine, based on the task and structural representation, a navigation action for the AI agent to perform on a component identified in the representation.);
performing, by the first Al agent, the determined action on the web page by interacting with the component identified in the programmatic structural representation (Kumar, paragraph [0064] “The prompts 304, 306 provide examples of a set of prompts (e.g., in the prompt store 118 of FIG. 1) that each define a topic of conversation and execute actions corresponding to each determined, corresponding intent determined by the router prompt 302.” Azose, paragraph [0028] “the drafting assistant may inspect the nodes (in the DOM, the accessibility tree, or both) related to the text box TB1 for attributes describing the text box, such as a character size limit, text describing the purpose of the text box TB1, etc.” Deshmukh, paragraph [0029] “Human interaction simulator 140 is configured to (1) receive voice signals generated by screen reader 114, while screen reader 114 reads webpage 112, (2) convert the voice signals into simulated user interactions, and (3) execute the simulated user interactions with webpage 112.” [0052] “ human interaction simulator 140 may wait until it receives a voice signal 144 identifying a link on webpage 112 through which a user may access the login page, and may then generate input commands configured to simulate an attempt by a user to access the link to the login page.” [0053] “In response to human interaction simulator 140 generating input commands configured to simulate the behavior of a user interacting with webpage 112, human interaction simulator 140 is configured to execute the input commands, thereby simulating the behavior of a low-vision user interacting with the webpage.” – Kumar teaches executing the action determined for the virtual agent. Azose teaches identifying a webpage component through a node in a DOM or accessibility tree. Deshmukh teaches executing the determined input command to interact with an identified webpage link. Accordingly, the combined teachings perform, through the first AI agent, the determined action on the component identified in the programmatic structural representation.);
updating, based on the at least one output, the programmatic structural representation to reflect a current interface state of the web page (Deshmukh, paragraph [0060] “ In step 408, tool 102 converts the voice signal 144 into an input command… In step 410, tool 102 executes the input command. In step 412 tool 102 monitors a behavior of browser 110 in response to tool 102 executing the input command… tool 102 determines the manner by which webpage 112 and/or any information displayed by webpage 112 changes, in response to tool 102 executing the input command.” [0054] “For example, in certain embodiments, monitoring the behavior of browser 110 includes monitoring the voice signals 144 received from screen reader 114 after human interaction simulator 140 has executed the input commands. As another example, in certain embodiments, monitoring the behavior of browser 110 includes analyzing the html code of the webpage displayed in browser 110 after human interaction simulator 140 has executed the input commands.” Azose, paragraph [0028] “In some implementations, the context is identified by analyzing the document object model (DOM) for text or metadata relevant to the text box TB1. In some implementations, the context is identified by analyzing an accessibility tree for text or metadata relevant to the text box TB1.” [0057] “At step 604, the system may obtain context identified using the web page… The context can be identified from a DOM tree. The context may be identified from an accessibility tree.” – Deshmukh teaches that the screen-reader output is converted to an input command and that, after executing the command, the system determines how the webpage changes and analyzes the resulting HTML. Azose teaches obtaining webpage context from a DOM or accessibility tree. Under BRI, obtaining the DOM/accessibility tree context after Deshmukh’s output-driven interaction corresponds to updating the programmatic structural representation to reflect the webpage’s current interface state.);
repeat the performance of the navigation-testing operation at least once, wherein, in each repetition, the language model determines a subsequent action based on at least one preceding output from the accessibility-technology and the updated programmatic structural representation (Kumar, paragraph [0076] “ the language model 112 thus receives inputs including, e.g., the router prompt 506, the user query 504, relevant state information, and relevant conversation history, and determines a “next prompt.”” [0078] “ the new prompt text, the user query 504, state information, and history information may be sent to the language model 112 for inference… the user query 504 may be directed back to the orchestration engine and routing may be repeated or a new input received until an appropriate prompt is found” Azose, paragraph [0057] “ At step 604, the system may obtain context identified using the web page… The context can be identified from a DOM tree. The context may be identified from an accessibility tree.” [0063] “ In response to the selection of the retry control the system may restart all or part of method 600... selection of retry control may generate a new prompt that tells the generative language model to “draft a different response” using the prior modified prompt and prior generated response may as context for the new prompt.” Deshmukh, paragraph [0059] “each voice signal 144 may be associated with one or more elements of the website and may provide information that may be used by users 104 to navigate to the one or more elements.” [0060] “ tool 102 converts the voice signal 144 into an input command… In step 410, tool 102 executes the input command… tool 102 determines the manner by which webpage 112 and/or any information displayed by webpage 112 changes, in response to tool 102 executing the input command.” – Kumar teaches repeating language-model processing and determining a next operation using prior state and history. Azose teaches restarting the model operation using prior response as context and obtaining a current webpage context from a DOM or accessibility tree. Deshmukh teaches using a preceding screen-reader output to determine a navigation command and monitoring the resulting webpage changes. Accordingly, the combined teachings repeat the navigation testing operation and determine a subsequent action using a preceding assistive-technology output and the refreshed DOM or accessibility tree representation of the current webpage state.); and
perform an accessibility test of the web page to determine, based on a sequence including the action and the subsequent action, corresponding outputs from the assistive-technology, and interface states reflected in the updated programmatic structural representation, an accessibility result indicating whether a sequence of the actions performed by the first Al agent reaches an interface state associated with completion of the task using the assistive-technology (Deshmukh, paragraph [0039] “ test scripts 128 may indicate the actions that user 104 should be able to perform with webpage 112 and/or the order by which elements of webpage 112 should be presented to user 104… test script 128 may indicate that when first landing on webpage 112, user 104 should be presented with the option of accessing a login page.” [0055] “Machine learning algorithm 136 may include any algorithm configured to identify ADA compliance issues associated with webpage 112 based on: (1) ADA rules 126, (2) voice signals 144 generated by screen reader 114 while reading webpage 112, (3) the behavior of browser 110 in response to human interaction simulator 140 executing an input command configured to simulate an interaction between a user and webpage 112, based on voice signals 144, (4) any information that may be provided in test scripts 128” [0057] “ compliance detector 142 may generate test results 302 indicating whether or not each such aspect is ADA compliant.” Kumar, paragraph [0021] “ one or more goal(s) of each unit of conversation may be specified for each prompt, so that each prompt may be considered complete (and a next or subsequent prompt predicted) once its corresponding goal or goals have been achieved.” Azose, paragraph [0028] “the context is identified by analyzing the document object model (DOM) for text or metadata relevant to the text box TB1. In some implementations, the context is identified by analyzing an accessibility tree for text or metadata relevant to the text box TB1.” – Deshmukh teaches testing a sequence of ordered webpage actions and determining an accessibility result based on screen-reader outputs, the simulated interactions, the resulting browser behavior, and the test-script objective, such as accessing a login page. Kumar teaches determining completion when the specified goal is achieved. Azose’s DOM or accessibility tree represents the interface states produced during the sequence. Accordingly, the combined teachings determine whether the assistive technology interaction sequence reaches the interface state associated with a completed task.).
Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, having a combination of Kumar, Deshmukh, and Azose before them, to incorporate Azose’s DOM or accessibility-tree webpage context into Kumar’s prompt orchestration virtual agent when performing the screen-reader based accessibility testing of Deshmukh. One would have been motivated to make such a combination in order to determine whether a blind or visually impaired user could reliably navigate a website and complete intended tasks using assistive technology, and to identify points at which inaccessible webpage elements prevent or hinder further navigation. This would allow website accessibility issues to be detected and corrected, thereby enabling uses who rely on assistive technology to navigate websites with greater confidence.
Regarding claim 13, Kumar teaches the following limitation:
A system for performing accessibility testing in web pages for users of an assistive-technology, comprising: a processing circuitry; a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system to: (Kumar, paragraph [0094] “Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. Elements of a computer may include at least one processor for executing instructions and one or more memory devices for storing instructions and data.”):
providing the prompt to the language model (Kumar, paragraph [0051] “The topic prompt and the request may then be provided to the language model (206).”);
However, Kumar does not teach but Kumar in view of Deshmukh teaches the following limitations:
generate a prompt, by a prompt engine, for a language model integrated with a first artificial intelligence (Al) agent, the prompt including a task for the first Al agent to complete on a web page using actions configured to mimic interactions expected of a user of an assistive-technology (Kumar, paragraph [0005] “process the request in response to a prompt to determine, from a language model, a topic prompt of a plurality of topic prompts associated with the request…to provide the topic prompt and the request to the language model, receive, from the language model and in response to the topic prompt and the request, a response to the request, and provide the response using the virtual agent.” Deshmukh, paragraph [0004] “ This disclosure contemplates a machine learning website accessibility testing tool…The tool is designed to operate in conjunction with a screen reader…The tool converts the audio output of the screen reader to simulated user interactions (for example, keystrokes)” [0029] “Human interaction simulator 140 is configured to (1) receive voice signals generated by screen reader 114, while screen reader 114 reads webpage 112, (2) convert the voice signals into simulated user interactions, and (3) execute the simulated user interactions with webpage 112.” [0031] “ This disclosure contemplates that the simulated user interactions may include any type of interactions that simulate the interactions that a user 104 may perform with webpages 112. For example, the simulated user interactions may include keystrokes that simulate a user 104 using a keyboard connected to device 106 to navigate webpage 112. As another example, the simulated user interactions may include information associated with cursor movements and/or clicks that simulate a user 104 using a mouse connected to device 106 to navigate webpage 112. As a further example, the simulated user interactions may include information that simulates one or more gestures that a user 104 may perform on display 108 to navigate webpage 112.” – Kumar teaches generating a prompt for a language model used by a virtual agent, where the prompt is used to process a request. Deshmukh teaches performing actions on a web page by converting outputs from a screen reader into simulated user interactions and executing those interactions on a web page, where the interactions simulate those performed by the user. Since a screen reader is an assistive technology, the simulated interactions correspond to actions configured to mimic interactions expected of a user of an assistive technology.)
receiving, by the first Al agent, at least one output from the assistive-technology in response to performing the determined action on the web page (Kumar, paragraph [0029] “an instance of a virtual agent 109 may be generated in response to each request for assistance received from the user 105 “ Deshmukh, paragraph [0059] “tool 102 launches screen reader 114. Screen reader 114 is configured to read and/or process the content of the webpage displayed in browser 110, in order to generate voice signals 144 that may be used by low-vision users 104 to navigate around the website…In step 406, tool 102 receives a voice signal 144 from screen reader 114…Accordingly, tool 102 may listen for a voice signal 144 associated with the login page. For example, tool 102 may listen for a voice signal 144 describing a link that leads to the login page.” – Kumar teaches a virtual agent. Deshmukh teaches that a screen reader, which is an assistive technology, generates voice signals associated with elements of a web page, and that the system receives the voice signals from the screen reader. The voice signals correspond to outputs from the assistive technology generated in response to navigation of a web page.);
However, Kumar in view of Deshmukh does not teach but Kumar in view of Deshmukh further in view of Azose teaches the following limitations:
a programmatic structural representation of the web page identifying components in a user interface of the web page (Azose, paragraph [0028] “the context is identified by analyzing the document object model (DOM) for text or metadata relevant to the text box TB1. In some implementations, the context is identified by analyzing an accessibility tree for text or metadata relevant to the text box TB1… the drafting assistant may inspect the nodes (in the DOM, the accessibility tree, or both) related to the text box TB1 for attributes describing the text box, such as a character size limit, text describing the purpose of the text box TB1, etc.” – Azose’s DOM or accessibility tree corresponds to the programmatic structural representation of the webpage, and the nodes associated with the text box TB1 identify a component in the webpage’s user interface.);
perform a navigation-testing operation, including: determining, by the language model and based on the task and the programmatic structural representation, an action for the first AI agent to perform on the webpage with respect to one of the components identified in the programmatic structural representation (Kumar, paragraph [0031] “a language model 112 represents one or more language models leveraged by the orchestration engine 102 to support operations of virtual agents, such as the virtual agent 109… the language model 112 may be implemented as an instruction-following large language model (LLM), which is trained on, and designed to reproduce, interactive and instruction-following behavior.” [0064] “The prompts 304, 306 provide examples of a set of prompts (e.g., in the prompt store 118 of FIG. 1) that each define a topic of conversation and execute actions corresponding to each determined, corresponding intent determined by the router prompt 302.” Azose, paragraph [0028] “ In some implementations, the context is identified by analyzing the document object model (DOM) for text or metadata relevant to the text box TB1. In some implementations, the context is identified by analyzing an accessibility tree for text or metadata relevant to the text box TB1… the drafting assistant may inspect the nodes (in the DOM, the accessibility tree, or both) related to the text box TB1 for attributes describing the text box… a DOM and/or an accessibility tree may be provided to the model and the model may determine the context and/or the instructions generated based on the context.” Deshmukh, paragraph [0052] “ test script 128a may include instructions indicating aspects of a website that webpage accessibility tool 102 should test for ADA compliance and/or a description of operations that user 104 should be able to perform with the website. As an example, test script 128a may indicate that the first action user 104 should be able to perform with webpage 112, after webpage 112 is loaded in browser 110, is to access a link on webpage 112 through which the user may navigate to a login page… human interaction simulator 140 may wait until it receives a voice signal 144 identifying a link on webpage 112 through which a user may access the login page, and may then generate input commands configured to simulate an attempt by a user to access the link to the login page.” – Kumar teaches an instruction-following language model that supports a virtual agent and executes actions corresponding to a determined intent. Azose teaches providing a DOM or accessibility tree to a model and inspecting nodes identifying webpage components, which corresponds to the claimed programmatic structural representation. Deshmukh teaches a navigation-testing task and generating an input command to access an identified webpage link. Accordingly, the combined teachings determine, based on the task and structural representation, a navigation action for the AI agent to perform on a component identified in the representation.);
performing, by the first Al agent, the determined action on the web page by interacting with the component identified in the programmatic structural representation (Kumar, paragraph [0064] “The prompts 304, 306 provide examples of a set of prompts (e.g., in the prompt store 118 of FIG. 1) that each define a topic of conversation and execute actions corresponding to each determined, corresponding intent determined by the router prompt 302.” Azose, paragraph [0028] “the drafting assistant may inspect the nodes (in the DOM, the accessibility tree, or both) related to the text box TB1 for attributes describing the text box, such as a character size limit, text describing the purpose of the text box TB1, etc.” Deshmukh, paragraph [0029] “Human interaction simulator 140 is configured to (1) receive voice signals generated by screen reader 114, while screen reader 114 reads webpage 112, (2) convert the voice signals into simulated user interactions, and (3) execute the simulated user interactions with webpage 112.” [0052] “ human interaction simulator 140 may wait until it receives a voice signal 144 identifying a link on webpage 112 through which a user may access the login page, and may then generate input commands configured to simulate an attempt by a user to access the link to the login page.” [0053] “In response to human interaction simulator 140 generating input commands configured to simulate the behavior of a user interacting with webpage 112, human interaction simulator 140 is configured to execute the input commands, thereby simulating the behavior of a low-vision user interacting with the webpage.” – Kumar teaches executing the action determined for the virtual agent. Azose teaches identifying a webpage component through a node in a DOM or accessibility tree. Deshmukh teaches executing the determined input command to interact with an identified webpage link. Accordingly, the combined teachings perform, through the first AI agent, the determined action on the component identified in the programmatic structural representation.);
updating, based on the at least one output, the programmatic structural representation to reflect a current interface state of the web page (Deshmukh, paragraph [0060] “ In step 408, tool 102 converts the voice signal 144 into an input command… In step 410, tool 102 executes the input command. In step 412 tool 102 monitors a behavior of browser 110 in response to tool 102 executing the input command… tool 102 determines the manner by which webpage 112 and/or any information displayed by webpage 112 changes, in response to tool 102 executing the input command.” [0054] “For example, in certain embodiments, monitoring the behavior of browser 110 includes monitoring the voice signals 144 received from screen reader 114 after human interaction simulator 140 has executed the input commands. As another example, in certain embodiments, monitoring the behavior of browser 110 includes analyzing the html code of the webpage displayed in browser 110 after human interaction simulator 140 has executed the input commands.” Azose, paragraph [0028] “In some implementations, the context is identified by analyzing the document object model (DOM) for text or metadata relevant to the text box TB1. In some implementations, the context is identified by analyzing an accessibility tree for text or metadata relevant to the text box TB1.” [0057] “At step 604, the system may obtain context identified using the web page… The context can be identified from a DOM tree. The context may be identified from an accessibility tree.” – Deshmukh teaches that the screen-reader output is converted to an input command and that, after executing the command, the system determines how the webpage changes and analyzes the resulting HTML. Azose teaches obtaining webpage context from a DOM or accessibility tree. Under BRI, obtaining the DOM/accessibility tree context after Deshmukh’s output-driven interaction corresponds to updating the programmatic structural representation to reflect the webpage’s current interface state.);
repeat the performance of the navigation-testing operation at least once, wherein, in each repetition, the language model determines a subsequent action based on at least one preceding output from the accessibility-technology and the updated programmatic structural representation (Kumar, paragraph [0076] “ the language model 112 thus receives inputs including, e.g., the router prompt 506, the user query 504, relevant state information, and relevant conversation history, and determines a “next prompt.”” [0078] “ the new prompt text, the user query 504, state information, and history information may be sent to the language model 112 for inference… the user query 504 may be directed back to the orchestration engine and routing may be repeated or a new input received until an appropriate prompt is found” Azose, paragraph [0057] “ At step 604, the system may obtain context identified using the web page… The context can be identified from a DOM tree. The context may be identified from an accessibility tree.” [0063] “ In response to the selection of the retry control the system may restart all or part of method 600... selection of retry control may generate a new prompt that tells the generative language model to “draft a different response” using the prior modified prompt and prior generated response may as context for the new prompt.” Deshmukh, paragraph [0059] “each voice signal 144 may be associated with one or more elements of the website and may provide information that may be used by users 104 to navigate to the one or more elements.” [0060] “ tool 102 converts the voice signal 144 into an input command… In step 410, tool 102 executes the input command… tool 102 determines the manner by which webpage 112 and/or any information displayed by webpage 112 changes, in response to tool 102 executing the input command.” – Kumar teaches repeating language-model processing and determining a next operation using prior state and history. Azose teaches restarting the model operation using prior response as context and obtaining a current webpage context from a DOM or accessibility tree. Deshmukh teaches using a preceding screen-reader output to determine a navigation command and monitoring the resulting webpage changes. Accordingly, the combined teachings repeat the navigation testing operation and determine a subsequent action using a preceding assistive-technology output and the refreshed DOM or accessibility tree representation of the current webpage state.); and
perform an accessibility test of the web page to determine, based on a sequence including the action and the subsequent action, corresponding outputs from the assistive-technology, and interface states reflected in the updated programmatic structural representation, an accessibility result indicating whether a sequence of the actions performed by the first Al agent reaches an interface state associated with completion of the task using the assistive-technology (Deshmukh, paragraph [0039] “ test scripts 128 may indicate the actions that user 104 should be able to perform with webpage 112 and/or the order by which elements of webpage 112 should be presented to user 104… test script 128 may indicate that when first landing on webpage 112, user 104 should be presented with the option of accessing a login page.” [0055] “Machine learning algorithm 136 may include any algorithm configured to identify ADA compliance issues associated with webpage 112 based on: (1) ADA rules 126, (2) voice signals 144 generated by screen reader 114 while reading webpage 112, (3) the behavior of browser 110 in response to human interaction simulator 140 executing an input command configured to simulate an interaction between a user and webpage 112, based on voice signals 144, (4) any information that may be provided in test scripts 128” [0057] “ compliance detector 142 may generate test results 302 indicating whether or not each such aspect is ADA compliant.” Kumar, paragraph [0021] “ one or more goal(s) of each unit of conversation may be specified for each prompt, so that each prompt may be considered complete (and a next or subsequent prompt predicted) once its corresponding goal or goals have been achieved.” Azose, paragraph [0028] “the context is identified by analyzing the document object model (DOM) for text or metadata relevant to the text box TB1. In some implementations, the context is identified by analyzing an accessibility tree for text or metadata relevant to the text box TB1.” – Deshmukh teaches testing a sequence of ordered webpage actions and determining an accessibility result based on screen-reader outputs, the simulated interactions, the resulting browser behavior, and the test-script objective, such as accessing a login page. Kumar teaches determining completion when the specified goal is achieved. Azose’s DOM or accessibility tree represents the interface states produced during the sequence. Accordingly, the combined teachings determine whether the assistive technology interaction sequence reaches the interface state associated with a completed task.).
Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, having a combination of Kumar, Deshmukh, and Azose before them, to incorporate Azose’s DOM or accessibility-tree webpage context into Kumar’s prompt orchestration virtual agent when performing the screen-reader based accessibility testing of Deshmukh. One would have been motivated to make such a combination in order to determine whether a blind or visually impaired user could reliably navigate a website and complete intended tasks using assistive technology, and to identify points at which inaccessible webpage elements prevent or hinder further navigation. This would allow website accessibility issues to be detected and corrected, thereby enabling uses who rely on assistive technology to navigate websites with greater confidence.
Regarding claim 15, Kumar in view of Deshmukh further in view of Azose teaches all the elements of claim 13, therefore is rejected for the same reasons as those presented for claim 13. The claim recites similar limitations corresponding to claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rationale.
Regarding claim 16, Kumar in view of Deshmukh further in view of Azose teaches all the elements of claim 15, therefore is rejected for the same reasons as those presented for claim 15. The claim recites similar limitations corresponding to claim 4 and is rejected for similar reasons as claim 4 using similar teachings and rationale.
Regarding claim 17, Kumar in view of Deshmukh further in view of Azose teaches all the elements of claim 15, therefore is rejected for the same reasons as those presented for claim 15. The claim recites similar limitations corresponding to claim 5 and is rejected for similar reasons as claim 5 using similar teachings and rationale.
Regarding claim 18, Kumar in view of Deshmukh further in view of Azose teaches all the elements of claim 13, therefore is rejected for the same reasons as those presented for claim 13. The claim recites similar limitations corresponding to claim 6 and is rejected for similar reasons as claim 6 using similar teachings and rationale.
Regarding claim 19, Kumar in view of Deshmukh further in view of Azose teaches all the elements of claim 13, therefore is rejected for the same reasons as those presented for claim 13. The claim recites similar limitations corresponding to claim 7 and is rejected for similar reasons as claim 7 using similar teachings and rationale.
Regarding claim 20, Kumar in view of Deshmukh further in view of Azose teaches all the elements of claim 19, therefore is rejected for the same reasons as those presented for claim 19. The claim recites similar limitations corresponding to claim 8 and is rejected for similar reasons as claim 8 using similar teachings and rationale.
Regarding claim 21, Kumar in view of Deshmukh further in view of Azose teaches all the elements of claim 13, therefore is rejected for the same reasons as those presented for claim 13. The claim recites similar limitations corresponding to claim 9 and is rejected for similar reasons as claim 9 using similar teachings and rationale.
Regarding claim 22, Kumar in view of Deshmukh further in view of Azose teaches all the elements of claim 13, therefore is rejected for the same reasons as those presented for claim 13. The claim recites similar limitations corresponding to claim 10 and is rejected for similar reasons as claim 10 using similar teachings and rationale.
Regarding claim 23, Kumar in view of Deshmukh further in view of Azose teaches all the elements of claim 13, therefore is rejected for the same reasons as those presented for claim 13. The claim recites similar limitations corresponding to claim 11 and is rejected for similar reasons as claim 11 using similar teachings and rationale.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Daravanh Phakousonh whose telephone number is (571)272-6324. The examiner can normally be reached Mon - Thurs 7 AM - 5 PM, Every other Friday 7 AM - 4PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li B Zhen can be reached at 571-272-3768. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Daravanh Phakousonh/Examiner, Art Unit 2121
/Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121