Prosecution Insights
Last updated: October 02, 2026
Application No. 18/773,784

OVERLAY APPLICATION AND TECHNIQUES FOR INTERFACING WITH A GENERATIVE RESPONSE ENGINE

Final Rejection §103
Filed
Jul 16, 2024
Priority
May 10, 2024 — provisional 63/645,438
Examiner
AUGUSTINE, NICHOLAS
Art Unit
2179
Tech Center
2100 — Computer Architecture & Software
Assignee
Openai Opco LLC
OA Round
2 (Final)
73%
Grant Probability
Favorable
3-4
OA Rounds
1y 5m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 73% — above average
73%
Career Allowance Rate
605 granted / 832 resolved
+17.7% vs TC avg
Strong +28% interview lift
Without
With
+28.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 8m
Avg Prosecution
27 currently pending
Career history
874
Total Applications
across all art units

Statute-Specific Performance

§101
10.5%
-29.5% vs TC avg
§103
37.9%
-2.1% vs TC avg
§102
48.4%
+8.4% vs TC avg
§112
1.9%
-38.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 832 resolved cases

Office Action

§103
DETAILED ACTION A. This action is in response to the following communications: Amendment filed: 08/07/2026. This action is made Final. B. Claims 1-20 and 22-25 remain pending. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-20 and 22-25 is/are rejected under 35 U.S.C. 103 as being unpatentable over Marks, Justin et al. (US Pub. 2024/0046142 A1), herein referred to as “Marks” in view of Bose, Gautam et al. (US Pub. 2024/0338248 A1), herein referred to as “Bose” in further view of Singh, Manbinder Pal et al. (US Pub. 2023/0236702 A1), herein referred to as “Singh”. As for claim 1, Marks teaches. A method of interacting with a generative response engine based on a scope identified by a user, comprising: Identifying, by an operator application, an application that receives a focus of an input device (par. 4 computer program is configured to cause at least one processor to run a clustering AI/ML model on vector representations of a sequence of screens pertaining to a captured task flow to produce a sequence of clusters in a trace to produce a trace including a sequence of clusters); displaying, by an operator application, an overlay over a window of the application, the overlay including an input component and receiving the focus of the input device (par. 135; fig. 8c overlay/popup menu for interaction with user to aid in user interaction with multiple applications of the computer terminal; par. 25 using computer vision to detect applications; par. 26 hooking into API to get application calls; par. 27 task mining recording; par. 138-140 The associated activities developed by the RPA developer may be stored as templates to be offered for activities for actions. For instance, if a user clicks a button, the logic created by the RPA developer may be used to propose button press activities for future RPA workflows.); Specification paragraph 59 notes definition of overlay as “Control engine 202 may also be configured to control view engine 210 for controlling the view of the operator. In one aspect, operator 200 is configured as a translucent overlay over an application and displays a minimal user interface over the application or just outside of the application”. The overlay application obtaining, by an operator application, first input from the input component, the first input comprising natural language (par. 28 text on the screen may be used and fit to a natural language processing (NLP) model such as word2vec, or a more advanced semantic NLP model such as GPT-3, to build a vector representation of the screen); obtaining, by an operator application, a response from a generative response engine based on at least one of the first input, the response describing taking an action in the application (par.29-31 training based upon user interaction and captured images of computer screen using computer vision to determine users intent and aid in user future interaction by taking action for the user; Some examples of actions that the task mining approach of some embodiments may understand include, but are not limited to, filling in a form, viewing a report, editing customer details, sending an email, Logging into SAP®, and logging in to a computing system, etc. Currently, AI/ML models are trained to follow an action, attempt to match the action to known tasks, and categorizing the action to the matched task.); and providing, by an operator application, one or more synthetic inputs into the application based on the response (par. 32 after identifying the user intent, the user action may be mapped to associated activities in an RPA designer application, such as UiPath Studio™. Current task mining techniques output a sequence of actions or a set of automatable tasks from recordings of user interactions that to which CV has been applied). Marks does not go into specific detail about overlay functionality; however in the same field of endeavor Bose teaches displaying, by the operator application, an overlay over a window of the application, the overlay including an input component and receiving the focus of the input device while the overlay is positioned over the window of the application; providing, by the operation application, one or more synthetic inputs into the application based on the response, wherein providing the one or more synthetic inputs comprise mapping a user interface input event received at the overlay to the window of the application; and based on the mapping, generating a new synthetic input effective to take the action in the application (par. 57 task model 210 determines tasks from visual information; Bose does not use the terminology overlay but strongly implies overlay-like mechanism that monitors user interaction with the task model receiving video of workflow, frames of UI, segments of frames containing UI elements, parameter information inside the UI elements, HTML and DOM structure of application monitored; by the Task Model 210 stating that it can determine which frames… include information about a task being performed… and determine the segment of a frame relevant to be performed task (e.g. UI element) that Task Model 210 is describing interaction monitoring overlay (watching UI, identify elements and tracks user action. Further thru at least a prompt by the user the Task model can generate new synthetic input effective to take action in the application to complete a task. “ In this variant, the task model 210 can determine a set of tasks 120 by decomposing a process description into tasks 120 to complete the process. In this variant, the task model 210 can include an LLM, MLM, or other model using a Seq2Seq, GRU, convolutional layers, transformers, HANs, translations and/or other suitable model architecture elements. In this variant, the input to the task model 210 can include unstructured text (e.g., a paragraph), structured text (e.g., questionnaire responses, a list of tasks 120, etc.), a set of instructions 35 (e.g., from a prior iteration of the method), HTML code, an HTML DOM, a native application underlying structure (e.g., application layout), and/or any combination of aforementioned information and/or other inputs. “ It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Bose into Marks because Bose suggests in par. 3-6 improvements to Marks RPA for automating repetitive manual tasks. Marks as modified by Bose teaches an overlay functionality as discussed but as an alternative Singh is introduced as an obvious addition in such that Singh teaches displaying, by the operator application, an overlay over a window of the application, the overlay including an input component and receiving the focus of the input device while the overlay is positioned over the window of the application (par. 74, 100, 111-112 using a translucent display technique for an overlay application with click through functionality such that the underlay application received user input while the overlay application can be used for other functionality such as computer vision and machine learning model training). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Singh into Marks as modified by Bose; this is true because Singh suggests in par. 4 Users often desire to view more than one window on a screen, but often must switch between the windows to view content of different applications. Moreover, application windows are often opaque such that one cannot view what is behind the window. As for claim 2, Marks teaches. The method of claim 1, further comprising: determining a control interface to interact with the application; and identifying, using the generative response engine, a tool to invoke the application based on the control interface (par. 32 For example, when a user is performing a task in Salesforce®, consider the case where the user is creating a new customer. Current task mining techniques would only observe and identify actions involved in the task. However, some embodiments further identify that this task pertains to the intent creating a new customer in Salesforce®. The identified intent is then matched with the associated activity in an RPA designer application. In certain embodiments, the AI/ML model may be further enhanced to automatically generate the associated activities in the RPA designer application). As for claim 3, Marks teaches. The method of claim 2, wherein the control interface comprises at least one of a document object model, an application programming interface of the application, or a computer vision for perceiving the application (par.25-26 computer vision and API hooking). As for claim 4, Marks teaches. The method of claim 3, further comprising: when the control interface comprises the document object model, extracting the document object model (par. 27 document object model (DOM) )associated with a current view of the application; and providing the document object model to the generative response engine (par. 27 Images captured during task mining recording and the associated API call information can be time synchronized, and the API information can be used to provide further understanding regarding what the user is doing in the screens. This may facilitate better understanding of user intent by matching sets of user actions to an activity, such as via image comparison techniques. For instance, CV (including OCR) may be used to extract information about a given screen and then a clustering algorithm may be used to match the extracted information to similar screens. Also, a full picture of what a user is doing may be understood by combining image analysis and API information collection, for example DOM). As for claim 5, Marks teaches. The method of claim 3, further comprising: obtaining an application state using the application programming interface, wherein the application state is provided to the generative response engine in connection with a task (par. 26 API hooking application calls of user interaction with applications). As for claim 6, Marks teaches. The method of claim 1, wherein the one or more synthetic inputs comprises: generating an event associated with a document object model based on the one or more synthetic inputs (par. 33 generative AI output from hyper-automation to complete tasks from tasks mining by means of computer vision and DOM). As for claim 7, Marks teaches. The method of claim 1, further comprising: in response to detecting an event or an input to cause a second application to receive the focus, identifying coordinates of a window of the second application; moving the overlay over the window of the second application; and applying the focus to the overlay (par. 4 computer program is configured to cause at least one processor to run a clustering AI/ML model on vector representations of a sequence of screens pertaining to a captured task flow to produce a sequence of clusters in a trace to produce a trace including a sequence of clusters; mention of screens denotes more than on application; par. 25 task mining applications). As for claim 8, Marks teaches. The method of claim 1, further comprising: in response to detecting hovering of the input device over the application, identifying a process identifier of the application (par. 27 detecting user interaction with computer vision and DOM). As for claim 9, Marks teaches. The method of claim 1, further comprising: identifying a file open in the application based on a process identifier of the application, wherein the file is provided to the generative response engine (par. 31 focus of user to what screens are open and interacted with to help determine user intent). As for claim 10, Marks teaches. The method of claim 1, further comprising: in response to a view control input, displaying an alternative view for interacting with the generative response engine, wherein a default view is partially superimposed over the application in the overlay (par. 135 UI is overlay on top of running application that accept user input to help hyper automate task of applications based upon detected intent of user through trained model by use of computer vision and DOM; par. 138-140 The associated activities developed by the RPA developer may be stored as templates to be offered for activities for actions. For instance, if a user clicks a button, the logic created by the RPA developer may be used to propose button press activities for future RPA workflows). As for claim 11, Marks teaches. The method of claim 1, further comprising: obtaining information from the generative response engine based on a first content in user input into the application and a context associated with the first content; displaying a notification in the information related to second content to replace the first content, wherein the second content includes an improvement or revision to the first content; and replacing at least the first content with the second content (par. 30-32 and 135 fig. 8c user can select submit button for accept automation help from detected user intent by the system; par. 138-140 The associated activities developed by the RPA developer may be stored as templates to be offered for activities for actions. For instance, if a user clicks a button, the logic created by the RPA developer may be used to propose button press activities for future RPA workflows). As for claim 12, Marks teaches. The method of claim 1, wherein providing the one or more synthetic inputs comprises: displaying a list of events to perform based on the first input corresponding to the text (par. 135 fig. 8D, 822 displayed list as exemplary;). As for claim 13, Marks teaches. The method of claim 12, further comprising receiving a second input to modify an event in the list of events, wherein the generative response engine is configured to use information in the second input to generate the one or more synthetic inputs (par. 135 fig. 8E modifying displayed list). As for claim 14, Marks teaches. The method of claim 12, wherein performing respective events in the list of events comprises: capturing one or more images associated with a respective event, wherein the one or more images correspond to different states associated with the respective event at different times; and capturing one or more states associated with the respective event (par. 27 Images captured during task mining recording and the associated API call information can be time synchronized, and the API information can be used to provide further understanding regarding what the user is doing in the screens). As for claim 15, Marks teaches. The method of claim 14, further comprising: displaying a time-series control to view the one or more states at different times based on the one or more images associated with the respective event; receiving an input to restore the application based on a first image in the one or more images; and identifying a state corresponding to the first image and restoring the state corresponding to the first image (par. 25-27, 135, 140 the intent of user is determined based upon images captured by computer vision and new action can take place based upon past captured images of applications therefore restoring based upon captured image). As for claim 16, Marks teaches. The method of claim 1, wherein providing the one or more synthetic inputs comprises: obtaining a local model for controlling the application based on instructions from the generative response engine, wherein the local model is configured to provide the one or more synthetic inputs (par. 31-32 AI/ML models are trained to follow an action, attempt to match the action to known tasks, and categorizing the action to the matched task. After identifying the user intent, the user action may be mapped to associated activit(ies) in an RPA designer application, such as UiPath Studio™. Current task mining techniques output a sequence of actions or a set of automatable tasks from recordings of user interactions that to which CV has been applied. As for claim 17, Marks teaches. A method of generating application-agnostic input, comprising: identifying a process identifier of an application window of an application to receive a focus of an input device (par. 4 computer program is configured to cause at least one processor to run a clustering AI/ML model on vector representations of a sequence of screens pertaining to a captured task flow to produce a sequence of clusters in a trace to produce a trace including a sequence of clusters); displaying, by the operator application, an overlay over the window of the application, the overlay receiving the focus of the input device (Fig. 8 810 popup window overlays the applications/ operating system beneath it; wherein the overview window that accepts user input to help with task associated with an application par. 4 computer program is configured to cause at least one processor to run a clustering AI/ML model on vector representations of a sequence of screens pertaining to a captured task flow to produce a sequence of clusters in a trace to produce a trace including a sequence of clusters; par. 25 computer vision to capture multiple applications focus or not in focus to determine users intent). detecting input into the application window, wherein the input comprises a first content (par. 26 hooking into API for information calls to applications to train the model in detection of intent of user); obtaining information from a machine learning model based on the first content and a context associated with the first content (par. 28 text on the screen may be used and fit to a natural language processing (NLP) model such as word2vec, or a more advanced semantic NLP model such as GPT-3, to build a vector representation of the screen); and obtaining, by the operator application, the first input in the input component, the first input being a prompt to induce a generative response engine to provide a response that pertains to taking an action in the application (par. 28 text on the screen may be used and fit to a natural language processing (NLP) model such as word2vec, or a more advanced semantic NLP model such as GPT-3, to build a vector representation of the screen); obtaining, by the operator application, a response from the generative response engine that is responsive to the prompt, and describes taking the action in the application (par.29-31 training based upon user interaction and captured images of computer screen using computer vision to determine users intent and aid in user future interaction); and providing, by the operator application, one or more synthetic inputs effective to take the action into the application based on the response (par. 32 after identifying the user intent, the user action may be mapped to associated activit(ies) in an RPA designer application, such as UiPath Studio™. Current task mining techniques output a sequence of actions or a set of automatable tasks from recordings of user interactions that to which CV has been applied). Marks does not go into specific detail about overlay functionality; however in the same field of endeavor Bose teaches displaying, by the operator application, an overlay over a window of the application, the overlay including an input component and receiving the focus of the input device while the overlay is positioned over the window of the application; providing, by the operation application, one or more synthetic inputs into the application based on the response, wherein providing the one or more synthetic inputs comprise mapping a user interface input event received at the overlay to the window of the application; and based on the mapping, generating a new synthetic input effective to take the action in the application (par. 57 task model 210 determines tasks from visual information; Bose does not use the terminology overlay but strongly implies overlay-like mechanism that monitors user interaction with the task model receiving video of workflow, frames of UI, segments of frames containing UI elements, parameter information inside the UI elements, HTML and DOM structure of application monitored; by the Task Model 210 stating that it can determine which frames… include information about a task being performed… and determine the segment of a frame relevant to be performed task (e.g. UI element) that Task Model 210 is describing interaction monitoring overlay (watching UI, identify elements and tracks user action. Further thru at least a prompt by the user the Task model can generate new synthetic input effective to take action in the application to complete a task. “ In this variant, the task model 210 can determine a set of tasks 120 by decomposing a process description into tasks 120 to complete the process. In this variant, the task model 210 can include an LLM, MLM, or other model using a Seq2Seq, GRU, convolutional layers, transformers, HANs, translations and/or other suitable model architecture elements. In this variant, the input to the task model 210 can include unstructured text (e.g., a paragraph), structured text (e.g., questionnaire responses, a list of tasks 120, etc.), a set of instructions 35 (e.g., from a prior iteration of the method), HTML code, an HTML DOM, a native application underlying structure (e.g., application layout), and/or any combination of aforementioned information and/or other inputs. “ It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Bose into Marks because Bose suggests in par. 3-6 improvements to Marks RPA for automating repetitive manual tasks. Marks as modified by Bose teaches an overlay functionality as discussed but as an alternative Singh is introduced as an obvious addition in such that Singh teaches displaying, by the operator application, an overlay over a window of the application, the overlay including an input component and receiving the focus of the input device while the overlay is positioned over the window of the application (par. 74, 100, 111-112 using a translucent display technique for an overlay application with click through functionality such that the underlay application received user input while the overlay application can be used for other functionality such as computer vision and machine learning model training). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Singh into Marks as modified by Bose; this is true because Singh suggests in par. 4 Users often desire to view more than one window on a screen, but often must switch between the windows to view content of different applications. Moreover, application windows are often opaque such that one cannot view what is behind the window. As for claim 18, Marks teaches. The method of claim 17, wherein the identifying the window of the application further comprises: detecting, by the operator application, a UI event received by a window manager of an operating system, wherein the UI event is the focus of the input device in the window of the application; and determining, by the operator application, that the window of the application is valid for receiving using inputs using a window selection heuristic, wherein the window selection heuristic discriminates between windows that do not accept user inputs (par. 25-27 using computer vision to detect applications, using a hook to obtain information on application calls and task mining to determine user intent between window applications recorded and stored in a time synchronized manner to facilitate better understanding of user intent to aid in user interaction with applications as shown in figure 8c). As for claim 19, Marks teaches. The method of claim 17, wherein the identifying the window of the application further comprises: capturing, by the operator application, at least one screen image (par. 25 computer vision); providing, by the operator application, the at least one screen image to the generative response engine; and receiving, by the operator application, an instruction to display an overlay over the window of the application, wherein the generative response engine determined that the window is valid for receiving using inputs from the at least one screen image (par. 26 hooking into API stack of applications to obtain API calls to determine user intent; Images captured during task mining recording and the associated API call information can be time synchronized, and the API information can be used to provide further understanding regarding what the user is doing in the screens). As for claim 20, Marks teaches. The method of claim 19, further comprising: providing, by the operator application, window data from a window manager of an operating system (par. 25-26 computer vision and hooking into API stack of applications to obtain API calls to determine user intent). As for claim 22, Marks teaches. The method of claim 21, further comprising: receiving a pointer event in the overlay over the window of the application which results in the display of the input component (par. 26 hooking into API to see users interactions from API calls, interactions include mouse events). As for claim 23, Marks teaches. The method of claim 19, wherein the response from the generative response engine describes taking the action in the application by including instructions for interacting with the window of the application that are effective to take the action (par. 135 action presented in popup window to help aid in user intent with application; par. 138-140 The associated activities developed by the RPA developer may be stored as templates to be offered for activities for actions. For instance, if a user clicks a button, the logic created by the RPA developer may be used to propose button press activities for future RPA workflows). As for claim 24, Marks teaches. The method of claim 23 further comprising: moving the window of the application off the visible portion of the screen; and generating pseudo-window manager events to cause the window of the application to receive inputs as if it were in focus on the visible portion of the screen, whereby the operator application can take the action by interacting with the window of the application while the window of the application is not on the visible portion of the screen (par. 4 computer program is configured to cause at least one processor to run a clustering AI/ML model on vector representations of a sequence of screens pertaining to a captured task flow to produce a sequence of clusters in a trace to produce a trace including a sequence of clusters; par. 25 computer vision to capture multiple applications focus or not in focus to determine users intent). As for claim 25, Marks teaches. The method of claim 24, wherein a user is concurrently interacting with a second window displayed on the visible portion of the screen (par. 25 computer vision can capture all visible and non-visible screens to determine user intent). (Note :) It is noted that any citation to specific, pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. In re Heck, 699 F.2d 1331, 1332-33, 216 USPQ 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006,1009, 158 USPQ 275, 277 (CCPA 1968)). Response to Arguments Applicant’s arguments with respect to claim(s) have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Inquires Any inquiry concerning this communication should be directed to NICHOLAS AUGUSTINE at telephone number (571)270-1056. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. PNG media_image1.png 213 559 media_image1.png Greyscale /NICHOLAS AUGUSTINE/Primary Examiner, Art Unit 2178 August 18, 2026
Read full office action

Prosecution Timeline

Jul 16, 2024
Application Filed
May 12, 2026
Non-Final Rejection mailed — §103
Jul 22, 2026
Applicant Interview (Telephonic)
Jul 22, 2026
Examiner Interview Summary
Aug 07, 2026
Response Filed
Aug 20, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749110
SYSTEM AND METHOD TO PROVIDE A VIRTUAL STORE-FRONT
3y 1m to grant Granted Sep 29, 2026
Patent 12750554
INTERACTION METHOD AND APPARATUS, ELECTRONIC DEVICE, AND COMPUTER READABLE STORAGE MEDIUM
3y 0m to grant Granted Sep 29, 2026
Patent 12743576
SYSTEM AND METHOD FOR CREATION AND HANDLING OF CONFIGURABLE APPLICATIONS FOR WEBSITE BUILDING SYSTEMS
3y 2m to grant Granted Sep 22, 2026
Patent 12725377
OBJECT MODELING BASED ON PROPERTIES AND IMAGES OF AN OBJECT
4y 2m to grant Granted Sep 01, 2026
Patent 12717594
MULTIMEDIA DEVICE AND CONTROL METHOD THEREFOR
2y 8m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
73%
Grant Probability
99%
With Interview (+28.2%)
3y 8m (~1y 5m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 832 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month