Prosecution Insights
Last updated: September 17, 2026
Application No. 18/275,100

METHOD AND SYSTEM FOR GENERATING EVENT IN OBJECT ON SCREEN BY RECOGNIZING SCREEN INFORMATION ON BASIS OF ARTIFICIAL INTELLIGENCE

Final Rejection §103§112
Filed
Jul 31, 2023
Priority
Feb 18, 2021 — RE 10-2021-0021501 +1 more
Examiner
STANLEY, JEREMY L
Art Unit
2127
Tech Center
2100 — Computer Architecture & Software
Assignee
Infofla Inc.
OA Round
2 (Final)
49%
Grant Probability
Moderate
3-4
OA Rounds
1m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 49% of resolved cases
49%
Career Allowance Rate
143 granted / 290 resolved
-5.7% vs TC avg
Strong +41% interview lift
Without
With
+41.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
26 currently pending
Career history
313
Total Applications
across all art units

Statute-Specific Performance

§101
10.6%
-29.4% vs TC avg
§103
54.4%
+14.4% vs TC avg
§102
13.9%
-26.1% vs TC avg
§112
16.6%
-23.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 290 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is responsive to the Application filed on June 23, 2026. Claims 14, 17, 22, 24, and 30 are amended. Claims 14-33 are pending in the case. Claims 14, 22, and 30 are the independent claims. This action is final. Applicant’s Response In the response filed on June 26, 2023, Applicant amended the claims and provided arguments in response to the rejections of the claims under 35 USC 101, 102, and 103 in the previous office action. Response to Argument/Amendment Applicant’s amendments to the claims in response to the rejection of the claims under 35 USC 101 are acknowledged, and Applicant’s associated arguments have been fully considered. Applicant’s arguments are persuasive at least to the extent that the amended claims appear to recite limitations which constitute an improvement in technology which is supported by the specification. Therefore the rejection is withdrawn. Applicant’s amendments to the claims in response to the rejections of the claims under 35 USC 102 and 103 are acknowledged, and Applicant’s associated arguments have been fully considered. Applicant’s arguments are persuasive at least to the extent that the cited references do not appear to explicitly disclose “the pre-trained AI model is trained to perform an object detector function that, in a single inference pass on the received screen image, provides information indicating what type of object is present and at which position the object is present on the screen image.” Therefore, the rejections are withdrawn. New grounds of rejection are provided below. Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 14-33 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. With respect to claims 14, 22, and 30, these claims recite “an agent installed on the user terminal…a pre-trained AI model provided in the server…the pre-trained AI model is trained to perform an object detector function that…provides information indicating what type of object is present and at which position the object is present on the screen image; and transmitting result data indicating a location of the at least one object to the agent…wherein the transmitted data causes the agent to generate an event for the at least one object…wherein the agent identifies the at least one object based solely on visual content of the screen image, independently of class ID metadata embedded in source code of the application displayed on the screen of the user terminal.” These combined limitations appear to recite that a pre-trained AI model on a server provides object type and position information, and that the server sends result information indicating a location of the object to the agent, where this information causes the agent to generate the event for the object. However, in addition to these recited interactions in which the server appears to perform the processing associated with obtaining object type and position/location and the agent appears to utilize the results provided by the server, the claims also appear to recite a separate act of identification, performed by the agent (i.e. at the client, instead of the AI model in the server), which is performed solely based on visual content of the screen image (i.e. and therefore potentially separately from the object detection functions performed by the AI model on the server). While the specification of the instant application does appear to describe the server including the AI screen 230 functionality (including object detection) (see Fig. 1) and the client/agent including the screen object detector 133 functionality (see Fig. 2), these appear to be alternate embodiments, i.e. in some embodiments, the screen object detection is performed at the server and in other embodiments, the screen object detection is performed at the client. The specification of the instant application does not appear to describe, or contemplate, any scenario in which object detection functionality is performed at both the AI model at the server and the object detector at the agent in the client device, such that “the pre-trained AI model is trained to perform an object detector function that, in a single inference pass on the received screen image, provides information indicating what type of object is present and at which position the object is present…” while the agent at the client device separately “identifies the at least one object based solely on visual content of the screen image, independently of class ID metadata embedded in source code of the application…,” as is recited in the independent claims. Therefore claims 14, 22, and 30 contain subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Claims 15-21, 23-29, and 31-33 each respectively depend upon claims 14, 22, and 30, and inherit the deficiencies identified above. Therefore, these dependent claims are rejected on the same basis as identified above with respect to the independent claims. With respect to claim 17, the claim recites “wherein the training data further comprises object type class labels associated with each labeled object in the full screen images, and wherein the training data is collected from a plurality of different software application environments.” However, the specification of the instant application does not appear to contain any description of at least collecting training data from a plurality of different software application environments. Therefore claim 17 contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. With respect to claim 24, the claim recites “wherein the generating the event comprises sequentially controlling the input device to perform the text data input or mouse button click for each of the plurality of objects identified by the object detector function in the result data, in an order determined based on the type and location information of each object.” However, the specification of the instant application does not appear to contain any description of at least sequentially controlling input devices to interact with each of a plurality of objects in an order determined based on type and location information of the objects. Therefore claim 17 contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 14-33 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. With respect to claims 14, 22, and 30, these claims recite “an agent installed on the user terminal…a pre-trained AI model provided in the server…the pre-trained AI model is trained to perform an object detector function that…provides information indicating what type of object is present and at which position the object is present on the screen image; and transmitting result data indicating a location of the at least one object to the agent…wherein the transmitted data causes the agent to generate an event for the at least one object…wherein the agent identifies the at least one object based solely on visual content of the screen image, independently of class ID metadata embedded in source code of the application displayed on the screen of the user terminal.” These combined limitations appear to recite that a pre-trained AI model on a server provides object type and position information, and that the server sends result information indicating a location of the object to the agent, where this information causes the agent to generate the event for the object. However, in addition to these recited interactions in which the server appears to perform the processing associated with obtaining object type and position/location and the agent appears to utilize the results provided by the server, the claims also appear to recite a separate act of identification, performed by the agent (i.e. at the client, instead of the AI model in the server), which is performed solely based on visual content of the screen image (i.e. and therefore potentially separately from the object detection functions performed by the AI model on the server). While the specification of the instant application does appear to describe the server including the AI screen 230 functionality (including object detection) (see Fig. 1) and the client/agent including the screen object detector 133 functionality (see Fig. 2), these appear to be alternate embodiments, i.e. in some embodiments, the screen object detection is performed at the server and in other embodiments, the screen object detection is performed at the client. The specification of the instant application does not appear to describe, or contemplate, any scenario in which object detection functionality is performed at both the AI model at the server and the object detector at the agent in the client device, such that “the pre-trained AI model is trained to perform an object detector function that, in a single inference pass on the received screen image, provides information indicating what type of object is present and at which position the object is present…” while the agent at the client device separately “identifies the at least one object based solely on visual content of the screen image, independently of class ID metadata embedded in source code of the application…,” as is recited in the independent claims. Therefore, it cannot be determined how the individual operations of the AI model at the server (providing object type and position information in a single inference pass), the server transmitting result data (indicating a location of the object, which may not necessarily be the same as the position of the object detected by the AI model), and the agent at the client device (identifying the object based solely on visual content of the screen image (and therefore apparently not using the provided result data and/or output of the AI model)) are combined to implement the claimed invention. Further it is unclear whether the act of identification, performed by the agent at the client device, is intended to be an independent and separate act of object detection, or if it is intended to further define the object detection performed by the AI model at the server. If the act of identification performed by the agent is intended to be an independent and separate act, it is unclear how the result of that identification is utilized with respect to the rest of the claimed invention. Moreover, it is unclear whether the result data indicating the location of the object is generated based on the object detection performed by the AI model in the sever, which results in information indicating the position and type of the object. In the interest of providing full examination on the merits, the claim is interpreted as requiring at least the AI model at the sever performing the object detection which is performed in a single inference pass and is based solely on visual content of the screen image, independently of class ID metadata embedded in source code of the application, and which results in the object type and position information, and where the transmitted result information indicating the location of the object is based at least in part on the position information from the AI model’s object detection. Claims 15-21, 23-29, and 31-33 each respectively depend upon claims 14, 22, and 30, and inherit the deficiencies identified above. Therefore, these dependent claims are rejected on the same basis as identified above with respect to the independent claims. With respect to claim 17, the claim recites “wherein the training data further comprises object type class labels associated with each labeled object in the full screen images, and wherein the training data is collected from a plurality of different software application environments.” However, the specification of the instant application does not appear to contain any description of at least collecting training data from a plurality of different software application environments. Therefore, the intended scope of the limitation “collected form a plurality of different software application environments” cannot be determined and the limitation is indefinite. In the interest of providing full examination on the merits, the limitation is interpreted as requiring at least training data from applications that are different in one or more ways. With respect to claim 24, the claim recites “wherein the generating the event comprises sequentially controlling the input device to perform the text data input or mouse button click for each of the plurality of objects identified by the object detector function in the result data, in an order determined based on the type and location information of each object.” However, the specification of the instant application does not appear to contain any description of at least sequentially controlling input devices to interact with each of a plurality of objects in an order determined based on type and location information of the objects. Therefore, the intended scope of the limitation “sequentially controlling the input device…in an order determined based on the type and location information of each object” cannot be determined (i.e. such as how the type and location information of the objects affects the determined sequence/order for input device control), and the limitation is indefinite. In the interest of providing full examination on the merits, the limitation is interpreted as referring to any sequential activation of user interface objects. Claim Rejections – 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 14, 15, 18, 22-27, 30, 32, and 33 are rejected under 35 U.S.C. 103 as being unpatentable over Petursson (US 20180197103 A1) in view of Bigham et al. (US 20210349587 A1). With respect to claim 14, Petursson teaches a method performed by a server for generating an event on an object on a screen based on Artificial Intelligence (AI), the method comprising: (a) receiving, via a communication interface, a screen image of a user terminal from an agent installed on the user terminal (e.g. paragraph 0054, Fig. 1B, learning engine/system taking screenshot of GUI 102 and storing in database; paragraph 0069, Fig. 2, AI learning engine consisting of client module 202 and server module 220; client 202 is in charge of collecting screenshots/screen captures of GUI of SUT and performing coordinate based GUI operations to monitor changes; paragraph 0071, client 202 starts/initiates SUT via start app function 204 which constructs the initial action array data structure containing actions performed by client 202 on GUI and the results of those actions in the form of corresponding screenshots/captures; it then communicates the initialized action array to server 220 by invoking send action array function 206 which communicates action array to screen store function 224 of server 220; paragraph 0072, sending empty action array with initial screenshot; paragraph 0073, sending new screenshot of detected changes in updated action array; paragraph 0081, action array construct communicated by client 202 to screen store function 224 of AI server 220); (b) inferring location information of at least one object on the screen image by inputting the received screen image into a pre-trained Al model provided in the server (e.g. paragraph 0055, Fig. 1B, detecting objects on the screenshot by detect objects from screenshot function that is a chain of successive algorithms applied to detect objects from the screenshot including use of canny edge detection, contour detection, and user supplied algorithms; paragraph 0072, Fig. 2, screenshot undergoes processing from various functional modules of the sever; paragraph 0084, triggering screen detection function 226; paragraph 0093, object detection function 228 (in server 202) performing zero-metadata detection by scanning the image/screenshot of the GUI and detecting the edges/contours of the object; paragraph 0094, once object is detected, storing X and Y coordinates of the location of the object on the screen and the height and width of the object; paragraph 0096, system/learning engine 200B learning objects having any shape, gathering more and more information/knowledge about screen objects until they are fully identified by object types with a high level of likelihood); and (c) transmitting result data indicating a location of the at least one object to the agent via the communication interface, wherein the transmitted result data causes the agent to generate an event for the at least one object on the screen of the user terminal, and wherein the event comprises controlling an input device to perform a text data input or a mouse button click at a screen coordinate corresponding to the location indicated by the result data (e.g. paragraph 0048, GUI including text input box; paragraph 0056, detected object may be a text input box; paragraph 0061, Fig. 1B, suggest next action functionality implemented by determining objects/object groups to utilize; checking to ensure object is indeed interactive or actionable and then determining number of actions applied to object or object group, and proceeding with object with the lowest action-count; paragraph 0062, determining weight/priority of object, selecting object with highest weight/priority, and invoking input generator to supply the input for the object; determining next action to perform on the object; paragraph 0064, performing the next action identified; paragraph 0070, simulating mouse and keyboard actions on GUI; paragraph 0072, Fig. 2, last of the sever side modules, action suggester function 246, sends populated action array with suggested actions to perform on various objects back to client 202; paragraph 0073, action execution function 210 performing suggested actions on the various objects; paragraph 0113-0116, after objects on pages/screens of GUI are detected, passing control to application mapping function 230 which models behavior of application as it changes states due to actions/events, including actions triggered by inputs in the form of GUI interactions; actions caused by action execution 210 which are automated/simulated actions by learning engine; paragraph 0144, action suggester 246 populates action array with actions to be performed by client 202, picking screen objects and populating action array fields object_id, object_info, an action_group_id, invoking guesser to determine what action to be guessed or tried on the screen object; paragraph 0200, multiple different types of mouse action available for single object; i.e. information about the detected object (such as location/coordinates) and a suggested action/event for the object are transmitted by the server back to the client, and the action/event is performed/generated for the object by the client’s action execution function 210, as shown in Fig. 2, where the action/event may include providing a mouse click or keyboard text input, as appropriate for a given UI object based on its type and location (such as providing a mouse click to a button, keyboard text input to a text input box, etc.)). Petursson does not explicitly disclose wherein the pre-trained AI model is trained using training data comprising full screen images and labeled location data of objects within the full screen images, and the pre-trained AI model is trained to perform an object detector function that, in a single inference pass on the received screen image, provides information indicating what type of object is present and at which position the object is present on the screen image, and wherein the agent identifies the at least one object based solely on visual content of the screen image, independently of class ID metadata embedded in source code of the application displayed on the screen of the user terminal. However, Bigham teaches wherein the pre-trained AI model is trained using training data comprising full screen images and labeled location data of objects within the full screen images (e.g. paragraph 0057, using object detector trained to detect UI elements from a screen shot to convert visual representation of an app (rendered screen) to a semantic one (location and types of UI elements); paragraph 0058, model trained using suitable number of application screens; each screen’s UI elements of the data set may be labeled as bounding boxes; i.e. the object detector is trained using application screens (i.e. fullscreen application images) which are labeled with bounding boxes for each object such that the object detector can predict both location and type of objects in the image/screen), and the pre-trained AI model is trained to perform an object detector function that, in a single inference pass on the received screen image, provides information indicating what type of object is present and at which position the object is present on the screen image (e.g. paragraphs 0056-0057, element detection model to semantically represent UI screen as list of bounding boxes; using object detector trained to detect UI elements from a screenshot to convert visual representation of app (rendered screen) to semantic one (location and types of UI element); the object detector is based on suitable architecture for single-shot object detection and returns bounding boxes for UI elements detected in input image; i.e. the pre-trained object detector model outputs a semantic representation of the app including locations and types of detected objects (i.e. analogous to providing information indicating what type of object is present and at which position the object is present) and is a single-shot model (i.e. performs the object detection and outputting of the object types and locations in a single inference pass on the received image)), and wherein the agent identifies the at least one object based solely on visual content of the screen image, independently of class ID metadata embedded in source code of the application displayed on the screen of the user terminal (e.g. paragraphs 0056-0057, Apps developed using default UI toolkits for OS of device may provide representations of screen layouts through metadata about on-screen elements including location and type; however, apps developed using other/third party development libraries and software development kit may not include this metadata; operating under minimal assumptions about the amount and type of information accessible to the device; regardless of what UI toolkit is used, all applications render content and controls to the screen for display to the user; converting visual representation of the app (i.e. the rendered screen) to a semantic one (location and types of UI elements) using single shot object detection; i.e. the object detection model/agent operates under minimal assumptions about available information and instead performs object detection using only rendered content of the application screen, without relying on/requiring that any metadata (such as class ID, object type, etc.) be available, such that object detection can be performed even for applications which are developed in environments not providing such metadata). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Petursson and Bigham in front of him to have modified the teachings of Petursson (directed to automatically testing/learning the behavior of a system under test, including graphical user interface behavior), to incorporate the teachings of Bigham (directed to pixel-based optimization for a user interface, such as by using a trained model) to include the capability to train the model using training data comprising full screen images and labeled location data objects within the images such that the model is trained to perform object detection in a single inference pass and provide position and type information for the object, based only on visual content of the screen image, without depending relevant metadata in application source code. One of ordinary skill would have been motivated to perform such a modification in order to permit the device, including an object detection model, to be operated under minimal assumptions about the amount and type of information accessible to the electronic device as described in Bingham (paragraph 0057). With respect to claims 22 and 30, Petursson teaches a non-transitory computer-readable storage medium storing a computer program, which when executed by a processor of a user terminal, causes the processor to perform operations comprising and method, and the method for generating an event on an object on a screen based on artificial intelligence (AI) (e.g. claim 1, program instructions stored in non-transitory storage medium), the method comprising: (a) transmitting, by an agent installed on a user terminal, a screen image of the user terminal to a server and (b) requesting, by the agent, object location information inferred from the screen image by a pre-trained Al model provided in the server (e.g. paragraph 0054, Fig. 1B, learning engine/system taking screenshot of GUI 102 and storing in database; paragraph 0069, Fig. 2, AI learning engine consisting of client module 202 and server module 220; client 202 is in charge of collecting screenshots/screen captures of GUI of SUT and performing coordinate based GUI operations to monitor changes; paragraph 0071, client 202 starts/initiates SUT via start app function 204 which constructs the initial action array data structure containing actions performed by client 202 on GUI and the results of those actions in the form of corresponding screenshots/captures; it then communicates the initialized action array to server 220 by invoking send action array function 206 which communicates action array to screen store function 224 of server 220; paragraph 0072, sending empty action array with initial screenshot; paragraph 0073, sending new screenshot of detected changes in updated action array; paragraph 0081, action array construct communicated by client 202 to screen store function 224 of AI server 220; i.e. where sending the screenshots to the server ultimately triggers the server sending back detected object information, including location information, as cited below, the transmission of the screenshot also acts as/constitutes as request for the location information); (c) receiving, by the agent via a communication interface, result data indicating a location of at least one object on the screen, wherein the result data is inferred by the Al model from the transmitted screen image (e.g. paragraph 0061, Fig. 1B, suggest next action functionality implemented by determining objects/object groups to utilize; checking to ensure object is indeed interactive or actionable and then determining number of actions applied to object or object group, and proceeding with object with the lowest action-count; paragraph 0062, determining weight/priority of object, selecting object with highest weight/priority, and invoking input generator to supply the input for the object; determining next action to perform on the object; paragraph 0072, Fig. 2, last of the sever side modules, action suggester function 246, sends populated action array with suggested actions to perform on various objects back to client 202; paragraph 0113-0116, after objects on pages/screens of GUI are detected, passing control to application mapping function 230 which models behavior of application as it changes states due to actions/events, including actions triggered by inputs in the form of GUI interactions; paragraph 0144, action suggester 246 populates action array with actions to be performed by client 202, picking screen objects and populating action array fields object_id, object_info, an action_group_id, invoking guesser to determine what action to be guessed or tried on the screen object; Table 1, showing that object_info field of the action array includes coordinates and size information of the objects; i.e. information about the detected object (such as location/coordinates) and a suggested action/event for the object are transmitted by the server back to the client, and the action/event is performed/generated for the object by the client’s action execution function 210, as shown in Fig. 2); and (d) generating, by the agent, an event for the at least one object on the screen of the user terminal based on the received result, wherein the generating the event comprises controlling an input device to perform a text data input or a mouse button click at a screen coordinate corresponding to the received location information (e.g. paragraph 0048, GUI including text input box; paragraph 0056, detected object may be a text input box; paragraph 0061, Fig. 1B, suggest next action functionality implemented by determining objects/object groups to utilize; checking to ensure object is indeed interactive or actionable and then determining number of actions applied to object or object group, and proceeding with object with the lowest action-count; paragraph 0062, determining weight/priority of object, selecting object with highest weight/priority, and invoking input generator to supply the input for the object; determining next action to perform on the object; paragraph 0064, performing the next action identified; paragraph 0070, simulating mouse and keyboard actions on GUI; paragraph 0072, Fig. 2, last of the sever side modules, action suggester function 246, sends populated action array with suggested actions to perform on various objects back to client 202; paragraph 0073, action execution function 210 performing suggested actions on the various objects; paragraph 0113-0116, after objects on pages/screens of GUI are detected, passing control to application mapping function 230 which models behavior of application as it changes states due to actions/events, including actions triggered by inputs in the form of GUI interactions; actions caused by action execution 210 which are automated/simulated actions by learning engine; paragraph 0144, action suggester 246 populates action array with actions to be performed by client 202, picking screen objects and populating action array fields object_id, object_info, an action_group_id, invoking guesser to determine what action to be guessed or tried on the screen object; paragraph 0200, multiple different types of mouse action available for single object; i.e. information about the detected object (such as location/coordinates) and a suggested action/event for the object are transmitted by the server back to the client, and the action/event is performed/generated for the object by the client’s action execution function 210, as shown in Fig. 2, where the action/event may include providing a mouse click or keyboard text input, as appropriate for a given UI object based on its type and location (such as providing a mouse click to a button, keyboard text input to a text input box, etc.)). Petursson does not explicitly disclose the pre-trained AI model is trained using training data comprising full screen images and labeled location data of objects within the full screen images, and the pre-trained AI model is trained to perform an object detector function that, in a single inference pass on the received screen image, provides information indicating what type of object is present and at which position the object is present on the screen image, and wherein the agent identifies the at least one object based solely on visual content of the screen image, independently of class ID metadata embedded in source code of the application displayed on the screen of the user terminal. However, Bigham teaches wherein the pre-trained AI model is trained using training data comprising full screen images and labeled location data of objects within the full screen images (e.g. paragraph 0057, using object detector trained to detect UI elements from a screen shot to convert visual representation of an app (rendered screen) to a semantic one (location and types of UI elements); paragraph 0058, model trained using suitable number of application screens; each screen’s UI elements of the data set may be labeled as bounding boxes; i.e. the object detector is trained using application screens (i.e. fullscreen application images) which are labeled with bounding boxes for each object such that the object detector can predict both location and type of objects in the image/screen), and the pre-trained AI model is trained to perform an object detector function that, in a single inference pass on the received screen image, provides information indicating what type of object is present and at which position the object is present on the screen image (e.g. paragraphs 0056-0057, element detection model to semantically represent UI screen as list of bounding boxes; using object detector trained to detect UI elements from a screenshot to convert visual representation of app (rendered screen) to semantic one (location and types of UI element); the object detector is based on suitable architecture for single-shot object detection and returns bounding boxes for UI elements detected in input image; i.e. the pre-trained object detector model outputs a semantic representation of the app including locations and types of detected objects (i.e. analogous to providing information indicating what type of object is present and at which position the object is present) and is a single-shot model (i.e. performs the object detection and outputting of the object types and locations in a single inference pass on the received image)), and wherein the agent identifies the at least one object based solely on visual content of the screen image, independently of class ID metadata embedded in source code of the application displayed on the screen of the user terminal (e.g. paragraphs 0056-0057, Apps developed using default UI toolkits for OS of device may provide representations of screen layouts through metadata about on-screen elements including location and type; however, apps developed using other/third party development libraries and software development kit may not include this metadata; operating under minimal assumptions about the amount and type of information accessible to the device; regardless of what UI toolkit is used, all applications render content and controls to the screen for display to the user; converting visual representation of the app (i.e. the rendered screen) to a semantic one (location and types of UI elements) using single shot object detection; i.e. the object detection model/agent operates under minimal assumptions about available information and instead performs object detection using only rendered content of the application screen, without relying on/requiring that any metadata (such as class ID, object type, etc.) be available, such that object detection can be performed even for applications which are developed in environments not providing such metadata). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Petursson and Bigham in front of him to have modified the teachings of Petursson (directed to automatically testing/learning the behavior of a system under test, including graphical user interface behavior), to incorporate the teachings of Bigham (directed to pixel-based optimization for a user interface, such as by using a trained model) to include the capability to train the model using training data comprising full screen images and labeled location data objects within the images such that the model is trained to perform object detection in a single inference pass and provide position and type information for the object, based only on visual content of the screen image, without depending relevant metadata in application source code. One of ordinary skill would have been motivated to perform such a modification in order to permit the device, including an object detection model, to be operated under minimal assumptions about the amount and type of information accessible to the electronic device as described in Bingham (paragraph 0057). With respect to claim 15, Petursson in view of Bigham teaches all of the limitations of claim 14 as previously discussed, and Petursson further taches the method further comprising: receiving a registration of a scheduler from the agent (e.g. paragraph 0064, indicating use of a predetermined timer for executing the process, including executing the process until the predetermined timer has run out; paragraph 0233, unmanned execution going until predetermined time has passed; i.e. a user or setting provides predetermined timing details for performing the execution/testing by the AI engine/system); and transmitting a start signal to the agent via the communication interface at a scheduled time, wherein the receiving of the screen image is performed in response to the start signal (e.g. paragraph 0064, indicating use of a predetermined timer for executing the process, including executing the process until the predetermined timer has run out; paragraph 0071, client 202 starts or initiates SUT via start app function; paragraph 0072, start app function initializing empty action array, taking first screenshot of GUI and sending to sever by send action array function; paragraph 0232-0233, system running unmanned, until predetermined time has passed; i.e. where a system running unmanned according to a predetermined timer, until a predetermined amount of time has passed would include beginning/starting the execution at a scheduled time, such as a designated time for the starting of the predetermined timer/time period). With respect to claim 23, Petursson in view of Bigham teaches all of the limitations of claim 22 as previously discussed, and Petursson further teaches the method further comprising: registering a scheduler with the server (e.g. paragraph 0064, indicating use of a predetermined timer for executing the process, including executing the process until the predetermined timer has run out; paragraph 0233, unmanned execution going until predetermined time has passed; i.e. a user or setting provides predetermined timing details for performing the execution/testing by the AI engine/system); and receiving a start signal from the server via the communication interface at a scheduled time, wherein the transmitting of the screen image is initiated in response to the start signal (e.g. paragraph 0064, indicating use of a predetermined timer for executing the process, including executing the process until the predetermined timer has run out; paragraph 0071, client 202 starts or initiates SUT via start app function; paragraph 0072, start app function initializing empty action array, taking first screenshot of GUI and sending to sever by send action array function; paragraph 0232-0233, system running unmanned, until predetermined time has passed; i.e. where a system running unmanned according to a predetermined timer, until a predetermined amount of time has passed would include beginning/starting the execution at a scheduled time, such as a designated time for the starting of the predetermined timer/time period). With respect to claim 18, Petursson in view of Bigham teaches all of the limitations of claim 14 as previously discussed, and Petursson further teaches wherein the at least one object includes at least one of a program window, a search bar of a browser, a login button, a company name, an ID input field, or a password input field (e.g. paragraphs 0076-0077, password field; paragraph 0077, login button). With respect to claim 24, Petursson in view of Bigham teaches all of the limitations of claim 22 as previously discussed, and Petursson further teaches wherein the generating the event comprises sequentially controlling the input device to perform a text data input or a mouse button click for each of a plurality of objects identified by the object detector function in the result data, in an order determined based on the type and location information of each object (e.g. paragraph 0048, GUI including text input box; paragraph 0056, detected object may be a text input box; paragraph 0061, Fig. 1B, suggest next action functionality implemented by determining objects/object groups to utilize; checking to ensure object is indeed interactive or actionable and then determining number of actions applied to object or object group, and proceeding with object with the lowest action-count; paragraph 0062, determining weight/priority of object, selecting object with highest weight/priority, and invoking input generator to supply the input for the object; determining next action to perform on the object; paragraph 0064, performing the next action identified; paragraph 0070, simulating mouse and keyboard actions on GUI; paragraph 0072, Fig. 2, last of the sever side modules, action suggester function 246, sends populated action array with suggested actions to perform on various objects back to client 202; paragraph 0073, action execution function 210 performing suggested actions on the various objects; paragraph 0077, grouping UI objects in order to execute multiple actions on UI objects successively (such as entering a username in an appropriate field, entering a password in the password field, and pressing/clicking a login button); paragraph 0113-0116, after objects on pages/screens of GUI are detected, passing control to application mapping function 230 which models behavior of application as it changes states due to actions/events, including actions triggered by inputs in the form of GUI interactions; actions caused by action execution 210 which are automated/simulated actions by learning engine; paragraph 0144, action suggester 246 populates action array with actions to be performed by client 202, picking screen objects and populating action array fields object_id, object_info, an action_group_id, invoking guesser to determine what action to be guessed or tried on the screen object; paragraph 0160, Tables 1 and 2, action on object including click of mouse, etc.; paragraph 0200, multiple different types of mouse action available for single object; i.e. using an action array and object grouping, the system can generate an event (such as logging in) by sequentially controlling related input devices to perform text input/mouse clicks for a plurality of objects, in an order based on the object types and locations, such as sequentially providing a user name text input to an associated UI object, providing a password text input to a password text input field, and providing a click/press to a login button, according to the action array/object grouping for the detected objects) With respect to claim 25, Petursson in view of Bigham teaches all of the limitations of claim 22 as previously discussed, and Petursson further teaches wherein the agent operates in an environment including at least one of a Web environment, a Command Line Interface (CLI) environment, or a Remote Desktop Protocol (RDP) environment (e.g. paragraph 0048, SUT/application implemented with thin client such as browser based client that communicates over web to backend webserver; i.e. the system operates in at least a web environment). With respect to claim 26, Petursson in view of Bigham teaches all of the limitations of claim 22 as previously discussed, and Petursson further teaches wherein the capturing the screen image and the generating the event are repeated to automatically perform a series of tasks performed by a user (e.g. as shown in Figs. 1B and 2, after taking the screenshot, sending to server, receiving suggested action, and performing suggested action, the process repeats; paragraphs 0075-0077, when more than one actions belong to group, action execution performs all actions in the group; grouping functionality useful for screens containing grouped GUI objects; in the absence of grouping functionality, multiple screenshots sent to server; as each action is executed, new screenshot sent to server and server determines next suggested action; object grouper groups username filed, password field, and login button into object group and then client executes all of the above described three actions successively, determines change thresholds are triggered, and takes screenshot and sends to the server; paragraphs 0148-0150, object grouper detecting which stored objects function together, groups them into functional sets, and assigns them a same action_group_id in action array; all actions having same action_group_id performed by action execution of client and then latest screenshot sent to sever side). With respect to claim 27, Petursson in view of Bigham teaches all of the limitations of claim 22 as previously discussed, and Petursson further teaches wherein the agent recognizes the at least one object even if a Class ID of the object in a source code is changed (e.g. paragraph 0093, object detection function performs zero-metadata detection; there is no prior knowledge or metadata needed about a screen object before it is detected, and detection is based solely on visual indicators in the GUI; i.e. because the detection/recognition occurs with no dependency on any metadata such as an object/class ID in source code, the detection/recognition will occur regardless of whether the Class ID of the object in source code is changed, or not changed). With respect to claim 32, Petursson in view of Bigham teaches all of the limitations of claim 30 as previously discussed, and Petursson further teaches wherein the operations further comprise creating a log upon completion of the event generation or upon occurrence of an error (e.g. paragraph 0061, determining total number of actions applied to object/object group thus far, referred to as the action-count; i.e. the system tracks the number of completed events/actions with respect to the different GUI objects, analogous to creating/keeping a log of completed events/actions). With respect to claim 33, Petursson in view of Bigham teaches all of the limitations of claim 30 as previously discussed, and Petursson further teaches wherein the user terminal includes any one of a desktop computer, a laptop, an IoT device, a connected car terminal, or a kiosk, and the operations are performed in a non-Windows Operating System (OS) environment or a remote terminal environment (e.g. paragraph 0049, computing platform may be a desktop computer, laptop computer, tablet, mobile/smartphone, etc.; paragraph 0160, discussing implementation on different operating systems, including non-Windows operating systems). Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Petursson in view of Bigham, further in view of Kessler et al. (US 20210026606 A1). With respect to claim 16, Petursson in view of Bigham teaches all of the limitations of claim 14 as previously discussed. Petursson does not explicitly disclose wherein the pre-trained Al model is trained using a Deep Learning algorithm including at least one of a Convolutional Neural Network (CNN), a Region-based Convolutional Neural Network (R-CNN), Fast R-CNN, Faster R- CNN, or Mask R-CNN. However, Kessler teaches wherein the pre-trained Al model is trained using a Deep Learning algorithm including at least one of a Convolutional Neural Network (CNN), a Region-based Convolutional Neural Network (R-CNN), Fast R-CNN, Faster R- CNN, or Mask R-CNN (e.g. paragraph 0018, using computer vision technique such as a trained convolutional neural network (CNN) to identify graphical user interface controls; paragraph 0080, CNN trained to analyze labeled digital images which depict GUI elements with associated component labels). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Petursson, Bigham, and Kessler in front of him to have modified the teachings of Petursson (directed to automatically testing/learning the behavior of a system under test, including graphical user interface behavior) and Bigham (directed to pixel-based optimization for a user interface, such as by using a trained model), to incorporate the teachings of Kessler (directed to visual programming methods, including use of ML models for identifying graphical user interface elements in screen images) to include the capability to implement the AI engine/component as a trained ML model such as a CNN. One of ordinary skill would have been motivated to perform such a modification in order to provide automated GUI control methods which allow enterprise employees to identify graphical user interface controls for use in voice control applications, add cloud-based voice functionality to business objects, retrofit legacy applications to include voice capability, automate applications using voice capabilities for robotic process automation, and create palettes of actions for intents as described in Kessler (paragraph 0018). Claims 17 and 31 are rejected under 35 U.S.C. 103 as being unpatentable over Petursson in view of Bigham, further in view of Kumar et al. (US 20190250891 A1). With respect to claim 17, Petursson in view of Bigham teaches all of the limitations of claim 14 as previously discussed, and Bigham further teaches wherein the training data further comprises object type class labels associated with each labeled object in the full screen images, and wherein the training data is collected from a plurality of different software application environments (e.g. paragraph 0056-0058, applications developed with default UI toolkits for operating systems; apps developed using third party app development libraries and software development kit; regardless of toolkit used all mobile applications render content and controls to the screen for display to the user; object detector trained to detect elements from screenshot, and convert visual representation of app to semantic one including location and types of UI elements; model trained using suitable number of application screens such as a dataset of 89,000 application screens, etc., with each screen’s UI elements of the dataset labeled as bounding boxes; i.e. the detector model is trained to detect object types and locations across a plurality of different software application environments, using a large training dataset including appropriately labeled application screenshots, i.e. training data including labeled screenshots and associated object type/class information for labeling objects across a wide variety of application types/environments such that objects can be detected regardless of the environment the application was developed in). Assuming arguendo that Petursson and Bigham do not explicitly disclose wherein the training data further comprises object type class labels associated with each labeled object in the full screen images, Kumar teaches this limitation (e.g. paragraph 0092, model trained using annotated training samples including images that include various UI components, where the annotations may include a label or tag that uniquely identifies each UI component, such as the location of each UI component within an image, type or class of the UI component, etc.; paragraph 0161, Fig. 14, showing example of input GUI screen image 1410). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Petursson, Bigham, and Kumar in front of him to have modified the teachings of Petursson (directed to automatically testing/learning the behavior of a system under test, including graphical user interface behavior) and Bigham (directed to pixel-based optimization for a user interface, such as by using a trained model), to incorporate the teachings of Kumar (directed to automated GUI development) to include the capability to train the model using full GUI screen images and labels including location data, type data, class data, etc. of objects/components within the GUI screen. One of ordinary skill would have been motivated to perform such a modification in order to automate development of a GUI for an application, avoiding tedious, time consuming, labor intensive, and expensive GUI development effort as described in Kumar (paragraph 0004-0005). With respect to claim 31, Petursson in view of Bigham teaches all of the limitations of claim 30 as previously discussed. Petursson does not explicitly disclose wherein the pre-trained AI model is updated by learning new training data comprising screen images and labeled object locations periodically or at a specific timing. However, Kumar teaches wherein the pre-trained AI model is updated by learning new training data comprising screen images and labeled object locations periodically or at a specific timing (e.g. paragraph 0024, user providing feedback such as information identifying new UI component types, etc.; retraining the machine learning based classifier based upon the user feedback; paragraph 0082, user providing feedback, using feedback to update reference information (training samples); updated reference information used for retraining the model; i.e. users may periodically provide feedback which is used to retrain/update the model, or the user may provide the feedback at a specific time, causing the model to be updated at that specific time). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Petursson, Bigham, and Kumar in front of him to have modified the teachings of Petursson (directed to automatically testing/learning the behavior of a system under test, including graphical user interface behavior) and Bigham (directed to pixel-based optimization for a user interface, such as by using a trained model), to incorporate the teachings of Kumar (directed to automated GUI development) to include the capability to update the model with new training data (including screen images and object location labels) periodically or at a specific time (i.e. based on either periodic user feedback or user feedback at a specific time). One of ordinary skill would have been motivated to perform such a modification in order to automate development of a GUI for an application, avoiding tedious, time consuming, labor intensive, and expensive GUI development effort as described in Kumar (paragraph 0004-0005). Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Petursson in view of Bigham, further in view of McKain et al. (US 20080295076 A1), further in view of Puszkiewicz et al. (US 20200159647 A1). With respect to claim 20, Petursson in view of Bigham teaches all of the limitations of claim 14 as previously discussed. Petursson does not explicitly disclose wherein the result data transmitted to the agent comprises coordinate data represented in a JSON (JavaScript Object Notation) or XML format. However, McKain teaches wherein the result data transmitted to the agent comprises coordinate data represented in a JSON (JavaScript Object Notation) or XML format (e.g. paragraph 0004, UI parser transforming text and data files associated with user interface into data format such as XML format making testing of UI data more efficient, including generating XML data for properties and components of the user interface including UI component location, etc.). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Petursson, Bigham, and McKain in front of him to have modified the teachings of Petursson (directed to automatically testing/learning the behavior of a system under test, including graphical user interface behavior) and Bigham (directed to pixel-based optimization for a user interface, such as by using a trained model), to incorporate the teachings of McKain (directed to graphical user interface testing) to include the capability to represent the transmitted coordinate data (as taught by Petursson) in an XML format. One of ordinary skill would have been motivated to perform such a modification in order to make testing of UI data more efficient as described in McKain (paragraph 0004). Petursson does not explicitly disclose wherein the result data transmitted to the agent comprises a confidence score indicating a probability of recognition for the at least one object. However, Puszkiewicz teaches wherein the result data transmitted to the agent comprises a confidence score indicating a probability of recognition for the at least one object (e.g. paragraph 0058, captured image representing application GUI comprising plurality of on-screen objects applied to model to analyze the image and identify and classify each object present in the image; paragraph 0063, identifying measure of confidence for each classified graphical object representing a confidence that the graphical object has been classified accurately). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Petursson, Bigham, McKain, and Puszkiewicz in front of him to have modified the teachings of Petursson (directed to automatically testing/learning the behavior of a system under test, including graphical user interface behavior), Bigham (directed to pixel-based optimization for a user interface, such as by using a trained model), and McKain (directed to graphical user interface testing), to incorporate the teachings of Puszkiewicz (directed to testing user interfaces using machine vision) to include the capability to include, in the transmitted result data (of Petursson), a confidence score indicating a probability of recognition of at least one object. One of ordinary skill would have been motivated to perform such a modification in order to enable validation of an application’s GUI in a dynamic fashion using a model which is universal to a plurality of computing platforms which can be used across different systems as described in Puszkiewicz (paragraph 0006). With respect to claim 21, Petursson in view of Bigham, further in view of McKain, further in view of Puszkiewicz teaches all of the limitations of claim 20 as previously discussed, and Puszkiewicz further teaches the method further comprising: determining whether the confidence score meets a predetermined threshold (e.g. paragraph 0063, determining whether measure of confidence exceeds threshold such as 90% or higher); and transmitting an error notification or a retry request to the agent if the confidence score is below the predetermined threshold (e.g. paragraph 0063, indicating validation has failed for graphical objects that were or unclassified or comprise a measure of confidence below the threshold value). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Petursson, Bigham, McKain, and Puszkiewicz in front of him to have modified the teachings of Petursson (directed to automatically testing/learning the behavior of a system under test, including graphical user interface behavior), Bigham (directed to pixel-based optimization for a user interface, such as by using a trained model), and McKain (directed to graphical user interface testing), to incorporate the teachings of Puszkiewicz (directed to testing user interfaces using machine vision) to include the capability to determine whether the confidence score is above or below a threshold and, when it is below the threshold, provide/transmit an indication/error notification that the validation has failed. One of ordinary skill would have been motivated to perform such a modification in order to enable validation of an application’s GUI in a dynamic fashion using a model which is universal to a plurality of computing platforms which can be used across different systems as described in Puszkiewicz (paragraph 0006). Claim 28 is rejected under 35 U.S.C. 103 as being unpatentable over Petursson in view of Bigham, further in view of Singh et al. (US 20210333983 A1). With respect to claim 28, Petursson in view of Bigham teaches all of the limitations of claim 22 as previously discussed. Petursson does not explicitly disclose the method further comprising: displaying a visual indicator, including a bounding box or a highlight overlay, on the at least one object on the screen of the user terminal based on the received result data prior to or simultaneously with the generating the event. However, Singh teaches the method further comprising: displaying a visual indicator, including a bounding box or a highlight overlay, on the at least one object on the screen of the user terminal based on the received result data prior to or simultaneously with the generating the event (e.g. paragraph 0041, object detection model receiving application UI images and outputting bounding boxes, class object labels, and confidence scores; detected UI control objects presented to user as part of the application and presentation of detected UI control objects is performed by highlighting the detected UI control objects contained in an application screen presented to the user; the user then performs the actions, the system captures the actions, etc.). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Petursson, Bigham, and Singh in front of him to have modified the teachings of Petursson (directed to automatically testing/learning the behavior of a system under test, including graphical user interface behavior) and Bigham (directed to pixel-based optimization for a user interface, such as by using a trained model), to incorporate the teachings of Singh (directed to learning user interfaces controls via incremental data synthesis) to include the capability for the model to output bounding boxes corresponding to detected UI components/controls and to highlight them on the GUI screen of the user terminal prior to generating the event/action. One of ordinary skill would have been motivated to perform such a modification in order to allow for automating tasks and application usage, including in legacy application programs, licensed applications, and other situations in which automated detection of controls is a challenge as described in Singh (paragraph 0002-0004). Claims 19 and 29 are rejected under 35 U.S.C. 103 as being unpatentable over Petursson in view of Bigham, further in view of Corwin et al. (US 20210150263 A1). With respect to claim 19, Petursson in view of Bigham teaches all of the limitations of claim 14 as previously discussed. Petursson does not explicitly disclose wherein the communication interface is configured to support TCP/IP socket communication. However, Corwin teaches wherein the communication interface is configured to support TCP/IP socket communication (e.g. paragraphs 0223-0224, training inputs including images such as training sketches, training images, and other information such as drawings, photographs, video files, screen recordings, images of existing GUIs, etc.; training information also including which GUIs are represented and which components or features are included in each GUI to provide baseline to be used for classifying the various images; paragraph 0247, user systems communicating with system using TCP/IP). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Petursson, Bigham, and Corwin in front of him to have modified the teachings of Petursson (directed to automatically testing/learning the behavior of a system under test, including graphical user interface behavior) and Bigham (directed to pixel-based optimization for a user interface, such as by using a trained model), to incorporate the teachings of Corwin (directed to machine learning for translating captured input images into interactive demonstrations for software products, including graphical user interfaces) to include the capability for the communication interface to support TCP/IP socket communication. One of ordinary skill would have been motivated to perform such a modification in order to provide the capability to quickly modify existing production screens for design iteration and testing, without writing code, in as high fidelity as possible as described in Corwin (paragraph 0080). With respect to claim 29, Petursson in view of Bigham teaches all of the limitations of claim 22 as previously discussed. Petursson does not explicitly disclose wherein the transmitting the screen image comprises converting the captured screen image into a grayscale image or compressing the screen image to a predetermined resolution to reduce network traffic load via the communication interface. However, Corwin teaches wherein the transmitting the screen image comprises converting the captured screen image into a grayscale image or compressing the screen image to a predetermined resolution to reduce network traffic load via the communication interface (e.g. paragraph 0050, image file format may store data in a compressed format, or can be PNG format that supports grayscale images; paragraph 0117, SVG images can be compressed; paragraph 0162, PNG graphics file which supports lossless data compression and grayscale; paragraph 0164, optimizing SVG images, including compression; paragraphs 0223-0224, training inputs including images such as training sketches, training images, and other information such as drawings, photographs, video files, screen recordings, images of existing GUIs, etc.; training information also including which GUIs are represented and which components or features are included in each GUI to provide baseline to be used for classifying the various images). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Petursson, Bigham, and Corwin in front of him to have modified the teachings of Petursson (directed to automatically testing/learning the behavior of a system under test, including graphical user interface behavior) and Bigham (directed to pixel-based optimization for a user interface, such as by using a trained model), to incorporate the teachings of Corwin (directed to machine learning for translating captured input images into interactive demonstrations for software products, including graphical user interfaces) to include the capability for the convert the captured image to grayscale or compressed format. One of ordinary skill would have been motivated to perform such a modification in order to provide the capability to quickly modify existing production screens for design iteration and testing, without writing code, in as high fidelity as possible as described in Corwin (paragraph 0080). It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. “The use of patents as references is not limited to what the patentees describe as their own inventions or to the problems with which they are concerned. They are part of the literature of the art, relevant for all they contain,” In re Heck, 699 F.2d 1331, 1332-33, 216 USPQ 1038, 1039 (Fed. Cir. 1983) (quoting in re Lemelson, 397 F.2d 1006, 1009, 158 USPQ 275, 277 (GCPA 1968)). Further, a reference may be relied upon for all that it would have reasonably suggested to one having ordinary skill the art, including nonpreferred embodiments. Merck & Co, v. Biocraft Laboratories, 874 F.2d 804, 10 USPQ2d 1843 (Fed. Cir.), cert, denied, 493 U.S. 975 (1989). See also Upsher-Smith Labs. v. Pamlab, LLC, 412 F,3d 1319, 1323, 75 USPQ2d 1213, 1215 (Fed. Cir, 2005): Celeritas Technologies Ltd. v. Rockwell International Corp., 150 F.3d 1354, 1361, 47 USPQ2d 1516, 1522-23 (Fed. Cir. 1998). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JEREMY L STANLEY whose telephone number is (469)295-9105. The examiner can normally be reached on Monday-Friday from 9:00 AM to 5:00 PM CST. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abdullah Al Kawsar, can be reached at telephone number (571) 270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from Patent Center and the Private Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from Patent Center or Private PAIR. Status information for unpublished applications is available through Patent Center and Private PAIR for authorized users only. Should you have questions about access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) Form at https://www.uspto.gov/patents/uspto-automated- interview-request-air-form. /JEREMY L STANLEY/ Primary Examiner, Art Unit 2127
Read full office action

Prosecution Timeline

Jul 31, 2023
Application Filed
Dec 09, 2025
Response after Non-Final Action
Mar 31, 2026
Non-Final Rejection mailed — §103, §112
Jun 15, 2026
Examiner Interview Summary
Jun 15, 2026
Applicant Interview (Telephonic)
Jun 23, 2026
Response Filed
Sep 10, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725028
NEURAL NETWORK OBTAINING METHOD AND RELATED DEVICE
5y 6m to grant Granted Sep 01, 2026
Patent 12688456
SECURE AND FAIR COMPETITIVE BIDDING
4y 5m to grant Granted Jul 21, 2026
Patent 12688443
INFORMATION PROCESSING METHOD AND APPARATUS, AND COMPUTER-READABLE STORAGE MEDIUM
4y 3m to grant Granted Jul 21, 2026
Patent 12670362
VIDEO SYNTHESIS WITHIN A MESSAGING SYSTEM
4y 9m to grant Granted Jun 30, 2026
Patent 12670198
Textual Summaries In Information Systems Based On Personalized Prior Knowledge
3y 4m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
49%
Grant Probability
90%
With Interview (+41.0%)
3y 3m (~1m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 290 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month