Prosecution Insights
Last updated: October 01, 2026
Application No. 18/934,902

ANALYZING GRAPHICAL USER INTERFACES TO FACILITATE AUTOMATIC INTERACTION

Non-Final OA §103
Filed
Nov 01, 2024
Priority
Jan 31, 2020 — nonprovisional of PCTUS2020016143 +2 more
Examiner
BADAWI, ANGIE M
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
59%
Grant Probability
Moderate
1-2
OA Rounds
2y 2m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 59% of resolved cases
59%
Career Allowance Rate
173 granted / 292 resolved
-0.8% vs TC avg
Strong +38% interview lift
Without
With
+37.6%
Interview Lift
resolved cases with interview
Typical timeline
4y 1m
Avg Prosecution
12 currently pending
Career history
307
Total Applications
across all art units

Statute-Specific Performance

§101
11.8%
-28.2% vs TC avg
§103
49.3%
+9.3% vs TC avg
§102
13.6%
-26.4% vs TC avg
§112
22.9%
-17.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 292 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 2, 7, 8, 13 & 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over KANG et al. (U.S. Pub 2020/0020334) hereinafter Kang, in view of Vangen et al. (U.S. Pub 2019/0243883) hereinafter Vangen. As per Claim 1, Kang teaches A method implemented using one or more processors, comprising: identifying a target visual cue to be located in a graphical user interface (“GUI”) comprising an interactive webpage accessible at a uniform resource locator (“URL”); (Fig. 5A, Fig. 5B, ¶106, ¶107, ¶108 wherien in a case where the first application program is a web browsing application, a screen downloaded from a server corresponding to an access URL may be transitorily or non-transitorily stored, and the first user interface may be included in the downloaded screen wherein the first user interface may be a text box provided to allow the user to input text. Or, the first user interface may include a keyboard (e.g., a virtual keyboard) for selecting a character input to, e.g., a text box wherien the electronic device 101 may determine whether the displayed screen includes the first user interface, such as a text box, using, e.g., identification information about the displayed screen or a result of analyzing screens displayed) the interactive webpage comprises one or more interactive elements; (Fig. 4A, Fig. 5A-5B, ¶106, , ¶108, ¶110 wherien in a case where the first application program is a web browsing application, a screen downloaded from a server corresponding to an access URL the electronic device 101 may display a second execution screen 510 corresponding to a particular URL, and the second execution screen 510 may include the first user interface, such as a text box 511 and a keyboard 512 for text entry to the text box 511) and automatically populating the one or more identified interactive elements with data. (Fig. 5A, ¶113 wherein the electronic device 101 may obtain the first user utterance 501 “Register Study schedule on February second.” The electronic device 101 may send the data about the first user utterance 501 to the external server, and the external server may apply ASR to the received data, thus obtaining the text “Register, Study schedule on February, second.” The external server may generate a command including tasks to execute a schedule management application and register the schedule of “Study” on February 2 on the schedule management application, corresponding to the first user utterance 501 from the obtained text using the intelligence system. The external server may send the generated command to the electronic device 101, and the electronic device 101 may perform the task included in the command. As shown on the right side of FIG. 5A, the electronic device 101 may display an execution screen 520 of the schedule management application and display the result of registering the schedule 522 of “Study” on the February 2 item 521) However, Kang does not explicitly teach obtaining a document object model (“DOM”) of the interactive webpage, wherein the DOM of the interactive webpage comprises one or more interactive elements; obtaining a bitmap screenshot of the GUI; applying features of the bitmap screenshot and the DOM as inputs across a machine learning model to generate output; based on the output, identifying one or more of the interactive elements of the GUI as corresponding to the target visual cue; Vangen teaches identifying a target visual cue to be located in a graphical user interface (“GUI”); (¶189 wherien use computer vision object detection and classification on the web page bitmap to extract the location of the play icon on the video player's player controls bar) obtaining a document object model (“DOM”) of the interactive webpage, wherein the DOM of the interactive webpage comprises one or more interactive elements; (¶53 wherein the Document Object Model (DOM) for the web page is analyzed to identify the present (and/or absence) of particular DOM items for the web page) obtaining a bitmap screenshot of the GUI; (¶47, ¶48 wherein The elements of interest in a web page can be identified from a visual analysis of the web page from a visual analysis of the bitmap representation of the rendered web page this can be done using appropriate computer vision techniques, such as pattern matching, optical character recognition, etc.) applying features of the bitmap screenshot and the DOM as inputs across a machine learning model to generate output; (¶37, ¶185 wherein the elements in a web page that can be identified by hyperlinks; (e.g. bit maps of) user interface icons such as play, pause, forward, rewind symbols, buttons, menu items, friends list, etc.; (e.g. bit maps of) text (words); data entry fields; areas or regions of a web page; HTML elements; CSS elements; images; videos; Adobe Flash, Microsoft Silverlight or other plug-ins; DOM trees wherein by parsing the DOM (Document Object Model) for the web page to detect the presence and visibility of particular DOM items and then determining the state based on the presence and visibility of (or lack of) one or more DOM items. The bitmap for the web page may also be processed (analyzed) with computer vision to detect and classify particular objects in the supplied bitmap, with a presence (or lack of) of one or more objects being used to determine the state of the web page.) based on the output, identifying one or more of the interactive elements of the GUI as corresponding to the target visual cue; (¶49, ¶186, ¶189 wherien such visual analysis can be performed as desired, e.g. by scanning the bitmap to be displayed for the web page, to identify, visual elements, such as text (words), symbols, icons, etc., that represent elements within the web page, to identify whether a web page contains any particular, visual elements (e.g. icons or text) that could correspond to desired elements of interest within the web page wherein a neural network or machine learning based classifier may be fed with data extracted using one or more of these methods and used to detect and classify the state of a web page wherein wherien use computer vision object detection and classification on the web page bitmap to extract the location of the play icon on the video player's player controls bar, and a user input actuator emulator to then move the mouse to the location of the play button and to emulate a left mouse click to actuate the play button) It would have been obvious to one having ordinary skill in the art at the time the invention was filed to utilize the teaching of a browser module configured to retrieve web pages from the Internet and an analysis module operable to analyze a retrieved web page to identify elements of interest in the web page of Vangen with the teaching of processing user speech of Kang because Vangen teaches an improved machine learning module that can learn from analysis of and interactions with web pages how to improve its operation, such as how to better identify elements of interest in a web page and/or how to better interact with a web page and analyze a user's interactions with web pages and web services that they access, and to correspondingly adapt and improve its operation based on that analysis. (¶109, ¶110) As per Claim 2, the rejection of claim 1 is hereby incorporated by reference; Kang as modified further teaches further comprising determining a user intent to interact with the GUI based at least in part on a free-form natural language input; and (Fig. 4, Fig. 5A, ¶112 wherien the electronic device 101 may receive a first user utterance 501 through the microphone 280 and provide data about the first user utterance 501 to an external server including an ASR system and an intelligence system. The intelligence system may apply natural-language understanding to a text obtained by, e.g., the ASR system and determine, e.g., the user's intent, thereby generating a command including a task corresponding thereto; as taught by Kang) based on the user intent, identifying the target visual cue. (Fig. 5A, Fig. 5B, ¶106, ¶107, ¶108, ¶116 wherien in a case where the first application program is a web browsing application, a screen downloaded from a server, and the first user interface may be included in the downloaded screen wherein the first user interface may be a text box provided to allow the user to input text wherien The external server may send the obtained text to the electronic device 101, and the electronic device 101 may display at least part 513 of the obtained text in the text box 511 as shown on the right side of FIG. 5B. According to various embodiments of the present invention, the electronic device 101 may be configured to input the text received from the external server to the first user interface based on the state information indicating that the first user interface is being displayed. According to various embodiments of the present invention, the external server may determine whether to obtain a text through ASR on the data received from the electronic device 101 and send the same to the electronic device 101 or to obtain a text through ASR on the data received from the electronic device 101, generate a command including a task from the obtained text, and send the same to the electronic device 101, according to the state information about the electronic device associated with whether the first user interface is displayed; as taught by Kang) Claim 7 is similar in scope to Claim 1; therefore, Claim 7 is rejected under the same rationale as Claim 1. Claim 8 is similar in scope to Claim 2; therefore, Claim 8 is rejected under the same rationale as Claim 2. Claim 13 is similar in scope to Claim 1; therefore, Claim 13 is rejected under the same rationale as Claim 1. Claim 14 is similar in scope to Claim 2; therefore, Claim 14 is rejected under the same rationale as Claim 2. Claims 3-5, 9-11 & 15-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kang in view of Vangen as applied to claim 1 above, and further in view of Nolan et al. (U.S. Pat 9,954,729) hereinafter Nolan. As per Claim 3, the rejection of claim 1 is hereby incorporated by reference; Kang as modified further teaches further comprising: generating, and storing in association with the URL of the interactive webpage, a script (Fig. 5A, ¶108, ¶185 wherein in a case where the first application program is a web browsing application program, a first execution screen 500 corresponding to a particular URL may be displayed wherein JavaScript may be injected in the web page and used to test for the presence of particular JavaScript and/or HTML elements (functions and/or objects), and their attributes, in order to determine their current state, with the presence and/or state of one or more of these elements being used to determine the particular state of the web page; as taught by Vangen) that is subsequently executable in association with the interactive webpage and a subsequent free-form natural language input to trigger subsequent automatic population of the one or more identified interactive elements with data determined from a subsequent user intent determined from the subsequent free-form natural language input. (Fig. 5B, ¶113,¶116 wherein the electronic device 101 may obtain the second user utterance 503 “Register Study schedule on February second.” The electronic device 101 may send the data about the second user utterance 503 to the external server, and the external server may apply ASR to the received data, the external server may include an automatic speech recognition (ASR) system capable of generating text using data about an utterance and an intelligence system capable of natural-language understanding text, grasping the meaning of the text, and generating a command corresponding to the text, thus obtaining the text “Register, Study schedule on February, second.” The external server may send the obtained text to the electronic device 101, and the electronic device 101 may display at least part 513 of the obtained text in the text box 511 as shown on the right side of FIG. 5B. According to various embodiments of the present invention, the electronic device 101 may be configured to input the text received from the external server to the first user interface based on the state information indicating that the first user interface is being displayed; as taught by Kang) However, Kang as modified does not explicitly teach further comprising: validating that submission of the data resulted in a next state of the interactive webpage; and in response to the validating, generating a script. (Fig. 5, col. 8 lines 38-67 wherein a test is conducted to determine whether the results of the validation rules is positive and if the results of the validation are positive, a user at the client 102 can optionally request the generation of the script subsequent to a confirmation of a valid processing) It would have been obvious to one having ordinary skill in the art at the time the invention was filed to utilize the teaching of provisioning and configuration of network infrastructure of Nolan with the teaching of processing user speech of Kang as modified because Nolan teaches an improved tool that utilizes pre-configured templates to collect information utilized in the configuration of the infrastructure equipment and automatically generate configuration scripts. The tool dramatically increases the ability to configure or re-configure infrastructure equipment. (Abstract) As per Claim 4, the rejection of claim 3 is hereby incorporated by reference; Kang as modified further teaches wherein the next state comprises a subsequent webpage that is generated at least in part on the data used to automatically populate the one or more identified interactive elements. (Fig. 5A, ¶113 wherien the electronic device 101 may obtain the first user utterance 501 “Register Study schedule on February second.” The electronic device 101 may send the data about the first user utterance 501 to the external server, and the external server may apply ASR to the received data, thus obtaining the text “Register, Study schedule on February, second.” The external server may generate a command including tasks to execute a schedule management application and register the schedule of “Study” on February 2 on the schedule management application, corresponding to the first user utterance 501 from the obtained text using the intelligence system. The external server may send the generated command to the electronic device 101, and the electronic device 101 may perform the task included in the command. As shown on the right side of FIG. 5A, the electronic device 101 may display an execution screen 520 of the schedule management application and display the result of registering the schedule 522 of “Study” on the February 2 item 521; as taught by Kang) As per Claim 5, the rejection of claim 4 is hereby incorporated by reference; Kang as modified further teaches wherein the validating comprises searching a URL of the subsequent webpage to determine the next state. (Fig. 5, col. 8 lines 38-67 wherein a test is conducted to determine whether the results of the validation rules is positive and if the results of the validation are positive, a user at the client 102 can optionally request the generation of the script subsequent to a confirmation of a valid processing; as taught by Nolan; ¶188 wherein identifiable web page states have associated with them one or more instructions for converting the web page state into another state. These instructions could be to perform relatively simple interactions, such as navigating the web page to a new url; as taught by Vangen) Claim 9 is similar in scope to Claim 3; therefore, Claim 9 is rejected under the same rationale as Claim 3. Claim 10 is similar in scope to Claim 4; therefore, Claim 10 is rejected under the same rationale as Claim 4. Claim 11 is similar in scope to Claim 5; therefore, Claim 11 is rejected under the same rationale as Claim 5. Claim 15 is similar in scope to Claim 3; therefore, Claim 15 is rejected under the same rationale as Claim 3. Claim 16 is similar in scope to Claim 4; therefore, Claim 16 is rejected under the same rationale as Claim 4. Claim 17 is similar in scope to Claim 5; therefore, Claim 17 is rejected under the same rationale as Claim 5. Claims 6, 12 & 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kang in view of Vangen as applied to claim 1 above, and further in view of PHAM et al. (U.S. Pub 2021/0081475) hereinafter Pham. As per Claim 6, the rejection of claim 1 is hereby incorporated by reference; Kang as modified previously taught the machine learning model. However, Kang as modified does not explicitly teach wherein the machine learning model comprises a convolutional neural network. Pham teaches wherein the machine learning model comprises a convolutional neural network. (Fig. 1, ¶46, ¶47, ¶62 wherein content generation subsystem 116 may train a prediction model, such as a machine learning model wherein the prediction model may be trained using training data including the initially accessed websites, the selected text from the initially accessed websites, and the subsequently accessed websites wherein the prediction model may include one or more neural networks wherein image item 204 may include an image of an object. Image item 204 may be analyzed using an object recognition computer vision model to determine the object included within the image, and a topic associated with the object may be determined by the object recognition computer vision model. In some embodiments, the object recognition computer vision model may be a convolutional neural network ) It would have been obvious to one having ordinary skill in the art at the time the invention was filed to utilize the teaching of integrating content into web pages of Pham with the teaching of processing user speech of Kang as modified because Pham teaches an improved apparatus and system for integrating content into one or more online resources wherein a first request for a reference identifier (e.g., URL or other reference identifier) to be embedded into first content on a first website may be obtained based on a first user accessing the first website. Based on the first request, first interaction data related to the first website may be retrieved. The first interaction data may indicate that a prior user interacted with text on the first website and subsequently accessed a second website. A first reference identifier related to the second website may be caused to be embedded into the text on the first website based on: (i) the second website comprising second content related to the text, (ii) the prior user interacting with the text on the first website, and (iii) the prior user accessing the second website after interacting with the first content on the first website. (¶4, ¶5) Claim 12 is similar in scope to Claim 6; therefore, Claim 12 is rejected under the same rationale as Claim 6. Claim 18 is similar in scope to Claim 6; therefore, Claim 18 is rejected under the same rationale as Claim 6. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANGIE BADAWI whose telephone number is (571)270-7590. The examiner can normally be reached Monday thru Wednesday 9:00am - 5:00pm EST with Thursdays and Fridays off. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Fred Ehichioya can be reached at (571) 272-4034. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ANGIE BADAWI/ Primary Examiner, Art Unit 2179
Read full office action

Prosecution Timeline

Nov 01, 2024
Application Filed
Sep 18, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748524
SCREENSHOT METHOD, ELECTRONIC DEVICE AND COMPUTER-READABLE STORAGE MEDIUM
3y 0m to grant Granted Sep 29, 2026
Patent 12717470
USER INTERFACE FOR DISPLAYING AND MANAGING WIDGETS
2y 11m to grant Granted Aug 25, 2026
Patent 12656940
DATA PROCESSING METHOD AND APPARATUS, DEVICE, MEDIUM, AND PROGRAM PRODUCT
2y 10m to grant Granted Jun 16, 2026
Patent 12638967
METHOD FOR DEVICE CONTROL, ELECTRONIC DEVICE, AND STORAGE MEDIUM
3y 5m to grant Granted May 26, 2026
Patent 12554394
SYSTEM AND METHOD FOR PROMOTING CONNECTIVITY BETWEEN A MOBILE COMMUNICATION DEVICE AND A VEHICLE TOUCH SCREEN
10y 5m to grant Granted Feb 17, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
59%
Grant Probability
97%
With Interview (+37.6%)
4y 1m (~2y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 292 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month