Detailed Action
Notice of Pre-AIA or AIA status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
This Final Office action is responsive to the communication filed under 37 C.F.R. § 1.111 on August 4, 2026 (hereafter “Response”). The amendments to the claims are acknowledged and have been entered.
The Specification is now amended.
Claims 1, 3, 7–9, 14, and 20 are now amended.
New claims 21 and 22 are now added.
Claims 1–22 are pending in the application, of which claims 7–13 and 17 are withdrawn from consideration.
Response to Arguments
Applicant’s arguments, see Response 9–10, filed August 4, 2026, with respect to the objections to the specification and claims (and the amendments correcting both) have been fully considered and are persuasive. The objections to the disclosure and claims 3–4 have been withdrawn.
Applicant’s arguments, see Response 10–15, with respect to the prior art rejections have been fully considered. The Examiner is persuaded that the Applicant narrowed the scope of all of the pending claims beyond what Chiou may fairly anticipate under 35 U.S.C. § 102. Therefore, all prior art rejections have been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of the obviousness of the pending claims over Chiou in view of newly cited prior art.
Therefore, the Applicant’s request for an allowance is respectfully denied, as is the corresponding request for a rejoinder, as none of the nonelected claims incorporate the limitations of any allowed claim.
Claim Objections
Claim 3 and 4 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim Rejections – 35 U.S.C. § 103
The following is a quotation of 35 U.S.C. § 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned at the time any inventions covered therein were effectively filed absent any evidence to the contrary. Applicant is advised of the obligation under 37 C.F.R. § 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned at the time a later invention was effectively filed in order for the examiner to consider the applicability of 35 U.S.C. § 102(b)(2)(C) for any potential 35 U.S.C. § 102(a)(2) prior art against the later invention.
Claims 1, 2, 5, 6, 14–16, and 18–22 are rejected under 35 U.S.C. § 103 as being unpatentable over U.S. Patent Application Publication No. 2023/0074584 A1 (“Chiou”) in view of U.S. Patent Application Publication No. 2017/0001308 A1 (“Bataller”).
Claim 1
Chiou teaches:
A method of analyzing an application screen, the method comprising:
“To model keyboard navigation flow of a web page, a Keyboard Navigation Flow Graph (KNFG) is defined.” Chiou ¶ 61.
generating a plurality of links for a plurality of user interface (UI) elements included in the application screen;
“A keyboard navigation flow of a page under test is represented by a set of KNFGs,” Chiou ¶ 61, where “[t]he node set of a KNFG, comprises a node for each HTML element in the page under test,” Chiou ¶ 62, and where “[e]ach of the edges represents a keyboard entry such as Tab that may allow movement between nodes.” Chiou ¶ 66.
generating a UI map for each of at least one primitive action, which is a user input for navigating the application screen, based on the plurality of links; and
As shown in FIG. 2A, a keyboard navigation flow graph 200 is generated by “iteratively exploring the page using only keyboard based actions (i.e.,
Φ
K
) until no new navigation information is found (i.e., the graph has reached a fixed point).” Chiou ¶ 69. Specifically, the set of keyboard actions
Φ
K
may include “all standard keyboard commands used to navigate a web application's user interface as defined by W3C and web accessibility testing communities.” Chiou ¶ 69.
identifying positions of a focus indicating UI elements with which a user is to interact among the plurality of UI elements,
“After each action, the process analyzes the page to determine the focus transition that occurred.” Chiou ¶ 69. “After triggering an action
ϕ
∈
Φ
K
on a node
v
i
, the process detects the focus change from
v
i
to
v
i
+
1
and creates an edge in the graph
v
i
,
v
i
+
1
,
ϕ
,
δ
,
V
s
, indicating that the browser focus could shift from a source node
v
i
to a target node
v
i
+
1
by pressing keystroke ϕ while
v
i
is in focus.” Chiou ¶ 70.
Chiou differs from claim 1 in that Chiou identifies the positions of focus live, rather than “using at least one of an application screen history or a primitive action history” as claimed.
Bataller, however, teaches a method comprising:
identifying positions of a focus indicating UI elements with which a user is to interact among the plurality of UI elements, wherein the identifying of the positions of the focus is performed using at least one of an application screen history or a primitive action history
As shown in FIG. 1, Bataller provides a system 100 that analyzes user interface activity from a plurality of users in order to find repeatable processes that the users perform with the user interface. Bataller ¶ 17. Bataller does this in both of the claimed ways: with respect to using an application screen history, an image capturer 110 records screenshots of the users’ activity on their respective computers, Bataller ¶¶ 18–19, and then provides those images to an activity identifier 120 that applies “computer vision techniques to the images received from the image capturer 110 to identify one or more activities associated with the process.” Bataller ¶ 22. “For example, the activity identifier 120 may determine that based on the difference between the first image where the ‘Yes’ and ‘No’ buttons are not highlighted and the second image where the ‘Yes’ button is highlighted, the ‘Yes’ button was touched on a touchscreen. In another example, the activity identifier 120 may determine that in a first image a mouse cursor is over a menu icon that opens a menu and that in a second image the mouse cursor is still over the menu icon and that the menu is now open, and in response, determine that the menu icon was left mouse button clicked.” Bataller ¶ 23.
Regarding the primitive action history, “the activity identifier 120 may additionally identify activities using inputs besides the images. For example, the activity identifier 120 may receive keyboard inputs from a keyboard driver of the computer and in response identify an activity of a key press, or receive mouse inputs from a mouse driver of the computer and in response identify an activity of a mouse click.” Bataller ¶ 25.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to improve Chiou’s method with Bataller’s technique of using past application interactions to harvest
ϕ
∈
Φ
K
for the UI map, rather than relying on a live crawl of the same. One would have been motivated to utilize Bataller’s technique of analyzing actual user interaction data because Bataller’s technique “may enable more accurate automation,” since computer vision techniques make it possible to “determine when and where interactions should automatically occur even if buttons, controls, windows, or other user interface elements appear differently than when a manual process was performed,” e.g., where “different computers have different screen resolutions” (among other differences). Bataller ¶ 9.
Claim 2
Chiou and Bataller teach the method of claim 1, wherein the generating of the plurality of links comprises:
identifying the plurality of UI elements, based on the application screen;
“The example process identifies the nodes by rendering the page under test in a browser and then analyzing the DOM of the page under test to identify each unique HTML element.” Chiou ¶ 62.
generating a plurality of nodes corresponding to the plurality of UI elements; and
“Each node is uniquely identified by its XPath in the DOM.” Chiou ¶ 62.
generating the plurality of links, based on a node feature of the plurality of nodes, and wherein the node feature comprises at least one of features regarding sizes, positions, content, images, names, or hierarchy of the plurality of UI elements.
As understood by the Examiner, the act of generating the links “based on” the node features does not require the links themselves to represent node features, because, in this application, a “link” refers to a navigational pathway for the input focus to move between two nodes in the UI map. (Spec. ¶ 68). Instead, the above claim language merely requires one or more of the node features listed in the claim to influence the method’s assessment of whether there is a navigational “link” between two nodes.
For its part, the process that Chiou uses to generate edges between nodes is influenced by at least the nodes’ sizes, contents, names, and hierarchy.
With respect to sizes, Chiou discloses that “[t]he first iteration of this process begins by interacting with each node
v
∈
V
s
.” Chiou ¶ 69. Vs is defined as a subset of all nodes in the page under test with special predefined characteristics, one of which is elements that are not rendered due to having “a height or width of zero pixels.” Chiou ¶ 68.
With respect to content, Chiou further discloses that Vs may further include strictly “non-disabled elements that do not exhibit a final computed DOM layout style of type=‘hidden’, visibility:hidden, [or] display:none.” Chiou ¶ 68.
With respect to names, Chiou further discloses that “[s]yntactically linked nodes such as a <label> and its bounded form element and elements wrapped within other inline control elements are grouped, since these nodes are intended to represent a single functionality.” Chiou ¶ 62. In other words, irrespective of membership in Vs, Chiou’s method ensures that nodes having certain names like “label” are not eligible to receive edges, and groups them together with other nodes such that edges can only be directed to/from the whole group, rather than the individual node.
Finally, with respect to hierarchy, Chiou discloses taking into consideration whether a node “inherit[ed] their ancestor’s rendered hidden properties” for inclusion in the V-s subset.
Claim 5
Chiou and Bataller teach the method of claim 2,
wherein the generating of the UI map comprises generating the UI map by using an edge labeler, and wherein the edge labeler is a model trained to receive an edge for a plurality of nodes including the plurality of links to output the UI map.
“After triggering an action
ϕ
∈
Φ
K
on a node υi, the process detects the focus change from
v
i
to
υ
i
+
1
and creates an edge in the graph
v
i
,
v
i
+
1
,
ϕ
,
δ
,
V
S
, indicating that the browser focus could shift from a source node
υ
i
to a target node
υ
i
+
1
by pressing keystroke ϕ while
υ
i
is in focus.” Chiou ¶ 70. To be clear, in this rejection, the edge labeler corresponds to the “process” described above, which, given an edge in the graph, labels that edge with
ϕ
for the primitive that caused the transition.
Claim 6
Chiou and Bataller teach the method of claim 1, wherein the generating of the UI map comprises:
mapping the at least one primitive action for each of the plurality of links; and
“For a given
v
, the process first sets the browser's focus on and then executes every action in ΦK on
υ
.” Chiou ¶ 69. “After triggering an action
ϕ
∈
Φ
K
on a node υi, the process detects the focus change from
v
i
to
υ
i
+
1
and creates an edge in the graph
v
i
,
v
i
+
1
,
ϕ
,
δ
,
V
S
, indicating that the browser focus could shift from a source node
υ
i
to a target node
υ
i
+
1
by pressing keystroke ϕ while
υ
i
is in focus.” Chiou ¶ 70.
generating the UI map, based on what primitive action has been mapped to each of the plurality of links.
“[T]he intra-state edges 262 in the example keyboard navigation flow graph 200 in FIG. 2A are identified by iteratively exploring the page using only keyboard based actions (i.e., ΦK) until no new navigation information is found (i.e., the graph has reached a fixed point).” Chiou ¶ 69.
Claims 14–16
Claims 14–16 recite an electronic device with general purpose computer hardware for performing exactly the same method as set forth in claims 1, 2, and 6. Therefore, the findings from the rejection of claims 1, 2, and 6 are hereby reincorporated by reference, as applied to claims 14–16, taken in conjunction with Chiou’s additional disclosure of the general purpose computer system. See Chiou ¶¶ 114–116.
Claim 18
Chiou teaches the electronic device of claim 14,
further comprising a communication interface, wherein the at least one processor, when executing the at least one instruction, is further configured to: receive the application screen from an external electronic device through the communication interface,
Throughout Chiou’s disclosure, Chiou explains that its process is performed on a web application. See Chiou Title, Abstract, and ¶¶ 60–67; see Chiou ¶¶ 124–125 (teaching the claimed communication interface). As such, Chiou at least teaches that the application being modeled is one received from an external electronic device (i.e., a web server).
That said, Chiou’s system does not transmit at least one of the UI map or the positions of the focus back to the external electronic device through the communication interface.
Bataller, however, teaches a system comprising:
a communication interface,
“Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network.” Bataller ¶ 74
wherein the at least one processor, when executing the at least one instruction, is further configured to: receive the application screen from an external electronic device through the communication interface,
“The image capturer 110 may provide the obtained images to the activity identifier 120.” Bataller ¶ 21. “[T]he image capturer 110 may be a software process running on the computer or on another computer that monitors video output from the computer to the display of the computer.” Bataller ¶ 18. In other words, some embodiments call for running image capturer 110 on the computer of the person being monitored (the external electronic device) and transmitting those images to system 100 (specifically, activity identifier 120).
and transmit at least one of the UI map or the positions of the focus to the external electronic device through the communication interface.
“The process 300 may include storing the process definition and later accessing the process definition and automatically instructing the robot to interact with a computer based on the activities and activity information indicated by the process definition.” Bataller ¶ 60. The computer that the robot later interacts with may be the same computer that was initially recorded by image capturer 110 in order to automate the task in the first place. See Bataller Claim 1 (“obtaining images taken of a display of the computer while the user is interacting with the computer . . . [and] generating a process definition for use in causing the robot to automatically perform the process by interacting with the computer”) (emphasis added).
Claim 19
Chiou and Bataller teach the electronic device of claim 18, further comprising:
a display, wherein the application screen is displayed on the display as an execution screen of a third-party application not provided with an application program interface.
“In some implementations, a server transmits data (e.g., an HTML page) to a client device (e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device).” Chiou ¶ 124. “A web page 310 that rendered in a browser on a display portion of a computing system is input for analysis.” Chiou ¶ 76.
Claim 20
Claim 20 recites a computer readable medium programmed with the same instructions as the memory of claim 14. Therefore, the findings set forth in the rejection of claim 14 for Chiou’s disclosure of the memory also anticipate the computer readable medium of claim 20 for the same reasons.
Claim 21
Chiou and Bataller teach the non-transitory computer-readable recording medium of claim 20,
wherein the application screen history includes a plurality of application screens stored at predetermined time intervals
“The image capturer 110 may obtain images at various times. For example, the image capturer 110 may obtain an image at predetermined intervals, e.g., every one, five, twenty five, one hundred milliseconds, or some other interval.” Bataller ¶ 19.
and a plurality of application screens stored every time a primitive action is performed.
“The activity information generator 130 may generate activity information associated with the activity. Activity information may be information that describes the activity. For example, the activity information for a screen touch may describe the coordinates for a touch on a touch screen, a snapshot, e.g., a twenty pixel by twenty pixel area, fifty pixel by fifty pixel area, or some other size area, centered around the coordinates that are touched right before the touch screen was touched, and a screenshot of the display after the touch screen was touched. In some implementations, the activity information generator 130 may generate snapshots using intelligent cropping that may automatically determine an optimal size of a snapshot. For example, the activity information generator 130 may determine that a logo for a program on a taskbar has been selected and, in response, identify just the portion of the taskbar that corresponds to the logo and generate a snapshot including just the identified portion of the taskbar.” Bataller ¶ 26.
Claim 22
Chiou and Bataller teach the non-transitory computer-readable recording medium of claim 20,
wherein the application screen history and the primitive action history have a mapping relationship therebetween.
“The activity information generator 130 may generate activity information associated with the activity.” Bataller ¶ 26. The activity links the screenshots to the user’s inputs. For example, “the activity information for a key press may describe what key was pressed, a screenshot of a display before the key was pressed, and a screenshot of the display after the key was pressed,” or “the activity information for a mouse click may describe what button of a mouse was clicked, the coordinates of the mouse cursor when the mouse was clicked, a snapshot centered around the coordinates of the mouse cursor right before the mouse was clicked, and a screenshot of the display after the mouse is clicked.” Bataller ¶ 27.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Justin R. Blaufeld whose telephone number is (571)272-4372. The examiner can normally be reached M-F 9:00am - 4:00pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, James K Trujillo can be reached at (571) 272-3677. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
Justin R. Blaufeld
Primary Examiner
Art Unit 2151
/Justin R. Blaufeld/Primary Examiner, Art Unit 2151