Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This communication is responsive to the amendment filed on 02/13/2026.
Status of claims:
Claims 3 and 13 are canceled.
Claims 1-2, 4-6, 8-12, 14-16 and 18-20 are amended.
Claims 1-2, 4-12 and 14-20 are pending for examination.
Remarks
Applicant argued regarding to the amended claims have been considered in view of the new ground(s) of rejection necessitated by amendment. Please refer the rejection below.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-2, 4-12 and 14-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract without significantly more.
Claims 1 and 11:Step 1: Statutory Category
The claims are directed to one of the four statutory categories of invention, i.e., process, machine, manufacture, or composition of matter.
Step 2A, Prong One
The claim recites the following limitations directed to an abstract idea:
“by the first or the second browser, determining whether the retrieved second content references the resource at the second address based on the retrieved second content comprising a file name associated with the resource” is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind, but for the recitation of generic computer components. One can mentally or manually with the aid of pen and paper determine whether the retrieved second content references the resource at the second address based on the retrieved second content comprising a file name associated with the resource.
“querying a local metadata database based on the file name to retrieve metadata associated with the resource, wherein the metadata indicated that the resource exists in the local file storage” is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind, but for the recitation of generic computer components. One can mentally or manually with the aid of pen and paper query a local metadata database based on the file name to retrieve metadata associated with the resource).
The limitation “determining, based on the metadata, whether and the resource is stored unexpired in the local file storage” is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind, but for the recitation of generic computer components. One can mentally or manually with the aid of pen and paper determine whether and the resource is stored unexpired in the local file storage.
At Step 2A, Prong Two:
The claim recites the following additional elements:
"implemented by a computing system" is a high-level recitation of a generic computer components and represents mere instructions to apply on a computer as in MPEP 2106.05(f), which does not provide integration into a practical application.
“in response to a first request to scrape, the first request comprising a first address of a target webpage on the Internet: by a first browser, retrieving a first content located at the first address, wherein the first content specifies a resource located at a second address, the resource needed to assemble the target webpage” amounts to data-gathering steps which is considered to be insignificant extra-solution activity, (See MPEP 2106.05(g));
“by the first browser, retrieving the resource from the Internet at the second address” amounts to data-gathering steps which is considered to be insignificant extra-solution activity, (See MPEP 2106.05(g)).
“by a second browser different from the first browser, retrieving a second content located at the first address” amounts to data-gathering steps which is considered to be insignificant extra-solution activity, (See MPEP 2106.05(g));
“when the resource is determined to be stored unexpired in the local file storage, retrieving the resource from the local file storage to avoid needing to retrieve the resource from the Internet to assemble the target webpage” amounts to data-gathering steps which is considered to be insignificant extra-solution activity, (See MPEP 2106.05(g)).
“storing the retrieved resource in a local file storage; and in response to a second request from a client device to scrape, the second request comprising the first address of the target webpage on the Internet” recites insignificant extra-solution activity such as mere outputting of the result. The mere outputting of data does not meaningfully limit the abstract idea. Viewing the additional limitations together and the claim as a whole, nothing provides integration into a practical application. (See MPEP 2106.05 (g)).
“transmitting the resource to the client device” recites insignificant extra-solution activity such as mere outputting of the result. The mere outputting of data does not meaningfully limit the abstract idea. Viewing the additional limitations together and the claim as a whole, nothing provides integration into a practical application. (See MPEP 2106.05 (g)).
“processor, memory and non-transitory computer-readable device” are recited at a high level of generality such that they amount to on more than mere instructions to apply the exception using a generic component. (see MPEP 2106.05(f)). These limitations can also be viewed as nothing more than an attempt to generally link the use of the judicial exception to the technological environment of a computer (see MPEP 2106.05(h)). Note, the mere instructions to apply an exception on a generic computer cannot integrate a judicial exception into a practical application.
Step 2B:
The conclusions for the mere implementation using a computer are carried over and does not provide significantly more.
With respect to the "retrieving… steps” identified as insignificant extra-solution activity above when re-evaluated this element is well-understood, routine, and conventional as evidenced by the court cases in MPEP 2106.05(d)(II), "i. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); … OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network);" and thus remains insignificant extra-solution activity that does not provide significantly more.
With respect to the “storing … step” identified as insignificant extra-solution activity above when re-evaluated this element is well-understood, routine, and conventional as evidenced by the court cases in MPEP 2106.05(d)(II), " iv. Presenting offers and gathering statistics, OIP Techs., 788 F.3d at 1362-63, 115 USPQ2d at 1092-93" and "i. … transmitting data over a network, …Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); … OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network)".
With respect to the “non-transitory computer readable storage medium” amount to elements that have been recognized as well-understood, routine, and conventional activity in particular fields, as demonstrate by: Relevant court decision: the followings are examples of court decisions demonstrating well-understood, routine and conventional activities, see e.g., MPEP 2106.05(d)(II) and MPEP 2106.05(f)(2): Computer readable storage media comprising instructions to implement a method, e.g., see Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015).
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements when considered both individually and as an ordered combination do not amount to significantly more than the abstract idea.
Looking at the claim as a whole does not change this conclusion and the claim appears to be ineligible.
Accordingly, claims 1 and 11 are directed to an abstract idea.
Claims 2, 4 and 7 recites the limitation, which is no additional elements recited so the claim does not provide a practical application and is not considered to be significantly. The same rationale applies claims 12, 14 and 17.
Claims 5-6 and 8-10 recite the additional limitations, which are recited at a high level of generality and would function in its ordinary capacity for re-retrieving the resource from the Internet at the second address to assemble the target webpage; and storing the re-retrieved resource in the local file storage, this additional element does not integrate the integrate the judicial exception into a practical application and does not amount to significantly more. The same rationale applies claims 15-16 and 18-20.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Vilcinskas et al., (US 2023/0018983), hereinafter “Vilcinskas”, in view of Suckel (US 2022/0321671), hereinafter “Suckel”.
Claim 1, Vilcinskas discloses a method for caching web resources during a scraping operation (par. [0058], web scraping system caches scraped data), comprising:
in response to a first request to scrape, the first request comprising a first address of a target webpage on the Internet (par. [0014], a web scraping request from a client computing device is received):
(a) by a first browser, retrieving a first content located at the first address, wherein the first content specifies a resource located at a second address, the resource needed to assemble the target webpage (par. [0042], [0052] and [0054], formulate the series of requests needed to obtain the desired content and it sends the request to a web proxy by accepting the request forwards the request to target web server);
(b) by the first browser, retrieving the resource from the Internet at the second address (par. [0042], and [0054], generating a response to the request, target web server 108 sends the response back to the web proxy that forwarded the request, which in turn forwards the response to web scraping system); and
(c) storing the retrieved resource in a local file storage (par. [0119]-[0120] and [0147], stores the results in job database), and
in response to a second request from a client device to scrape (par. [0046], generating a second request from other data received in response to the first request),
the second request comprising the first address of the target webpage on the Internet (par. [0047], reproducing a series of HTTP requests and responses to scrape data from the target website):
(d) by a second browser different from the first browser, retrieving a second content located at the first address (par. [0080] and [0148], generating the next request in the sequence or requests needed to retrieve the requested content).
Vilcinskas fails to discloses “by the first or the second browser, determining whether the retrieved second content references the resource at the second address based on the retrieved second content comprising a file name associated with the resource; querying a local metadata database based on the file name to retrieve metadata associated with the resource, wherein the metadata indicated that the resource exists in the local file storage; and determining, based on the metadata, whether the resource is stored unexpired in the local file storage"
Meanwhile, Suckel discloses (e) by the first or the second browser, determining whether the retrieved second content references the resource at the second address based on the retrieved second content comprising a file name associated with the resource (par. [0018], web scraping includes the following steps—a) retrieving Hypertext Markup Language (HTML) data from a website; b) parsing the data for target information; c) saving target information; d) repeating the process if needed on another page. A program that is designed to do all of these steps is called a web scraper. Another related program known as the web crawler (also known as a web spider) is a program or an automated script which performs the first task, i.e. it navigates the web in an automated manner to retrieve raw HTML data of the accessed web sites (the process also known as indexing);
(f) querying a local metadata database based on the file name to retrieve metadata associated with the resource, wherein the metadata indicated that the resource exists in the local file storage (par. [0025], [0073] and [0086]-[0087], Upon receiving the request from User Device 102, FE Proxy 106 forwards the request to Proxy Supernode 108, which checks the request and chooses a suitable exit node pool by accessing the Pool Database 110. After choosing a suitable pool, Proxy Supernode 108 retrieves and checks the metadata of exit nodes belonging to the chosen exit node pool. The retrieved metadata contains the quality rates (Q.sub.r) and available capacity (C.sub.avail) values of each exit node in the respective pool. Proxy Supernode 108 analyzes the retrieved metadata to select an exit node to service the user request. In one of the embodiments, from the retrieved metadata, Proxy Supernode 108 identifies the exit nodes with greater than zero available capacity (C.sub.avail) values, i.e., C.sub.avail>0. After which, Proxy Supernode 108 arranges the identified exit nodes according to their respective quality rating (Q.sub.r) values in a descending order, i.e., beginning with the highest Q.sub.r value. By identifying and arranging the exit nodes with available capacity (C.sub.avail) values greater than zero, Proxy Supernode 108 can isolate the exit nodes with zero available capacity (C.sub.avail) values. Proxy Supernode 108 selects an exit node with the highest quality rate (Q.sub.r) value from the arranged list of exit nodes. If there are multiple exit nodes with the highest quality rate (Q.sub.r) value, then Proxy Supernode 108 selects an exit node with the highest quality rate (Q.sub.r) values at random);
(g) determining, based on the metadata, whether the resource is stored unexpired in the local file storage (par. [0101], [0104], [0108] and [0146]-[0150], determine the number of requests that can still be executed concurrently by each exit node while avoiding potential failures or being blocked by the target. Therefore, after computing available capacity (C.sub.avail) values, in step 407 Proxy Supernode 108 reports the computed available capacity (C.sub.avail) values for each exit node according to their pool classification to Pool Database; and choosing a suitable exit node pool conforming to requirements of the request; [0149] retrieving and analyzing metadata of exit nodes belonging to the chosen suitable exit node pool, wherein the metadata retrieved contains quality rates (Q.sub.r) and available capacity values (C.sub.avail) of each exit node in the pool);
(h) when the resource is determined to be stored unexpired in the local file storage, retrieving the resource from the local file storage to avoid needing to retrieve the resource from the Internet to assemble the target webpage (par. [0156]-[0162], choosing a suitable exit node pool conforming to the requirements of the user request; retrieving and checking metadata of exit nodes belonging to the chosen suitable exit node pool, wherein the metadata retrieved contains quality rates (Q.sub.r) and available capacity values (C.sub.avail) of each exit node in the pool; analyzing the quality rate (Q.sub.r) values; analyzing the available capacity (C.sub.avail) values; selecting the exit node from the chosen pool with a highest quality rate (Q.sub.r) and a highest available capacity (C.sub.avail) value); and
(i) transmitting the resource to the client device (par. [0083], [0087] and [0090]-[0091], select proxy servers to route user requests for data extraction, by evaluating proxy servers' performance quality and the capacity to execute concurrent connections).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Vilcinskas with for the proxy server resources of Suckel, in order to improve proxy services, especially to select proxy servers to route user requests for data extraction, thereby evaluating proxy servers' performance quality and the capacity to execute concurrent connections.
Claim 2, the combination of Vilcinskas and Suckel discloses the invention as claimed. In addition, Suckel discloses the determining (e) comprises determining that the retrieved second content refers to a file name stored for the resource in the local file storage (par [0018], web scraping includes the following steps—a) retrieving Hypertext Markup Language (HTML) data from a website; b) parsing the data for target information; c) saving target information; d) repeating the process if needed on another page. A program that is designed to do all of these steps is called a web scraper. Another related program known as the web crawler (also known as a web spider) is a program or an automated script which performs the first task, i.e. it navigates the web in an automated manner to retrieve raw HTML data of the accessed web sites (the process also known as indexing).
Claim 4, the combination of Vilcinskas and Suckel discloses the invention as claimed. In addition, Suckel discloses the first request is specified by a first client and wherein the determining (g) comprises determining whether a time period associated with the first client has elapsed since the resource was retrieved in (b) (par. [0019], websites use to stop or slow down a web scraper since scraping may overload the website, but to identify the web scraper's IP address and block it to prevent further access by the bot. To do that, the website needs to identify the bot-like behavior of the web scraper and to identify its IP address).
Claim 5, the combination of Vilcinskas and Suckel discloses the invention as claimed. In addition, Suckel discloses when the time frame is expired:
(j) re-retrieving the resource from the Internet at the second address to assemble the target webpage (par. [0019], websites use to stop or slow down a web scraper since scraping may overload the website. For example, they may try to identify the web scraper's IP address and block it to prevent further access by the bot. To do that, the website needs to identify the bot-like behavior of the web scraper and to identify its IP address); and
(k) storing the re-retrieved resource in the local file storage (par. [0018], saving target information).
Claim 6, the combination of Vilcinskas and Suckel discloses the invention as claimed. In addition, Suckel discloses the querying (f) comprises the retrieved second content references another resource not stored in the local file storage (par. [0014], Web scraping is usually accomplished by a program that queries a web server and requests data automatically, then parses the data to extract the requested information),
(j)retrieving the other resource from the Internet to assemble the target webpage (par. [0019], websites use to stop or slow down a web scraper since scraping may overload the website. For example, they may try to identify the web scraper's IP address and block it to prevent further access by the bot. To do that, the website needs to identify the bot-like behavior of the web scraper and to identify its IP address); and
(k) storing the other resource in the local file storage (par. [0018], saving target information).
Claim 7, the combination of Vilcinskas and Suckel discloses the invention as claimed. In addition, Suckel discloses the resource is at least one of javascript, a stylesheets, a font, an image, or a video file (par. [0007]-[0015], Classifications of proxy servers are also based on protocols on which a particular proxy may operate. For instance, HTTP proxies, SOCKS proxies and FTP proxies are some of the protocol-based proxy categories. The term HTTP stands for Hypertext Transfer Protocol, the foundation for any data exchange on the Internet. Over the years, HTTP has evolved and extended, making it an inseparable part of the Internet. HTTP allows file transfers over the Internet and, in essence, initiates the communication between a client/user and a server. HTTP remains a crucial aspect of the World Wide Web because HTTP enables the transfer of audio, video, images, and other files over the Internet. HTTP is a widely adopted protocol currently available in two different versions—HTTP/2 and the latest one—HTTP/3).
Claim 8, the combination of Vilcinskas and Suckel discloses the invention as claimed. In addition, Suckel discloses the retrieving (a) and the retrieving (d) each occur through a proxy server (par. [0004]-[0021], employ proxy servers to maintain better network performance. Proxy servers can cache common web resources, so when a user requests a particular web resource, the proxy server will check to see if it has the most recent copy of the web resource, and then sends the user the cached copy. This can help reduce latency and improve overall network performance to a certain extent. Here, latency refers specifically to delays that take place within a network. In simpler terms, latency is the time between user action and the website's response or application to that action—for instance, the delay between when a user clicks a link to a webpage and when the browser displays that webpage).
Claim 9, the combination of Vilcinskas and Suckel discloses the invention as claimed. In addition, Suckel discloses the retrieving (a) and the retrieving (d) each occur through a residential proxy server (par. [0004]-[0021], employ proxy servers to maintain better network performance. Proxy servers can cache common web resources, so when a user requests a particular web resource, the proxy server will check to see if it has the most recent copy of the web resource, and then sends the user the cached copy. This can help reduce latency and improve overall network performance to a certain extent. Here, latency refers specifically to delays that take place within a network. In simpler terms, latency is the time between user action and the website's response or application to that action—for instance, the delay between when a user clicks a link to a webpage and when the browser displays that webpage).
Claim 10, the combination of Vilcinskas and Suckel discloses the invention as claimed. In addition, Suckel discloses the storing (c) comprises placing a request to store the retrieved resource in a queue for storage in the local file storage (par. [0004]-[0021], employ proxy servers to maintain better network performance. Proxy servers can cache common web resources, so when a user requests a particular web resource, the proxy server will check to see if it has the most recent copy of the web resource, and then sends the user the cached copy. This can help reduce latency and improve overall network performance to a certain extent. Here, latency refers specifically to delays that take place within a network. In simpler terms, latency is the time between user action and the website's response or application to that action—for instance, the delay between when a user clicks a link to a webpage and when the browser displays that webpage).
Claims 11-20, claims 11-20 are non-transitory computer-readable storage medium having stored therein instruction for executing the method of 1-10 above. They are rejected under the same rationale.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Loan T. Nguyen whose telephone number is (571) 270-3103. The examiner can normally be reached on Monday from 10:00 am - 6:00 pm, Thursday-Friday from 10:00 am - 2:00 pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aleksandr Kerzhner can be reached on (571) 270-1760. The fax phone number for the organization where this application or proceeding is assigned is 571-270-4103. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
07/10/2026
/LOAN T NGUYEN/Examiner, Art Unit 2165
/ALEKSANDR KERZHNER/Supervisory Patent Examiner, Art Unit 2165