Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Status of the Application
The following is a Final Office Action.
In response to Examiner's communication of 2/12/2026, Applicant responded on 5/29/2026. Amended claim 1, 9, 11, 19, and 20.
Claims 1-20 are pending in this application and have been examined.
Response to Amendment
Applicant's amendments to claims 1, 9, 11, 19, and 20 are not sufficient to overcome the 35 USC 101 rejections set forth in the previous action.
Applicant's amendments to claims 1, 9, 11, 19, and 20 are not sufficient to overcome the prior art rejections set forth in the previous action.
Response to Arguments – 35 USC § 101
Applicant’s arguments with respect to the rejections have been fully considered, but they are not persuasive.
Applicant submits, “…the claims recite a specific machine-learning architecture involving processor- generated encoded representations derived from activated leaf-node states across ensembles of decision trees. The claims do not merely recite evaluating information conceptually or making generalized predictions. Rather, the claims recite a specific computational architecture involving processor-based traversal of multiple decision trees, identification of activated leaf nodes across the decision-tree ensemble, generation of concatenated leaf-based encodings, generation of encoded representations associated with the activated leaf nodes, and calibration-model processing using those encoded representations.…the foregoing operations are not practically performable mentally. The claims require traversal and state identification across a plurality of decision trees for each data sample, generation of encoded representations derived from activated leaf-node states across the decision- tree ensemble, and concatenation of leaf-based encodings into machine-generated encoded representations used for downstream calibration processing. Such operations cannot realistically be performed mentally or with pen and paper...Here, the claimed architecture requires machine-learning ensemble traversal, encoded representation generation, and calibration-model processing that fundamentally depend upon processor-based computational operations…the claims recite an integration of the alleged abstract idea into a practical application…The specification explains that conventional machine-learning systems experience technical deficiencies associated with calibration of tree-based ensemble models. In particular, conventional systems are unable to effectively utilize structural information associated with activated leaf-node states across decision-tree ensembles for downstream calibration processing. The specification therefore discloses generation of encoded representations from activated leaf- node states across decision-tree ensembles to improve downstream calibration-model operations and calibrated prediction generation…In USPTO Example 48, the USPTO explained that claims directed to transforming intermediate machine-learning outputs into improved downstream machine-learning processing structures constitute patent-eligible practical applications. U.S. Pat. & Trademark Off., July 2024 Subject Matter Eligibility Examples, Ex. 48 (July 17, 2024). Here, similarly, the claims transform activated leaf-node states into encoded machine representations used for downstream calibration- model processing and calibrated prediction generation. The claims therefore recite a practical technological application…In Ex parte Desjardins, the PTAB further explained that claims directed to specific machine-learning architectural improvements are patent eligible where the claims reflect the disclosed technical solution identified in the specification. Ex parte Desjardins, Appeal No. 2024- 000567 (P.T.A.B. Sept. 26, 2025) (precedential). Here, the specification identifies technical deficiencies in calibration of tree-based ensemble systems and discloses the presently claimed encoded-representation architecture as the technical solution. The claims expressly recite those architectural improvements…the July 2024 AI SME Guidance explains that AI optimizations improving downstream technical outputs constitute practical applications under Step 2A, Prong 2. Here, the claimed encoding-generation architecture improves downstream calibration-model operation and calibrated prediction generation, thereby improving machine-learning system functionality itself…The claims recite a specialized machine-learning calibration architecture involving processor-based traversal of decision-tree ensembles, activated leaf-node identification, concatenated leaf-based encoding generation, generation of encoded representations derived from activated leaf-node states, and calibration-model processing using those encoded representations. The claims therefore recite significantly more than merely "applying" an abstract idea using generic computers…The Office Action provides no evidentiary support establishing that the claimed concatenated leaf-based encodings, activated leaf-node encoding representations, or calibration architectures were well-understood, routine, or conventional. Under Berkheimer, such conventionality findings require evidentiary support. Berkheimer v. HP Inc., 881 F.3d 1360 (Fed. Cir. 2018)…Under BASCOM, such non- conventional arrangements of otherwise known components constitute significantly more than an abstract idea. BASCOM Global Internet Services, Inc. v. AT&T Mobility LLC, 827 F.3d 1341 (Fed. Cir. 2016). For at least the foregoing reasons, the claims do not recite an abstract idea under Step 2B...” The Examiner respectfully disagrees.
While Applicant’ amendments furthers prosecution, unlike Example 48, Desjardins, July 2024 AI SME Guidance, BASCOM, by Applicant’s own admission, the claims indeed recite and direct to, …deficiencies associated with calibration of tree-based ensemble models…to effectively utilize structural information associated with activated leaf-node states across decision-tree ensembles for downstream calibration processing…generation of encoded representations from activated leaf-node states across decision-tree ensembles to improve downstream calibration-model operations and calibrated prediction generation…transforming intermediate…outputs into improved downstream…processing structures…transform activated leaf-node states into encoded…representations used for downstream calibration-model processing and calibrated prediction generation…, which is a problem directed to, a mental process, organizing human activity, as established in Step 2A Prong 1. Generating mathematical decision tree models, and calibrating mathematical decision tree models using a downstream mathematical calibration models to improve mathematical models’ predictions, is a mental and mathematical problem. This problem does not specifically arise in the realm of computer technology, but rather, this problem existed and was addressed long before the advent of computers. Thus, the claims do not recite a technical improvement to a technical problem. Additionally, pursuant to the broadest reasonable interpretation, as an ordered combination, each of the additional elements are computing elements recited at high level of generality implementing the abstract idea, and thus, are no more than applying the abstract idea with generic computer components, i.e. computer, machine encoding, performing extra solution activities, gathering data and outputting data, and generally linked to a technical environment, i.e. computer, machine encoding. Therefore, as a whole, the additional elements do not integrate the abstract ideas into a practical application in Step 2A Prong 2 (apply it and general link) or amount to significantly more in Step 2B (apply it and WURC).
Even novel and newly discovered judicial exceptions are still exceptions, despite their novelty. July 2015 Update, p. 3; see SAP America Inc. v. Investpic, LLC, No. 2017-2081, slip op. at 2 (Fed Cir. May 15, 2018).
Simply reciting specific limitations that narrow the abstract idea does not make an abstract idea non-abstract. 79 Fed. Reg. 74631; buySAFE Inc. v. Google, Inc., 765 F.3d 1350, 1355 (2014); see SAP America at p. 12. As discussed in SAP America, no matter how much of an advance the claims recite, when “the advance lies entirely in the realm of abstract ideas, with no plausibly alleged innovation in the non-abstract application realm,” “[a]n advance of that nature is ineligible for patenting.” Id. at p. 3.
Claims can recite a mental process even if they are claimed as being performed on a computer. The Supreme Court recognized this in Benson, determining that a mathematical algorithm for converting binary coded decimal to pure binary within a computer’s shift register was an abstract idea. The Court concluded that the algorithm could be performed purely mentally even though the claimed procedures “can be carried out in existing computers long in use, no new machinery being necessary.” 409 U.S at 67, 175 USPQ at 675. See also Mortgage Grader, 811 F.3d at 1324, 117 USPQ2d at 1699 (concluding that concept of “anonymous loan shopping” recited in a computer system claim is an abstract idea because it could be “performed by humans without a computer”).
Use of a computer or other machinery in its ordinary capacity for economic or other tasks (e.g., to receive, store, or transmit data) or simply adding a general purpose computer or computer components after the fact to an abstract idea (e.g., a fundamental economic practice or mathematical equation) does not integrate a judicial exception into a practical application or provide significantly more. See Affinity Labs v. DirecTV, 838 F.3d 1253, 1262, 120 USPQ2d 1201, 1207 (Fed. Cir. 2016) (cellular telephone); TLI Communications LLC v. AV Auto, LLC, 823 F.3d 607, 613, 118 USPQ2d 1744, 1748 (Fed. Cir. 2016) (computer server and telephone unit). Similarly, “claiming the improved speed or efficiency inherent with applying the abstract idea on a computer” does not integrate a judicial exception into a practical application or provide an inventive concept. Intellectual Ventures I LLC v. Capital One Bank (USA), 792 F.3d 1363, 1367, 115 USPQ2d 1636, 1639 (Fed. Cir. 2015).
TLI Communications provides an example of a claim invoking computers and other machinery merely as a tool to perform an existing process. The court stated that the claims describe steps of recording, administration and archiving of digital images, and found them to be directed to the abstract idea of classifying and storing digital images in an organized manner. 823 F.3d at 612, 118 USPQ2d at 1747. The court then turned to the additional elements of performing these functions using a telephone unit and a server and noted that these elements were being used in their ordinary capacity (i.e., the telephone unit is used to make calls and operate as a digital camera including compressing images and transmitting those images, and the server simply receives data, extracts classification information from the received data, and stores the digital images based on the extracted information). 823 F.3d at 612-13, 118 USPQ2d at 1747-48. In other words, the claims invoked the telephone unit and server merely as tools to execute the abstract idea. Thus, the court found that the additional elements did not add significantly more to the abstract idea because they were simply applying the abstract idea on a telephone network without any recitation of details of how to carry out the abstract idea.
Response to Arguments – Prior Art
Applicant’s arguments with respect to the rejections have been fully considered, but they are not persuasive. However, Applicant’s remarks are moot in light of new grounds of rejections necessitated by Applicant’s amendments.
Claim Rejections – 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 is/are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter.
Claim 1 (similarly 11, 20) recite, “A computer-implemented method for generating a prediction for an event, comprising:
accessing, by a …, a feature set corresponding to each data sample in an input dataset from a …;
generating, by a set of decision trees associated with the …, an intermediate prediction for the event based, at least in part, on applying the feature set of each data sample on the set of decision trees, each decision tree comprising a plurality of nodes;
identifying, by the …, an activated leaf node from each decision tree based, at least in part, on the intermediate prediction, the activated leaf node indicating a leaf node of one or more leaf nodes of the corresponding decision tree contributing to the intermediate prediction;
generating, by the …, a tree-based encoding indicating an encoded representation of the set of decision trees for each data sample based, at least in part, on an encoding type and the activated leaf node from each decision tree, wherein the tree-based encoding comprises a concatenated leaf-based encoding generated from activated leaf nodes of the set of decision trees, the tree-based encoding includes a position-type leaf-based encoding or a type-type leaf-based encoding, and the tree-based encoding comprises … to identify activated leaf nodes across the set of decision trees for the data sample; and
generating, by the …, the prediction for the event, wherein the prediction is calibrated by a calibration model associated with the … based, at least in part, on the tree-based encoding for each data sample.”
Analyzing under Step 2A, Prong 1:
The limitations regarding, …accessing, by a …, a feature set corresponding to each data sample in an input dataset from a …; generating, by a set of decision trees associated with the …, an intermediate prediction for the event based, at least in part, on applying the feature set of each data sample on the set of decision trees, each decision tree comprising a plurality of nodes; identifying, by the …, an activated leaf node from each decision tree based, at least in part, on the intermediate prediction, the activated leaf node indicating a leaf node of one or more leaf nodes of the corresponding decision tree contributing to the intermediate prediction; generating, by the …, a tree-based encoding indicating an encoded representation of the set of decision trees for each data sample based, at least in part, on an encoding type and the activated leaf node from each decision tree, wherein the tree-based encoding comprises a concatenated leaf-based encoding generated from activated leaf nodes of the set of decision trees, the tree-based encoding includes a position-type leaf-based encoding or a type-type leaf-based encoding, and the tree-based encoding comprises … to identify activated leaf nodes across the set of decision trees for the data sample; and generating, by the …, the prediction for the event, wherein the prediction is calibrated by a calibration model associated with the … based, at least in part, on the tree-based encoding for each data sample…, under the broadest reasonable interpretation, can include a human using their mind and using pen and paper to perform the above identified limitations, therefore, the claims recite a mental process.
Further, …accessing, by a …, a feature set corresponding to each data sample in an input dataset from a …; generating, by a set of decision trees associated with the …, an intermediate prediction for the event based, at least in part, on applying the feature set of each data sample on the set of decision trees, each decision tree comprising a plurality of nodes; identifying, by the …, an activated leaf node from each decision tree based, at least in part, on the intermediate prediction, the activated leaf node indicating a leaf node of one or more leaf nodes of the corresponding decision tree contributing to the intermediate prediction; generating, by the …, a tree-based encoding indicating an encoded representation of the set of decision trees for each data sample based, at least in part, on an encoding type and the activated leaf node from each decision tree, wherein the tree-based encoding comprises a concatenated leaf-based encoding generated from activated leaf nodes of the set of decision trees, the tree-based encoding includes a position-type leaf-based encoding or a type-type leaf-based encoding, and the tree-based encoding comprises … to identify activated leaf nodes across the set of decision trees for the data sample; and generating, by the …, the prediction for the event, wherein the prediction is calibrated by a calibration model associated with the … based, at least in part, on the tree-based encoding for each data sample…, are humans predicting events associated with a plurality of human users in Claim 10, which are managing interactions and relationship between people, therefore the claims recite certain methods of organizing human activities.
Accordingly, the claims are directed to a mental process, certain methods of organizing human activities, and thus, the claims are directed to an abstract idea under the first prong of Step 2A.
Analyzing under Step 2A, Prong 2:
This judicial exception is not integrated into a practical application under the second prong of Step 2A.
In particular, the claims recite the additional elements beyond the recited abstract idea identified under Step 2A, Prong 1, such as:
Claim 1, 11, 20: server system, database associated with the server system, A server system, comprising: a communication interface; a memory comprising executable instructions; and a processor communicably coupled to the communication interface and the memory, the processor configured to cause the server system, A non-transitory computer-readable storage medium comprising computer-executable instructions that, when executed by at least a processor of a server system, cause the server system to perform, encoded, a machine-generated encoded representation stored in memory and generated using a processor
, and pursuant to the broadest reasonable interpretation, as an ordered combination, each of the additional elements are computing elements recited at high level of generality implementing the abstract idea, and thus, are no more than applying the abstract idea with generic computer components.
Further, these additional elements generally link the abstract idea to a technical environment, namely the environment of a computer.
Additionally, with respect to, “…accessing…”, “…identifying…”, “…generating…”, these elements do not add a meaningful limitations to integrate the abstract idea into a practical application because they are extra-solution activity, pre and post solution activity - i.e. data gathering – “…accessing…”, “…identifying…”, data output – “…generating…”
Analyzing under Step 2B:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under Step 2B.
As noted above, the aforementioned additional elements beyond the recited abstract idea are not sufficient to amount to significantly more than the recited abstract idea because, as an order combination, the additional elements are no more than mere instructions to implement the idea using generic computer components (i.e. apply it).
Additionally, as an order combination, the additional elements append the recited abstract idea to well-understood, routine, and conventional activities in the field as individually evinced by the applicant’s own disclosure, as required by the Berkheimer Memo, in at least:
[0055]In one embodiment, the server system 102 is used by a managing entity to train the ML model such as the set of decision trees and use it for generating predictions related to a downstream task. In a non-limiting implementation, the managing entity may be any individual, representative of a person, an institution, an organization, a corporate entity, a non-profit organization, a financial institution, a bank, medical facilities (e.g., hospitals, laboratories, etc.), educational institutions, government agencies, telecom industries, weather forecast agency, or the like. In an example, the managing entity may be an administrator of the server system 102.
[0056]Examples of the downstream task include, but are not limited to, weather forecasting, speech recognition, image classification, email spam detection, performing medical diagnosis, fraud detection, risk management, charge-back decision-making systems, payment authorization systems, data analytics, credit card scoring systems, cross-border transaction management systems, consumer segmenting, or the like.
[0057]In a specific embodiment, the users (e.g., users 104) correspond to individuals whose data is used for training the models. For instance, the users 104 may be patients who are undergoing treatment for certain diseases. Data generated corresponding to such patients can be used to learn and understand the experience of the patients at a particular clinical center. Thus, such data is used to train AI or ML models to identify diseases and diagnoses. For example, classifying different diseases, such as cancer using images, predicting the progression of pre-diabetes, predicting response to depression treatment, etc. In another instance of a weather forecasting application (as shown in FIG. 12), the users 104 may correspond to individuals that provide information, such as location, date and time, preferences, alerts, activities, and the like. The information provided by such individuals can be used to generate predictions related to weather that are more personalized and actionable. For example, preferences influence how the weather forecast data is presented to the user (e.g., the user 104(1)), while activity details can enable the application to highlight relevant weather conditions (e.g., rain or wind for outdoor plans) for the user 104(1). In yet another instance of a payment industry (as shown in FIG. 11), the users 104 may be cardholders, account holders, merchants, consumers, issuers, acquirers, banks, third-party users, financial institutions, or the like. Data related to such individuals include historical financial transaction-related data, income-related data, expenditure-related data, and the like. Such data can be used to train AI or ML models to predict the income of an individual, predict financial frauds and risks, perform payment authorization operations, and the like.
[0058]In some embodiments, the users 104 may use their corresponding electronic devices (not shown in figures) to access a mobile application or a website associated with the issuing bank, or any third-party payment application to perform a payment transaction. In various non-limiting examples, the electronic devices may refer to any electronic devices, such as, but not limited to, Personal Computers (PCs), tablet devices, smart wearable devices, Personal Digital Assistants (PDAs), voice-activated assistants, Virtual Reality (VR) devices, smartphones, laptops, and the like.
[0059]Further, in a specific embodiment, the data sources 106 may correspond to various data sources that may be associated with the users 104 and are responsible for collecting and collating data related to the users 104 and their surroundings. Herein, the data sources 106 may act as a data provider for the server system 102 and a data collector for the users 104. In one embodiment, the data sources 106 can collect data from the users 104 via the network 110. It may be noted that the data sources 106 may be independent institutions that are independent of the users 104 or may be associated with the users 104. In another embodiment, the data sources 106 can include satellites, data gathering stations, sensors, and the like. Thus, the data sources 106 may directly collect data from its environment and provide it to the server system 102 upon receiving a request. Further, in an embodiment, the data sources 106 may be local storage units or cloud/remote storage units. In another embodiment, the data sources 106 may be owned by third-party organizations.
[00205]The disclosed method 1300 with reference to FIG. 13, or one or more operations of the server system 200 may be implemented using software including computer-executable instructions stored on one or more computer-readable media (e.g., non-transitory computer-readable media, such as one or more optical media discs, volatile memory components (e.g., Dynamic Random Access Memory (DRAM) or Statis Random Access Memory (SRAM)), or nonvolatile memory or storage components (e.g., hard drives or solid-state nonvolatile memory components, such as Flash memory components) and executed on a computer (e.g., any suitable computer, such as a laptop computer, netbook, Web book, tablet computing device, smartphone, or other mobile computing devices). Such software may be executed, for example, on a single local computer or in a network environment (e.g., via the Internet, a wide-area network, a local-area network, a remote web-based server, a client-server network (such as a cloud computing network), or other such networks) using one or more network computers. Additionally, any of the intermediate or final data created and used during the implementation of the disclosed methods or systems may also be stored on one or more computer-readable media (e.g., non-transitory computer-readable media) and are considered to be within the scope of the disclosed technology. Furthermore, any of the software-based embodiments may be uploaded, downloaded, or remotely accessed through a suitable communication mode. Such a suitable communication modes include, for example, the Internet, the World Wide Web, an intranet, software applications, cable (including fiber optic cable), magnetic communications, electromagnetic communications (including RF, microwave, and infrared communications), electronic communications, or other such communication modes.
[00206]Although the invention has been described with reference to specific exemplary embodiments, it is noted that various modifications and changes may be made to these embodiments without departing from the broad scope of the invention. For example, the various operations, blocks, etc., described herein may be enabled and operated using hardware circuitry (for example, Complementary Metal Oxide Semiconductor (CMOS) based logic circuitry), firmware, software, and/or any combination of hardware, firmware, and/or software (for example, embodied in a machine-readable medium). For example, the apparatuses and methods may be embodied using transistors, logic gates, and electrical circuits (for example, Application-Specific Integrated Circuit (ASIC) circuitry and/or in Digital Signal Processor (DSP) circuitry).
[00207]Particularly, the server system 200 and its various components may be enabled using software and/or using transistors, logic gates, and electrical circuits (for example, integrated circuit circuitry such as ASIC circuitry). Various embodiments of the invention may include one or more computer programs stored or otherwise embodied on a computer-readable medium, wherein the computer programs are configured to cause a processor or the computer to perform one or more operations. A computer-readable medium storing, embodying, or encoded with a computer program, or similar language, may be embodied as a tangible data storage device storing one or more software programs that are configured to cause a processor or computer to perform one or more operations. Such operations may be, for example, any of the steps or operations described herein. In some embodiments, the computer programs may be stored and provided to a computer using any type of non-transitory computer-readable media. Non-transitory computer-readable media includes any type of tangible storage media. Examples of non-transitory computer-readable media include magnetic storage media (such as floppy disks, magnetic tapes, hard disk drives, etc.), optical magnetic storage media (e.g. magneto-optical disks), Compact Disc Read-Only Memory (CD-ROM), Compact Disc Recordable CD-R, Compact Disc Rewritable CD-R/W), Digital Versatile Disc (DVD), and semiconductor memories (such as mask ROM, programmable ROM (PROM), Erasable PROM (EPROM), flash memory, Random Access Memory (RAM), etc.). Additionally, a tangible data storage device may be embodied as one or more volatile memory devices, one or more non-volatile memory devices, and/or a combination of one or more volatile memory devices and non-volatile memory devices. In some embodiments, the computer programs may be provided to a computer using any type of transitory computer-readable media. Examples of transitory computer-readable media include electric signals, optical signals, and electromagnetic waves. Transitory computer-readable media can provide the program to a computer via a wired communication line (e.g., electric wires, and optical fibers) or a wireless communication line.
[00208]Various embodiments of the invention, as discussed above, may be practiced with steps and/or operations in a different order, and/or with hardware elements in configurations, which are different from those which, are disclosed. Therefore, although the invention has been described based on these exemplary embodiments, it is noted that certain modifications, variations, and alternative constructions may be apparent and well within the scope of the invention.
[00209]Although various exemplary embodiments of the invention are described herein in a language specific to structural features and/or methodological acts, the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as exemplary forms of implementing the claims.
Furthermore, as an ordered combination, these elements amount to generic computer components receiving or transmitting data over a network, performing repetitive calculations, electronic record keeping, and storing and retrieving information in memory, which, as held by the courts, are well-understood, routine, and conventional. See MPEP 2106.05(d).
Moreover, the remaining elements of dependent claims do not transform the recited abstract idea into a patent eligible invention because these remaining elements merely recite further abstract limitations that provide nothing more than simply a narrowing of the abstract idea recited in the independent claims.
Looking at these limitations as an ordered combination adds nothing additional that is sufficient to amount to significantly more than the recited abstract idea because they simply provide instructions to use a generic arrangement of generic computer components to “apply” the recited abstract idea, perform insignificant extra-solution activity, and generally link the abstract idea to a technical environment. Thus, the elements of the claims, considered both individually and as an ordered combination, are not sufficient to ensure that the claim as a whole amounts to significantly more than the abstract idea itself. Since there are no limitations in these claims that transform the exception into a patent eligible application such that these claims amount to significantly more than the exception itself, claims 1-20 are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter.
Claim Rejections – 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
Determining the scope and contents of the prior art.
Ascertaining the differences between the prior art and the claims at issue.
Resolving the level of ordinary skill in the pertinent art.
Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-7, 9-17, 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable by US Patent Publication to US20190122139A1 to Perez, (hereinafter referred to as “Perez”) in view of US Patent Publication to US20230342348A1 to Wang, (hereinafter referred to as “Wang”)
As per Claim 1, Perez teaches: A computer-implemented method for generating a prediction for an event, comprising: ([0032]-[0048])
accessing, by a server system, a feature set corresponding to each data sample in an input dataset from a database associated with the server system; (in at least [0023][0024] Process 400 may begin with operation 402, where data is retrieved. The data retrieved may come in the form of a user data set, user transaction information, account information, buyer information, seller information, etc. In some instances, the data retrieved is historical data, while in other instances, the data may be current. Following retrieval of the data, the data retrieved may then be preprocessed at operation 404. Preprocessing of the data may include formatting or organizing the data in such a way that it may be input into a machine learning model for analysis and prediction.)
generating, by a set of decision trees associated with the server system, an intermediate prediction for the event based, at least in part, on applying the feature set of each data sample on the set of decision trees, each decision tree comprising a plurality of nodes; (in at least [0016] FIG. 2 illustrates a decision tree structure applicable to machine learning. In particular, FIG. 2 provides an exemplary decision tree 200 illustrating the mapping of relationships and observations to arrive at a conclusion(s). Generally, decision trees can be described in terms of three components, a root node, leaves, and branches. The root node can be the starting point of the tree. As illustrated in FIG. 2 this can include root 202, the top most node which indicates the start of the decision tree. In most instances, the root 202 may be used to indicate the main idea or main criteria that should be met to arrive at a decision. [0025] Once the data has been preprocessed, process 400 continues to train the machine learning model used for processing the data and making predictions. As previously indicated, large data is constantly collected by devices that oftentimes needs to be organized and analyzed. Such data can include the preprocessed data for training and using with the machine learning model. Oftentimes, the large data is organized using decision tree structures which can be used by the machine learning algorithms to extract the information of interest. In one embodiment, the decision tree structures may be used in training the machine learning model. As indicated, machine learning techniques using decision tree ensembles may be used. Decision tree ensembles such as Gradient Boosting Trees and Random Forest may be used for such regression and classification problems and predictions. As such, at operation 408, a decision is made as to whether the model has converged.)
identifying, by the server system, an … leaf node from each decision tree based, at least in part, on the intermediate prediction, the … leaf node indicating a leaf node of one or more leaf nodes of the corresponding decision tree contributing to the intermediate prediction; (in at least [0017] once the main criteria or root 202 has been established, branches extended from root node 202 can be used to indicate the various options available from the root 202. For example, the root may include a criteria which can have one of two outcomes which can be answered by a “yes” or “no.” FIG. 2, illustrates such example, where the root 202 provides two possible options indicated by the two branches 210 extending from the root 202. The branches can then attach to other (child) nodes 204 which are related to the root 202. Again branches 210 may then be extended from each node as more decisions (e.g. nodes 204) and possible outcomes (e.g., branches 210) exist. The decision tree 200 may continue to grow as more and more decisions are made until a final outcome is reached and represented by a leaf/label 206. As illustrated in FIG. 2, one or more leafs 206 are possible from a single node and in some instances, the leafs 206 may appear as early as the second node 204, while in other instances, the nodes 204 may extend several layers before arriving at a leaf 206. [0021] consider an instance where a prediction is needed regarding the risk associated with a seller having limited or no prior transaction history. To make a prediction, information known about the seller (e.g., buyers associated with the seller) may be used to help make a determination as to whether or not open an account, complete a transaction, provide credit, etc. to the seller. Thus, the first decision tree may be used to classify the buyer into a good or bad category. These categories and corresponding buyers can be analyzed and a determination can be made regarding the errors made and location where they occurred. Thus, a second decision tree may be trained to try and “fix” the mistakes. A third decision tree then continues by focusing on any errors on the second decision tree. The process continues until model converges. [0025] Once the data has been preprocessed, process 400 continues to train the machine learning model used for processing the data and making predictions. As previously indicated, large data is constantly collected by devices that oftentimes needs to be organized and analyzed. Such data can include the preprocessed data for training and using with the machine learning model. Oftentimes, the large data is organized using decision tree structures which can be used by the machine learning algorithms to extract the information of interest. In one embodiment, the decision tree structures may be used in training the machine learning model. As indicated, machine learning techniques using decision tree ensembles may be used. Decision tree ensembles such as Gradient Boosting Trees and Random Forest may be used for such regression and classification problems and predictions. As such, at operation 408, a decision is made as to whether the model has converged. [0026] If the model has not yet converged, then process 400 returns to operation 406 for further training (e.g., further processing of additional decision trees). Alternatively, if the model is trained, the model continues to operation 410 for evaluation. Training model evaluation may include a comparison of known labels or results with those output by the model. In other words, a data set may be input into the trained model and its result compared to a known label or result for accuracy.)
generating, by the server system, a tree-based encoding indicating an encoded representation of the set of decision trees for each data sample based, at least in part, on an encoding type and the … leaf node from each decision tree, wherein the tree-based encoding comprises a …, the tree-based encoding includes a position-type leaf-based encoding or a type-type leaf-based encoding, and the tree-based encoding comprises a machine-generated encoded representation stored in memory and generated using a processor to … the set of decision trees for the data sample; and (in at least [0041][0019] FIG. 3 for example, illustrates a gradient boosting tree boosting structure applicable to machine learning. Gradient boosting trees is a machine learning technique that uses regression for classification and prediction.[0020] A GBT works in a serial manner such that a first decision tree is trained and a target is obtained. When the decision tree is trained however, it is recognized that the prediction is weak and/or errors may exist. Therefore, the model can take the error encountered and used the data to train another decision tree based on the error. The process continues and in some instances, a chart may be generated which graphs the errors until the error graph flattens out or converges. The chart can include a bar chart, a line graph, a statistical plot (e.g., bell shaped curve) or the like, where the errors are mapped. [0027] Once the model is fully trained and working properly, the model is ready for placement into production. To add the model into production, in a first embodiment, an SQL query is introduced which to enables the use of the machine learning model by translating the model into SQL query so that the model may be used in production for calculating and making predictions. In one embodiment, a python library is implemented to convert the machine learning model to legible SQL. For example, in order to run arbitrary Python code on Teradata, a python library may be implemented to convert a GBT and/or Random Forest model to legible SQL code. Therefore, process 400 continues to operation 412 where the model is translated for use with SQL. Process 400 then concludes with the run of the now integrated ML model in production at operation 414.)
generating, by the server system, the prediction for the event, wherein the prediction is calibrated by a calibration model associated with the server system based, at least in part, on the tree-based encoding for each data sample. (in at least [0029] In order to ensure that process 400 was successfully implemented, a sample run production is presented in FIGS. 5A-5B. In particular, a validation run was performed, with FIG. 5A illustrating the accuracy of the model by presenting that the SQL model and the original python model achieved results that are closely correlated. As illustrated, the SQL model predictions are correlated to the python model prediction with correlation=99.99%. [0030] Therefore, process 400 may be successfully implemented to run various data collection analytics. As an example, process 400 can be implemented to provide an indication of the risk involved with a new individual selling on a site provided that only minimal information is known about the seller. The information may be limited if the seller is fairly new to the system and/or is a casual buyer where sales are not habitual. In this example, process 400 can be implemented to collect information about the buyer provide a prediction as to the risk associated with the seller. The seller can be provided a code indicating the risk. FIG. 5B illustrates such use of the model where a seller is assessed and scored. Note that other analysis and predictions are possible using process 400. As an example, payment flows used, bank account information, and seller information may be used.)
Although implied, Perez does not expressly disclose the following limitations, which however, are taught by Wang,
…activated leaf node… (in at least [0065] To compute the position of the next symbol, the decoder network uses a stack-based algorithm similar to that used during encoding. The algorithm works by maintaining a stack of partially constructed nodes in the output operator tree. Each node on the stack represents an operator that has been generated but whose arguments have not yet been generated. The top node on the stack is always the current node being constructed. [0108] Nodes 802 and edges 804 carry additional associations. Namely, every edge is associated with a numerical value. The edge numerical values, or even the edges 804 themselves, are often referred to as “weights” or “parameters”. While training a neural network, numerical values are assigned to each edge 804. Additionally, every node 802 is associated with a numerical variable and an activation function. Activation functions are not limited to any functional class, but traditionally follow the form [0109] When the neural network receives an input, the input is propagated through the network according to the activation functions and incoming node 802 values and edge 804 values to compute a value for each node 802.)
…wherein the tree-based encoding comprises a concatenated leaf-based encoding generated from activated leaf nodes of the set of decision trees, the tree-based encoding includes a position-type leaf-based encoding or a type-type leaf-based encoding, and the tree-based encoding comprises a machine-generated encoded representation stored in memory and generated using a processor to identify activated leaf nodes across the set of decision trees for the data sample…(in at least [0050] Turning to FIG. 2 , in one or more embodiments, the encoding process includes formula tree traversal and node embedding. In this process, each formula tree in depth first search (DFS) order is traversed to extract node symbols. This step returns a DFS-ordered list of nodes. [0051] Continuing with FIG. 2 , a two-step method is used to extract the structure of a formula tree. In the first step, a position of each node may be calculated based on the position of a parent node. In the second step, positions obtained from the previous step may be embedded into fixed dimensional tree positional embeddings. [0052] In addition, an embedding function may be used to transform the formula tree into its embedding. The function includes concatenating the node and tree positional embeddings such that the encoder is aware of both the nodes and their positions. [0053] Indeed, the encoding process is critical for machine learning tasks involving mathematical formulas, as it provides a structured representation of the formula that can capture its underlying semantics and relationships between symbols. The use of a tree-based representation allows the model to leverage the hierarchical structure of formulas, which can help improve its accuracy and ability to generalize to new formulas. [0065] To compute the position of the next symbol, the decoder network uses a stack-based algorithm similar to that used during encoding. The algorithm works by maintaining a stack of partially constructed nodes in the output operator tree. Each node on the stack represents an operator that has been generated but whose arguments have not yet been generated. The top node on the stack is always the current node being constructed. [0081] Step 103 includes converting, with the processor, the tree format of the first formula into a plurality of lists. The tree format of the first formula is then converted into multiple lists using the same processor. Each list corresponds to one level of the tree and contains information about the nodes at that level (e.g., their type, position in the tree). The plurality of lists includes a list of nodes and a list of positions.)
At the time the invention was filed, it would have been obvious for one of ordinary skill in the art to have modified the teachings of Perez, as taught by Wang above, with a reasonable expectation of success if arriving at the claimed invention. One of ordinary skill in the art would have been motivated to make this modification to the teachings of Perez with the motivation of, …Such a model can significantly improve the performance of mathematical language processing applications, including but not limited to, automated theorem proving, mathematical formula recognition and retrieval, and natural language generation from mathematical expressions…improve the quality of formula tree generation, a novel tree beam search algorithm that is of independent scientific interest has been developed…The use of a tree-based representation allows the model to leverage the hierarchical structure of formulas, which can help improve its accuracy and ability to generalize to new formulas.…the accuracy level of the present application is improved...the framework has potential to serve as a drop-in replacement in some existing math retrieval systems to significantly improve their performance…automatically generate formula can serve as an intelligent assistant for authoring mathematical content (e.g., providing suggestions on what formula to write next, potentially improving the efficiency in the generation of such content)…the discussed framework can also be used for (1) automatic grading students’ mathematical responses to STEM questions involving formulae; (2) automated derivation and verification of simple mathematical steps or proofs; (3) automatic detection of cheating on answers to math questions involving formulae; (4) automatic generation of formulae from text and vice versa; (5) automatic generation of math practice problems in STEM disciplines with different contexts and numeric values to adapt to interest of different students and teachers; (6) improving math knowledge tracing with models such as item response theory (IRT) by explicitly taking into account of the mathematical content in the question and students’ answers…, as recited in Wang.
As per Claim 2, Perez teaches: The computer-implemented method as claimed in claim 1, wherein generating the tree-based encoding for each data sample comprises:
generating, by the server system, a set of leaf-based encodings for the set of decision trees based, at least in part, on the encoding type and the one or more leaf nodes in each decision tree, each leaf-based encoding indicating a leaf-level encoded representation of each decision tree for each data sample; and (in at least [0017] once the main criteria or root 202 has been established, branches extended from root node 202 can be used to indicate the various options available from the root 202. For example, the root may include a criteria which can have one of two outcomes which can be answered by a “yes” or “no.” FIG. 2, illustrates such example, where the root 202 provides two possible options indicated by the two branches 210 extending from the root 202. The branches can then attach to other (child) nodes 204 which are related to the root 202. Again branches 210 may then be extended from each node as more decisions (e.g. nodes 204) and possible outcomes (e.g., branches 210) exist. The decision tree 200 may continue to grow as more and more decisions are made until a final outcome is reached and represented by a leaf/label 206. As illustrated in FIG. 2, one or more leafs 206 are possible from a single node and in some instances, the leafs 206 may appear as early as the second node 204, while in other instances, the nodes 204 may extend several layers before arriving at a leaf 206. [0018] the decision tree 200 may grow and can extend significantly as more data is received and decisions are to be made. FIG. 2 is for illustrative purposes only and more or less nodes, branches, and leaves may be added. In addition, FIG. 2 is designed to illustrate the intricacies of a decision tree and how quickly, the tree may become large. Further, FIG. 2 illustrates how the decision tree may arrive at a leaf/label 206 right away of after several iterations. In some instances, the decision reached may be incorrect and the model may need to be trained further to achieve competent results.)
concatenating, the server system, the set of leaf-based encodings to obtain the tree-based encoding for each data sample. (in at least [0018] the decision tree 200 may grow and can extend significantly as more data is received and decisions are to be made. FIG. 2 is for illustrative purposes only and more or less nodes, branches, and leaves may be added. In addition, FIG. 2 is designed to illustrate the intricacies of a decision tree and how quickly, the tree may become large. Further, FIG. 2 illustrates how the decision tree may arrive at a leaf/label 206 right away of after several iterations. In some instances, the decision reached may be incorrect and the model may need to be trained further to achieve competent results.)
As per Claim 3, Perez teaches: The computer-implemented method as claimed in claim 2, wherein generating a leaf-based encoding for a decision tree comprises:
in response to determining that the encoding type is a position-based encoding, generating, by the server system, a position-type leaf-based encoding for the decision tree, wherein the position-type leaf-based encoding is the leaf-based encoding. (in at least [0016] FIG. 2 illustrates a decision tree structure applicable to machine learning. In particular, FIG. 2 provides an exemplary decision tree 200 illustrating the mapping of relationships and observations to arrive at a conclusion(s). Generally, decision trees can be described in terms of three components, a root node, leaves, and branches. The root node can be the starting point of the tree. As illustrated in FIG. 2 this can include root 202, the top most node which indicates the start of the decision tree. In most instances, the root 202 may be used to indicate the main idea or main criteria that should be met to arrive at a decision. [0017] once the main criteria or root 202 has been established, branches extended from root node 202 can be used to indicate the various options available from the root 202. For example, the root may include a criteria which can have one of two outcomes which can be answered by a “yes” or “no.” FIG. 2, illustrates such example, where the root 202 provides two possible options indicated by the two branches 210 extending from the root 202. The branches can then attach to other (child) nodes 204 which are related to the root 202. Again branches 210 may then be extended from each node as more decisions (e.g. nodes 204) and possible outcomes (e.g., branches 210) exist. The decision tree 200 may continue to grow as more and more decisions are made until a final outcome is reached and represented by a leaf/label 206. As illustrated in FIG. 2, one or more leafs 206 are possible from a single node and in some instances, the leafs 206 may appear as early as the second node 204, while in other instances, the nodes 204 may extend several layers before arriving at a leaf 206.)
As per Claim 4, Perez teaches: The computer-implemented method as claimed in claim 3, wherein generating the position-type leaf-based encoding comprises:
assigning, by the server system, a first label to the … leaf node and a second label to each of remaining leaf nodes of the one or more leaf nodes of the decision tree; and (in at least [0017] The decision tree 200 may continue to grow as more and more decisions are made until a final outcome is reached and represented by a leaf/label 206. As illustrated in FIG. 2, one or more leafs 206 are possible from a single node and in some instances, the leafs 206 may appear as early as the second node 204, while in other instances, the nodes 204 may extend several layers before arriving at a leaf 206. [0018] the decision tree 200 may grow and can extend significantly as more data is received and decisions are to be made. FIG. 2 is for illustrative purposes only and more or less nodes, branches, and leaves may be added. In addition, FIG. 2 is designed to illustrate the intricacies of a decision tree and how quickly, the tree may become large. Further, FIG. 2 illustrates how the decision tree may arrive at a leaf/label 206 right away of after several iterations. In some instances, the decision reached may be incorrect and the model may need to be trained further to achieve competent results.)
concatenating, by the server system, the first label and the second label of each of the remaining leaf nodes based, at least on a position of each of the one or more leaf nodes in the decision tree to obtain the position-type leaf-based encoding of the decision tree. (in at least [0016] FIG. 2 illustrates a decision tree structure applicable to machine learning. In particular, FIG. 2 provides an exemplary decision tree 200 illustrating the mapping of relationships and observations to arrive at a conclusion(s). Generally, decision trees can be described in terms of three components, a root node, leaves, and branches. The root node can be the starting point of the tree. As illustrated in FIG. 2 this can include root 202, the top most node which indicates the start of the decision tree. In most instances, the root 202 may be used to indicate the main idea or main criteria that should be met to arrive at a decision. [0017] The decision tree 200 may continue to grow as more and more decisions are made until a final outcome is reached and represented by a leaf/label 206. As illustrated in FIG. 2, one or more leafs 206 are possible from a single node and in some instances, the leafs 206 may appear as early as the second node 204, while in other instances, the nodes 204 may extend several layers before arriving at a leaf 206. [0018] the decision tree 200 may grow and can extend significantly as more data is received and decisions are to be made. FIG. 2 is for illustrative purposes only and more or less nodes, branches, and leaves may be added. In addition, FIG. 2 is designed to illustrate the intricacies of a decision tree and how quickly, the tree may become large. Further, FIG. 2 illustrates how the decision tree may arrive at a leaf/label 206 right away of after several iterations. In some instances, the decision reached may be incorrect and the model may need to be trained further to achieve competent results. [0026] If the model has not yet converged, then process 400 returns to operation 406 for further training (e.g., further processing of additional decision trees). Alternatively, if the model is trained, the model continues to operation 410 for evaluation. Training model evaluation may include a comparison of known labels or results with those output by the model. In other words, a data set may be input into the trained model and its result compared to a known label or result for accuracy.)
Although implied, Perez does not expressly disclose the following limitations, which however, are taught by Wang,
…activated leaf node… (in at least [0065] To compute the position of the next symbol, the decoder network uses a stack-based algorithm similar to that used during encoding. The algorithm works by maintaining a stack of partially constructed nodes in the output operator tree. Each node on the stack represents an operator that has been generated but whose arguments have not yet been generated. The top node on the stack is always the current node being constructed. [0108] Nodes 802 and edges 804 carry additional associations. Namely, every edge is associated with a numerical value. The edge numerical values, or even the edges 804 themselves, are often referred to as “weights” or “parameters”. While training a neural network, numerical values are assigned to each edge 804. Additionally, every node 802 is associated with a numerical variable and an activation function. Activation functions are not limited to any functional class, but traditionally follow the form [0109] When the neural network receives an input, the input is propagated through the network according to the activation functions and incoming node 802 values and edge 804 values to compute a value for each node 802.)
The reason and rationale to combine Perez and Wang is the same recited above.
As per Claim 5, Perez teaches: The computer-implemented method as claimed in claim 2, wherein generating a leaf-based encoding for a decision tree comprises:
in response to determining that the encoding type is a …-based encoding, generating, by the server system, a …-type leaf-based encoding for the decision tree, wherein the …-type leaf-based encoding is the leaf-based encoding. (in at least [0016] FIG. 2 illustrates a decision tree structure applicable to machine learning. In particular, FIG. 2 provides an exemplary decision tree 200 illustrating the mapping of relationships and observations to arrive at a conclusion(s). Generally, decision trees can be described in terms of three components, a root node, leaves, and branches. The root node can be the starting point of the tree. As illustrated in FIG. 2 this can include root 202, the top most node which indicates the start of the decision tree. In most instances, the root 202 may be used to indicate the main idea or main criteria that should be met to arrive at a decision. [0026] If the model has not yet converged, then process 400 returns to operation 406 for further training (e.g., further processing of additional decision trees). Alternatively, if the model is trained, the model continues to operation 410 for evaluation. Training model evaluation may include a comparison of known labels or results with those output by the model. In other words, a data set may be input into the trained model and its result compared to a known label or result for accuracy.)
Although implied, Perez does not expressly disclose the following limitations, which however, are taught by Wang,
…weight… (in at least [0108] Nodes 802 and edges 804 carry additional associations. Namely, every edge is associated with a numerical value. The edge numerical values, or even the edges 804 themselves, are often referred to as “weights” or “parameters”. While training a neural network, numerical values are assigned to each edge 804. Additionally, every node 802 is associated with a numerical variable and an activation function. Activation functions are not limited to any functional class, but traditionally follow the form)
The reason and rationale to combine Perez and Wang is the same recited above.
As per Claim 6, Perez teaches: The computer-implemented method as claimed in claim 5, wherein generating the weight-type leaf-based encoding comprises:
extracting, by the server system, a … parameter associated with the … leaf node from the decision tree; and (in at least [0017] once the main criteria or root 202 has been established, branches extended from root node 202 can be used to indicate the various options available from the root 202. For example, the root may include a criteria which can have one of two outcomes which can be answered by a “yes” or “no.” FIG. 2, illustrates such example, where the root 202 provides two possible options indicated by the two branches 210 extending from the root 202. The branches can then attach to other (child) nodes 204 which are related to the root 202. Again branches 210 may then be extended from each node as more decisions (e.g. nodes 204) and possible outcomes (e.g., branches 210) exist. The decision tree 200 may continue to grow as more and more decisions are made until a final outcome is reached and represented by a leaf/label 206. As illustrated in FIG. 2, one or more leafs 206 are possible from a single node and in some instances, the leafs 206 may appear as early as the second node 204, while in other instances, the nodes 204 may extend several layers before arriving at a leaf 206. [0018] the decision tree 200 may grow and can extend significantly as more data is received and decisions are to be made. FIG. 2 is for illustrative purposes only and more or less nodes, branches, and leaves may be added. In addition, FIG. 2 is designed to illustrate the intricacies of a decision tree and how quickly, the tree may become large. Further, FIG. 2 illustrates how the decision tree may arrive at a leaf/label 206 right away of after several iterations. In some instances, the decision reached may be incorrect and the model may need to be trained further to achieve competent results. [0025] the large data is organized using decision tree structures which can be used by the machine learning algorithms to extract the information of interest. In one embodiment, the decision tree structures may be used in training the machine learning model.)
assigning, by the server system, the extracted … parameter to the leaf-based encoding of the decision tree to obtain the … type leaf-based encoding. (in at least [0017] once the main criteria or root 202 has been established, branches extended from root node 202 can be used to indicate the various options available from the root 202. For example, the root may include a criteria which can have one of two outcomes which can be answered by a “yes” or “no.” FIG. 2, illustrates such example, where the root 202 provides two possible options indicated by the two branches 210 extending from the root 202. The branches can then attach to other (child) nodes 204 which are related to the root 202. Again branches 210 may then be extended from each node as more decisions (e.g. nodes 204) and possible outcomes (e.g., branches 210) exist. The decision tree 200 may continue to grow as more and more decisions are made until a final outcome is reached and represented by a leaf/label 206. As illustrated in FIG. 2, one or more leafs 206 are possible from a single node and in some instances, the leafs 206 may appear as early as the second node 204, while in other instances, the nodes 204 may extend several layers before arriving at a leaf 206. [0018] the decision tree 200 may grow and can extend significantly as more data is received and decisions are to be made. FIG. 2 is for illustrative purposes only and more or less nodes, branches, and leaves may be added. In addition, FIG. 2 is designed to illustrate the intricacies of a decision tree and how quickly, the tree may become large. Further, FIG. 2 illustrates how the decision tree may arrive at a leaf/label 206 right away of after several iterations. In some instances, the decision reached may be incorrect and the model may need to be trained further to achieve competent results. [0025] the large data is organized using decision tree structures which can be used by the machine learning algorithms to extract the information of interest. In one embodiment, the decision tree structures may be used in training the machine learning model.)
Although implied, Perez does not expressly disclose the following limitations, which however, are taught by Wang,
…activated leaf node… (in at least [0065] To compute the position of the next symbol, the decoder network uses a stack-based algorithm similar to that used during encoding. The algorithm works by maintaining a stack of partially constructed nodes in the output operator tree. Each node on the stack represents an operator that has been generated but whose arguments have not yet been generated. The top node on the stack is always the current node being constructed. [0108] Nodes 802 and edges 804 carry additional associations. Namely, every edge is associated with a numerical value. The edge numerical values, or even the edges 804 themselves, are often referred to as “weights” or “parameters”. While training a neural network, numerical values are assigned to each edge 804. Additionally, every node 802 is associated with a numerical variable and an activation function. Activation functions are not limited to any functional class, but traditionally follow the form [0109] When the neural network receives an input, the input is propagated through the network according to the activation functions and incoming node 802 values and edge 804 values to compute a value for each node 802.)
…weight… (in at least [0108] Nodes 802 and edges 804 carry additional associations. Namely, every edge is associated with a numerical value. The edge numerical values, or even the edges 804 themselves, are often referred to as “weights” or “parameters”. While training a neural network, numerical values are assigned to each edge 804. Additionally, every node 802 is associated with a numerical variable and an activation function. Activation functions are not limited to any functional class, but traditionally follow the form)
The reason and rationale to combine Perez and Wang is the same recited above.
As per Claim 7, Perez teaches:The computer-implemented method as claimed in claim 1, further comprising:
extracting, by the server system, an intermediate predicted probability … associated with the intermediate prediction for each data sample from the set of decision trees, the intermediate predicted probability … indicating a likelihood of the event to take place; (in at least [0021] consider an instance where a prediction is needed regarding the risk associated with a seller having limited or no prior transaction history. To make a prediction, information known about the seller (e.g., buyers associated with the seller) may be used to help make a determination as to whether or not open an account, complete a transaction, provide credit, etc. to the seller. Thus, the first decision tree may be used to classify the buyer into a good or bad category. These categories and corresponding buyers can be analyzed and a determination can be made regarding the errors made and location where they occurred. Thus, a second decision tree may be trained to try and “fix” the mistakes. A third decision tree then continues by focusing on any errors on the second decision tree. The process continues until model converges. [0029] In order to ensure that process 400 was successfully implemented, a sample run production is presented in FIGS. 5A-5B. In particular, a validation run was performed, with FIG. 5A illustrating the accuracy of the model by presenting that the SQL model and the original python model achieved results that are closely correlated. As illustrated, the SQL model predictions are correlated to the python model prediction with correlation=99.99%. [0030] process 400 may be successfully implemented to run various data collection analytics. As an example, process 400 can be implemented to provide an indication of the risk involved with a new individual selling on a site provided that only minimal information is known about the seller. The information may be limited if the seller is fairly new to the system and/or is a casual buyer where sales are not habitual. In this example, process 400 can be implemented to collect information about the buyer provide a prediction as to the risk associated with the seller. The seller can be provided a code indicating the risk. FIG. 5B illustrates such use of the model where a seller is assessed and scored. Note that other analysis and predictions are possible using process 400. As an example, payment flows used, bank account information, and seller information may be used.)
extracting, by the server system, a predicted probability … associated with the prediction for each data sample from the calibration model, the predicted probability … indicating a calibrated likelihood of the event to take place; (in at least [0021] consider an instance where a prediction is needed regarding the risk associated with a seller having limited or no prior transaction history. To make a prediction, information known about the seller (e.g., buyers associated with the seller) may be used to help make a determination as to whether or not open an account, complete a transaction, provide credit, etc. to the seller. Thus, the first decision tree may be used to classify the buyer into a good or bad category. These categories and corresponding buyers can be analyzed and a determination can be made regarding the errors made and location where they occurred. Thus, a second decision tree may be trained to try and “fix” the mistakes. A third decision tree then continues by focusing on any errors on the second decision tree. The process continues until model converges. [0029] a sample run production is presented in FIGS. 5A-5B. In particular, a validation run was performed, with FIG. 5A illustrating the accuracy of the model by presenting that the SQL model and the original python model achieved results that are closely correlated. As illustrated, the SQL model predictions are correlated to the python model prediction with correlation=99.99%. [0030] process 400 may be successfully implemented to run various data collection analytics. As an example, process 400 can be implemented to provide an indication of the risk involved with a new individual selling on a site provided that only minimal information is known about the seller. The information may be limited if the seller is fairly new to the system and/or is a casual buyer where sales are not habitual. In this example, process 400 can be implemented to collect information about the buyer provide a prediction as to the risk associated with the seller. The seller can be provided a code indicating the risk. FIG. 5B illustrates such use of the model where a seller is assessed and scored. Note that other analysis and predictions are possible using process 400. As an example, payment flows used, bank account information, and seller information may be used.)
accessing, by the server system, one or more actual behavior parameters related to the event from the database; (in at least [0026] if the model is trained, the model continues to operation 410 for evaluation. Training model evaluation may include a comparison of known labels or results with those output by the model. In other words, a data set may be input into the trained model and its result compared to a known label or result for accuracy.)
computing, by the server system, a first calibration error for each data sample based, at least in part, on the intermediate predicted probability … for each data sample and the one or more actual behavior parameters; and (in at least [0020] When the decision tree is trained however, it is recognized that the prediction is weak and/or errors may exist. Therefore, the model can take the error encountered and used the data to train another decision tree based on the error. The process continues and in some instances, a chart may be generated which graphs the errors until the error graph flattens out or converges. The chart can include a bar chart, a line graph, a statistical plot (e.g., bell shaped curve) or the like, where the errors are mapped. [0021] consider an instance where a prediction is needed regarding the risk associated with a seller having limited or no prior transaction history. To make a prediction, information known about the seller (e.g., buyers associated with the seller) may be used to help make a determination as to whether or not open an account, complete a transaction, provide credit, etc. to the seller. Thus, the first decision tree may be used to classify the buyer into a good or bad category. These categories and corresponding buyers can be analyzed and a determination can be made regarding the errors made and location where they occurred. Thus, a second decision tree may be trained to try and “fix” the mistakes. A third decision tree then continues by focusing on any errors on the second decision tree. The process continues until model converges. [0026] if the model is trained, the model continues to operation 410 for evaluation. Training model evaluation may include a comparison of known labels or results with those output by the model. In other words, a data set may be input into the trained model and its result compared to a known label or result for accuracy.)
computing, by the server system, a second calibration error for each data sample based, at least in part, on the predicted probability … for each data sample and the one or more actual behavior parameters. (in at least [0020] When the decision tree is trained however, it is recognized that the prediction is weak and/or errors may exist. Therefore, the model can take the error encountered and used the data to train another decision tree based on the error. The process continues and in some instances, a chart may be generated which graphs the errors until the error graph flattens out or converges. The chart can include a bar chart, a line graph, a statistical plot (e.g., bell shaped curve) or the like, where the errors are mapped. [0021] consider an instance where a prediction is needed regarding the risk associated with a seller having limited or no prior transaction history. To make a prediction, information known about the seller (e.g., buyers associated with the seller) may be used to help make a determination as to whether or not open an account, complete a transaction, provide credit, etc. to the seller. Thus, the first decision tree may be used to classify the buyer into a good or bad category. These categories and corresponding buyers can be analyzed and a determination can be made regarding the errors made and location where they occurred. Thus, a second decision tree may be trained to try and “fix” the mistakes. A third decision tree then continues by focusing on any errors on the second decision tree. The process continues until model converges. [0026] if the model is trained, the model continues to operation 410 for evaluation. Training model evaluation may include a comparison of known labels or results with those output by the model. In other words, a data set may be input into the trained model and its result compared to a known label or result for accuracy.)
Although implied, Perez does not expressly disclose the following limitations, which however, are taught by Wang,
… intermediate predicted probability score associated with the intermediate prediction…intermediate predicted probability score… (in at least [0071] The TBS algorithm works by maintaining a priority queue of candidate trees sorted by their scores (i.e., how well they match the desired output). At each time step, TBS generates new candidate trees by expanding each tree in the current beam with all possible next symbols and adding them to the priority queue. The beam is then updated by selecting the top-k trees from the priority queue based on their scores, where k is equal to the beam size. [0089] Step 205 includes scoring each candidate child node based on its likelihood of generating a valid formula and similarity to the input information using a scoring function, which corresponds to how the tree beam search algorithm scores each candidate node based on its probability of generating valid formulas and similarity to input formulas. In other words, the scoring function is a function that evaluates a likelihood of generating a valid formula and similarly to the input information. [0090] Step 206 includes selecting a plurality of most probable partial trees based on their scores, which corresponds to how the tree beam search algorithm selects multiple partial trees based on their scores and keeps them in the beam.)
…predicted probability score… (in at least [0071] The TBS algorithm works by maintaining a priority queue of candidate trees sorted by their scores (i.e., how well they match the desired output). At each time step, TBS generates new candidate trees by expanding each tree in the current beam with all possible next symbols and adding them to the priority queue. The beam is then updated by selecting the top-k trees from the priority queue based on their scores, where k is equal to the beam size. [0089] Step 205 includes scoring each candidate child node based on its likelihood of generating a valid formula and similarity to the input information using a scoring function, which corresponds to how the tree beam search algorithm scores each candidate node based on its probability of generating valid formulas and similarity to input formulas. In other words, the scoring function is a function that evaluates a likelihood of generating a valid formula and similarly to the input information. [0090] Step 206 includes selecting a plurality of most probable partial trees based on their scores, which corresponds to how the tree beam search algorithm selects multiple partial trees based on their scores and keeps them in the beam.)
The reason and rationale to combine Perez and Wang is the same recited above.
As per Claim 9, Perez teaches: The computer-implemented method as claimed in claim 1, further comprising:
accessing, by the server system, a training feature set for each training data sample of a training dataset from the database, the training feature set comprising … labels; (in at least [0025] Once the data has been preprocessed, process 400 continues to train the machine learning model used for processing the data and making predictions. As previously indicated, large data is constantly collected by devices that oftentimes needs to be organized and analyzed. Such data can include the preprocessed data for training and using with the machine learning model. Oftentimes, the large data is organized using decision tree structures which can be used by the machine learning algorithms to extract the information of interest. In one embodiment, the decision tree structures may be used in training the machine learning model. [0026] If the model has not yet converged, then process 400 returns to operation 406 for further training (e.g., further processing of additional decision trees). Alternatively, if the model is trained, the model continues to operation 410 for evaluation. Training model evaluation may include a comparison of known labels or results with those output by the model. In other words, a data set may be input into the trained model and its result compared to a known label or result for accuracy.)
generating, by the server system, a plurality of training tree-based encodings for a plurality of training data samples, respectively of the training dataset; and (in at least [0025] Once the data has been preprocessed, process 400 continues to train the machine learning model used for processing the data and making predictions. As previously indicated, large data is constantly collected by devices that oftentimes needs to be organized and analyzed. Such data can include the preprocessed data for training and using with the machine learning model. Oftentimes, the large data is organized using decision tree structures which can be used by the machine learning algorithms to extract the information of interest. In one embodiment, the decision tree structures may be used in training the machine learning model. [0026] If the model has not yet converged, then process 400 returns to operation 406 for further training (e.g., further processing of additional decision trees). Alternatively, if the model is trained, the model continues to operation 410 for evaluation. Training model evaluation may include a comparison of known labels or results with those output by the model. In other words, a data set may be input into the trained model and its result compared to a known label or result for accuracy.)
training, by the server system, the calibration model, …, to generate the prediction for the event that is calibrated, wherein training the calibration model comprises performing iteratively until convergence criteria are met, a set of operations comprising: (in at least [0025] Once the data has been preprocessed, process 400 continues to train the machine learning model used for processing the data and making predictions. As previously indicated, large data is constantly collected by devices that oftentimes needs to be organized and analyzed. Such data can include the preprocessed data for training and using with the machine learning model. Oftentimes, the large data is organized using decision tree structures which can be used by the machine learning algorithms to extract the information of interest. In one embodiment, the decision tree structures may be used in training the machine learning model. [0026] If the model has not yet converged, then process 400 returns to operation 406 for further training (e.g., further processing of additional decision trees). Alternatively, if the model is trained, the model continues to operation 410 for evaluation. Training model evaluation may include a comparison of known labels or results with those output by the model. In other words, a data set may be input into the trained model and its result compared to a known label or result for accuracy.)
initializing the calibration model based, at least in part, on one or more calibration model parameters; (in at least [0029] In order to ensure that process 400 was successfully implemented, a sample run production is presented in FIGS. 5A-5B. In particular, a validation run was performed, with FIG. 5A illustrating the accuracy of the model by presenting that the SQL model and the original python model achieved results that are closely correlated. As illustrated, the SQL model predictions are correlated to the python model prediction with correlation=99.99%.)
generating, by the calibration model, a calibrated predicted probability score for each training data sample based, at least in part, on the plurality of training tree-based encodings and the one or more calibration model parameters, the calibrated predicted probability score indicating a likelihood of an occurrence of the event; (in at least [0029] In order to ensure that process 400 was successfully implemented, a sample run production is presented in FIGS. 5A-5B. In particular, a validation run was performed, with FIG. 5A illustrating the accuracy of the model by presenting that the SQL model and the original python model achieved results that are closely correlated. As illustrated, the SQL model predictions are correlated to the python model prediction with correlation=99.99%. [0030] provide an indication of the risk involved with a new individual selling on a site provided that only minimal information is known about the seller. The information may be limited if the seller is fairly new to the system and/or is a casual buyer where sales are not habitual. In this example, process 400 can be implemented to collect information about the buyer provide a prediction as to the risk associated with the seller.)
generating, by the calibration model, the prediction for the event based, at least in part, on the calibrated predicted probability score and an event threshold, the prediction comprising a label associated with the event; (in at least [0029] In order to ensure that process 400 was successfully implemented, a sample run production is presented in FIGS. 5A-5B. In particular, a validation run was performed, with FIG. 5A illustrating the accuracy of the model by presenting that the SQL model and the original python model achieved results that are closely correlated. As illustrated, the SQL model predictions are correlated to the python model prediction with correlation=99.99%. [0030] FIG. 5B illustrates such use of the model where a seller is assessed and scored. Note that other analysis and predictions are possible using process 400. As an example, payment flows used, bank account information, and seller information may be used.)
computing, by the calibration model, a regularized … for each training data sample based, at least in part, on the calibrated predicted probability score, the … labels, and a regularized … function; and (in at least [0029] In order to ensure that process 400 was successfully implemented, a sample run production is presented in FIGS. 5A-5B. In particular, a validation run was performed, with FIG. 5A illustrating the accuracy of the model by presenting that the SQL model and the original python model achieved results that are closely correlated. As illustrated, the SQL model predictions are correlated to the python model prediction with correlation=99.99%.)
optimizing the one or more calibration model parameters based, at least in part, on the regularized …. (in at least [0026] If the model has not yet converged, then process 400 returns to operation 406 for further training (e.g., further processing of additional decision trees). Alternatively, if the model is trained, the model continues to operation 410 for evaluation. Training model evaluation may include a comparison of known labels or results with those output by the model. In other words, a data set may be input into the trained model and its result compared to a known label or result for accuracy. [0029] In order to ensure that process 400 was successfully implemented, a sample run production is presented in FIGS. 5A-5B. In particular, a validation run was performed, with FIG. 5A illustrating the accuracy of the model by presenting that the SQL model and the original python model achieved results that are closely correlated. As illustrated, the SQL model predictions are correlated to the python model prediction with correlation=99.99%.)
Although implied, Perez does not expressly disclose the following limitations, which however, are taught by Wang,
…ground truth labels…(in at least [0078] The encoder is trained using a loss function that measures the difference between the predicted embeddings and the ground truth embeddings for each node in the formula tree. The loss function encourages the encoder to learn representations that capture important structural features of formulas and are useful for downstream tasks. [0111] Training data is provided to the neural network. Generally, training data consists of pairs of inputs and associated targets. The targets represent the “ground truth”, or the otherwise desired output, upon processing the inputs )
…regularized loss… regularized loss function…(in at least [0078] The encoder is trained using a loss function that measures the difference between the predicted embeddings and the ground truth embeddings for each node in the formula tree. The loss function encourages the encoder to learn representations that capture important structural features of formulas and are useful for downstream tasks. [0111] The comparison of the neural network output to the target is typically performed by a so-called “loss function”; although other names for this comparison function such as “error function”, “misfit function”, and “cost function” are commonly employed. Many types of loss functions are available, such as the mean-squared-error function, however, the general characteristic of a loss function is that the loss function provides a numerical evaluation of the similarity between the neural network output and the associated target. The loss function may also be constructed to impose additional constraints on the values assumed by the edges 804, for example, by adding a penalty term, which may be physics-based, or a regularization term. Generally, the goal of a training procedure is to alter the edge 804 values to promote similarity between the neural network output and associated target over the training data. Thus, the loss function is used to guide changes made to the edge 804 values, typically through a process called “backpropagation”. [0113] Once the edge 804 values have been updated, or altered from their initial values, through a backpropagation step, the neural network will likely produce different outputs. Thus, the procedure of propagating at least one input through the neural network, comparing the neural network output with the associated target with a loss function, computing the gradient of the loss function with respect to the edge 804 values, and updating the edge 804 values with a step guided by the gradient, is repeated until a termination criterion is reached. Common termination criteria are: reaching a fixed number of edge 804 updates, otherwise known as an iteration counter; a diminishing learning rate; noting no appreciable change in the loss function between iterations; reaching a specified performance metric as evaluated on the data or a separate hold-out data set. Once the termination criterion is satisfied, and the edge 804 values are no longer intended to be altered, the neural network 801 is said to be “trained.”)
…using the plurality of training tree-based encodings generated from the plurality of activated leaf nodes…(in at least [0050] Turning to FIG. 2 , in one or more embodiments, the encoding process includes formula tree traversal and node embedding. In this process, each formula tree in depth first search (DFS) order is traversed to extract node symbols. This step returns a DFS-ordered list of nodes. [0051] Continuing with FIG. 2 , a two-step method is used to extract the structure of a formula tree. In the first step, a position of each node may be calculated based on the position of a parent node. In the second step, positions obtained from the previous step may be embedded into fixed dimensional tree positional embeddings. [0052] In addition, an embedding function may be used to transform the formula tree into its embedding. The function includes concatenating the node and tree positional embeddings such that the encoder is aware of both the nodes and their positions. [0065] To compute the position of the next symbol, the decoder network uses a stack-based algorithm similar to that used during encoding. The algorithm works by maintaining a stack of partially constructed nodes in the output operator tree. Each node on the stack represents an operator that has been generated but whose arguments have not yet been generated. The top node on the stack is always the current node being constructed. [0075] Step 101 includes training, with a server, a model by using a machine learning algorithm with a data set that includes a plurality of formulae. A machine learning algorithm is used to train a model on a dataset that includes multiple formulae. The goal of the training process is to learn how to encode formulae into vector representations that capture their semantic meaning. the trained model captures both local and global dependencies between nodes in the tree format for accurate encoding of formulas. [0076] The model may be trained using a supervised learning approach. Specifically, the model may be trained to predict an embedding vector for each node in a given formula tree. [0077] During training, the encoder takes as input a formula tree and generates an embedding vector for each node in the tree. The embedding vectors capture both semantic and syntactic information about each node and its relationships with other nodes in the tree. These embedding vectors are then used as input to downstream tasks such as similar formula retrieval. [0108] Nodes 802 and edges 804 carry additional associations. Namely, every edge is associated with a numerical value. The edge numerical values, or even the edges 804 themselves, are often referred to as “weights” or “parameters”. While training a neural network, numerical values are assigned to each edge 804. Additionally, every node 802 is associated with a numerical variable and an activation function. Activation functions are not limited to any functional class, but traditionally follow the form [0109] When the neural network receives an input, the input is propagated through the network according to the activation functions and incoming node 802 values and edge 804 values to compute a value for each node 802.)
The reason and rationale to combine Perez and Wang is the same recited above.
As per Claim 10, Perez teaches: The computer-implemented method as claimed in claim 1, further comprising:
accessing, by the server system, the input dataset from the database, the input dataset comprising the plurality of data samples associated with a plurality of users; (in at least [0021] consider an instance where a prediction is needed regarding the risk associated with a seller having limited or no prior transaction history. To make a prediction, information known about the seller (e.g., buyers associated with the seller) may be used to help make a determination as to whether or not open an account, complete a transaction, provide credit, etc. to the seller. Thus, the first decision tree may be used to classify the buyer into a good or bad category. These categories and corresponding buyers can be analyzed and a determination can be made regarding the errors made and location where they occurred. Thus, a second decision tree may be trained to try and “fix” the mistakes. A third decision tree then continues by focusing on any errors on the second decision tree. [0024] operation 402, where data is retrieved. The data retrieved may come in the form of a user data set, user transaction information, account information, buyer information, seller information, etc. In some instances, the data retrieved is historical data, while in other instances, the data may be current. Following retrieval of the data, the data retrieved may then be preprocessed at operation 404. Preprocessing of the data may include formatting or organizing the data in such a way that it may be input into a machine learning model for analysis and prediction.)
generating, by the server system, the feature set for each data sample of the plurality of data samples based, at least in part, on the input dataset; and (in at least [0021] consider an instance where a prediction is needed regarding the risk associated with a seller having limited or no prior transaction history. To make a prediction, information known about the seller (e.g., buyers associated with the seller) may be used to help make a determination as to whether or not open an account, complete a transaction, provide credit, etc. to the seller. Thus, the first decision tree may be used to classify the buyer into a good or bad category. These categories and corresponding buyers can be analyzed and a determination can be made regarding the errors made and location where they occurred. Thus, a second decision tree may be trained to try and “fix” the mistakes. A third decision tree then continues by focusing on any errors on the second decision tree. [0024] Process 400 may begin with operation 402, where data is retrieved. The data retrieved may come in the form of a user data set, user transaction information, account information, buyer information, seller information, etc. In some instances, the data retrieved is historical data, while in other instances, the data may be current. Following retrieval of the data, the data retrieved may then be preprocessed at operation 404. Preprocessing of the data may include formatting or organizing the data in such a way that it may be input into a machine learning model for analysis and prediction. [0025] Once the data has been preprocessed, process 400 continues to train the machine learning model used for processing the data and making predictions. As previously indicated, large data is constantly collected by devices that oftentimes needs to be organized and analyzed. Such data can include the preprocessed data for training and using with the machine learning model. Oftentimes, the large data is organized using decision tree structures which can be used by the machine learning algorithms to extract the information of interest. In one embodiment, the decision tree structures may be used in training the machine learning model.)
storing, by the server system, the feature set for each data sample in the database. (in at least [0023] the machine learning models as described above and in conjunction with FIGS. 2-3, may be integrated into the system using an SQL query for generating SQL support for tree ensemble classifiers, FIG. 4 is introduced which illustrates example process 400 that may be implemented on a system, such as system 600 in FIG. 6. In particular, FIG. 4 illustrates a flow diagram illustrating how to generate and use an SQL query to translate a machine learning model for use in production. According to some embodiments, process 400 may include one or more of operations 402-414, which may be implemented, at least in part, in the form of executable code stored on a non-transitory, tangible, machine readable media that, when run on one or more hardware processors, may cause a system to perform one or more of the operations 402-414. [0040] Application servers 630 of network-based system 610 may be a server that provides various services to clients including, but not limited to, data analysis, geofence management, order processing, checkout processing, and/or the like. Application server 630 of network-based system 610 may provide services to a third party merchants such as real time consumer metric visualizations, real time purchase information, and/or the like. Application servers 630 may include an account server 632, device identification server 634, payment server 636, queue analysis server 638, purchase analysis server 640, user ID server 642, feedback server 644, and/or content statistics server 646. These servers, which may be in addition to other servers, may be structured and arranged to configure the system for monitoring queues as well as running and storing learning information for the decision ensemble tree processing.)
As per Claim 11-17, 19 for A server system (see at least Perez [0032]-[0048]), substantially recite the subject matter of Claim 1-7, 9 and are rejected based on the same reasoning and rationale.
As per Claim 20 for A non-transitory computer-readable storage medium (see at least Perez [0045]), substantially recite the subject matter of Claim 1 and are rejected based on the same reasoning and rationale.
Claims 8, 18 is/are rejected under 35 U.S.C. 103 as being unpatentable by US Patent Publication to US20190122139A1 to Perez, (hereinafter referred to as “Perez”) in view of US Patent Publication to US20230342348A1 to Wang, (hereinafter referred to as “Wang”) in view of US Patent Publication to US20210303986A1 to Saha et al., (hereinafter referred to as “Saha”)
As per Claim 8, Perez teaches: The computer-implemented method as claimed in claim 7, further comprising:
computing, by the server system, an … for each data sample based, at least in part, on the first calibration error and the second calibration error, the … of a positive impact on calibration of the intermediate prediction. (in at least [0021] To make a prediction, information known about the seller (e.g., buyers associated with the seller) may be used to help make a determination as to whether or not open an account, complete a transaction, provide credit, etc. to the seller. Thus, the first decision tree may be used to classify the buyer into a good or bad category. These categories and corresponding buyers can be analyzed and a determination can be made regarding the errors made and location where they occurred. Thus, a second decision tree may be trained to try and “fix” the mistakes. A third decision tree then continues by focusing on any errors on the second decision tree. The process continues until model converges.)
Although implied, Perez in view of Wang does not expressly disclose the following limitations, which however, are taught by Saha,
…improvisation factor…improvisation factor indicating an extent of a positive impact…(in at least [0113] The decision tree classifier 700 depicted in FIG. 7 may be an exemplary pre-trained classifier of the set of classifiers 106, however, the scope of the disclosure may not be so limited. The set of classifiers 106 may include other classifiers including, but not limited to, a Support Vector Machine (SVM) classifier, a Naïve Bayes classifier, a Logistic Regression classifier, or a k-nearest neighbor classifier. Hence, the decision tree classifier 700 (i.e. one of the set of classifiers 106) may be trained by the disclosed electronic device 102 on the decision conditions (at each of a root node 702 and the set of internal nodes 704A-704C) based on the correlation established between the features (i.e. set of feature values) and the corresponding weights (i.e. extracted from the DNN 104 based on CAM process) and the correct or incorrect prediction (i.e. indicated by the set of leaf nodes 706A-706D). Such trained classifier (such as the decision tree classifier 700 orthogonal to the DNN 104) may store or learn all the correct predictions and related correct features (i.e. success patterns) and/or incorrect predictions and related wrong features (i.e. error pattern) of the DNN 104, to further validate the correctness of the DNN 104 and further enhance the prediction accuracy of the DNN 104 in various real-time application.)
At the time the invention was filed, it would have been obvious for one of ordinary skill in the art to have modified the teachings of Perez in view of Wang, as taught by Saha above, with a reasonable expectation of success if arriving at the claimed invention. One of ordinary skill in the art would have been motivated to make this modification to the teachings of Perez in view of Wang with the motivation of, …the DNN may be improved by configuring a computing system in a manner the computing system is able to effectively determine an accuracy or a reliability of a prediction result of the DNN. The computing system may include at least one classifier (i.e. pre-trained on a certain class associated with the DNN) which may validate a prediction result of the DNN to further improve an accuracy of the prediction, as compared to other conventional systems which may validate the prediction result of classical machine learning models.…This reduction in the dimensionality of the huge datasets and features in the DNN 104 may further improve the efficiency of the classifier that may be used for the validation of the predicted first class of the first data point. Further, the selection of features (i.e. the first set of features) having weights above a certain weight threshold or the selection of top weighted features from the second set of features may ensure that the selected first set of features may be relevant and as well as important for the accurate validation of the first class predicted by the DNN 104 of the first data point...achieved a good classification accuracy in various classification tasks, however, in certain applications, such as autonomous driving or medical diagnosis, an incorrect detection or mis-classifications (in even a fraction of cases) may lead to financial or physical losses…to effectively determine the reliability of the DNN due to high dimensional nature of the data used in the DNN..., as recited in Saha.
As per Claim 18 for A server system (see at least Perez [0032]-[0048]), substantially recite the subject matter of Claim 8 and are rejected based on the same reasoning and rationale.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PO HAN MAX LEE whose telephone number is (571)272-3821. The examiner can normally be reached on Mon-Thurs 8:00 am - 7:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Rutao Wu can be reached on (571) 272-6045. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PO HAN LEE/Primary Examiner, Art Unit 3623