DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-20 are pending. Claims 1, 13, and 17 are independent.
This Application was published as US 20250315595.
Apparent priority: 4-8-2024
Information Disclosure Statement
The IDS dated 1/24/2025 has been considered and placed in the application file.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claim(s) 5 are rejected under 35 U.S.C. 112(b), as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention.
Claim(s) 5 recite “continuously:” It is unclear when the system fully satisfies the condition continuously. We are interoperating continuously to be the operations of claim 2.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The independent Claims are directed to statutory categories:
Claim 1 is a system claim and directed to the machine or manufacture category of patentable subject matter.
Claim 13 is a system claim and directed to the machine or manufacture category of patentable subject matter.
Claim 17 is a method claim and directed to the process category of patentable subject matter.
Step 2A, Prong One: Does the Claim recite a Judicially Recognized Exception? Abstract Idea? Are these Claims nevertheless considered Abstract as a Mathematical Concept (mathematical relationships, mathematical formulas or equations, mathematical calculations), Mental Process (concepts performed in the human mind (including an observation, evaluation, judgment, opinion), or Certain Methods of Organizing Human Activity (1-fundamental economic principles or practices (including hedging, insurance, mitigating risk), 2-commercial or legal interactions (including agreements in the form of contracts; legal obligations; advertising, marketing or sales activities or behaviors; business relations), 3- managing personal behavior or relationships or interactions between people (including social activities, teaching, and following rules or instructions) and fall under the judicial exception to patentable subject matter?)
The rejected Claims recite Mental Processes or Methods of Organizing Human Activity such as a human receiving data and modifying it and further making a report.
Step 2A, Prong Two: Additional Elements that Integrate the Judicial Exception into a Practical Application? Identifying whether there are any additional elements recited in the claim beyond the judicial exception(s), and evaluating those additional elements to determine whether they integrate the exception into a practical application of the exception. “Integration into a practical application” requires an additional element(s) or a combination of additional elements in the claim to apply, rely on, or use the judicial exception in a manner that imposes a meaningful limit on the judicial exception, such that the claim is more than a drafting effort designed to monopolize the exception. Uses the considerations laid out by the Supreme Court and the Federal Circuit to evaluate whether the judicial exception is integrated into a practical application.
1. A system comprising:
one or more processors; and
a memory in communication with the one or more processors and storing instructions that, when executed by the one or more processors, are configured to cause the system to:
receive first data comprising one or more first text threads;
transform the first data into modified first data by:
inserting a grammatical pattern into the one or more first text threads; and inserting one or more text phrases into the one or more first text threads adjacent to the grammatical pattern;
train a first language model to identify one or more first features from the modified first data to create a trained first language model;
receive second data comprising one or more second text threads;
transform the second data into modified second data by inserting the grammatical pattern into the one or more second text threads;
identify, via the trained first language model, the one or more first features from a first portion of the modified second data;
dynamically map the first portion of the modified second data to one or more first categories; and
generate a first customized report based on one or more of the modified second data, the one or more first features, the one or more first categories, or combinations thereof.
13. A system comprising:
one or more processors; and
a memory in communication with the one or more processors and storing instructions that, when executed by the one or more processors, are configured to cause the system to:
receive first data;
transform the first data into modified first data;
identify, via a first language model, one or more first features from a first portion of the modified first data, wherein the first language model is trained to identify the one or more first features from the modified first data based on the modified first data comprising the first data and a grammatical pattern inserted into the first data;
dynamically map the first portion of the modified first data to one or more first categories; and
generate a first customized report based on one or more of the modified first data,
the one or more first features, the one or more first categories, or combinations thereof.
17. A method of training a first language model to identify one or more first features from modified first data, the method comprising:
collecting first data comprising one or more text threads;
transforming the first data into the modified first data by:
inserting a grammatical pattern into the one or more text threads; and
inserting one or more first text phrases into the one or more text threads adjacent to the grammatical pattern;
creating a first training set comprising the first data and the modified first data; and
training the first language model using the first training set.
The limitation of “receive …”, “transform …”, “receive…”, “transform …”, “identify…”, “dynamically…”, and “generate …” , as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, a person receives business messages regarding payments, billing, disputes, or returns. When he goes through these messages, he modifies it by adding a pattern to know right away what type of message it is, such as payment, refund, cancelation, or complaint. Further he makes a report based upon this. Recap: going through messages identifying a feature, assigning them into categories, further generating a report is a mental process.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea.
This judicial exception is not integrated into a practical application. In particular, claim 1 and 13 only recites additional elements that are computer components “processors” (paragraph 59), “memory” (paragraphs 59), and “model” (paragraph 69) recited at a high-level of generality such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are directed to an abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using the computer components amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claims are not patent eligible.
Claims 2 additionally recite 2. The system of claim 1, wherein the instructions are further configured to cause the system to: determine whether the trained first language model identifies the one or more first features from a second portion of the modified second data; responsive to determining the trained first language model identifies the one or more first features from the second portion of the modified second data: dynamically map the second portion of the modified second data to the one or more first categories; and calculate one or more first statistical metrics associated with a third portion of the modified second data; responsive to determining the trained first language model fails to identify the one or more first features from the second portion of the modified second data: calculate the one or more first statistical metrics associated with the second portion of the modified second data; and dynamically map the third portion of the modified second data to the one or more first categories; and generate a second customized report based on one or more of the modified second data, the one or more first features, the one or more first categories, the one or more first statistical metrics, or combinations thereof. However, this limitation does not prevent a human from performing the steps mentally as described above. Further, he would analyze the financial transaction by identifying characteristics/transaction. Categorizing identified transaction statistically analyzing transaction for which characteristics were not identified the first time. Further generating a report. Thus, these claims are directed towards a mental process. The claim 2 recites additional elements that are computer components “model” (paragraph 69) recited at a high-level of generality such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are directed to an abstract idea.
Claims 3 additionally recites 3. The system of claim 2, wherein calculating the one or more first statistical metrics comprises transforming the second or third portion of the modified second data into a frequency space via a Fourier Transformation. However, these limitations encompass a person tallying up the unclassified information. Thus, these claims are directed towards a mental process. Similar to above, no additional limitations are provided that provide a practical application, or amount to significantly more than the abstract idea. Therefore, the claims are not patent eligible.
Claims 4 additionally recites 4. The system of claim 2, wherein the one or more first statistical metrics comprise one or more of recurring inflows, non-recurring inflows, recurring outflows, non-recurring outflows, or combinations thereof. However, these limitations encompass a person keeping a tally for each of the category’s it recognizes. Thus, these claims are directed towards a mental process. Similar to above, no additional limitations are provided that provide a practical application, or amount to significantly more than the abstract idea. Therefore, the claims are not patent eligible.
Claims 5 additionally recites 5. The system of claim 2, wherein the instructions are further configured to cause the system to: continuously: receive third data; transform the third data into modified third data; Identify, via the trained first language model, the one or more first features from a fourth portion of the modified third data; dynamically map the fourth portion of the modified third data to the one or more first categories; automatically update the first customized report in real-time based on one or more of the modified third data, the one or more first features, the one or more first categories, or combinations thereof; determine whether the trained first language model identifies the one or more first features from a fifth portion of the modified third data; responsive to determining the trained first language model identifies the one or more first features from the fifth portion of the modified third data: dynamically map the fifth portion of the modified third data to the one or more first categories; and calculate the one or more first statistical metrics associated with a sixth portion of the modified third data; responsive to determining the trained first language model fails to identify the one or more first features from the fifth portion of the modified third data: calculate the one or more first statistical metrics associated with the fifth portion of the modified third data; and dynamically map the sixth portion of the modified third data to the one or more first categories; and automatically update the second customized report in real-time based on one or more of the modified third data, the one or more first features, the one or more first categories, the one or more first statistical metrics, or combinations thereof. However, these limitations encompass a person continuously receiving messages and analyzing it for transaction data based on the features. Furthermore, creating a financial report. Thus, these claims are directed towards a mental process. The claim recites additional elements that are computer components “model” (paragraph 69) recited at a high-level of generality such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are directed to an abstract idea.
Claims 6 additionally recites 6. The system of claim 1, wherein the first and second data comprise transaction data. However, these limitations encompass a person receiving data/ modifying data that deals with transaction. Thus, these claims are directed towards a mental process. Similar to above, no additional limitations are provided that provide a practical application, or amount to significantly more than the abstract idea. Therefore, the claims are not patent eligible.
Claims 7 additionally recites 7. The system of claim 1, wherein the grammatical pattern comprises one or more characters, one or more symbols, or both. However, these limitations encompass a person receiving data/ modifying data that deals with transactions. Modifying it using the greater than and less than symbol. Thus, these claims are directed towards a mental process. Similar to above, no additional limitations are provided that provide a practical application, or amount to significantly more than the abstract idea. Therefore, the claims are not patent eligible.
Claims 8 additionally recites 8. The system of claim 7, wherein the one or more symbols comprise an equals sign, a greater-than sign, or both. However, these limitations encompass a person receiving data/ modifying data that deals with transactions. Modifying it using the greater than and less than symbol. Thus, these claims are directed towards a mental process. Similar to above, no additional limitations are provided that provide a practical application, or amount to significantly more than the abstract idea. Therefore, the claims are not patent eligible.
Claims 9 additionally recite 9. The system of claim 1, wherein the one or more first features comprise one or more of a second category, a counterparty, a payment channel, or combinations thereof. However, these limitations encompass a person analyzing the data and based on the features categorizing the data. Thus, the claim is directed towards a mental process. Similar to above, no additional limitations are provided that provide a practical application, or amount to significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Claims 10 additionally recite 10. The system of claim 1, wherein the one or more first categories comprise Profit and Loss Statement (P&L) categories. However, these limitations encompass a person analyzing the data and having a category for profit and loss data. Thus, the claim is directed towards a mental process. Similar to above, no additional limitations are provided that provide a practical application, or amount to significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Claims 11 additionally recite 11. The system of claim 1, wherein the instructions are further configured to cause the system to: retrieve third data associated with a business; and train a second language model to identify one or more second features associated with the business from the third data, wherein training the first language model to identify the one or more first features from the modified first data is based on the one or more second features associated with the business, and wherein the first customized report is unique to the business. However, these limitations encompass a person receiving data about a business identifying data about it (profits and loss), further categorizing that data. Further generating a report. Thus, these claims are directed towards a mental process. The claim recites additional elements that are computer components “model” (paragraph 69) recited at a high-level of generality such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are directed to an abstract idea.
Claims 12 additionally recite 12. The system of claim 11, wherein retrieving the third data is conducted via a search engine, a web-scraper, or both. However, these limitations encompass a person receiving the data. Thus, the claim is directed towards a mental process. Similar to above, no additional limitations are provided that provide a practical application, or amount to significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Claim 13 contains limitations similar to those found in claim 1 and therefore are not patent eligible for the same reasons.
Claim 14 contains limitations similar to those found in claim 2 and therefore are not patent eligible for the same reasons.
Claim 15 contains limitations similar to those found in claim 11 and therefore are not patent eligible for the same reasons.
Claim 16 contains limitations similar to those found in claim 11 and therefore are not patent eligible for the same reasons.
Claim 17 contains limitations similar to those found in claim 1 and therefore are not patent eligible for the same reasons.
Claim 18 contains limitations similar to those found in claim 1 and therefore are not patent eligible for the same reasons.
Claim 19 contains limitations similar to those found in claim 15 and therefore are not patent eligible for the same reasons.
Claim 20 contains limitations similar to those found in claim 7 and therefore are not patent eligible for the same reasons.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 6-10, 13, are rejected under 35 U.S.C. 103 as obvious over SHEN (US 20210035556) in view of Fehling (US 20200233857).
Claim 1, 13
Regarding Claim 1, SHEN teaches
SHEN:
PNG
media_image1.png
534
339
media_image1.png
Greyscale
1. A system comprising:
one or more processors; and
(“[0091] While the reader will appreciate that the above embodiments are applicable to any computing system, a typical computing system is illustrated in FIG. 4, which provides means capable of putting an embodiment, as described herein, into effect. As illustrated, the computing system 400 comprises a processor 401 coupled to a mass storage unit 403 and accessing a working memory 405. As illustrated, a language model (LM) controller 407 is represented as a software product stored in working memory 405. However, it will be appreciated that elements of the LM controller 407 may, for convenience, be stored in the mass storage unit 403.”)
a memory in communication with the one or more processors and storing instructions that, when executed by the one or more processors, are configured to cause the system to:
(“[0091] While the reader will appreciate that the above embodiments are applicable to any computing system, a typical computing system is illustrated in FIG. 4, which provides means capable of putting an embodiment, as described herein, into effect. As illustrated, the computing system 400 comprises a processor 401 coupled to a mass storage unit 403 and accessing a working memory 405. As illustrated, a language model (LM) controller 407 is represented as a software product stored in working memory 405. However, it will be appreciated that elements of the LM controller 407 may, for convenience, be stored in the mass storage unit 403.”)
receive first data comprising one or more first text threads;
(“[0057] The training data received is labelled training data suitable for supervised learning to train the system to perform the specific task required. This task may be any supervised learning task, such as any classification task (e.g., sentiment analysis, intent recognition or inference). Accordingly, the training data includes labelled observations, each including an observation (a set of text for input) and a label (an appropriate output according to the specific task).”)
transform the first data into modified first data by:
(“[0059] Given some task with some supervised dataset of (input, output) pairs, the method produces a dataset suitable for language modelling by, for each pair of input and output, concatenating the input with a task trigger corresponding to the task (the task linking the input and output) and the output for the input (in that order). That is, the task trigger is concatenated to the end of the input and the output is concatenated to the end of the task trigger. This forms a processed observation.”)
inserting a grammatical pattern into the one or more first text threads;
(“[0059] Given some task with some supervised dataset of (input, output) pairs, the method produces a dataset suitable for language modelling by, for each pair of input and output, concatenating the input with a task trigger corresponding to the task (the task linking the input and output) and the output for the input (in that order). That is, the task trigger is concatenated to the end of the input and the output is concatenated to the end of the task trigger. This forms a processed observation.”)
and inserting one or more text phrases into the one or more first text threads adjacent to the grammatical pattern;
(“[0059] Given some task with some supervised dataset of (input, output) pairs, the method produces a dataset suitable for language modelling by, for each pair of input and output, concatenating the input with a task trigger corresponding to the task (the task linking the input and output) and the output for the input (in that order). That is, the task trigger is concatenated to the end of the input and the output is concatenated to the end of the task trigger. This forms a processed observation.”)
train a first language model to identify one or more first features from the modified first data to create a trained first language model;
(“[0075] Once the processed training data has been produced, the language model is trained based on the processed training data 206. That is, the weights of the language model are updated based on an objective function applied to the processed training data. General unsupervised training can be used. The method for updating the weights of the language model may be the same as that used to train the initial language model. The only difference in this case is the training data. As the training data has been encoded with the task trigger(s) and outputs, the model learns to predict outputs when prompted with an input (or delimited inputs) and a task trigger.”)
receive second data comprising one or more second text threads;
(“[0078] Given a fine-tuned model produced from the training method above (see FIG. 2), this can be used to predict outputs for novel (unlabelled) inputs by providing the appropriate trigger.
[0079] FIG. 3 shows a method for predicting an output based on an input using a language model trained using the method of FIG. 2.
[0080] The method starts by obtaining a fine-tuned language model, along with some input for processing and a task trigger representing the task to be performed 302. The language model may be accessed from storage, received from an external source, or may be obtained through training in accordance with FIG. 2. Regardless of how the language model is obtained, it is a language model that has been trained for a specific task in accordance with the methods described herein.
[0081] The task trigger may be received with the language model (e.g., from storage, from an external source or during training), may be preconfigured (e.g., where the language model has been trained for only a single task), may be received with the input, or may be input by the user or selected by the user when prompting a task to be performed. The input includes natural language data for processing in accordance with the task indicated by the task trigger.”)
transform the second data into modified second data by inserting the grammatical pattern into the one or more second text threads;
(“[0081] The task trigger may be received with the language model (e.g., from storage, from an external source or during training), may be preconfigured (e.g., where the language model has been trained for only a single task), may be received with the input, or may be input by the user or selected by the user when prompting a task to be performed. The input includes natural language data for processing in accordance with the task indicated by the task trigger.
[0082] The input is then processed by concatenating the corresponding task trigger to the end of the input to form a processed input 304. For the running example of sentiment analysis, if the new input is “I loved this”, the processed input produced would be:
[0083] “I loved this <sentiment>””
Second text thread= I love thisgrammatical pattern =<sentiment>
insert the pattern = paragraph 59“…the task trigger is concatenated to the end of the input, and the output is concatenated to the end of the task trigger…”)
identify, via the trained first language model, the one or more first features from a first portion of the modified second data;
(“[0086] The processed input is then input into the fine-tuned language model 306. This produces a set of probabilities for the next token, as described with regard to FIG. 1. If multiple tokens are to be predicted, the language model may be applied multiple times and a prediction for the set of next tokens made (as described with regard to FIG. 1).
[0087] The output is then selected based on the one or more sets of probabilities produced by the fine-tuned language model 308. As described with regard to FIG. 1, the top prediction for what each token should be can be selected (based on probabilities for every token in the vocabulary). Alternatively, if this is too noisy and the number of possible outputs is relatively small, only the probabilities of each possible output (as judged by the model) can be considered. That is, the most probable output is selected from the set of potential outputs for the specific task, rather than selecting the most probable token from the dictionary. This can help to ensure that the output is constrained to the required set of outputs for the task at hand.
[0088] In light of the above, it can be seen that language models trained on non-specific natural language data can be easily and efficiently trained to perform specific tasks usually reserved for supervised training systems through the application of unsupervised learning on training data that has been specifically processed to encode task triggers and outputs (labels). This is achieved without any change to the architecture of the language model (or the addition of any further layers to the model) and, potentially, without any change to the unsupervised training method for the language model.”
Identify one or more features= Classification characteristic/label)
[dynamically map the first portion of the modified second data to one or more first categories; and
generate a first customized report based on one or more of the modified second data, the one or more first features, the one or more first categories, or combinations thereof.]
SHEN does not explicitly teach all of the dynamically map the first portion of the modified second data to one or more first categories; and
generate a first customized report based on one or more of the modified second data, the one or more first features, the one or more first categories, or combinations thereof.
Fehling :
PNG
media_image2.png
512
795
media_image2.png
Greyscale
However, Fehling teaches dynamically map the first portion of the modified second data to one or more first categories; and
(“[0050] FIG. 6 depicts how a model 600 can be used to select appropriate categories for the prediction set of clusters. The model receives the training data 522 which has been categorized by the base mapping operations 518 and the high-value tagging 516. The training data is used to determine model parameters used for determining the category of a log or cluster based on associated data in the MDS. Once the model is trained and applied to the prediction set 520, it provides category predictions 602. The sets 520 and 522 provide the model with the log vector for each log. In some embodiments, only a subset of log data is provided to the model, and in some cases, additional spend data in the cost database can be provided to the model. In some cases, logs might also be withdrawn from the training set during model calibration to balance the likelihood of the categories from which it will learn (e.g., avoid that a specific category is over-represented).”
“[0027] In generating the CDS, various natural language processing (NLP) steps are applied to the cost data to aid subsequent analysis of each transaction. These operations can aid in determining relevant keywords and in determining relationships between transactions for clustering. Some of these NLP operations include (1) conversion to lowercase text, (2) removing duplicate words, (3) removing punctuation, (4) removing non-alphanumeric characters, (5) removing numbers (in some cases, only from certain fields), (6) removing (in some cases) words that are less than a threshold number of characters (e.g., 2 characters), (7) removing codes identified as combinations of letters and numbers, (8) translating text to a single language (e.g., English), (9) lemmatizing words by converting words to their base dictionary form (e.g. “expenses” becomes “expense”), (10) removing month names and abbreviations, (11) removing stop words such as “for”, “the”, (12) removing city names, (13) removing proper nouns and names, (14) substituting the supplier family name if there is no supplier field, (15) removing a supplier name when present in full description, (16) selecting keywords based on predetermined lists or ad-hoc analysis like their occurrence of appearance in one/several categories, (17) using informative scoring like term frequency-inverse document frequency (TF-IDF), and (18) using Machine Learning models for Named Entity Recognition. It should be understood that there may be additional or fewer NLP operations applied to each log. Additionally, some NLP operations may only be applied to certain fields or portions of a transaction. For example, in some cases, NPL operations are only applied to invoice description fields and not to fields listing, e.g., a supplier name or a supplier's contact information.”)
generate a first customized report based on one or more of the modified second data, the one or more first features, the one or more first categories, or combinations thereof.
(“[0016] FIG. 1A depicts an application for presenting categorized spend data to a user, often referred to as a spend cube. In the depicted interface 100, expense totals 102 for various level 1 categories 104a are depicted, as well as visual data 106 (graphs, charts, diagrams, and the like) for conveying the breakdown of categorized spend data to a user. An interface may provide a variety of user selectable features for allowing a user to explore the spend data—allowing a user to quickly understand the state of the organization's spending over a selected period (e.g., a 1 month, 6 months, 1 year, or another selected amount for which spend data has been provided). The spend cube allows a user to see how the spend is distributed across and between organizational units of a company, providing transparency on the most granular level across all cost packages.
[0017] FIG. 1B depicts a screenshot 101 provided by the application showing a higher resolution view of categorized spend data and thus allowing for more in-depth analysis than shown by the screenshot in FIG. 1A. The additional detail is provided using a hierarchical category structure in which various expense logs are categorized. In the depicted interface, each level 1 category 104a is made of one or more level 2 categories 104b, and each level 2 category is in turn made of one or more level 3 categories 104c. While a three-tier hierarchical structure is depicted, there may be more or fewer tiers. In some cases, spend data may be searched or displayed based on one or more log or invoice categories that do not fall within the hierarchical category structure. This can be done in the context of spend categories 104a-c. For instance, the property expenses 108 for an international organization may be broken down by countries in which an organization operates. In some cases, a user may filter spend data based on expense logs that satisfy date criteria, location criteria, department criteria, or any other selected criteria that may be useful for classifying expense data. In other cases, a user can search spend data based on an alternate hierarchical category structure. For instance, a user may be able to choose an alternate hierarchal category structure to see how costs are distributed amongst various organizational units of the company such as business units, sub-business units or regions, countries, and the like.
[0018] In some cases, an application may be used to calculate cost metrics based on the categorized spend data. In some cases, an application may provide a user with recommendations or warnings based on spend data. For instance, an application might alert a user that a particular spend category has been highly variable over past budgeting periods and that the user or organization should plan accordingly or investigate the source of variability. It is appreciated that many software tools for presenting and further analyzing spend data known now or later developed may be used with categorized data sets produced by the methods described herein. For the sake of brevity, such known tools for analyzing and presenting spend data are not discussed in great detail here.”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified SHEN to incorporate the teachings of Fehling to provide “the dynamically map the first portion of the modified second data to one or more first categories; and generate a first customized report based on one or more of the modified second data, the one or more first features, the one or more first categories, or combinations thereof.” Doing so would help the user quickly understand the state of the organization, as recognized by Fehling . (paragraph 16).
Claim 13 contains limitations similar to those found in claim 1 and is rejected under similar rationale. Claim 13 is broader than claim 1 and does not require inserting a pattern and inserting a text phrase.
Claim 6
Regarding Claim 6, SHEN in view of Fehling teaches the limitation of claim 1.
Further Fehling teaches 6. The system of claim 1, wherein the first and second data comprise transaction data.
(“[0013] In particular, various embodiments described herein provide methods for categorizing spend data which may include general ledger (GL), accounts payable (AP), purchase order (PO) information, including but not limited to transactions, invoices, expenditure receipts, supplier-based data sets and other documented expenses, herein collectively referred to as spend logs (or simply logs). After collecting spend data from all relevant data systems and/or sources, the data is processed and consolidated to generate a cleaned data set (CDS). The CDS includes spend data that has been filtered to remove less important information and/or processed to standardize information used for log categorization. In some cases, the CDS includes an organized structure that breaks spend information down by field types (e.g., total cost, vendor, transaction date, etc.). In some cases, standardizing spend data involves applying natural language processing operations to text information associated with logs. Logs from the CDS are then clustered into groups based on a similarity of words, costs, dates, or other patterns and features, which generates a new data set of smaller size: the minimal data set (MDS). The MDS constitutes groups of logs representing the same type of transaction (e.g. “Taxi fare” and “Supplier A”). Based on clustering operations, each log within the same cluster can be mapped to the same cost category.”)
See claim 1 for rationale.
Claim 7
Regarding Claim 7, SHEN in view of Fehling teaches the limitation of claim 1.
Further SHEN teaches 7. The system of claim 1, wherein the grammatical pattern comprises one or more characters, one or more symbols, or both.
(“[0009] A task trigger may be any string or token that uniquely identifies the task being trained. The task trigger provides an indication to the language model that it is to predict an output.
[0010] According to an embodiment, the training does not adjust the architecture of the language model such that the updated language model has the same architecture as the language model. The training may instead simply update (fine-tune) the weights of the language model. The language model may be a neural network that models a probability distribution over sequences of tokens. A token may be a word, a symbol (such as for punctuation), or any other string that is specified in a dictionary (or vocabulary) for the language model.”
“[0022] A token can be considered a potential string according to a predefined dictionary. This can include words and strings of one or more characters, such as punctuation. The method may select the most probable output (most probable token) based on the set of probabilities. This may be the most probable individual token.”)
Claim 8
Regarding Claim 8, SHEN in view of Fehling teaches the limitation of claim 1.
Further SHEN teaches 8. The system of claim 7, wherein the one or more symbols comprise an equals sign, a greater-than sign, or both.
(“[0059] Given some task with some supervised dataset of (input, output) pairs, the method produces a dataset suitable for language modelling by, for each pair of input and output, concatenating the input with a task trigger corresponding to the task (the task linking the input and output) and the output for the input (in that order). That is, the task trigger is concatenated to the end of the input and the output is concatenated to the end of the task trigger. This forms a processed observation.
[0060] For example, if the task is sentiment classification (with a task trigger of “<sentiment>”), and one input is “This movie is terrible” with a label (output) of “negative”, the method produces the below processed observation:
[0061] “The movie is terrible <sentiment> negative””)
Claim 9
Regarding Claim 9, SHEN in view of Fehling teaches the limitation of claim 1.
Further SHEN teaches 9. The system of claim 1, wherein the one or more first features comprise one or more of a second category, a counterparty, a payment channel, or combinations thereof.
(“[0013] According to an embodiment, the one or more natural language processing tasks are one or more classification tasks and the training outputs set are labels for corresponding training inputs in the natural language training data set. The one or more natural language processing tasks may comprise one or more of a sentiment analysis task, an intent recognition task and an inference task.”)
Claim 10
Regarding Claim 10, SHEN in view of Fehling teaches the limitation of claim 1.
Further Fehling teaches 10. The system of claim 1, wherein the one or more first categories comprise Profit and Loss Statement (P&L) categories.
(“[0040] FIG. 5 depicts aspects of clustering and categorizing expense logs when a three-tier category hierarchy is used such depicted in FIG. 1B. While explained in the context of a three-tier category hierarchy, it should be appreciated that the described process is also applicable to other category structures. Block 500 represents the consolidated cleaned data set (CDS) for expense logs in the cost database. Depending on the client and the period that the logs represent, this can represent millions of expense logs. Clustering (502) the logs from the CDS results in the minimal data set (MDS) 512 which constitutes cluster groups 506, 508, and 510. After the clustering, the log groups are analyzed with their respective client's account structure from the profit & loss statement, general ledger, AP and PO Accrual Reconciliation Report. Depending on the result of that analysis, the clusters are then split into cluster groups 506, 508, and 510 to facilitate the categorization effort. Logs in block 504 represent logs which cannot be clustered generally due to missing data.”)
See claim 1 for rationale.
Claims 11, 15, 16 are rejected under 35 U.S.C. 103 as obvious over SHEN in view of Fehling in further view of LU (US 20250173782)
Claim 11
Regarding Claim 11, SHEN, do not explicitly teach 11. The system of claim 1, wherein the instructions are further configured to cause the system to:
retrieve third data associated with a business; and
train a second language model to identify one or more second features associated with the business from the third data,
wherein training the first language model to identify the one or more first features from the modified first data is based on the one or more second features associated with the business, and
wherein the first customized report is unique to the business.
However, Fehling teaches 11. The system of claim 1, wherein the instructions are further configured to cause the system to:
retrieve third data associated with a business; and
(“[0061] In phase 702 the raw spend data is received from a client and consolidated in the cost database. As discussed, this may include operations such as digitizing and/or recognizing text in documents, joining data fields, conversion into a standardized target schema and reconciliation of negative and out-of-scope spend. The data is then consolidated, cleaned (e.g., corrupt and duplicate logs are removed), and various natural language processing operations are applied to recognize text, resulting in the consolidated cleaned data set (CDS). In phase 704 the minimal data set (MDS) is generated using various clustering techniques on the CDS multi-dimensional logs. In operation 706, base mapping rules are applied to automatically map logs to level 3 cost categories. In some cases, this phase can account for, categorizing about 10%-20% of the total spend. In cases, where an in-depth knowledge of a client's practice is known or where the disclosed categorization processes have been performed for a prior budgeting period, a higher percentage of the total spend may be tagged for during this phase. “)
wherein the first customized report is unique to the business.
(“[0014] In some cases, a hierarchical category structure is determined at least in part on the clustering structure, and in some cases, a category structure is based on particular client needs. The logs are then tagged or categorized in phases, where one or more representative logs from each cluster are used to determine category information for each of the logs associated with the cluster...”
“[0016] FIG. 1A depicts an application for presenting categorized spend data to a user, often referred to as a spend cube. In the depicted interface 100, expense totals 102 for various level 1 categories 104a are depicted, as well as visual data 106 (graphs, charts, diagrams, and the like) for conveying the breakdown of categorized spend data to a user. An interface may provide a variety of user selectable features for allowing a user to explore the spend data—allowing a user to quickly understand the state of the organization's spending over a selected period (e.g., a 1 month, 6 months, 1 year, or another selected amount for which spend data has been provided). The spend cube allows a user to see how the spend is distributed across and between organizational units of a company, providing transparency on the most granular level across all cost packages.”
“[0019] As used herein, a client may be any organization, business, individual, group, or entity which maintains logs of business expenses (also referred to as the client's spend)...”)
See claim one for rationale.
SHEN in view of Fehling, do not explicitly teach train a second language model to identify one or more second features associated with the business from the third data,
wherein training the first language model to identify the one or more first features from the modified first data is based on the one or more second features associated with the business, and
However, LU teaches train a second language model to identify one or more second features associated with the business from the third data,
(“[0023] The industry embedding model may be trained using historical business transactions involving a plurality of different payors and a plurality of different payees. In one example, the industry embedding model comprises a Bidirectional Encoder Representations from Transformer (BERT) model, which involves the use of masked language modeling to determine embeddings. In a particular example, the embedding model comprises a Sentence-BERT model. In other embodiments, the embedding model may involve embedding techniques such as Word2Vec and GloVe embeddings. These are included as examples, and other techniques for generating embeddings are possible.
[0024] In some embodiments, the industry embedding model may be configured to convert the historical business transactions into a vector representation. For example, in some embodiments, each of the historical business transactions may be provided to the industry embedding model as a string of text (e.g., a sentence) that may be converted into the vector representation. In alternative embodiments, the industry embedding model may, itself, generate the sentence for each of the historical business transactions by concatenating the industry name descriptor for the payor to the unique payor/payee combination.”
“[0041] It should be appreciated that an industry name embedding generally refers to a vector representation of a particular industry in n-dimensional space. Furthermore, the particular industry that the industry name embedding represents may be included in a list of different industries recognized by an entity (e.g., government agency). In operation, the industry embedding model 120 may generate a first industry name embedding representative of a first industry (e.g., transportation) included in the list of different industries for a transaction in the historical data 114 that involves a payor having an industry description (e.g., rideshare driver) associated with the first industry. Additionally, the industry embedding model 120 may generate a second industry name embedding representative of a second industry included in the list of different industries for a transaction in the historical data 114 that involves a payor having an industry description associated with the second industry.”)
wherein training the first language model to identify the one or more first features from the modified first data is based on the one or more second features associated with the business, and
(“[0043] The training pipeline 110 may include a transaction embedding model 124. The transaction embedding model 124 may be trained using the historical data 114 retrieved from the data store 106. For instance, transactions (that is, both business and personal) included in the historical data 114 may be provided to the transaction embedding model 124. In some embodiments, each of the transaction may be provided as a sentence that includes a plurality of different features. For example, the features may include, without limitation, a description of the payee for a particular transaction.
[0044] In alternative embodiments, the transaction embedding model 124 may simply be provided all the transactions included in the historical data 114 and may, itself, generate a sentence for each of the transactions. For example, the transaction embedding model 124 may generate a sentence for a particular transaction by concatenating a description for the payee associated with a particular transaction to a unique payor/payee combination associated with the particular transaction.
[0045] It should be appreciated that the input features provided to the transaction embedding model 124 may include more or fewer features than discussed above. For instance, in some embodiments, the industry name embeddings 122 generated by the industry embedding model 120 may be an input feature for the transaction embedding model 124.
[0046] To generate transaction embeddings 126, the transaction embedding model 124 may perform one or more operations. For example, the transaction embedding model 124 may convert each of the transactions into a suitable form (e.g., sentences) for use as training data for the transaction embedding model 124. Alternatively, or additionally, the transaction embedding model 124 may strip out sensitive information in each of the transactions as appropriate (e.g., by using machine learning, rules, and/or regular expressions to identify such sensitive information and removing identified sensitive information and/or replacing such sensitive information with dummy information).”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified SHEN in view of Fehling to incorporate the teachings of LU to provide a “train a second language model to identify one or more second features associated with the business from the third data, wherein training the first language model to identify the one or more first features from the modified first data is based on the one or more second features associated with the business, and” Doing so would improve predictions for classifications of transactions, as recognized by LU. (paragraph 64).
Claim 15 contains limitations similar to those found in claim 11 and is rejected under similar rationale.
Claim 16 contains limitations similar to those found in claim 11 and is rejected under similar rationale.
Claims 17 and 20 are rejected under 35 U.S.C. 103 as obvious over SHEN in view of Stabler ( US 20210224486 )
Claim 17
Regarding Claim 17, SHEN, teaches 17. A method of training a first language model to identify one or more first features from modified first data, the method comprising:
collecting first data comprising one or more text threads;
(“[0057] The training data received is labelled training data suitable for supervised learning to train the system to perform the specific task required. This task may be any supervised learning task, such as any classification task (e.g., sentiment analysis, intent recognition or inference). Accordingly, the training data includes labelled observations, each including an observation (a set of text for input) and a label (an appropriate output according to the specific task).”)
transforming the first data into the modified first data by:
(“[0059] Given some task with some supervised dataset of (input, output) pairs, the method produces a dataset suitable for language modelling by, for each pair of input and output, concatenating the input with a task trigger corresponding to the task (the task linking the input and output) and the output for the input (in that order). That is, the task trigger is concatenated to the end of the input and the output is concatenated to the end of the task trigger. This forms a processed observation.”)
inserting a grammatical pattern into the one or more text threads; and
(“[0059] Given some task with some supervised dataset of (input, output) pairs, the method produces a dataset suitable for language modelling by, for each pair of input and output, concatenating the input with a task trigger corresponding to the task (the task linking the input and output) and the output for the input (in that order). That is, the task trigger is concatenated to the end of the input and the output is concatenated to the end of the task trigger. This forms a processed observation.”)
inserting one or more first text phrases into the one or more text threads adjacent to the grammatical pattern;
(“[0059] Given some task with some supervised dataset of (input, output) pairs, the method produces a dataset suitable for language modelling by, for each pair of input and output, concatenating the input with a task trigger corresponding to the task (the task linking the input and output) and the output for the input (in that order). That is, the task trigger is concatenated to the end of the input and the output is concatenated to the end of the task trigger. This forms a processed observation.”)
[creating a first training set comprising the first data and the modified first data; and
training the first language model using the first training set.]
SHEN does not explicitly teach creating a first training set comprising the first data and the modified first data; and
training the first language model using the first training set.
However, Stabler teaches
creating a first training set comprising the first data and the modified first data; and
(“[0072] In some embodiments, the components of the functional architecture 300 may operate according to the following pseudocode. It should be noted, however, that other implementations of the functional architecture 300 are also possible. The operations of the machine learning algorithm 308 may be expressed as follows. [0073] Data:=Seed Data, (Input, Value) Pairs [0074] Model:=ML(Data) [0075] While Accuracy(Model)<Requirement: [0076] Data+=Worst Case Adversaries(Ext, Rate, Model) [0077] Model:=ML(Data)”
“[0068] The symmetries selected with the symmetries definition process 312 are used by an adversarial sample generation process 314, which represents an automated process that implements the symmetries selected by the symmetries definition process 312. In this example, the adversarial sample generation process 314 receives an intermediate model 316 (denoted Model.sub.i) generated by the machine learning algorithm 308 in one iteration of the training process. The adversarial sample generation process 314 uses the intermediate model 316 to produce additional training data 318 (denoted Data.sub.i+=1) for use during a subsequent iteration of the training process. The additional training data 318 includes additional linguistic samples, which represent initial linguistic samples from the initial training data 306 that have been modified in accordance with one or more of the symmetries selected by the symmetries definition process 312. At least some of the additional linguistic samples selected (based on the model 316) for use in the subsequent training iteration are adversarial examples, meaning the additional linguistic samples are selected based on their likelihood of being misclassified by the machine learning algorithm 308 during the subsequent iteration of the training process.
[0069] Note that the adversarial sample generation process 314 can make one or multiple changes to the initial linguistic samples from the initial training data 306 in order to generate the additional linguistic samples in the additional training data 318. In some embodiments, for example, the adversarial sample generation process 314 may make single changes to the initial linguistic samples in order to generate additional linguistic samples. If more adversarial examples are needed, the adversarial sample generation process 314 may then make two changes to the initial linguistic samples in order to generate more additional linguistic samples. The number of changes may continue to increase until a desired number of adversarial examples is obtained or some threshold number of changes is met. Note, however, that the additional linguistic samples may be generated in any other suitable manner based on any suitable number of changes to the initial linguistic samples. Also note that the additional linguistic samples selected for use in the additional training data 318 may represent adversarial examples that are as close as possible to the initial linguistic samples from the initial training data 306 while still causing the machine learning algorithm 308 to misclassify the additional linguistic samples.”)
training the first language model using the first training set.
(“[0066] A machine learning algorithm 308 is executed and used to train a machine learning model using (among other things) the initial training data 306, plus additional training data that is generated as described below. The training performed using the machine learning algorithm 308 is typically iterative in nature. In this type of training process, the machine learning algorithm 308 receives training data and generates an intermediate language model based on that training data, and a determination is made whether the intermediate language model is adequately accurate. If not, the machine learning algorithm 308 performs another training iteration (possibly using more or different training data) to generate another intermediate language model, and a determination is made whether that intermediate language model is adequately accurate. Accuracy here can be measured in any suitable manner, such as by using F.sub.1 scores. This process can be repeated over any number of iterations, typically until a language model is trained that has at least some desired level of accuracy. The language model can then be output as a final machine learning model 310, and the model 310 may then be used in a desired natural language application. The model 310 may be used in any suitable natural language application, such as a conversational assistant, a question answering (QA) system, an IoT home automation interface, a video game system, or an educational system. Note that any suitable machine learning algorithm 308 (now known or later developed) may be used here depending on the application.”
“[0072] In some embodiments, the components of the functional architecture 300 may operate according to the following pseudocode. It should be noted, however, that other implementations of the functional architecture 300 are also possible. The operations of the machine learning algorithm 308 may be expressed as follows. [0073] Data:=Seed Data, (Input, Value) Pairs [0074] Model:=ML(Data) [0075] While Accuracy(Model)<Requirement: [0076] Data+=Worst Case Adversaries(Ext, Rate, Model) [0077] Model:=ML(Data)”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified SHEN to incorporate the teachings of Stabler to provide a “creating a first training set comprising the first data and the modified first data; and training the first language model using the first training set.” Doing so would increase the models accuracy and be more effective than standard data augmentation, as recognized by Stabler . (paragraph 124).
Claim 20
Regarding Claim 20, SHEN in view of Stabler teaches the limitation of claim 17.
Further SHEN teaches 20. The method of claim 17, wherein the grammatical pattern comprises one or more characters, one or more symbols, or both.
(“[0009] A task trigger may be any string or token that uniquely identifies the task being trained. The task trigger provides an indication to the language model that it is to predict an output.
[0010] According to an embodiment, the training does not adjust the architecture of the language model such that the updated language model has the same architecture as the language model. The training may instead simply update (fine-tune) the weights of the language model. The language model may be a neural network that models a probability distribution over sequences of tokens. A token may be a word, a symbol (such as for punctuation), or any other string that is specified in a dictionary (or vocabulary) for the language model.”
“[0022] A token can be considered a potential string according to a predefined dictionary. This can include words and strings of one or more characters, such as punctuation. The method may select the most probable output (most probable token) based on the set of probabilities. This may be the most probable individual token.”)
Claims 18 are rejected under 35 U.S.C. 103 as obvious over SHEN in view of Stabler in further view of Vu (US 20230153687)
Claim 18
Regarding Claim 18, SHEN in view of Stabler teaches the limitations of claim 17.
Further SHEN teaches [18. The method of claim 17, further comprising:
determining whether the first data comprises one or more additional features;
responsive to determining the first data comprises the one or more additional features:]
transforming the first data into modified second data by:
(“[0014] According to an embodiment, the method further comprises training the updated language model to perform one or more further tasks. This comprises: obtaining a further training data set comprising further training inputs and corresponding further training outputs, wherein each output represents a result of a mapping from a corresponding training input via a corresponding further task of the one or more further tasks; combining each further training input with its corresponding further training output and a further task trigger representing its corresponding further task to form a further set of processed training inputs; and training the updated language model to perform the one or more further tasks. The training produces a further updated language model configured to perform any one of the one or more further tasks to predict an output through processing of an input and the task trigger for one of the one or more further tasks, wherein the training of the updated language model applies unsupervised learning to the further set of processed training inputs to further update weights of the updated language model.”)
inserting the grammatical pattern into the one or more text threads; and inserting one or more second text phrases into the one or more text threads adjacent to the grammatical pattern;
(“[0059] Given some task with some supervised dataset of (input, output) pairs, the method produces a dataset suitable for language modelling by, for each pair of input and output, concatenating the input with a task trigger corresponding to the task (the task linking the input and output) and the output for the input (in that order). That is, the task trigger is concatenated to the end of the input and the output is concatenated to the end of the task trigger. This forms a processed observation.”)
[creating a second training set comprising the first data and the modified second data; and
training the first language model using the second training set.]
SHEN do not explicitly teach
18. The method of claim 17, further comprising:
determining whether the first data comprises one or more additional features;
responsive to determining the first data comprises the one or more additional features:
creating a second training set comprising the first data and the modified second data; and
training the first language model using the second training set.
However, Stabler teaches
creating a second training set comprising the first data and the modified second data; and
(“[0072] In some embodiments, the components of the functional architecture 300 may operate according to the following pseudocode. It should be noted, however, that other implementations of the functional architecture 300 are also possible. The operations of the machine learning algorithm 308 may be expressed as follows. [0073] Data:=Seed Data, (Input, Value) Pairs [0074] Model:=ML(Data) [0075] While Accuracy(Model)<Requirement: [0076] Data+=Worst Case Adversaries(Ext, Rate, Model) [0077] Model:=ML(Data)”)
training the first language model using the second training set
(“[0072] In some embodiments, the components of the functional architecture 300 may operate according to the following pseudocode. It should be noted, however, that other implementations of the functional architecture 300 are also possible. The operations of the machine learning algorithm 308 may be expressed as follows. [0073] Data:=Seed Data, (Input, Value) Pairs [0074] Model:=ML(Data) [0075] While Accuracy(Model)<Requirement: [0076] Data+=Worst Case Adversaries(Ext, Rate, Model) [0077] Model:=ML(Data)”
“[0066] A machine learning algorithm 308 is executed and used to train a machine learning model using (among other things) the initial training data 306, plus additional training data that is generated as described below. The training performed using the machine learning algorithm 308 is typically iterative in nature. In this type of training process, the machine learning algorithm 308 receives training data and generates an intermediate language model based on that training data, and a determination is made whether the intermediate language model is adequately accurate. If not, the machine learning algorithm 308 performs another training iteration (possibly using more or different training data) to generate another intermediate language model, and a determination is made whether that intermediate language model is adequately accurate. Accuracy here can be measured in any suitable manner, such as by using F.sub.1 scores. This process can be repeated over any number of iterations, typically until a language model is trained that has at least some desired level of accuracy. The language model can then be output as a final machine learning model 310, and the model 310 may then be used in a desired natural language application. The model 310 may be used in any suitable natural language application, such as a conversational assistant, a question answering (QA) system, an IoT home automation interface, a video game system, or an educational system. Note that any suitable machine learning algorithm 308 (now known or later developed) may be used here depending on the application.”
“[0070] The adversarial sample generation process 314 may be performed in any suitable manner. For example, the adversarial sample generation process 314 may be implemented using software instructions that are executed by the processor 120 of the electronic device 101, server 106, or other component(s) in FIG. 1. Note, however, that the adversarial sample generation process 314 may be performed using any other suitable component(s) in any suitable device(s) or system(s). Also, the adversarial sample generation process 314 may occur once or multiple times, possibly depending on how many iterations are performed as part of the training process using the machine learning algorithm 308. During each iteration where the adversarial sample generation process 314 generates additional training data 318, the adversarial sample generation process 314 can receive the most-recent intermediate model 316 generated by the machine learning algorithm 308, which helps the adversarial sample generation process 314 identify adversarial examples likely to be misclassified.”)
See claim 17 for rationale.
SHEN in view of Stabler, do not explicitly teach 18. The method of claim 17, further comprising:
determining whether the first data comprises one or more additional features;
responsive to determining the first data comprises the one or more additional features:
However, Vu teaches
18. The method of claim 17, further comprising:
determining whether the first data comprises one or more additional features;
(“[0031] In various embodiments, a computer-implemented method is provided that includes: obtaining an original labeled data set for training a machine learning model to classify sentiment, where each example from the original labeled data set is labeled with at least a sentiment classification; preparing a list of named entities is using one or more data sources; for each example in the original labeled data set with one or more named entities, replacing each named entity with a corresponding entity type tag to generate a labeled template data set; executing a sampling process for each entity type t within the labeled template data set to generate a first augmented invariance data set comprising one or more invariance groups having labeled examples for each entity type t, wherein the sampling process comprises: (i) selecting an example from the labeled template data set comprising an entity type tag of entity type t; and (ii) generating an invariance group by iteratively replacing the entity type tag in the labeled example with a named entity selected from the list of named entities; and training the machine learning model using the labeled examples from the first augmented invariance data set.”
“[0175] The tag replacement technique: [0176] Use a NER model to detect all named entities of interest such as PERSON, LOCATION, ORGANIZATION, etc. in each sentiment training example, [0177] Keep only high-quality named entities, e.g., those that are most frequent and with high confidence scores, [0178] For each training example with named entities, replace each named entity by the corresponding entity type tag <PER>, <LOC>, <ORG>, etc. to generate modified training examples. Table 6 shows examples of the tag replacement technique. [0179] Train sentiment analysis model using the modified training examples.”)
responsive to determining the first data comprises the one or more additional features:
(“[0175] The tag replacement technique: [0176] Use a NER model to detect all named entities of interest such as PERSON, LOCATION, ORGANIZATION, etc. in each sentiment training example, [0177] Keep only high-quality named entities, e.g., those that are most frequent and with high confidence scores, [0178] For each training example with named entities, replace each named entity by the corresponding entity type tag <PER>, <LOC>, <ORG>, etc. to generate modified training examples. Table 6 shows examples of the tag replacement technique. [0179] Train sentiment analysis model using the modified training examples.”
“[0226] At 515, for each example in the original labeled data set with one or more named entities, each named entity is replaced by a corresponding entity type tag to generate a labeled template data set. In some embodiments, for each example in the original unlabeled data set with one or more named entities, each named entity is replaced by a corresponding entity type tag to generate an unlabeled template data set. For example, for each labeled or unlabeled example, replace each of its named entities by the corresponding tag or class such as <LOC>, <PER>, or <ORG> to create a new template.
[0227] At 520, a sampling process is executed for each entity type t within the labeled template data set to generate a first augmented invariance data set comprising one or more invariance groups having labeled examples for each entity type t. The sampling process comprises: (i) selecting an example from the labeled template data set comprising an entity type tag of entity type t; and (ii) generating an invariance group by iteratively replacing the entity type tag in the labeled example with a named entity selected from the list of named entities. The labeled example may be selected randomly or based on a predefined selection protocol. The named entity may be selected randomly or based on a predefined selection protocol.”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified SHEN in view of Stabler to incorporate the teachings of Vu to provide a “18. The method of claim 17, further comprising: determining whether the first data comprises one or more additional features; responsive to determining the first data comprises the one or more additional features:” Doing so would improve the sentiment analysis, as recognized by Vu . (paragraph 30).
Claims 19 are rejected under 35 U.S.C. 103 as obvious over SHEN in view of Stabler in view of Vu in further view of Luong (US 20220383206)
Claim 19
Regarding Claim 19, SHEN in view of Stabler teaches the limitations of claim 18.
Further SHEN teaches 19. The method of claim 18, further comprising:
collecting second data;
(“[0014] According to an embodiment, the method further comprises training the updated language model to perform one or more further tasks. This comprises: obtaining a further training data set comprising further training inputs and corresponding further training outputs, wherein each output represents a result of a mapping from a corresponding training input via a corresponding further task of the one or more further tasks; combining each further training input with its corresponding further training output and a further task trigger representing its corresponding further task to form a further set of processed training inputs; and training the updated language model to perform the one or more further tasks. The training produces a further updated language model configured to perform any one of the one or more further tasks to predict an output through processing of an input and the task trigger for one of the one or more further tasks, wherein the training of the updated language model applies unsupervised learning to the further set of processed training inputs to further update weights of the updated language model.”)
[identifying, via a second language model, one or more second features from the second data;
creating a third training set comprising the second data and the one or more second features; and
training the first language model using the third training set.]
SHEN in view of Stabler in view of Vu, do not explicitly teach identifying, via a second language model, one or more second features from the second data;
creating a third training set comprising the second data and the one or more second features; and
training the first language model using the third training set.
However, Luong teaches
identifying, via a second language model, one or more second features from the second data;
(“[0070] At FIG. 3C, a third machine-learned language model is shown being used to filter the set of synthetic training data. For example using the third machine-learned model to filter the plurality of different synthetic strings of tokens can include the following for each pair of unlabeled string of tokens and synthetic string of tokens: processing the pair of unlabeled string of tokens and synthetic string of tokens with the third machine-learned model to generate a predicted label; and determining whether the predicted label matches the supplied label that was supplied to generate the synthetic string of tokens. In some implementations, if the predicted label matches the supplied label, then the pair can be retained. Conversely, if the predicted label does not match the supplied label, then the pair can be discarded.”)
creating a third training set comprising the second data and the one or more second features; and
(“[0076] Turning to FIG. 4B, the operations can include accessing a set of unlabeled training data associated with the target task. The set of unlabeled training data can include unlabeled training examples that are in-domain for the target task. The operations can include processing each unlabeled training data with the current student model to respectively generate a synthetic label for each unlabeled training data. In particular, the unlabeled training examples and synthetic labels can be combined or associated to form a set of self-labeled training data.”)
training the first language model using the third training set.
(“[0072] At FIG. 3D, a second machine-learned model is shown being trained using the set of synthetic training data (e.g., the data FIG. 3B or the filtered data from FIG. 3C). In particular, the second machine-learned language model can process the unlabeled string and the synthetic string to generate an output. The output can be compared to a label (e.g., the supplied label or the predicted label) to train the second machine-learned language model. At FIG. 3E, the second machine-learned model is shown being further trained on a set of labeled training data that is in-domain relative to the target task.”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified SHEN in view of Stabler in view of Vu to incorporate the teachings of Luong to provide a “identifying, via a second language model, one or more second features from the second data; creating a third training set comprising the second data and the one or more second features; and training the first language model using the third training set.” Doing so would improve learning with few training examples, as recognized by Luong. (paragraph 5).
Claims 12 are rejected under 35 U.S.C. 103 as obvious over SHEN in view of Fehling in view of LU in further view of Crabtree (US 20240195833)
Claim 12
Regarding Claim 12, SHEN in view of Fehling in further view of LU , do not explicitly teach 12. The system of claim 11, wherein retrieving the third data is conducted via a search engine, a web-scraper, or both.
However, Crabtree teaches 12. The system of claim 11, wherein retrieving the third data is conducted via a search engine, a web-scraper, or both.
(“[0024] According to a preferred embodiment of the invention, 1. A system for fully integrated collection of business impacting data, analysis of that data and generation of both analysis driven business decisions and analysis driven simulations of alternate candidate business decision comprising: a business data retrieval engine stored in a memory of and operating on a processor of a computing device, a business data analysis engine stored in a memory of and operating on a processor of a computing device and a business decision and business action path simulation engine stored in a memory of and operating on a processor of one of more computing devices. The business information retrieval engine: retrieves a plurality of business related data from a plurality of sources, …
[0025] According to another embodiment of the invention, the system's business information retrieval engine a stored in the memory of and operating on a processor of a computing device, employs a portal for human interface device input at least a portion of which are business related data and at least another portion of which are commands and parameters related to the conduct of a current business analysis campaign. The business information retrieval engine employs a high volume deep web scraper stored in the memory of an operating on a processor of a computing device, …”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified SHEN in view of Fehling in view of LU to incorporate the teachings of Crabtree to provide a “12. The system of claim 11, wherein retrieving the third data is conducted via a search engine, a web-scraper, or both.” Doing so would retrieve business information from many sources in a better way, as recognized by Crabtree. (paragraph 47).
Allowable Subject Matter
Claim 2, 3, 4, 5, and 14 if rewritten to overcome the rejection(s) under 35 U.S.C. 101, and if rewritten in independent form including all of the limitations of the base claim and all limitations of any intervening claims, would comprise a particular combination of elements, which is neither taught nor suggested by the prior art.
Reference Cited
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure.
US 20190035032 to Soufiani discloses a Fourier transform on transaction data to identify recurring versus non-recurring transactions patterns.
US 20130013469 to Krakowiecki discloses retrieving transaction data categorizing them and further generating a report.
US 20250094707 to Portisch discloses identifies entities, retrieves related contextual facts to insert it to create a modified text.
US 20230289538 to Goel discloses a first model to generate labeled/parsed training example for a second model.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALI M HASSAN whose telephone number is (571)272-5331. The examiner can normally be reached Monday - Friday 8:00am - 4:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras Shah can be reached at (571)270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALI M HASSAN/Examiner, Art Unit 2653
/Paras D Shah/Supervisory Patent Examiner, Art Unit 2653
09/19/2026