DETAILED ACTION
Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
2. Receipt of Applicant’s Preliminary Amendment filed on 12/04/2023 is acknowledged. The preliminary amendment includes the amending of the specification.
Priority
3. Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Information Disclosure Statement
4. The information disclosure statements (IDS) submitted on 12/04/2023 and 02/17/2025 have been received, entered into the record, and considered. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner.
Election/Restrictions
5. Applicant’s election without traverse of Group II (Claims 9-12 and 13-17) in the reply filed on 07/08/2026 is acknowledged.
Claim Objections
6. Claim 13 is objected to because of the following informalities: The limitation “determining a designated of vectorizer functions for each feature based on a set vectorizer function decision rule and a feature attribute” and should be replaced with “determining designated vectorizer functions for each feature based on a set vectorizer function decision rule and a feature attribute”. Appropriate correction is required.
Dependent claims 14-17 are objected for incorporating the deficiencies of independent claim 13.
Claim Rejections - 35 USC § 112
7. The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
8. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
9. Claim 9 recites the limitation "determining a set of vectorizer functions for each eature based on a designated vectorizer function decision rule and a feature attribute" in Page 29. There is insufficient antecedent basis for this limitation in the claim as no “eature” is claimed earlier in the claims. The examiner suggests that the applicants replace the term “eature” with “feature” to absolve this issue.
Dependent claims 10-12 are rejected for incorporating the deficiencies of independent claim 9.
Claim Rejections - 35 USC § 101
10. 35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
11. Independent claim 13 is rejected under 35 U.S.C 101 because the claimed invention is directed to the non-statutory subject area of electro-magnetic signals and carrier waves. Claim 13 is directed towards a “computer program product stored in a computer-readable storage medium”. The examiner interprets a “computer-readable storage medium” as a medium defined by the characteristics in Page 25 of the applicant’s specification. According to Page 25 of the applicant’s specification, “A computer program may include instructions executed by the processor 310, and may be stored on a non-transitory computer readable storage medium, and the instructions cause the processor 310 to execute the operation of the present disclosure” (Page 25, lines 13-15). Thus, independent claim 13 is rejected for containing nonstatutory subject matter of carrier/propagated signals/waves because the instant claim is not limited to a non-transitory computer readable storage medium. The examiner interprets a computer-readable storage medium as being directed towards the non-statutory subject matter of carrier waves. The examiner suggests that applicants replace “computer program product stored in a computer-readable storage medium” with “computer program product stored in a non-transitory computer-readable storage medium” to absolve this issue.
Dependent claims 14-17 are rejected for incorporating the deficiencies of independent claim 13.
12. Claims (13-15) are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more.
Under the 2019 PEG, when considering subject matter eligibility under 35 U.S.C. § 101, it must be determined whether the claim is directed to one of the four statutory categories of invention, i.e., process, machine, manufacture, or composition of matter (step 1). If the claim does fall within one of the statutory categories, it must then be determined whether the claim is directed to a judicial exception (i.e., law of nature, natural phenomenon, and abstract idea) (step 2A prong 1), and if so, it must additionally be determined whether the claim is integrated into a practical application (step 2A prong 2). If an abstract idea is present in the claim without integration into a practical application, any element or combination of elements in the claim must be sufficient to ensure that the claim amounts to significantly more than the abstract idea itself (step 2B).
In the instant case, claims (13-15) are directed to a computer program. Thus, each of the claims falls within one of the four statutory categories. However, the claims also fall within the judicial exception of an abstract idea.
Under Step 2A Prong 1, the test is to identify whether the claims are “directed to” a judicial exception. The examiner notes that the claimed invention is directed to an abstract idea in that the instant application is directed to mental processes, specifically transforming data.
The examiner further notes that claims (13-15) recite a computer program for transforming data which is similar to themes defined above of method of mental processes such as performing the transformation of data, and is similar to the abstract idea identified in the 2019 PEG in grouping “c” in that the claims recite certain methods of mental processes such as performing the transformation of data. The limitations, substantially comprising the body of the claim, recite a process of transforming data. The examiner notes that the claimed invention transforms data. Because the limitations above closely follow the steps in transforming data, and the steps of the claims involve mental processes, the claim recites an abstract idea consistent with the “mental processes” grouping set forth in the 2019 PEG.
Claim 13:
A computer program stored in a computer-readable storage medium and comprising instructions for causing at least one processor to execute: receiving patient-specific medical data and storing feature information including values of features included in the medical data in a feature data table;
in the feature data table, checking at least one feature to be transformed; and
looking up a feature type of each feature by referring to a feature metadata store;
looking up vectorizer functions mapped to the feature type by referring a vectorizer store; and
determining a designated of vectorizer functions for each feature based on a set vectorizer function decision rule and a feature attribute;
generating transformed data by applying at least one specified vectorizer function to the feature to be transformed, according to a transformation condition set for each vectorizer function; and
generating input data for an artificial intelligence model by using the generated transformed data.
These limitations, as drafted, is an apparatus that, under its broadest reasonable interpretation, covers the performance of mental processes specifically transforming data. Transforming data has long before the modern computer was invented, and continues to be predominantly a product of human endeavor. The instant application is directed to transforming data. Moreover, the checking of at least one feature to be transformed can be performed by a human via their mind and/or pen & paper. Additionally, the lookup of a feature type for a feature can be performed by a human via their mind and/or pen & paper. Furthermore, the lookup of vectorizer functions for a feature type can be performed by a human via their mind and/or pen & paper. Moreover, the determination of designated vectorizer functions for features based on a rule and an attribute can be performed by a human via their mind and/or pen & paper. Additionally, the generation of transformed data via an applied vectorizer function can be performed by a human via their mind and/or pen & paper. Furthermore, the generation of input data can be performed by a human via their mind and/or pen & paper. Because the limitations above closely follow the steps of transforming data, and the steps involved human judgments, observations and evaluations that can be practically or reasonably performed in the human mind and/or pen & paper, the claim recites an abstract idea consistent with the “mental process” grouping set forth in the 2019 PEG.
The mere nominal recitation of generic computing components such as a “computer-readable storage medium”, “at least one processor”, “feature data table”, “feature metadata store”, and “vectorizer store” do not take the claim out of certain methods of mental processes grouping. Therefore, the limitation is directed to an abstract idea.
If the claims are directed toward the judicial exception of an abstract idea, it must then be determined under Step 2A Prong 2 whether the judicial exception is integrated into a practical application. The Examiner notes that considerations under Step 2A Prong 2 comprise most the consideration previously evaluated in the context of Step 2B. The Examiner submits that the considerations discussed previously determined that the claim does not recite “significantly more” at Step 2B would be evaluated the same under Step 2A Prong 1 and result in the determination that the claim does not integrate the abstract idea into a practical application. Specifically, the receiving of patient medical data is simply a data gathering step that is an insignificant extra-solution activity and does not integrate the abstract idea into a practical application. Moreover, the storage of patient medical data is simply a data storage step that is an insignificant extra-solution activity and does not integrate the abstract idea into a practical application.
The instant application fails to integrate the judicial exception into a practical application because the instant application merely recites words “apply it” (or an equivalent) with the judicial exception or merely includes instructions to implement an abstract idea. The instant application is directed to an apparatus instructing the reader to implement the identified apparatus of mental processes of transforming data. The elements of the claim do not themselves amount to an improvement to the computer, to a technology or another technical field.
Here, the claim elements entirely comprise the abstract idea, leaving little if any aspects of the claim for further consideration under Step 2A Prong 2. In short, the claims have failed to integrate a practical application (see at least 84 Fed. Reg. (4) at 55). Under the 2019 PEG, this supports the conclusion that the claim is directed to an abstract idea, and the analysis proceeds to Step 2B.
While many considerations in Step 2A need not be reevaluated in Step 2B because the outcome will be the same. Here, on the basis of the additional elements other than the abstract idea, considered individually and in combination as discussed above, the Examiner respectfully submits that claim 13 does not contain any additional elements that individually or as an ordered combination amount to an inventive concept and the claims are ineligible.
With respect to the dependent claims do not recite anything that is found to render the abstract idea as being transformed into a patent eligible invention. The dependent claims are merely reciting further embellishments of the abstract idea and do not claim anything that amounts to significantly more than the abstract idea itself.
With respect to the dependent claims, they have been considered and are not found to be reciting anything that amounts to being significantly more than the abstract idea. Claims 14-15 are directed to further embellishments of the central theme of the abstract idea in that the claims are directed to further embellishments of the transforming data of the steps of claim 13 and do not amount to significantly more.
Specifically, claim 14 recites the storage of feature types, vectorizer functions, and a transformation condition which is simply a data storage step that is an insignificant extra-solution activity and does not integrate the abstract idea into a practical application.
Furthermore, claim 15 recites the receiving of feedback data which is simply a data gathering step that is an insignificant extra-solution activity and does not integrate the abstract idea into a practical application. Furthermore, the updating of a rule can be performed by a human via their mind and/or pen & paper. Furthermore, the storage of generated AI models is simply a data storage step that is an insignificant extra-solution activity and does not integrate the abstract idea into a practical application.
Moreover, claim 16 is not directed towards an abstract idea as the specifically defined queuing process for subsequent vectorization is significantly more than the abstract idea.
Additionally, claim 17 is not directed towards an abstract idea for depending on claim 16.
Claim Rejections - 35 USC § 103
10. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
11. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
12. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
13. Claims 13-15 are rejected under 35 U.S.C. 103 as being unpatentable over Yang et al. (Article entitled “Understanding Traditional Chinese Medicine Via Statistical Learning of Expert-Specific Electronic Medical Records”, dated 26 March 2019), in view of McPherson et al. (U.S. PGPUB 2014/0059038), and further in view of Shaked et al. (U.S. PGPUB 2017/0300814).
14. Regarding claim 13, Yang teaches a computer program comprising:
A) instructions for causing at least one processor to execute: receiving patient-specific medical data and storing feature information including values of features included in the medical data in a feature data table (Abstract, Page 213);
B) in the feature data table, checking at least one feature to be transformed (Page 222).
The examiner notes that Yang teaches “instructions for causing at least one processor to execute: receiving patient-specific medical data and storing feature information including values of features included in the medical data in a feature data table” as “Processing the text data in the archive via a series of data processing steps, we transformed the semi structured EMRs in the archive to a well-structured feature table” (Abstract) and “Figure 1 shows the route map of the data processing procedure, which digests the original archive as the input, and returns the following outputs: (i) a feature codebook F which encodes all features generated from the archive, (ii) a term dictionary D which fully covers the vocabulary specific to the archive (including all background words, common TCM terms and special terms used by Prof. Zhou), (iii) a term-feature map M which links terms in D and the standard feature codes they correspond to, and the most importantly, (iv) a well-organized structured feature table T with columns for different features and rows for different records. Different from the raw data in the archive, which delivers information via semi-structured and unstructured texts, the transformed two-dimensional feature table T encodes information with a well-designed data format and coding system” (Page 213). The examiner further notes that the EMR data of a patient (i.e. the claimed patient-specific medical data) is transformed into a feature table (i.e. the claimed feature data table) that is stored. The examiner further notes that Yang teaches “in the feature data table, checking at least one feature to be transformed” as “we analyzed the structured feature tables obtained from the Zhou Archive 222 from an alternative perspective via embedding methods [32–34]. Different from previous correlation and enrichment analyses, embedding analysis considers co-occurrence patterns of different features globally, and embeds features with no geometric meanings (e.g., symptoms and herbs) into a linear space with geometric interpretation” (Page 222). The examiner further notes that performing embedding analysis on the feature table entails the checking of at least one feature to be transformed.
Yang does not explicitly teach:
C) looking up a feature type of each feature by referring to a feature metadata store.
McPherson, however, teaches “looking up a feature type of each feature by referring to a feature metadata store” as “the data access application 110 may utilize the metadata 116 retrieved from or associated with the database 114 to determine data types for the columns in the result set 204, and then utilize data type interpretation to column data types in the filter rules 122 to determine the columns to which each filter term 210 is to be applied” (Paragraph 34).
The examiner further notes that McPherson teaches the concept of determining a data type (i.e. a feature type) from stored metadata (i.e. the claimed feature metadata store). The combination would result in looking up the feature types for the features of Yang.
It would have been obvious to one of ordinary skill in the art before the effective filing date of instant invention to combine the teachings of the cited references because teaching McPherson’s would have allowed Yang to provide a method for accessing metadata to determine the structure of structured data, as noted by McPherson (Paragraph 18).
Yang and McPherson do not explicitly teach:
D) looking up vectorizer functions mapped to the feature type by referring a vectorizer store; and
E) determining a designated of vectorizer functions for each feature based on a set vectorizer function decision rule and a feature attribute;
F) generating transformed data by applying at least one specified vectorizer function to the feature to be transformed, according to a transformation condition set for each vectorizer function;
G) generating input data for an artificial intelligence model by using the generated transformed data.
Shaked, however, teaches “looking up vectorizer functions mapped to the feature type by referring a vectorizer store” as “The system processes a first set of features from the obtained features using a deep machine learning model to generate a deep model intermediate predicted output (step 204). As described above, the deep machine learning model includes a deep neural network and an embedding layer that includes embedding functions. In some implementations, the system applies the embedding layer to a subset of the first set of features. In particular, the system uses each of the embedding functions for each of the feature type of the features in the subset to generate a numeric embedding, e.g., a floating-point vector representation, of the feature. Depending on the feature type and on the implementation, the embedding function for a given feature type can be any of a variety of embedding functions” (Paragraph 43), “For example, for a feature type whose features consist of a single token, the embedding function may be a simple embedding function. A simple embedding function maps a single token to a floating point vector, i.e., a vector of floating point values. For example, the simple embedding function may map a token “cat” to a vector [0.1, 0.5, 0.2] and the word “iPod” to a vector [0.3, 0.9, 0.0], based on current parameter values, e.g., using a particular lookup table” (Paragraph 44), and “As another example, for a feature type whose features can potentially consist of a list of two or more tokens, the embedding function may be a parallel embedding function. A parallel embedding function maps each token in a list of tokens to a respective floating point vector and outputs a single vector that is the concatenation of the respective floating point vectors. For example, for an ordered list of tokens {“Atlanta”, “Hotel”}, the parallel embedding function may map “Atlanta” to a vector [0.1, 0.2, 0.3] and “Hotel” to [0.4, 0.5, 0.6], and then output [0.1, 0.2, 0.3, 0.4, 0.5, 0.6]. In order to identify the respective floating point vectors, the parallel embedding function may use a single lookup table or multiple different look up tables” (Paragraph 45), “determining a designated of vectorizer functions for each feature based on a set vectorizer function decision rule and a feature attribute” as “The system processes a first set of features from the obtained features using a deep machine learning model to generate a deep model intermediate predicted output (step 204). As described above, the deep machine learning model includes a deep neural network and an embedding layer that includes embedding functions. In some implementations, the system applies the embedding layer to a subset of the first set of features. In particular, the system uses each of the embedding functions for each of the feature type of the features in the subset to generate a numeric embedding, e.g., a floating-point vector representation, of the feature. Depending on the feature type and on the implementation, the embedding function for a given feature type can be any of a variety of embedding functions” (Paragraph 43), “For example, for a feature type whose features consist of a single token, the embedding function may be a simple embedding function. A simple embedding function maps a single token to a floating point vector, i.e., a vector of floating point values. For example, the simple embedding function may map a token “cat” to a vector [0.1, 0.5, 0.2] and the word “iPod” to a vector [0.3, 0.9, 0.0], based on current parameter values, e.g., using a particular lookup table” (Paragraph 44), and “As another example, for a feature type whose features can potentially consist of a list of two or more tokens, the embedding function may be a parallel embedding function. A parallel embedding function maps each token in a list of tokens to a respective floating point vector and outputs a single vector that is the concatenation of the respective floating point vectors. For example, for an ordered list of tokens {“Atlanta”, “Hotel”}, the parallel embedding function may map “Atlanta” to a vector [0.1, 0.2, 0.3] and “Hotel” to [0.4, 0.5, 0.6], and then output [0.1, 0.2, 0.3, 0.4, 0.5, 0.6]. In order to identify the respective floating point vectors, the parallel embedding function may use a single lookup table or multiple different look up tables” (Paragraph 45), “generating transformed data by applying at least one specified vectorizer function to the feature to be transformed, according to a transformation condition set for each vectorizer function” as “The system processes a first set of features from the obtained features using a deep machine learning model to generate a deep model intermediate predicted output (step 204). As described above, the deep machine learning model includes a deep neural network and an embedding layer that includes embedding functions. In some implementations, the system applies the embedding layer to a subset of the first set of features. In particular, the system uses each of the embedding functions for each of the feature type of the features in the subset to generate a numeric embedding, e.g., a floating-point vector representation, of the feature. Depending on the feature type and on the implementation, the embedding function for a given feature type can be any of a variety of embedding functions” (Paragraph 43), “For example, for a feature type whose features consist of a single token, the embedding function may be a simple embedding function. A simple embedding function maps a single token to a floating point vector, i.e., a vector of floating point values. For example, the simple embedding function may map a token “cat” to a vector [0.1, 0.5, 0.2] and the word “iPod” to a vector [0.3, 0.9, 0.0], based on current parameter values, e.g., using a particular lookup table” (Paragraph 44), and “As another example, for a feature type whose features can potentially consist of a list of two or more tokens, the embedding function may be a parallel embedding function. A parallel embedding function maps each token in a list of tokens to a respective floating point vector and outputs a single vector that is the concatenation of the respective floating point vectors. For example, for an ordered list of tokens {“Atlanta”, “Hotel”}, the parallel embedding function may map “Atlanta” to a vector [0.1, 0.2, 0.3] and “Hotel” to [0.4, 0.5, 0.6], and then output [0.1, 0.2, 0.3, 0.4, 0.5, 0.6]. In order to identify the respective floating point vectors, the parallel embedding function may use a single lookup table or multiple different look up tables” (Paragraph 45), and “generating input data for an artificial intelligence model by using the generated transformed data” as “The deep neural network 130 receives the numeric embeddings from the embedding layer and, optionally, other input features (e.g., feature 108) as an input” (Paragraph 36), “The system processes a first set of features from the obtained features using a deep machine learning model to generate a deep model intermediate predicted output (step 204). As described above, the deep machine learning model includes a deep neural network and an embedding layer that includes embedding functions. In some implementations, the system applies the embedding layer to a subset of the first set of features. In particular, the system uses each of the embedding functions for each of the feature type of the features in the subset to generate a numeric embedding, e.g., a floating-point vector representation, of the feature. Depending on the feature type and on the implementation, the embedding function for a given feature type can be any of a variety of embedding functions” (Paragraph 43), “For example, for a feature type whose features consist of a single token, the embedding function may be a simple embedding function. A simple embedding function maps a single token to a floating point vector, i.e., a vector of floating point values. For example, the simple embedding function may map a token “cat” to a vector [0.1, 0.5, 0.2] and the word “iPod” to a vector [0.3, 0.9, 0.0], based on current parameter values, e.g., using a particular lookup table” (Paragraph 44), and “As another example, for a feature type whose features can potentially consist of a list of two or more tokens, the embedding function may be a parallel embedding function. A parallel embedding function maps each token in a list of tokens to a respective floating point vector and outputs a single vector that is the concatenation of the respective floating point vectors. For example, for an ordered list of tokens {“Atlanta”, “Hotel”}, the parallel embedding function may map “Atlanta” to a vector [0.1, 0.2, 0.3] and “Hotel” to [0.4, 0.5, 0.6], and then output [0.1, 0.2, 0.3, 0.4, 0.5, 0.6]. In order to identify the respective floating point vectors, the parallel embedding function may use a single lookup table or multiple different look up tables” (Paragraph 45).
The examiner further notes that Shaked teaches the use of stored multiple embedding functions (i.e. vectorizer functions) for subsequent determination of which function should be applied based off of attributes and rules. Moreover, generated embeddings are input into a deep learning (AI) model. The combination would result in the use of such functions to embed the features of Yang for subsequent input into an AI model.
It would have been obvious to one of ordinary skill in the art before the effective filing date of instant invention to combine the teachings of the cited references because teaching Shaked’s would have allowed Yang and McPherson’s to provide a method for analyzing unseen feature combinations, as noted by Shaked (Paragraph 13).
Regarding claim 14, Yang does not explicitly teach a computer program comprising:
A) wherein the feature metadata store stores the feature type of each feature as at least one of categorical, numerical, timedelta, Boolean, and date/time.
McPherson, however, teaches “wherein the feature metadata store stores the feature type of each feature as at least one of categorical, numerical, timedelta, Boolean, and date/time” as “In further embodiments, if a filter term 210 can be interpreted as one or more non-text data types, such as a number, date, time, datetime, or the like, then the data access application 110 may apply that filter term to both columns of text data type and columns of data types that can be compared to the interpreted data type. For example, the filter term 210B comprising a value of "3-22-07" may be applied as a text value to the COMPANY, PO#, and SALESPERSON columns (text data type), and as a date value to the PURCH_DT column (datetime data type), resulting in the inclusion of both rows 206A and 206D in the displayed result set 204” (Paragraph 25).
The examiner further notes that McPherson teaches the concept of data types (i.e. feature types) can include date/time. The combination would result in looking up the feature types (which would be date/time) for the features of Yang.
It would have been obvious to one of ordinary skill in the art before the effective filing date of instant invention to combine the teachings of the cited references because teaching McPherson’s would have allowed Yang to provide a method for accessing metadata to determine the structure of structured data, as noted by McPherson (Paragraph 18).
Yang and McPherson do not explicitly teach:
B) wherein the vectorizer store stores a plurality of vectorizer functions available for each feature type and a transformation condition for transforming a feature for each vectorizer function.
Shaked, however, teaches “wherein the vectorizer store stores a plurality of vectorizer functions available for each feature type and a transformation condition for transforming a feature for each vectorizer function” as “The system processes a first set of features from the obtained features using a deep machine learning model to generate a deep model intermediate predicted output (step 204). As described above, the deep machine learning model includes a deep neural network and an embedding layer that includes embedding functions. In some implementations, the system applies the embedding layer to a subset of the first set of features. In particular, the system uses each of the embedding functions for each of the feature type of the features in the subset to generate a numeric embedding, e.g., a floating-point vector representation, of the feature. Depending on the feature type and on the implementation, the embedding function for a given feature type can be any of a variety of embedding functions” (Paragraph 43), “For example, for a feature type whose features consist of a single token, the embedding function may be a simple embedding function. A simple embedding function maps a single token to a floating point vector, i.e., a vector of floating point values. For example, the simple embedding function may map a token “cat” to a vector [0.1, 0.5, 0.2] and the word “iPod” to a vector [0.3, 0.9, 0.0], based on current parameter values, e.g., using a particular lookup table” (Paragraph 44), and “As another example, for a feature type whose features can potentially consist of a list of two or more tokens, the embedding function may be a parallel embedding function. A parallel embedding function maps each token in a list of tokens to a respective floating point vector and outputs a single vector that is the concatenation of the respective floating point vectors. For example, for an ordered list of tokens {“Atlanta”, “Hotel”}, the parallel embedding function may map “Atlanta” to a vector [0.1, 0.2, 0.3] and “Hotel” to [0.4, 0.5, 0.6], and then output [0.1, 0.2, 0.3, 0.4, 0.5, 0.6]. In order to identify the respective floating point vectors, the parallel embedding function may use a single lookup table or multiple different look up tables” (Paragraph 45).
The examiner further notes that Shaked teaches the use of stored multiple embedding functions (i.e. vectorizer functions) for based on feature types for subsequent vectorization based off of rules (i.e. the undefined condition in the broadest reasonable interpretation). The combination would result in the use of such functions to embed the features of differing types of Yang.
It would have been obvious to one of ordinary skill in the art before the effective filing date of instant invention to combine the teachings of the cited references because teaching Shaked’s would have allowed Yang and McPherson’s to provide a method for analyzing unseen feature combinations, as noted by Shaked (Paragraph 13).
Regarding claim 15, Yang and McPherson do not explicitly teach a computer program comprising:
A) instructions for causing the at least one processor to execute: receiving feedback on prediction performance of the artificial intelligence model trained by using the input data; and
B) updating the vectorizer function decision rule to determine a set of vectorizer functions of features for optimizing the prediction performance; and
C) storing different types of artificial intelligence models generated with input data of various structures and generation information of each artificial intelligence model.
Shaked, however, teaches “instructions for causing the at least one processor to execute: receiving feedback on prediction performance of the artificial intelligence model trained by using the input data” as “The deep neural network 130 receives the numeric embeddings from the embedding layer and, optionally, other input features (e.g., feature 108) as an input” (Paragraph 36), “The system processes a first set of features from the obtained features using a deep machine learning model to generate a deep model intermediate predicted output (step 204). As described above, the deep machine learning model includes a deep neural network and an embedding layer that includes embedding functions. In some implementations, the system applies the embedding layer to a subset of the first set of features. In particular, the system uses each of the embedding functions for each of the feature type of the features in the subset to generate a numeric embedding, e.g., a floating-point vector representation, of the feature. Depending on the feature type and on the implementation, the embedding function for a given feature type can be any of a variety of embedding functions” (Paragraph 43), “The system then trains the combined model by, for each of the training inputs, processing the features of the training input using the deep machine learning model to generate a deep model intermediate predicted output for the training input in accordance with current values of parameters of the deep machine learning model (step 304)” (Paragraph 58), “The system then processes the deep model intermediate predicted output and the wide model intermediate predicted output for the training input using the combining layer to generate a predicted output for the training input (step 308)” (Paragraph 60), and “The system then determines an error between the predicted output for the training input and the known output for the training input” (Paragraph 61), “updating the vectorizer function decision rule to determine a set of vectorizer functions of features for optimizing the prediction performance” as “The deep neural network 130 receives the numeric embeddings from the embedding layer and, optionally, other input features (e.g., feature 108) as an input” (Paragraph 36), “The system processes a first set of features from the obtained features using a deep machine learning model to generate a deep model intermediate predicted output (step 204). As described above, the deep machine learning model includes a deep neural network and an embedding layer that includes embedding functions. In some implementations, the system applies the embedding layer to a subset of the first set of features. In particular, the system uses each of the embedding functions for each of the feature type of the features in the subset to generate a numeric embedding, e.g., a floating-point vector representation, of the feature. Depending on the feature type and on the implementation, the embedding function for a given feature type can be any of a variety of embedding functions” (Paragraph 43), “The system then trains the combined model by, for each of the training inputs, processing the features of the training input using the deep machine learning model to generate a deep model intermediate predicted output for the training input in accordance with current values of parameters of the deep machine learning model (step 304)” (Paragraph 58), “The system then processes the deep model intermediate predicted output and the wide model intermediate predicted output for the training input using the combining layer to generate a predicted output for the training input (step 308)” (Paragraph 60), and “The system then determines an error between the predicted output for the training input and the known output for the training input. In addition, the system backpropagates a gradient determined from the error through the combining layer to the wide machine learning model and the deep machine learning model to jointly adjust the current values of the parameters of the deep machine learning model and the wide machine learning model in a direction that reduces the error (step 310). Furthermore, through the method of backpropagation, the system can send an error signal to the deep learning model, which allows the deep learning model to adjust the parameters of its internal components, e.g., the deep neural network and the set of embedding functions, though successive stages of backpropagation. The system can also send an error signal to the wide learning model to allow the wide learning model to adjust the parameters of the generalized linear model” (Paragraph 61), and “storing different types of artificial intelligence models generated with input data of various structures and generation information of each artificial intelligence model” as “FIG. 1 is a block diagram of an example of a wide and deep machine learning model 102 that includes a deep machine learning model 104, a wide machine learning model 106, and a combining layer 134. The wide and deep machine learning model 102 receives a model input including multiple features, e.g. features 108-122, and processes the features to generate a predicted output, e.g., predicted output 136, for the model input” (Paragraph 19).
The examiner further notes that Shaked teaches the concept of updating stored multiple models (which include embedding functions) based off of “feedback” from predicted outputs. The combination would result in the use of such models to embed the features of differing types of Yang.
It would have been obvious to one of ordinary skill in the art before the effective filing date of instant invention to combine the teachings of the cited references because teaching Shaked’s would have allowed Yang and McPherson’s to provide a method for analyzing unseen feature combinations, as noted by Shaked (Paragraph 13).
Allowable Subject Matter
17. Claim 9 would be allowable if rewritten or amended to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action.
Specifically, although the prior art (See Yang) teaches the receipt of patient data, McPherson teaches the use of a metadata store, and Shaked teaches the use of different vector functions based on feature types, the detailed language directed towards the specifically defined queuing process is not found in the prior art, in conjunction with the rest of the limitations of the independent claim.
Dependent claims 10-12 are deemed allowable for depending on the deemed allowable subject matter of independent claim 9.
Conclusion
18. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
U.S. PGPUB 2021/0304895 issued to Buscemi et al. on 30 September 2021. The subject matter disclosed therein is pertinent to that of claims 9-17 (e.g., methods to store patient data).
U.S. PGPUB 2020/0194112 issued to Izadpanah et al. on 18 June 2010. The subject matter disclosed therein is pertinent to that of claims 9-17 (e.g., methods to store patient data).
Contact Information
19. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Mahesh Dwivedi whose telephone number is (571) 272-2731. The examiner can normally be reached on Monday to Friday 8:20 am – 4:40 pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Charles Rones can be reached (571) 272-4085. The fax number for the organization where this application or proceeding is assigned is (571) 273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see 20. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
Mahesh Dwivedi
Primary Examiner
Art Unit 2168
August 17, 2026
/MAHESH H DWIVEDI/Primary Examiner, Art Unit 2168