DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of the Claims
Claims 1-20 are pending for examination.
Claims 1 and 18 are independent Claims.
Claims 1-20 are rejected under 35 U.S.C. §§ 101, 103.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Independent Claims
As Claims 1:
Step 1: Are the Claims to a process, machine, manufacture or composition of matter? Yes.
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. See the analysis below.
The Claim recites:
A system comprising:
one or more computer processors;
a non-transitory computer-readable storage medium communicatively coupled to the one or more computer processors; and
a machine learning model for multi-label clinical document classification stored on the storage medium, the model comprising:
an input layer that automatically transforms one or more documents into a plurality of word embeddings;
a deep convolutional-based encoder that automatically combines information of adjacent words present in the documents and learns one or more representations of the words present in the documents;
an attention component that automatically selects one or more document features and generates label-specific representations for a plurality of identified labels; and
an output layer that produces zero or more final classification predictions.
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Regarding the non-emphasized limitations:
Step 2A prong 1:
“an input layer that automatically transforms one or more documents into a plurality of word embeddings;
a deep convolutional-based encoder that automatically combines information of adjacent words present in the documents and learns one or more representations of the words present in the documents;
an attention component that automatically selects one or more document features and generates label-specific representations for a plurality of identified labels; and ” is/are directed to a mental processes group of abstract idea. Mental processes are defined as concepts that can practically be performed in the human mind, or by a human using pen and paper as a physical aid. Examples of mental processes include observations, evaluations, judgements and opinions.
These steps are considered mental processes group of abstract idea.
Step 2A prong 2:
Limitations “a system comprising:
one or more computer processors;
a non-transitory computer-readable storage medium communicatively coupled to the one or more computer processors; and
a machine learning model for multi-label clinical document classification stored on the storage medium, the model comprising:
an output layer that produces zero or more final classification predictions.” are insignificant extra solution activity. See MPEP §2106.05(g).
The Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that integrate the Judicial Exception into a practical application? No.
Limitation “a system comprising:
one or more computer processors;
a non-transitory computer-readable storage medium communicatively coupled to the one or more computer processors; and
a machine learning model for multi-label clinical document classification stored on the storage medium, the model comprising:” was considered insignificant extra solution activity in Step 2A, and thus it is reevaluated in Step 2B to determine if it is more than what is well-understood, routine, conventional activity in the field. The addition of insignificant extra-solution activity does not amount to an inventive concept, particularly when the activity is well-understood or conventional (MPEP 2106.05(d)).
Limitation “an output layer that produces zero or more final classification predictions” in Step 2A, and thus it is reevaluated in Step 2B to determine if it is more than what is well-understood, routine, conventional activity in the field. The addition of insignificant extra-solution activity does not amount to an inventive concept, particularly when the activity is well-understood or conventional (MPEP 2106.05(d)). This appears to be well-understood, routine, conventional as evidenced by court case cited in 2106.05(d) as evidenced by HAWK TECHNOLOGY SYSTEMS, LLC, v. CASTLE RETAIL, LLC.
The claim is directed to mental processes group of abstract idea. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claim is not patent eligible.
As Claims 18:
Step 1: Are the Claims to a process, machine, manufacture or composition of matter? Yes.
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. See the analysis below.
The Claim recites:
A method for training a neural network comprising:
(a) receiving a plurality of documents;
(b) providing the received documents to a neural network model comprising a deep convolutional-based encoder including a plurality of squeeze-and-excitation (SE) and residual convolutional modules that form a plurality of SE/residual convolutional block pairs,
(c) determining a word embedding matrix for the plurality of documents;
(d) providing input including at least one or more word embeddings in the word embedding matrix to the encoder;
(e) generating one or more label-specific representations based on the output of the plurality of SE/residual convolutional block pairs;
(f) computing a probability of a label being present in the one or more documents given the one or more label specific representations; and
(g) using a first loss function to train the model for frequently occurring labels and a second loss function to train the model for rarely occurring labels, wherein the first loss function is used until the model performance saturates using the first loss function before training the model using the second loss function.
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Regarding the non-emphasized limitations:
Step 2A prong 1:
“(c) determining a word embedding matrix for the plurality of documents;
(e) generating one or more label-specific representations based on the output of the plurality of SE/residual convolutional block pairs; ” is/are directed to a mental processes group of abstract idea. Mental processes are defined as concepts that can practically be performed in the human mind, or by a human using pen and paper as a physical aid. Examples of mental processes includes observations, evaluations, judgements and opinions.
“(f) computing a probability of a label being present in the one or more documents given the one or more label specific representations; and ” is directed to a mathematical concepts group of abstract ideas. Mathematical concepts are defined as mathematical relationships, mathematical formulas or equations, or mathematical calculations.
These steps are considered mental processes group of abstract idea.
Step 2A prong 2:
Limitations “a method for training a neural network comprising:
(a) receiving a plurality of documents;
(b) providing the received documents to a neural network model comprising a deep convolutional-based encoder including a plurality of squeeze-and-excitation (SE) and residual convolutional modules that form a plurality of SE/residual convolutional block pairs,
(d) providing input including at least one or more word embeddings in the word embedding matrix to the encoder;” are insignificant extra solution activity. See MPEP §2106.05(g).
Limitations “(g) using a first loss function to train the model for frequently occurring labels and a second loss function to train the model for rarely occurring labels, wherein the first loss function is used until the model performance saturates using the first loss function before training the model using the second loss function” are mere instruction to apply an exception. See MPEP §2106.05(f).
The Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that integrate the Judicial Exception into a practical application? No.
Limitation “(a) receiving a plurality of documents;
(b) providing the received documents to a neural network model comprising a deep convolutional-based encoder including a plurality of squeeze-and-excitation (SE) and residual convolutional modules that form a plurality of SE/residual convolutional block pairs,
(d) providing input including at least one or more word embeddings in the word embedding matrix to the encoder;” step was considered to be extra-solution activity in Step 2A, and thus it is reevaluated in Step 2B to determine if it is more than what is well-understood, routine, conventional activity in the field. The addition of insignificant extra-solution activity does not amount to an inventive concept, particularly when the activity is well-understood or conventional (MPEP 2106.05(d)). This appears to be well-understood, routine, conventional as evidenced by MPEP 2106.05(d)(II)(i. Receiving or transmitting data over a network, e.g., using the Internet to gather data”).
The claim is directed to mental processes group of abstract idea. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claim is not patent eligible.
Dependent Claims
As Claim 2, the Claim recites “wherein the word embeddings are pretrained word embeddings.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the word embeddings are pretrained word embeddings.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claim 3, the Claim recites “wherein the pretrained word embeddings are context insensitive.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the pretrained word embeddings are context insensitive.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claim 4, the Claim recites “where the pretrained word embeddings are context sensitive.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “where the pretrained word embeddings are context sensitive.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claim 5, the Claim recites “wherein the pretrained word embeddings are determined using a skip-gram technique.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the pretrained word embeddings are determined using a skip-gram technique.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claim 6, the Claim recites “wherein the documents are unstructured documents.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the documents are unstructured documents.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claim 7, the Claim recites “wherein a final classification prediction includes one or more probabilities that one or more respective labels in the plurality of labels are present in the text of the one or more documents.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein a final classification prediction includes one or more probabilities that one or more respective labels in the plurality of labels are present in the text of the one or more documents.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claim 8, the Claim recites “wherein the deep convolutional-based encoder further comprises a plurality of squeeze-and-excitation (SE) and residual convolutional modules.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the deep convolutional-based encoder further comprises a plurality of squeeze-and-excitation (SE) and residual convolutional modules.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claim 9, the Claim recites “wherein the SE and residual convolutional blocks form an SE/residual convolutional block pair, and the SE and residual convolutional blocks in each pair are independent but receive the same data.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the SE and residual convolutional blocks form an SE/residual convolutional block pair, and the SE and residual convolutional blocks in each pair are independent but receive the same data.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claim 10, the Claim recites “wherein the SE convolutional module comprises an SE network followed by a layer normalization component.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the SE convolutional module comprises an SE network followed by a layer normalization component.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claim 11, the Claim recites “wherein the SE network comprises:
one or more one-dimensional convolutional layers;
a global average pooling component; a squeeze component; and an excitation component.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the SE network comprises:
one or more one-dimensional convolutional layers;
a global average pooling component; a squeeze component; and an excitation component.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claim 12, the Claim recites “wherein the deep convolutional-based encoder comprises a plurality of encoding blocks.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the deep convolutional-based encoder comprises a plurality of encoding blocks.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claim 13, the Claim recites “wherein the attention component is configured to extract all outputs from the plurality of encoding blocks.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the attention component is configured to extract all outputs from the plurality of encoding blocks.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claim 14, the Claim recites “wherein the attention component automatically selects the most important text features.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the attention component automatically selects the most important text features.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claim 15, the Claim recites “wherein the model identifies both frequently occurring and rarely occurring labels in the documents using a first determination for the frequently occurring labels and a second determination for the rarely occur labels.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the model identifies both frequently occurring and rarely occurring labels in the documents using a first determination for the frequently occurring labels and a second determination for the rarely occur labels.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claim 16, the Claim recites “wherein the first determination is a binary cross entropy loss determination and the second determination is a focal loss determination.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the first determination is a binary cross entropy loss determination and the second determination is a focal loss determination.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claim 17, the Claim recites “wherein the one or more documents is represented by a word embedding matrix that includes each of the transferred word embeddings.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the one or more documents is represented by a word embedding matrix that includes each of the transferred word embeddings.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claim 19, the Claim recites “wherein the first loss function is a binary cross entropy loss function and the second loss function is a focal loss function.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the first loss function is a binary cross entropy loss function and the second loss function is a focal loss function.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
As Claim 20, the Claim recites “wherein the encoder is a plurality of SE/residual module pairs and providing each word embedding further comprises providing each word embedding to at least one SE/residual module pair.”
The non-emphasized limitations describe abstract processes while emphasized limitations recited additional limitation(s).
Step 2A: Are the Claims directed to a law of nature, a natural phenomenon (product of nature) or an abstract idea? Yes, the Claims is an abstract idea. Prong 1: There are no additional abstract idea(s). Prong 2: The limitation “wherein the encoder is a plurality of SE/residual module pairs and providing each word embedding further comprises providing each word embedding to at least one SE/residual module pair.” are Mere instruction to apply an exception. See MPEP §2106.05(f). Claim(s) does not recite additional elements that integrate the judicial exception into a practical application.
Step 2B: Does the Claim recite additional elements that amount to significantly more than the Judicial Exception? No. The Claim is not patent eligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-4, 6-7 and 12-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Elkind et al. (U.S. 2019/0273509 hereinafter Elkind) in view of Brandes et al. (U.S. 2020/0364520 herein after Brandes).
As Claim 1, Elkind teaches a system comprising:
one or more computer processors (Elkind (¶0041 last 4 lines), processor 212);
a non-transitory computer-readable storage medium communicatively coupled to the one or more computer processors (Elkind (¶0041 last 4 lines), a removable storage); and
a machine learning model for multi-label (Bhaskar (¶0087 line 8, “protected health information”, ¶0097 line 9-13), “policy server 506 can maintain one set of trained models for organizations associated with the healthcare sector and a different set of trained models for organizations associated with the finance sector.”) stored on the storage medium (Elkind (¶0030 line 1-4), “The neural network classifying system may be deployed in various architectures. For example, the neural network classifying system can be deployed in a cloud based system that is accessed by other computing devices”), the model comprising:
an input layer that automatically transforms one or more documents into a plurality of word embeddings (Elkind (¶0138 line 2-9, fig. 10), “Unordered discrete inputs include tokenized text strings. Example tokenized string text include command line text and natural language text. Input in other examples can include ordered discrete inputs or other types of input. In FIG. 10, the tokenized inputs 1005 may be represented to computational layers as real-valued vectors or initial "embeddings" representing the input”);
a deep convolutional-based encoder that automatically combines information of adjacent words present in the documents and learns one or more representations of the words present in the documents (Elkind (¶0140 last 4 lines), “Convolutional layer 1020 extracts information from the embedded tokenized inputs (words present in the documents) based on its convolution function to create a convolved representation (one or more representations) of the embedded tokenized inputs 1005”);
an attention component (Elkind (¶0132 line 8-14), “Decoder RNN 730 can consist of any number of layers of one or more types of RNN cells, each layer consisting of an LSTM, GRU, or other RNN cell type. Additionally, each layer can have multiple RNN cells, the output of which is combined in some way before being sent to the next layer up if there is one”) that automatically selects one or more document features (Elkind (¶0143 middle portion), “the decoder RNN is trained such that the output matches approximate the output of any embedding layer in embedder 1010. In other examples, the embedding for tokens into a vector space is also learned.”) and generates label-specific representations for a plurality of identified labels (Elkind (¶0144 line1-6), “the decoder RNN is trained such that the output matches approximate the output of any embedding layer in embedder 1010. In other examples, the embedding for tokens into a vector space is also learned.”); and
an output layer that produces zero or more final classification predictions (Elkind (¶0144 line1-6), “the decoder RNN is trained such that the output matches approximate the output of any embedding layer in embedder 1010. In other examples, the embedding for tokens into a vector space is also learned.”).
Elkind may not explicitly disclose:
clinical document classification
Brandes teaches:
clinical document classification (Brandes (¶0055 last 3 lines), “health data”)
Elkind disclose a system and method to classify documents. Brandes teaches a system/method to include health document in the classification. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify document of Elkind instead be a document taught by Brandes, with a reasonable expectation of success. The motivation would be to allow “the proposed concept is suitable for any kind of classification ( sound, text, video, health data, stock market data, just to name a few application areas)” (Brandes (¶0055 last 4 lines)).
As Claim 2, besides Claim 1, Elkind in view of Brandes teaches wherein the word embeddings are pretrained word embeddings (Elkind (¶0138 last 3 lines), “These initial embeddings can be created using word2vec, one-hot encoding, feature hashing, or latent semantic analysis, or other techniques”).
As Claim 3, besides Claim 2, Elkind in view of Brandes teaches wherein the pretrained word embeddings are context insensitive (Bhaskar (¶0096 line -14), “may include a machine-learning (ML) engine configured to classify content as sensitive or not sensitive using one or more trained ML models”).
As Claim 4, besides Claim 2, Elkind in view of Brandes teaches where the pretrained word embeddings are context sensitive (Bhaskar (¶0096 line -14), “may include a machine-learning (ML) engine configured to classify content as sensitive or not sensitive using one or more trained ML models”).
As Claim 6, besides Claim 1, Elkind in view of Brandes teaches wherein the documents are unstructured documents (Elkind (¶0139 line 1-2), “depicts an example network that can be used for classifying unordered discrete inputs”).
As Claim 7, besides Claim 1, Elkind in view of Brandes teaches wherein a final classification prediction includes one or more probabilities that one or more respective labels in the plurality of labels are present in the text of the one or more documents (Elkind (¶0142 line 3-5), “if the output for "clean" is 0. 7 and the output for "not clean" is 0.3, the tokenized input can be classified as clean.”).
As Claim 12, besides Claim 1, Elkind in view of Brandes teaches wherein the deep convolutional-based encoder comprises a plurality of encoding blocks (Elkind (¶0111 line 4-6), “encoding recurrent neural network layers 940 receive samples as input, generate output, and update their state values”).
As Claim 13, besides Claim 12, Elkind in view of Brandes teaches wherein the attention component is configured to extract all outputs from the plurality of encoding blocks (Elkind (¶0111 last 4 lines), “a statistical analysis of one or more output values and/or state values generated by encoding RNN layer 940 is used as the output of the encoding portion of the system.” Output of the encoder is input of the attention component).
As Claim 14, besides Claim 1, Elkind in view of Brandes teaches wherein the attention component automatically selects the most important text features (Elkind (¶0132, middle), “RNN 730 can consist of any number of layers of one or more types of RNN cells, each layer consisting of an LSTM, GRU, or other RNN cell type. Additionally, each layer can have multiple RNN cells, the output (best features) of which is combined in some way before being sent to the next layer up if there is one”).
As Claim 15, besides Claim 1, Elkind in view of Brandes teaches wherein the model identifies both frequently occurring and rarely occurring labels in the documents using a first determination for the frequently occurring labels and a second determination for the rarely occur labels (Brandes (¶0057 line 1-4), label is associate the classification as rare or frequent case).
As Claim 16, besides Claim 15, Elkind in view of Brandes teaches wherein the first determination is a binary cross entropy loss determination (Brandes (¶0057 line 1-4), “If the case is not a rare case, i.e., the confidence value is good enough (above a predefined confidence level threshold value or a maximum number of iterations has been reached) (cross entropy loss determination)”) and the second determination is a focal loss determination (Brandes (¶0066 line 1-5), “evaluates the confidence level of all five additional images. At least one is not classified as a rare case and sent as output to the right, i.e., to the valid output box 306 because of its relative confidence level which is above a predefined threshold value”).
As Claim 17, besides Claim 14, Elkind in view of Brandes teaches
wherein the one or more documents is represented by a word embedding matrix (Elkind (¶0073), “convolution filter F may be applied to the voxels of a color photograph. In this case, the voxels are represented by a matrix for each component of a photograph's color model (RGB, HSY, HSL, CIE XYZ, CIELUV, L*a*b*, YIQ, for example), so a convolutional filter may be a 2-tensor sequentially applied over the color space or a 3-tensor, depending on the application”) that includes each of the transferred word embeddings (Elkind (¶0027 line 3-end), “The disclosed systems and methods provide various examples of embeddings, including multiple embeddings. For example, "initial embeddings" describing relationships between source data elements may be initially generated. These "initial embeddings" may be further analyzed to create additional sets of embeddings describing relation ships between source data elements. It is understood that the disclosed systems and methods can be applied to generate an arbitrary number of embeddings describing the relationships between the source data elements. These relationships may be used to classify the source data according to a criterion, such as malicious or not.”).
Claim(s) 8-11 and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Elkind in view of Brandes in further view of Roy et al (“FuSENet: fused squeeze-and-excitation network for spectral-spatial hyperspectral image classification” hereinafter Roy).
As Claim 8, besides Claim 14, Elkind in view of Brandes does not explicitly disclose:
wherein the deep convolutional-based encoder further comprises a plurality of squeeze-and-excitation (SE) and residual convolutional modules.
Roy teaches:
wherein the deep convolutional-based encoder further comprises a plurality of squeeze-and-excitation (SE) and residual convolutional modules (Roy (pg. 1654, last 2 line of left column), “Every residual block is followed by the proposed fused squeeze-and-excitation”).
Elkind discloses a system and method to classify documents using machine learning. Roy teaches a machine learning system/method using squeeze and excitation with residual convolution module. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify machine learning system of Elkind in view of Brandes instead be a machine learning system taught by Roy, with a reasonable expectation of success. The motivation would be to allow “superiority of the proposed FuSENet method with respect to the state-of-the-art methods” (Roy (abstract)).
As Claim 9, besides Claim 8, Elkind in view of Brandes in further view of Roy teaches wherein the SE and residual convolutional blocks form an SE/residual convolutional block pair, and the SE and residual convolutional blocks in each pair are independent (Roy (pg. 1654, last 2 line of left column), “Every residual block is followed by the proposed fused squeeze-and-excitation”) but receive the same data (Roy (pg. 1654, last 2 line of left column), “Every residual block is followed by the proposed fused squeeze-and-excitation (FuSE) block (as depicted by residual FuSE block in Fig. 1c). The output of the FuSE block is used to re-calibrate the input channels of that block.” Since the blocks are in series, only one input data is required).
As Claim 10, besides Claim 8, Elkind in view of Brandes in further view of Roy teaches wherein the SE convolutional module comprises an SE network followed by a layer normalization component (Roy (pg. 1654, bottom of left column, fig. 1a), “The batch normalisation (BN) [37] is used followed by a 3D convolutional layer within the residual blocks. The conventional residual blocks shown in Fig. 1b and can be formulated as …”, fig. 1a shows plurality of operation in series. After SE block is another residual bock which includes a batch nomalisation).
As Claim 11, besides Claim 10, Elkind in view of Brandes in further view of Roy teaches
wherein the SE network comprises:
one or more one-dimensional convolutional layers (Roy (pg. 1654, bottom of left column, fig. 1a), “The batch normalisation (BN) [37] is used followed by a 3D convolutional layer within the residual blocks. The conventional residual blocks shown in Fig. 1b and can be formulated as …”);
a global average pooling component; a squeeze component; and an excitation component (Roy (pg. 1654, first paragraph of right column), “the squeeze operation extracts the channel-wise information. Moreover, the global pooling will retail the information in global context, whereas the max pooling will retain the information in local context. The excitation networks are used to prioritise the features extracted by the squeeze operation”).
As Claim 18, Elkind teaches a method for training a neural network comprising:
(a) receiving a plurality of documents (Elkind (¶0021, bottom), “source data may comprise executables or executable files, including files in the aforementioned executable code format. Other source data may include command line data, a registry key, a registry key value, a file name, a domain name, a Uniform Resource Identifier, script code, a word processing file, a portable document format file, or a spreadsheet. Furthermore, source data can include document files, PDF files, image files, images, or other non-executable formats.”);
(c) determining a word embedding matrix for the plurality of documents (Elkind (¶0073), “convolution filter F may be applied to the voxels of a color photograph. In this case, the voxels are represented by a matrix for each component of a photograph's color model (RGB, HSY, HSL, CIE XYZ, CIELUV, L*a*b*, YIQ, for example), so a convolutional filter may be a 2-tensor sequentially applied over the color space or a 3-tensor, depending on the application”);
(f) computing a probability of a label being present in the one or more documents given the one or more label specific representations (Elkind (¶0142 line 3-5), “if the output for "clean" is 0. 7 and the output for "not clean" is 0.3, the tokenized input can be classified as clean.”);
Elkind may not explicitly disclose:
(g) using a first loss function to train the model for frequently occurring labels and a second loss function to train the model for rarely occurring labels, wherein the first loss function is used until the model performance saturates using the first loss function before training the model using the second loss function.
Brandes teaches:
(g) using a first loss function to train the model for frequently occurring labels and a second loss function to train the model for rarely occurring labels (Brandes (¶0057 line 1-4), label is associate the classification as rare or frequent case), wherein the first loss function is used until the model performance saturates using the first loss function (Brandes (¶0057 line 1-4), “If the case is not a rare case, i.e., the confidence value is good enough (above a predefined confidence level threshold value or a maximum number of iterations has been reached)”) before training the model (Brandes (¶0058 line 1-5), “If it is determined by the evaluator engine 306 that the case is a rare case, the input data are forwarded to the rare case extractor 310. This module is used to potentially enlarge the corpus of training data with potentially related or similar images”) using the second loss function (Brandes (¶0066 line 1-5), “evaluates the confidence level of all five additional images. At least one is not classified as a rare case and sent as output to the right, i.e., to the valid output box 306 because of its relative confidence level which is above a predefined threshold value”).
Elkind disclose a system and method to classify documents. Brandes teaches a system/method to include health document in the classification. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify document of Elkind instead be a document taught by Brandes, with a reasonable expectation of success. The motivation would be to allow “the proposed concept is suitable for any kind of classification ( sound, text, video, health data, stock market data, just to name a few application areas)” (Brandes (¶0055 last 4 lines)).
Elkind in view of Brandes does not explicitly disclose:
(b) providing the received documents to a neural network model comprising a deep convolutional-based encoder including a plurality of squeeze-and-excitation (SE) and residual convolutional modules that form a plurality of SE/residual convolutional block pairs,
(e) generating one or more label-specific representations based on the output of the plurality of SE/residual convolutional block pairs;
Roy teaches:
(b) providing the received documents to a neural network model (Roy (pg. 1655, left column line 1-30, “As shown in Fig. 1c, the final scaling factor is used to scale the FuSE block input that can re-formulate the residual block”) comprising a deep convolutional-based encoder including a plurality of squeeze-and-excitation (SE) and residual convolutional modules that form a plurality of SE/residual convolutional block pairs (Roy (pg. 1654, last 2 line of left column), “Every residual block is followed by the proposed fused squeeze-and-excitation”),
(e) generating one or more label-specific representations based on the output of the plurality of SE/residual convolutional block pairs (Roy (pg. 1656, second to last paragraph of left column), “To explore the robustness of the proposed FuSENet, Table 4 shows the classification performance of in terms of OA, FuSENet Kappa, and AA using varying training samples 10 and 20% over IP, UP, and SA datasets, respectively”);
Elkind discloses a system and method to classify documents using machine learning. Roy teaches a machine learning system/method using squeeze and excitation with residual convolution module. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify machine learning system of Elkind in view of Brandes instead be a machine learning system taught by Roy, with a reasonable expectation of success. The motivation would be to allow “superiority of the proposed FuSENet method with respect to the state-of-the-art methods” (Roy (abstract)).
As Claim 19, besides Claim 18, Elkind in view of Brandes in further view of Roy teaches wherein the first loss function is a binary cross entropy loss function (Brandes (¶0057 line 1-4), “If the case is not a rare case, i.e., the confidence value is good enough (above a predefined confidence level threshold value or a maximum number of iterations has been reached) (cross entropy loss determination)”) and the second loss function is a focal loss function (Brandes (¶0066 line 1-5), “evaluates the confidence level of all five additional images. At least one is not classified as a rare case and sent as output to the right, i.e., to the valid output box 306 because of its relative confidence level which is above a predefined threshold value”).
As Claim 20, besides Claim 18, Elkind in view of Brandes in further view of Roy teaches wherein the encoder is a plurality of SE/residual module pairs (Roy (pg. 1654, last 2 line of left column), “Every residual block is followed by the proposed fused squeeze-and-excitation”) and providing each word embedding further comprises providing each word embedding to at least one SE/residual module pair (Roy (pg. 1655, lines 1-3 of left column), “As shown in Fig. 1c, the final scaling factor is used to scale the FuSE block input that can re-formulate the residual block as”).
Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Elkind in view of Brandes in further view of Iso-Sipila et al. (U.S. 2022/0188520 hereinafter Iso-Sipila).
As Claim 5, besides Claim 2, Elkind in view of Brandes does not explicitly disclose:
wherein the pretrained word embeddings are determined using a skip-gram technique.
Iso-Sipila teaches:
wherein the pretrained word embeddings are determined using a skip-gram technique (Iso-Sipila (¶0139), skip-gram technique).
Elkind discloses a system and method to classify documents using machine learning. Iso-Sipila teaches a machine learning system/method using squeeze and excitation with residual convolution module. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify machine learning system of Elkind in view of Brandes instead be a machine learning system taught by Iso-Sipila, with a reasonable expectation of success. The motivation would be to allow “to reliably and efficiently generate and/or build entity dictionaries and/or augment existing entity dictionaries of an NER dictionary based systems based on a corpus of text” (Iso-Sipilar (¶0007 last 4 lines)).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Bhaskar et al. (U.S. 2021/0194888) discloses a system/method to restrict access to sensitive classified content.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NHAT HUY T NGUYEN whose telephone number is (571)270-7333. The examiner can normally be reached M-F: 12:00-8:00 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at 571-270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NHAT HUY T NGUYEN/ Primary Examiner, Art Unit 2147