DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 1 is/are rejected under 35 U.S.C. 103 as being unpatentable over Harrington et al., (US 2025/0021870 A1, hereinafter Harrington) in view of Wang et al., (US 20220414737 A1, hereinafter Wang).
Regarding claim 1:
Harrington shows:
“A method for modeling variable-distanced input dependencies,” (Paragraph [0076]: “the transformer 700 uses self-attention mechanisms to selectively weigh the importance of different parts of an input sequence during processing and allows the model to attend to different parts of the input sequence while generating the output. The input sequence is first embedded into vectors and then passed through multiple layers of self-attention and feed-forward networks. The transformer 700 may process input sequences of variable length, making it well-suited for natural language processing tasks where input lengths may vary greatly. Additionally, the self-attention mechanism allows the transformer 700 to capture long-range dependencies between words in the input sequence, which is difficult for recurrent neural networks (RNNs) and CNNs. The transformer with self-attention has achieved results in several natural language processing tasks that are beyond the capabilities of other neural networks and has become a popular choice for language and text applications. For example, the various large language models, such as a generative pretrained transformer (e.g., ChatGPT, etc.) and other current models are types of transformer networks.”)
“comprising: providing non-linear readouts using attentional neural networks to replace the linear readouts;” (Paragraph [0080]: “The neural network encoder 850 is a trained neural network that has learned mappings between a high-dimensional input into a lower-dimensional space. The neural network encoder 850 includes an encoder and may include a decoder. Each encoder of the neural network encoder 850 includes several layers of artificial neurons that perform a non-linear transformation on the input data and reduce high-dimensional data to lower data by learning based on various techniques, such as backpropagation. The neural network encoder may be trained using various optimization techniques to minimize a loss function that measures the difference between the original high-dimensional data and the reconstructed data. The neural network encoder 850 provides flexibility and ability to learn complex and non-linear mappings between the input data and the encoding result but requires large amounts of training data, computational resources, and careful tuning of the network architecture and hyperparameters.”)
“learning, via the non-linear readout reservoir, sample dependencies in the complete dataset, wherein, the learning complements the transformer that only handles the dependencies within a sample in a short context;” (Paragraph [0076]: “the transformer 700 uses self-attention mechanisms to selectively weigh the importance of different parts of an input sequence during processing and allows the model to attend to different parts of the input sequence while generating the output. The input sequence is first embedded into vectors and then passed through multiple layers of self-attention and feed-forward networks. The transformer 700 may process input sequences of variable length, making it well-suited for natural language processing tasks where input lengths may vary greatly. Additionally, the self-attention mechanism allows the transformer 700 to capture long-range dependencies between words in the input sequence, which is difficult for recurrent neural networks (RNNs) and CNNs. The transformer with self-attention has achieved results in several natural language processing tasks that are beyond the capabilities of other neural networks and has become a popular choice for language and text applications. For example, the various large language models, such as a generative pretrained transformer (e.g., ChatGPT, etc.) and other current models are types of transformer networks.” And in paragraph [0078]: “The matrix encoder 810 identifies the most important features of data (e.g., most important embeddings) and reduces the features into a lower dimensional representation. Non-limiting examples of techniques incorporated into the matrix encoder 810 include singular value decomposition (SVD), principal component analysis (PCA), or autoencoders to perform the transformation. For example, the matrix encoder 810 converts a matrix 812 into a column 814 of components and a row 816 of features associated with the components. The lower-dimension representation of the matrix 812 may be used to assist in clustering, classification, and visualization, as well as improve the efficiency of computations.” And in paragraph [0080]: “The neural network encoder 850 is a trained neural network that has learned mappings between a high-dimensional input into a lower-dimensional space. The neural network encoder 850 includes an encoder and may include a decoder. Each encoder of the neural network encoder 850 includes several layers of artificial neurons that perform a non-linear transformation on the input data and reduce high-dimensional data to lower data by learning based on various techniques, such as backpropagation. The neural network encoder may be trained using various optimization techniques to minimize a loss function that measures the difference between the original high-dimensional data and the reconstructed data. The neural network encoder 850 provides flexibility and ability to learn complex and non-linear mappings between the input data and the encoding result but requires large amounts of training data, computational resources, and careful tuning of the network architecture and hyperparameters.”)
“and where the learning long-sequential inputs improves ... performance and significantly increases prediction accuracy in language modeling, text classification, and dialogue modelling tasks over the state-of-the-art.” (Paragraph [0076]: “the transformer 700 uses self-attention mechanisms to selectively weigh the importance of different parts of an input sequence during processing and allows the model to attend to different parts of the input sequence while generating the output. The input sequence is first embedded into vectors and then passed through multiple layers of self-attention and feed-forward networks. The transformer 700 may process input sequences of variable length, making it well-suited for natural language processing tasks where input lengths may vary greatly. Additionally, the self-attention mechanism allows the transformer 700 to capture long-range dependencies between words in the input sequence, which is difficult for recurrent neural networks (RNNs) and CNNs. The transformer with self-attention has achieved results in several natural language processing tasks that are beyond the capabilities of other neural networks and has become a popular choice for language and text applications. For example, the various large language models, such as a generative pretrained transformer (e.g., ChatGPT, etc.) and other current models are types of transformer networks.” And in paragraph [0078]: “The matrix encoder 810 identifies the most important features of data (e.g., most important embeddings) and reduces the features into a lower dimensional representation. Non-limiting examples of techniques incorporated into the matrix encoder 810 include singular value decomposition (SVD), principal component analysis (PCA), or autoencoders to perform the transformation. For example, the matrix encoder 810 converts a matrix 812 into a column 814 of components and a row 816 of features associated with the components. The lower-dimension representation of the matrix 812 may be used to assist in clustering, classification, and visualization, as well as improve the efficiency of computations.” And in paragraph [0080]: “The neural network encoder 850 is a trained neural network that has learned mappings between a high-dimensional input into a lower-dimensional space. The neural network encoder 850 includes an encoder and may include a decoder. Each encoder of the neural network encoder 850 includes several layers of artificial neurons that perform a non-linear transformation on the input data and reduce high-dimensional data to lower data by learning based on various techniques, such as backpropagation. The neural network encoder may be trained using various optimization techniques to minimize a loss function that measures the difference between the original high-dimensional data and the reconstructed data. The neural network encoder 850 provides flexibility and ability to learn complex and non-linear mappings between the input data and the encoding result but requires large amounts of training data, computational resources, and careful tuning of the network architecture and hyperparameters.”)
But Harrington does not appear to explicitly recite the use of “BERT and Blenderbot” models.
However, Wang teaches the use of “BERT and Blenderbot” models. (Paragraph [0037]: “The result of the training can be a generic query representation generator 250 capable of producing a query representation for any query input. The generic query representation generator 250 can include a dimensionality including dimensions for inputs of grouped query data and dimensions for output query representations. In implementations, outputs of the generic query representation generator 250 can be of a predefined dimension, perhaps a vector of a predefined dimensionality (e.g., conforming to a predefined length and/or conforming values to a predefined range). The product representations can be of a same dimensionality as the query product representation, making the representations input-string-length-invariant output vectors.” And in paragraph [0038]: “In implementations, the generic query representation generator 250 can be an inference model and/or a machine learning model. In this specification, examples of machine learning or inference models can include, without limitation, … bidirectional transformation, unidirerctional transformation, gradient descent, autoregression, autoencoding, permutation language modeling, two-stream self attenuation, federated learning, absorbing transformer-XL, natural language processing (NLP), bidirectional encoder representations from transformers (BERT) models and variants (e.g., RoBERTa, XLM-RoBERTa, and DistilBERT, ALBERT, CamemBERT, ConvBERT, DeBERTA, DeBERTA-v2, FlauBERT, I-BERT, herBERT, BertGeneration, BertJapanese, Bertweet, MegatronBERT, PhoBERT, MobileBERT, SqueezeBERT, … Blenderbot, Blenderbot Small”)
Harrington and Wang are analogous in the arts because both Harrington and Wang describe the use of input data for artificial intelligence models.
Therefore, it would be obvious to one of ordinary skill in the art at the filing date of the instant application, having the teachings of Harrington and Wang before him or her, to modify the teachings of Harrington to include the teachings of Wang in order to support additional models like BERT and Blenderbot (see Wang paragraphs [0038] and [0087]) to increase marketability and accuracy of Harrington.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Hajarnis et al., (US 2022/0051080 A1), part of the prior art made of record, teaches the use of a BERT model, transformer, and short context of claim 1 in paragraph [0024] through the use of context with transformer layers with the use of input data from BERT neural networks.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHANE D WOOLWINE whose telephone number is (571)272-4138. The examiner can normally be reached M-F 9:30-6:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MIRANDA HUANG can be reached at (571) 270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
SHANE D. WOOLWINE
Primary Examiner
Art Unit 2124
/SHANE D WOOLWINE/Primary Examiner, Art Unit 2124