CTNF 18/990,968 CTNF 96826 DETAILED ACTION Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Specification 07-29 AIA The disclosure is objected to because of the following informalities: In paragraph 0006, lines 6-7, “can configured” should read “can be configured”. In paragraph 0040, line 3, “cam be determined” should read “can be determined”. In paragraph 0089, line 3, “first textual block representations 405B” should read “first textual block representations 405A”. Figure 4B element 452 is cites as “textual blocks 452” in paragraph 0091, line 5, and paragraph 0092, lines 3, 5, and 8, and as “token embeddings 452” in paragraph 0091, line 8, and paragraph 0092, lines 1-2. In paragraph 0097, line 5, “representations520B-520D” should read “representations 520B-520D” . Appropriate correction is required. Claim Objections 07-29-01 AIA Claim 36 is objected to because of the following informalities: In line 1, “The computing system of claim 21” should read “The computing system of claim 29” . Appropriate correction is required. Claim Rejections - 35 USC § 112 07-30-02 AIA The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. 07-34-01 Claims 21 – 40 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 21 recites the limitations "the first document encoding" in lines 18-19 and "the second document encoding" in line 19. There is insufficient antecedent basis for these limitations in the claim. This rejection can be overcome by changing "the first document encoding" to "the first document representation" and changing "the second document encoding" to "the second document representation". Claims 22 – 24 are also rejected as they depend from claim 21 and thus recite the limitations of claim 21, and do not resolve the indefinite language from claim 21. Claim 25 is also rejected as it depends from claim 21 and thus recite the limitations of claim 21, and does not resolve the indefinite language from claim 21. Claim 25 also recites the limitations "the contextual block representation" in lines 8-9 and "the first document encoding" in line 10. There is insufficient antecedent basis for these limitations in the claim. This rejection can be overcome by changing "the contextual block representation" to "a contextual block representation" and changing "the first document encoding" to "the first document representation". Claims 26 – 27 are also rejected as they depend from claim 21 and thus recite the limitations of claim 21, and do not resolve the indefinite language from claim 21. Claim 28 is also rejected as it depends from claim 21 and thus recite the limitations of claim 21, and does not resolve the indefinite language from claim 21. Claim 28 also recites the limitation "the first document encoding" in line 2. There is insufficient antecedent basis for these limitations in the claim. This rejection can be overcome by changing "the first document encoding" to "the first document representation". Claim 29 recites the limitations "the first document encoding" in line 21 and "the second document encoding" in lines 21-22. There is insufficient antecedent basis for these limitations in the claim. This rejection can be overcome by changing "the first document encoding" to "the first document representation" and changing "the second document encoding" to "the second document representation". Claims 30 – 32 are also rejected as they depend from claim 29 and thus recite the limitations of claim 29, and do not resolve the indefinite language from claim 29. Claim 33 is also rejected as it depends from claim 29 and thus recite the limitations of claim 29, and does not resolve the indefinite language from claim 29. Claim 33 also recites the limitations "the contextual block representation" in line 7 and "the first document encoding" in lines 8-9. There is insufficient antecedent basis for these limitations in the claim. This rejection can be overcome by changing "the contextual block representation" to "a contextual block representation" and changing "the first document encoding" to "the first document representation". Claims 34 – 35 are also rejected as they depend from claim 29 and thus recite the limitations of claim 29, and do not resolve the indefinite language from claim 29. Claim 36 is also rejected as it depends from claim 29 and thus recite the limitations of claim 29, and does not resolve the indefinite language from claim 29. Claim 36 also recites the limitation "the first document encoding" in line 2. There is insufficient antecedent basis for this limitation in the claim. This rejection can be overcome by changing "the first document encoding" to "the first document representation". Claim 37 recites the limitations "the first document encoding" in line 19 and "the second document encoding" in lines 19-20. There is insufficient antecedent basis for these limitations in the claim. This rejection can be overcome by changing "the first document encoding" to "the first document representation" and changing "the second document encoding" to "the second document representation". Claims 38 – 40 are also rejected as they depend from claim 37 and thus recite the limitations of claim 37, and do not resolve the indefinite language from claim 37. Claim Rejections - 35 USC § 103 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-21-aia AIA Claim s 21 – 27, 29 – 35 and 37 – 40 are rejected under 35 U.S.C. 103 as being unpatentable over Jiang et al. ("Semantic Text Matching for Long-Form Documents"), hereinafter Jiang, in view of Zhang et al. ("HIBERT: Document Level Pre-training of Hierarchical Bidirectional Transformers for Document Summarization"), hereinafter Zhang . Regarding claim 21, as best understood based on the 35 U.S.C. 112(b) issues identified above, Jiang discloses a computer-implemented method, comprising: obtaining, by a computing system comprising one or more computing devices, a first document comprising a plurality of first textual blocks and a second document comprising a plurality of second textual blocks (Section 3.1, lines 20-24, "Given a source document d s and a set of candidate documents D c , our goal is to estimate semantic similarity ŷ = Sim(d s , d c ) between the source document d s and every candidate document d c ∈ D c so that the target documents semantically matched to the source document have higher semantic similarity scores."; Section 3.1, lines 4-6, "To facilitate readability, we assume that there are three levels in hierarchy – paragraphs, sentences and words."; A source document containing paragraphs and sentences reads on a first document comprising a plurality of first textual blocks, and a candidate document containing paragraphs and sentences reads on a second document comprising a plurality of second textual blocks.); processing, by the computing system, the plurality of first textual blocks with [a first hierarchical transformer model of] a machine-learned semantic document encoding model to obtain a respective plurality of first textual block representations (Section 3.2, lines 7-12, "For each level, an attention-based hierarchical RNN (with corresponding level depth) is constructed as an encoder to generate representations for that level. For example, the paragraph-level encoder produces paragraph-level representations with a depth-3 encoder while the sentence-level encoder produces sentence-level representations with a depth-2 encoder."; The paragraph-level representations and sentence-level representations read on the textual block representations.); processing, by the computing system, the plurality of second textual blocks with [the first hierarchical transformer model of] the machine-learned semantic document encoding model to obtain a respective plurality of second textual block representations (Section 3.2, lines 7-12, "For each level, an attention-based hierarchical RNN (with corresponding level depth) is constructed as an encoder to generate representations for that level. For example, the paragraph-level encoder produces paragraph-level representations with a depth-3 encoder while the sentence-level encoder produces sentence-level representations with a depth-2 encoder."; The paragraph-level representations and sentence-level representations read on the textual block representations.); processing, by the computing system, the plurality of first textual block representations with [a second hierarchical transformer model of] the machine-learned semantic document encoding model to obtain a first document representation (Section 3.2, lines 4-15, "For each document, MASH RNN derives an informative representation based on the knowledge from different levels of document structure. For each level, an attention-based hierarchical RNN (with corresponding level depth) is constructed as an encoder to generate representations for that level. For example, the paragraph-level encoder produces paragraph-level representations with a depth-3 encoder while the sentence-level encoder produces sentence-level representations with a depth-2 encoder. The final document representation is then acquired by concatenating the representations of different levels, comprehensively covering the knowledge in all document structure levels."; Concatenating the paragraph-level representations and sentence-level representations to acquire the document representation reads on processing the textual block representations to obtain a document representation.); processing, by the computing system, the plurality of second textual block representations with [the second hierarchical transformer model of] the machine-learned semantic document encoding model to obtain a second document representation (Section 3.2, lines 4-15, "For each document, MASH RNN derives an informative representation based on the knowledge from different levels of document structure. For each level, an attention-based hierarchical RNN (with corresponding level depth) is constructed as an encoder to generate representations for that level. For example, the paragraph-level encoder produces paragraph-level representations with a depth-3 encoder while the sentence-level encoder produces sentence-level representations with a depth-2 encoder. The final document representation is then acquired by concatenating the representations of different levels, comprehensively covering the knowledge in all document structure levels."; Concatenating the paragraph-level representations and sentence-level representations to acquire the document representation reads on processing the textual block representations to obtain a document representation.); and determining, by the computing system, a similarity metric descriptive of a semantic similarity between the first document and the second document based on the first document encoding and the second document encoding (Section 3.2, lines 15-21, "To estimate semantic similarity for semantic text matching, SMASH RNN adopts the Siamese structure with two MASH RNN towers. Given representations generated by MASH RNN for both the source and target documents, a fully-connected layer with nonlinearity infers a probabilistic score to examine the semantic relation between two documents with a sigmoid function."; Estimating semantic similarity for semantic text matching between the source and target documents reads on determining a similarity metric descriptive of a semantic similarity between the first document and the second document.). Jiang does not specifically disclose: processing, by the computing system, the plurality of first textual blocks with a first hierarchical transformer model of a machine-learned semantic document encoding model to obtain a respective plurality of first textual block representations; processing, by the computing system, the plurality of second textual blocks with the first hierarchical transformer model of the machine-learned semantic document encoding model to obtain a respective plurality of second textual block representations; processing, by the computing system, the plurality of first textual block representations with a second hierarchical transformer model of the machine-learned semantic document encoding model to obtain a first document representation; processing, by the computing system, the plurality of second textual block representations with the second hierarchical transformer model of the machine-learned semantic document encoding model to obtain a second document representation. Zhang teaches: processing, by the computing system, the plurality of first textual blocks with a first hierarchical transformer model of a machine-learned semantic document encoding model to obtain a respective plurality of first textual block representations (Section 1, lines 77-84, "In this paper, we propose HIBERT, which stands for HIerachical Bidirectional Encoder Representations from Transformers. We design an unsupervised method to pre-train HIBERT for document modeling. We apply the pre-trained HIBERT to the task of document summarization and achieve state-of-the-art performance on both the CNN/Dailymail and New York Times dataset."; Section 3.1, lines 6-14, “To obtain the representation of D, we use two encoders: a sentence encoder to transform each sentence in D to a vector and a document encoder to learn sentence representations given their surrounding sentences as context. Both the sentence encoder and document encoder are based on the Transformer encoder described in Vaswani et al. (2017). As shown in Figure 1, they are nested in a hierarchical fashion.”; A sentence encoder reads on a first hierarchical transformer model of a machine-learned semantic document encoding model.); processing, by the computing system, the plurality of second textual blocks with the first hierarchical transformer model of the machine-learned semantic document encoding model to obtain a respective plurality of second textual block representations (Section 1, lines 77-84, "In this paper, we propose HIBERT, which stands for HIerachical Bidirectional Encoder Representations from Transformers. We design an unsupervised method to pre-train HIBERT for document modeling. We apply the pre-trained HIBERT to the task of document summarization and achieve state-of-the-art performance on both the CNN/Dailymail and New York Times dataset."; Section 3.1, lines 6-14, “To obtain the representation of D, we use two encoders: a sentence encoder to transform each sentence in D to a vector and a document encoder to learn sentence representations given their surrounding sentences as context. Both the sentence encoder and document encoder are based on the Transformer encoder described in Vaswani et al. (2017). As shown in Figure 1, they are nested in a hierarchical fashion.”; A sentence encoder reads on a first hierarchical transformer model of a machine-learned semantic document encoding model.); processing, by the computing system, the plurality of first textual block representations with a second hierarchical transformer model of the machine-learned semantic document encoding model to obtain a first document representation (Section 1, lines 77-84, "In this paper, we propose HIBERT, which stands for HIerachical Bidirectional Encoder Representations from Transformers. We design an unsupervised method to pre-train HIBERT for document modeling. We apply the pre-trained HIBERT to the task of document summarization and achieve state-of-the-art performance on both the CNN/Dailymail and New York Times dataset."; Section 3.1, lines 6-14, “To obtain the representation of D, we use two encoders: a sentence encoder to transform each sentence in D to a vector and a document encoder to learn sentence representations given their surrounding sentences as context. Both the sentence encoder and document encoder are based on the Transformer encoder described in Vaswani et al. (2017). As shown in Figure 1, they are nested in a hierarchical fashion.”; A document encoder reads on a second hierarchical transformer model of a machine-learned semantic document encoding model.); processing, by the computing system, the plurality of second textual block representations with the second hierarchical transformer model of the machine-learned semantic document encoding model to obtain a second document representation (Section 1, lines 77-84, "In this paper, we propose HIBERT, which stands for HIerachical Bidirectional Encoder Representations from Transformers. We design an unsupervised method to pre-train HIBERT for document modeling. We apply the pre-trained HIBERT to the task of document summarization and achieve state-of-the-art performance on both the CNN/Dailymail and New York Times dataset."; Section 3.1, lines 6-14, “To obtain the representation of D, we use two encoders: a sentence encoder to transform each sentence in D to a vector and a document encoder to learn sentence representations given their surrounding sentences as context. Both the sentence encoder and document encoder are based on the Transformer encoder described in Vaswani et al. (2017). As shown in Figure 1, they are nested in a hierarchical fashion.”; A document encoder reads on a second hierarchical transformer model of a machine-learned semantic document encoding model.). Zhang is considered to be analogous to the claimed invention because it is in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jiang to incorporate the teachings of Zhang to generate a document representation using a sentence encoder transformer model and a document encoder transformer model nested in a hierarchical fashion. Doing so would allow for performing document summarization (Zhang; Section 1, lines 77-84). Regarding claim 22, as best understood based on the 35 U.S.C. 112(b) issues identified above, Jiang in view of Zhang discloses the computer-implemented method as claimed in claim 21. Zhang further teaches: wherein the first encoding model comprises a multi-head self-attention mechanism (Section 3.1, lines 14-19, “A transformer encoder usually has multiple layers and each layer is composed of a multi-head self attentive sub-layer followed by a feed-forward sub-layer with residual connections (He et al., 2016) and layer normalizations (Ba et al., 2016).”). Zhang is considered to be analogous to the claimed invention because it is in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jiang in view of Zhang to further incorporate the teachings of Zhang to implement an encoder with a multi-head self-attention mechanism. Doing so would allow for performing document summarization (Zhang; Section 1, lines 77-84). Regarding claim 23, as best understood based on the 35 U.S.C. 112(b) issues identified above, Jiang in view of Zhang discloses the computer-implemented method as claimed in claim 21. Zhang further teaches: wherein processing each of the plurality of first textual blocks with the sentence encoding portion of the first encoding submodel comprises: for each of the plurality of first textual blocks: processing, by the computing system, a respective first textual block with the first hierarchical transformer model to obtain sentence tokens respectively corresponding to words in each of one or more first sentences of the respective first textual block (Section 3.1, lines 1-3, “Let D = (S 1 , S 2 , ..., S |D| ) denote a document, where S i = (w 1 i , w 2 i , ..., w |Si| i ) is a sentence in D and w j i a word in S i .”; Section 3.1, lines 21-38, “To learn the representation of S i , S i = (w 1 i , w 2 i , ..., w |Si| i ) is first mapped into continuous space E i = (e 1 i , e 2 i , ..., e |Si| i ) where e j i = e(w j i ) + p j where e(w j i ) and p j are the word and positional embeddings of w j i , respectively. The word embedding matrix is randomly initialized and we adopt the sine-cosine positional embedding (Vaswani et al., 2017). Then the sentence encoder (a Transformer) transforms E i into a list of hidden representations (h 1 i , h 2 i , ..., h |Si| i ). We take the last hidden representation h |Si| i (i.e., the representation at the EOS token) as the representation of sentence S i . Similar to the representation of each word in S i , we also take the sentence position into account. The final representation of S i is ĥ i = h |Si| i + p i = (S 1 , S 2 , ..., S |D| )”; Learning sentence representations for the sentences in a document from the word embeddings of the words in the sentences reads on obtaining sentence tokens respectively corresponding to words in each of one or more first sentences of the respective first textual block.); and concatenating, by the computing system, a first sentence token of the sentence tokens with a position embedding corresponding to the first sentence token to obtain a first textual block representation for the respective first textual block (Section 3.1, lines 21-38, “To learn the representation of S i , S i = (w 1 i , w 2 i , ..., w |Si| i ) is first mapped into continuous space E i = (e 1 i , e 2 i , ..., e |Si| i ) where e j i = e(w j i ) + p j where e(w j i ) and p j are the word and positional embeddings of w j i , respectively. The word embedding matrix is randomly initialized and we adopt the sine-cosine positional embedding (Vaswani et al., 2017). Then the sentence encoder (a Transformer) transforms E i into a list of hidden representations (h 1 i , h 2 i , ..., h |Si| i ). We take the last hidden representation h |Si| i (i.e., the representation at the EOS token) as the representation of sentence S i . Similar to the representation of each word in S i , we also take the sentence position into account. The final representation of S i is ĥ i = h |Si| i + p i ”; Determining the final sentence representations by concatenating the sentence representations with the positional embedding of the sentences reads on concatenating a first sentence token of the sentence tokens with a position embedding corresponding to the first sentence token to obtain a first textual block representation for the respective first textual block.). Zhang is considered to be analogous to the claimed invention because it is in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jiang in view of Zhang to further incorporate the teachings of Zhang to learn sentence representations for the sentences in a document from the word embeddings of the words in the sentences, and determine the final sentence representations by concatenating the sentence representations with the positional embedding of the sentences. Doing so would allow for performing document summarization (Zhang; Section 1, lines 77-84). Regarding claim 24, as best understood based on the 35 U.S.C. 112(b) issues identified above, Jiang in view of Zhang discloses the computer-implemented method as claimed in claim 23. Jiang further discloses: wherein each of the sentence tokens comprises an attentional weight (Section 3.3, lines 64-65, “With the attention mechanism, the importance of each sentence can be measured”; Section 3.3, lines 68-70, “Finally, the representation of the k-th paragraph can be represented as p k p = ∑ j α j , k p h j , k p .”; The parameter α j , k p reads on an attention weight for each sentence.). Regarding claim 25, as best understood based on the 35 U.S.C. 112(b) issues identified above, Jiang in view of Zhang discloses the computer-implemented method as claimed in claim 23. Jiang further discloses: wherein processing the plurality of first textual block representations with the second hierarchical transformer model of the machine-learned semantic document encoding model to obtain the first document representation comprises: determining, by the computing system based at least in part on the attentional weights of the sentence tokens of each of the plurality of first textual blocks, a weighted sum of the plurality of first textual block representations (Section 3.3, lines 64-65, “With the attention mechanism, the importance of each sentence can be measured”; Section 3.3, lines 68-70, “Finally, the representation of the k-th paragraph can be represented as p k p = ∑ j α j , k p h j , k p .”; The parameter α j , k p reads on attentional weights of the sentence tokens, and p k p reads on a weighted sum of the plurality of first textual block representations.). Zhang further teaches: concatenating, by the computing system, the weighted sum and the contextual block representation associated with at least one first textual block of the plurality of first textual blocks to determine the first document encoding (Section 3.1, lines 21-38, “To learn the representation of S i , S i = (w 1 i , w 2 i , ..., w |Si| i ) is first mapped into continuous space E i = (e 1 i , e 2 i , ..., e |Si| i ) where e j i = e(w j i ) + p j where e(w j i ) and p j are the word and positional embeddings of w j i , respectively. The word embedding matrix is randomly initialized and we adopt the sine-cosine positional embedding (Vaswani et al., 2017). Then the sentence encoder (a Transformer) transforms E i into a list of hidden representations (h 1 i , h 2 i , ..., h |Si| i ). We take the last hidden representation h |Si| i (i.e., the representation at the EOS token) as the representation of sentence S i . Similar to the representation of each word in S i , we also take the sentence position into account. The final representation of S i is ĥ i = h |Si| i + p i ”; Section 3.1, lines 41-47, “In analogy to the sentence encoder, as shown in Figure 1, the document encoder is yet another Transformer but applies on the sentence level. After running the Transformer on a sequence of sentence representations (ĥ 1, ĥ 2, …, ĥ |D| ), we obtain the context sensitive sentence representations (d 1, d 2, …, d |D| ).”; The document encoder being analogous to the sentence encoder and applying a transformer on the sentence level, where the sentence encoder determines the final sentence representations by concatenating the sentence representations with the positional embedding of the sentences, reads on concatenating the weighted sum and the contextual block representation associated with at least one first textual block of the plurality of first textual blocks to determine the first document encoding.). Zhang is considered to be analogous to the claimed invention because it is in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jiang in view of Zhang to further incorporate the teachings of Zhang to implement a document encoder analogous to a sentence encoder that applies a transformer on the sentence level, where the sentence encoder determines the final sentence representations by concatenating the sentence representations with the positional embedding of the sentences. Doing so would allow for performing document summarization (Zhang; Section 1, lines 77-84). Regarding claim 26, as best understood based on the 35 U.S.C. 112(b) issues identified above, Jiang in view of Zhang discloses the computer-implemented method as claimed in claim 21. Jiang further discloses: wherein the method further comprises: evaluating, by the computing system, a loss function that evaluates a difference between the similarity metric and ground truth data associated with the first document and the second document; and adjusting, by the computing system, one or more parameters of the machine-learned semantic document encoding model based at least in part on the loss function (Section 3.5, lines 1-6, “The task of semantic text matching can be modeled as a binary classification problem. Given a tuple of training data (d s ,d c ,y), where y is a Boolean value showing whether two documents are semantically matched, SMASH RNN optimizes the binary cross entropy [20] between the estimated probabilistic score ŷ and the gold standard y.”; Section 4.1, lines 18-19, "The Adam optimizer [26] is applied to optimize the parameters with an initial learning rate of 10 −5 ."); Optimizing the binary cross entropy between the estimated probabilistic score and the gold standard reads on a loss function that evaluates a difference between the similarity metric and ground truth data associated with the first document and the second document and adjusting one or more parameters of the machine-learned semantic document encoding model based at least in part on the loss function, where the binary cross-entropy reads on the loss function, the probabilistic score ŷ reads on the similarity metric, and the gold standard y reads on the ground truth data.). Regarding claim 27, as best understood based on the 35 U.S.C. 112(b) issues identified above, Jiang in view of Zhang discloses the computer-implemented method as claimed in claim 21. Jiang further discloses: wherein the similarity metric comprises: a binary prediction whether the first document and the second document are semantically similar; or a predicted level of semantic similarity between the two documents (Section 3.1, lines 20-24, "Given a source document d s and a set of candidate documents D c , our goal is to estimate semantic similarity ŷ = Sim(d s , d c ) between the source document d s and every candidate document d c ∈ D c so that the target documents semantically matched to the source document have higher semantic similarity scores."; Estimating semantic similarity between a source document and a candidate document reads on a similarity metric comprising a predicted level of semantic similarity between the first document and the second document.). Regarding claim 29, arguments analogous to claim 21 are applicable. In addition, Jiang discloses a computing system, comprising: one or more processors; and one or more tangible, non-transitory computer readable media storing computer-readable instructions (Section 3.3, lines 10-11, “In this paper, we propose to model documents with information from different document structure levels.”; Section 3.3, lines 14-16, “The computation of encoders in MASH RNN follows a bottom-up principle with bidirectional recurrent neural networks (Bi-RNNs) with attention.”; Section 3.3, lines 30-32, “The backward pass processes the input sequence in reverse order and generates the backward hidden states”; Implementing a recurrent neural network, performing computations, and processing demonstrates the use of a processor executing instructions from memory.) that when executed by the one or more processors cause the one or more processors to perform the steps of claim 21. Regarding claim 30, arguments analogous to claim 22 are applicable. Regarding claim 31, arguments analogous to claim 23 are applicable. Regarding claim 32, arguments analogous to claim 24 are applicable. Regarding claim 33, arguments analogous to claim 25 are applicable. Regarding claim 34, arguments analogous to claim 26 are applicable. Regarding claim 35, arguments analogous to claim 27 are applicable. Regarding claim 37, arguments analogous to claim 21 are applicable. In addition, Jiang discloses one or more tangible, non-transitory computer readable media storing computer-readable instructions that when executed by one or more processors cause the one or more processors to perform operations (Section 3.3, lines 10-11, “In this paper, we propose to model documents with information from different document structure levels.”; Section 3.3, lines 14-16, “The computation of encoders in MASH RNN follows a bottom-up principle with bidirectional recurrent neural networks (Bi-RNNs) with attention.”; Section 3.3, lines 30-32, “The backward pass processes the input sequence in reverse order and generates the backward hidden states”; Implementing a recurrent neural network, performing computations, and processing demonstrates the use of a processor executing instructions from memory.), the operations comprising the steps of claim 21. Regarding claim 38, arguments analogous to claim 22 are applicable. Regarding claim 39, arguments analogous to claim 23 are applicable. Regarding claim 40, arguments analogous to claim 24 are applicable . 07-21-aia AIA Claim 28 and 36 are rejected under 35 U.S.C. 103 as being unpatentable over Jiang in view of Zhang, and further in view of Huang et al. (“Embedding-based Retrieval in Facebook Search”), hereinafter Huang . Regarding claim 28, as best understood based on the 35 U.S.C. 112(b) issues identified above, Jiang in view of Zhang discloses the computer-implemented method as claimed in claim 21, but does not specifically disclose: wherein the method further comprises indexing, by the computing system for a search system, the first document encoding as a representation of the first document. Huang teaches: indexing, by the computing system for a search system, the first document encoding as a representation of the first document (Abstract, lines 9-12, "We introduce the unified embedding framework developed to model semantic embeddings for personalized search, and the system to serve embedding-based retrieval in a typical search system based on an inverted index."; Section 2.3, lines 1-6, "To learn embeddings that are optimizing the triplet loss, our model comprises three major components: a query encoder EQ = f (Q) which produces a query embedding, a document encoder ED = g(D) which produces a document embedding, and a similarity function S(EQ, ED) which produces a score between query Q and document D."; The document embedding reads on the document encoding.). Huang is considered to be analogous to the claimed invention because it is in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jiang in view of Zhang to incorporate the teachings of Huang to implement a search system that indexes document embeddings. Doing so would allow for providing semantic matching in search retrieval (Huang; Section 7, lines 1-9). Regarding claim 36, arguments analogous to claim 28 are applicable . Conclusion 07-96 AIA The prior art made of record and not relied upon is considered pertinent to applicant's disclosure : Lu et al. (“TwinBERT: Distilling Knowledge to Twin-Structured BERT Models for Efficient Retrieval”) teaches a model for effective and efficient retrieval with Bidirectional Encoder Representations from Transformers (BERT) encoders to represent query and document respectively and a crossing layer to combine the embeddings and produce a similarity score. The following art made of record and not relied upon is considered pertinent to applicant's disclosure: Yang et al. (“Beyond 512 Tokens: Siamese Multi-depth Transformer-based Hierarchical Encoder for Long-Form Document Matching”) teaches the use of a Siamese multi-depth attention-based hierarchical recurrent neural network for learning long document representations for document matching. Any inquiry concerning this communication or earlier communications from the examiner should be directed to James Boggs whose telephone number is (571)272-2968. The examiner can normally be reached M-F 8:00 AM - 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571)272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JAMES BOGGS/Examiner, Art Unit 2657 Application/Control Number: 18/990,968 Page 2 Art Unit: 2657 Application/Control Number: 18/990,968 Page 3 Art Unit: 2657 Application/Control Number: 18/990,968 Page 4 Art Unit: 2657 Application/Control Number: 18/990,968 Page 5 Art Unit: 2657 Application/Control Number: 18/990,968 Page 7 Art Unit: 2657 Application/Control Number: 18/990,968 Page 8 Art Unit: 2657 Application/Control Number: 18/990,968 Page 9 Art Unit: 2657 Application/Control Number: 18/990,968 Page 10 Art Unit: 2657 Application/Control Number: 18/990,968 Page 11 Art Unit: 2657 Application/Control Number: 18/990,968 Page 12 Art Unit: 2657 Application/Control Number: 18/990,968 Page 13 Art Unit: 2657 Application/Control Number: 18/990,968 Page 14 Art Unit: 2657 Application/Control Number: 18/990,968 Page 15 Art Unit: 2657 Application/Control Number: 18/990,968 Page 16 Art Unit: 2657 Application/Control Number: 18/990,968 Page 17 Art Unit: 2657 Application/Control Number: 18/990,968 Page 18 Art Unit: 2657 Application/Control Number: 18/990,968 Page 19 Art Unit: 2657 Application/Control Number: 18/990,968 Page 20 Art Unit: 2657 Application/Control Number: 18/990,968 Page 21 Art Unit: 2657 Application/Control Number: 18/990,968 Page 22 Art Unit: 2657 Application/Control Number: 18/990,968 Page 23 Art Unit: 2657