CTNF 18/457,002 CTNF 101456 DETAILED ACTION This action is in response to the original filing on August 28, 2023. Claims 1-20 are pending and have been considered below. Claims 1, 8, and 15 are independent claims. Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Information Disclosure Statement The information disclosure statement (IDS) submitted on August 28, 2023 is being considered by the examiner. Drawings 06-22 AIA The drawings are objected to because of the following informalities: Fig. 2 does not include “precompute module 205” described in page 14, paragraph 53 of the description Fig. 2C, described in page 17, paragraph 58 of the description, is missing Fig. 4 does not include “sparsity modules 450” described in page 22, paragraph 71 of the description Fig. 7 does not include “sparsity bitmap 735” described in page 20, paragraph 68 of the description In Fig. 8, “Fx” and “Fy” should read “W f ” and “H f ,” respectively, as described in page 42, paragraph 128 of the description Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Specification 07-29 AIA The disclosure is objected to because of the following informalities: On page 15, paragraph 54, “230and” should read “230 and” and “data transfer is in being” should read “data transfer is being” On page 21, paragraph 70, “may determine a position the activation” should read “may determine a position of the activation” On page 26, paragraph 84, “centering matric” should read “centering matrix” On page 28, paragraph 88, “negative predicts” should read “negative predictions” On page 31, paragraph 97, “eq. 7 and eq. 8” are missing On page 35, paragraph 109, “444 is a duplication of the auxiliary layer set 444” should read “444 is a duplication of the auxiliary layer set 442” On page 40, paragraph 120, “input tensor 510” should read “input tensor 610” On pages 41-42, paragraph 127, “7x7 3D matrix” should read “7x7 2D matrix” On page 44, paragraph 133, “multiple the activation and the weight” should read “multiply the activation and the weight” On page 44, paragraph 134, “PE array 250 in Fig. 3” should read “PE array 250 in Fig. 2” On page 45, paragraph 138, “local memory 240 in Fig. 3” should read “local memory 240 in Fig. 2” On page 45, paragraph 139, “in conjunction with Fig. 3” should read “in conjunction with Fig. 2” On page 50, paragraph 154, “tensor computed based on” should read “tensor is computed based on” Appropriate correction is required. Claim Rejections - 35 USC § 112 07-30-02 AIA The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. 07-34-01 Claims 1-2, 5-6, 8-9, 12-13, 15-16, and 18-19 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The term the one or more deep learning operations that compute a second detection tensor in claim 1 lacks sufficient antecedent basis as there is no prior reference to the one or more deep learning operations that compute a second detection tensor made in claim 1. For examination purposes, this term is interpreted to mean the one or more deep learning operations that compute a second detection tensor . The term the one or more other deep learning operations that compute a fourth detection tensor in claim 2 lacks sufficient antecedent basis as there is no prior reference to the one or more other deep learning operations that compute a fourth detection tensor made in these claims. For examination purposes, this term is interpreted to mean the one or more other deep learning operations that compute a fourth detection tensor . The term the one or more other operations in claim 5 lacks sufficient antecedent basis as claim 2 includes both one or more other deep learning operations that compute a third detection tensor and the one or more other deep learning operations that compute a fourth detection tensor and it is unclear which operations the applicant is referring to. For examination purposes this term is interpreted to mean “one or more other deep learning operations that compute a third detection tensor” in claim 2. The term the one or more other operations in claim 6 lacks sufficient antecedent basis as claim 2 includes both one or more other deep learning operations that compute a third detection tensor and the one or more other deep learning operations that compute a fourth detection tensor and claim 5 includes the one or more other operations and it is unclear which operations the applicant is referring to. For examination purposes this term is interpreted to mean “the one or more other operations” in claim 5. Claims 8-9 and 12-13 recite a non-transitory computer-readable media that parallels the method claims of 1-2 and 5-6, respectively. Therefore, the analysis discussed above with respect to claims 1-2 and 5-6 applies to claims 8-9 and 12-13, respectively. Accordingly, claims 8-9 and 12-13 are rejected based on substantially the same rationale as set forth above with respect to claims 1-2 and 5-6, respectively. Claims 15-16 and 18-19 recite an apparatus that parallels the method claims of 1-2 and 5-6, respectively. Therefore, the analysis discussed above with respect to claims 1-2 and 5-6 applies to claims 15-16 and 18-19, respectively. Accordingly, claims 15-16 and 18-19 are rejected based on substantially the same rationale as set forth above with respect to claims 1-2 and 5-6, respectively. Claim Rejections - 35 USC § 101 07-04-01 AIA 07-04 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 1 : Step 1 – Claim 1 is directed to a method: A method of training a neural network … Step 2A, Prong 1 – A judicial exception is recited in this claim as it recites mathematical concepts and a mental process (see MPEP 2106.04(a)(2)(I) and MPEP 2106.04(a)(2)(III)): selecting a layer in a backbone of the neural network, the layer outputting an intermediate tensor based on the training dataset … A person can reasonably perform “selecting a layer in a backbone of the neural network” given a small enough number of layers in a neural network. Hence, “selecting a layer in a backbone of the neural network, the layer outputting an intermediate tensor based on the training dataset” is a mental process. determining a loss for the neural network, the loss comprising a diversity loss, the diversity loss indicating a measurement of similarity between the first detection tensor and the second detection tensor … To determine a “loss comprising a diversity loss… indicating a measurement of similarity” involves using a formula to calculate a similarity between two tensors, which is a mathematical concept. Hence, “determining a loss for the neural network, the loss comprising a diversity loss, the diversity loss indicating a measurement of similarity between the first detection tensor and the second detection tensor” is a mathematical concept. training the neural network by adjusting one or more weights in the backbone, first head, or second head based on the loss … To train a neural network “by adjusting one or more weights” involves following an algorithm, for example gradient descent, to modify the weights of a neural network, which is a mathematical concept. Hence, “training the neural network by adjusting one or more weights in the backbone, first head, or second head based on the loss” is a mathematical concept. Step 2A, Prong 2 – The following limitations are additional elements without significantly more than the abstract idea: inputting a training dataset into the neural network … “inputting a training dataset” amounts to insignificant extra-solution activity of necessary data gathering that does not add a meaningful limitation to the “method” (see MPEP 2106.05(g)). inputting the intermediate tensor into a first head of the neural network, the first head comprising one or more deep learning operations that compute a first detection tensor … “inputting the intermediate tensor” amounts to insignificant extra-solution activity of necessary data gathering that does not add a meaningful limitation to the “method” (see MPEP 2106.05(g)). inputting the intermediate tensor into a second head of the neural network, the second head comprising the one or more deep learning operations that compute a second detection tensor that is different from the first detection tensor… “inputting the intermediate tensor” amounts to insignificant extra-solution activity of necessary data gathering that does not add a meaningful limitation to the “method” (see MPEP 2106.05(g)). Step 2B – These elements are recited at such a high level of generality that they fail to integrate the abstract idea into a practical application, since they only amount to data gathering or outputting without significantly more (MPEP 2106.05(g)). These limitations, either taken alone or in combination, fail to provide an inventive concept. Thus, the claim is not patent eligible. Claims 2-7 recite limitations which further narrow the abstract ideas of claim 1 by specifying more details of the mathematical concepts that occur: Regarding claim 2 , this claim further limits the abstract ideas of claim 1 to be based on a mathematical concept and mental process: selecting another layer in the backbone of the neural network, the another layer generating another intermediate tensor based on the training dataset … Given a small enough number of layers in the neural network, selecting another layer can reasonably be performed in the human mind, which is a mental process. Furthermore, wherein the loss further comprises another diversity loss that indicates a measurement of similarity between the third detection tensor and the fourth detection tensor … Another diversity loss that indicates a measurement of similarity between two tensors involves using a mathematical formula to calculate a measurement of similarity, which is a mathematical concept. Furthermore, specifying inputting the another intermediate tensor into a third head of the neural network, the third head comprising one or more other deep learning operations that compute a third detection tensor … and inputting the another intermediate tensor into a fourth head of the neural network, the fourth head comprising the one or more other deep learning operations that compute a fourth detection tensor that is different from the third detection tensor … is still insignificant extra-solution activity of necessary data gathering (see MPEP 2106.05(g)). Regarding claim 3 , specifying wherein the another intermediate tensor has a different size from the intermediate tensor in this manner does not overcome the rejection of claim 1 as modifying the another intermediate tensor does not make the abstract ideas of claim 1 to not be mathematical concepts or mental processes. Regarding claim 4 , specifying inputting the first detection tensor into the third head, the third detection tensor computed based on the first detection tensor and the another intermediate tensor … and inputting the second detection tensor into the fourth head, the fourth detection tensor computed based on the second detection tensor and the another intermediate tensor … is still insignificant extra-solution activity of necessary data gathering (see MPEP 2106.05(g)). Regarding claim 5 , this claim further limits the abstract ideas of claim 4 to be based on a mathematical concept: wherein the one or more other operations comprise an upsampling operation that computes an upsampled tensor by increasing one or more dimensions of the first detection tensor … An upsampling operation that computes an upsampled tensor involves padding the values in a tensor with additional zeroes or values, which is a mathematical concept. Regarding claim 6 , this claim further limits the abstract ideas of claim 5 to be based on a mathematical concept: wherein the one or more other operations further comprise a concatenation operation that computes a tensor by concatenating the upsampled tensor with the another intermediate tensor … A concatenation operation that computes a tensor involves appending two tensors together to create a single, larger tensor, which is a mathematical concept. Regarding claim 7 , this claim further limits the abstract ideas of claim 1 to be based on a mathematical concept: wherein the measurement of similarity comprises a measurement of centered kernel alignment similarity between the first detection tensor and the second detection tensor or a measurement of cosine similarity between the first detection tensor and the second detection tensor … A measurement of centered kernel alignment similarity or a measure of cosine similarity both involve calculating the similarity between two sets of data, which is a mathematical concept. Regarding claim 8 : Step 1 – Claim 8 is directed to a product: One or more non-transitory computer-readable media … Step 2A, Prong 1 – A judicial exception is recited in this claim as it recites mathematical concepts and a mental process (see MPEP 2106.04(a)(2)(I) and MPEP 2106.04(a)(2)(III)): selecting a layer in a backbone of the neural network, the layer outputting an intermediate tensor based on the training dataset … A person can reasonably perform “selecting a layer in a backbone of the neural network” given a small enough number of layers in a neural network. Hence, “selecting a layer in a backbone of the neural network, the layer outputting an intermediate tensor based on the training dataset” is a mental process. determining a loss for the neural network, the loss comprising a diversity loss, the diversity loss indicating a measurement of similarity between the first detection tensor and the second detection tensor … To determine a “loss comprising a diversity loss… indicating a measurement of similarity” involves using a formula to calculate a similarity between two tensors, which is a mathematical concept. Hence, “determining a loss for the neural network, the loss comprising a diversity loss, the diversity loss indicating a measurement of similarity between the first detection tensor and the second detection tensor” is a mathematical concept. training the neural network by adjusting one or more weights in the backbone, first head, or second head based on the loss … To train a neural network “by adjusting one or more weights” involves following an algorithm, for example gradient descent, to modify the weights of a neural network, which is a mathematical concept. Hence, “training the neural network by adjusting one or more weights in the backbone, first head, or second head based on the loss” is a mathematical concept. Step 2A, Prong 2 – The following limitations are additional elements without significantly more than the abstract idea: One or more non-transitory computer-readable media storing instructions executable to perform operations for training a neural network … one or more non-transitory computer-readable media used as mere tools to apply an exception are generic elements for performing or applying the abstract idea using a generic computing environment (see MPEP 2106.05(f)). inputting a training dataset into the neural network … “inputting a training dataset” amounts to insignificant extra-solution activity of necessary data gathering that does not add a meaningful limitation to the “non-transitory computer-readable media” (see MPEP 2106.05(g)). inputting the intermediate tensor into a first head of the neural network, the first head comprising one or more deep learning operations that compute a first detection tensor … “inputting the intermediate tensor” amounts to insignificant extra-solution activity of necessary data gathering that does not add a meaningful limitation to the “non-transitory computer-readable media” (see MPEP 2106.05(g)). inputting the intermediate tensor into a second head of the neural network, the second head comprising the one or more deep learning operations that compute a second detection tensor that is different from the first detection tensor… “inputting the intermediate tensor” amounts to insignificant extra-solution activity of necessary data gathering that does not add a meaningful limitation to the “non-transitory computer-readable media” (see MPEP 2106.05(g)). Step 2B – These elements are recited at such a high level of generality that they fail to integrate the abstract idea into a practical application, since they provide nothing more than mere instructions to implement an abstract idea on a generic computer (MPEP 2106.05(f)) or only amount to data gathering or outputting without significantly more (MPEP 2106.05(g)). These limitations, either taken alone or in combination, fail to provide an inventive concept. Thus, the claim is not patent eligible. Claims 9-14 recite a non-transitory computer-readable media that parallels the method claims of 2-7, respectively. Therefore, the analysis discussed above with respect to claims 2-7 applies to claims 9-14, respectively. Accordingly, claims 9-14 are rejected based on substantially the same rationale as set forth above with respect to claims 2-7, respectively. Regarding claim 15 : Step 1 – Claim 15 is directed to an apparatus: An apparatus … Step 2A, Prong 1 – A judicial exception is recited in this claim as it recites mathematical concepts and a mental process (see MPEP 2106.04(a)(2)(I) and MPEP 2106.04(a)(2)(III)): selecting a layer in a backbone of the neural network, the layer outputting an intermediate tensor based on the training dataset … A person can reasonably perform “selecting a layer in a backbone of the neural network” given a small enough number of layers in a neural network. Hence, “selecting a layer in a backbone of the neural network, the layer outputting an intermediate tensor based on the training dataset” is a mental process. determining a loss for the neural network, the loss comprising a diversity loss, the diversity loss indicating a measurement of similarity between the first detection tensor and the second detection tensor … To determine a “loss comprising a diversity loss… indicating a measurement of similarity” involves using a formula to calculate a similarity between two tensors, which is a mathematical concept. Hence, “determining a loss for the neural network, the loss comprising a diversity loss, the diversity loss indicating a measurement of similarity between the first detection tensor and the second detection tensor” is a mathematical concept. training the neural network by adjusting one or more weights in the backbone, first head, or second head based on the loss … To train a neural network “by adjusting one or more weights” involves following an algorithm, for example gradient descent, to modify the weights of a neural network, which is a mathematical concept. Hence, “training the neural network by adjusting one or more weights in the backbone, first head, or second head based on the loss” is a mathematical concept. Step 2A, Prong 2 – The following limitations are additional elements without significantly more than the abstract idea: a computer processor for executing computer program instructions … a computer processor used as a mere tool to apply an exception is a generic element for performing or applying the abstract idea using a generic computing environment (see MPEP 2106.05(f)). a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations for training a neural network … a non-transitory computer-readable memory used as a mere tool to apply an exception is a generic element for performing or applying the abstract idea using a generic computing environment (see MPEP 2106.05(f)). inputting a training dataset into the neural network … “inputting a training dataset” amounts to insignificant extra-solution activity of necessary data gathering that does not add a meaningful limitation to the “apparatus” (see MPEP 2106.05(g)). inputting the intermediate tensor into a first head of the neural network, the first head comprising one or more deep learning operations that compute a first detection tensor … “inputting the intermediate tensor” amounts to insignificant extra-solution activity of necessary data gathering that does not add a meaningful limitation to the “apparatus” (see MPEP 2106.05(g)). inputting the intermediate tensor into a second head of the neural network, the second head comprising the one or more deep learning operations that compute a second detection tensor that is different from the first detection tensor… “inputting the intermediate tensor” amounts to insignificant extra-solution activity of necessary data gathering that does not add a meaningful limitation to the “apparatus” (see MPEP 2106.05(g)). Step 2B – These elements are recited at such a high level of generality that they fail to integrate the abstract idea into a practical application, since they provide nothing more than mere instructions to implement an abstract idea on a generic computer (MPEP 2106.05(f)) or only amount to data gathering or outputting without significantly more (MPEP 2106.05(g)). These limitations, either taken alone or in combination, fail to provide an inventive concept. Thus, the claim is not patent eligible. Claims 16-20 recite an apparatus that parallels the method claims of 2 and 4-7, respectively. Therefore, the analysis discussed above with respect to claims 2 and 4-7 applies to claims 9-14, respectively. Accordingly, claims 9-14 are rejected based on substantially the same rationale as set forth above with respect to claims 2 and 4-7, respectively. Claim Rejections - 35 USC § 103 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-21-aia AIA Claim s 1-3, and 7 are rejected under 35 U.S.C. 103 as being unpatentable over Ro et al. (“Heterogeneous Double-Head Ensemble for Deep Metric Learning,” 2020, hereinafter Ro) in view of Li et al. (“On the diversity of multi-head attention, 2021, hereinafter Li) . Regarding claim 1 , Ro teaches a method of training a neural network, comprising: inputting a training dataset into the neural network (Page 1, Col. 1, Section 1, ¶1 “ Deep metric learning refers to the design of feature extracting functions with deep neural networks so that the features of semantically similar images are close to others. Ensemble is a method of ensuring robust performance by training diverse models and aggregating their prediction results,” Page 2, Col. 2, Section 3, ¶1 “The multi-head structure for deep metric learning ,” Page 5, Col. 2, Section IV, ¶1 “we provide experimental results for evaluating the proposed structure , especially in deep metric learning for image retrieval tasks,” Page 3, Col. 1, ¶1 “For the baseline network, we adopt the ResNet-50 network … All experiments covered in this section are conducted on validating the CUB-200–2011 dataset . The CUB-200–2011 training dataset is divided into training and validation sets equally,” Page 7, Col. 2, Section D, ¶1 “The proposed HDhE structure trained by only softmax loss with a random sampling manner… a simple and fast way to train the model ”). Ro further teaches selecting a layer in a backbone of the neural network, the layer outputting an intermediate tensor based on the training dataset (Page 2, Col. 2, Section III, ¶1 “the feature vector should be of a small dimension… the structure that shares low-level layers should save memory and not degrade performance,” Page 3, Col. 1, ¶1 “For the baseline network, we adopt the ResNet-50 network, most popular backbone model ,” Page 1, Fig. 1 depicts “Shared layers,” or a backbone of the neural network , Page 3, Col. 1, Section A, ¶1 “we set up multiple variants of the last feature block in the direction in order to reduce feature dimension… Then, we select one of the variants ,” Page 3, Col. 2, Section B, ¶2 “The layers of ResNet-50 network can be divided into… blocks… We have simply redefined it as {B1, B2, B3, B4} ,” one of ordinary skill in the art would recognize that individual blocks of a “ResNet-50 network” produce tensors as their outputs, hence outputting an intermediate tensor based on the training dataset is implicit when generating a tensor from an intermediate block, for example B3, in a shared “ResNet-50 network,” or backbone , that receives a training dataset , Page 5, Fig. 3 – (b), Page 5, Col. 2, Section D, ¶2 “we can customize the multi-head structure in various views. For example in Figure 3-(b), a memory-saving version of HDhE (Ms-HDhE) can be suggested as follows. In Ms-HDhE, the shared body is enlarged Until B3 ,” wherein “customizing” the last block in the shared backbone , for example B3, encompasses selecting a layer in a backbone of the neural network , as the last layer in a selected block is implied to be the layer outputting an intermediate tensor based on the training dataset ). Ro further teaches inputting the intermediate tensor into a first head of the neural network, the first head comprising one or more deep learning operations that compute a first detection tensor (Page 1, Col. 1, Section I, ¶2 “A variety of feature vectors can be obtained by semantically diverse attention of the input image ,” wherein “feature vectors” generated from an “input image” encompasses detection tensors , Page 2, Col. 1, ¶2 “we search for an effective multi-head ensemble structure referred to as Heterogeneous Double-head Ensemble (HDhE), as shown in Figure 1… we consider… designing a dimension-reduced feature vector, determining the shared body part, and designing diverse heads to generate diverse feature vectors ,” Page 1, Fig. 1 depicts a first head of the neural network computing a first detection tensor , “feature vector f1,” after receiving an input from a preceding block; one of ordinary skill in the art would recognize that inputting the intermediate tensor into a first head of the neural network is implicit when using a multi-head architecture that receives an input from an intermediate block of a ResNet-50 backbone , as explained above, Page 5, Fig. 3 – (b) and Col. 2, Section D, ¶2 “the heads can be chosen by Res4_2 and Res4_1 with their learning rates of 0.5× and 4×, respectively,” wherein a first head with a unique block architecture and “learning rate” encompasses one or more deep learning operations that compute a first detection tensor ). Ro further teaches inputting the intermediate tensor into a second head of the neural network, the second head comprising the one or more deep learning operations that compute a second detection tensor that is different from the first detection tensor (Page 2, Col. 1, ¶2 “designing diverse heads to generate diverse feature vectors,” Page 1, Fig. 1 depicts a second head of the neural network computing a second detection tensor that is different from the first detection tensor , “feature vector f2,” after receiving an input from a preceding block; one of ordinary skill in the art would recognize that inputting the intermediate tensor into a second head of the neural network is implicit when using a multi-head architecture that receives an input from an intermediate block of a ResNet-50 backbone , as explained above, Page 5, Fig. 3 – (b) and Col. 2, Section D, ¶2 “the heads can be chosen by Res4_2 and Res4_1 with their learning rates of 0.5× and 4× , respectively,” wherein a second head with a different block architecture and “learning rate” from the first head encompasses one or more deep learning operations that compute a second detection tensor that is different from the first detection tensor ). Ro teaches determining a loss for the neural network (Page 2, Col. 1, ¶3 “To show the advantages of our HDhE structure… we show that the HDhE is valid for other types of loss”) and the first detection tensor and the second detection tensor (Page 1, Fig. 1 – f1-f2, Page 2, Col. 1, ¶2 “designing diverse heads to generate diverse feature vectors”). However, Ro fails to teach the loss comprising a diversity loss, the diversity loss indicating a measurement of similarity between the first detection tensor and the second detection tensor . Li, in the same field of endeavor, teaches the loss comprising a diversity loss, the diversity loss indicating a measurement of similarity between outputs of a first head and a second head (Page 3, Col. 1, Section 3.1, ¶1 “To further guarantee the diversity , we enlarge the distances among multiple attention heads with disagreement regularization . To this end, we introduce an auxiliary regularization term in order to encourage the diversity among multiple attention heads,” Page 3, Col. 2, Equation (7) and ¶2 “ Disagreement on Outputs. This disagreement directly applies regularization on the outputs of each attention head, by maximizing the difference among them … we employ negative cosine similarity to measure the distance ,” Page 5, Col. 1, Section 3.3, ¶1 “disagreement regularization focuses on adjusting the training objective, i.e. the loss function ”). Ro teaches training the neural network by adjusting one or more weights in the backbone, first head, or second head based on the loss (Page 6, Col. 1, Section B, ¶1 “our HDhE was trained along with classifier by softmax (cross-entropy) loss. The leaning rates were initially set to 0.001 for convolutional blocks (B1, B2, B3, and B4) and 0.01 for the classifier… The optimizer used in this paper was stochastic gradient descent (SGD) with nesterov momentum. The initial momentum rate and the weight decay were set to 0.9 and 5×10 −4 , respectively”). Ro and Li are analogous to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the diversity loss of Li with the methodology of Ro. The motivation to do so is to “further improve… model performance” (Li, Abstract). Regarding claim 2 , Ro in view of Li teaches the method of claim 1 (and thus the rejection of claim 1 is incorporated). Ro teaches selecting another layer in the backbone of the neural network, the another layer generating another intermediate tensor based on the training dataset (Page 3, Col. 1, ¶1 “For the baseline network, we adopt the ResNet-50 network, most popular backbone model … All experiments covered in this section are conducted on validating the CUB-200–2011 dataset . The CUB-200–2011 training dataset is divided into training and validation sets equally,” Page 1, Fig. 1 , Page 3, Col. 1, Section A, ¶1 “we set up multiple variants of the last feature block in the direction in order to reduce feature dimension … Then, we select one of the variants ,” Page 3, Col. 2, Section B, ¶2 “The layers of ResNet-50 network can be divided into five blocks… We have simply redefined it as {B1, B2, B3, B4} ,” Page 5, Fig. 3 – (a), Page 4, Col. 1, ¶3 “the Until B2 is selected for our proposed structure,” wherein selecting another block “B2” encompasses selecting another layer in the backbone of the neural network , as explained above with respect to claim 1; one of ordinary skill in the art would recognize that the another layer generating another intermediate tensor based on the training dataset is implicit when generating an output from an intermediate block, for example B2, in a shared “ResNet-50 network,” or backbone , that receives a training dataset ). Ro teaches inputting the another intermediate tensor into a third head of the neural network, the third head comprising one or more other deep learning operations that compute a third detection tensor (Page 7, Table 8 and Col. 1, Section 4, ¶1 “Table 8 shows the evaluation results of the model using triple-heads with Until B2 shared structure ,” one of ordinary skill in the art would recognize that inputting the another intermediate tensor into a third head of the neural network is implicit when using a multi-head architecture that receives an output from an intermediate block “B2” of a ResNet-50 backbone , as explained above with respect to claim 1 , Page 1, Fig. 1 depicts multiple heads of the neural network outputting diverse vectors after receiving an input , Page 4, Col. 1, Section C, ¶1 “The third design concept of the proposed multi-head structure is to produce diverse feature vectors from multi-heads… the diverse designs of multi-heads can yield diverse feature vectors ,” Page 7, Col. 1, Section 4, ¶1 “The triple-heads were designed as Res4_3, Res4_2, and Res4_1 of Last layer and 0.5× , 2× , and 4× of the default learning rate,” wherein a third head with a different block architecture and “learning rate” from the other two heads encompasses the third head comprising one or more other deep learning operations that compute a third detection tensor ). Ro teaches inputting the another intermediate tensor into a fourth head of the neural network, the fourth head comprising the one or more other deep learning operations that compute a fourth detection tensor that is different from the third detection tensor (Page 7, Col. 2, ¶1 “four-head structure,” one of ordinary skill in the art would recognize that inputting the another intermediate tensor into a fourth head of the neural network is implicit when using a multi-head architecture that receives an output from an intermediate block “B2” of a ResNet-50 backbone , as explained above with respect to claim 1, Page 1, Fig. 1 depicts multiple heads of the neural network outputting diverse vectors after receiving an input , Page 4, Col. 1, Section C, ¶1 “The third design concept of the proposed multi-head structure is to produce diverse feature vectors from multi-heads… the diverse designs of multi-heads can yield diverse feature vectors ,” wherein a fourth head of the neural network, the fourth head comprising the one or more other deep learning operations that compute a fourth detection tensor that is different from the third detection tensor is implicit when describing a “four-head structure” with “diverse designs” for each head to “yield” a “diverse feature vector” when compared to the outputs from the other three heads ). Ro teaches the third detection tensor and the fourth detection tensor (Page 2, Col. 1, ¶2 “we consider… designing diverse heads to generate diverse feature vectors,” Page 7, Col. 2, ¶1 “the triple-head or four-head structure,” wherein the third detection and the fourth detection tensor are implicit when generating “diverse feature vectors” from “the triple-head or four-head structure” ). However, Ro fails to teach wherein the loss further comprises another diversity loss that indicates a measurement of similarity between the third detection tensor and the fourth detection tensor . Li teaches wherein the loss further comprises another diversity loss that indicates a measurement of similarity between outputs of a third and fourth head (Page 3, Col. 1, Section 3.1, ¶1 “To further guarantee the diversity , we enlarge the distances among multiple attention heads with disagreement regularization . To this end, we introduce an auxiliary regularization term in order to encourage the diversity among multiple attention heads ,” Page 3, Col. 2, Equation (7) and ¶2 “ Disagreement on Outputs. This disagreement directly applies regularization on the outputs of each attention head ,” Page 5, Col. 1, Section 3.3, ¶1 “disagreement regularization focuses on adjusting the training objective, i.e. the loss function ,” Page 5, Col. 2, ¶1 “the number of attention head is 8,” wherein the calculation of the “disagreement regularization” or another diversity loss that indicates a measurement of similarity is implied to occur between “the outputs of each attention head,” or between any of the 8 heads, which includes a third and fourth head and their respective outputs ). Ro and Li are analogous to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the another diversity loss of Li with the methodology of Ro. The motivation to do so is to “further improve… model performance” (Li, Abstract). Regarding claim 3 , Ro in view of Li teaches the method of claim 2 (and thus the rejection of claim 2 is incorporated). Ro teaches wherein the another intermediate tensor has a different size from the intermediate tensor (Page 3, Col. 1, ¶1 “we adopt the ResNet-50 network, most popular backbone model ,” Col. 2, Section B, ¶2 “The layers of ResNet-50 network can be divided into five blocks as {conv1, Res1_x, Res2_x, Res3_x, Res4_x}. We have simply redefined it as {B1, B2, B3, B4} by combining conv1 into Res1_x,” Page 1, Fig. 1 depicts block “B2” producing an output to two heads; one of ordinary skill in the art would recognize that it is implicit that the output is the another intermediate tensor , as explained above with respect to claim 1, Page 5, Fig. 3 – (b) depicts block “B3” producing an output to two heads; one of ordinary skill in the art would recognize that it is implicit that the output is the intermediate tensor , as explained above with respect to claim 1; although Ro does not explicitly state wherein the another intermediate tensor has a different size from the intermediate tensor , one of ordinary skill in the art would recognize that different “blocks” in a “ResNet-50 network” such as “Res2_x, Res3_x,” which has been redefined as blocks “B2, B3,” output tensors of different sizes, hence wherein the another intermediate tensor has a different size from the intermediate tensor is implicit ). Regarding claim 7 , Ro in view of Li teaches the method of claim 1 (and thus the rejection of claim 1 is incorporated). Ro teaches the first detection tensor and the second detection tensor (Page 1, Fig. 1 – f1-f2, Page 2, Col. 1, ¶2 “designing diverse heads to generate diverse feature vectors”). However, Ro fails to teach wherein the measurement of similarity comprises a measurement of centered kernel alignment similarity between the first detection tensor and the second detection tensor or a measurement of cosine similarity between the first detection tensor and the second detection tensor . Li teaches wherein the measurement of similarity comprises a measurement of centered kernel alignment similarity between the first detection tensor and the second detection tensor or a measurement of cosine similarity between outputs of a first head and a second head (Page 3, Col. 2, Equation (7) and ¶2 “ Disagreement on Outputs. This disagreement directly applies regularization on the outputs of each attention head, by maximizing the difference among them… we employ negative cosine similarity to measure the distance ”). Ro and Li are analogous to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the measurement of cosine similarity of Li with the methodology of Ro. The motivation to do so is to “further improve… model performance” (Li, Abstract) . 07-21-aia AIA Claim s 4-6 are rejected under 35 U.S.C. 103 as being unpatentable over Ro in view of Li, and further in view of Lai et al. (“DCPNet: A Densely Connected Pyramid Network for Monocular Depth Estimation,” 2021, hereinafter Lai) . Regarding claim 4 , Ro in view of Li teaches the method of claim 2 (and thus the rejection of claim 2 is incorporated). Regarding the limitation inputting the first detection tensor into the third head, the third detection tensor computed based on the first detection tensor and the another intermediate tensor , Ro teaches the another intermediate tensor (Page 3, Col. 1, ¶1 “ResNet-50 network… backbone model,” Page 3, Col. 2, Section B, ¶2 “The layers of ResNet-50 network can be divided into… blocks… {B1, B2, B3, B4},” Page 1, Fig. 1 depicts multiple heads receiving an output from block “B2,” one of ordinary skill in the art would recognize that the output from “B2” is the another intermediate tensor , as explained above with respect to claim 1 ). However, Ro fails to teach inputting the first detection tensor into the third head, the third detection tensor computed based on the first detection tensor and the another intermediate tensor . Lai, in the same field of endeavor, teaches inputting the first detection tensor into the third head, the third detection tensor computed based on the first detection tensor and another tensor (Page 5, Section 3.1, ¶1 “The encoder part is a traditional pre-trained backbone such as ResNet… The input color image is highly compressed as a very dense feature (i.e., 𝑆 /32 feature block in the illustration) through the encoder part, which contains a large amount of deeply stacked convolution blocks. In the encoder process, the intermediate features with the size of 𝑆 /2, 𝑆 /4, 𝑆 /8 and 𝑆 /16 are preserved to be connected with the decoder part,” ¶2 “The proposed decoder is a pyramid-like structure which contains six floors (i.e., 𝐹 1˜ 𝐹 6 from top to bottom). Each floor contains six layers (i.e., 𝐿 1˜ 𝐿 6 from left to right),” one of ordinary skill in the art would recognize that a “floor” that produces “features” at the end of a “backbone” functions in substantially the same manner as a “head,” Page 6, ¶1 “The features generated… are preserved to be fused with features in lower floors,” ¶2 “a dense connection module (DCM) is used to fuse features on higher floors with features on the current floor,” Page 6, Fig. 2, Caption: “The resolution of the features on the last layer of every floor is S … the same as the input image,” one of ordinary skill in the art would recognize that “features on the last layer of every floor” generated after receiving an “input color image” encompass detection tensors ; Figure 2 depicts inputting the first detection tensor “D1” from the first head “F1” into the third head “F3,” which computes the third detection tensor “D3” based on the first detection tensor “D1” and other tensors from previous layers in third head “F3” ). Regarding the limitation and inputting the second detection tensor into the fourth head, the fourth detection tensor computed based on the second detection tensor and the another intermediate tensor , Ro teaches the another intermediate tensor (Page 3, Col. 1, ¶1, Page 3, Col. 2, Section B, ¶2, Page 1, Fig. 1 as explained above ). However, Ro fails to teach and inputting the second detection tensor into the fourth head, the fourth detection tensor computed based on the second detection tensor and the another intermediate tensor . Lai teaches and inputting the second detection tensor into the fourth head, the fourth detection tensor computed based on the second detection tensor and another tensor (Page 5, Section 3.1, ¶1 “The encoder part is a… backbone such as ResNet… which contains a large amount of deeply stacked convolution blocks… the intermediate features… are preserved to be connected with the decoder part,” ¶2 “The… decoder is a pyramid-like structure which contains six floors… Each floor contains six layers,” Page 6, ¶1 “The features generated… are preserved to be fused with features in lower floors,” ¶2 “a dense connection module (DCM) is used to fuse features on higher floors with features on the current floor,” Page 6, Fig. 2, Caption: “The resolution of the features on the last layer of every floor is S … the same as the input image,” Figure 2 depicts and inputting the second detection tensor “D2” from the second head “F2” into the fourth head “F4,” which computes the fourth detection tensor “D4” based on the second detection tensor “D2” and other tensors from previous layers in fourth head “F4” ). Ro and Lai are analogous to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the multi-head neural network architecture of Lai with the another intermediate tensor of Ro. The motivation to do so is to offer “a more efficient way to fuse features from multiple scales” (Lai, Abstract). Regarding claim 5 , Ro in view of Li and further in view of Lai teaches the method of claim 4 (and thus the rejection of claim 4 is incorporated). Regarding the limitation wherein the one or more other operations comprise an upsampling operation that computes an upsampled tensor by increasing one or more dimensions of the first detection tensor , Ro teaches the first detection tensor (Page 1, Fig. 1 – “feature vector” f1). However, Ro fails to teach wherein the one or more other operations comprise an upsampling operation that computes an upsampled tensor by increasing one or more dimensions of the first detection tensor . Lai teaches wherein the one or more other operations comprise an upsampling operation that computes an upsampled tensor by increasing one or more dimensions of a tensor (Page 5, Section 3.1, ¶2 “The upscale block is a sequence of a convolution operation and a nearest upsample operation … the latter enlarges the size of the feature … The upscale blocks enlarge the size of the features by a ratio of 2,” Page 6, Fig. 2 depicts a third head “F3” receiving a “feature” from the “Encoder backbone” that is passed through an “Upscale block,” wherein the one or more other operations comprises an upsampling operation that computes an upsampled tensor by increasing one or more dimensions of a tensor ). Ro and Lai are analogous to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the upsampling operation of Lai with the first detection tensor of Ro. The motivation to do so is to offer “a more efficient way to fuse features from multiple scales” (Lai, Abstract). Regarding claim 6 , Ro in view of Li and further in view of Lai teaches the method of claim 5 (and thus the rejection of claim 5 is incorporated). Regarding the limitation wherein the one or more other operations further comprise a concatenation operation that computes a tensor by concatenating the upsampled tensor with the another intermediate tensor , Ro teaches the another intermediate tensor (Page 3, Col. 1, ¶1, Page 3, Col. 2, Section B, ¶2, Page 1, Fig. 1 as explained above with respect to claim 4 ). However, Ro fails to teach wherein the one or more other operations further comprise a concatenation operation that computes a tensor by concatenating the upsampled tensor with the another intermediate tensor . Lai teaches wherein the one or more other operations further comprise a concatenation operation that computes a tensor by concatenating the upsampled tensor with another tensor (Page 7, Fig. 3, Caption: “the symbols ©… indicate concatenation ”, Page 6, Section 3.2, ¶1 “Figure 3 shows the dense connection module… 𝑓 𝑖 , 𝑗 denotes one of the features in floor i as well as layer j , which is obtained from an adjacent upscale block and will be sent to next dense connection module. In the DCM, feature 𝑓 𝑖 , 𝑗 is concatenated with features obtained from DCMs at the same layer j from higher floors ,” Page 5, Section 3.1, ¶2 “The upscale block… enlarges the size of the feature,” Page 6, Fig. 2 depicts third head “F3” with multiple “DCM” wherein the one or more other operations further comprise a concatenation operation , the upsampled tensor calculated after the first “Upscale block” in third head “F3” is concatenated with other tensors at each “DCM” in third head “F3” to compute a tensor ). Ro and Lai are analogous to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the concatenation operation and upsampled tensor of Lai with the another intermediate tensor of Ro. The motivation to do so is to offer “a more efficient way to fuse features from multiple scales” (Lai, Abstract) . 07-21-aia AIA Claim s 8 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Ro in view of Li, and further in view of Rosewarne (US 20240242481 A1, hereinafter Rosewarne) . Regarding claim 8 : Regarding the limitation one or more non-transitory computer-readable media storing instructions executable to perform operations for training a neural network, the operations comprising: inputting a training dataset into the neural network , Ro teaches operations for training a neural network, the operations comprising: inputting a training dataset into the neural network (Page 1, Col. 1, Section 1, ¶1, Page 2, Col. 2, Section 3, ¶1, Page 5, Col. 2, Section IV, ¶1, Page 3, Col. 1, ¶1, Page 7, Col. 2, Section D, ¶1). However, Ro fails to teach one or more non-transitory computer-readable media storing instructions executable to perform operations … Rosewarne, in the same field of endeavor, teaches one or more non-transitory computer-readable media storing instructions executable to perform operations (Fig. 2A – 200, ¶81 “the software can also be loaded into the computer system 200 from other computer readable media. Computer readable storage media refers to any non-transitory tangible storage medium that provides recorded instructions and/or data to the computer system 200 for execution and/or processing”). Ro further teaches selecting a layer in a backbone of the neural network, the layer outputting an intermediate tensor based on the training dataset (Page 2, Col. 2, Section III, ¶1, Page 3, Col. 1, ¶1, Page 1, Fig. 1, Page 3, Col. 1, Section A, ¶1, Page 3, Col. 2, Section B, ¶2, Page 5, Fig. 3 – (b), Page 5, Col. 2, Section D, ¶2 all as described above with respect to claim 1 ). Ro further teaches inputting the intermediate tensor into a first head of the neural network, the first head comprising one or more deep learning operations that compute a first detection tensor (Page 1, Col. 1, Section I, ¶2, Page 2, Col. 1, ¶2, Page 1, Fig. 1, Page 5, Fig. 3 – (b) and Col. 2, Section D, ¶2 all as described above with respect to claim 1 ). Ro further teaches inputting the intermediate tensor into a second head of the neural network, the second head comprising the one or more deep learning operations that compute a second detection tensor that is different from the first detection tensor (Page 2, Col. 1, ¶2, Page 1, Fig. 1 , Page 5, Fig. 3 – (b) and Col. 2, Section D, ¶2 all as described above with respect to claim 1 ). Ro teaches determining a loss for the neural network (Page 2, Col. 1, ¶3) and the first detection tensor and the second detection tensor (Page 1, Fig. 1 – f1-f2, Page 2, Col. 1, ¶2). However, Ro fails to teach the loss comprising a diversity loss, the diversity loss indicating a measurement of similarity between the first detection tensor and the second detection tensor . Li, in the same field of endeavor, teaches the loss comprising a diversity loss, the diversity loss indicating a measurement of similarity between outputs of a first head and a second head (Page 3, Col. 1, Section 3.1, ¶1, Page 3, Col. 2, Equation (7) and ¶2, Page 5, Col. 1, Section 3.3, ¶1). Ro teaches training the neural network by adjusting one or more weights in the backbone, first head, or second head based on the loss (Page 6, Col. 1, Section B, ¶1). Ro, Li, and Rosewarne are analogous to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the non-transitory computer-readable media of Rosewarne and the diversity loss of Li with the teachings of Ro. The motivation to do so is to “further improve… model performance” (Li, Abstract) and to achieve high compression efficiency when encoding and decoding video and image signals (Rosewarne, ¶215 “The arrangements described are applicable to the computer and data processing industries and particularly for the digital signal processing for the encoding and decoding of signals such as video and image signals, achieving high compression efficiency”). Claims 9-14 recite a non-transitory computer-readable media that parallels the method claims of 2-7, respectively. Therefore, the analysis discussed above with respect to claims 2-7 applies to claims 9-14, respectively. Accordingly, claims 9-14 are rejected based on substantially the same rationale as set forth above with respect to claims 2-7, respectively. Regarding claim 15 : Ro teaches an apparatus, comprising: a computer processor for executing computer program instructions (Page 1, Col. 1, Section I, ¶1 “Deep metric learning has been successfully used in various applications related to computer vision , such as image retrieval , person re-identification, and face verification,” while Ro does not explicitly describe an apparatus, comprising: a computer processor for executing computer program instructions , one of ordinary skill in the art would recognize that an apparatus, comprising: a computer processor for executing computer program instructions is implicit when implementing “various applications related to computer vision” ). Regarding the limitation a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations for training a neural network, the operations comprising: inputting a training dataset into the neural network , Ro teaches operations for training a neural network, the operations comprising: inputting a training dataset into the neural network (Page 1, Col. 1, Section 1, ¶1, Page 2, Col. 2, Section 3, ¶1, Page 5, Col. 2, Section IV, ¶1, Page 3, Col. 1, ¶1, Page 7, Col. 2, Section D, ¶1). However, Ro fails to teach a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations … Rosewarne, in the same field of endeavor, teaches a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations (Fig. 2A – 200-201, 205, ¶75 “The computer module 201 typically includes at least one processor unit 205 ,” ¶81 “the software can also be loaded into the computer system 200 from other computer readable media. Computer readable storage media refers to any non-transitory tangible storage medium that provides recorded instructions and/or data to the computer system 200 for execution and/or processing,” ¶88 “the processor 205 is given a set of instructions which are executed therein”). Ro further teaches selecting a layer in a backbone of the neural network, the layer outputting an intermediate tensor based on the training dataset (Page 2, Col. 2, Section III, ¶1, Page 3, Col. 1, ¶1, Page 1, Fig. 1, Page 3, Col. 1, Section A, ¶1, Page 3, Col. 2, Section B, ¶2, Page 5, Fig. 3 – (b), Page 5, Col. 2, Section D, ¶2 all as described above with respect to claim 1 ). Ro further teaches inputting the intermediate tensor into a first head of the neural network, the first head comprising one or more deep learning operations that compute a first detection tensor (Page 1, Col. 1, Section I, ¶2, Page 2, Col. 1, ¶2, Page 1, Fig. 1, Page 5, Fig. 3 – (b) and Col. 2, Section D, ¶2 all as described above with respect to claim 1 ). Ro further teaches inputting the intermediate tensor into a second head of the neural network, the second head comprising the one or more deep learning operations that compute a second detection tensor that is different from the first detection tensor (Page 2, Col. 1, ¶2, Page 1, Fig. 1 , Page 5, Fig. 3 – (b) and Col. 2, Section D, ¶2 all as described above with respect to claim 1 ). Ro teaches determining a loss for the neural network (Page 2, Col. 1, ¶3) and the first detection tensor and the second detection tensor (Page 1, Fig. 1 – f1-f2, Page 2, Col. 1, ¶2). However, Ro fails to teach the loss comprising a diversity loss, the diversity loss indicating a measurement of similarity between the first detection tensor and the second detection tensor . Li, in the same field of endeavor, teaches the loss comprising a diversity loss, the diversity loss indicating a measurement of similarity between outputs of a first head and a second head (Page 3, Col. 1, Section 3.1, ¶1, Page 3, Col. 2, Equation (7) and ¶2, Page 5, Col. 1, Section 3.3, ¶1). Ro teaches training the neural network by adjusting one or more weights in the backbone, first head, or second head based on the loss (Page 6, Col. 1, Section B, ¶1). Ro, Li, and Rosewarne are analogous to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the non-transitory computer-readable memory of Rosewarne and the diversity loss of Li with the teachings of Ro. The motivation to do so is to “further improve… model performance” (Li, Abstract) and to achieve high compression efficiency when encoding and decoding video and image signals (Rosewarne, ¶215 “The arrangements described are applicable to the computer and data processing industries and particularly for the digital signal processing for the encoding and decoding of signals such as video and image signals, achieving high compression efficiency”). Claims 16-20 recite an apparatus that parallels the method claims of 2 and 4-7, respectively. Therefore, the analysis discussed above with respect to claims 2 and 4-7 applies to claims 9-14, respectively. Accordingly, claims 9-14 are rejected based on substantially the same rationale as set forth above with respect to claims 2 and 4-7, respectively. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to WILLIAM M LEE whose telephone number is (571)272-4761. The examiner can normally be reached Mon-Fri. 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571)272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WILLIAM MICHAEL LEE/ Examiner, Art Unit 2145 /CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145 Application/Control Number: 18/457,002 Page 2 Art Unit: 2145 Application/Control Number: 18/457,002 Page 3 Art Unit: 2145 Application/Control Number: 18/457,002 Page 4 Art Unit: 2145 Application/Control Number: 18/457,002 Page 5 Art Unit: 2145 Application/Control Number: 18/457,002 Page 6 Art Unit: 2145 Application/Control Number: 18/457,002 Page 7 Art Unit: 2145 Application/Control Number: 18/457,002 Page 8 Art Unit: 2145 Application/Control Number: 18/457,002 Page 9 Art Unit: 2145 Application/Control Number: 18/457,002 Page 10 Art Unit: 2145 Application/Control Number: 18/457,002 Page 11 Art Unit: 2145 Application/Control Number: 18/457,002 Page 12 Art Unit: 2145 Application/Control Number: 18/457,002 Page 13 Art Unit: 2145 Application/Control Number: 18/457,002 Page 14 Art Unit: 2145 Application/Control Number: 18/457,002 Page 15 Art Unit: 2145 Application/Control Number: 18/457,002 Page 16 Art Unit: 2145 Application/Control Number: 18/457,002 Page 17 Art Unit: 2145 Application/Control Number: 18/457,002 Page 18 Art Unit: 2145 Application/Control Number: 18/457,002 Page 19 Art Unit: 2145 Application/Control Number: 18/457,002 Page 20 Art Unit: 2145 Application/Control Number: 18/457,002 Page 21 Art Unit: 2145 Application/Control Number: 18/457,002 Page 22 Art Unit: 2145 Application/Control Number: 18/457,002 Page 23 Art Unit: 2145 Application/Control Number: 18/457,002 Page 24 Art Unit: 2145 Application/Control Number: 18/457,002 Page 25 Art Unit: 2145 Application/Control Number: 18/457,002 Page 26 Art Unit: 2145 Application/Control Number: 18/457,002 Page 27 Art Unit: 2145 Application/Control Number: 18/457,002 Page 28 Art Unit: 2145 Application/Control Number: 18/457,002 Page 29 Art Unit: 2145 Application/Control Number: 18/457,002 Page 30 Art Unit: 2145 Application/Control Number: 18/457,002 Page 31 Art Unit: 2145 Application/Control Number: 18/457,002 Page 32 Art Unit: 2145 Application/Control Number: 18/457,002 Page 33 Art Unit: 2145