DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response to amendments filed October 30th, 2025. The status of the claims is as follows. Claims 1, 6, 10 and 15 are amended and Claims 5 and 14 are cancelled. Claims 1-4, 6-13 and 15-19 are currently pending.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-2, 6, 9; 10, 11, 15, 18; 19 are rejected under 35 U.S.C. 103 as being unpatentable over Cai et al. (US20200279156A1, hereinafter “Cai”) in view of Zhao et al. (US20200356802A1, hereinafter “Zhao”).
Regarding Claim 1,
Cai discloses An electronic apparatus for performing a preset task by using a deep neural network (DNN) (Cai [0069]; “Example 3 may include the elements of example 2, wherein the deep learning models include at least one of convolutional neural networks (CNNs) or recurrent neural networks (RNNs).”
Cai [0067]; “According to example 1, there is provided an apparatus capable of fusing features from a plurality of data modalities. The apparatus may comprise a processor, network interface circuitry to receive a sample dataset and a test dataset, training logic to determine one or more handcrafted features for each data modality based on the sample dataset, train a handcrafted model for each data modality based on the corresponding handcrafted features and the sample dataset, and train a plurality of deep learning model sets based on the sample dataset and the handcrafted models, the training including feed-forward training based on the sample dataset, determining error information, and updating parameters of the deep learning model sets based on the determined error information, and runtime prediction logic to predict a label based on the deep learning model sets, the handcrafted models, and the test dataset.” which includes an apparatus implemented through deep learning models read on as a DNN implemented apparatus)
the electronic apparatus comprising: an input interface configured to receive input data of a first type and input data of a second type (Cai [0067]; “According to example 1, there is provided an apparatus capable of fusing features from a plurality of data modalities. The apparatus may comprise a processor, network interface circuitry to receive a sample dataset and a test dataset, training logic to determine one or more handcrafted features for each data modality based on the sample dataset, train a handcrafted model for each data modality based on the corresponding handcrafted features and the sample dataset, and train a plurality of deep learning model sets based on the sample dataset and the handcrafted models, the training including feed-forward training based on the sample dataset, determining error information, and updating parameters of the deep learning model sets based on the determined error information, and runtime prediction logic to predict a label based on the deep learning model sets, the handcrafted models, and the test dataset.” which includes an interface for receiving sample datasets read on as inputs of data types)
a memory storing one or more instructions; and a processor configured to execute the one or more instructions stored in the memory (Cai [0066]; “The following examples of the present disclosure may comprise subject material such as an apparatus, a method, at least one machine-readable medium for storing instructions that when executed cause a machine to perform acts based on the method, means for performing acts based on the method and/or a system to integrate correlated cues from a multi-modal data set. According to example 1, there is provided an apparatus capable of fusing features from a plurality of data modalities. The apparatus may comprise a processor, network interface circuitry to receive a sample dataset and a test dataset, training logic to determine one or more handcrafted features for each data modality based on the sample dataset, train a handcrafted model for each data modality based on the corresponding handcrafted features and the sample dataset, and train a plurality of deep learning model sets based on the sample dataset and the handcrafted models, the training including feed-forward training based on the sample dataset, determining error information, and updating parameters of the deep learning model sets based on the determined error information, and runtime prediction logic to predict a label based on the deep learning model sets, the handcrafted models, and the test dataset”)
obtain first sub-feature information corresponding to the input data of the first type and second sub-feature information corresponding to the input data of the second type (Cai [Fig. 5];
PNG
media_image1.png
134
406
media_image1.png
Greyscale
)
obtain feature information from each of a plurality of layers of the DNN by inputting the first sub-feature information and the second sub-feature information into the DNN (Cai [Fig. 6];
PNG
media_image2.png
134
414
media_image2.png
Greyscale
Cai [0036];
“Operations also include sending the test data to deep learning and handcrafted models 604. These models may include models trained via training logic 408, as described in FIG. 5, to output feature vectors.” wherein the feature vectors output by the plurality of deep learning and handcrafted model layers obtained through inputting multi-modal test data reads on obtaining feature information from each of a plurality of layers of the DNN by inputting first and second sub-feature information)
calculate a weight for each type corresponding to each of the plurality of layers, based on the feature information, the first sub-feature information, and the second sub-feature information (Cai [0031]; “Training logic 408 is further configured to perform or cause feed-forward training and back-propagation parameter revision of deep learning models 114 and 134. For each data modality, the feed-forward phase may include, for example, passing sample data through deep learning models (e.g., 204A-204N) to produce feature vectors (e.g., 210A-210N), concatenating the feature vectors (e.g., 210A-210N), and passing the concatenated vectors to early an abstraction layer (e.g., 140A). Training logic 408 may further concatenate the output of each early abstraction layer (e.g., 140A, 140B, etc.) and each feature vector (e.g., 220) output from handcrafted model 118, passing the concatenated output to late abstraction layer 150. Training logic 408 may additionally determine an output feature vector 160 based on the late abstraction layer 150.” wherein for each data type, feed-forward training conducted based on concatenated feature vectors reads on calculating weights for each type corresponding to each of the layers based on the feature information and its associated first and second sub-feature information)
obtain a final output value corresponding to the preset task by applying the weight for each type, in each of the plurality of layers (Cai [0036]; ”Operations further include determining an output vector of predictions 614. This may include a weighted prediction value corresponding to each of a plurality of options (e.g., emotions expressed in a sample video clip).” wherein the prediction vector weighted for each of the plurality of options is read as a final output value corresponding to the preset task of prediction that is obtained by applying weights for different types)
Cai fails to explicitly disclose but Zhao discloses obtain first query information corresponding to each of the plurality of layers, based on the first sub-feature information and a pre-trained query matrix corresponding to each of the plurality of layers wherein the first query information indicates a weight of the first sub-feature information (Zhao [0061]; “Therefore, in the embodiments of the present application, the feature map is processed by using two branches separately, to obtain a first weight vector with respect to the inward reception weights of each of the multiple feature points included in the feature map, and a second weight vector with respect to the outward transmission weights of at least one of the multiple feature points” wherein the first weight vector is read as first query information obtained through the first sub-feature information (one of multiple feature points) as well as a pre-trained query matrix (feature map))
obtain second query information corresponding to each of the plurality of layers, based on the second sub-feature information and the pre-trained query matrix, wherein the second query information indicates a weight of the second sub-feature information (Zhao [0061]; “Therefore, in the embodiments of the present application, the feature map is processed by using two branches separately, to obtain a first weight vector with respect to the inward reception weights of each of the multiple feature points included in the feature map, and a second weight vector with respect to the outward transmission weights of at least one of the multiple feature points” wherein the second weight vector is read as second query information obtained through the second sub-feature information (one of multiple feature points) as well as a pre-trained query matrix (feature map))
wherein the pre-trained query matrix comprises parameters related to the first sub-feature information and the second sub-feature information (Zhao [0050]; “feature extraction is performed on a to-be-processed image to generate a feature map of the image, a feature weight corresponding to each of multiple feature points included in the feature map is determined” wherein each of multiple feature points included in the feature map reads on the pre-trained query matrix comprising parameters of the extracted sub-features’ information.
Zhao [0148]; “The feature extraction network involved in the embodiments can be pre-trained or untrained” wherein the feature map understood to be pre-trained reads on a pre-trained query matrix)
It would have been obvious to use Zhao’s method of feature analysis through extracted weight features on the feature information output by the neural network output of Cai. One would have been motivated to obtain weight vectors because: “in order to obtain comprehensive weight information corresponding to each feature point in the feature map, it is necessary to obtain weights used by the feature point to transmit information to surrounding locations” (Zhao [0093]).
Regarding Claim 2,
Cai/Zhao teaches the method of Claim 1 (and thus the rejection of Claim 1 is incorporated). Cai/Zhao further discloses to obtain the first sub-feature information by inputting the input data of the first type into a pre-trained first sub-network; and obtain the second sub-feature information by inputting the input data of the second type into a pre-trained second sub-network (Cai [0016]; “System 100 generally receives video input 110 and audio input 132, trains itself based on sample data, extracts features from the inputs using both deep learning models (e.g., deep learning based video models 114 and deep learning based audio models)
Regarding Claim 6,
The combination of Cai/Zhao teaches the method of Claim 1 (and thus the rejection of Claim 1 is incorporated). The combination of Cai/Zhao fails to explicitly disclose but Zhao discloses to obtain key information corresponding to each of the plurality of layers, based on the feature information extracted from each of the plurality of layers and a pre-trained key matrix corresponding to each of the plurality of layers (Zhao [0207]; “In the embodiments, feature information received by a feature point in the feature map is obtained by using the first weight vector and the feature map, and feature information transmitted by a feature point in the feature map is obtained by using the second weight vector and the feature map. That is, feature information of bi-direction transmission is obtained. The enhanced feature map including more information can be obtained based on the feature information of bi-direction transmission and the original feature map.” wherein an enhanced feature map derived through the first and second weight vectors as well as the feature map reads on key information based on extracted feature information and a pre-trained key matrix)
It would have been obvious to use Zhao’s method of obtaining key information through feature analysis of extracted weight features on the feature information output by the neural network output of Cai/Zhao. One would have been motivated to obtain key information (enhanced feature map) through the feature information weights and feature map key matrix because by obtaining a feature-enhanced feature map, “context information can be better used, and the feature-enhanced feature map includes more information” (Zhao [0093]).
Regarding Claim 9,
Cai/Zhao teaches the method of Claim 1 (and thus the rejection of Claim 1 is incorporated). Cai/Zhao further discloses wherein the input data of the first type and the input data of the second type comprise at least one of image data, text data, sound data, or video data (Cai [0016]; “System 100 generally receives video input 110 and audio input 132, trains itself based on sample data, extracts features from the inputs using both deep learning models (e.g., deep learning based video models 114 and deep learning based audio models 134) and “handcrafted” models (e.g., handcrafted video models 118 and handcrafted audio models 138).”)
Claims 10-11, 15, 18 recite the same method performed by the electronic apparatus of Claims 1-2, 6, 9. Thus, Claims 10-11, 15, 18 are rejected for reasons set forth in the rejection of Claims 1-2, 6, 9, respectively.
Claim 19 recites a non-transitory medium storing instructions to perform the same method of Claim 10. Thus, Claim 19 is rejected for reasons set forth in the rejection of Claim 10.
Claims 3-4; 12-13 are rejected under 35 U.S.C. 103 as being unpatentable over Cai et al. (US20200279156A1, hereinafter “Cai”) in view of Zhao et al. (US20200356802A1, hereinafter “Zhao”) in view of Chan et al. (US20210406266A1, hereinafter “Chan”).
Regarding Claim 3,
Cai/Zhao teaches the method of Claim 1 (and thus the rejection of Claim 1 is incorporated). Cai/Zhao further discloses to input the … first sub-feature information and the … second sub- feature information to the DNN (Cai [0016]; “System 100 generally receives video input 110 and audio input 132, trains itself based on sample data, extracts features from the inputs using both deep learning models (e.g., deep learning based video models 114 and deep learning based audio models)
Cai/Zhao fails to explicitly disclose but Chan teaches to encode, based on type identification information that distinguishes a type of the input data, the first sub-feature information and the second sub-feature information (Chan [0027]; “Particular embodiments derive (e.g., receive or generate) a feature vector that represents the set of features of the first cell based on the extracting of the set of features. A “feature vector” (also referred to as a “vector”) as described herein includes one or more real numbers, such as a series of floating values or integers (e.g., [0, 1, 0, 0]) that represent one or more other real numbers, a natural language (e.g., English) word and/or other character sequence (e.g., a symbol (e.g., @, !, #), a phrase, and/or sentence, etc.). Such natural language words and/or character sequences correspond to the set of features and are encoded or converted into corresponding feature vectors so that computers can process the corresponding extracted features.” wherein the feature vectors of an integer type as well as language type read on first and second sub-feature information
Chan [0092]; “The feature the feature vector 404 represents a shape vector. A “shape vector” is a feature vector that indicates the type or class of one or more characters of a particular element. For instance, the shape vector can indicate whether a character is a letter, number (e.g., an integer or other real number), and/or a symbol or type of symbol (e.g., picture, exclamation point, question mark, etc.).” wherein the shape vector indicating the type of the input data reads on type identification information
Chan [0028]; “Particular embodiments of the present disclosure model or process these data by sequentially encoding particular feature vectors based on using one or more machine learning models. For example, some embodiments convert each feature vector of a row in a table from left to right in an ordered fashion into another concatenated feature vector. In some embodiments, such sequential encoding includes using a 1-dimensional and/or 2-dimensional bi-directional Long Short Term Memory (LSTM) model to encode sequential data into a concatenated or aggregated feature vector of multiple values representing multiple cells in a table.” wherein encoding is performed through concatenation of the feature vector information)
Cai/Zhao reads on input[ting] the … first sub-feature information and the … second sub- feature information to the DNN. Although Cai/Zhao doesn’t read on input[ting] the encoded first sub-feature information and the encoded second sub- feature information to the DNN, Chan reads on the encoded first and second sub-feature. Therefore, by using Chan’s encoded sub-features in the method of Cai/Zhao, the combination reads on input[ting] the encoded first sub-feature information and the encoded second sub- feature information to the DNN
It would have been obvious to use Chan’s encoded sub-features in Cai/Zhao’s method of inputting features into the DNN. One would have been motivated to encode the inputted features because: “Such natural language words and/or character sequences correspond to the set of features and are encoded or converted into corresponding feature vectors so that computers can process the corresponding extracted features” (Chan [0027]).
Regarding Claim 4,
The combination of Cai/Zhao/Chan teaches the method of Claim 3 (and thus the rejection of Claim 3 is incorporated). The combination already discloses to encode the first sub-feature information and the second sub-feature information by concatenating the first sub-feature information and the second sub- feature information (Chan [0028]; “Particular embodiments of the present disclosure model or process these data by sequentially encoding particular feature vectors based on using one or more machine learning models. For example, some embodiments convert each feature vector of a row in a table from left to right in an ordered fashion into another concatenated feature vector. In some embodiments, such sequential encoding includes using a 1-dimensional and/or 2-dimensional bi-directional Long Short Term Memory (LSTM) model to encode sequential data into a concatenated or aggregated feature vector of multiple values representing multiple cells in a table.” wherein encoding is performed through concatenation of sequential feature vectors thus reading on encoding the first and second sub-feature information through concatenation)
Claims 12-13 recite the same method performed by the electronic apparatus of Claims 3-4. Thus, Claims 12-13 are rejected for reasons set forth in the rejection of Claims 3-4, respectively.
Claims 7-8; 16-17 are rejected under 35 U.S.C. 103 as being unpatentable over Cai et al. (US20200279156A1, hereinafter “Cai”) in view of Zhao et al. (US20200356802A1, hereinafter “Zhao”) and further in view of Wang et al. (US20220343638A1, hereinafter “Wang”).
Regarding Claim 7,
The combination of Cai/Zhao teaches the method of Claim 6 (and thus the rejection of Claim 6 is incorporated). Cai/Zhao fails to explicitly disclose but Wang teaches to obtain first context information corresponding to each of the plurality of layers, the first context information indicating a correlation between the first query information and the key information; obtain second context information corresponding to each of the plurality of layers, the second context information indicating a correlation between the second query information and the key information (Wang [0048]; “perform channel dimension reduction on the first feature map through the second-order pooling module in the classifier model to obtain a dimension-reduced second feature map”Wang [0114]; “S10323, calculate a weight vector corresponding to the second feature map. The terminal calculates the weight vector corresponding to the second feature map through the trained classifier model. Please refer to FIG. 3, specifically, covariance information of each two channels between different channels in the dimension-reduced second feature map is calculated to obtain a covariance matrix; and the weight vector having the same channel number with that of the four-dimensional feature map is acquired through a grouping convolution and the 1×1×1 convolution according to the covariance matrix.” wherein the covariance matrix is representative of the magnitude of correlation between the feature map’s values and weights across channels, thus interpreted as context information indicating a correlation between first/second query information and key information)
It would have been obvious to perform Wang’s method to determine first and second context information representing the correlation between query information (Cai/Zhao’s first and second feature type weights) and key information (Cai/Zhao’s feature map) in Cai/Zhao’s apparatus for multimodal sub-feature extraction and evaluation because “correlation information between different channels of high-order features, makes the weight of the important feature channel larger and the weight of the unimportant feature channel smaller” (Wang [0169])
Regarding Claim 8,
The combination of Cai/Zhao/Wang teaches the method of Claim 7 (and thus the rejection of Claim 7 is incorporated). The combination already discloses to calculate the weight for each type corresponding to each of the plurality of layers, based on the first context information and the second context information corresponding to each of the plurality of layers (Wang [0114]; “S10323, calculate a weight vector corresponding to the second feature map. The terminal calculates the weight vector corresponding to the second feature map through the trained classifier model. Please refer to FIG. 3, specifically, covariance information of each two channels between different channels in the dimension-reduced second feature map is calculated to obtain a covariance matrix; and the weight vector having the same channel number with that of the four-dimensional feature map is acquired through a grouping convolution and the 1×1×1 convolution according to the covariance matrix.” wherein the weight vector corresponding to the covariance matrix reads on calculating the weight for multiple features based on first and second context information)
Claims 16-17 recite the same method performed by the electronic apparatus of Claims 7-8. Thus, Claims 16-17 are rejected for reasons set forth in the rejection of Claims 7-8, respectively.
Response to Arguments
The Examiner acknowledges the Applicant’s amendments to Claims 1, 6, 10 and 15.
Applicant’s arguments filed October 30th, 2025, traversing the rejection of claims 1-19 under 35 U.S.C. § 101 have been fully considered, and are found to be fully persuasive.
Applicant’s arguments regarding the 35 U.S.C. § 103 rejection of Claims 1-19 have been considered, but are not fully persuasive.
Applicant alleges on Page 17-19 of the Remarks that Claims 1, 10 have been amended so that the combination of Cai/Zhao does not teach all of the limitations of the newly amended Claims 1, 10.
The amended limitations are described below:
obtain first query information corresponding to each of the plurality of layers, based on the first sub-feature information and a pre-trained query matrix corresponding to each of the plurality of layers wherein the first query information indicates a weight of the first sub-feature information
obtain second query information corresponding to each of the plurality of layers, based on the second sub-feature information and the pre-trained query matrix, wherein the second query information indicates a weight of the second sub-feature information
wherein the pre-trained query matrix comprises parameters related to the first sub-feature information and the second sub-feature information
Applicant alleges that Zhao is generally directed to a "Point-Wise Spatial Attention (PSA) solution" in which "each feature point in the feature map can not only collect information about other points to help the prediction of the current point, but also distribute information about the current point to help the prediction of other points." Zhao, [0052]. In particular, Zhao discloses that a "feature map is processed by using two branches separately, to obtain a first weight vector with respect to the inward reception weights of each of the multiple feature points included in the feature map, and a second weight vector with respect to the outward transmission weights of at least one of the multiple feature points." [0061]. Zhao further discloses that "feature extraction is performed on a to-be-processed image to generate a feature map of the image, a feature weight corresponding to each of multiple feature points included in the feature map is determined" [0050]. According to Zhao, the "feature extraction network involved in the embodiments can be pre-trained or untrained." [0148].
That is, at best, Zhao discloses a point-wise spatial attention (PSA) mechanism for processing a single feature map in which its first and second weight vectors are for inward reception weights and outward transmission weights, respectively, and thus, relate to spatial relationships within the single feature map.
To the extent the Office is attempting to equate the first weight vector of Zhao to the claimed "first query information", the second weight vector of Zhao to the claimed "second query information", and the untrained feature map of Zhao to the claimed "pre-trained query matrix", claim 1 recites to "obtain first query information corresponding to each of the plurality of layers, based on the first sub-feature information and a pre-trained query matrix corresponding to each of the plurality of layers, wherein the first query information indicates a weight of the first sub-feature information [corresponding to the input data of the first type]", to "obtain second query information corresponding to each of the plurality of layers, based on the second sub-feature information and the pre-trained query matrix, wherein the second query information indicates a weight of the second sub-feature information [corresponding to the input data of the second type]", and that "the pre-trained query matrix comprises parameters related to the first sub-feature information [corresponding to the input data of the first type] and the second sub-feature information [corresponding to the input data of the second type]". Zhao does not disclose that its first and second weight vectors or its feature map correspond to any such input data. Therefore, Zhao's purported disclosure of first and second weight vectors and a feature map does not disclose the claimed "first query information", the claimed "second query information", or the claimed "pre-trained query matrix".
That is, Zhao is generally directed to a PSA mechanism for processing a single feature map, whereas claim 1 recites a modal or type-based attention mechanism that processes two distinct inputs (e.g., the claimed "first sub-feature information" and the claimed "second sub- feature information") against a single common "pre-trained query matrix" to determine the relative importance of each data type. Thus, Zhao's spatial attention mechanism for a single input is not interchangeable with the claimed modal attention mechanism for multiple, different inputs. Consequently, Zhao is unable to disclose or suggest at least the cited features recited in claim 1.
Examiner respectfully disagrees. The combination of Cai/Zhao discloses obtain first query information corresponding to each of the plurality of layers, based on the first sub-feature information and a pre-trained query matrix corresponding to each of the plurality of layers wherein the first query information indicates a weight of the first sub-feature information (Zhao [0061]; “Therefore, in the embodiments of the present application, the feature map is processed by using two branches separately, to obtain a first weight vector with respect to the inward reception weights of each of the multiple feature points included in the feature map, and a second weight vector with respect to the outward transmission weights of at least one of the multiple feature points” wherein the first weight vector is read as first query information obtained through the first sub-feature information (one of multiple feature points) as well as a pre-trained query matrix (feature map))
obtain second query information corresponding to each of the plurality of layers, based on the second sub-feature information and the pre-trained query matrix, wherein the second query information indicates a weight of the second sub-feature information (Zhao [0061]; “Therefore, in the embodiments of the present application, the feature map is processed by using two branches separately, to obtain a first weight vector with respect to the inward reception weights of each of the multiple feature points included in the feature map, and a second weight vector with respect to the outward transmission weights of at least one of the multiple feature points” wherein the second weight vector is read as second query information obtained through the second sub-feature information (one of multiple feature points) as well as a pre-trained query matrix (feature map))
wherein the pre-trained query matrix comprises parameters related to the first sub-feature information and the second sub-feature information (Zhao [0050]; “feature extraction is performed on a to-be-processed image to generate a feature map of the image, a feature weight corresponding to each of multiple feature points included in the feature map is determined” wherein each of multiple feature points included in the feature map reads on the pre-trained query matrix comprising parameters of the extracted sub-features’ information.
Zhao [0148]; “The feature extraction network involved in the embodiments can be pre-trained or untrained” wherein the feature map understood to be pre-trained reads on a pre-trained query matrix)
Although applicant’s recited invention may prove to overcome the existing Cai/Zhao combination of prior art, the broadest reasonable interpretation of existing amended claim language recites a broader scope than applicant’s intentions. Applicant’s claim language makes no mention that the plurality of first and second sub-feature information cannot be derived from Zhao’s single feature map. The claim language simply recites that first and second sub-feature information corresponding to input data of a first and second type need be obtained. Additionally, examiner notes that “input data of a first type and input data of a second type” can, under broadest reasonable interpretation, both be interpreted under Zhao’s single feature map’s weight vectors of spatial data types since there is no mention that the two data types must be different from one another. In the case of Zhao, the singular feature map of Zhao containing weight vectors of spatial data types thus is interpretable under the broadest reasonable interpretation of Claim 1. Thus, Zhao’s spatial attention mechanism for a single input is interpreted to be interchangeable with the claimed attention mechanism.
Applicant alleges that the Office does not establish a prima facie case of obviousness at least because the Office does not provide an adequate explanation as to why it would have been obvious to one of ordinary skill in the art to have combined the cited references to purportedly arrive at the claimed invention. An obviousness rejection requires more than just a capability to combine teachings from multiple prior art references; instead, there must be a rationale, a reason, why the particular combination would have been predictable. See MPEP §§ 2143.01(III)-(IV). In addition, even if the Office is able to provide some reasoning, the Federal Circuit has held that it is not obvious to combine references when "the proposed modifications would merely have been expected to have the same functional properties as the prior art product." See MPEP § 2143(I)(A) Example 3 (citing In re Omeprazole Patent Litigation, 536 F.3d 1361, 8719.
Examiner respectfully disagrees. The rationale behind going through the extra work and expense to take Zhao's spatial attention mechanism (for processing a single feature map) and applying it to Cai's different problem of fusing two different sub-feature types using a query matrix with a result that purportedly provides the same functionality as the original lies in how comprehensive weight information is obtained. That is, previously incomprehensible or difficult to understand sub-feature type information within the query matrix may undergo analysis to prove easier for one to process. As such, a prima facie case of obviousness is established.
The rejection of Claim 1 under 35 U.S.C. § 103 has been maintained. Similarly, the rejection of Claim 10 under 35 U.S.C. § 103 has been maintained.
The rejection of Claims 2-4 and 6-9, under 35 U.S.C. § 103, which depend directly or indirectly from Claim 1, have been maintained. The rejection of Claims 11-13 and 15-19, under 35 U.S.C. § 103, which depend directly or indirectly from Claim 10, have been maintained.
Conclusion
Applicant’s amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JONATHAN J KIM whose telephone number is (571)272-0523. The examiner can normally be reached 8-6.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matt Ell can be reached on (571) 270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JONATHAN J KIM/Examiner, Art Unit 2141
/MATTHEW ELL/Supervisory Patent Examiner, Art Unit 2141