DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response to a request for continued examination filed May 13th, 2026, in which claims 1, 3, 13-14 and 16 have been amended, no claims have been added, and claims 2 and 5 have been cancelled. The amendments have been entered, and claims 1, 3, 7, 9-11, 13-14, and 16-17 are currently pending in the case. Claims 1, 13, and 14 are independent claims.
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on May 13th, 2026 has been entered.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1, 3, 7, 9-11, 13-14, and 16-17 are rejected under 35 U.S.C. § 112(b) or 35 U.S.C. § 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1, 13 and 14 recites the limitation "indicating an importance of each of the plurality of modalities to attribute prediction". There is insufficient antecedent basis for the term “attribute prediction” in the claim. For examination purposes, this limitation has been interpreted as “indicating an importance of each of the plurality of modalities to an attribute prediction”.
Claims 3, 7, 9-11, and 16-17 are rejected for being dependent on a rejected base claim without curing any of the deficiencies.
Claim Rejections - 35 USC § 101
35 U.S.C. § 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1, 3, 7, 9-11, 13-14, and 16-17 are rejected under 35 U.S.C. § 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1:
Step 1: Claim 1 is directed to an apparatus, therefore it falls under the statuary category of a machine.
Step 2A Prong 1: The claim recites, in part:
“apply the plurality of modalities to…generate feature values for each of the plurality of modalities” this limitation is a mathematical concept.
“apply the feature values output from the fully connected network to…derive attention weights corresponding to each of the plurality of modalities based on the feature values for each of the plurality of modalities and the information identifying the object, the attention weights indicating an importance of each of the plurality of modalities to attribute prediction” this encompasses the mental derivation of weights for modalities based on observed feature values and observed identifying information, the attention weights indicating an observed importance of the modalities to attribute prediction.
“predict an attribute of the object from a concatenated value of the feature values for each of the plurality of modalities, weighted by the corresponding weights” this encompasses the mental prediction of attributes of an object from observed feature values, and weighting that prediction by observed weights. Further, this limitation is a mathematical concept.
“for each feature value of the feature values, multiply the feature value by a corresponding attention weight to generate a weighted feature value of a plurality of weighted feature values, concatenate the plurality of weighted feature values to generate a concatenated value” this limitation is a mathematical concept.
“apply the concatenated value to…predict at least one attribute of the object based on the concatenated value” this encompasses the mental prediction of attributes based on observed values. Further, this limitation is a mathematical concept.
Step 2A Prong 2: The judicial exception is not integrated into a practical application; the remaining limitations of the claim are as follows: “at least one processor configured to operate as instructed by the program code”, “acquisition code configured to cause the at least one processor to, “feature generation code configured to cause the at least one processor to”, “weight derivation code configured to cause the at least one processor to”, “classification code configured to cause the at least one processor to”, “the fully connected network is configured to”, “the neural network is configured to”, “a deep neural network configured to” the limitations are an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP § 2106.05(f)(2). “acquire a plurality of modalities associated with an object and information identifying the object, wherein the plurality of modalities include image data corresponding to an image of the object and text data describing the object”, “receive the feature values for each of the plurality of modalities and information identifying the object” these limitations are an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g). “at least one memory configured to store program code”, “a fully connected network”, “a neural network” these limitations are an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP § 2106.05(h).
Step 2B: The claim does not contain significantly more than the judicial exception. The limitations “at least one processor configured to operate as instructed by the program code”, “acquisition code configured to cause the at least one processor to, “feature generation code configured to cause the at least one processor to”, “weight derivation code configured to cause the at least one processor to”, “classification code configured to cause the at least one processor to”, “the fully connected network is configured to”, “the neural network is configured to”, “a deep neural network configured to” the limitations are an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP § 2106.05(f)(2). “acquire a plurality of modalities associated with an object and information identifying the object, wherein the plurality of modalities include image data corresponding to an image of the object and text data describing the object”, “receive the feature values for each of the plurality of modalities and information identifying the object” these limitations are an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g). Furthermore the additional element is directed to receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d as well as storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). See MPEP § 2106.05(d)/(II). . “at least one memory configured to store program code”, “a fully connected network”, “a neural network” these limitations are an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP § 2106.05(h).
Regarding claim 3, the rejection of claim 1 is incorporated and further:
Step 2A Prong 1: The claim recites, in part: “the attention weights of the plurality of modalities are 1 in total” a continuation of the abstract idea identified in the parent claim.
Step 2A Prong 2: The claim does not recite any additional limitations, thus does not further recite any additional elements that integrates the judicial exception into a practical application or amount to significantly more.
Regarding claim 7, the rejection of claim 4 is incorporated and further:
Step 2A Prong 1: The claim recites, in part:
“as input, the plurality of modalities and outputs the feature values of the plurality of modalities by mapping them to a latent space common to the plurality of modalities” this limitation is a mathematical concept.
Step 2A Prong 2: The judicial exception is not integrated into a practical application; the remaining limitations of the claim are as follows: “the fully connected network takes” the limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h).
Step 2B: The claim does not contain significantly more than the judicial exception. The limitations “the fully connected network takes” the limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h).
Regarding claim 9, the rejection of claim 1 is incorporated and further:
Step 2A Prong 1: The claim recites, in part:
“encode the plurality of modalities to acquire a plurality of encoded modalities” this limitation is a mathematical concept.
“generate the feature values for each of the plurality of encoded modalities” This limitation is a mathematical concept.
Step 2A Prong 2: The judicial exception is not integrated into a practical application; the remaining limitations of the claim are as follows: “the acquisition code is further configured to cause the at least one processor to”, “the feature generation fully connected network is further configured to” the limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h).
Step 2B: The claim does not contain significantly more than the judicial exception. The limitations “the acquisition code is further configured to cause the at least one processor to”, “the feature generation fully connected network is further configured to” the limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h).
Regarding claim 10, the rejection of claim 1 is incorporated and further:
Step 2A Prong 1: The claim recites, in part:
“the object is a commodity, and the plurality of modalities includes two or more of data of an image representing the commodity, data of text describing the commodity, and data of sound describing the commodity” a continuation of the abstract idea identified in the parent claim.
Step 2A Prong 2: The claim does not recite any additional limitations, thus does not further recite any additional elements that integrates the judicial exception into a practical application or amount to significantly more.
Regarding claim 11, the rejection of claim 1 is incorporated and further:
Step 2A Prong 1: The claim recites, in part:
“the attribute of the object includes color information of a product” a continuation of the abstract idea identified in the parent claim.
Step 2A Prong 2: The claim does not recite any additional limitations, thus does not further recite any additional elements that integrates the judicial exception into a practical application or amount to significantly more.
Regarding claim 13:
Step 1: Claim 13 is directed to a method, therefore it falls under the statuary category of a process.
Step 2A Prong 1: The claim recites, in part:
“applying the plurality of modalities to…generate feature values for each of the plurality of modalities” this limitation is a mathematical concept.
“applying the feature values output from the fully connected network to…derive attention weights corresponding to each of the plurality of feature values” this encompasses the mental derivation of weights for modalities based on observed feature values and observed identifying information.
“for each feature value of the feature values, multiply the feature value by a corresponding attention weight to generate a weighted feature value of a plurality of weighted feature values, concatenate the plurality of weighted feature values to generate a concatenated value” this limitation is a mathematical concept.
“applying the concatenated value to…predict at least one attribute of the object based on the concatenated value” this encompasses the mental prediction of attributes based on observed values. Further, this limitation is a mathematical concept.
Step 2A Prong 2: The judicial exception is not integrated into a practical application; the remaining limitations of the claim are as follows: “a fully connected network, wherein the fully connected network is configured to”, “a neural network configured to”, “a deep neural network configured to” the limitations are an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP § 2106.05(f)(2). “acquiring a plurality of modalities associated with an object and information identifying the object, wherein the plurality of modalities includes image data corresponding to an image of the object and text data describing the object”, “receive the feature values for each of the plurality of modalities and information identifying the object” these limitations are an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g).
Step 2B: The claim does not contain significantly more than the judicial exception. The limitations “a neural network configured to”, “a neural network configured to”, “a deep neural network configured to” the limitations are an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP § 2106.05(f)(2). “acquiring a plurality of modalities associated with an object and information identifying the object, wherein the plurality of modalities includes image data corresponding to an image of the object and text data describing the object”, “receive the feature values for each of the plurality of modalities and information identifying the object” these limitations are an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g). Furthermore the additional element is directed to receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d as well as to storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). See MPEP § 2106.05(d)/(II).
Regarding claim 14:
Step 1: Claim 14 is directed to a non-transitory computer-readable storage medium, therefore it falls under the statuary category of a manufacture.
Step 2A Prong 1: The claim recites, in part:
“applying the feature values output from the fully connected network to…generate feature values for each of the plurality of modalities” this limitation is a mathematical concept.
“applying the feature values output from the fully connected network to…derive attention weights corresponding to each of the plurality of modalities based on the feature values for each of the plurality of modalities and the information identifying the object, the attention weights indicating an importance of each of the plurality of modalities to attribute prediction” this encompasses the mental derivation of weights for modalities based on observed feature values and observed identifying information.
“for each feature value of the feature values, multiply the feature value by a corresponding attention weight to generate a weighted feature value of a plurality of weighted feature values, concatenate the plurality of weighted feature values to generate a concatenated value” this limitation is a mathematical concept.
“applying the concatenated value to…predict at least one attribute of the object based on the concatenated value” this encompasses the mental prediction of attributes based on observed values. Further, this limitation is a mathematical concept.
Step 2A Prong 2: The judicial exception is not integrated into a practical application; the remaining limitations of the claim are as follows: “a neural network configured to”, “a neural network configured to”, “a deep neural network configured to” the limitations are an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP § 2106.05(f)(2). “acquiring a plurality of modalities associated with an object and information identifying the object, wherein the plurality of modalities includes image data corresponding to an image of the object and text data describing the object”, “receive the feature values for each of the plurality of modalities and information identifying the object” these limitations are an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g).
Step 2B: The claim does not contain significantly more than the judicial exception. The limitations “a neural network configured to”, “a neural network configured to”, “a deep neural network configured to” the limitations are an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP § 2106.05(f)(2). “acquiring a plurality of modalities associated with an object and information identifying the object, wherein the plurality of modalities includes image data corresponding to an image of the object and text data describing the object”, “receive the feature values for each of the plurality of modalities and information identifying the object” these limitations are an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g). Furthermore the additional element is directed to receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d as well as to storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). See MPEP § 2106.05(d)/(II). .
Regarding claim 16, the rejection of claim 1 is incorporated and further:
Step 2A Prong 1: The claim recites, in part:
“derive the attention weights according to a formula
σ
(
W
[
f
θ
h
ⅈ
j
,
f
θ
h
t
j
]
+
b
)
, where
a
j
is an attention weight for the object j,
h
i
j
is the image data,
h
t
j
is the text data,
σ
is an activation function,
f
θ
h
ⅈ
j
is the feature value with respect to the image data,
f
θ
h
t
j
is the feature value with respect to the text data,
W
[
f
θ
h
ⅈ
j
,
f
θ
h
t
j
]
is the concatenated value in which a weight coefficient W is applied to the feature values
f
θ
h
i
j
and
f
θ
h
t
j
, and b is a bias value” this limitation is a mathematical concept.
Step 2A Prong 2: The judicial exception is not integrated into a practical application; the remaining limitations of the claim are as follows: “the neural network is configured to” the limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP § 2106.05(f)(2).
Step 2B: The claim does not contain significantly more than the judicial exception. The limitations “the neural network is configured to” the limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP § 2106.05(f)(2).
Regarding claim 17, the rejection of claim 1 is incorporated and further:
Step 2A Prong 1: a continuation of the abstract idea identified in the parent claim.
Step 2A Prong 2: The judicial exception is not integrated into a practical application; the remaining limitations of the claim are as follows: “output code configured to cause at least one of the last least one processor to” the limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP § 2106.05(f)(2). “output the predicted at least one attribute” the limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g).
Step 2B: The claim does not contain significantly more than the judicial exception. The limitations “output code configured to cause at least one of the last least one processor to” the limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP § 2106.05(f)(2). “output the predicted at least one attribute” the limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g). Furthermore the additional element is directed to storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). See MPEP § 2106.05(d)/(II).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. § 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 7, 9, 13, and 14 are rejected under 35 U.S.C. § 103 as being unpatentable over Nesta et al. (US20190354797A1) (as cited in the IDS, hereinafter “Nesta”) in view of Li et al. (US 2021/0233511 A1) (hereinafter “Li”) in further view of Bi et al. (“A Multimodal Late Fusion Model for E-Commerce Product Classification”, Bi et al., 14 Aug 2020) (hereinafter “Bi”).
Regarding claim 1:
Nesta teaches [a]n information processing apparatus comprising:
at least one memory configured to store program code (Nesta, claim 16 “a memory storing instructions;”); and
at least one processor configured to operate as instructed by the program code (Nesta, claim 16 “a processor coupled to the memory and configured to execute the instructions to cause the system to perform operations comprising:”), the program code comprising:
acquisition code configured to cause the at least one processor to acquire a plurality of modalities associated with an object and information identifying the object (Nesta, ¶1 “The system may use a variety of input modalities including images, video and/or audio, and the expert modules may comprise a corresponding image expert, a corresponding video expert and/or a corresponding audio expert.” In light of ¶26 of the specification, modalities associated with an object include images “An example of a plurality of modalities, which are information associated with an object, is image data showing an image of a product (hereinafter simply referred to as image data) and text data describing the product (hereinafter simply referred to as text data)”);
feature generation code configured to cause the at least one processor to apply the plurality of modalities to a fully connected network (Nesta, ¶29 “A preprocessing network (such as networks 223 and 224) may be used for a first feature transformation (e.g. by using an Inception V3 network or a VGG16 network), and the output layer is then fed to fully connected layers 225 to produce a logistic classification 226.”), wherein the fully connected network is configured to generate feature values for each of the plurality of modalities (Nesta, ¶12 “The system may be further configured to accept a variety of input modalities including images, video and/or audio, and the operations performed by the processor may further include extracting features associated with each input modality by a process that includes inputting the corresponding data stream to a trained neural network” here, the features extracted associated with each input modality can be considered the generated feature for a plurality of modalities);
weight derivation code configured to cause the at least one processor to apply the feature values output from the fully connected network to a neural network (Nesta, ¶26 “A Gate Recurrent Neural Network 140 receives the feature vectors zn(l) as an input (e.g., as a single stacked vector z(l)=[zn (l); . . . ; zN(l)]) and is trained to produce fusion weights wn to assign a relative weight to each input modality.”), wherein the neural network is configured to:
receive the feature values for each of the plurality of modalities and information identifying the object (Nesta, ¶26 “A Gate Recurrent Neural Network 140 receives the feature vectors zn(l) as an input (e.g., as a single stacked vector z(l)=[zn (l); . . . ; zN(l)]) and is trained to produce fusion weights wn to assign a relative weight to each input modality 110.” here, the received feature vectors which relate to each modality can be considered the received feature values for each of the plurality of modalities), and
derive attention weights corresponding to each of the plurality of modalities based on the feature values for each of the plurality of modalities and the information identifying the object (Nesta, ¶7 “A gate expert is configured to receive the extracted features from the plurality of expert modules and output a set of weights for each of the input modalities.” Here, the weights output for each input modalities from extracted features can be considered the derived attention weights corresponding to the features), the attention weights indicating an importance of each of the plurality of modalities to attribute prediction (Nesta, ¶26 “A Gate Recurrent Neural Network 140 receives the feature vectors zn(l) as an input (e.g., as a single stacked vector z(l)=[zn (l); . . . ; zN(l)]) and is trained to produce fusion weights wn to assign a relative weight to each input modality 110.” Here, the weight assigned to each input modalities can be considered an importance of each modalities to attribute prediction); and
Nesta does not teach “classification code configured to cause the at least one processor to:
for each feature value of the feature values, multiply the feature value by a corresponding attention weight to generate a weighted feature value of a plurality of weighted feature values ,
concatenate the plurality of weighted feature values to generate a concatenated value , and
apply the concatenated value”
However Li teaches classification code configured to cause the at least one processor to:
for each feature value of the feature values, multiply the feature value by a corresponding attention weight to generate a weighted feature value of a plurality of weighted feature values (Li, ¶126 “Step 131: the contribution weight is calculated for each frame-level shallow-layer feature vector by processing each frame-level shallow-layer feature vectors on the neural networks with attention mechanism.” Further Li, ¶168 “wherein the shallow-layer features are input into the neural network structure of frame-level deep integrated feature vector 703 respectively to be weighted by the corresponding contribution weights which are calculated with the attention mechanism”),
concatenate the plurality of weighted feature values to generate a concatenated value (Li, ¶36 “In some specific embodiments, the weighted integration processing comprises: weighting the shallow-layer features in frame level with the corresponding contribution weight, and performing concatenation or accumulation processing. The mathematical formula of the concatenation processing is as follows:
I=Concat(a 1 F 1 ,a 2 F 2 , . . . ,a N F N)”), and
apply the concatenated value(Li, ¶137 “Step 133: Perform dimension reduction or normalization processing on the frame-level preliminary integrated feature vector to obtain a frame-level deep integrated feature vector.”) configured to predict at least one attribute of the object based on the concatenated value (Li, ¶143 “Step 141: Inputting the frame-level deep integrated feature vectors into the neural networks for specific speech task.”).
Nesta and Li are analogous art because both references concern methods for multimodal networks. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Nesta’s multimodal network to incorporate the weighted concatenated features fed into a deep neural network taught by Li. The motivation for doing so would have been to make full use of the distinction and complementarity between different types of features as stated in Li, ¶57 “By making full use of the distinction and complementarity between different types of acoustic features, the entire deep neural networks are joint optimized with the acoustic feature integration process, to obtain the frame-level or segment-level deep integrated feature vectors of the task-related adaptation.”
Nesta in view of Li does not teach “wherein the plurality of modalities include image data corresponding to an image of the object and text data describing the object”
However, Li teaches wherein the plurality of modalities include image data corresponding to an image of the object and text data describing the object (Bi, page 1, col 2, section 2, “The dataset consists of product titles, descriptions, images and their corresponding product type codes.”)
Nesta in view of Li and Bi are analogous art because both references concern methods for multimodal networks. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Nesta’s multimodal network to incorporate the weighted concatenated features fed into a deep neural network taught by Li. The motivation for doing so would have been to outperform other unimodal and multimodal methods as stated in Bi, page 1, col 1, abstract, ¶1 “Experimental results on Multimodal Product Classification Task of SIGIR 2020 E-Commerce Workshop Data Challenge1 demonstrate the superiority and effectiveness of our proposed method compared with unimodal and other multimodal methods.”
Regarding claim 7:
Nesta in view of Li in further view of Bi teaches [t]he information processing apparatus according to Claim 1, wherein the fully connected network takes (Nesta, ¶29 “A preprocessing network (such as networks 223 and 224) may be used for a first feature transformation (e.g. by using an Inception V3 network or a VGG16 network), and the output layer is then fed to fully connected layers 225 to produce a logistic classification 226.”), as input, the plurality of modalities and outputs the feature values of the plurality of modalities by mapping them to a latent space common to the plurality of modalities (Nesta, ¶6 “In another embodiment a co-learning framework is defined to encourage co-adaptation of a subset of latent variables belonging to an expert network related to different modalities.” Here, the co-learning framework can be considered the first machine learning model. A co-adaption of a subset of latent variables can be considered a common latent space as they share common latent variables. The adaption of these variables can be considered a mapping).
Regarding claim 9:
Nesta in view of Li in further view of Bi teaches [t]he information processing apparatus according to Claim 1, wherein the acquisition code is further configured to cause the at least one processor to encode the plurality of modalities to acquire a plurality of encoded modalities (Nesta, ¶24 “The feature extraction may be obtained, for example, by using the encoding part of a neural network trained on a classification task related to the specific modality, using an autoencoder (AE) compression scheme, and/or other feature extraction process” here, using the encoding part of a neural network related to the modality can be considered the acquisition of encoded modalities and the features are extracted from those encoded modalities), and the fully connected network is further configured to generate the feature values for each of the plurality of encoded modalities (Nesta, ¶29 “A preprocessing network (such as networks 223 and 224) may be used for a first feature transformation (e.g. by using an Inception V3 network or a VGG16 network), and the output layer is then fed to fully connected layers 225 to produce a logistic classification 226.”).
Regarding claim 13:
Nesta teaches [a]n information processing method comprising:
acquiring a plurality of modalities associated with an object and information identifying the object (Nesta, ¶1 “The system may use a variety of input modalities including images, video and/or audio, and the expert modules may comprise a corresponding image expert, a corresponding video expert and/or a corresponding audio expert.” In light of ¶26 of the specification, modalities associated with an object include images “An example of a plurality of modalities, which are information associated with an object, is image data showing an image of a product (hereinafter simply referred to as image data) and text data describing the product (hereinafter simply referred to as text data)”);
applyinng the plurality of modalities to a fully connected network (Nesta, ¶29 “A preprocessing network (such as networks 223 and 224) may be used for a first feature transformation (e.g. by using an Inception V3 network or a VGG16 network), and the output layer is then fed to fully connected layers 225 to produce a logistic classification 226.”), wherein the fully connected network is configured to generate feature values for each of the plurality of modalities (Nesta, ¶12 “The system may be further configured to accept a variety of input modalities including images, video and/or audio, and the operations performed by the processor may further include extracting features associated with each input modality by a process that includes inputting the corresponding data stream to a trained neural network” here, the features extracted associated with each input modality can be considered the generated feature for a plurality of modalities);
applying the feature values output from the fully connected network to a neural network (Nesta, ¶26 “A Gate Recurrent Neural Network 140 receives the feature vectors zn(l) as an input (e.g., as a single stacked vector z(l)=[zn (l); . . . ; zN(l)]) and is trained to produce fusion weights wn to assign a relative weight to each input modality.”), wherein the neural network is configured to:
receive the feature values for each of the plurality of modalities and information identifying the object (Nesta, ¶26 “A Gate Recurrent Neural Network 140 receives the feature vectors zn(l) as an input (e.g., as a single stacked vector z(l)=[zn (l); . . . ; zN(l)]) and is trained to produce fusion weights wn to assign a relative weight to each input modality 110.” here, the received feature vectors which relate to each modality can be considered the received feature values for each of the plurality of modalities), and
derive attention weights corresponding to each of the plurality of modalities based on the feature values for each of the plurality of modalities and the information identifying the object (Nesta, ¶7 “A gate expert is configured to receive the extracted features from the plurality of expert modules and output a set of weights for each of the input modalities.” Here, the weights output for each input modalities from extracted features can be considered the derived attention weights corresponding to the features), the attention weights indicating an importance of each of the plurality of modalities to attribute prediction (Nesta, ¶26 “A Gate Recurrent Neural Network 140 receives the feature vectors zn(l) as an input (e.g., as a single stacked vector z(l)=[zn (l); . . . ; zN(l)]) and is trained to produce fusion weights wn to assign a relative weight to each input modality 110.” Here, the weight assigned to each input modalities can be considered an importance of each modalities to attribute prediction); and
Nesta does not teach “classification code configured to cause the at least one processor to:
for each feature value of the feature values, multiply the feature value by a corresponding attention weight to generate a weighted feature value of a plurality of weighted feature values ,
concatenating the plurality of weighted feature values to generate a concatenated value , and
applying the concatenated value”
However Li teaches classification code configured to cause the at least one processor to:
for each feature value of the feature values, multiply the feature value by a corresponding attention weight to generate a weighted feature value of a plurality of weighted feature values (Li, ¶126 “Step 131: the contribution weight is calculated for each frame-level shallow-layer feature vector by processing each frame-level shallow-layer feature vectors on the neural networks with attention mechanism.” Further Li, ¶168 “wherein the shallow-layer features are input into the neural network structure of frame-level deep integrated feature vector 703 respectively to be weighted by the corresponding contribution weights which are calculated with the attention mechanism”),
concatenate the plurality of weighted feature values to generate a concatenated value (Li, ¶36 “In some specific embodiments, the weighted integration processing comprises: weighting the shallow-layer features in frame level with the corresponding contribution weight, and performing concatenation or accumulation processing. The mathematical formula of the concatenation processing is as follows:
I=Concat(a 1 F 1 ,a 2 F 2 , . . . ,a N F N)”), and
apply the concatenated value(Li, ¶137 “Step 133: Perform dimension reduction or normalization processing on the frame-level preliminary integrated feature vector to obtain a frame-level deep integrated feature vector.”) configured to predict at least one attribute of the object based on the concatenated value (Li, ¶143 “Step 141: Inputting the frame-level deep integrated feature vectors into the neural networks for specific speech task.”).
Nesta and Li are analogous art because both references concern methods for multimodal networks. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Nesta’s multimodal network to incorporate the weighted concatenated features fed into a deep neural network taught by Li. The motivation for doing so would have been to make full use of the distinction and complementarity between different types of features as stated in Li, ¶57 “By making full use of the distinction and complementarity between different types of acoustic features, the entire deep neural networks are joint optimized with the acoustic feature integration process, to obtain the frame-level or segment-level deep integrated feature vectors of the task-related adaptation.”
Nesta in view of Li does not teach “wherein the plurality of modalities include image data corresponding to an image of the object and text data describing the object”
However, Li teaches wherein the plurality of modalities include image data corresponding to an image of the object and text data describing the object (Bi, page 1, col 2, section 2, “The dataset consists of product titles, descriptions, images and their corresponding product type codes.”)
Nesta in view of Li and Bi are analogous art because both references concern methods for multimodal networks. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Nesta’s multimodal network to incorporate the weighted concatenated features fed into a deep neural network taught by Li. The motivation for doing so would have been to outperform other unimodal and multimodal methods as stated in Bi, page 1, col 1, abstract, ¶1 “Experimental results on Multimodal Product Classification Task of SIGIR 2020 E-Commerce Workshop Data Challenge1 demonstrate the superiority and effectiveness of our proposed method compared with unimodal and other multimodal methods.”
Regarding claim 14:
Nesta teaches [a] non-transitory computer-readable storage medium storing computer executable instructions for causing a computer to implement an information processing method (Nesta, ¶41 “Software, in accordance with the present disclosure, such as program code and/or data, may be stored on one or more computer readable mediums.”) the information processing method comprising:
acquiring a plurality of modalities associated with an object and information identifying the object (Nesta, ¶1 “The system may use a variety of input modalities including images, video and/or audio, and the expert modules may comprise a corresponding image expert, a corresponding video expert and/or a corresponding audio expert.” In light of ¶26 of the specification, modalities associated with an object include images “An example of a plurality of modalities, which are information associated with an object, is image data showing an image of a product (hereinafter simply referred to as image data) and text data describing the product (hereinafter simply referred to as text data)”);
applyinng the plurality of modalities to a fully connected network (Nesta, ¶29 “A preprocessing network (such as networks 223 and 224) may be used for a first feature transformation (e.g. by using an Inception V3 network or a VGG16 network), and the output layer is then fed to fully connected layers 225 to produce a logistic classification 226.”), wherein the fully connected network is configured to generate feature values for each of the plurality of modalities (Nesta, ¶12 “The system may be further configured to accept a variety of input modalities including images, video and/or audio, and the operations performed by the processor may further include extracting features associated with each input modality by a process that includes inputting the corresponding data stream to a trained neural network” here, the features extracted associated with each input modality can be considered the generated feature for a plurality of modalities);
applying the feature values output from the fully connected network to a neural network (Nesta, ¶26 “A Gate Recurrent Neural Network 140 receives the feature vectors zn(l) as an input (e.g., as a single stacked vector z(l)=[zn (l); . . . ; zN(l)]) and is trained to produce fusion weights wn to assign a relative weight to each input modality.”), wherein the neural network is configured to:
receive the feature values for each of the plurality of modalities and information identifying the object (Nesta, ¶26 “A Gate Recurrent Neural Network 140 receives the feature vectors zn(l) as an input (e.g., as a single stacked vector z(l)=[zn (l); . . . ; zN(l)]) and is trained to produce fusion weights wn to assign a relative weight to each input modality 110.” here, the received feature vectors which relate to each modality can be considered the received feature values for each of the plurality of modalities), and
derive attention weights corresponding to each of the plurality of modalities based on the feature values for each of the plurality of modalities and the information identifying the object (Nesta, ¶7 “A gate expert is configured to receive the extracted features from the plurality of expert modules and output a set of weights for each of the input modalities.” Here, the weights output for each input modalities from extracted features can be considered the derived attention weights corresponding to the features), the attention weights indicating an importance of each of the plurality of modalities to attribute prediction (Nesta, ¶26 “A Gate Recurrent Neural Network 140 receives the feature vectors zn(l) as an input (e.g., as a single stacked vector z(l)=[zn (l); . . . ; zN(l)]) and is trained to produce fusion weights wn to assign a relative weight to each input modality 110.” Here, the weight assigned to each input modalities can be considered an importance of each modalities to attribute prediction); and
Nesta does not teach “classification code configured to cause the at least one processor to:
for each feature value of the feature values, multiply the feature value by a corresponding attention weight to generate a weighted feature value of a plurality of weighted feature values ,
concatenating the plurality of weighted feature values to generate a concatenated value , and
applying the concatenated value”
However Li teaches classification code configured to cause the at least one processor to:
for each feature value of the feature values, multiply the feature value by a corresponding attention weight to generate a weighted feature value of a plurality of weighted feature values (Li, ¶126 “Step 131: the contribution weight is calculated for each frame-level shallow-layer feature vector by processing each frame-level shallow-layer feature vectors on the neural networks with attention mechanism.” Further Li, ¶168 “wherein the shallow-layer features are input into the neural network structure of frame-level deep integrated feature vector 703 respectively to be weighted by the corresponding contribution weights which are calculated with the attention mechanism”),
concatenate the plurality of weighted feature values to generate a concatenated value (Li, ¶36 “In some specific embodiments, the weighted integration processing comprises: weighting the shallow-layer features in frame level with the corresponding contribution weight, and performing concatenation or accumulation processing. The mathematical formula of the concatenation processing is as follows:
I=Concat(a 1 F 1 ,a 2 F 2 , . . . ,a N F N)”), and
apply the concatenated value(Li, ¶137 “Step 133: Perform dimension reduction or normalization processing on the frame-level preliminary integrated feature vector to obtain a frame-level deep integrated feature vector.”) configured to predict at least one attribute of the object based on the concatenated value (Li, ¶143 “Step 141: Inputting the frame-level deep integrated feature vectors into the neural networks for specific speech task.”).
Nesta and Li are analogous art because both references concern methods for multimodal networks. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Nesta’s multimodal network to incorporate the weighted concatenated features fed into a deep neural network taught by Li. The motivation for doing so would have been to make full use of the distinction and complementarity between different types of features as stated in Li, ¶57 “By making full use of the distinction and complementarity between different types of acoustic features, the entire deep neural networks are joint optimized with the acoustic feature integration process, to obtain the frame-level or segment-level deep integrated feature vectors of the task-related adaptation.”
Nesta in view of Li does not teach “wherein the plurality of modalities include image data corresponding to an image of the object and text data describing the object”
However, Li teaches wherein the plurality of modalities include image data corresponding to an image of the object and text data describing the object (Bi, page 1, col 2, section 2, “The dataset consists of product titles, descriptions, images and their corresponding product type codes.”)
Nesta in view of Li and Bi are analogous art because both references concern methods for multimodal networks. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Nesta’s multimodal network to incorporate the weighted concatenated features fed into a deep neural network taught by Li. The motivation for doing so would have been to outperform other unimodal and multimodal methods as stated in Bi, page 1, col 1, abstract, ¶1 “Experimental results on Multimodal Product Classification Task of SIGIR 2020 E-Commerce Workshop Data Challenge1 demonstrate the superiority and effectiveness of our proposed method compared with unimodal and other multimodal methods.”
Claim 2 is rejected under 35 U.S.C. § 103 as being unpatentable over Nesta in view of Li in view of Bi in further view of Gao et al. (“Attention driven multi-modal similarity learning”, Gao et al., March 2018) (hereinafter “Gao”).
Regarding claim 3:
Nesta in view of Li in further view of Bi in further view of Gao teaches [t]he information processing apparatus according to Claim 1, wherein the attention weights of the plurality of modalities are 1 in total (Gao, pages 5-6, section 3.1.2, ¶3 “In Eqs. (6) and (7), the softmax function is used to generate positive attention weights that sum to 1”).
It would have been obvious to combine the teachings of Nesta in view of Li in further view of Bi and Gao for the reasons set forth in connection with claim 2 above.
Nesta in view of Li in further view of Bi and Gao are analogous art because both references concern multimodal learning. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Nesta/Li/Bi’s recurrent multimodal system to incorporate the attention as a measure of importance taught by Gao. The motivation for doing so would have been to improve the accuracy of similarity learning as stated in Gao, page 2, ¶2 “we propose an interaction-oriented attention mechanism to improve the accuracy of similarity learning, and meanwhile show that the attention weights returned by the mechanism are able to improve the model interpretability”.
Claims 10, 11 and 16 are rejected under 35 U.S.C. § 103 as being unpatentable over Nesta in view of Li in view of Bi in further view of Zhu et al. (“Multimodal Joint Attribute Prediction and Value Extraction for E-commerce Product”, Zhu et al., 15 Sep 2020) (as cited in the IDS, hereinafter “Zhu”).
Regarding claim 10:
Nesta in view of Li in further view of Bi teaches [t]he information processing apparatus according to Claim 1
Nesta in view of Li in further view of Bi does not teach “wherein the object is a commodity, and the plurality of modalities includes two or more of data of an image representing the commodity, data of text describing the commodity”
However, Zhu teaches wherein the object is a commodity, and the plurality of modalities includes two or more of data of an image representing the commodity, data of text describing the commodity, (Zhu, page 2, col 1, ¶2 “Furthermore, beyond the textual product descriptions, product images can provide additional clues for the attribute prediction and value extraction tasks.”) and data of sound describing the commodity. It is noted the claim recites alternative language, and Zhu teaches at least one of the alternatives.
Nesta in view of Li in further view of Bi and Zhu are analogous art because both references concern multimodal learning. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Nesta/Li/Bi’s recurrent multimodal system to incorporate the image and text data as taught by Zhu. The motivation for doing so would have been to more accurately extract attribute values Zhu, page 2, col 1, ¶1 “Given a textual product description, we can extract attribute values more accurately with a known product attribute”
Regarding claim 11:
Nesta in view of Li in further view of Bi teaches [t]he information processing apparatus according to Claim 1
Nesta in view of Li in further view of Bi does not teach “wherein the attribute of the object includes color information of a product”
However, Zhu teaches wherein the attribute of the object includes color information of a product (Zhu, page 5, col 1, ¶1 “Finally, we obtained 87,194 text-image instances consisting of the following categories of products: Clothes, Pants, Dresses, Shoes, Boots, Luggage, and Bags, and involving 26 types of product attributes such as “Material”, “Collar Type”, “Color”, etc.”).
Nesta in view of Li in further view of Bi and Zhu are analogous art because both references concern multimodal learning. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Nesta/Li/Bi’s recurrent multimodal system to incorporate the color attributes as taught by Zhu. The motivation for doing so would have been to add a crucial indication for attribute values as stated in Zhu, page 4, col 1, section 2.6, ¶2 “We argue that the product attributes can provide crucial indications for the attribute values. For example, given a sentence “The red collar and golden buttons in the shirt form a colorful fashion topic” and the predicted product attribute “Color”, it is easy to recognize the value “golden” corresponding to attribute “Color” instead of “Material”. Thus, we incorporate the result of the product attribute prediction
y
^
a
to improve the value extraction.”
Regarding claim 16:
Nesta in view of Li in further view of Bi teaches [t]he information processing apparatus according to Claim 1,
Nesta in view of Li in further view of Bi does not teach “wherein the neural network is configured to derive the weights according to the formula
σ
(
W
[
f
θ
h
ⅈ
j
,
f
θ
h
t
j
]
+
b
)
, where
a
j
is an attention weight for the object j,
h
i
j
is image data,
h
t
j
is text data,
σ
is an activation function,
f
θ
h
ⅈ
j
is a feature value with respect to the image data,
f
θ
h
t
j
is a feature value with respect to the text data,
W
[
f
θ
h
ⅈ
j
,
f
θ
h
t
j
]
is a concatenated value in which a weight coefficient W is applied to the feature values
f
θ
h
i
j
and
f
θ
h
t
j
, and b is a bias value”
Nesta in view of Li in further view of Bi teaches wherein the neural network is configured to derive the weights according to the formula
σ
(
W
[
f
θ
h
ⅈ
j
,
f
θ
h
t
j
]
+
b
)
, where
a
j
is an attention weight for the object j,
h
i
j
is image data,
h
t
j
is text data,
σ
is an activation function,
f
θ
h
ⅈ
j
is a feature value with respect to the image data,
f
θ
h
t
j
is a feature value with respect to the text data,
W
[
f
θ
h
ⅈ
j
,
f
θ
h
t
j
]
is a concatenated value in which a weight coefficient W is applied to the feature values
f
θ
h
i
j
and
f
θ
h
t
j
, and b is a bias value (Zhu, page 4, col 2, section 2.4, ¶3 “Specifically, we feed the text and image representations hi and vk into the global-gated cross modality attention layer…” and Zhu, page 4-5, col 2-1, section 2.4, ¶4 “The global visual gate
g
i
G
is determined by the representation of the sentence and the image, which are obtained by the text encoder and the image encoder, respectively, as follows:
PNG
media_image1.png
45
308
media_image1.png
Greyscale
where W1 and W2 are weight matrices.” Here, the global visual gate can be considered equivalent to the formula given, wherein the weight matrices have been distributed. Hi and vG are the text and image data, respectively).
Nesta in view of Li in further view of Bi and Zhu are analogous art because both references concern multimodal learning. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Nesta/Li/Bi’s recurrent multimodal system to incorporate the global visual gate as taught by Zhu. The motivation for doing so would have been to benefit attribute prediction task with visually grounded semantics Zhu, page 2, col 2, ¶1 “First, we selectively enhance the semantic representation of the textual product descriptions with a global gated cross-modality attention module that is anticipated to benefit attribute prediction task with visually grounded semantics.”
Claim 17 is rejected under 35 U.S.C. § 103 as being unpatentable over Nesta in view of Li in view of Bi in further view of Tzirakis et al. (“End-to-End Multimodal Emotion Recognition using Deep Neural Networks”, Tzirakis et al., 27 April 2017) (hereinafter “Tzirakis”).
Regarding claim 17:
Nesta in view of Li in further view of Bi teaches [t]he information processing apparatus according to Claim 1
Nesta in view of Li in further view of Bi does not teach “further comprising output code configured to cause at least one of the last least one processor to output the predicted at least one attribute”
However, Tzirakis teaches further comprising output code configured to cause at least one of the last least one processor to output the predicted at least one attribute (Tzirakis, page 7, col 1, ¶1 “Finally, to further demonstrate the benefits of our model for automatic prediction of arousal and valence Figure 3 illustrates results for single test subject from RECOLA.”). Here, the predicted results can be considered the outputted predicted attributes).
Nesta in view of Li in further view of Bi and Tzirakis are analogous art because both references concern methods for multimodal networks. Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Nesta/Li/Bi’s multimodal network to incorporate the weighted concatenated features fed into a deep neural network taught by Tzirakis. The motivation for doing so would have been to outperform traditional networks as stated in Tzirakis, page 1, abstract “The system is then trained in an end-to-end fashion where– by also taking advantage of the correlations of the each of the streams– we manage to significantly outperform the traditional approaches based on auditory and visual handcrafted features for the prediction of spontaneous and natural emotions on the RECOLA database of the AVEC 2016 research challenge on emotion recognition.”.
Response to Arguments
Applicant's arguments filed May 13th, 2026 (hereinafter “Remarks”) have been fully considered but they are not persuasive.
Applicant’s arguments regarding the 35 U.S.C. 112(b) rejections of the previous office action have been fully considered. However, the amendments have required additional indefiniteness rejections to be made in this action.
Applicant’s arguments with respect to the prior art rejections have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Rejections under 35 U.S.C. § 101:
Argument 1:
“The improvement therefore lies in the ML-based architecture itself. More specifically, the claimed apparatus introduces a multi-stage neural architecture in which one network generates feature values, another network derives object-specific attention weights for the modalities based on those feature values and object-identifying information, and a downstream deep neural network performs classification using the weighted multimodal representation. By incorporating this dedicated weighting mechanism into the processing pipeline, the system modifies how the attribute prediction apparatus itself processes multimodal input data. This is a technical improvement to ML-based attribute prediction systems because the claimed architecture dynamically adjusts modality contributions on an object-specific basis, thereby improving predictive performance relative to conventional architectures that concatenate modalities uniformly… These limitations recite a specific ML pipeline and a particular arrangement of interacting model components. The claim does not merely recite the result of predicting an attribute or the generalized concept of considering some information more important than other information. Rather, the claim recites an architecture in which a fully connected network generates modality- specific feature values, a neural network derives object-specific attention weights using those feature values and identifying information, and a deep neural network performs classification on a weighted concatenated representation. The claimed advance is therefore technological because it improves the manner in which the ML-based attribute prediction system itself operates.” (Remarks, page 12-13).
Examiners Response:
Examiner respectfully disagrees, the MPEP states “it is important to keep in mind that an improvement in the abstract idea itself (e.g. a recited fundamental economic concept) is not an improvement in technology.” See MPEP § 2106.05(a)(II). Here the improvement to product searching and identification is an improvement to a mental process which can be practically performed in the human mind. The improvement to the way attributes are predicted, through importance scores, is an improvement in an abstract idea. A person could observe products and predict various attributes by assigning importance to various modalities. The use of neural networks and computing components amount an additional element that generally links the use of the judicial exception to a particular technological environment or field of use, or amounts to adding the words “apply it” (or an equivalent) with the judicial exception. See MPEP §§ 2106.05(f)(2), 2106.05(h). Further, the additional elements are recited at a high level of generality, and even when viewed in combination, these additional elements do not integrate the recited judicial exception into a practical application.
Argument 2:
“This characterization oversimplifies the claims and abstracts away the very claim features that supply the technological improvement. The claims are not directed merely to the concept of “assigning importance” to different information. Instead, the claims recite a specific ML-based architecture that implements object-specific modality weighting through particular model components arranged in a particular sequence. The recited improvement is thus not the abstract idea itself, but an improvement to the operation of ML-based attribute prediction systems through the addition of a neural network based weighting stage and a downstream deep neural network classification stage that together process multimodal data in a manner different from conventional systems. Put differently, the alleged abstract idea is not what improves the system. The improvement is the claimed architecture that changes how the system generates feature values, derives attention weights, forms a weighted multimodal representation, and performs classification…Accordingly, when considered as a whole, claim 1 integrates any alleged abstract idea into a practical application because it recites a specific ML-based technical architecture that improves the operation of attribute prediction systems themselves, rather than merely improving an alleged judicial exception.” (Remarks, page 14-15).
Examiners Response:
Examiner respectfully disagrees, The courts consider a mental process (thinking) that “can be performed in the human mind, or by a human using a pen and paper” to be an abstract idea. CyberSource Corp. v. Retail Decisions, Inc., 654 F.3d 1366, 1372, 99 USPQ2d 1690, 1695 (Fed. Cir. 2011).” See MPEP § 2106.04(a)(2)(III). The improvement to predictions is an improvement in a mental process. A person could observe products and predict various attributes by assigning importance to various modalities. While the claim does require certain computing components, the additional elements are recited at a high level of generality, and even when viewed in combination, these additional elements do not integrate the recited judicial exception into a practical application.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Zahavy et al. (“Is a picture worth a thousand words? A Deep Multi-Modal Fusion Architecture for Product Classification in e-commerce”, Zahavy et al., 29 Nov 2016) discloses a decision level fusion approach for multi-modal product classification using text and image inputs.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JACOB Z SUSSMAN MOSS whose telephone number is (571) 272-1579. The examiner can normally be reached Monday - Friday, 9 a.m. - 5 p.m. ET
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached on (571) 272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.S.M./Examiner, Art Unit 2122
/KAKALI CHAKI/Supervisory Patent Examiner, Art Unit 2122