Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response to the application and claims filed 03/06/2024. Claims 1-21 are
pending and have been examined. Claims 1-21 are rejected.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 5/27/2026 and 5/27/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Priority
The examiner acknowledges the priority benefit to U.S. Provisional Application No. 63/450,551 filed on 03/07/2023. The present application claims priority to U.S. Provisional Application No. 63/450,551 filed on 03/07/2023.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-21 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1, 10 and 16 recite the limitation "the latent variable" in two instances in each claim. There is insufficient antecedent basis for this limitation in the claim. Claims 2-9, 11-15, and 17-21 are rejected for being dependent on claim 1, 10, and 16.
Claim 1, 10 and 16 recite the limitation "the examples" in three instances in each claim. There is insufficient antecedent basis for this limitation in the claim. Claims 2-9, 11-15, and 17-21 are rejected for being dependent on claim 1, 10, and 16.
Claim Objections
Claim 1, 10, 16 objected to because of the following informalities:
In the generating an attention-based encoder limitation “generates … for the samples” add ; and after “for the samples” and before “uses”
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-21 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract
idea without significantly more.
Claim 1
Step 1: The claim recites a method; therefore, it is directed to the statutory category of processes.
Step 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. The claim recites the following abstract ideas:
“(...) processing the training data to extract, for each modality in the full set of modalities, a fixed-dimensional input vector format representing that modality (...);” (This limitation is a mental process. A person mentally or with a pen and paper can look at the training data for each modality, e.g., an image, a trajectory or a body pose, and write down a list of numbers having a fixed length, i.e., a vector of a fixed dimension, that represents that modality.)
“(...) generates, from the training vectors for the samples, a fixed-dimensional vector representation template for the prediction target, wherein the number of dimensions in the fixed-dimensional vector representation template is constant and is independent of the number of modalities represented by the training vectors for the samples” (This limitation is a mental process. A person mentally or with a pen and paper can look at the training vectors for the samples and draw a template for the prediction target, e.g., a table having a fixed number of entries. The wherein clause merely specifies that the person chooses a constant number of entries for the template regardless of how many modalities are present in the training vectors, which is a choice that can be made in the mind.)
“(...) uses the samples and the fixed-dimensional vector representation template to generate, from the training vectors for the samples, a latent distribution;” (This limitation falls within the mathematical concepts grouping because generating a latent distribution from the training vectors amounts to calculating the parameters of a probability distribution, e.g., the mean μ and the variance σ of a Gaussian distribution, see specification paragraphs [0028], [0040] and [0056] and equation (19), which is a mathematical calculation.)
“(...) generate, from the representations of the examples of the prediction target according to the fixed-dimensional input vector format and the latent variable from the latent distribution, predictions for the examples of the prediction target.” (This limitation is a mental process. A person mentally or with a pen and paper can look at the representations of an example, e.g., the past positions of a pedestrian, together with a value drawn from the latent distribution and predict an outcome, e.g., where the pedestrian will be in the future.)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
“A method for training a first machine learning model to handle multimodal data including examples with missing modalities, the method comprising:” (The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)). -- Examiner’s Note (EN): The preamble merely limits the recited abstract ideas to the field of multimodal machine learning and to the technological environment of a machine learning model. The instant specification characterizes the technology as “confined to multimodal machine learning applications” (paragraph [00112]), i.e., a field of use.)
“receiving a plurality of multimodal training data, the training data comprising a plurality of samples of a prediction target, wherein: each sample includes at least a subset of a full set of modalities, wherein the full set of modalities is a plurality of modalities; and the samples collectively include instances of each modality within the full set of modalities;” (Data Gathering - Mere data gathering recited at a high level of generality, and thus is insignificant extra-solution activity (MPEP 2106.05(g)). -- Examiner’s Note (EN): This limitation amounts to receiving the training data on which the recited abstract ideas are performed. The wherein clauses merely describe the characteristics of the data received, i.e., that a sample may be missing one or more modalities and that every modality appears in at least one sample, and this doesn’t change the limitation from being mere data gathering.)
“using the training data as input to a first attention-based neural network, comprising:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based neural network as a tool to perform the recited abstract ideas. No details of the first attention-based neural network are recited beyond that it is attention-based.)
“(...) to generate a respective feature encoder for each modality in the full set of modalities;” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): This limitation amounts to generic training of a generic feature encoder for each modality using the extracted vectors. The specification at paragraph [0064] indicates that the feature encoders are generic, off the shelf networks, e.g., a ConvNet for an image and a gated recurrent unit (GRU) for a sequence of vectors.)
“generating an attention-based encoder that:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based encoder as a tool to perform the recited abstract ideas. The specification describes the attention-based encoder as a transformer encoder that adopts the known Set Transformer.)
“receives sets of training vectors in the fixed-dimensional input vector format, wherein each set of training vectors represents one of the samples; and” (Data Gathering - Mere data gathering recited at a high level of generality, and thus is insignificant extra-solution activity (MPEP 2106.05(g)). -- Examiner’s Note (EN): This limitation amounts to receiving the vectors on which the recited abstract ideas are performed, which is mere data gathering.)
“the method further comprising using: representations of the samples of the prediction target according to the fixed-dimensional input vector format; and the latent variable from the latent distribution; as input to a second attention-based neural network to generate an attention-based decoder;” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): This limitation amounts to generic training of a generic attention-based decoder by applying data as input to a generic second attention-based neural network. The claim denotes generic training and a generic attention-based neural network with no additional details or limitations beyond a generic, off the shelf attention-based network; the specification at paragraphs [0048] and [0069] describes the decoder as a transformer decoder made up of MLPs and multihead attention layers.)
“wherein the attention-based decoder is adapted to:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based decoder as a tool to perform the recited abstract idea of generating predictions.)
“receive representations of the examples of the prediction target according to the fixed-dimensional input vector format; and” (Data Gathering - Mere data gathering recited at a high level of generality, and thus is insignificant extra-solution activity (MPEP 2106.05(g)). -- Examiner’s Note (EN): This limitation amounts to receiving the representations on which the recited abstract idea of generating predictions is performed, which is mere data gathering.)
Step 2B: The claim further recites:
“A method for training a first machine learning model to handle multimodal data including examples with missing modalities, the method comprising:” (The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)). -- Examiner’s Note (EN): The preamble merely limits the recited abstract ideas to the field of multimodal machine learning and to the technological environment of a machine learning model. Applicant’s own specification characterizes the technology as “confined to multimodal machine learning applications” (paragraph [00112]), i.e., a field of use.)
“receiving a plurality of multimodal training data, the training data comprising a plurality of samples of a prediction target, wherein: each sample includes at least a subset of a full set of modalities, wherein the full set of modalities is a plurality of modalities; and the samples collectively include instances of each modality within the full set of modalities;” (MPEP 2106.05(d)(II) indicates that merely receiving or gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
“using the training data as input to a first attention-based neural network, comprising:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based neural network as a tool to perform the recited abstract ideas. No details of the first attention-based neural network are recited beyond that it is attention-based.)
“(...) to generate a respective feature encoder for each modality in the full set of modalities;” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): This limitation amounts to generic training of a generic feature encoder for each modality using the extracted vectors. The specification at paragraph [0064] indicates that the feature encoders are generic, off the shelf networks, e.g., a ConvNet for an image and a gated recurrent unit (GRU) for a sequence of vectors.)
“generating an attention-based encoder that:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based encoder as a tool to perform the recited abstract ideas. The specification at paragraphs [0036], [0065] and [0067] describes the attention-based encoder as a transformer encoder that adopts the known Set Transformer (reference [15]).)
“receives sets of training vectors in the fixed-dimensional input vector format, wherein each set of training vectors represents one of the samples; and” (MPEP 2106.05(d)(II) indicates that merely receiving or gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer. )
“the method further comprising using: representations of the samples of the prediction target according to the fixed-dimensional input vector format; and the latent variable from the latent distribution; as input to a second attention-based neural network to generate an attention-based decoder;” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): This limitation amounts to generic training of a generic attention-based decoder by applying data as input to a generic second attention-based neural network. The claim denotes generic training and a generic attention-based neural network with no additional details or limitations beyond a generic, off the shelf attention-based network; the specification at paragraphs [0048] and [0069] describes the decoder as a transformer decoder made up of MLPs and multihead attention layers.)
“wherein the attention-based decoder is adapted to:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based decoder as a tool to perform the recited abstract idea of generating predictions.)
“receive representations of the examples of the prediction target according to the fixed-dimensional input vector format; and” (MPEP 2106.05(d)(II) indicates that merely receiving or gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 2
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 1 above, which claim 2 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
“wherein the attention-based decoder is part of the first machine learning model;” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The limitation merely specifies that the generic attention-based decoder and the generic attention-based encoder are components of the same generic machine learning model, which is an arrangement of generic computer components and does not impose any meaningful limits on practicing the abstract ideas.)
“the method further comprising using the training data to train the attention-based decoder jointly with generating the attention-based encoder.” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): This limitation amounts to generic training of the generic attention-based decoder, performed at the same time as the generic attention-based encoder is generated. The claim denotes generic joint training with no additional details or limitations beyond generic, off the shelf training of a neural network.)
Step 2B: The claim further recites:
“wherein the attention-based decoder is part of the first machine learning model;” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The limitation merely specifies that the generic attention-based decoder and the generic attention-based encoder are components of the same generic machine learning model, which is an arrangement of generic computer components and does not impose any meaningful limits on practicing the abstract ideas.)
“the method further comprising using the training data to train the attention-based decoder jointly with generating the attention-based encoder.” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): This limitation amounts to generic training of the generic attention-based decoder, performed at the same time as the generic attention-based encoder is generated. The claim denotes generic joint training with no additional details or limitations beyond generic, off the shelf training of a neural network.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 3
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 1 above, which claim 3 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
“the attention-based decoder is part of a second machine learning model that is different from the first machine learning model; and” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The limitation merely specifies that the generic attention-based decoder is a component of a second generic machine learning model, which is a generic computer component.)
“the second machine learning model is trained independently in a separate operation from generating the attention-based encoder.” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): This limitation amounts to generic training of a generic machine learning model, performed as a separate operation. Specifying when the generic training is performed does not impose any meaningful limits on practicing the abstract ideas.)
Step 2B: The claim further recites:
“the attention-based decoder is part of a second machine learning model that is different from the first machine learning model; and” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The limitation merely specifies that the generic attention-based decoder is a component of a second generic machine learning model, which is a generic computer component.)
“the second machine learning model is trained independently in a separate operation from generating the attention-based encoder.” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): This limitation amounts to generic training of a generic machine learning model, performed as a separate operation. Specifying when the generic training is performed does not impose any meaningful limits on practicing the abstract ideas.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 4
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 1 above, which claim 4 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
“wherein the attention-based encoder comprises a plurality of transformer layers.” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim only recites generic, off the shelf transformer layers.)
Step 2B: The claim further recites:
“wherein the attention-based encoder comprises a plurality of transformer layers.” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim only recites generic, off the shelf transformer layers.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 5
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 4 above, which claim 5 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
“wherein each transformer layer comprises a multihead self-attention (MSA) portion, a layer normalization (LN) portion and a multilayer perceptron (MLP) portion applied using residual connections.” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim merely recites the generic, off the shelf components of a standard transformer layer.)
Step 2B: The claim further recites:
“wherein each transformer layer comprises a multihead self-attention (MSA) portion, a layer normalization (LN) portion and a multilayer perceptron (MLP) portion applied using residual connections.” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim merely recites the generic, off the shelf components of a standard transformer layer.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 6
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 1 above, which claim 6 depends on. Claim 6 further recites:
“wherein the fixed-dimensional vector representation template has a dimensionality that is greater than a number of the full set of modalities.” (This limitation is a mental process. A person mentally or with a pen and paper can choose to draw the template with a number of entries that is greater than the number of modalities, e.g., a table having six entries when there are five modalities.)
Step 2A Prong 2: The claim does not recite any additional elements that integrate the judicial exception into a practical application.
Step 2B: The claim does not recite any additional elements that amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 7
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 1 above, which claim 7 depends on. Claim 7 further recites:
“wherein the fixed-dimensional vector representation template has a dimensionality that is fewer than a number of the full set of modalities.” (This limitation is a mental process. A person mentally or with a pen and paper can choose to draw the template with a number of entries that is fewer than the number of modalities, e.g., a table having four entries when there are five modalities.)
Step 2A Prong 2: The claim does not recite any additional elements that integrate the judicial exception into a practical application.
Step 2B: The claim does not recite any additional elements that amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 8
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 1 above, which claim 8 depends on. Claim 8 further recites:
“wherein the fixed-dimensional vector representation template has a dimensionality that is equal to a number of the full set of modalities.” (This limitation is a mental process. A person mentally or with a pen and paper can choose to draw the template with a number of entries that is equal to the number of modalities, e.g., a table having five entries when there are five modalities.)
Step 2A Prong 2: The claim does not recite any additional elements that integrate the judicial exception into a practical application.
Step 2B: The claim does not recite any additional elements that amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 9
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 1 above, which claim 9 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
“wherein the plurality of modalities is at least three modalities.” (Data Gathering - Mere data gathering recited at a high level of generality, and thus is insignificant extra-solution activity (MPEP 2106.05(g)). -- Examiner’s Note (EN): The limitation merely specifies the amount of data gathered, i.e., that the received training data spans at least three modalities, and does not change the receiving step of claim 1 as mere data gathering.)
Step 2B: The claim further recites:
“wherein the plurality of modalities is at least three modalities.” (MPEP 2106.05(d)(II) indicates that merely receiving or gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer. )
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 10
Step 1: The claim recites a computer program product comprising a non-transitory computer-readable medium; therefore, it is directed to the statutory category of manufacture.
Step 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. The claim recites the following abstract ideas:
“(...) processing the training data to extract, for each modality in the full set of modalities, a fixed-dimensional input vector format representing that modality (...);” (This limitation is a mental process. A person mentally or with a pen and paper can look at the training data for each modality, e.g., an image, a trajectory or a body pose, and write down a list of numbers having a fixed length, i.e., a vector of a fixed dimension, that represents that modality.)
“(...) generates, from the training vectors for the samples, a fixed-dimensional vector representation template for the prediction target, wherein the number of dimensions in the fixed-dimensional vector representation template is constant and is independent of the number of modalities represented by the training vectors for the samples” (This limitation is a mental process. A person mentally or with a pen and paper can look at the training vectors for the samples and draw a template for the prediction target, e.g., a table having a fixed number of entries. The wherein clause merely specifies that the person chooses a constant number of entries for the template regardless of how many modalities are present in the training vectors, which is a choice that can be made in the mind.)
“(...) uses the samples and the fixed-dimensional vector representation template to generate, from the training vectors for the samples, a latent distribution;” (This limitation falls within the mathematical concepts grouping because generating a latent distribution from the training vectors amounts to calculating the parameters of a probability distribution, e.g., the mean μ and the variance σ of a Gaussian distribution, see specification paragraphs [0028], [0040] and [0056] and equation (19), which is a mathematical calculation.)
“(...) generate, from the representations of the examples of the prediction target according to the fixed-dimensional input vector format and the latent variable from the latent distribution, predictions for the examples of the prediction target.” (This limitation is a mental process. A person mentally or with a pen and paper can look at the representations of an example, e.g., the past positions of a pedestrian, together with a value drawn from the latent distribution and predict an outcome, e.g., where the pedestrian will be in the future.)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
“A computer program product comprising at least one tangible non-transitory computer-readable medium embodying instructions which, when implemented by at least one processor of a computer, cause the computer to carry out a method (...), the method comprising:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf computer readable medium, processor and computer as tools to perform the recited abstract ideas, see FIG. 6.)
“(...) a method for training a first machine learning model to handle multimodal data including examples with missing modalities (...)” (The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)). -- Examiner’s Note (EN): The preamble merely limits the recited abstract ideas to the field of multimodal machine learning and to the technological environment of a machine learning model. Applicant’s own specification characterizes the technology as “confined to multimodal machine learning applications” (paragraph [00112]), i.e., a field of use.)
“receiving a plurality of multimodal training data, the training data comprising a plurality of samples of a prediction target, wherein: each sample includes at least a subset of a full set of modalities, wherein the full set of modalities is a plurality of modalities; and the samples collectively include instances of each modality within the full set of modalities;” (Data Gathering - Mere data gathering recited at a high level of generality, and thus is insignificant extra-solution activity (MPEP 2106.05(g)). -- Examiner’s Note (EN): This limitation amounts to receiving the training data on which the recited abstract ideas are performed. The wherein clauses merely describe the characteristics of the data received, i.e., that a sample may be missing one or more modalities and that every modality appears in at least one sample, and do not change the nature of the step as mere data gathering.)
“using the training data as input to a first attention-based neural network, comprising:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based neural network as a tool to perform the recited abstract ideas. No details of the first attention-based neural network are recited beyond that it is attention-based.)
“(...) to generate a respective feature encoder for each modality in the full set of modalities;” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): This limitation amounts to generic training of a generic feature encoder for each modality using the extracted vectors. The specification at paragraph [0064] indicates that the feature encoders are generic, off the shelf networks, e.g., a ConvNet for an image and a gated recurrent unit (GRU) for a sequence of vectors.)
“generating an attention-based encoder that:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based encoder as a tool to perform the recited abstract ideas. The specification at paragraphs [0036], [0065] and [0067] describes the attention-based encoder as a transformer encoder that adopts the known Set Transformer (reference [15]).)
“receives sets of training vectors in the fixed-dimensional input vector format, wherein each set of training vectors represents one of the samples; and” (Data Gathering - Mere data gathering recited at a high level of generality, and thus is insignificant extra-solution activity (MPEP 2106.05(g)). -- Examiner’s Note (EN): This limitation amounts to receiving the vectors on which the recited abstract ideas are performed, which is mere data gathering.)
“the method further comprising using: representations of the samples of the prediction target according to the fixed-dimensional input vector format; and the latent variable from the latent distribution; as input to a second attention-based neural network to generate an attention-based decoder;” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“wherein the attention-based decoder is adapted to:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based decoder as a tool to perform the recited abstract idea of generating predictions.)
“receive representations of the examples of the prediction target according to the fixed-dimensional input vector format; and” (Data Gathering - Mere data gathering recited at a high level of generality, and thus is insignificant extra-solution activity (MPEP 2106.05(g)). -- Examiner’s Note (EN): This limitation amounts to receiving the representations on which the recited abstract idea of generating predictions is performed, which is mere data gathering.)
Step 2B: The claim further recites:
“A computer program product comprising at least one tangible non-transitory computer-readable medium embodying instructions which, when implemented by at least one processor of a computer, cause the computer to carry out a method (...), the method comprising:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf computer readable medium, processor and computer as tools to perform the recited abstract ideas, see FIG. 6).
“(...) a method for training a first machine learning model to handle multimodal data including examples with missing modalities (...)” (The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)). -- Examiner’s Note (EN): The preamble merely limits the recited abstract ideas to the field of multimodal machine learning and to the technological environment of a machine learning model. Applicant’s own specification characterizes the technology as “confined to multimodal machine learning applications” (paragraph [00112]), i.e., a field of use.)
“receiving a plurality of multimodal training data, the training data comprising a plurality of samples of a prediction target, wherein: each sample includes at least a subset of a full set of modalities, wherein the full set of modalities is a plurality of modalities; and the samples collectively include instances of each modality within the full set of modalities;” (MPEP 2106.05(d)(II) indicates that merely receiving or gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
“using the training data as input to a first attention-based neural network, comprising:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based neural network as a tool to perform the recited abstract ideas. No details of the first attention-based neural network are recited beyond that it is attention-based.)
“(...) to generate a respective feature encoder for each modality in the full set of modalities;” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): This limitation amounts to generic training of a generic feature encoder for each modality using the extracted vectors. The specification at paragraph [0064] indicates that the feature encoders are generic, off the shelf networks, e.g., a ConvNet for an image and a gated recurrent unit (GRU) for a sequence of vectors.)
“generating an attention-based encoder that:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based encoder as a tool to perform the recited abstract ideas. The specification at paragraphs [0036], [0065] and [0067] describes the attention-based encoder as a transformer encoder that adopts the known Set Transformer (reference [15]).)
“receives sets of training vectors in the fixed-dimensional input vector format, wherein each set of training vectors represents one of the samples; and” (MPEP 2106.05(d)(II) indicates that merely receiving or gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
“the method further comprising using: representations of the samples of the prediction target according to the fixed-dimensional input vector format; and the latent variable from the latent distribution; as input to a second attention-based neural network to generate an attention-based decoder;” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).)
“wherein the attention-based decoder is adapted to:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based decoder as a tool to perform the recited abstract idea of generating predictions.)
“receive representations of the examples of the prediction target according to the fixed-dimensional input vector format; and” (MPEP 2106.05(d)(II) indicates that merely receiving or gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Therefore, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 11
Claim 11 is a computer program product claim that recites the same limitations as claim 2, therefore claim 11 is rejected using the same rationale as claim 2.
Claim 12
Claim 12 is a computer program product claim that recites the same limitations as claim 3, therefore claim 12 is rejected using the same rationale as claim 3.
Claim 13
Claim 13 is a computer program product claim that recites the same limitations as claim 6, therefore claim 13 is rejected using the same rationale as claim 6.
Claim 14
Claim 14 is a computer program product claim that recites the same limitations as claim 7, therefore claim 14 is rejected using the same rationale as claim 7.
Claim 15
Claim 15 is a computer program product claim that recites the same limitations as claim 8, therefore claim 15 is rejected using the same rationale as claim 8.
Claim 16
Step 1: The claim recites a data processing system; therefore, it is directed to the statutory category of machine.
Step 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. The claim recites the following abstract ideas:
“(...) processing the training data to extract, for each modality in the full set of modalities, a fixed-dimensional input vector format representing that modality (...);” (This limitation is a mental process. A person mentally or with a pen and paper can look at the training data for each modality, e.g., an image, a trajectory or a body pose, and write down a list of numbers having a fixed length, i.e., a vector of a fixed dimension, that represents that modality.)
“(...) generates, from the training vectors for the samples, a fixed-dimensional vector representation template for the prediction target, wherein the number of dimensions in the fixed-dimensional vector representation template is constant and is independent of the number of modalities represented by the training vectors for the samples” (This limitation is a mental process. A person mentally or with a pen and paper can look at the training vectors for the samples and draw a template for the prediction target, e.g., a table having a fixed number of entries. The wherein clause merely specifies that the person chooses a constant number of entries for the template regardless of how many modalities are present in the training vectors, which is a choice that can be made in the mind.)
“(...) uses the samples and the fixed-dimensional vector representation template to generate, from the training vectors for the samples, a latent distribution;” (This limitation falls within the mathematical concepts grouping because generating a latent distribution from the training vectors amounts to calculating the parameters of a probability distribution, e.g., the mean μ and the variance σ of a Gaussian distribution, see specification paragraphs [0028], [0040] and [0056] and equation (19), which is a mathematical calculation.)
“(...) generate, from the representations of the examples of the prediction target according to the fixed-dimensional input vector format and the latent variable from the latent distribution, predictions for the examples of the prediction target.” (This limitation is a mental process. A person mentally or with a pen and paper can look at the representations of an example, e.g., the past positions of a pedestrian, together with a value drawn from the latent distribution and predict an outcome, e.g., where the pedestrian will be in the future.)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
“A data processing system comprising at least one processor and memory embodying instructions which, when implemented by the at least one processor, cause the data processing system to carry out a method (...), the method comprising:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf data processing system, processor and memory as tools to perform the recited abstract ideas, see FIG. 6.)
“(...) a method for training a first machine learning model to handle multimodal data including examples with missing modalities (...)” (The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)). -- Examiner’s Note (EN): The preamble merely limits the recited abstract ideas to the field of multimodal machine learning and to the technological environment of a machine learning model. Applicant’s own specification characterizes the technology as “confined to multimodal machine learning applications” (paragraph [00112]), i.e., a field of use.)
“receiving a plurality of multimodal training data, the training data comprising a plurality of samples of a prediction target, wherein: each sample includes at least a subset of a full set of modalities, wherein the full set of modalities is a plurality of modalities; and the samples collectively include instances of each modality within the full set of modalities;” (Data Gathering - Mere data gathering recited at a high level of generality, and thus is insignificant extra-solution activity (MPEP 2106.05(g)). -- Examiner’s Note (EN): This limitation amounts to receiving the training data on which the recited abstract ideas are performed. The wherein clauses merely describe the characteristics of the data received, i.e., that a sample may be missing one or more modalities and that every modality appears in at least one sample, and do not change the nature of the step as mere data gathering.)
“using the training data as input to a first attention-based neural network, comprising:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based neural network as a tool to perform the recited abstract ideas. No details of the first attention-based neural network are recited beyond that it is attention-based.)
“(...) to generate a respective feature encoder for each modality in the full set of modalities;” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): This limitation amounts to generic training of a generic feature encoder for each modality using the extracted vectors. The specification at paragraph [0064] indicates that the feature encoders are generic, off the shelf networks, e.g., a ConvNet for an image and a gated recurrent unit (GRU) for a sequence of vectors.)
“generating an attention-based encoder that:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based encoder as a tool to perform the recited abstract ideas. The specification at paragraphs [0036], [0065] and [0067] describes the attention-based encoder as a transformer encoder that adopts the known Set Transformer (reference [15]).)
“receives sets of training vectors in the fixed-dimensional input vector format, wherein each set of training vectors represents one of the samples; and” (Data Gathering - Mere data gathering recited at a high level of generality, and thus is insignificant extra-solution activity (MPEP 2106.05(g)). -- Examiner’s Note (EN): This limitation amounts to receiving the vectors on which the recited abstract ideas are performed, which is mere data gathering.)
“the method further comprising using: representations of the samples of the prediction target according to the fixed-dimensional input vector format; and the latent variable from the latent distribution; as input to a second attention-based neural network to generate an attention-based decoder;” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“wherein the attention-based decoder is adapted to:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based decoder as a tool to perform the recited abstract idea of generating predictions.)
“receive representations of the examples of the prediction target according to the fixed-dimensional input vector format; and” (Data Gathering - Mere data gathering recited at a high level of generality, and thus is insignificant extra-solution activity (MPEP 2106.05(g)). -- Examiner’s Note (EN): This limitation amounts to receiving the representations on which the recited abstract idea of generating predictions is performed, which is mere data gathering.)
Step 2B: The claim further recites:
“A data processing system comprising at least one processor and memory embodying instructions which, when implemented by the at least one processor, cause the data processing system to carry out a method (...), the method comprising:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf data processing system, processor and memory as tools to perform the recited abstract ideas, see FIG. 6.)
“(...) a method for training a first machine learning model to handle multimodal data including examples with missing modalities (...)” (The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)). -- Examiner’s Note (EN): The preamble merely limits the recited abstract ideas to the field of multimodal machine learning and to the technological environment of a machine learning model. Applicant’s own specification characterizes the technology as “confined to multimodal machine learning applications” (paragraph [00112]), i.e., a field of use.)
“receiving a plurality of multimodal training data, the training data comprising a plurality of samples of a prediction target, wherein: each sample includes at least a subset of a full set of modalities, wherein the full set of modalities is a plurality of modalities; and the samples collectively include instances of each modality within the full set of modalities;” (MPEP 2106.05(d)(II) indicates that merely receiving or gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Therefore, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
“using the training data as input to a first attention-based neural network, comprising:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based neural network as a tool to perform the recited abstract ideas. No details of the first attention-based neural network are recited beyond that it is attention-based.)
“(...) to generate a respective feature encoder for each modality in the full set of modalities;” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): This limitation amounts to generic training of a generic feature encoder for each modality using the extracted vectors. The specification at paragraph [0064] indicates that the feature encoders are generic, off the shelf networks, e.g., a ConvNet for an image and a gated recurrent unit (GRU) for a sequence of vectors.)
“generating an attention-based encoder that:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based encoder as a tool to perform the recited abstract ideas. The specification at paragraphs [0036], [0065] and [0067] describes the attention-based encoder as a transformer encoder that adopts the known Set Transformer (reference [15]).)
“receives sets of training vectors in the fixed-dimensional input vector format, wherein each set of training vectors represents one of the samples; and” (MPEP 2106.05(d)(II) indicates that merely receiving or gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
“the method further comprising using: representations of the samples of the prediction target according to the fixed-dimensional input vector format; and the latent variable from the latent distribution; as input to a second attention-based neural network to generate an attention-based decoder;” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“wherein the attention-based decoder is adapted to:” (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner’s Note (EN): The claim recites a generic, off the shelf attention-based decoder as a tool to perform the recited abstract idea of generating predictions.)
“receive representations of the examples of the prediction target according to the fixed-dimensional input vector format; and” (MPEP 2106.05(d)(II) indicates that merely receiving or gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 17
Claim 17 is a system claim that recites the same limitations as claim 2, therefore claim 17 is rejected using the same rationale as claim 2.
Claim 18
Claim 18 is a system claim that recites the same limitations as claim 3, therefore claim 18 is rejected using the same rationale as claim 3.
Claim 19
Claim 19 is a system claim that recites the same limitations as claim 6, therefore claim 19 is rejected using the same rationale as claim 6.
Claim 20
Claim 20 is a system claim that recites the same limitations as claim 7, therefore claim 20 is rejected using the same rationale as claim 7.
Claim 21
Claim 21 is a system claim that recites the same limitations as claim 8, therefore claim 21 is rejected using the same rationale as claim 8.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Examiner’s Note: Some rejections will include an Examiner’s Note (labeled ‘EN’) to provide additional context or rationale explaining the basis for the rejection.
Claims 1, 2, 4-6, 9-11, 13, 16, 17 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Girgis et al., "Latent Variable Sequential Set Transformers for Joint Multi-Agent Motion Prediction," (hereinafter "Girgis"), in view of Shvetsova et al., "Everything at Once - Multi-modal Fusion Transformer for Video Retrieval," (hereinafter "Shvetsova").
Regarding claim 1, Girgis teaches:
"A method for training a first machine learning model to handle multimodal data ..., the method comprising:" (Page 16, "We train the model using the Adam optimizer"; Page 5, "the benchmark provides access to all neighbouring agents' past trajectories and a birds-eye-view RGB image of the road network" -- EN: the trained AutoBot encoder-decoder network is the first machine learning model, trained on two kinds of input, agent trajectories and a map image, each a modality under BRI in light of spec, so the data is multimodal.)
"receiving a plurality of multimodal training data, the training data comprising a plurality of samples of a prediction target, wherein:" (Page 4, "Given a dataset D = {(X_1, . . . , X_T)}^N_{i=1} consisting of N sequences ... our goal is to maximize the likelihood of the future trajectory of all elements in the set given the input sequence of sets" -- EN: the N training scenes are the samples, each with past trajectories and a map image; the future trajectory is the prediction target.)
"each sample includes at least a subset of a full set of modalities, wherein the full set of modalities is a plurality of modalities; and" (Page 5, same passage; Page 3, Figure 2, the agent tensor "K, M, t" enters the encoder and the "map" enters the decoder -- EN: the full set is {trajectories, map image}, a plurality. Every scene includes both, and a sample with the full set includes "at least a subset" of it; Shvetsova, below, adds samples with only some.)
"the samples collectively include instances of each modality within the full set of modalities;" (Page 5, same passage; Page 17, "We used mirroring of all trajectories and the map information as a data augmentation method" -- EN: every training scene supplies both modalities, so the training set as a whole contains instances of each.)
"using the training data as input to a first attention-based neural network, comprising:" (Page 1, "The encoder is a stack of interleaved temporal and social multi-head self-attention (MHSA) modules"; Page 3, Figure 2, and Page 16, Figure 5, the rFFN embedding, MAB encoder, map CNN and MABD/MAB decoder of one network -- EN: the AutoBot network, built on MHSA and trained end to end on the scenes, is the first attention-based neural network; the two steps below are parts of that training.)
"processing the training data to extract, for each modality in the full set of modalities, a fixed-dimensional input vector format representing that modality to generate a respective feature encoder for each modality in the full set of modalities;" (Page 3, "A row-wise feed-forward network (rFFN) is applied to each row ... transforming vectors of dimension K to d_K"; Page 4, "Additional contextual information (such as a rasterized image of the environment) is encoded using a convolutional neural network to produce a vector of features m_i" -- EN: each modality has its own learned encoder, rFFN for trajectories and CNN for map, each yielding a fixed-length vector. Both are generated by the end-to-end training of the attention network (Pages 4-5, Eqs. 3-5), which is all the claim requires: it does not place the feature encoders inside the attention layers, and the map CNN, whose output enters on the decoder side (Figure 2), is generated by that same training.)
"generating an attention-based encoder that: receives sets of training vectors in the fixed-dimensional input vector format, wherein each set of training vectors represents one of the samples; and" (Page 3, "AutoBot takes as an input a sequence of sets ... the state of a scene evolving over t timesteps"; Page 3, "S_m = MAB(rFFN(X_m))" -- EN: the MAB stack, with the seed-query block of Appendix C.2 that outputs the latent distribution, is the attention-based encoder; it receives the d_K-dimensional trajectory vectors of a scene, which represent that scene, one scene per sample. The map vector enters at the decoder (Figure 2); the combination below feeds it to the encoder too.)
"generates, from the training vectors for the samples, a fixed-dimensional vector representation template for the prediction target, wherein the number of dimensions in the fixed-dimensional vector representation template is constant and is independent of the number of modalities represented by the training vectors for the samples" (Page 16, "we employ c learnable vectors (one for each mode) concatenated into the matrix P ∈ R^(d_K,c), which behave like seeds"; Page 16, "F = MABD(P_{1:c}, C)"; Page 17, Table 4, "c Number of discrete latent variables" = 10, 4, 10, 6, 5, 6, and "d_k Hidden dimension"; Page 13, Eq. 7, "Attn(Q, K, V) = softmax(QK^T)V" -- EN: the learnable seed matrix P, learned during training on the scene vectors, is the fixed-dimensional template for the future-trajectory modes. Under BRI in light of the specification, the "number of dimensions" (dimensionality) of the template is the number of query vectors it holds, i.e., its number of rows.)
"uses the samples and the fixed-dimensional vector representation template to generate, from the training vectors for the samples, a latent distribution;" (Page 16, "Our model also computes the distribution P(Z|X_{1:t}) of discrete random variables"; Page 16, "p(Z|X_{1:t}) = softmax(rLin(F)), where F = MABD(P_{1:c}, C)"; Page 3, "a context tensor C ∈ R^(d_K,M,t) summarizing the entire input scene" -- EN: the seed template P attends over the scene context C, built from the training vectors, to output p(Z|X), the latent distribution over the discrete latent variable Z.)
"the method further comprising using: representations of the samples of the prediction target according to the fixed-dimensional input vector format; and the latent variable from the latent distribution;" (Page 4, "conditioning on the encoder's output C, and the encoded seed parameters and environmental information in H"; Page 4, "each matrix of learnable seed parameters corresponds to a setting of the discrete latent variable in AutoBot" -- EN: C and the map features m_i are the d_K-dimensional representations of the samples; each value of the latent variable Z is supplied to the decoder as its seed Q_i.)
"as input to a second attention-based neural network to generate an attention-based decoder;" (Page 1, "The decoder employs learnable seed parameters in combination with temporal and social MHSA modules"; Page 4, "H'_m = MABD(H_m, C_m)"; Page 5, Eq. 5, one objective over both log p_θ(Y|Z, X_{1:t}) and log p_θ(Z|X_{1:t}) -- EN: the MABD/MAB decoder stack is the second attention-based neural network; training it on C, m_i and each latent setting generates the attention-based decoder. Claim 1 does not require the two networks to be disjoint; claim 2 confirms the decoder may be part of the same model.)
"wherein the attention-based decoder is adapted to: receive representations of the examples of the prediction target according to the fixed-dimensional input vector format; and" (Page 7, Table 2, Argoverse test-set results; Page 16, "this context is copied across the M and T dimensions, and concatenated with the decoder seed parameters during the sequence generation process" -- EN: at test time, held-out scenes (the examples) are encoded into the same d_K-dimensional C and m_i and fed to the decoder.)
"generate, from the representations of the examples of the prediction target according to the fixed-dimensional input vector format and the latent variable from the latent distribution, predictions for the examples of the prediction target." (Page 4, "The output of the decoder is a tensor O ∈ R^(d_K,M,T,c) which can then be processed ... using a neural network φ(.)"; Page 4, "φ produces the parameters of a bivariate Gaussian distribution" -- EN: the predicted future-trajectory distributions are the predictions, one per latent setting, and p(Z|X) identifies the most likely ones.)
Girgis does not explicitly teach:
[to handle multimodal data] ... "including examples with missing modalities,"
"each sample includes at least a subset of a full set of modalities" -- to the extent this requires samples missing a modality. Every Girgis scene has both the trajectories and the map, and the model needs both: to run without the map, Girgis trained a separate "No Map Information" model (Page 18, Table 5).
[a template independent of] "the number of modalities represented by the training vectors for the samples" [where those vectors span more than one modality] -- EN: Girgis's encoder receives only trajectory vectors, the map entering at the decoder (Figure 2), so its fixed-size template never attends over vectors of more than one modality, or of a changing number.
However Shvetsova teaches:
[to handle multimodal data] "including examples with missing modalities," (Page 1, "At test time, the resulting model can process and fuse any number of input modalities"; Page 6, Table 1, rows "Ours tva t→v" and "Ours tva t→va", one model trained on text, video and audio, tested with and without audio -- EN: the trained model is used on examples that lack one of the modalities.)
"each sample includes at least a subset of a full set of modalities" [including samples with missing modalities] (Page 4, "we apply it to joint sets of input tokens from all possible combinations of modalities: singles - t, v, a and pairs - (t, v), (v, a), (t, a)"; Page 5, "Since not every video clip has all three modalities, we computed NCE only over non-empty embeddings"; Page 5, "If the sampled clip contains narration (95% of all clips)" -- EN: the full set is {text, video, audio}; some clips lack the text, so a sample holds two or three modalities, while all three appear across the training set.)
"the number of modalities represented by the training vectors for the samples" [spanning one, two or more modalities] (Page 4, "we create a joint list of tokens, e.g. for v_i a_i: [ν_i1, ..., ν_im, α_i1, ..., α_in]. We apply the transformer to this input"; Page 2, "share key, query, and value weights to all tokens, agnostic of their input modality" -- EN: all vectors of a sample's available modalities go together, as one set, into one attention block whose weights doesn’t depend on modality. The set spans one or two modalities from input to input. In the combination, Girgis's map vector m_i and trajectory vectors enter the encoder as one such set, over which seed matrix P attends and by Eq. 7 the output keeps one row per seed however many vectors, of however many modalities, the set holds.)
Motivation: Before the effective filing date, it would have been obvious to a person having ordinary skill in the art to combine Girgis's attention-based encoder-decoder for motion prediction WITH Shvetsova's way of feeding the vectors of every available modality together into one attention block and training that block on every combination of modalities, so that one trained model accepts any subset. In the combination the map vector m_i joins the trajectory vectors at the input of Girgis's encoder, and training scenes come with and without the map. The motivation for doing so would have been, first, to obtain a single trained model that works with whatever inputs are available, instead of a separate model per input configuration (Girgis trained a separate "No Map Information" model, Page 18, Table 5), the benefit Shvetsova discloses: "the resulting model can process and fuse any number of input modalities" (Page 1). Second, to let the two modalities inform each other inside the encoder, since Shvetsova's block "allows modalities to attend to each other" (Page 2) and its ablation finds "the best performance is achieved in the fusion transformer + comb. loss setup" (Page 8; Page 7, Table 4).
Regarding claim 2, Girgis in view of Shvetsova teaches all the limitations of claim 1. Girgis further teaches:
"wherein the attention-based decoder is part of the first machine learning model;" (Page 1, "We propose Latent Variable Sequential Set Transformers which are encoder-decoder architectures"; Page 3, "Latent Variable Sequential Set Transformers is a class of encoder-decoder architectures that process sequences of sets (see Fig 2)"; Page 17, Table 4, "d_k Hidden dimension throughout all parts of the model" -- EN: the encoder and the decoder are the two halves of the one AutoBot model, which is the first machine learning model.)
"the method further comprising using the training data to train the attention-based decoder jointly with generating the attention-based encoder." (Page 5, Eq. 5, "Q(θ, θ_old) = ... = Σ_Z p_θold(Z|Y, X_{1:t}) [log p_θ(Y|Z, X_{1:t}) + log p_θ(Z|X_{1:t})]"; Page 5, "θ_old corresponds to AutoBot's parameters before performing the parameter update"; Page 16, "We train the model using the Adam optimizer with an initial learning rate" -- EN: one set of parameters θ and one training objective cover both the decoder term log p(Y|Z, X) and the encoder's latent term log p(Z|X), so the decoder is trained together with the encoder in the same training run on the same training data.)
Regarding claim 4, Girgis in view of Shvetsova teaches all the limitations of claim 1. Girgis further teaches:
"wherein the attention-based encoder comprises a plurality of transformer layers." (Page 2, "Multi-Head Attention Blocks (MAB) resemble the encoder proposed in Vaswani et al. (2017)"; Page 3, "the encoder passes the tensor through L repeated layers of multi-head attention blocks (MAB) that are applied to the time axis (time encoding) and the agent axis (social encoding)"; Page 17, Table 4, "L_enc Number of stacked social/Temporal blocks in the encoder" = 2 for Nuscenes, Argoverse, TrajNet and TrajNet++; Page 20, Table 7, "L_enc = 2" and "L_enc = 3" -- EN: each MAB is a transformer layer. The encoder stacks a temporal MAB and a social MAB and repeats the pair two or three times, i.e., a plurality of transformer layers.)
Shvetsova also teaches a fusion transformer built from a plurality of transformer blocks (Page 13, Table 10, "#blocks" = 2 and 4, "#blocks stands for a number of transformer blocks"; Page 13, "the model can further boost performance by increasing the number of transformer blocks").
Regarding claim 5, Girgis in view of Shvetsova teaches all the limitations of claim 4. Girgis further teaches:
"wherein each transformer layer comprises a multihead self-attention (MSA) portion, a layer normalization (LN) portion and a … [feed-forward] portion applied using residual connections." (Page 2, "they consist of the MHSA operation described in Eq. 8 (appendix) followed by a row-wise feed-forward neural network (rFFN), with residual connections and layer normalization (LN) (Ba et al., 2016) after each block"; Page 2, Eq. 1, "MAB(X) = LN(H + rFFN(H)), where H = LN(X + MHSA(X, X, X))" -- EN: MHSA(X, X, X) is the multihead self-attention portion, LN is the layer normalization portion, and "X + MHSA(...)" and "H + rFFN(H)" are the residual connections. The feed-forward sub-block rFFN sits where the claimed MLP sits.)
Girgis does not explicitly teach:
A "multilayer perceptron (MLP) portion" [as the feed-forward sub-block] -- EN: Girgis calls its feed-forward sub-block a "row-wise feed-forward neural network (rFFN)" (Page 2)
However Shvetsova teaches:
"multilayer perceptron (MLP) portion" (Page 4, "we adopt regular transformer blocks [50]. Each transformer block consist of a multiheaded self-attention and a multilayer perceptron (MLP) with two LayerNorm (LN) transforms before them along with two residual connections, as illustrated in Figure 2"; Page 5, "we use one transformer block with a hidden size of 4096, 64 heads, and an MLP size of 4096"; Page 4, Figure 2, LN → Multi-Head Attention → LN → MLP, each with a skip connection -- EN: the transformer block has all three portions, MSA, LN and MLP, joined by two residual connections.)
Motivation: Before the effective filing date, it would have been obvious to a person having ordinary skill in the art to combine Girgis's attention encoder layers, each with multihead self-attention, layer normalization, a feed-forward sub-block and residual connections, WITH Shvetsova's regular transformer block, in which that feed-forward sub-block is a multilayer perceptron. The motivation for doing so would have been to build the encoder from the standard, off-the-shelf transformer block, which needs no custom design and is enough for the network to learn to combine its inputs, as Shvetsova discloses: "we are not proposing a new transformer architecture, but rather refer to it as a transformer that is trained in a way that enables fusion without any need for changes to the self-attention mechanism" (Page 2) and "the resulting fusion can actually be learned by a vanilla transformer block, if it is specifically trained for this task" (Page 4).
Regarding claim 6, Girgis in view of Shvetsova teaches all the limitations of claim 1. Girgis further teaches:
"wherein the fixed-dimensional vector representation template has a dimensionality that is greater than a number of the full set of modalities." (Page 16, "we employ c learnable vectors (one for each mode) concatenated into the matrix P ∈ R^(d_K,c)"; Page 17, Table 4, "c Number of discrete latent variables" = 10, 4, 10, 6, 5, 6 -- EN: as read for claim 1, the dimensionality of the template is its number of seed (query) vectors. P holds c seed vectors, so its dimensionality is c = 4 to 10, which is greater than Girgis's two modalities (trajectories and map) and than Shvetsova's three (text, video and audio). The hidden size d_K of each seed vector (64 or 128, Table 4) is the width of the vectors, not the number of dimensions of the template, and is not relied on.)
Regarding claim 9, Girgis in view of Shvetsova teaches all the limitations of claim 1.
Shvetsova further teaches:
"wherein the plurality of modalities is at least three modalities." (Page 1, "multiple modalities, such as video, audio, and text"; Page 3, "we consider three modalities, video, audio, and text (corresponding ASR caption or linguistic narration); but the proposed method can be extended to more modalities"; Page 4, "in case of four modalities, we would consider all combinations up to a triplet (t, v, a) during training" -- EN: this denotes three modalities, with more than 3 modalities mentioned.)
Motivation: Before the effective filing date, it would have been obvious to a person having ordinary skill in the art to combine Girgis's motion prediction model, which draws on two kinds of input (past trajectories and a map), WITH Shvetsova's fusion of three or more modalities in one attention model. The motivation for doing so would have been to let the prediction model use every kind of input that is available for a scene, which Shvetsova shows improves results when a third modality is added on top of two: "after finetuning on the MSR-VTT, the model greatly improves performance by utilizing an audio channel" (Page 7). Shvetsova states the underlying benefit: "Humans capture their world in various ways, combining different sensory input modalities such as vision, sound, touch, and more, to make sense of their environment" (Page 1).
Regarding claim 10,
Girgis teaches:
"A computer program product comprising at least one tangible non-transitory computer-readable medium embodying instructions which, when implemented by at least one processor of a computer, cause the computer to carry out a method for training a first machine learning model ..." (Page 16, "We implemented our model using the Pytorch open-source framework"; Page 17, "approximately 3 hours of compute time on a single Nvidia Geforce GTX 1080Ti GPU, using approximately 2 GB of VRAM"; Page 1, code released at https://fgolemo.github.io/autobots/ -- EN: the model is software (PyTorch code) that is stored and run on a computer with a processor (the GPU) and memory (VRAM); the released code is instructions stored on a computer-readable medium. Shvetsova likewise releases its software and runs it on GPUs (Page 1, footnote 1, "Our code for this work is also available"; Page 15, "four Nvidia V100 32GB GPUs")
The remaining limitations of claim 10 are substantially the same as those of claim 1 and therefore are rejected for the same rationale.
Motivation: Before the effective filing date, it would have been obvious to a person having ordinary skill in the art to combine Girgis's attention-based encoder-decoder for motion prediction,implemented as PyTorch instructions run by a GPU-equipped computer, which encodes its trajectory and map inputs with separate feature encoders, attends over them with a fixed-size learnable seed matrix to compute a distribution over future modes, and decodes future trajectories, WITH Shvetsova's way of training an attention model on every combination of its modalities so that one trained model accepts any subset of them. The motivation for doing so would have been to obtain a single trained model that works with whatever inputs happen to be available, instead of a separate model for each input configuration (Girgis needed a separately trained "No Map Information" model to run without the map, Page 18, Table 5). This is the benefit Shvetsova discloses: "At test time, the resulting model can process and fuse any number of input modalities" (Page 1), curing the shortcoming that "so far, none of these transformers allow for adaption to any given number of input modalities" (Pages 1-2), in a setting where, as in real data, "not every video clip has all three modalities" (Page 5).
Regarding claim 11, Girgis in view of Shvetsova teaches all the limitations of claim 10. Claim 11 recites substantially the same limitations of claim 2 in computer program product form and is rejected under the same rationale as claim 2.
Regarding claim 13, Girgis in view of Shvetsova teaches all the limitations of claim 10. Claim 13 recites substantially the same limitations of claim 6 in computer program product form and is rejected under the same rationale as claim 6.
Regarding claim 16, Girgis in view of Shvetsova teaches all the limitations of claim 1 (the method). Girgis further teaches:
"A data processing system comprising at least one processor and memory embodying instructions which, when implemented by the at least one processor, cause the data processing system to carry out a method for training a first machine learning model ..." (Page 16, "We implemented our model using the Pytorch open-source framework"; Page 17, "a single Nvidia Geforce GTX 1080Ti GPU, using approximately 2 GB of VRAM" -- EN: a computer with a processor (the GPU) and memory (VRAM) holding and executing the software that trains the model is a data processing system.)
The remaining limitations of claim 16 are substantially the same as those of claim 1 and therefore are rejected for the same rationale.
Motivation: Before the effective filing date, it would have been obvious to a person having ordinary skill in the art to combine Girgis's attention-based encoder-decoder for motion prediction,implemented as a data processing system, which encodes its trajectory and map inputs with separate feature encoders, attends over them with a fixed-size learnable seed matrix to compute a distribution over future modes, and decodes future trajectories, WITH Shvetsova's way of training an attention model on every combination of its modalities so that one trained model accepts any subset of them. The motivation for doing so would have been to obtain a single trained model that works with whatever inputs happen to be available, instead of a separate model for each input configuration (Girgis needed a separately trained "No Map Information" model to run without the map, Page 18, Table 5). This is the benefit Shvetsova discloses: "At test time, the resulting model can process and fuse any number of input modalities" (Page 1), curing the shortcoming that "so far, none of these transformers allow for adaption to any given number of input modalities" (Pages 1-2), in a setting where, as in real data, "not every video clip has all three modalities" (Page 5).
Regarding claim 17, Girgis in view of Shvetsova teaches all the limitations of claim 16. Claim 17 recites substantially the same limitations of claim 2 in system form and is rejected under the same rationale as claim 2.
Regarding claim 19, Girgis in view of Shvetsova teaches all the limitations of claim 16. Claim 19 recites substantially the same limitations of claim 6 in system form and is rejected under the same rationale as claim 6.
Claims 3, 12 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Girgis in view of Shvetsova, and further in view of Lin et al., "Frozen CLIP Models are Efficient Video Learners," (hereinafter "Lin").
Regarding claim 3, Girgis in view of Shvetsova teaches all the limitations of claim 1.
Girgis does not explicitly teach:
"wherein: the attention-based decoder is part of a second machine learning model that is different from the first machine learning model; and the second machine learning model is trained independently in a separate operation from generating the attention-based encoder." (Girgis trains its encoder and decoder together as one model, Pages 4-5, Eqs. 3-5. Shvetsova has no prediction decoder; its embeddings are used by dot-product retrieval, Page 6.)
However Lin teaches:
"wherein: the attention-based decoder is part of a second machine learning model that is different from the first machine learning model; and the second machine learning model is trained independently in a separate operation from generating the attention-based encoder." (Abstract, "an efficient framework for directly training high-quality video recognition models with frozen CLIP features"; Abstract, "we employ a lightweight Transformer decoder and learn a query token to dynamically collect frame-level spatial features from the CLIP image encoder"; Section 3.1, "The overall structure of EVL, as illustrated in Fig. 2, is a multi-layer spatiotemporal Transformer decoder on top of a fixed CLIP backbone"; Section 3.1, "A linear layer projects the output of the last decoder block to class predictions"; Section 3.1, "the back-propagation stops at image features X and no weight in the image encoder is updated"; Section 2, "CLIP [36] and ALIGN [24] pretrain vision-language models with a contrastive loss on large-scale datasets consisting of open-vocabulary image-text pairs"; Table 4, "ViT-B/16" backbone -- EN: the CLIP image encoder (a Vision Transformer, i.e., an attention-based encoder) was generated earlier, in its own contrastive pre-training, and is kept fixed. The Transformer decoder that produces the predictions is a different model, trained afterwards in a separate operation in which no weight of the encoder is changed.)
Motivation: Before the effective filing date, it would have been obvious to a person having ordinary skill in the art to combine Girgis's attention-based encoder and attention-based decoder for motion prediction (as modified by Shvetsova) WITH Lin's practice of training the Transformer decoder as a separate model on top of a previously trained, fixed attention encoder. The motivation for doing so would have been to cut the memory and time needed to train the decoder, because no gradients have to flow back through the encoder, a benefit Lin discloses: "we can completely avoid back-propagation through the backbone. This vastly reduces both the memory consumption and the time per training iteration" (Section 3.3), with Lin's model reaching the same accuracy as an end-to-end model at 60 GPU-hours instead of 5,000 (Table 4).
Regarding claim 12, Girgis in view of Shvetsova teaches all the limitations of claim 10. Claim 12 recites the limitations of claim 3 in computer program product form and is rejected over Girgis in view of Shvetsova and further in view of Lin for the same reasons given for claim 3.
Regarding claim 18, Girgis in view of Shvetsova teaches all the limitations of claim 16. Claim 18 recites the limitations of claim 3 in system form and is rejected over Girgis in view of Shvetsova and further in view of Lin for the same reasons given for claim 3.
Claims 7, 8, 14, 15, 20 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Girgis in view of Shvetsova, and further in view of Lee et al., "Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks," (hereinafter "Lee").
Regarding claim 7, Girgis in view of Shvetsova teaches all the limitations of claim 1.
Girgis does not explicitly teach:
"wherein the fixed-dimensional vector representation template has a dimensionality that is fewer than a number of the full set of modalities." (Read as in claims 1 and 6, the dimensionality of Girgis's template P is its number of seed vectors, c = 4 to 10, Page 17, Table 4, which is more than its two modalities and more than Shvetsova's three.)
However Lee teaches:
"wherein the fixed-dimensional vector representation template has a dimensionality that is fewer than a number of the full set of modalities." (Section 3.2, "applying multihead attention on a learnable set of k seed vectors S ∈ R^(k×d)"; Section 3.2, "Pooling by Multihead Attention (PMA) with k seed vectors is defined as PMA_k(Z) = MAB(S, rFF(Z))"; Section 3.2, "Note that the output of PMA_k is a set of k items. We use one seed vector (k = 1) in most cases"; Section 1, "such a model should be able to process input sets of any size"; Section 3, "our aggregating function pool(·) is parameterized and can thus adapt to the problem at hand" -- EN: Lee's seed matrix S is the same kind of fixed-size learnable query template as Girgis's P; Girgis itself calls P "learnable vectors ... which behave like seeds" (Page 16) and cites Lee. Lee's S ∈ R^(k×d) has the same shape as the applicant's Q ∈ R^(N×d) Para [0031]): k seed vectors, each of width d. Read as in claims 1 and 6, the dimensionality of the template is therefore k. Lee uses k = 1 in most cases, a dimensionality of one, which is fewer than any plurality of modalities in the full set (two in Girgis, three in Shvetsova), whatever the width d of that one vector.)
Motivation: Before the effective filing date, it would have been obvious to a person having ordinary skill in the art to combine Girgis's seed-query attention block, which turns the encoded scene into a latent distribution (as modified by Shvetsova), WITH Lee's single-seed (k = 1) pooling. The motivation for doing so would have been to size the template by the number of outputs actually needed, here one pooled summary of the input set from which the distribution over modes is read out (Girgis, Page 16, "rLin is a row-wise linear projection layer to a vector of size c"), while still handling input sets of any size with a pooling that adapts to the data, as Lee discloses: "We use one seed vector (k = 1) in most cases" (section 3.2) and "our aggregating function pool(·) is parameterized and can thus adapt to the problem at hand" (section 3).
Regarding claim 8, Girgis in view of Shvetsova teaches all the limitations of claim 1.
Girgis does not explicitly teach:
"wherein the fixed-dimensional vector representation template has a dimensionality that is equal to a number of the full set of modalities." (Read as in claims 1 and 6, the dimensionality of Girgis's template is its number of seed vectors, c = 4 to 10, Page 17, Table 4, against two modalities.)
However Lee teaches:
"wherein the fixed-dimensional vector representation template has a dimensionality that is equal to a number of the full set of modalities." (Section 3.2, "Note that the output of PMA_k is a set of k items"; Section 3.2, "for problems such as amortized clustering which requires k correlated outputs, the natural thing to do is to use k seed vectors"; Section 5.3, "We used four seed vectors for the PMA (S ∈ R^(4×d)) so that each seed vector generates the parameters of a cluster" -- EN: Lee teaches that the template holds one seed vector per output wanted, so its dimensionality (its number of seed vectors, read as in claims 1 and 6) is set by the number of outputs. Where one output per modality of the full set is wanted, the template has exactly as many seed vectors as there are modalities, i.e., a dimensionality equal to the number of the full set of modalities (k = 2 for Girgis's two, k = 3 for Shvetsova's three).)
Motivation: Before the effective filing date, it would have been obvious to a person having ordinary skill in the art to combine Girgis's seed-query attention block (as modified by Shvetsova) WITH Lee's rule of using one seed vector per output wanted, giving the template one seed vector for each modality of the full set. The motivation for doing so would have been to obtain one pooled representation per modality, which Shvetsova already produces and uses: "As a result, we obtain a vector representation for each modality included in this computation" (Page 5); Lee discloses that when k correlated outputs are wanted "the natural thing to do is to use k seed vectors" (Section 3.2), and that the parameterized pooling "can thus adapt to the problem at hand" (Section 3).
Regarding claims 14 and 15, Girgis in view of Shvetsova teaches all the limitations of claim 10. Claims 14 and 15 recite the limitations of claims 7 and 8 in computer program product form and are rejected over Girgis in view of Shvetsova and further in view of Lee for the same reasons given for claims 7 and 8.
Regarding claims 20 and 21, Girgis in view of Shvetsova teaches all the limitations of claim 16. Claims 20 and 21 recite the limitations of claims 7 and 8 in system form and are rejected over Girgis in view of Shvetsova and further in view of Lee for the same reasons given for claims 7 and 8.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NAYMUR RAHMAN ALI whose telephone number is (571)272-0007. The examiner can normally be reached Mon-Fri. 9:30-6:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NAYMUR RAHMAN ALI/Examiner, Art Unit 2123
/ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123