Prosecution Insights
Last updated: August 16, 2026
Application No. 17/972,703

METHOD AND APPARATUS FOR ANALYZING MULTIMODAL DATA

Non-Final OA §101§103
Filed
Oct 25, 2022
Priority
Oct 26, 2021 — RE 10-2021-0143791
Examiner
RODEN, DONALD THOMAS
Art Unit
2128
Tech Center
2100 — Computer Architecture & Software
Assignee
Samsung SDS Co., Ltd.
OA Round
2 (Non-Final)
20%
Grant Probability
At Risk
2-3
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants only 20% of cases
20%
Career Allowance Rate
1 granted / 5 resolved
-35.0% vs TC avg
Strong +100% interview lift
Without
With
+100.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 9m
Avg Prosecution
21 currently pending
Career history
30
Total Applications
across all art units

Statute-Specific Performance

§101
34.0%
-6.0% vs TC avg
§103
48.9%
+8.9% vs TC avg
§102
5.0%
-35.0% vs TC avg
§112
5.7%
-34.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 5 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is made final. This office action is in response to the amendments filed on March 13, 2026. Claims 1, 3, 4, 6, 7, 8, 11, 13, 14, 16, 17, and 18 have been amended. Claims 2, 5, 9, 10, 12, 15, 19, and 20 have been cancelled. Response to Amendment The amendments filed March 13, 2026 has been entered. Claims 1, 3, 4, 6, 7, 8, 11, 13, 14, 16, 17, and 18 remain pending in the case. Response to Arguments Regarding the 101 Arguments Applicant's arguments filed March 13, 2026 have been fully considered but they are not persuasive. Applicant argues: Step 2A, Prong One Argument Regarding Step 2A, Prong One, "Reminders on evaluating subject matter eligibility of claims under 35 U.S.C. 101" (August 4, 2025, available at https://www.uspto.gov/sites/default/files/documents/memo-101-20250804.pdf) (hereinafter "Reminders") specifically recite, "The mental process grouping is not without limits. Examiners are reminded not to expand this grouping in a manner that encompasses claim limitations that cannot practically be performed in the human mind. The MPEP and the AI-SME Update provide examples of claim limitations that cannot be practically performed in the human mind. Claim limitations that encompass AI in a way that cannot be practically performed in the human mind do not fall within this grouping." (Emphasis added). Thus, the Examiner's arguments about the claimed AI-based inventions are entirely different from the above specific instructions of the Reminders. Examiner Response: The rejection does not assert that the specific recited AI techniques themselves are performed mentally. Rather, the claims are directed to concepts of generating, selecting, ranking, and combining data representations, which are forms of data analysis that can be performed mentally or with pen and paper. The recited neural networks and self-attention are generic computational tools used to implement these abstract data processing concepts. Accordingly, the claims still recite a mental process or an abstract idea under Step 2A, Prong One. Applicant argues: Further, the Reminders states, "Examiners should be careful to distinguish claims that recite an exception (which require further eligibility analysis) from claims that merely involve an exception (which are eligible and do not require further eligibility analysis)" "Consider for example, the published USPTO examples 39, which illustrates claim limitations that merely involve an abstract idea, and 47, which shows limitations that recite an abstract idea. The claim limitation "training the neural network in a first stage using the first training set" of example 39 does not recite a judicial exception. Even though "training the neural network" involves a broad array of techniques and/or activities that may involve or rely upon mathematical concepts, the limitation does not set forth or describe any mathematical relationships, calculations, formulas, or equations using words or mathematical symbols." Claim 1 includes elements related to training the image processor, the text processor, and the encoder, which are similar in nature to the limitation of training the neural network of example 39 explained in the Reminders. Thus, Applicant respectfully submits that the Examiner's arguments should be withdrawn in view of the Reminders. Examiner Response: The rejection is not based solely on the presence of a “training” limitation. Rather, the claims as a while recite generating embeddings, selecting and ranking features, concatenating vectors, and producing representations based on relationships between data, which constitute abstract data processing. Example 39 is fact-specific and does not stand for the proposition that all claims involving training are per se non-abstract. Here, the training is part of a broader abstract data analysis framework and does not remove the claims form the abstract idea category. Applicant argues: Second, Applicant cites Ex Partes Desjardins (September 26, 2025, available at https://www.uspto.gov/sites/default/files/documents/202400567-arp-rehearing-decision- 20250926.pdf) (hereinafter "Desjardins"). In Desjardins, the USPTO director Squires, along with the Acting Commissioner Wallace, and Judge Kim, specifically held that, "Categorically excluding Al innovations from patent protection in the United States jeopardizes America's leadership in this critical emerging technology. Yet, under the panel's reasoning, many Al innovations are potentially unpatentable- even if they are adequately described and nonobvious-because the panel essentially equated any machine learning with an unpatentable "algorithm" and the remaining additional elements as "generic computer components," without adequate explanation. [] Examiners and panels should not evaluate claims at such a high level of generality." Here, just as in Desjardins, the Examiner evaluates the claims of the present application "at such a high level of generality" without adequate explanation. Thus, based on Desjardins, Applicant respectfully submits that the Examiner's § 101 rejections should be withdrawn. Examiner Response: The rejection is based on the actual recited claim limitations, including generating activation maps, calculating feature values, selecting data based on those values, concatenating embeddings, and generating a representation using self-attention. These limitations were considered collectively and are reasonably characterized as stat analysis and manipulation of information. The identification of the abstract idea is therefore not an overgeneralization, but a fair characterization of the claimed subject matter. Applicant argues: Step 2A, Prong Two Argument In view of the 2019 Revised Patent Subject Matter Eligibility Guidance (hereinafter, "Revised Guidance") and MPEP § 2106.05(a), Applicant respectfully submits that the present claims are not under the "Mental Processes" grouping as alleged by the Examiner. Rather, the claims recite "an additional element reflects an improvement in the functioning of a computer, or an improvement to other technology or technical field." The application explains in detail in the Background section the technical problems and challenges in the field of a technology for analyzing multimodal data. See paras. 0003-0004 of the specification as filed reproduced below. BACKGROUND [0003] In multimodal representation learning, according to the related art, object detection (for example, R-CNN) is mainly utilized to extract features based on region of interest (Rol) of objects included in an image, and the extracted features are used in image embedding. [0004] However, such a method is significantly dependent on the object detection, so that R-CNN trained for each domain is required. In this case, a label (for example, a bounding box) for an object detection test is additionally required to train R-CNN. To solve this problem, the claimed invention provides an apparatus and a method for analyzing multimodal data, represented by independent claims. For example, claim 1 recites (emphasis added): An apparatus for analyzing multimodal data, the apparatus comprising: an image processor configured to generate an activation embedding vector based on an index of an activation map obtained from image data through a convolutional neural network; a text processor configured to receive text data to generate a text embedding vector; a vector concatenator configured to concatenate the activation embedding vector and the text embedding vector to each other to generate a concatenated embedding vector; and an encoder configured to generate a multimodal representation vector in consideration of an influence between elements constituting the concatenated embedding vector based on self-attention, wherein at least one of the image processor, the text processor, the vector concatenator, and the encoder comprises a hardware, wherein the image processor is configured to generate the activation embedding vector by performing: generating an activation map set comprising a plurality of activation maps for the image data, using a synthetic neural network; calculating a feature value for each of the plurality of activation maps, and generating an index vector including indices of one or more activation maps selected among the plurality of activation maps in an order of a descending feature value; and embed the index vector to generate the activation embedding vector, wherein the image processor, the text processor, and the encoder are configured to be trained based on the same loss function, and wherein the loss function is calculated based on an image-text matching (ITM) loss function, a masked language modeling (MLM) loss function, and a masked activation modeling (MAM) loss function. Based on the above features, example embodiments of claimed invention achieve the technical benefits of performing more delicate multimodal expression at a higher speed than the prior art approaches. See e.g., para. 0085 of the specification as filed. Because the additional elements of the claims provide a solution to the technical problem explained in the application, Applicant respectfully submits that the claims integrate the alleged abstract idea into a practical application, and are patent eligible under Prong Two of Step 2A. Examiner Response: The alleged improvement is directed to how data is processed and represented (i.e., generating and combining embeddings) rather than an improvement to the functioning of a computer or another technology. The claims do not recite a specific improvement to the functioning of a computer or another technology. The claims do not recite a specific improvement to computer architecture, processor operation, or neural network structure itself. Instead, they use known components (e.g., CNNs, self-attention mechanisms, and loss functions) as tools to perform abstract data processing. Accordingly, the claims do not integrate the abstract idea into a practical application under Step 2A, Prong Two. The additional elements, including processors, neural networks, and training mechanisms, are recited at a high level of generality and perform their conventional functions. Therefore, the claims do not include significantly more than the abstract idea. Regarding the 103 arguments Applicant's arguments filed March 13, 2026 have been fully considered but they are not persuasive. Applicant argues: The Examiner asserts that Kiros at Col. 5, lines 12-49 and FIG. 2 discloses generating embedding vectors corresponding to images selected form a ranked list of image search results, wherein each embedding correspond to a respective index position, which corresponds to "generate an activation embedding vector based on an index vector." However, Kiros does not disclose anything about an index of an item or an index position, much less "an index of an activation map" or "indices of one or more activation maps." Further, the ranked list of image search results is based on how image search results are classified by the image search engine as being responsive to the search query, that is, the search algorithm of the image search engine. It does not have anything to do with calculating a feature value for each of the plurality of activation maps and selecting one or more activation maps among the plurality of activation maps in an order of a descending feature value. Examiner Response: This argument is not persuasive because the rejection does not rely on Kiros alone for the entirety of these limitations. Kiros is relied upon for teaching generation of embeddings form CNN-derived image information while preserving ordering/index information, including ordered image search results and concatenation of image embeddings according to that ordering. Zhou is relied upon for teaching activation maps, global average pooling to obtain feature values, and class-specific importance of feature maps/units. Thus, the combined teachings of Kiros and Zhou, in view of Liu’s multimodal image/text processing system, teach or suggest generating an activation embedding vector based on ordered/indexed activation-map information, including selection based on feature/importance values. Applicant argues: In addition, the Examiner takes the position that Zhou at pages 2-3, Section 2 describes identifying the activation maps with the largest pooled values Fk and corresponding importance weights, selecting the top contributing activation maps for generating the class activation map. See page 19 in the rejection of claim 4. However, this is clearly wrong as there is no such disclosure at all in Zhou. Rather, Zhou discloses at pages 2-3, Section 2 that for a given class c, the input to the softmax, Sc, is as follows (see also FIG. 2 reproduced below): That is, in Zhou, the activation maps with the largest pooled values Fk are not selected but a unit activated by some visual pattern (or weighted activation maps) is selected as the class activation map. See e.g., discriminative image regions such as the head of the animal or the plates in barbell as shown in FIG. 3. That is, Zhou uses the softmax weights W1 ... W to rank the units for a given class. See page 8, Section 5. Selecting an activation map with the largest pooled value Fk, as alleged by the Examiner, would forfeit the purpose of Zhou as it cannot accurately represent the presence of the visual pattern for the given class. Examiner Response: Applicant argues that Zhou does not disclose selecting activation maps in descending feature-value order and that selecting activation maps with the largest pooled values would forfeit the purpose of Zhou. This argument is not persuasive. Zhou teaches generating class activation maps from convolutional feature maps using global average pooling, where the result of global average pooling, Fk, is obtained for each unit/feature map, and where class specific weights indicate the importance of Fk for a given class. Zhou further teaches that the weighted activation maps identify discriminative image regions used by the CNN for classification. This, Zhou teaches calculating values for activation maps and determining the relative importance/contribution of such maps. In view of Kiro’s teaching of preserving ordered/indexed visual embedding information, it would have been obvious to select and order activation maps according to their calculated feature/importance values to generate an index-based activation embedding vector. Applicant argues: Furthermore, claim 1 recites that the image processor, the text processor, and the encoder are configured to be trained based on the same loss function, and the loss function is calculated based on an image-text matching (ITM) loss function, a masked language modeling (MLM) loss function, and a masked activation modeling (MAM) loss function. This subject matter incorporates claim 10, which is now cancelled. In the rejection of claim 10, the Examiner simply relied on the same premise upon which claims 6-8 are rejected. However, the subject matter of claim 10 cannot be equated as claims 6-8 because claim 10 requires jointly training the image processor, the text processor, and the encoder based on the same loss function. The Examiner cites to Reed at para. 0065 as teaching claim 9, which recited "wherein the image processor, the text processor, and the encoder are configured to be trained based on the same loss function." However, Reed at para. 0065 merely discloses that in addition to training the encoder neural network 102, the training system 100 may jointly train: (i) the context neural network 112, and (ii) the parameter values of the transformations applied to the context latent representations. This is not equivalent to "wherein the image processor, the text processor, and the encoder are configured to be trained based on the same loss function." Furthermore, Reed does not disclose "wherein the image processor, the text processor, and the encoder are configured to be trained based on the same loss function, which is calculated based on the ITM loss function, the MLM loss function, and the MAM loss function." Accordingly, Applicant respectfully submits that allowance of claim 1 and all claims dependent therefrom is warranted for these reasons. Allowance of independent claim 11 is warranted for reasons similar to those provided with respect to claim 1. Examiner Response: Applicants’ argument does not persuasive because the rejection does not rely on Reed alone to teach the entire loss function limitation. Reed is relied upon for teaching joint training of neural network components using a common objective/loss. Reed discloses that, in addition to training the encoder neural network, the system may jointly train the context neural network and transformation parameters using the loss/objective function. Lee teaches a multimodal vision language transformer trained using masked language modeling and sentence image prediction/image text matching objectives. Bao teacher masked image modeling, in which image patches are masked and the model is trained to recover corresponding visual tokens from the corrupted image. Therefore, the combined teachings of Reed, Li, and Bao teach or at least suggest training multimodal neural network components using a shared combined lost calculated Based on ITM, MLM, and MAM objectives, One of ordinary skill in the art at the time of the claimed invention would have been motivated to use these teachings in a common training objected for the joint trained multimodal components as each objective trains a complementary aspect of the multimodal representation such as cross modal image text alignment, text side mass prediction, and image side mass prediction. Combining these known pre training objectives would have predictably improved the learning multimodal representation. Independent claim 11 recites method limitations corresponding to the apparatus limitations of claim 1. Accordingly, applicants’ arguments regarding claim 11 are not necessarily set for the same reasons as described above Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. To determine if a claim is directed to patent ineligible subject matter, the Court has guided the Office to apply the Alice/Mayo test, which requires: Step 1: Determining if the claim falls within a statutory category. Step 2A: Determining if the claim is directed to a patent ineligible judicial exception consisting of a law of nature, a natural phenomenon, or abstract idea; and Step 2A is a two prong inquiry. MPEP 2106.04(II)(A). Under the first prong, examiners evaluate whether a law of nature, natural phenomenon, or abstract idea is set forth or described in the claim. Abstract ideas include mathematical concepts, certain methods of organizing human activity, and mental processes. MPEP 2104.04(a)(2). The second prong is an inquiry into whether the claim integrates a judicial exception into a practical application. MPEP 2106.04(d). Step 2B: If the claim is directed to a judicial exception, determining if the claim recites limitations or elements that amount to significantly more than the judicial exception. (See MPEP 2106). Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: Claims 1-10 are directed to an apparatus, comprising computer hardware and software (a machine), and Claims 11-20 are directed to a method (a process). Therefore, Claims 1-20 are directed to a process, machine or manufacture or composition of matter. Regarding claim 1 Step 2A Prong 1 Claim 1 recites the following mathematical concepts, that in each case under the broadest reasonable interpretation, covers performance of mathematical relationships, mathematical formulas or equations, and mathematical calculations but for recitation of generic computer components (e.g., an image processor”, “activation embedding vector “, “convolutional neural network”, “text processor”, “vector concatenator”, “encoder”) [see MPEP 2106.04(a)(2)(I)]. “an image processor configured to generate an activation embedding vector”( e.g., a mathematical technique of linear algebra and numeric manipulation of data) “generate a text embedding vector” (e.g., a mathematical technique to assign strings to a numeric representation) “a vector concatenator configured to concatenate the activation embedding vector and the text embedding vector to each other to generate a concatenated embedding vector”(e.g., combining two vectors to generate a third) “an encoder configured to generate a multimodal representation vector in consideration of an influence between elements constituting the concatenated embedding vector based on self-attention,” (e.g., processing numeric vectors using mathematical formulas (self-attention)) “generating an activation map set comprising a plurality of activation maps for the image data, using a synthetic neutral network” (e.g., applying a neural network function to compute multidimensional numeric arrays) “calculating a feature value for each of the plurality of activation maps” (e.g., constructing an index vector by mapping a subset selection) “generating an index vector including indices of one or more activation maps selected among the plurality of activation maps in an order of a descending feature value” (e.g., constructing an index vector by mapping a subset selection and ranking/sorting data according to calculated values) “wherein the image processor is configured to embed the index vector to generate an activation embedding vector” (e.g., mathematical operation that maps discrete indices to a real-valued numeric vector) Accordingly, at Step 2A, prong one, the claim recites an abstract idea. Step 2A Prong 2 The judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of “an image processor”, “activation embedding vector “, “convolutional neutral network”, “text processor”, “vector concatenator”, “encoder”, and “wherein at least one of the image processor, the text processor, the vector concatenator, and the encoder comprises a hardware” which are recited at a high-level of generality such that they amount to no more than mere instructions to apply the exception using generic computer components (See MPEP 2106.05(f)). The Examiner notes that this is used throughout the claim limitations, and is rejected thusly for each claim which recites the same language. Regarding the “based on an index of an activation map obtained from image data through a convolutional neutral network”, and “a text processor configured to receive text data” limitations, these additional elements are recited at a high-level of generality and amounts to extra-solution activity of obtaining data to input for a model, i.e., pre-solution activity of data gathering (e.g., obtaining information for processing in a computer system (see MPEP 2106.05(g)). Regarding the “wherein the image processor, the text processor, and the encoder are configured to be trained based on the same loss function” which is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components (See MPEP 2106.05(f)). Regarding the “wherein the loss function is calculated based on an image-text matching (ITM) loss function, a masked language modeling (MLM) loss function, and a masked activation modeling (MAM) loss function” which is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components (See MPEP 2106.05(f)). Accordingly, at Step 2A, prong two, the additional elements individually or in combination do not integrate the judicial exception into a practical application. Step 2B In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. As discussed above, the additional element of a “an image processor”, “activation embedding vector “, “convolutional neutral network”, “text processor”, “vector concatenator”, “encoder”, and “wherein at least one of the image processor, the text processor, the vector concatenator, and the encoder comprises a hardware” which are recited at a high-level of generality such that they amount to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)). Regarding the “based on an index of an activation map obtained from image data through a convolutional neutral network”, and “a text processor configured to receive text data” limitations, these additional elements are recited at a high-level of generality and amounts to extra-solution activity of obtaining data to input for a model, i.e., pre-solution activity of data gathering. The courts have found limitations directed to obtaining information electronically, recited at a high-level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory"). Regarding the “wherein the image processor, the text processor, and the encoder are configured to be trained based on the same loss function”, limitation, the additional element is which is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components (See MPEP 2106.05(f)). Regarding the “wherein the loss function is calculated based on an image-text matching (ITM) loss function, a masked language modeling (MLM) loss function, and a masked activation modeling (MAM) loss function” which is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components (See MPEP 2106.05(f)). Accordingly, at Step 2B, the additional element individually or in combination does not amount to significantly more than the judicial exception. Regarding claim 2 (Cancelled) Regarding claim 3 Step 2A Prong 1 Claim 3 recites the following mathematical concepts, that in each case under the broadest reasonable interpretation, covers performance of mathematical relationships, mathematical formulas or equations, and mathematical calculations but for recitation of generic computer components (e.g., an image processor”, “activation embedding vector “, “convolutional neutral network”, “text processor”, “vector concatenator”, “encoder”) [see MPEP 2106.04(a)(2)(I)]. “wherein the image processor is configured to perform global average pooling on the plurality of activation maps to calculate the feature value for each of the plurality of activation maps” (e.g., computing the mean of all numbers in a matrix) Accordingly, at Step 2A, prong one, the claim recites an abstract idea. Step 2A Prong 2 In accordance with Step 2A, Prong 2, the claim does not include any additional elements and the judicial exception is not integrated into a practical application. Step 2B In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Regarding claim 4 Step 2A Prong 1 Claim 4 recites the following mathematical concepts, that in each case under the broadest reasonable interpretation, covers performance of mathematical relationships, mathematical formulas or equations, and mathematical calculations but for recitation of generic computer components (e.g., an image processor”, “activation embedding vector “, “convolutional neutral network”, “text processor”, “vector concatenator”, “encoder”) [see MPEP 2106.04(a)(2)(I)]. “each of the selected one or more activation maps has the feature value greater than each of non-selected activation maps” (e.g., finding the highest value between different maps) The examiner notes that this could also be a mental concept, as a human can look at a list of values of maps and determine which one has the higher value. Accordingly, at Step 2A, prong one, the claim recites an abstract idea. Step 2A Prong 2 The judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of “wherein the image processor is configured to select one or more activation maps among the plurality of activation maps” which is recited at a high level of generality and amounts to extra-solution activity of selecting a subset of numeric feature maps forma set of maps, i.e. post-solution activity of selecting a particular data source or type of data to be manipulated (see MPEP 2106.05(g)). Accordingly, at Step 2A, prong two, the additional elements individually or in combination do not integrate the judicial exception into a practical application. 2B In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. As discussed above, the additional elements of “wherein the image processor is configured to select one or more activation maps among the plurality of activation maps”, limitation, the additional element is recited at a high-level of generality and amounts to extra-solution activity of selecting a particular data type, i.e., post-solution activity of selecting a particular data source or type of data to be manipulated. The courts have found limitations directed to obtaining information electronically, recited at a high-level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory"). Accordingly, at Step 2B, the additional elements individually or in combination do not integrate the judicial exception into a practical application. Regarding claim 5 (Cancelled) Regarding claim 6 Step 2A Prong 1 Claim 6 recites the following mathematical concepts, that in each case under the broadest reasonable interpretation, covers performance of mathematical relationships, mathematical formulas or equations, and mathematical calculations but for recitation of generic computer components (e.g., an image processor”, “activation embedding vector “, “convolutional neutral network”, “text processor”, “vector concatenator”, “encoder”) [see MPEP 2106.04(a)(2)(I)]. “determine whether the text embedding vector and the activation embedding vector constituting the concatenated embedding vector match each other” (e.g., mathematical equation to compare two values to see if they match) The examiner notes that this could also be a mental concept, as a human can look at values and compare them to determine if they are equal to one another. “be trained based on the image-text matching (ITM) loss function calculated based on whether a result of the determination is correct” (e.g., mathematical optimization of adjusting parameters to minimize a numeric objective) Accordingly, at Step 2A, prong one, the claim recites an abstract idea. Step 2A Prong 2 In accordance with Step 2A, Prong 2, the claim does not include any additional elements and the judicial exception is not integrated into a practical application. Step 2B In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Regarding claim 7 Step 2A Prong 1 Claim 7 recites the following mathematical concepts, that in each case under the broadest reasonable interpretation, covers performance of mathematical relationships, mathematical formulas or equations, and mathematical calculations but for recitation of generic computer components (e.g., an image processor”, “activation embedding vector “, “convolutional neutral network”, “text processor”, “vector concatenator”, “encoder”) [see MPEP 2106.04(a)(2)(I)]. “be trained based on the masked language modeling (MLM) loss function calculated based on similarity between a masked element of the text mask concatenated embedding vector and an element, corresponding to the masked element, among elements of the text mask multimodal representation vector” (e.g., computing a numeric loss from similarity between predicted and true elements and adjust model parameters accordingly) Accordingly, at Step 2A, prong one, the claim recites an abstract idea. Step 2A Prong 2 The judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of “generate a text mask multimodal representation vector for a text mask concatenated embedding vector generated by masking at least one element, among elements of the text embedding vector constituting the concatenated embedding vector” which is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components (See MPEP 2106.05(f)). Accordingly, at Step 2A, prong two, the additional elements individually or in combination do not integrate the judicial exception into a practical application. 2B In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. As discussed above, the additional elements of “generate a text mask multimodal representation vector for a text mask concatenated embedding vector generated by masking at least one element, among elements of the text embedding vector constituting the concatenated embedding vector”, limitation, the additional element is which is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components (See MPEP 2106.05(f)). Accordingly, at Step 2B, the additional elements individually or in combination do not integrate the judicial exception into a practical application. Regarding claim 8 Step 2A Prong 1 Claim 8 recites the following mathematical concepts, that in each case under the broadest reasonable interpretation, covers performance of mathematical relationships, mathematical formulas or equations, and mathematical calculations but for recitation of generic computer components (e.g., an image processor”, “activation embedding vector “, “convolutional neutral network”, “text processor”, “vector concatenator”, “encoder”) [see MPEP 2106.04(a)(2)(I)]. “be trained based on the masked activation modeling (MAM) loss function calculated based on similarity between the masked element of the image mask concatenated embedding vector and an element, corresponding to the masked element, among elements of the image mask multimodal representation vector” (e.g., calculating a masked-activation loss from similarity scores and using it to train the model) Accordingly, at Step 2A, prong one, the claim recites an abstract idea. Step 2A Prong 2 The judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of “generate an image mask multimodal representation vector for an image mask concatenated embedding vector generated by masking an element among elements of an activation embedding vector constituting a concatenated embedding vector” which is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components (See MPEP 2106.05(f)). Accordingly, at Step 2A, prong two, the additional elements individually or in combination do not integrate the judicial exception into a practical application. 2B In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. As discussed above, the additional elements of “generate an image mask multimodal representation vector for an image mask concatenated embedding vector generated by masking an element among elements of an activation embedding vector constituting a concatenated embedding vector”, limitation, the additional element is which is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components (See MPEP 2106.05(f)). Accordingly, at Step 2B, the additional elements individually or in combination do not integrate the judicial exception into a practical application. Regarding claim 9 (Cancelled) Regarding claim 10 (Cancelled) Regrading claims 11-20 Claims 11-20 recites a method, which corresponds directly to the apparatus steps of claims 1-10, respectively, with the addition of software instructions which are insufficient to render the claims subject matter eligible for the same reasons described above. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 3, 4, 6-8, 11, 13, 14, and 16-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (US 20210216862 A1, referred to as Liu), in view of Kiros et al. (US 11003856 B2, referred to as Kiros), in view of Zhou et al. (“Learning Deep Features for Discriminative Localization”, referred to as Zhou), in view of Reed et al. (US 20200104680 A1, referred to as Reed), in view of Li et al. ("Visualbert: A simple and performant baseline for vision and language.", referred to as Li), in view of Bao et al. ("Beit: Bert pre-training of image transformers.", referred to as Bao). Regarding claim 1, Liu teaches, an apparatus for analyzing multimodal data, the apparatus comprising ([0085]: Describes an apparatus, comprising computer hardware components, which are configured to run the methods described within. [0091]: Describes a method for multimodal reasoning and matching of data.): an image processor configured to generate an activation embedding vector based on an index of an activation map obtained from image data through a convolutional neural network ([0124-0125]: Describes a context type module which receives multimedia context including image data, rotes image context to a CNN module. The CNN module performs convolution on the image to extract visual features from the hidden layers, splitting the image into multiple local regions (indexed positions) and produces corresponding region feature vectors for each region, CNN-derived activation features for later embedding.); Liu further teaches, a text processor configured to receive text data to generate a text embedding vector ([0126-01127]: Describes a text one-hot encoder and word embedding module that receives textual context. It converts each word of the text into one-hot vectors and then embed those vectors into context text embeddings.); a vector concatenator configured to concatenate the activation embedding vector and the text embedding vector to each other to generate a concatenated embedding vector ([0128-0131]: Describes a context concatenating module that receives context features vectors and context embeddings vectors then concatenates these vectors to produce a concatenated context vectors.; [0151-0155]: Describes “Procedure 360” which the multimedia context encoder concatenates that context feature vector and the context embedding vectors to obtain concatenated context vectors 430, providing the combined multimodal representation downstream to the attention module.); and an encoder configured to generate a multimodal representation vector in consideration of an influence between elements constituting the concatenated embedding vector based on self-attention ([0131-0136: Describes a dual attention fusion module 150 receives the concatenated task and context vectors. The computing attention weights via SoftMax (C_enc Q_encT) and SoftMax(Q_enc C_encT), representing influence between elements of the concatenated vectors based on attention mechanism, self-attention. The final encoding module 160 outputs final context and task encodings (C_com and Q_com), which function as multimodal representation vectors; [0155-0158]: Describes generating a fusion representation from the concatenated vectors and forwarding it to the final encoding module.), wherein at least one of the image processor, the text processor, the vector concatenator, and the encoder comprises a hardware(FIG. 1 and [0111-0116]: Describes computing device 110 comprises, “a processor 112, a memory 114, and a storage device 116, and optionally a database 190.” For executing multimedia data analysis and further explains that execution modules may be implemented as a circuit corresponding to at least one processing component implemented in the hardware.) Although Liu extracts feature vectors, it does not generate an activation embedding vector based on an index vector. Kiros (Col. 5, lines 12-49 and FIG. 2: Describes generating embedding vectors corresponding to images selected form a tanked list of image search results, wherein each embedding correspond to a respective index position. The system processes each indexed item through a CNN to obtain a numeric embedding and concatenates these embeddings according to their indices.) It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to combine Liu’s multimodal CNN fusion system with Kiros’s indexed-based activation embeddings. Doing so would have enabled the system to use stable positional indexing on CNN derived features. Allowing the system to be more consistent with alignment comparison and fusion of image features with text embeddings. Kiros further teaches, is configured to generate an index vector including indices of the selected one or more activation maps (Col. 4, lines 56-67. Cont. Col 5, line 1-58, and FIG. 2 : Describes selecting the top results form an ordered set and generating a vector structure that preserves the indices by concatenating embeddings according to their rank, corresponding to generating an index vector including indices selected items.) Kiros teaches, embed the index vector to generate an activation embedding vector(Col. 4, lines 56-67. Cont. Col 5, line 1-58, and FIG. 2 : Describes selecting a set of items having associated indices (ranked image results) and generating a corresponding embedding for each index, then constructing a final embedding vector by concatenating the embeddings according to their index ordering, corresponding to embedding an index vector to generate an activation embedding vector.) generating an index vector including indices of one or more activation maps selected among the plurality of activation maps in an order of a descending feature value (Describes that image search results ordered form most responsive to least responsive, selecting top results according to the order, processing each result using a CNN to generate image numeric embeddings, and concatenating the image embeddings according to the order of the corresponding image search results.); Although Liu in view of Kiros teaches, an apparatus for analyzing multimodal data, the apparatus comprising: an image processor configured to generate an activation embedding vector based on an index of an activation map obtained from image data through a convolutional neural network; a text processor configured to receive text data to generate a text embedding vector; a vector concatenator configured to concatenate the activation embedding vector and the text embedding vector to each other to generate a concatenated embedding vector; and an encoder configured to generate a multimodal representation vector in consideration of an influence between elements constituting the concatenated embedding vector based on self- attention, wherein at least one of the image processor, the text processor, the vector concatenator, and the encoder comprises a hardware . They do not teach wherein the image processor is configured to generate an activation map set comprising a plurality of activation maps for the image data, using a synthetic neutral network; calculating a feature value for each of the plurality of activation maps, and generating an index vector including indices of one or more activation maps selected among the plurality of activation maps in an order of a descending feature value. Zhou teaches, wherein the image processor is configured to generate the activation embedding vector by performing: generating an activation map set comprising a plurality of activation maps for the image data, using a synthetic neutral network (Page 2-3, Section 2: Describes that a convolutional neural network generates a plurality of activation maps in the last convolutional layer, where each unit k produces its own activation map, and these collectively form the set of activation maps used to compute class activation maps. This corresponds to generating an activation map set comprising a plurality of activation maps for image data using a neural network.). It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to combine Liu’s in view of Kiros’s apparatus with Zhou’s activation mappings. Doing so would have enabled the system to improve downstream fusion and attention mechanisms by providing more structured, informative CNN features for alignment with the text embeddings, increasing accuracy and efficiency of the multimodal analysis. Zhou further teaches, calculating a feature value for each of the plurality of activation maps (Pages 2-3, Section 2: Describes performing global average pooling on each activation map in the last convolutional layer, where the global average pooled value is computed for each map, thereby calculating a feature value for each of a plurality of activation maps.), and generating an index vector including indices of one or more activation maps selected among the plurality of activation maps in an order of a descending feature value (Zhou teaches calculating feature/importance values for activation maps or feature map units, including using global average pooling to obtain Fk values and using class specific weights indicating the importance of Fk for a class. Zhou further teaches ranking units based on such importance to identify the most discriminative units for a class.; As described above Kirs teaches selecting ordered CNN derived image information and preserving the order when generating an embedding, including concatenating image embeddings according to the ordering of selected image search results. It would have been obvious to generate an index vector including indices of selected activation maps ordered by descending feature/importance value so that the generated embedding vector preserves the order of the most informative activation maps. Doing so would enable the system to improve visual representation quality for downstream machine learning tasks). Although Liu in view of Kiros, in view of Zhou teaches, an apparatus for analyzing multimodal data, the apparatus comprising: an image processor configured to generate an activation embedding vector based on an index of an activation map obtained from image data through a convolutional neural network; a text processor configured to receive text data to generate a text embedding vector; a vector concatenator configured to concatenate the activation embedding vector and the text embedding vector to each other to generate a concatenated embedding vector; and an encoder configured to generate a multimodal representation vector in consideration of an influence between elements constituting the concatenated embedding vector based on self- attention, wherein at least one of the image processor, the text processor, the vector concatenator, and the encoder comprises a hardware wherein the image processor is configured to generate an activation map set comprising a plurality of activation maps for the image data, using a synthetic neutral network; calculating a feature value for each of the plurality of activation maps, and generating an index vector including indices of one or more activation maps selected among the plurality of activation maps in an order of a descending feature value. They do not teach, wherein the image processor, the text processor, and the encoder are configured to be trained based on the same loss function and wherein the loss function is calculated based on an image-text matching (ITM) loss function. Reed teaches, wherein the image processor, the text processor, and the encoder are configured to be trained based on the same loss function ([0065]: Describes that the encoder neural network the context network, and the associated transformation modules are jointly trained using a single objective (loss) function whose gradients update all component parameters, corresponding to training the image processor, text processor, and encoder based on the same loss function.) It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to combine Liu in view of Kiros, in view of Zhou’s apparatus with Reed’s matching determination and loss function. Doing so would have enabled the system to more accurately align and filter multimodal embeddings by training the encoder to score matching pairs higher than mismatched ones, which would improve accuracy and efficiency. Reed further teaches wherein the loss function is calculated based on an image-text matching (ITM) loss function([0013-0015]: Describes that a gradient of a loss function is determined, where the loss characterizes the accuracy of estimates of the latent representations, and in some implementations the loss function is a noise contrastive estimation loss that uses similarities between an estimate and positive/negative latent representations.; [0065-0067]: Describes that the training engine computes gradients of an objective loss function that characterizes the accuracy of predications and specifics the NCE objective in terms of similarities for the positive vs negative pairs.; [0116-0117]: Describes determining gradient of a loss function (optionally NCE) with respect to encoder parameters and updating the encoder based on that loss.) Although Liu in view of Kiros, in view of Zhou, in view of Reed teaches, an apparatus for analyzing multimodal data, the apparatus comprising: an image processor configured to generate an activation embedding vector based on an index of an activation map obtained from image data through a convolutional neural network; a text processor configured to receive text data to generate a text embedding vector; a vector concatenator configured to concatenate the activation embedding vector and the text embedding vector to each other to generate a concatenated embedding vector; and an encoder configured to generate a multimodal representation vector in consideration of an influence between elements constituting the concatenated embedding vector based on self- attention, wherein at least one of the image processor, the text processor, the vector concatenator, and the encoder comprises a hardware wherein the image processor is configured to generate an activation map set comprising a plurality of activation maps for the image data, using a synthetic neutral network; calculating a feature value for each of the plurality of activation maps, and generating an index vector including indices of one or more activation maps selected among the plurality of activation maps in an order of a descending feature value, wherein the image processor, the text processor, and the encoder are configured to be trained based on the same loss function and wherein the loss function is calculated based on an image-text matching (ITM) loss function. They do not teach, a masked language modeling (MLM) loss function. Li teaches, a masked language modeling (MLM) loss function(Page 3-4, Section 3.2-3.3: Describes training the multimodal transformer using masked language modeling, in which some text tokens are masked and the model is trained to predict the correct token. The masked text tokens must be predicted from the multimodal encoder output, and the model is trained using a masked language modeling objective that compares the predicted representation at the masked position with the correct token. i.e., the encoder is trained based on a masked language modeling loss function (MLM) and the loss is calculated based on similarity between the masked element and the corresponding predicted element in the multimodal representation. ). It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to combine Liu in view of Kiros, in view of Zhou, in view of Reed’s apparatus with Li’s masked multimodal training. Doing so would have enabled the system to create more accurate and efficient multimodal representation enabling the encoder to learn text-image correspondences directly form the concatenated embeddings. Although Liu in view of Kiros, in view of Zhou, in view of Reed, in view of Li teaches, an apparatus for analyzing multimodal data, the apparatus comprising: an image processor configured to generate an activation embedding vector based on an index of an activation map obtained from image data through a convolutional neural network; a text processor configured to receive text data to generate a text embedding vector; a vector concatenator configured to concatenate the activation embedding vector and the text embedding vector to each other to generate a concatenated embedding vector; and an encoder configured to generate a multimodal representation vector in consideration of an influence between elements constituting the concatenated embedding vector based on self- attention, wherein at least one of the image processor, the text processor, the vector concatenator, and the encoder comprises a hardware wherein the image processor is configured to generate an activation map set comprising a plurality of activation maps for the image data, using a synthetic neutral network; calculating a feature value for each of the plurality of activation maps, and generating an index vector including indices of one or more activation maps selected among the plurality of activation maps in an order of a descending feature value, wherein the image processor, the text processor, and the encoder are configured to be trained based on the same loss function and wherein the loss function is calculated based on an image-text matching (ITM) loss function and a masked language modeling (MLM) loss function. They do not teach, a masked activation modeling (MAM) loss function. Bao teaches, a masked activation modeling (MAM) loss function(Page 3-4, Section 2.3: Describes computing a masked image modeling loss by comparing the predicted representation at each masked patch position with the corresponding true visual token, corresponding to calculating a similarity based loss function between the masked element and the predicted element, teaching a masked activation modeling loss function.). It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to combine Liu in view of Kiros, in view of Zhou, in view of Reed, in view of Li’s apparatus with Bao’s masked image modeling. Doing so would have enabled the system to learn richer visual representations by training the encoder to reconstruct masked portions for the activation embedding vector, enhancing downstream fusion. Regarding claim 2 (Cancelled) Regarding claim 3, Liu in view of Kiros, in view of Zhou, in view of Reed, in view of Li teaches, the apparatus of claim 1, wherein the image processor is configured to perform global average pooling on the plurality of activation maps to calculate a feature value for each of the plurality of activation maps (Zhou Pages 2-3, Section 2: Describes performing global average pooling on each activation map in the last convolutional layer, where the global average pooled value is computed for each map, thereby calculating a feature value for each of a plurality of activation maps.). Regarding claim 4, Liu in view of Kiros, in view of Zhou, in view of Reed, in view of Li teaches, the apparatus of claim 1 wherein the image processor is configured to select one or more activation maps among the plurality of activation maps (Zhou Pages 2-3, Section 2: Describes determining the importance of each activation map via the value Fk and weight wck which identifies (selects) the activation maps that contribute most strongly to the class score.), and each of the selected one or more activation maps has the feature value greater than each of non-selected activation maps (Pages 2-4, Section 2 and Section 3: Describes identifying the activation m aps with the largest pooled values Fk and corresponding importance weights, selecting the top contributing activation maps for generating the class activation map, corresponding to selecting activation maps whose feature values exceed those of non-selected maps). Regarding claim 6, Liu in view of Kiros, in view of Zhou, in view of Reed, in view of Li teaches, the apparatus of claim 1, wherein the encoder is configured to: determine whether the text embedding vector and the activation embedding vector constituting the concatenated embedding vector match each other (Reed [0013-0015]: Describes that the system uses a noise contrastive estimation (NCE) where the system compares the estimate to a positive latent representation and to multiple negative latent representation, using a similarity measure to distinguish which latent representation the estimate should match.;[0063-0065]: Describes generating predictions of latent representations for subsequent observations from a context latent representation and then measuring how accurate those predictions are with a loss. ;[0067]: Describes the noise contrastive estimation objective, where for a prediction p of a latent representation z, the system gets a set of negative latent representations, computes similarity for the positive pair and for each negative and defines the loss in terms of these similarities so the correct (matching) pair is distinguished form the non-matching ones.); and be trained based on the image-text matching (ITM) loss function calculated based on whether a result of the determination is correct ([0013-0015]: Describes that a gradient of a loss function is determined, where the loss characterizes the accuracy of estimates of the latent representations, and in some implementations the loss function is a noise contrastive estimation loss that uses similarities between an estimate and positive/negative latent representations.; [0065-0067]: Describes that the training engine computes gradients of an objective loss function that characterizes the accuracy of predications and specifics the NCE objective in terms of similarities for the positive vs negative pairs.; [0116-0117]: Describes determining gradient of a loss function (optionally NCE) with respect to encoder parameters and updating the encoder based on that loss.). Regarding claim 7, Liu in view of Kiros teaches, the apparatus of claim 1, wherein the encoder is configured to: generate a text mask multimodal representation vector for a text mask concatenated embedding vector generated by masking at least one element, among elements of the text embedding vector constituting the concatenated embedding vector (Li Page 3-4, Section 3.2-3.3: Describes forming a joint input sequence by concatenating text embeddings with visual embeddings for image regions, and then randomly masking some elements of the text input (while leaving the image region vectors unmasked), corresponding to generating a text-masked concatenated embedding vector by masking at least one element among the text embedding elements of the concatenated embedding vector. It further describes a transformer encoder that receives the concatenated sequence of text and image embeddings, with some text tokens masked, and produces contextualized joint representations over all positions (including the mased text position) corresponding to generating a multimodal representation vector for the text-mask concatenated embedding vector.); and be trained based on the masked language modeling (MLM) loss function calculated based on similarity between a masked element of the text mask concatenated embedding vector and an element, corresponding to the masked element, among elements of the text mask multimodal representation vector (Li Page 3-4, Section 3.2-3.3: Describes training the multimodal transformer using masked language modeling, in which some text tokens are masked and the model is trained to predict the correct token. The masked text tokens must be predicted from the multimodal encoder output, and the model is trained using a masked language modeling objective that compares the predicted representation at the masked position with the correct token. i.e., the encoder is trained based on a masked language modeling loss function (MLM) and the loss is calculated based on similarity between the masked element and the corresponding predicted element in the multimodal representation. ). It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to combine Liu’s in view of Kiros’s apparatus with Li’s masked multimodal training. Doing so would have enabled the system to create more accurate and efficient multimodal representation enabling the encoder to learn text-image correspondences directly form the concatenated embeddings. Regarding claim 8, Liu in view of Kiros, in view of Zhou, in view of Reed, in view of Li teaches, the apparatus of claim 1, wherein the encoder is configured to: generate an image mask multimodal representation vector for an image mask concatenated embedding vector generated by masking an element among elements of an activation embedding vector constituting a concatenated embedding vector (Bao Page 2 Section 1: “We randomly mask some proportion of image patches (gray patches in the figure) and replace them with a special mask embedding [M]. Then the patches are fed to a backbone vision Transformer.” This describes randomly ,asking image patch embeddings and replacing the masked embeddings with a learned mask token [M], generating a masked image embedding vector that is then concatenated and input to the transformer encoder.; Page 3-4, Section 2.3: Describes feeding the masked image embedding sequence into a multi-layer transformer encoder, which outputs contextualized hidden vectors representing the masked-image input sequence, corresponding to generating an image mask multimodal representation vector.); and be trained based on a masked activation modeling (MAM) loss function calculated based on similarity between the masked element of the image mask concatenated embedding vector and an element, corresponding to the masked element, among elements of the image mask multimodal representation vector (Bao Page 3-4, Section 2.3: Describes computing a masked image modeling loss by comparing the predicted representation at each masked patch position with the corresponding true visual token, corresponding to calculating a similarity based loss function between the masked element and the predicted element, teaching a masked activation modeling loss function.). Regarding claim 11, which recites substantially the same limitations as claim 1. Claim 1 further recites a method for analyzing multimodal data, the method performed by a computing device comprising a processor and a computer-readable storage medium storing a program comprising a computer-executable command executed by the processor to perform operations (Liu [0004-0030]: Describes a method executed on computer hardware to execute multimodal data processing and analysis.) to perform the apparatus steps of claim 1, respectively, and is therefore rejected on the same premise. Regarding claims 12-15, which recites substantially the same limitations as claim 2-5. Claims 12-15 further recites a method (as described in claim 11) to perform the apparatus steps of claims 2-5, respectively, and is therefore rejected on the same premise. Regarding claim 16, which recites substantially the same limitations as claim 6. Claim 16 further recites a method (as described in claim 11) to perform the apparatus steps of claim 6, respectively, and is therefore rejected on the same premise. Regarding claim 17, which recites substantially the same limitations as claim 7. Claim 17 further recites a method (as described in claim 11) to perform the apparatus steps of claim 7, respectively, and is therefore rejected on the same premise. Regarding claim 18, which recites substantially the same limitations as claim 8. Claim 18 further recites a method (as described in claim 11) to perform the apparatus steps of claim 8, respectively, and is therefore rejected on the same premise. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DONALD T RODEN whose telephone number is (571)272-6441. The examiner can normally be reached Mon-Thur 8:00-5:00 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Omar Fernandez Rivas can be reached at (571) 272-2589. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /D.T.R./Examiner, Art Unit 2128 /OMAR F FERNANDEZ RIVAS/Supervisory Patent Examiner, Art Unit 2128
Read full office action

Prosecution Timeline

Oct 25, 2022
Application Filed
Dec 19, 2025
Non-Final Rejection mailed — §101, §103
Mar 09, 2026
Applicant Interview (Telephonic)
Mar 09, 2026
Examiner Interview Summary
Mar 13, 2026
Response Filed
May 15, 2026
Final Rejection mailed — §101, §103
Jul 15, 2026
Response after Non-Final Action

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
20%
Grant Probability
99%
With Interview (+100.0%)
3y 9m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 5 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month