Prosecution Insights
Last updated: October 04, 2026
Application No. 18/179,398

METHOD FOR OPTIMIZING WORKFLOW-BASED NEURAL NETWORK INCLUDING ATTENTION LAYER

Final Rejection §103§112
Filed
Mar 07, 2023
Examiner
AGRAWAL, SHISHIR
Art Unit
2123
Tech Center
2100 — Computer Architecture & Software
Assignee
Hong Kong Applied Science and Technology Research Institute Company Limited
OA Round
2 (Final)
8%
Grant Probability
At Risk
3-4
OA Rounds
5m
Est. Remaining
24%
With Interview

Examiner Intelligence

Grants only 8% of cases
8%
Career Allowance Rate
2 granted / 24 resolved
-46.7% vs TC avg
Strong +15% interview lift
Without
With
+15.4%
Interview Lift
resolved cases with interview
Typical timeline
4y 0m
Avg Prosecution
13 currently pending
Career history
49
Total Applications
across all art units

Statute-Specific Performance

§101
23.9%
-16.1% vs TC avg
§103
40.0%
+0.0% vs TC avg
§102
6.8%
-33.2% vs TC avg
§112
29.4%
-10.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 24 resolved cases

Office Action

§103 §112
DETAILED ACTION Status of Claims This Office action is responsive to communications filed on 2026-08-06. Claim(s) 3 and 8-16 was/were cancelled and claim(s) 17-19 was/were added. Claim(s) 1-2, 4-7 and 17-19 is/are pending and is/are examined herein. Claim(s) 1-2, 4-7 and 17-19 is/are objected to. Claim(s) 1-2, 4-7 and 17-19 invoke(s) interpretation under 35 USC 112(f). Claim(s) 1-2, 4-7 and 17-19 is/are rejected under 35 USC 112(b). Claim(s) 1-2, 4-7 and 17-19 is/are rejected under 35 USC 112(a). Claim(s) 1-2, 4-7 and 17-19 is/are rejected under 35 USC 103. Notice of Pre-AIA or AIA Status The present application, filed on or after 2013-03-16, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Regarding objections for informalities and rejections under 35 USC 112(b), the applicant’s remarks have been fully considered. The amendments resolve some, but not all, of the issues raised in the previous Office action. They also raise a number of new issues. Unresolved issues in the pending claims are given below. Regarding rejections under 35 USC 103, the applicant’s remarks have been fully considered. Regarding claim 1, the applicant makes numerous assertions regarding Galassi in view of Wu not disclosing certain purportedly claimed features [remarks, pages 8-9]. However, the relevance of these remarks to the pending claims is tenuous since the features recited in the applicant’s remarks do not in fact appear in the pending claims. For example, the applicant asserts this with respect to a step of “extracting attention maps with their elements and features as interpretable patterns” but no such step is recited in either the pending claims or the originally filed specification. Similarly, the applicant asserts this with respect to a “deviation function (part of internal attention correction mechanism)” [remarks, page 9] but not internal attention correction mechanism is recited in either the pending claims or the originally filed specification. Regarding claim 1, the applicant makes remarks regarding the conditional limitations [remarks, pages 8-9]. The remarks are unpersuasive. The examiner notes that the limitations remain conditional. The claim remains a genus claim (now encompassing at least 3 species; cf. examiner’s remarks), and that the prior art made of record discloses at least one of these species in the largely the same way described in the rejection of claim 3 in the previous Office action. The applicant is invited to consult the complete prior art mapping of claim 1 as given below. Regarding claim 2, the applicant asserts that Galassi in view of Wu does not disclose a step of visualizing the attention layer [remarks, page 9]. However, these remarks are unpersuasive since, as indicated in the previous Office action, Galassi and Wu include numerous visualizations of attention layers (e.g., [Galassi, figures 3-7] and [Wu, figure 2]). The applicant is also invited to consult [Xu, figures 3 and 5-15] as cited in the prior art mapping below. Regarding claim 19, the applicant’s remarks have been fully considered but they are moot in view of the updated rejection as given below. The complete prior art mapping, updated in view of the applicant’s amendments, is given below. Examiner’s Remarks MPEP 2111.04(II) indicates that the “broadest reasonable interpretation of a method (or process) claim having contingent limitations requires only that those steps that must be performed and does not include steps that are not required to be performed because the condition(s) are not met”. Claims 1-2 recites such contingent limitations: [Claim 1] if the attention mask updating function can be created to fulfill the attention mask pattern proposal, creating an attention mask updating function; else: determining whether a new attention pattern proposal can be obtained; if a new attention pattern proposal can be obtained, repeating the method steps starting from step b; else, adopting a reinforcement learning (RL) model as the attention mask updating function; [Claim 2] creating the proposed attention mask pattern based on one or more desired attention placements and corresponding desired attention weights on the input elements if the attentions and the corresponding attention weights are not being placed on the identified elements of the input elements that correspond with the original prediction results. The claims do not positively recite that the conditions on which these actions are contingent are satisfied, so, in keeping with the interpretation of contingent limitations as described in MPEP, the broadest reasonable interpretation of these claims does not include the recited contingent actions. For the purpose of compact prosecution, claim 1 is interpreted herein as an attempt to formulate a genus claim having at least three species: (1) a species in which the attention mask updating function can be created, (2) a species in which it cannot, but a new attention pattern proposal can be obtained, and (3) a species in which the attention mask cannot be created and a new attention pattern proposal cannot be obtained. Claim Interpretation – 35 USC 112(f) The following is a quotation of 35 USC 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 USC 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 USC 112(f) or pre-AIA 35 USC 112, sixth paragraph, is invoked. As explained in MPEP 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 USC 112(f) or pre-AIA 35 USC 112, sixth paragraph: the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 USC 112(f) or pre-AIA 35 USC 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 USC 112(f) or pre-AIA 35 USC 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 USC 112(f) or pre-AIA 35 USC 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 USC 112(f) or pre-AIA 35 USC 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 USC 112(f) or pre-AIA 35 USC 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 USC 112(f) or pre-AIA 35 USC 112, sixth paragraph, except as otherwise indicated in an Office action. Claim 1 recites numerous “steps” (e.g., a “step a” involving “training the workflow-based neural network to predict one or more results from one or more input elements under a prediction model…”). Moreover, the claims do not recite structure that is sufficient for performing the recited steps. The specification indicates that the methods disclosed therein “may be implement using computing devices, computer processors, or electronic circuits” [specification, 0061] and “may be executed in one or more general purpose or computing devices” [specification, 0062], so the claims are interpreted accordingly. Dependent claims 2, 4-7, and 17-19 inherit these limitations and their interpretations. If applicant does not intend to have this/these limitation(s) interpreted under 35 USC 112(f) or pre-AIA 35 USC 112, sixth paragraph, applicant may: amend the claim limitation(s) to avoid it/them being interpreted under 35 USC 112(f) or pre-AIA 35 USC 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 USC 112(f) or pre-AIA 35 USC 112, sixth paragraph. Claim Rejections - 35 USC 112(b) The following is a quotation of 35 USC 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 USC 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim(s) 1-2, 4-7 and 17-19 is/are rejected under 35 USC 112(b) or 35 USC 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 USC 112, the applicant), regards as the invention. Claim 1 is indefinite for at least the following reasons: It recites the following ordered steps but this phrase lacks antecedent basis. Alternative language is advised. It includes multiple recitations of the attention pattern proposal but this phrase lacks antecedent basis. It should be “the attention mask pattern proposal” for consistent nomenclature and proper antecedent basis. It recites determining whether an attention mask updating function can be created to fulfill the attention pattern proposal and if the attention mask updating function can be created to fulfill the attention mask pattern proposal [emphasis added]. However, the underlined clauses are indefinite because it is not clear what it means to “fulfill” a proposal. It is possible to fulfill, for example, a requirement, but a proposal is not a requirement. The applicant is invited to consult a rejection of a possibly related indefinite limitation given below. For the purpose of compact prosecution, the limitation is interpreted broadly as requiring that the attention mask updating function be related in some way to the proposed attention mask pattern. It recites creating an attention mask updating function [emphasis added] but this should be “creating the attention mask updating function” for proper antecedent basis (since the claim previously introduces “an attention mask updating function”). It recites if the attention mask updating function can be created to fulfill the attention mask pattern proposal, creating an attention mask updating function [emphasis added]. However, this limitation as recited conflicts with the specification, which recites instead “if the attention mask updating function D206 can be created to fulfill the proposed attention mask pattern D204, creating the attention mask updating function D206” [specification, 0053; emphasis added]. The examiner notes that the pending claims recite both an “attention mask pattern proposal” as well as a “proposed attention mask pattern” in apparently distinct contexts, and the conflict between the verbiage of the claims and the specification renders unclear which of these the attention mask updated function is intended to “fulfill”. It also renders unclear whether or not the “attention mask pattern proposal” refers to the same entity as the “proposed attention mask pattern” of the claim. MPEP 2173.03 indicates that a claim is “indefinite when a conflict or inconsistency between the claimed subject matter and the specification disclosure renders the scope of the claim uncertain as inconsistency with the specification disclosure or prior art teachings may make an otherwise definite claim take on an unreasonable degree of uncertainty” so this inconsistency with the specification renders the claim indefinite. It recites if a new attention pattern proposal can be obtained, repeating the method steps starting from step b [emphasis added]. However, this is indefinite for at least the following reasons: First, the phrase “the method steps starting from step b” lacks antecedent basis. Alternative language is advised. Second, the limitation as recited conflicts with the specification, which instead recites “determining whether a new attention pattern proposal can be obtained, and if so, reiterating the executions of the process steps beginning from the process step S20[6]4” [specification, 0055] where step S2064 refers to the step of “determining whether an attention mask updating function can be created to fulfill the attention pattern proposal”, which is “step d” in the verbiage of the claims. This conflict between the claims and the specification renders unclear which steps are supposed to be repeated. In view of the guidance of MPEP 2173.03 as quoted above, the claim is consequently indefinite. It recites wherein the fulfillment of the attention pattern proposal being the original attention mask pattern is manipulatable by the attention mask updating function to approach the proposed attention mask pattern [emphasis added]. However, this limitation is indefinite for at least the following reasons: First, the phrase “the fulfillment” lacks antecedent basis. Alternative language is advised. Second, the grammar of the entire wherein clause is ambiguous. It is not clear which entity recited by the limitation is “being the original attention mask pattern” (whether it the “attention pattern proposal” or the “fulfillment of the attention pattern proposal”). The grammar of the clause also suggests that the entity that is “manipulatable by the attention mask updating function to approach the proposed attention mask pattern” is the “fulfillment of the attention pattern proposal” but it is not clear what it means for a “fulfillment” to be “manipulatable” in the manner recited. As best understood by the examiner in view of the applicant’s remarks [remarks, page 7], this limitation appears to be a grammatically ambiguous attempt to define what it means to fulfill the attention pattern proposal, and as noted above, for the purpose of compact prosecution, the limitation is interpreted broadly as requiring that the attention mask updating function be related in some way to the proposed attention mask pattern. Dependent claims 2, 4-7, and 17-19 inherit the rejection. Claim 2 is indefinite for at least the following reasons: It recites one or more attentions and corresponding attention weights [emphasis added]. However, the underlined phrase is duplicate nomenclature, since the parent claim already recites “one or more attention placements and corresponding attention weights”. This duplication of nomenclature makes unclear whether the “corresponding attention weights” are bound in scope by those recited in the parent claim, and it also renders ambiguous the intended antecedent of all subsequent recitations of “the corresponding attention weights”. It also raises a question regarding whether or not the “one or more attentions” of this dependent claim might be bound in scope by the “one or more attention placements” of the parent claim. For the purpose of compact prosecution, the claim is interpreted broadly so that the “corresponding attention weights” of the dependent claim are not necessarily bound in scope by those recited by the parent claim. It includes multiple recitations of the identified elements but these recitations lack antecedent basis. They should be “the one or more identified elements” for consistent nomenclature and proper antecedent basis (since the claim previously recites “identifying one or more elements”). Dependent claims 18-19 inherit the rejections. Claim 19 is indefinite for at least the following reasons: It recites domain knowledge is injected into the optimization of the workflow-based neural network [emphasis added] but the underlined phrase lacks antecedent basis. The examiner suggests “domain knowledge is injected into the optimizing of the workflow-based neural network” for proper antecedent basis and nomenclature that is consistent with that is used in the parent claim (cf. “optimizing a workflow-based neural network” [claim 1]). Claim(s) 1-2, 4-7 and 17-19 recite(s) elements which invoke interpretation under 35 USC 112(f). As indicated above, these limitations are interpreted according to the specification as being implemented on a general purpose computer [specification, 0061-0062]. However, MPEP 2181(II)(B) indicates that “the structure be more than simply a general purpose computer or microprocessor and that the specification must disclose an algorithm for performing the claimed function”, that “[a]n algorithm is defined, for example, as ‘a finite sequence of steps for solving a logical or mathematical problem or performing a task’” and that “a rejection under 35 USC 112(b) or pre-AIA 35 USC 112, second paragraph is appropriate if the specification discloses no corresponding algorithm associated with a computer or microprocessor”. In the present instance, no explicit sequence of steps is provided in the specification for providing each of the “steps” of the claim (for example, the specification provides no algorithm or sequence of steps for the step of “training the workflow-based neural network” of the claim). Consequently, the claims are rejected under 35 USC 112(b) for failing to disclose sufficient structure. Claim Rejections - 35 USC 112(a) The following is a quotation of the first paragraph of 35 USC 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 USC 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claim(s) 1-2, 4-7 and 17-19 is/are rejected under 35 USC 112(a) or 35 USC 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 USC 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Claim(s) 1-2, 4-7 and 17-19 recite(s) limitations invoked 35 USC 112(f) and are rejected under 35 USC 112(b) for failing to disclose sufficient structure. MPEP 2181(II)(B) indicates that “[w]hen a claim containing a computer-implemented 35 USC 112(f) claim limitation is found to be indefinite under 35 USC 112(b) for failure to disclose sufficient corresponding structure (e.g., the computer and the algorithm) in the specification that performs the entire claimed function, it will also lack written description under 35 USC 112(a)”. Consequently, the claim(s) is/are rejected under 35 USC 112(a) for lack of written description. Claim 2 recites a step of determining, based on the visualization, whether the attentions and the corresponding attention weights are being placed on the identified elements of the input elements and that correspond with the original prediction results [emphasis added]. As best understood by the examiner in view of the applicant’s remarks [remarks, pages 6-7] and by comparing against the originally filed specification [specification, 0041], this limitation appears to be an attempt to claim what it means for the original attention mask pattern to make “intuitive sense” and/or for the attentions/weights to be “placed correctly”. This interpretation is further reinforced by a limitation in dependent claim 19, which indicates that this step is to be performed “by a human observer” and the only place in which a human is mentioned in the specification is, again, [specification, 0041]. However, this paragraph of the specification merely recites “whether attentions are being placed correctly and with appropriate attention weights in relation to the identified elements of the input elements D201 and the original prediction results D209” [specification, 0041] and does not specifically describe the attentions/weights being “placed on the identified elements” as recited by the amended claim. Consequently, the claim as amended is rejected for inadequate written description. Dependent claims 18-19 inherit the rejection. Claim Rejections - 35 USC 103 The following is a quotation of 35 USC 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 USC 102(b)(2)(C) for any potential 35 USC 102(a)(2) prior art against the later invention. Claim(s) 1, 4-7, and 17 is/are rejected under 35 USC 103 as being unpatentable over Andrea GALASSI et al. (Attention in Natural Language Processing, published 2020-05-28; hereafter, “Galassi”) in view of Hongqiu WU et al. (Not All Attention Is All You Need, published 2021-06-01; hereafter, “Wu”). Claim 1 A method for optimizing a workflow-based neural network including an attention layer, the method comprising the following ordered steps: ([Galassi, abstract]: Galassi discloses the use of attention mechanisms in neural architectures [Galassi, abstract]. Such a neural network maps to the “workflow-based neural network” of the claim, and the attention mechanism to the “attention layer” of the claim. The examiner notes that each step as mapped below is “ordered” in the sense that it has a regular arrangement/order, i.e., each step has elements that are arranged/ordered according to certain rules.) step a: training the workflow-based neural network to predict one or more results from one or more input elements under a prediction model ([Galassi, sections II.A and IV.A]: Galassi discloses a use of the neural network having an attention mechanism to produce an output sequence based on an input sequence x [Galassi, section II.A; see also, figure 4]. Galassi also discloses that attention mechanisms are trained alongside the rest of the neural architecture [Galassi, section IV.A first paragraph]. The neural network also maps to the “prediction model” of the claim; its training maps to the “training” step of the claim, the input sequence to the “input elements” of the claim, and the output sequence to the “one or more results” of the claim.) with the attention layer assigning one or more attention placements and corresponding attention weights, ([Galassi, figure 3 and section II.B]: The attention mechanism disclosed in Galassi produces a vector of “attention weights” [Galassi, figure 3]. The vector of attention weights maps to the “one or more attention placements and corresponding attention weights” of the claim (the elements of the vector being the “attention weights” and the indices at which they are located being the “attention placements”). The examiner notes that Galassi indicates that attention weights are computed as a = g(e) where e = f(q, K) [Galassi, section II.B equations (6-7)], where g is a distribution function, f is a compatibility function, q is a query vector, and K are the keys [Galassi, figure 3; see also, table II].) based on an original attention function of the attention layer, ([Galassi, figure 4 and section II.B]: Galassi indicates that attention weights are combined with V to obtain the context vector [Galassi, figure 4 and section II.B equations (8-9)]. The overall attention model, which takes keys, queries, and values as input and outputs a context vector [Galassi, figure 4], maps to the “original attention function” of the claim.) to the input elements ([Galassi, figure 4 and section II.B]: In the attention mechanism [Galassi, figure 4], the “one or more attention placements and weights” as mapped above are assigned to the “input elements” as mapped above.) until the prediction model converges; ([Galassi, section IV.A]: As noted above, Galassi discloses training the neural network [Galassi, section IV.A first paragraph]. The point when training concludes is when the “prediction model converges” as recited by the claim. The examiner notes that the reference Wu used in the combination as proposed below discusses convergence more explicitly; see, for instance, [Wu, figure 4 and section 6.2].) Galassi might not distinctly disclose: step b: generating an attention mask pattern proposal to obtain an original attention mask pattern and a proposed attention mask pattern; step c: designing a deviation function representing a quantifiable deviation between the original attention mask pattern and the proposed attention mask pattern; step d: determining whether an attention mask updating function can be created to fulfill the attention pattern proposal; if the attention mask updating function can be created to fulfill the attention mask pattern proposal, creating an attention mask updating function; else: determining whether a new attention pattern proposal can be obtained; if a new attention pattern proposal can be obtained, repeating the method steps starting from step b; else, adopting a reinforcement learning (RL) model as the attention mask updating function; wherein the fulfillment of the attention pattern proposal being the original attention mask pattern is manipulatable by the attention mask updating function to approach the proposed attention mask pattern; step e: combining the attention mask updating function with the original attention function to form an updated attention function of the attention layer. Wu is in the field of machine learning. Like Galassi, it discusses attention mechanisms [Wu, abstract] and more specifically mentions scaled dot-product attention [Galassi, table IV and section IV.C; Wu, section 3.1 equation (1)] (cf. “original attention function” of the claim). Moreover, Galassi in view of Wu discloses: step b: generating an attention mask pattern proposal to obtain an original attention mask pattern and a proposed attention mask pattern; ([Wu, sections 1 and 3]: Wu discloses a method which “dynamically generate[s] dropout patterns for each attention layer” [Wu, section 1 paragraph beginning “Dropout”]. Each dropout pattern is determined by a mask matrix M [Wu, section 3 first paragraph; see also, section 4 paragraph beginning “Generator”], and the mechanism performs “element-wise multiplication” of Softmax(QK^T/sqrt{d_k}) and M [Wu, section 3.2 paragraph beginning “Weights Dropout” and equation (2)]. The mask matrix M maps to the “attention mask pattern proposal” of the claim, the matrix Softmax(QK^T/sqrt{d_k}) maps to the “original attention mask pattern” of the claim, and the result of element-wise multiplication to the “proposed attention mask pattern” of the claim.) step c: designing a deviation function representing a quantifiable deviation between the original attention mask pattern and the proposed attention mask pattern; ([Wu, section 4.1]: As noted above, Wu discloses computing rewards for the G-Net based on a comparison between the A-Net (which uses the mask generated by the G-Net) and the D-Net (which does not) [Wu, section 4.1]. In other words, the reward maps to the “quantifiable deviation between the original attention mask pattern and the proposed attention mask pattern” (with the “original attention mask pattern” and the “proposed attention mask pattern” being as mapped above). The function computing the reward is the “deviation function” of the claim.) step d: determining whether an attention mask updating function can be created to fulfill the attention pattern proposal; ([Wu, algorithm 1, sections 3.2 and 4.1]: Element-wise multiplication by M [Wu, section 3.2] maps to the “attention mask updating function” of the claim. Wu discloses an iterative process in which, “for each training step” [Wu, algorithm 1 line 2], the algorithm uses the G-Net to produce a mask to apply to the A-Net [Wu, algorithm 1 line 4 and section 4.1]. The algorithm stops creating masks when this “for” loop [Wu, algorithm 1 lines 2-11] terminates; i.e., when the “for” loop has not yet terminated, a mask and thus the “attention mask updating function” can be created, and when the “for” loop has terminated, a mask and thus the “attention mask updating function” cannot be created. In other words, the check regarding whether the “for” loop has terminated [Wu, algorithm 1 line 2] falls under the broadest reasonable interpretation of “determining whether an attention mask updating function can be created” as recited by the claim. Since the mask matrix M maps to the “attention [mask] pattern proposal” of the claim, and element-wise multiplication by M maps to the “attention mask updating function” of the claim, the “attention mask updating function” as mapped here has the property that it “fulfill[s] the attention [mask] pattern proposal” in the sense that it is related to the “attention [mask] pattern proposal” as mapped above; cf. 112(b) rejections.) if the attention mask updating function can be created to fulfill the attention mask pattern proposal, creating an attention mask updating function; else: determining whether a new attention pattern proposal can be obtained; if a new attention pattern proposal can be obtained, repeating the method steps starting from step b; else, adopting a reinforcement learning (RL) model as the attention mask updating function; ([Wu, section 3.2]: As noted under examiner’s remarks, this recites conditional limitations: the claim does not require either that the attention mask updating function can be created or that it cannot; and if it cannot, the claim also does not require either that a new attention pattern proposal can be obtained or that it cannot. The claim thus also does not require any steps that are not required to be performed because conditions are not met. As noted above, Wu already discloses “creating the attention mask updating function”, which means that it also discloses “creating attention mask updating function if the attention mask updating function can be created to fulfill the attention [mask] pattern proposal” as required by the first of the at least three species encompassed by these limitations as best understood by the examiner (cf. examiner’s remarks). MPEP 2131.02(I) indicates that a species always anticipates a genus, so the genus claim as presented cannot be allowed over the species disclosed by Galassi in view of Wu. The examiner notes that Wu also discloses iteratively generating new actions a_t, i.e., new mask matrices [Wu, section 4], and that the rewards mechanism of Wu is a reinforcement learning system.) wherein the fulfillment of the attention pattern proposal being the original attention mask pattern is manipulatable by the attention mask updating function to approach the proposed attention mask pattern; ([Wu, algorithm 1, sections 3.2 and 4.1]: Since the mask matrix M maps to the “attention [mask] pattern proposal” of the claim, and element-wise multiplication by M maps to the “attention mask updating function” of the claim, the “attention mask updating function” as mapped here has the property that it “fulfill[s] the attention [mask] pattern proposal” in the sense that it is related to the “attention [mask] pattern proposal” as mapped above; cf. 112(b) rejections.) step e: combining the attention mask updating function with the original attention function to form an updated attention function of the attention layer. ([Galassi, figure 4; Wu, section 3]: As noted above, the overall attention model taking keys, queries, and values as input and outputting a context vector [Galassi, figure 4; see also, table IV] corresponds to scaled dot-product attention [Wu, section 3.1 equation (1)] of Wu and maps to the “original attention function” of the claim. Moreover, element-wise multiplication by M [Wu, section 3.2] maps to the “attention mask updating function” of the claim, and scaled dot-product attention. Wu discloses combining these [Wu, section 3.2 equation (2)]; the resulting combination maps to the “updated attention function” of the claim.) Before the effective filing date of the invention, it would have been obvious to a person of ordinary skill in the art to combine attention mechanisms as disclosed with Galassi with the method of generating dropout masks disclosed by Wu because the “proposed approach is universal and qualified to enable more robust task-specific tuning” [Wu, section 7], thereby resulting in a more effective system overall. Claim 4 Galassi in view of Wu discloses the elements of the parent claim(s). It also discloses: [The method according to claim 1, wherein the combining of the attention mask updating function with the original attention function to form the updated attention function of the attention layer comprises:] training the workflow-based neural network to learn through backpropagation with an auxiliary loss function defined with the original attention function and the attention mask updating function. ([Galassi, section IV.A; Wu, section 4]: As noted under the parent claim, Galassi discloses training the neural network [Galassi, section IV.A first paragraph]. Wu discusses more details about training [Wu, section 4]. The function J(theta_G) [Wu, section 4.2 second displayed equation] maps to the “auxiliary loss function” of the claim; it is defined in terms of the rewards r_t provided to G-Net [Wu, section 4.2 first paragraph], which are based on the comparison between A-Net and D-Net [Wu, section 4.1 paragraph beginning “Generator”], which means that J(theta_G) is “defined with the original attention function and the attention mask updating function” as required by the claim. The use of “policy gradient to update theta_G” [Wu, section 4.2 first paragraph] (i.e., using [Wu, section 4.2 equation (5)]) falls under the broadest reasonable interpretation of “learn[ing] through backpropagation with an auxiliary loss function” as recited by the claim.) The same motivation to combine applies. Claim 5 Galassi in view of Wu discloses the elements of the parent claim(s). It also discloses: [The method according to claim 1, wherein the combining of the attention mask updating function with the original attention function to form the updated attention function of the attention layer comprises:] directly applying the attention mask updating function to the original attention mask pattern in the attention layer. ([Wu, section 3.2]: In Wu, element-wise multiplication by M (i.e., the “attention mask updating function” of the claim) is applied directly to the softmax expression (i.e., the “original attention mask pattern” of the claim).) The same motivation to combine applies. Claim 6 Galassi in view of Wu discloses the elements of the parent claim(s). It also discloses: [The method according to claim 1, wherein] the original attention function is a scaled dot-product attention function ([Wu, section 3.1]: Wu discloses scaled-dot product attention [Wu, section 3.1 equation (1)]. The examiner notes that scaled dot-product attention can also be found in Galassi by taking the compatibility function to be the “scaled multiplicative” one [Galassi, table IV] and the distribution function to be the “softmax function” [Galassi, section IV.C second paragraph].) having input including a query matrix of queries and a key matrix of keys, both of a dimension d_k, and a value matrix of values; ([Wu, section 3.1]: The matrices Q, K, and V of Wu map, respectively, to the “key matrix of queries”, the “key matrix of keys”, and the “value matrix of values” of the claim. The variable d_k of Wu corresponds to the identically named variable of the claim. The examiner notes that Wu cites the reference Vaswani for details about scaled dot-product attention [Wu, section 3 first paragraph]; the reference Vaswani, made of record in the conclusion of this Office action, indicates more explicitly that d_k is the dimension of Q and K.) wherein the attention weights are obtained by computing dot products of the queries and the keys, then dividing each of the dot products by a square root of d_k; ([Wu, section 3.1]: The computation of QK^T in [Wu, section 3.1 equation (1)] is a computation of “dot products of the queries and the keys” as recited by the claim, and the division by the square root of d_k in [Wu, section 3.1 equation (1)] is “dividing each of the dot products by a square root of d_k” as recited by the claim.) wherein the obtained attention weights are applied to the values to obtain the attention function. ([Wu, section 3.1]: The dot product with V [Wu, section 3.1 equation (1)] maps to “appl[ying the obtained attention weights] to the values” as recited by the claim.) The same motivation to combine applies. Claim 7 Galassi in view of Wu discloses the elements of the parent claim(s). It also discloses: [The method according to claim 6, wherein] the original attention function is expressed as: PNG media_image1.png 34 267 media_image1.png Greyscale wherein Q represents the query matrix of queries, K represents the key matrix of keys, d_k represents a dimension of both the query matrix of queries and the key matrix of keys, and V represents the value matrix of values; ([Wu, section 3.1]: The formula [Wu, section 3.1 equation (1)] is identical to the formula appearing in the claim.) and wherein the updated attention function is expressed as: PNG media_image2.png 39 288 media_image2.png Greyscale wherein f() represents the attention mask updating function. ([Wu, section 3.2]: As noted under the parent claims, element-wise multiplication by the mask matrix M maps to the “attention mask updating function” of the claim. In other words, taking the function f of the claim to be element-wise multiplication by M, the formula [Wu, section 3.2 equation (2)] is identical to the formula appearing in the claim.) The same motivation to combine applies. Claim 17 Galassi in view of Wu discloses the elements of the parent claim(s). It also discloses: [The method according to claim 1, wherein] the prediction model is for prediction of words and phrases, prediction of textual character per timestep, or prediction of natural language translation. ([Galassi, section I]: Galassi discloses the use of neural networks for numerous natural language processing tasks, such as “machine translation” (cf. “prediction of natural language translation” as recited in the claim) and “visual question-answering” (cf. “prediction of words and phrases” as recited in the claim).) The same motivation to combine applies. Claim(s) 2 and 18-19 is/are rejected under 35 USC 103 as being unpatentable over Galassi in view of Wu, further in view of Kelvin XU et al. (Show, Attend and Tell: Neural Image Caption Generation with Visual Attention, published 2016-04-19; hereafter, “Xu”). Claim 2 Galassi in view of Wu discloses the elements of the parent claim(s). It also discloses: [The method according to claim 1, wherein the generating of the attention mask pattern proposal comprises:] identifying one or more elements of the input elements based on one or more features extracted from the input elements; ([Galassi, section IV.C]: Galassi indicates that masks are used to “to focus the attention only on a specific portion of the input” because “relevant features are found in a neighborhood of a certain position” [Galassi, section IV.C paragraph beginning “In some tasks”]. The specific portion of the input on which attention is focused using the masks of Wu maps to the “one or more elements” of the claim, and the relevant features map to the “one or more features” of the claim.) visualizing the attention layer to identify the original attention mask pattern based on one or more attentions and corresponding attention weights being placed by the attention layer in relation to the identified elements of the input elements and one or more original prediction results; ([Galassi, figures 3-7; Wu, figure 2]: The figures [Galassi, figures 3-7] and [Wu, figure 2] are all visualizations of attention mechanisms, including the attention weights and their relationship to the inputs. The applicant is also invited to consult [Xu, figures 2-3 and 5-15] in the combination as proposed below.) and creating the proposed attention mask pattern based on one or more desired attention placements and corresponding desired attention weights on the input elements if the attentions and the corresponding attention weights are not being placed on the identified elements of the input elements that correspond with the original prediction results. ([Wu, sections 1 and 3]: This recites a conditional limitation, since the claim does not require the attentions and corresponding attention weights “not being placed on the identified elements” as recited by the claim. Consequently, the broadest reasonable interpretation of the claim also does not require any actions that are contingent on this condition. Nonetheless, since the result of element-wise multiplication [Wu, section 3.2 equation (2)] was mapped above to the “proposed attention mask pattern” of the claim, the process of performing this multiplication maps to the “creating” step of the claim. The mask matrix, for example, maps to the “one or more desired attention placements and corresponding desired attention weights” of the claim (the indices of the matrix mapping to the “desired attention placements” and the entries to the “corresponding desired attention weights”).) The same motivation to combine applies. Galassi in view of Wu might not distinctly disclose: and determining, based on the visualization, whether the attentions and the corresponding attention weights are being placed on the identified elements of the input elements and that correspond with the original prediction results; Xu is in the field of machine learning. It describes an image captioning system with an attention mechanism [Xu, abstract]. Moreover, Galassi in view of Wu and Xu discloses: and determining, based on the visualization, whether the attentions and the corresponding attention weights are being placed on the identified elements of the input elements and that correspond with the original prediction results; ([Xu, figures 3 and 5-15]: Xu describes “visualizing ‘where’ and ‘what’ the attention focused on” [Xu, section 1 last paragraph second bullet point]. In the combination, the elements where attention is focused correspond to the “identified [one or more] elements” of the claim as mapped above. Examples of such visualizations are given in [Xu, figures 3 and 5-15] and include both “[e]xamples of attending to the correct object” [Xu, figure 3] and “[e]xamples of mistakes” [Xu, figure 5]. A determination regarding whether or not the model is attending to the correct object maps to the “determining” step of the claim, since this is a determination “based on the visualization” regarding “whether the attentions and the corresponding attention weights are being placed on the identified elements”, as recited by the claim.) Before the effective filing date of the invention, it would have been obvious to a person of ordinary skill in the art to combine the attention mechanisms of Galassi in view of Wu with the attention visualization mechanisms disclosed in Xu because it allows for “gain[ing] insight and interpret[ing] the results of this framework” [Xu, section 1 last paragraph second bullet point], thereby resulting in a more interpretable and transparent system. (The examiner notes that Galassi itself cites Xu as reference [16].) Claim 18 Galassi in view of Wu and Xu discloses the elements of the parent claim(s). It also discloses: [The method according to claim 2, wherein] the features extracted from the input elements comprise one or more of spatial features, temporal features, or contextual features. ([Galassi, section IV.C]: As noted under the parent claim, Galassi indicates that masks are used to “to focus the attention only on a specific portion of the input” because “relevant features are found in a neighborhood of a certain position” [Galassi, section IV.C paragraph beginning “In some tasks”]. In other words, the “features” as mapped under the parent claim fall under the broadest reasonable interpretation of at least of either the “spatial features” or the “contextual features” as recited by the claim. The applicant is invited to consult [Galassi, section II.A] regarding “temporal features” of the claim.) The same motivation to combine applies. Claim 19 Galassi in view of Wu and Xu discloses the elements of the parent claim(s). It also discloses: [The method according to claim 2, wherein] the determination of whether the attentions and corresponding attention weights are being placed on the identified elements of the input elements that correspond with the original prediction results is performed by a human observer ([Xu, figures 3 and 5-15; section 5.4]: As noted under the parent claim, a determination regarding whether or not the model is attending to the correct object maps to the “determining” step of the claim. These determinations are made by humans, since the determination is regarding whether alignments correspond with “human intuition” [Xu, section 5.4 last paragraph]. For example, a human observer of the system who sorted examples between [Xu, figure 3] and [Xu, figure 5] (e.g., one of the human authors of the paper) could map to the “human observer” of the claim. Alternatively, a human reader of the paper studying the examples of [Xu, figures 6-15] could also map to a “human observer” of the claim when they determine, for example, that the attention placement of [Xu, figure 8(b) image corresponding to “dog”] is correct while that of [Xu, figure 8(a) image corresponding to “book”] is not.) such that domain knowledge is injected into the optimization of the workflow-based neural network. ([Galassi, section III-V]: Galassi discloses training neural networks [Galassi, section III paragraph beginning “For tasks”; section IV.A first two paragraphs; section V.A, etc]. Training a neural network is a form of “optimization of the workflow-based neural network” and the training data maps to the “domain knowledge” of the claim.) The same motivation to combine applies. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Meng-Hao GUO et al. (Attention mechanisms in computer vision: A survey, published 2022; hereafter, “Guo”) surveys numerous attention mechanisms in computer vision, including the mechanism described in Xu (as reference [35]). The applicant is invited to consult [Guo, section 3.3] in particular. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Shishir AGRAWAL whose telephone number is +1 703-756-1183. The examiner can normally be reached Monday through Thursday, 08:30-14:30 Pacific Time. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey SHMATOV can be reached on +1 571-270-3428. The fax phone number for the organization where this application or proceeding is assigned is +1 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at +1 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call +1 800-786-9199 (IN USA OR CANADA) or +1 571-272-1000. /S.A./Examiner, Art Unit 2123 /ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123
Read full office action

Prosecution Timeline

Mar 07, 2023
Application Filed
May 08, 2026
Non-Final Rejection mailed — §103, §112
Aug 06, 2026
Response Filed
Sep 15, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725051
RANKING DATA SLICES USING MEASURES OF INTEREST
4y 7m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 1 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
8%
Grant Probability
24%
With Interview (+15.4%)
4y 0m (~5m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 24 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month