DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Amendment
Applicant submitted amendments on 06/16/2026. The Examiner acknowledges the amendment and has reviewed the claims accordingly.
Priority
Applicant claims the benefit of US Provisional Application No. 63/509,590, filed 06/22/2023. Claims 1-20 have been afforded the benefit of this filing date.
Information Disclosure Statement
The IDS dated 03/22/2024 and 11/01/20247 have been considered and placed in the application file.
Applicant Arguments:
In regards to the argument on Argument 1, Applicant/s state/s “To expedite prosecution, the claims have been amended to recite, in relevant part, a "spatial resolution" of the transformed version of image pixels. Applicant therefore respectfully requests withdrawal of the rejections.” (See Remarks Pg 6, paragraph 5). Therefore the 35 U.S.C 112(b) rejection should be withdrawn.
In regards to the argument on Argument 2, Applicant/s state/s “The memorandum titled "Reminders on evaluating subject matter eligibility of claims under 35 U.S.C. 101," dated August 4, 2025 (hereinafter the "August Memorandum"), reiterates that "Examiners should be careful to distinguish claims that recite an exception (which require further eligibility analysis) from claims that merely involve an exception (which are eligible and do not require further eligibility analysis)." Page 3 ( emphasis added). Applicant respectfully submits that the present claims, at most, "merely involve" an exception, and clearly do not "recite" such an exception.” (See Remarks Pg 7, paragraph 3). Therefore U.S.C 101 rejection should be withdrawn.
In regards to the argument on Argument 3, Applicant/s state/s “Applicant respectfully submits that even if these elements "involve" or "relate to" mathematical concepts, such concepts are plainly not recited in the present claims. For example, even if applying attention operations involves use of "matrix multiplications," the claims do not recite any such multiplications. Similarly, even if selecting a number of operations to apply may "involve" using mathematical relationships, the claims do not recite any such relationship. In this regard, the present claims are substantially similar to those discussed in the August Memorandum, which at most "involve" or "rely upon" mathematical concepts without actually reciting them (e.g., without specifying "specific mathematical calculations," as noted in the August Memorandum).” (See Remarks Pg 8, paragraph 2). Therefore U.S.C 101 rejection should be withdrawn.
In regards to the argument on Argument 4, Applicant/s state/s “Here, even if the amended claims are found to recite a judicial exception (which they do not), they are eligible because they "reflect[] an improvement in the functioning of a computer, or an improvement to other technology or technical field," and further "integrate" the alleged "judicial exception into a practical application of the exception. "the cited references, taken alone or in combination, fail to disclose, suggest, or otherwise render obvious the features recited in the independent claims. Thus, Applicant submits that the independent claims are allowable. Applicant further submits that the dependent claims are allowable at least by virtue of their dependence on the independent claims, and for the additional features recited.” (See Remarks Pg 9, paragraph 2). Therefore U.S.C 101 rejection should be withdrawn.
In regards to the argument on Argument 5, Applicant/s state/s “Applicant respectfully submits that the present claims are clearly eligible as discussed above, and that the present claims at least reflect a "close call" such that unpatentability cannot be established by a preponderance of the evidence, as required.” (See Remarks Pg 10, paragraph 2). Therefore U.S.C 101 rejection should be withdrawn.
In regards to the argument on Argument 6, Applicant/s state/s “Applicant respectfully submits that the present rejections similarly "evaluate [the] claims at [] a high level of generality" without adequate explanation.” (See Remarks Pg 10, paragraph 3). Therefore U.S.C 101 rejection should be withdrawn.
In regards to the argument on Argument 7, Applicant/s state/s “Applicant respectfully submits that selectively dropping tokens between attention operations, as discussed in Rao, clearly does not teach or suggest "select[ing] a number of local attention operations to apply, in one transformer," as recited in the pre-amended claims. That is, Rao contemplates processing a selected subset of tokens using the entire sequence of attention operations ( e.g., as 12-layer transformer), but does not teach or suggest selecting a subset of that sequence of operations. Further, claim 1 has been amended to recite, in part, "select a subset of local attention operations, from the sequence of local attention operations, to apply to the transformed version of image pixels." Rao does not teach or suggest this element of the claims.” (See Remarks Pg 11, paragraph 4). Therefore the U.S.C 35 103 rejection of Claim 1 should be withdrawn.
In regards to the argument on Argument 8, Applicant/s state/s “Applicant respectfully submits that Wang appears to be entirely silent with respect to selecting any "number of local attention operations," despite the Office's assertion, much less an architecture where the number of local attention operations "scales proportionally with the spatial size of the input feature map." The only reference to "proportional" scales in Wang is a brief mention that the "computational budget of a convolutional layer" is "proportional to" the "input/output dimension." Page 4, Column 2, paragraph 2. However, this clearly does not teach or suggest the "number of local attention operations" being proportional to such dimensionality.” (See Remarks Pg 12, paragraph 2). Therefore the U.S.C 35 103 rejection of Claim 1 should be withdrawn.
In regards to the argument on Argument 9, Applicant/s state/s “Applicant assumes that the Office is referring to Wang's discussion of "groups" and performing "self-attention within each group," where the "group size" can be enlarged progressively in deeper layers. Page 2, Column 1, Paragraph 4, and Page 3, Column 1, Paragraph 1. That is, in theory, using a given group size to process input having a larger spatial resolution would necessitate a larger number of groups, thereby resulting in a larger number of "self-attention" operations ( one for each group). However, Applicant respectfully submits that even this charitable interpretation does not teach or suggest selecting a subset of local attention operations, "from the sequence of local attention operations" in a transformer, to apply based on the input size, as recited in the amended claims.” (See Remarks Pg 12, paragraph 4). Therefore the U.S.C 35 103 rejection of Claim 1 should be withdrawn.
In regards to the argument on Argument 10, Applicant/s state/s “Applicant respectfully submits that the cited references neither teach nor suggest at least these elements of claim 1. For example, Liu describes a "hierarchical transformer" that applies "self-attention computation to non-overlapping local windows" (see Abstract), but does not teach or suggest selecting and applying a subset of local attention operations from a sequence of local attention operations in a transformer. Similarly, as discussed above, Rao describes pruning redundant tokens between attention operations, but does not teach or suggest selecting and applying a subset of such attention operations. Additionally, as discussed above, Wang contemplates progressively expanding the "group size" of self-attention operations, but does not teach or suggest selecting a subset of local attention operations "based on a spatial resolution" of the input.” (See Remarks Pg 13, paragraph 3). Therefore the U.S.C 35 103 rejection of Claim 1 should be withdrawn.
Examiner’s Responses:
In response to Argument 1, Applicant’s arguments, see Remarks, filed 06/16/2026, with respect to the 35 U.S.C 112(b) rejection pertaining to “size”, the amendments and arguments have been considered persuasive and the 112(b) rejection pertaining to “size” has been withdrawn.
In response to Arguments 2-6, Applicant’s arguments, see Remarks, filed 06/16/2026, with respect to the U.S.C 101 rejection of Claim 1-19 have been considered and are persuasive. Therefore, the U.S.C 101 rejection has been withdrawn due to amendments.
In response to Arguments 7-10, Applicant’s arguments, see Remarks, filed 06/16/2026, with respect to the U.S.C 103 rejections of Claim 1 have been considered but are moot in view of new ground(s) of rejection caused by the amendments. A new ground(s) of rejection is made for claims 1- 20 under 35 U.S.C. 103 in view of Lui et al (Swin transformer: Hierarchical vision transformer using shifted windows." Proceedings of the IEEE/CVF international conference on computer vision. 2021, hereafter referred to as Lui) in view of Rao et al (Dynamicvit: Efficient vision transformers with dynamic token sparsification." Advances in neural information processing systems 34 (2021): 13937-13949, hereafter referred to as Rao) in further view of Wang et at (Wang, Wenxiao, et al. "CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention." arXiv e-prints (2021): arXiv-2108, hereafter referred to as Wang) in further view of Hu et al (Hu, Yunqing, et al. "Rams-trans: Recurrent attention multi-scale transformer for fine-grained image recognition." Proceedings of the 29th ACM international conference on multimedia. 2021. Hereafter referred to as Hu).
The Examiner finds that Rao teaches on the amended claim language “selecting a subset” in amended claim 1.
Specifically, Rao teaches dynamically selecting attention operations to apply in Section 3.1 ¶01 and Figure 2. Applicant argues that “Applicant respectfully submits that selectively dropping tokens between attention operations, as discussed in Rao, clearly does not teach or suggest "select[ing] a number of local attention operations to apply, in one transformer," as recited in the pre-amended claims. That is, Rao contemplates processing a selected subset of tokens using the entire sequence of attention operations ( e.g., as 12-layer transformer), but does not teach or suggest selecting a subset of that sequence of operations. Further, claim 1 has been amended to recite, in part, "select a subset of local attention operations, from the sequence of local attention operations, to apply to the transformed version of image pixels." Rao does not teach or suggest this element of the claims.”. However, we determine claim scope not solely on the basis of claim language, but also on giving claims their broadest reasonable construction in light of the specification as it would be interpreted by one of ordinary skill in the art. In re Am. Acad. of Sci. Tech. Ctr., 367 F.3d 1359, 1364 (Fed. Cir. 2004). See also Superguide Corp. v. DirecTV Enterprises, Inc., 358 F.3d 870, 875 (Fed. Cir. 2004) (“Though understanding the claim language may be aided by explanations contained in the written description, it is important not to import into a claim limitations that are not part of the claim.”). The Examiner interprets that under broadest reasonable interpretation “subset” has no special definition in the claims, and therefore can be interpreted as any local attention operation, therefore this can be interpreted as selecting an attention operation. Therefore, the Examiner interprets that Lui teaches the main concept of a machine learning model transforming image pixels, the additional details of the functions of the main concepts as stated above by the applicant in the amendments is taught by Rao, Wang, and Hu in the details of the rejection below. The Examiner will maintain prior art Rao and details of the rejection are below.
The Examiner finds that Wang teaches on the amended claim language “based at least in part on a spatial resolution” in amended claim 1.
Specifically, Wang teaches that the local attention operations are related to the spatial resolution in Pg 2 ¶03-¶04 and ¶07. Applicant argues that “Applicant respectfully submits that Wang appears to be entirely silent with respect to selecting any "number of local attention operations," despite the Office's assertion, much less an architecture where the number of local attention operations "scales proportionally with the spatial size of the input feature map." The only reference to "proportional" scales in Wang is a brief mention that the "computational budget of a convolutional layer" is "proportional to" the "input/output dimension." Page 4, Column 2, paragraph 2. However, this clearly does not teach or suggest the "number of local attention operations" being proportional to such dimensionality.” And “Applicant assumes that the Office is referring to Wang's discussion of "groups" and performing "self-attention within each group," where the "group size" can be enlarged progressively in deeper layers. Page 2, Column 1, Paragraph 4, and Page 3, Column 1, Paragraph 1. That is, in theory, using a given group size to process input having a larger spatial resolution would necessitate a larger number of groups, thereby resulting in a larger number of "self-attention" operations ( one for each group). However, Applicant respectfully submits that even this charitable interpretation does not teach or suggest selecting a subset of local attention operations, "from the sequence of local attention operations" in a transformer, to apply based on the input size, as recited in the amended claims.” The Examiner respectfully disagrees. The Examiner made a proper determination of obviousness under 35 U.S.C. §103, and also provided an appropriate supporting rationale in view of the decision by the Supreme Court in KSR International Co. v. Teleflex Inc. (KSR), 550 U.S. 398, 82 USPQ2d 1385 (2007). The Examiner’s rational are based on the Office’s current understanding of the law, and are believed to be fully consistent with the binding precedent of the Supreme Court. Furthermore, the Examiner supported the rejection under 35 U.S.C. §103 via making the clear articulation of the reason(s) why the claimed invention would have been obvious by citing the specific areas in the prior art references. Further the Examiner, clearly stating the modification of the inventions, supported the rejection under 35 U.S.C. §103 by making the analysis explicit. Last, the Examiner did not make conclusory statements. The Court quoting In re Kahn, 441 F.3d 977, 988, 78 USPQ2d 1329, 1336 (Fed. Cir. 2006), stated that “‘[R]ejections on obviousness cannot be sustained by mere conclusory statements; instead, there must be some articulated reasoning with some rational underpinning to support the legal conclusion of obviousness.’” KSR, 550 U.S. at ___, 82 USPQ2d at 1396. Therefore, the Examiner has established a proper 35 U.S.C. §103 rejection with Lui, in view of Rao, in view of Wang, in view of Hu, which is disclosed in detail below. Therefore, the Examiner interprets that Lui teaches the main concept of a machine learning model transforming image pixels, the additional details of the functions of the main concepts as stated above by the applicant in the amendments is taught by Rao, Wang, and Hu in the details of the rejection below. The Examiner will maintain prior art Rao and Wang and details of the rejection are below.
The Examiner finds that Hu teaches on the amended claim language “the transformer comprising a sequence of local attention operations and a global attention operation” and of local attention operations, from the sequence of local attention operation” in amended claim 1.
Specifically, Hu teaches the transformer comprising a sequence of local attention operations and a global attention operation in and Figure 2 and Section 2.2. Hu also teaches based on spatial resolution in Section 3.1, adding additional support even though the examiner finds that this limitation is also taught by Wang . Hu is brought in as an additional reference based on the amended claims and remedies the deficiencies of Lui, Rao, and Wang. Therefore, the Examiner interprets that Lui teaches the main concept of a machine learning model transforming image pixels, the additional details of the functions of the main concepts as stated above by the applicant in the amendments is taught by Rao, Wang, and Hu in the details of the rejection below. The Examiner will maintain prior art Lui, Rao, and Wang and details of the rejection are below.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1’s term “subset” renders claim 1 indefinite. The specification (US 20240428576) uses “subset” to refer only to including use of positional embeddings to a subset of the matrices in ¶0064. There is no mention of subset in the specification in relation to how the local attention operations are selected or how many are selected. The only contextual guidance offered in the selection discussion, is performing the local attention in a specific arrangement based on use of positional embeddings to a subset of matrixes at ¶0064. Because the specification does not resolve which quantities constitutes the operative “subset” for the purposes of the selection step, the metes and bounds of claim 1 cannot be determined with reasonable certainty. Nautilus, Inc. v. Biosig Instruments, Inc., 572 U.S. 898, 910 (2014).
Claim(s) 2-20 depend either directly or indirectly from the rejection(s) of claim(s) 1, therefore they are also rejected. Appropriate correction is required.
Claim Interpretation
The claims in this application are given their broadest reasonable interpretation using the
plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification.
Under MPEP 2143.03, "All words in a claim must be considered in judging the patentability of that claim against the prior art." In re Wilson, 424 F.2d 1382, 1385, 165 USPQ 494, 496 (CCPA 1970). As a general matter, the grammar and ordinary meaning of terms as understood by one having ordinary skill in the art used in a claim will dictate whether, and to what extent, the language limits the claim scope. Language that suggests or makes a feature or step optional but does not require that feature or step does not limit the scope of a claim under the broadest reasonable claim interpretation. In addition, when a claim requires selection of an element from a list of alternatives, the prior art teaches the element if one of the alternatives is taught by the prior art. See, e.g., Fresenius USA, Inc. v. Baxter Int’l, Inc., 582 F.3d 1288, 1298, 92 USPQ2d 1163, 1171 (Fed. Cir. 2009).
Claim 19 recite “at least one of” then listing “a depth map, a classification, or a segmentation map.”. Since “at least one” is disjunctive, any one of the elements found in the prior art is sufficient to reject the claim. While citations have been provided for completeness and rapid prosecution, only one element is required. Because, on balance, it appears the disjunctive interpretation enjoys the most specification support and for that reason the disjunctive interpretation (one of A, B OR C) is being adopted for the purposes of this Office Action. Applicant’s comments and/or amendments relating to this issue are invited to clarify the claim language and the prosecution history.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-20 are rejected under 35 U.S.C. 103 as obvious over Lui et al (Swin transformer: Hierarchical vision transformer using shifted windows." Proceedings of the IEEE/CVF international conference on computer vision. 2021, hereafter referred to as Lui) in view of Rao et al (Dynamicvit: Efficient vision transformers with dynamic token sparsification." Advances in neural information processing systems 34 (2021): 13937-13949, hereafter referred to as Rao) in further view of Wang et at (Wang, Wenxiao, et al. "CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention." arXiv e-prints (2021): arXiv-2108, hereafter referred to as Wang) in further view of Hu et al (Hu, Yunqing, et al. "Rams-trans: Recurrent attention multi-scale transformer for fine-grained image recognition." Proceedings of the 29th ACM international conference on multimedia. 2021. Hereafter referred to as Hu).
Regarding Claim 1, Lui teaches a processing system in a device (Lui Pg 3 Col 1 ¶02 and Pg 2 Col 2 ¶03 discloses implementation of the method in general hardware which is well understood to include processing), comprising:
memory configured to store machine learning model parameters (Lui Pg 2, Col 1, ¶02 discloses all query patches within a window share the same key set, which facilitates memory access in hardware); and
one or more processors, coupled to the memory (Lui Pg 2, Col 1, ¶02 discloses all query patches within a window share the same key set, which facilitates memory access in hardware), configured to:
access a transformed version of image pixels (Lui Fig 1 and Fig 2 disclose the transformed version of image pixels by merging image patches in a window partitioning scheme)
of a machine learning model (Lui Fig 2 and Pg 2 Col 1 ¶2 and Table 4 discloses self-attention in computed at every window and provides connections among the windows in the model)
to the transformed version of image pixels (Lui Fig 1 and Fig 2 disclose the transformed version of image pixels by merging image patches in a window partitioning scheme)
of the transformed version of image pixels (Lui Fig 1 and Fig 2 disclose the transformed version of image pixels by merging image patches in a window partitioning scheme); and
generate a transformer output (Lui Fig 2, Pg 2 Col 1 ¶1 discloses the output of the transformer being new windows where an additional self attention layer can be applied)
of the machine learning model (Lui Fig 2 and Pg 2 Col 1 ¶2 and Table 4 discloses self-attention in computed at every window and provides connections among the windows in the model)
to the transformed version of image pixels (Lui Fig 1 and Fig 2 disclose the transformed version of image pixels by merging image patches in a window partitioning scheme).
Lui does not explicitly teach select a subset, based on applying the selected subset and the global attention operation
Rao is in the same field of use same hierarchical vision transformer design and employing compatible attention mechanisms. Further, Rao teaches select a subset (Rao Section 3.1 ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply) based on applying the selected subset (Rao Section 3.1 ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply) and the global attention operation (Rao ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply Pg 4 Section 3.2 discloses the local and global attention operations being applied).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Lui by incorporating the runtime dynamic selection mechanism as taught by Rao, to make an invention that optimizes computational efficiency for variable resolution inputs; thus, one of ordinary skilled in the art would be motivated to combine the references since an object of the present inventions reduce complexity of vision transformers while maintaining accuracy. (Rao, Pg 2 ¶02).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
Lui and Rao in combination so not explicitly disclose based at least in part on a spatial resolution.
Wang is in the same field of use same hierarchical vision transformer design and employs compatible attention mechanisms. Further, Wang teaches based at least in part on a spatial resolution (Wang Pg 2 ¶03-04 and ¶07 and discloses a vision transformer where the number of local attention operations scales proportionally with the spatial size of the input feature map).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Lui in view of Rao by incorporating size-proportional attention scaling as taught by Wang, to make an invention that optimizes computational efficiency for variable resolution inputs; thus, one of ordinary skilled in the art would be motivated to combine the references since an object of the present inventions reduce complexity of vision transformers while maintaining accuracy. (Wang, Abstract and Introduction).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
Lui and Rao and Wang in combination so not explicitly disclose as input to a transformer, the transformer comprising a sequence of local attention operations and a global attention operation, of local attention operations from the sequence of local attention operations, based at least in part on a spatial resolution, for the transformer, of local attention operations, from the sequence of local attention operations.
Hu is in the same field of use same hierarchical vision transformer design and employs compatible attention mechanisms. Further, Hu teaches as input to a transformer (Hu Fig 2 discloses the input of the transformer being based off of transformed image patches input into the machine learning model) the transformer comprising a sequence of local attention operations (Hu Figure 2 and Section 2.2 disclose the use of local attention operations and use of attention weights for amplification and reuse of region attentions, meaning they are used multiple times which the examiner is interpreting as a sequence in the transformer machine learning model) and a global attention operation (Hu Fig 2 discloses the global attention operation also being a part of the transformer model)
of local attention operations from the sequence of local attention operations (Hu Figure 2 and Section 2.2 disclose the use of local attention operations and use of attention weights for amplification and reuse of region attentions, meaning they are used multiple times which the examiner is interpreting as a sequence in the transformer machine learning model),
based at least in part on a spatial resolution (Hu Section 3.1 discloses that the patch tokens input to the subsequent transformer are position-agnostic, and the image processing depends on the spatial information of each pixel)
for the transformer (Hu Fig 2 discloses the input of the transformer being based off of transformed image patches input into the machine learning model)
of local attention operations, from the sequence of local attention operations (Hu Figure 2 and Section 2.2 disclose the use of local attention operations and use of attention weights for amplification and reuse of region attentions, meaning they are used multiple times which the examiner is interpreting as a sequence in the transformer machine learning model).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Lui in view of Rao in view of Wang by incorporating a sequence of local attention operations and global attention operations to comprise the transformer as taught by Hu, to make an invention that optimizes computational efficiency for variable resolution inputs; thus, one of ordinary skilled in the art would be motivated to combine the references since an object of the present invention is to introduce attention to local regions in the model, that is, extend the effective receptive field, the recognition performance of the model is likely to be further improved. (Hu, Abstract and Introduction).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
Regarding Claim 2, Lui in view of Rao in further view of Wang in further view of Hu teaches the processing system of claim 1, wherein the one or more processors (Lui Pg 2, Col 1, ¶02 discloses all query patches within a window share the same key set, which facilitates memory access in hardware which is well understood would include processors) are configured to:
generate a saliency map (Rao Abstract, Section 3.1 ¶1 and Section 4.3 ¶05 disclose importance score of each token given the current features being used to determine the selected pixels) and based on the transformed version of image pixels (Lui Fig 1 and Fig 2 disclose the transformed version of image pixels by merging image patches in a window partitioning scheme); and
determine a semantic complexity (Lui Pg 2 Col 1 ¶01 and Pg 3 Col 2 ¶01 discloses semantic segmentation to require dense prediction at the pixel level) of the transformed version of image pixels (Lui Fig 1 and Fig 2 disclose the transformed version of image pixels by merging image patches in a window partitioning scheme) based on the saliency map. See Claim 1 for rationale, its parent claim.
Regarding Claim 3, Lui in view of Rao in further view of Wang in further view of Hu teaches the processing system of claim 2, wherein, to select the subset of local attention operations (Rao Section 3.1 ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply), the one or more processors (Lui Pg 2, Col 1, ¶02 discloses all query patches within a window share the same key set, which facilitates memory access in hardware which is well understood would include processors) are configured to select the subset of local attention operations (Rao Section 3.1 ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply) based on a number of contextual objects (Rao Section 3.2 discloses the global features containing the context of the whole image) indicated in the saliency map (Rao Abstract, Section 3.1 ¶1 and Section 4.3 ¶05 disclose importance score of each token given the current features being used to determine the selected pixels). See Claim 1 for rationale, its parent claim.
Regarding Claim 4, Lui in view of Rao in further view of Wang in view of Hu teaches the processing system of claim 3, wherein, to select the subset of local attention operations (Rao Section 3.1 ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply), the one or more processors (Lui Pg 2, Col 1, ¶02 discloses all query patches within a window share the same key set, which facilitates memory access in hardware which is well understood would include processors) are configured to compare the number of contextual objects against one or more thresholds (Wang Pg 4-5 Section 3.2.1 and discloses the attention operations being proportional to the features including the contextual relationship of the small scale embeddings to the large scale embeddings) to select the subset of local attention operations (Rao Section 3.1 ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply). See Claim 1 for rationale, its parent claim.
Regarding Claim 5, Lui in view of Rao in further view of Wang in view of Hu teaches the processing system of claim 3, wherein a number of the selected subset of local attention operations (Rao Section 3.1 ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply), is directly proportional to the number of contextual objects (Wang Pg 4-5 Section 3.2.1 and discloses the attention operations being proportional to the features including the contextual relationship of the features). See Claim 1 for rationale, its parent claim.
Regarding Claim 6, Lui in view of Rao in further view of Wang in view of Hu teaches the processing system of claim 3, wherein, to select the subset of local attention operations, (Rao Section 3.1 ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply), the one or more processors (Lui Pg 2, Col 1, ¶02 discloses all query patches within a window share the same key set, which facilitates memory access in hardware which is well understood would include processors) are configured to select at least two local attention operations (Wang Pg 4-5 Section 3.2.1 and discloses the attention operations being proportional to the features including the contextual relationship of the features including the small scale embeddings to the large scale embeddings) based on a determination that the number of contextual objects satisfies a defined threshold (Wang Pg 4-5 Section 3.2.1 and discloses the attention operations being proportional to the features including the contextual relationship of the features including the small scale embeddings to the large scale embeddings therefore if there are two contextual features there will be two attention operators). See Claim 1 for rationale, its parent claim.
Regarding Claim 7, Lui in view of Rao in further view of Wang in view of Hu teaches the processing system of claim 3, wherein, to select the subset of local attention operations, (Rao Section 3.1 ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply), the one or more processors (Lui Pg 2, Col 1, ¶02 discloses all query patches within a window share the same key set, which facilitates memory access in hardware which is well understood would include processors) are configured to:
obtain a display resolution of a display device (Lui Pg 4 Col 1 ¶01 discloses outputting a high resolution image) included in the processing system (Lui Pg 3 Col 1 ¶02 and Pg 2 Col 2 ¶03 discloses implementation of the method in general hardware which is well understood to include processing); and
select three local attention operations (Wang Pg 4-5 Section 3.2.1 and discloses the attention operations being proportional to the features including the contextual relationship of the features including the small scale embeddings to the large scale embeddings) , in the transformer, when a display resolution is set to at least a maximum spatial resolution (Hu Section 3.1 discloses that the patch tokens input to the subsequent transformer are position-agnostic, and the image processing depends on the spatial information of each pixel) (Lui Pg 4 Col 1 ¶01 discloses outputting a high resolution image) of the transformed version of image pixels (Lui Fig 1 and Fig 2 disclose the transformed version of image pixels by merging image patches in a window partitioning scheme) and the number of contextual objects is three or more (Wang Pg 4-5 Section 3.2.1 and discloses the attention operations being proportional to the features including the contextual relationship of the features including the small scale embeddings to the large scale embeddings therefore if there are three contextual features there will be three attention operators). See Claim 1 for rationale, its parent claim.
Regarding Claim 8, Lui in view of Rao in further view of Wang in view of Hu teaches the processing system of claim 3, wherein, to select the subset of local attention operations, (Rao Section 3.1 ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply), the one or more processors (Lui Pg 2, Col 1, ¶02 discloses all query patches within a window share the same key set, which facilitates memory access in hardware which is well understood would include processors) are configured to :
obtain a display resolution of a display device (Lui Pg 4 Col 1 ¶01 discloses outputting a high resolution image) included in the processing system (Lui Pg 3 Col 1 ¶02 and Pg 2 Col 2 ¶03 discloses implementation of the method in general hardware which is well understood to include processing); and
select two local attention operations, in the transformer (Wang Pg 4-5 Section 3.2.1 and discloses the attention operations being proportional to the features including the contextual relationship of the features including the small scale embeddings to the large scale embeddings, when a display resolution is set to less than a maximum spatial resolution (Hu Section 3.1 discloses that the patch tokens input to the subsequent transformer are position-agnostic, and the image processing depends on the spatial information of each pixel) (Lui Pg 4 Col 1 ¶01 discloses outputting a high resolution image) of the transformed version of image pixels (Lui Fig 1 and Fig 2 disclose the transformed version of image pixels by merging image patches in a window partitioning scheme) and the number of contextual objects is two (Wang Pg 4-5 Section 3.2.1 and discloses the attention operations being proportional to the features including the contextual relationship of the features including the small scale embeddings to the large scale embeddings therefore if there are two contextual features there will be two attention operators). See Claim 1 for rationale, its parent claim.
Regarding Claim 9, Lui in view of Rao in further view of Wang in view of Hu teaches the processing system of claim 3, wherein, to select the subset of local attention operations, (Rao Section 3.1 ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply), the one or more processors (Lui Pg 2, Col 1, ¶02 discloses all query patches within a window share the same key set, which facilitates memory access in hardware which is well understood would include processors) are configured to :
obtain a display resolution of a display device (Lui Pg 4 Col 1 ¶01 discloses outputting a high resolution image) included in the processing system (Lui Pg 3 Col 1 ¶02 and Pg 2 Col 2 ¶03 discloses implementation of the method in general hardware which is well understood to include processing); and
select one local attention operation, in the transformer(Wang Pg 4-5 Section 3.2.1 and discloses the attention operations being proportional to the features including the contextual relationship of the features including the small scale embeddings to the large scale embeddings) , when a display resolution is set to less than a maximum spatial resolution (Hu Section 3.1 discloses that the patch tokens input to the subsequent transformer are position-agnostic, and the image processing depends on the spatial information of each pixel) (Lui Pg 4 Col 1 ¶01 discloses outputting a high resolution image) of the transformed version of image pixels (Lui Fig 1 and Fig 2 disclose the transformed version of image pixels by merging image patches in a window partitioning scheme) and the number of contextual objects is one (Wang Pg 4-5 Section 3.2.1 and discloses the attention operations being proportional to the features including the contextual relationship of the features including the small scale embeddings to the large scale embeddings therefore if there is one contextual features there will be one attention operators). See Claim 1 for rationale, its parent claim.
Regarding Claim 10, Lui in view of Rao in further view of Wang in view of Hu teaches the processing system of claim 3, wherein, to select the subset of local attention operations, (Rao Section 3.1 ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply), the one or more processors (Lui Pg 2, Col 1, ¶02 discloses all query patches within a window share the same key set, which facilitates memory access in hardware which is well understood would include processors) are configured to :
obtain a display resolution of a display device (Lui Pg 4 Col 1 ¶01 discloses outputting a high resolution image) included in the processing system (Lui Pg 3 Col 1 ¶02 and Pg 2 Col 2 ¶03 discloses implementation of the method in general hardware which is well understood to include processing); and
select one local attention operation, in the transformer (Wang Pg 4-5 Section 3.2.1 and discloses the attention operations being proportional to the features including the contextual relationship of the features including the small scale embeddings to the large scale embeddings)when a display resolution is set to a smallest spatial resolution (Hu Section 3.1 discloses that the patch tokens input to the subsequent transformer are position-agnostic, and the image processing depends on the spatial information of each pixel) (Lui Pg 4 Col 1 ¶01 discloses outputting a high resolution image) of the transformed version of image pixels (Lui Fig 1 and Fig 2 disclose the transformed version of image pixels by merging image patches in a window partitioning scheme) and the number of contextual objects is one (Wang Pg 4-5 Section 3.2.1 and discloses the attention operations being proportional to the features including the contextual relationship of the features including the small scale embeddings to the large scale embeddings therefore if there is one contextual features there will be one attention operators). See Claim 1 for rationale, its parent claim.
Regarding Claim 11, Lui in view of Rao in further view of Wang in view of Hu teaches the processing system of claim 1, wherein a number of the selected subset of local attention operations (Rao Section 3.1 ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply) is directly proportional (Wang Pg 2 ¶03-04 and ¶07 and discloses a vision transformer where the number of local attention operations scales proportionally with the spatial size of the input feature ) to the spatial resolution(Hu Section 3.1 discloses that the patch tokens input to the subsequent transformer are position-agnostic, and the image processing depends on the spatial information of each pixel) of the transformed version of image pixels (Lui Fig 1 and Fig 2 disclose the transformed version of image pixels by merging image patches in a window partitioning scheme, which in this case is the input). See Claim 1 for rationale, its parent claim.
Regarding Claim 12, Lui in view of Rao in further view of Wang in view of Hu teaches the processing system of claim 11, wherein, wherein, to select the subset of local attention operations, (Rao Section 3.1 ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply), the one or more processors (Lui Pg 2, Col 1, ¶02 discloses all query patches within a window share the same key set, which facilitates memory access in hardware which is well understood would include processors) are configured to select at least two local attention operations (Wang Pg 4-5 Section 3.2.1 and discloses the attention operations being proportional to the features including the contextual relationship of the features including the small scale embeddings to the large scale embeddings) based on a determination that the spatial resolution (Hu Section 3.1 discloses that the patch tokens input to the subsequent transformer are position-agnostic, and the image processing depends on the spatial information of each pixel) satisfies a defined threshold (Wang Pg 2 ¶03-04 and ¶07 and discloses a vision transformer where the number of local attention operations scales proportionally with the spatial size of the input feature ). See Claim 1 for rationale, its parent claim.
Regarding Claim 13, Lui in view of Rao in further view of Wang in view of Hu teaches the processing system of claim 1, wherein the subset of local attention operations is selected (Wang Pg 4-5 Section 3.2.1 and discloses the attention operations being proportional to the features including the contextual relationship of the features including the small scale embeddings to the large scale embeddings) based further on a resolution of a display (Lui Pg 4 Col 1 ¶01 discloses outputting a high resolution image) that will be used to display output of the machine learning model (Lui Pg 3 Col 1 ¶02 and Pg 2 Col 2 ¶03 discloses implementation of the method in general hardware which is well understood to include processing and displaying).See Claim 1 for rationale, its parent claim.
Regarding Claim 14, Lui in view of Rao in further view of Wang in view of Hu teaches the processing system of claim 13, a number of the selected subset of local attention operations (Rao Section 3.1 ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply) is directly proportional (Wang Pg 4-5 Section 3.2.1 and discloses the attention operations being proportional to the features including the contextual relationship of the features including the small scale embeddings to the large scale embeddings) to the resolution (Lui Pg 3 Col 2 ¶01discloses multiple resolution feature maps). See Claim 1 for rationale, its parent claim.
Regarding Claim 15, Lui in view of Rao in further view of Wang in view of Hu teaches the processing system of claim 1, further comprising a camera coupled to the one or more processors (Lui Pg 3 Col 1 ¶02 and Pg 2 Col 2 ¶03 discloses implementation of the method using general hardware which is well understood to include processing and image capture since images are the subject of the processing), wherein the camera is configured to capture image data (Lui Fig 2 disclose image pixels), and wherein the one or more processors (Lui Pg 2, Col 1, ¶02 discloses all query patches within a window share the same key set, which facilitates memory access in hardware which is well understood would include processors) are configured to transform the image data to generate the transformed version of image pixels (Lui Fig 1 and Fig 2 disclose the transformed version of image pixels by merging image patches in a window partitioning scheme). See Claim 1 for rationale, its parent claim.
Regarding Claim 16, Lui in view of Rao in further view of Wang in view of Hu teaches the processing system of claim 1, further comprising a transmitter coupled to the one or more processors (Lui Pg 3 Col 1 ¶02 and Pg 2 Col 2 ¶03 discloses implementation of the method using general hardware which is well understood to include processing and a transmitter), wherein the transmitter is configured to transmit the transformer output (Lui Fig 2, Pg 2 Col 1 ¶1 discloses the output of the transformer being new windows where an additional self attention layer can be applied) to a receiver (Lui Pg 3 Col 1 ¶02 and Pg 2 Col 2 ¶03 discloses implementation of the method using general hardware which is well understood to include processing and a receiver). See Claim 1 for rationale, its parent claim.
Regarding Claim 17, Lui in view of Rao in further view of Wang in view of Hu teaches the processing system of claim 1, wherein the one or more processors (Lui Pg 3 Col 1 ¶02 and Pg 2 Col 2 ¶03 discloses implementation of the method using general hardware which is well understood to include processing including processors) are configured to generate an output prediction of the machine learning model (Rao Section 3.4 discloses prediction models to determine the most influential tokens on the model) based at least in part on the transformer output (Lui Fig 2, Pg 2 Col 1 ¶1 discloses the output of the transformer being new windows where an additional self attention layer can be applied). See Claim 1 for rationale, its parent claim.
Regarding Claim 18, Lui in view of Rao in further view of Wang in view of Hu teaches the processing system of claim 17, further comprising a display coupled to the one or more processors (Lui Pg 3 Col 1 ¶02 and Pg 2 Col 2 ¶03 discloses implementation of the method in general hardware which is well understood to include processing and displaying), wherein the display is configured to display (Lui Pg 3 Col 1 ¶02 and Pg 2 Col 2 ¶03 discloses implementation of the method in general hardware which is well understood to include a display) the output prediction (Rao Section 3.4 discloses prediction models to determine the most influential tokens on the model). See Claim 1 for rationale, its parent claim.
Regarding Claim 19, Lui in view of Rao in further view of Wang in view of Hu teaches the processing system of claim 17, wherein the output prediction (Rao Section 3.4 discloses prediction models to determine the most influential tokens on the model) comprises at least one of: a depth map (Rao Fig 5 discloses the result of scarification of tokens), a classification (Lui Pg 3 Col 1-2 ¶04 discloses image classification), or a segmentation map (Lui Pg 3 Col 1-2 ¶04 discloses segmentation).See Claim 1 for rationale, its parent claim.
Regarding Claim 20, Lui in view of Rao in further view of Wang in view of Hu teaches the processing system of claim 1, wherein, to generate the transformer output (Lui Fig 2, Pg 2 Col 1 ¶1 discloses the output of the transformer being new windows where an additional self attention layer can be applied), the one or more processors (Lui Pg 2, Col 1, ¶02 discloses all query patches within a window share the same key set, which facilitates memory access in hardware which is well understood would include processors) are configured to:
generate a first local attention output (Lui Fig 2, Pg 2 Col 1 ¶1 discloses the output of the transformer being new windows where an additional self attention layer can be applied)based on processing the transformed version of image pixels (Lui Fig 1 and Fig 2 disclose the transformed version of image pixels by merging image patches in a window partitioning scheme) using a first sliced (Wang Section 3.2.1 discloses splitting local attention and global attention) local attention operation (Rao Section 3.1 ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply) at a first scale (Wang Pg 2 ¶02 and Pg discloses local attention and global attention split at each scale);
generate a second local attention output (Lui Fig 2, Pg 2 Col 1 ¶1 discloses the output of the transformer being new windows where an additional self attention layer can be applied) based on the first local attention output (Lui Fig 2, Pg 2 Col 1 ¶1 discloses the output of the transformer being new windows where an additional self attention layer can be applied) and a second (Wang Section 3.2.1 discloses splitting local attention and global attention) sliced local attention operation (Rao Section 3.1 ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply) at a second scale (Wang Pg 2 ¶02 and Pg discloses local attention and global attention split at each scale and the scale depends on the input);
generate a global attention output (Rao Section 3.2 discloses the global features containing the context of the whole image) based on the second local attention output (Lui Fig 2, Pg 2 Col 1 ¶1 discloses the output of the transformer being new windows where an additional self attention layer can be applied) and a global attention operation(Rao ¶01 and Figure 2 discloses a vision transformer that dynamically selects, at runtime, which attention operation to apply Pg 4 Section 3.2 discloses the local and global attention operations being applied) ; and
generate the transformer output (Lui Fig 2, Pg 2 Col 1 ¶1 discloses the output of the transformer being new windows where an additional self attention layer can be applied) based on the first local attention output (Lui Fig 2, Pg 2 Col 1 ¶1 discloses the output of the transformer being new windows where an additional self attention layer can be applied) , the second local attention output(Lui Fig 2, Pg 2 Col 1 ¶1 discloses the output of the transformer being new windows where an additional self attention layer can be applied) , and the global attention output (Rao Section 3.2 discloses the global features containing the context of the whole image). See Claim 1 for rationale, its parent claim.
Reference Cited
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure.
US-20230401716-A1 to Wang discloses a system and method for image segmentation using convolutional self-attention models.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RACHEL ROBERTS whose telephone number is (571)272-6413. The examiner can normally be reached Monday- Friday 7:30am- 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Oneal Mistry can be reached on (313) 446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RACHEL L ROBERTS/Examiner, Art Unit 2674
/ONEAL R MISTRY/Supervisory Patent Examiner, Art Unit 2674