Prosecution Insights
Last updated: October 01, 2026
Application No. 18/968,277

EXPANDING TOKEN LENGTHS IN TRANSFORMER ENCODERS

Non-Final OA §103
Filed
Dec 04, 2024
Priority
Dec 05, 2023 — provisional 63/606,219 +2 more
Examiner
DUFFY, CAROLINE TABANCAY
Art Unit
Tech Center
Assignee
NEC Laboratories America Inc.
OA Round
1 (Non-Final)
80%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
77 granted / 96 resolved
+20.2% vs TC avg
Strong +18% interview lift
Without
With
+18.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
16 currently pending
Career history
104
Total Applications
across all art units

Statute-Specific Performance

§101
13.6%
-26.4% vs TC avg
§103
58.9%
+18.9% vs TC avg
§102
7.7%
-32.3% vs TC avg
§112
17.5%
-22.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 96 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Applicant’s claim for the benefit of prior-filed application 63/606219 filed 12/05/2023 under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged. Information Disclosure Statement The information disclosure statement (IDS) submitted on 12/04/2024 is being considered by the examiner. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1, 6, 8, 9, 11, 16, 18, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation, published 2022), in view of Zhu et al. (Deformable DETR: Deformable Transformers for end-to-End Object Detection, published 2021). Regarding Claim 1, Li teaches “A computer implemented method for image segmentation, comprising: generating features at a plurality of scales from an input image using a backbone model” (Li, Section 3.3 discloses “Mask DINO adopts the same architecture design for detection as in DINO with minimal modifications.” Li, Section 3.1 discloses “DINO is a typical DETR-like model, which is composed of a backbone, a Transformer encoder, and a Transformer decoder”; where a backbone is a backbone model; where, in Figure 1 “Multi-Scale Features” are features at a plurality of scales; see leftmost input image of Figure 1); “encoding the features using a transformer encoder that creates a per-pixel embedding map from a high-resolution scale of the plurality of scales (Li, Section 3.4 discloses “To perform mask classification, we adopt a key idea from Mask2Former [3] to construct a pixel embedding map which is obtained from the backbone and Transformer encoder features.” Li, Section 3.1 also discloses “Note that DINO uses multi-scale features with deformable attention [40]. Therefore, the updated anchor boxes are also used to constrain deformable attention in a sparse and soft way”; see Figure 1, where 1/4 resolution (high-resolution scale) feature map is used to produce a “Pixel embedding map” in combination with encoder output; where multi-scale features 1/4, 1/8, 1/16, and 1/32 in Figure 1 are progressively higher-resolution scales); “and decoding the features using a transformer decoder to generate a segmentation mask” (Li, Section 3.4 discloses “Then we dot-product each content query embedding qc from the decoder with the pixel embedding map to obtain an output mask m”; see mask in Figure 1; where output mask is a segmentation mask; where decoder layers in Figure 1 is a transformer decoder). PNG media_image1.png 599 1301 media_image1.png Greyscale Figure 1 of Li Li does not explicitly teach “encoding the features using a transformer encoder that creates a per-pixel embedding map from a high-resolution scale of the plurality of scales using deformable attention layers that operate on progressively higher-resolution scales of the plurality of scales.” However, in an analogous field of endeavor, Zhu teaches “encoding the features using a transformer encoder that creates a per-pixel embedding map from a high-resolution scale of the plurality of scales using deformable attention layers that operate on progressively higher-resolution scales of the plurality of scales” (Zhu, Section 1,paragraph 3 discloses “In Deformable DETR , we utilize (multi-scale) deformable attention modules to replace the Transformer attention modules processing feature maps, as shown in Fig. 1”; see Fig. 1). PNG media_image2.png 649 915 media_image2.png Greyscale Figure 1 of Zhu It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Li to incorporate the teachings of Zhu by using deformable self-attention in an encoder. Li references directly the incorporation of Zhu in the decoder layer: Li, Section 3.1 discloses “Note that DINO uses multi-scale features with deformable attention [40]. Therefore, the updated anchor boxes are also used to constrain deformable attention in a sparse and soft way” (reference [40] in Li is Zhu, Deformable detr). Additionally, applying a known technique to a known method ready for improvement yields predictable results. The prior art Li contained a ‘base’ method upon which the claimed invention can be seen as an ‘improvement.’ Li teaches an object segmentation method using a transformer architecture on multi-scale features. The claimed invention also recites using deformable layers in the encoding step, operating on progressively higher-resolution scales. The prior art contained a known technique that is applicable to the base method. Zhu teaches a transformer architecture that uses multi-scale deformable self-attention on progressively higher-resolution scales (see Figure 1 of Zhu). One of ordinary skill in the art would have recognized that applying the known technique would have yielded predictable results and resulted in an improved system by reducing computational and memory complexities (Zhu, Section 1, paragraph 2 discloses “Despite its interesting design and good performance, DETR has its own issues:… the attention weights computation in Transformer encoder is of quadratic computation w.r.t. pixel numbers. Thus, it is of very high computational and memory complexities to process high-resolution feature maps.” Zhu, Section 1, paragraph 4 discloses “Deformable DETR opens up possibilities for us to exploit variants of end-to-end object detectors, thanks to its fast convergence, and computational and memory efficiency”). Additionally, Zhu teaches an embodiment using an “encoder-only Deformable DETR for region proposal generation” (Zhu, Section 4.2). Thus, it would be obvious to one of ordinary skill in the art to apply the deformable self-attention encoder method to Li as each element performs the same function as it does separately. Accordingly, the combination of Li and Zhu discloses the invention of Claim 1. Regarding Claim 6, the combination of Li and Zhu teaches “The method of claim 1, further comprising performing light-pixel embedding on features of the high-resolution scale to generate a per-pixel embedding map” (Li, Section 3.4 discloses “As shown in Fig. 1, the pixel embedding map is obtained by fusing the 1/4 resolution feature map Cb from the backbone with an upsampled 1/8 resolution feature map Ce from the Transformer encoder”; where generating a pixel embedding map using upsampled 1/8 resolution feature map is light-pixel embedding), “and generating the segmentation mask by combining the per-pixel embedding map with an output of the transformer decoder” (Li, Section 3.4 discloses “Then we dot-product each content query embedding qc from the decoder with the pixel embedding map to obtain an output mask m"; see also Figure 1 of Li). Regarding Claim 8, the combination of Li and Zhu teaches “The method of claim 6, wherein combining the per-pixel embedding map with the output of the transformer decoder includes an element-wise multiplication” (Li, Section 3.4 discloses “Then we dot-product each content query embedding qc from the decoder with the pixel embedding map to obtain an output mask m"; where determining the dot-product of each content query embedding with the pixel embedding map is element-wise multiplication). Regarding Claim 9, the combination of Li and Zhu teaches “The method of claim 1, further comprising performing object detection in the input image using the segmentation mask” (Li, Section 3.4 discloses “To perform mask classification, we adopt a key idea from Mask2Former [3] to construct a pixel embedding map which is obtained from the backbone and Transformer encoder features”; where mask classification is object detection; see Figure 1). Regarding Claims 11, 16, 18, and 19, Claims 11, 16, 18, and 19 recite a system with elements corresponding to the steps recited in Claims 1, 6, 8, and 9, respectively. Therefore, the recited elements of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim. Additionally, the rationale and motivation to combine the Li and Zhu references, presented in rejection of Claim 1, apply to this claim. Finally, the combination of Li and Zhu references discloses a “A system for image segmentation, comprising: a hardware processor; and a memory that stores a computer program” (Li, Appendix B.1 discloses “Under the ResNet-50 backbone, we use 4 A100 GPUs each with 40GB memory for all tasks”; where GPU is a hardware processor). Claims 10 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation, published 2022) in view of Zhu et al. (Deformable DETR: Deformable Transformers for end-to-End Object Detection, published 2021), further in view of Raichelgauz et al. (US 2021/0053573 A1). Regarding Claim 10, the combination of Li and Zhu do not explicitly teach the method of Claim 10. However, in an analogous field of endeavor, Raichelgauz teaches “The method of claim 9, further comprising automatically performing a driving action in an autonomous vehicle responsive to the object detection” (Raichelgauz, [0841] discloses “Various examples of responses to object detection and/or responding to an outcome of an object behavior estimation are provided in FIGS. 3A-44. For example—the responding may include at least one out of (i) performing an obstacle avoidance movement in response to the outcome of the object detection and/or the object behavior estimation, (ii) autonomously driving an autonomous vehicle in response to the outcome of the object detection and/or the object behavior estimation”; where performing obstacle avoidance is performing a driving action). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination of Li and Zhu to incorporate the teachings of Raichelgauz by implementing obstacle avoidance movements on an autonomous vehicle in response to the outcome of object detection. One of ordinary skill in the art would be motivated to combine the Li, Zhu, and Raichelgauz references in order to reduce risk in driving environments: Raichelgauz, [0007] discloses “There is a growing need to reduce the risk imposed by dangerous vehicles.” It would have been obvious to one of ordinary skill in the art that the object detection method of Raichelgauz may be simply substituted for the object detection methods of Li and Zhu with predictable results of automatically responding to detected objects in autonomous vehicle applications. Accordingly, the combination of Li, Zhu, and Raichelgauz discloses the invention of Claim 10. Regarding Claim 20, Claim 20 recites a system with elements corresponding to the steps recited in Claim 10. Therefore, the recited elements of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim. Additionally, the rationale and motivation to combine the Li, Zhu, and Raichelgauz references, presented in rejection of Claim 10, apply to this claim. Finally, the combination of Li, Zhu, and Raichelgauz references discloses a “A system for image segmentation, comprising: a hardware processor; and a memory that stores a computer program” (Li, Appendix B.1 discloses “Under the ResNet-50 backbone, we use 4 A100 GPUs each with 40GB memory for all tasks”; where GPU is a hardware processor). Allowable Subject Matter Claims 2-5, 7, 12-15, and 17 are objected to as being dependent upon a rejected base claim but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: Regarding Claim 2, the combination of Li and Zhu does not explicitly teach the method of Claim 2. Li recites tokens and a pixel embedding map (Li, Section 3.5 discloses “The classification score of each token is considered as the confidence to select top-ranked features and feed them to the decoder as content queries.” Li, Section 3.4 discloses “As shown in Fig. 1, the pixel embedding map is obtained by fusing the 1/4 resolution feature map Cb from the backbone with an upsampled 1/8 resolution feature map Ce from the Transformer encoder.”) However, Li and Zhu do not explicitly teach recalibrating tokens on features at lower-resolution scales based on an attention map, nor does the cited prior art explicitly teach an attention map from a higher-resolution scale. That is, although the elements of tokens of features at multiple resolutions are known in the art, as taught by Li, and attention maps are also known in the art, the token recalibration step of Claim 2 is not explicitly taught by the cited prior art. Thus, none of the previously cited prior art references, alone or in combination, provides a motivation to teach the ordered combination of “The method of claim 1, further comprising performing token recalibration on features at lower-resolution scales based on an attention map from a higher-resolution scale to generate recalibrated tokens.” Claims 3-5 depend from Claim 2 and thus contain all allowable subject matter of Claim 2. Regarding Claim 7, the combination of Li and Zhu does not explicitly teach the method of Claim 6. Although Li teaches performing light-pixel embedding (Li, Section 3.4 discloses “As shown in Fig. 1, the pixel embedding map is obtained by fusing the 1/4 resolution feature map Cb from the backbone with an upsampled 1/8 resolution feature map Ce from the Transformer encoder”), Li does not explicitly teach performing light-pixel embedding using a max pooling layer with a particular pooling kernel size. Ren et al. (CN 115049584 A) teaches a maximum pooling layer of kernel size 3 (Ren, Section 3.1.3 discloses “Five maximum pool layers (e.g., Pool1, Pool2, Pool3, Pool4, Pool 5), two full connection layers, and a softmax output layer (softmax in Fig. 3) of Fig. 3. wherein, all the convolution kernel is 3 * 3 * 3, all the pool nucleation is 2 * 2 * 2, stride is 2 * 2 * 2 except the first pool layer, the pool core is 1*2 * 2, stride is 1*2 * 2.”) However, Ren does not teach the max pooling layer in connection with light-pixel embedding as required by Claim 7. Thus, none of the previously cited prior art, alone or in combination provides a motivation to teach the ordered combination of “The method of claim 6, wherein light-pixel embedding includes a max pooling layer with a pooling kernel size of 3.” Regarding Claims 12-15 and 17, claims 12-15 and 17 recite a system with elements corresponding to the method steps of claims 2-5 and 7, respectively, and thus contain all allowable subject matter of Claims 2-5 and 7. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Thawakar et al. (US 2024/0161334 A1) discloses a multi-scale spatio-temporal split attention transformer for use in autonomous vehicle applications. Kidd et al. (US 2025/0005733 A1) discloses a method of analyzing poultry carcasses for defects using a machine learning algorithm including a multi-scale transformer encoder model. Any inquiry concerning this communication or earlier communications from the examiner should be directed to CAROLINE TABANCAY DUFFY whose telephone number is (703)756-1859. The examiner can normally be reached Monday - Friday 8:00 am - 5:30 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached at 5712723382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CAROLINE TABANCAY DUFFY/Examiner, Art Unit 2662 /AMANDEEP SAINI/Supervisory Patent Examiner, Art Unit 2662
Read full office action

Prosecution Timeline

Dec 04, 2024
Application Filed
Sep 03, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749194
DEVICE, SYSTEM, AND METHOD FOR ANALYZING BONE GROWTH STAGE
2y 10m to grant Granted Sep 29, 2026
Patent 12737905
DEPTH DATA MEASUREMENT HEAD, DEPTH DATA COMPUTING DEVICE, AND CORRESPONDING METHOD
2y 11m to grant Granted Sep 15, 2026
Patent 12738051
AUTOMATIC SEGMENTATION OF GEOSPATIAL AIRPORT DATA
3y 0m to grant Granted Sep 15, 2026
Patent 12737895
SYSTEM AND METHODS FOR MONITORING AIRBORNE TARGETS
2y 5m to grant Granted Sep 15, 2026
Patent 12711601
SYSTEM AND METHOD FOR GENERATING TRAINING IMAGE DATA FOR SUPERVISED MACHINE LEARNING, AND NON-TRANSITORY RECORDING MEDIUM
3y 0m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
80%
Grant Probability
98%
With Interview (+18.3%)
2y 10m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 96 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month