Prosecution Insights
Last updated: August 18, 2026
Application No. 18/815,057

METHOD AND ANALYSIS DEVICE THEREOF FOR ANALYZING DEGREE OF CORRELATION BETWEEN IMAGE GENERATED BY GENERATIVE MODEL AND TEXT PROMPT

Final Rejection §102
Filed
Aug 26, 2024
Priority
Nov 21, 2023 — RE 10-2023-0162432
Examiner
SHAH, ANTIM G
Art Unit
2693
Tech Center
2600 — Communications
Assignee
Electronics and Telecommunications Research Institute
OA Round
2 (Final)
74%
Grant Probability
Favorable
3-4
OA Rounds
1y 3m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
437 granted / 588 resolved
+12.3% vs TC avg
Strong +39% interview lift
Without
With
+39.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
24 currently pending
Career history
604
Total Applications
across all art units

Statute-Specific Performance

§101
7.8%
-32.2% vs TC avg
§103
49.6%
+9.6% vs TC avg
§102
20.2%
-19.8% vs TC avg
§112
13.2%
-26.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 588 resolved cases

Office Action

§102
CTFR 18/815,057 CTFR 84978 DETAILED ACTION Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Response to Amendment Applicants’ amendment filed on 5/27/26 has been entered. Claims 1, 5-6, 10 have been amended. Claims 3-4, 8-9 have been canceled. No new claims have been added. Claims 1-2, 5-7, 10 are still pending in this application, with claims 1 and 6 being independent. Claim Rejections - 35 USC § 102 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-07-aia AIA 07-07 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – 07-08-aia AIA (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. 07-12-aia AIA (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. 07-15 AIA Claim s 1-2, 5-7, 10 are rejected under 35 U.S.C. 102( a)(1 ) as being anticipated by Non-Patent Literature “Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models” (May 31, 2023) to Chefer et al . (“ Chefer ”) . As to claims 1 and 6 , Chefer discloses a method and an analysis device for analyzing a degree of correlation between an image and a text prompt, which are generated by a generative model [ Chefer pages 1-24], the method comprising: receiving the text prompt by an analysis device [page 1:3, Fig. 3: “A lion with a crown”]; inputting the text prompt into the text-to-image generative model by the analysis device [page 1:3, Fig. 3, pages 1:3-1:5: sections 3 and 4]; generating, by the analysis device during a denoising process of the text-to-image generative model, across-attention map for each of text elements constituting the text prompt based on cross attention between conditional latent vectors generated in a process of generating the image and the text elements; distinguishing, by the analysis device from the cross-attention map, a specific area of the image correlated with a corresponding text element [pages 1:3, 1:4: Fig. 3: specific area the image is correlated with Lion and crown ( 𝐴 𝑡 2 , 𝐴 𝑡 5 )]; binarizing, by the analysis device, the cross-attention map such that the specific area appears white on a black background [pages 1:3, 1:4: Fig. 3]; quantifying, by the analysis device, a degree of correlation between the corresponding text element and the specific area based on a number of white pixels or a density of the white pixels in the binarized cross-attention map [pages 1:3, 1:4: Fig. 3: Lion and crown ( 𝐴 𝑡 2 , 𝐴 𝑡 5 ) are elected based on white pixels in the cross-attention map]; and electing, by the analysis device, at least one text element among the text elements significant to the image or an object in the image when the degree of correlation for the at least one text element is greater than or equal to a threshold value [pages: 1:3-1:4: “Text-Conditioning via Cross-Attention”, see section 4 on pages 1:3-1:4 and algorithm 1, “set of thresholds…”, page 1:10, Appendix A.1: “until the specified threshold value is attained”, page 1:11, Appendix B: “the iterative refinement process is not always applied (e.g., when the threshold is already met for all subject tokens]; wherein the conditional latent vectors are generated in a process of generating the image by the text-to-image generative model [Fig. 3, pages 1:3-1:5]. As to claims 2 and 7 , Chefer discloses wherein the text-to-image generative model is a Stable Diffusion model [page 1:3, section 3: “we apply our method over the state-of-the-art Stable Diffusion mode (SD)”]. As to claims 5 and 10 , Chefer discloses wherein the analysis device sets a bounding box for the specific area and annotates the bounding box with the corresponding text element, or distinguishes or segments an area where the at least one text element is positioned in the image on the basis of the attention maps by performing pixel-level segmentation for the specific area based on the cross-attention map [pages 1:3-1:4, Figs. 3-4, Fig. 10] . Response to Arguments 07-37 AIA Applicant's arguments filed on 5/27/26 have been fully considered but they are not persuasive. On pages 8-10 of applicant’s remark, the applicant argues the following: “It is respectfully submitted that there are functional differences between the threshold values of the present invention and those in Chefer and that the thresholds serve different purposes. The limitation in amended independent claims 1 and 6 reciting "selecting, by the analysis device, at least one text element among the text elements significant to the image or an object in the image when the degree of correlation for the at least one text element is greater than or equal to a threshold value" is not disclosed in Chefer.” “In contrast, Chefer discloses a structure in which a relationship between each text token and an image region is analyzed based on an attention map, and latent updating is repeatedly performed until an attention value for a specific token reaches a sufficient level. In this regard, the threshold value used in Chefer is merely used as a criterion for determining whether to terminate the iterative latent refinement process and is not used as a criterion for selecting some of a plurality of text elements or for determining whether a specific text element should be reflected in image generation” “Chefer merely adjusts attention and improves generation results on the premise that all text tokens are used, without performing a process of selecting certain text elements based on their relative importance. On the other hand, the present invention differs in that specific text elements are selected using the degree of correlation and the threshold value and are utilized as a basis for image generation. Accordingly, the present invention and Chefer differ in both configuration and technical effect” Examiner respectfully disagrees with Applicant's arguments for the following reasons: Chefer clearly discloses using set of thresholds (T 1 …T K ) as an input during denoising step using Attend-and-Excite method [page 1:3-1:4, Algorithm 1]. Chefer also clearly discloses that when the threshold is already met for all subject tokens, the iterative refinement process is not always applied [page 1:11, Appendix B: column 2 lines 6-10]. Chefer uses either iterative latent refinement process OR stop when the threshold value is met (similar to selecting text element significant to the image when d for one or more text element is greater than or equal a threshold value). Thus, Chefer clearly discloses that the threshold value is used for selecting some of a plurality of text elements or for determining whether a specific text element should be reflected in image generation. Therefore, Chefer discloses all the limitation of claims 1 and 6 including “selecting, by the analysis device, at least one text element among the text elements significant to the image or an object in the image when the degree of correlation for the at least one text element is greater than or equal to a threshold value”. See prior art rejection for more detail. On pages 10 of applicant’s remark, the applicant argues the following (with respect to claims 5 and 10): “Dependent claims 5 and 10 likewise are not anticipated because Chefer does not disclose setting a bounding box for the specific area and annotating the bounding box with the corresponding text element or performing pixel-level segmentation for the specific area based on the cross-attention map, as now expressly recited. Applicants' support for those operations appears in the specification's discussion of FIGS. 4 and 5” Examiner respectfully disagrees with Applicant's arguments for the following reasons: Claims 5 and 10 now require “wherein the analysis device sets a bounding box for the specific area and annotates the bounding box with the corresponding text element, or distinguishes or segments an area where the at least one text element is positioned in the image on the basis of the attention maps by performing pixel-level segmentation for the specific area based on the cross-attention map”. Chefer discloses performing pixel-level segmentation for the specific area based on the cross-attention map [See Figs. 3-4, pages 1:3 and 1:4 : “maximally-activated patch or “max patch” shown in Fig. 3]. As shown in Fig. 3, pixel-level segmentation is performed at the specific area based on cross-attention map [Fig. 3: “.Given a prompt( e.g. “A lion with a crown”), we extract the subject t tokens (lion, crown), and their corresponding attention maps( 𝐴 2 𝑡 , 𝐴 5 𝑡 )”]. Thus, Chefer clearly discloses all the limitation of claims 5 and 10 including “performing pixel-level segmentation for the specific area based on the cross-attention map. See prior art rejection for more detail”. Conclusion 07-39 AIA THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANTIM G SHAH whose telephone number is (571)270-5214. The examiner can normally be reached Mon-Fri 7:30am-4pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ahmad Matar can be reached at 571-272-7488. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ANTIM G SHAH/Primary Examiner, Art Unit 2693 Application/Control Number: 18/815,057 Page 2 Art Unit: 2693 Application/Control Number: 18/815,057 Page 3 Art Unit: 2693 Application/Control Number: 18/815,057 Page 4 Art Unit: 2693 Application/Control Number: 18/815,057 Page 5 Art Unit: 2693 Application/Control Number: 18/815,057 Page 6 Art Unit: 2693 Application/Control Number: 18/815,057 Page 7 Art Unit: 2693 Application/Control Number: 18/815,057 Page 8 Art Unit: 2693
Read full office action

Prosecution Timeline

Aug 26, 2024
Application Filed
Mar 05, 2026
Non-Final Rejection mailed — §102
May 27, 2026
Response Filed
Jun 03, 2026
Final Rejection mailed — §102 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705437
EVALUATION FRAMEWORK FOR LLM-BASED NETWORK TROUBLESHOOTING AND MONITORING AGENTS
2y 9m to grant Granted Aug 11, 2026
Patent 12707005
DYNAMIC MODIFICATION OF ERROR MAPPING OPERATIONS DURING COMMUNICATION OPERATIONS
2y 7m to grant Granted Aug 11, 2026
Patent 12705231
CHATBOT ASSISTANT POWERED BY ARTIFICIAL INTELLIGENCE FOR TROUBLESHOOTING ISSUES BASED ON HISTORICAL RESOLUTION DATA
2y 6m to grant Granted Aug 11, 2026
Patent 12707207
MODULAR HEARING INSTRUMENT COMPRISING ELECTRO-ACOUSTIC CALIBRATION PARAMETERS
2y 4m to grant Granted Aug 11, 2026
Patent 12707226
EFFICIENT ORIENTATION TRACKING WITH FUTURE ORIENTATION PREDICTION
2y 5m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
74%
Grant Probability
99%
With Interview (+39.0%)
3y 2m (~1y 3m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 588 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month