Prosecution Insights
Last updated: August 17, 2026
Application No. 18/593,834

SUMMARY PAGE GENERATION USING DOCUMENTS

Final Rejection §103
Filed
Mar 01, 2024
Examiner
MERCADO, GABRIEL S
Art Unit
2171
Tech Center
2100 — Computer Architecture & Software
Assignee
Adobe Inc.
OA Round
2 (Final)
42%
Grant Probability
Moderate
3-4
OA Rounds
1y 0m
Est. Remaining
69%
With Interview

Examiner Intelligence

Grants 42% of resolved cases
42%
Career Allowance Rate
87 granted / 206 resolved
-12.8% vs TC avg
Strong +27% interview lift
Without
With
+26.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
30 currently pending
Career history
249
Total Applications
across all art units

Statute-Specific Performance

§101
15.0%
-25.0% vs TC avg
§103
47.2%
+7.2% vs TC avg
§102
10.3%
-29.7% vs TC avg
§112
24.3%
-15.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 206 resolved cases

Office Action

§103
DETAILED ACTION This office action is responsive to communication(s) filed on 1/28/2026. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Election/Restrictions Applicant has failed to affirm the telephone provisional election by Matthew Rojanakiathavorn on 10/24/2025, which was made with traverse of Group 1, claims 1-14. However, because applicant did not distinctly and specifically point out any supposed errors in the restriction requirement, the provisional election is being treated as a final election without traverse (MPEP § 818.01(a)). Claims Status Claims 1-18 and 21-22 are pending, of which Claims 1-14 and 21-22 are currently being examined. Claims 1, 9 and 15 are independent. Claims 15-18 are withdrawn for being directed to a non-elected invention. Claims 19-20 are newly canceled. Claims 21-22 are newly added. Claim 1, 4, 8-9, and 12 are newly amended. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 5, 7, 8, 9, 13 and 21-22 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gourley; Sean William-Joseph et al. (hereinafter Gourley – US 20190130031 A1) in view of Zhang; Lvmin et al. (hereinafter Zhang, Non-Patent Literature [NPL], “Adding Conditional Control to Text-to-Image Diffusion Models”, Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 3836-3847) and Religa; Tomasz L. et al. (hereinafter Religa – US 20230315969 A1). Independent Claim 1: The rejection of claim 1 is incorporated. Gourley teaches: A method comprising: receiving a text document; (an object, such as text news story, a document, etc., about an event are obtained [receiving] by summary service 110, ¶ 17 and fig. 1), generating, by a document summarizer model, a text summary based on the text document and a structured representation of the text summary, (a text based summary of the event is provided [generating], ¶¶ 22-24. The summary is generated based on characteristics [structured representation of the text summary] of objects that are determined to identify an event(s) that is being summarized, Par 27. This process involves an intermediate, underlying logical step of distilling the original text into its core components [structured representation] before rendering the final output in the desired format.) […]; and generating […] a multimedia summary document corresponding to the text document, (“This summary may include data points and other information derived from the data objects associated with the event, wherein the summary may comprise a text based summary, a graph based summary, an image based summary, or some other type of summary including combinations thereof”, ¶ 36. That is, the summary may include a text based summary combined with an image based summary, that is, a multimedia summary document) wherein the multimedia summary document includes a […] background imagery […], (supplemental sources are used to provide background information for the event, ¶ 20, and this information can include images [background imagery], ¶ 60, as mentioned above, these images can be combined with the text summary) and wherein the multimedia summary document includes at least a portion of the text summary which is placed within the multimedia summary document, (¶ 36. That is, the summary may include a text based summary combined with an image based summary, that is, a multimedia summary document) […]. Gourley does not appear to expressly teach, but Zhang teaches: generating, by a prompt generator, an image generation prompt based on the text summary and the structured representation of the text summary (a diffusion model is used to transform text inputs [based on the text summary] into latent vectors [structured representation of the text summary] to generate state-of-the art images, Page 3838. It was well within the capabilities of a person having ordinary skill in the art to have realized that the textual summary of Gourley may be used as the text input for generating the image based portions of the summary.). that the generating of the multimedia summary document is done using “using a diffusion model and the image generation prompt” (the features are achieved by a diffusion model in a neural network that uses the vector form of the text as input [prompt] for a machine learning model for generating and outputting images, Pages 3836-3838) that the background imagery is a “generated” background imagery “based on the text summary” (a machine learning model output images based on the vector representation of the text, Pages 3837-3838) Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, to modify the method of Gourley to include generating, by a prompt generator, an image generation prompt based on the text summary and the structured representation of the text summary, that the generating of the multimedia summary document is done using “using a diffusion model and the image generation prompt”, that the background imagery is a “generated” background imagery “based on the text summary”, as taught by Zhang. One would have been motivated to make such a combination in order to improve efficiency and quality of the method by generating high-quality images and reducing the computation cost of generating the images, Zhang Page 3836. Gourley further teaches that the summary service receives objects from information sources, including webpages, ¶ 20, And that textual summary is placed/inserted into the mixed output based on the identified event(s), which is/are identified based characteristics of the objects [structured representation], ¶ 27 Gourley-Zhang does not appear to expressly teach, but Religa teaches: the structured representation of the text summary including tags indicating characteristics of text elements included in the structured representation of the text summary; (a document conversion engine can identify section breaks in an original document, identify section heading titles based on the section breaks [characteristics of text elements], generates a summary of each section, and forms a slide based on the titles and summaries, ¶ 27 and fig. 9. Wherein the document can be in HTML format, can contain a combination of text, images, diagrams, and tables, and wherein the tags can be used to detect the section breaks, ¶¶ 18 and 43. Because the process explicitly relies on a detection module to find and utilize HTML heading tags (such as <h1>, <h2>, or <h3>) to identify section breaks and titles, the underlying parsing system must necessarily generate a structured representation that preserves these HTML tags.) and wherein a layout of the portion of the text summary within the multimedia summary document is determined based on the tags indicating the characteristics of the text elements included in the structured representation of the text summary (A layout of the slide is determined by analyzing text characteristics, such as paragraph styles, HTML tags, or font modifications, ¶¶ 43 and 45, to automatically select appropriate templates based on content structure and summarization metrics, ¶ 53) Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, to further modify the method of Gourley to include the structured representation of the text summary including tags indicating characteristics of text elements included in the structured representation of the text summary; and wherein a layout of the portion of the text summary within the multimedia summary document is determined based on the tags indicating the characteristics of the text elements included in the structured representation of the text summary, as taught by Religa. One would have been motivated to make such a combination in order to facilitate a summarization process of documents received from sources in another formats, e.g., HTML format, Religa ¶¶ 3 and 23 and Gourley ¶¶ 20 and 27. Claim 5: The rejection of claim 1 is incorporated. Gourley-Zhang further teaches: wherein generating, using the diffusion model and the image generation prompt, the multimedia summary document corresponding to the text document, further comprises: generating a canvas using the structured representation; (as explained above, Zhang teaches that the diffusion model is used to transform text inputs [based on the text summary] into latent vectors [structured representation of the text summary] to generate state-of-the art images, Page 3838, and Gourley teaches that the summary may include a text based summary combined with an image based summary, that is, a multimedia summary document, as explained above. The use of a diffusion model to transform textual content into latent vectors, which then generates an image based summary combined with the text based summary reflects generating a canvas using a structured representation, e.g., in the background workspace representation of a canvas wherein the combination of the two occurs, because the latent vector serves as a compressed, organized blueprint of the text's meaning, which then systematically guides the image generation process) and generating, by the diffusion model, the multimedia summary document using the canvas and the image generation prompt. (as explained above, the use of a diffusion model to transform textual content into latent vectors, which then generates an image based summary combined with the text based summary reflects generating a canvas using a structured representation, e.g., in the background workspace representation of a canvas wherein the combination of the two occurs, because the latent vector serves as a compressed, organized blueprint of the text's meaning, which then systematically guides the image generation process. This image generation occurs based on image generation prompt, as explained above) Claim 7: The rejection of claim 5 is incorporated. Zhang further teaches: wherein the diffusion model is a ControlNet diffusion model. (Zhang Page 3836) Claim 8: The rejection of claim 1 is incorporated. Religa further teaches: wherein the structured representation is an HTML. (¶ 43) Independent Claim 9: Claim(s) 9 is directed to a medium for accomplishing the steps of the method in claim 1, and is rejected using similar rationale(s). Claim 13: The rejection of claim 9 is incorporated. Claim(s) 13 is directed to a medium for accomplishing the steps of the method in claim 5, and is rejected using similar rationale(s). Claims 21: The rejection of claim 1 is incorporated. Religa further teaches: wherein the characteristics of the text elements included in the structured representation of the text summary include text element types and a reading order indicated by the tags. (Religa teaches text element types and reading order by detailing how metadata [e.g., location marker, location pin, position marker, a place marker, a hidden user interface element, a paragraph identifier, or any other location identifiers], and document control items map out a document's structure, identify section breaks and headings, and establish functional navigation pathways that guide a system or user through the content logically, ¶ 31) Claim 22: The rejection of claim 9 is incorporated. Claim(s) 22 is directed to a medium for accomplishing the steps of the method in claim 21, and is rejected using similar rationale(s). Claim(s) 2 and 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gourley (US 20190130031 A1) in view of Zhang (NPL, “Adding Conditional Control to Text-to-Image Diffusion Models”) and Religa (US 20230315969 A1), as applied to claims 1 and 9 above, and further in view of Amit; Aviel et al. (hereinafter Amit – US 20060103667 A1). Claim 2: The rejection of claim 1 is incorporated. Gourley-Zhang further teaches: wherein generating, using the diffusion model and the image generation prompt, the multimedia summary document corresponding to the text document, further comprises: generating, by the diffusion model, the generated background imagery using the image generation prompt; (efficient generating of images using the diffusion model, as explained above for claim 1) Gourley-Zhang does not appear to expressly teach, but Amit teaches: determining, using a genetic algorithm, a position of the portion of the text summary; (a layout [position] of visual media, ¶ 1, is determined/selected using a selection engine such as a genetic algorithm, ¶ 36, and includes dynamic templates, ¶ 154, for consistency of the design, ¶ 170) determining a font color of the portion of the text summary; (layout templates also include font style and size information, text color information [font color], white space or gutter color information, gutter size information, background color information, ¶ 169) determining a font style of the portion of the text summary; (layout templates also include font style and size information, text color information [font color], white space or gutter color information, gutter size information, background color information, ¶ 169) and determining a font size of the portion of the text summary (layout templates also include font style and size information [font size], text color information [font color], white space or gutter color information, gutter size information, background color information, ¶ 169). Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, to further modify the method of Gourley to include determining, using a genetic algorithm, a position of the portion of the text summary; determining a font color of the portion of the text summary; determining a font style of the portion of the text summary; and determining a font size of the portion of the text summary, as taught by Amit. One would have been motivated to make such a combination in order to improve the flexibility and functionalities of the method by including the ability to consistent, dynamic media layout in an easy manning that preserve the creative concept of the designer, Amit ¶¶ 14, 18 and 170. Claim 10: The rejection of claim 9 is incorporated. Claim(s) 10 is directed to a medium for accomplishing the steps of the method in claim 2, and is rejected using similar rationale(s). Claim(s) 3 and 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gourley (US 20190130031 A1) in view of Zhang (NPL, “Adding Conditional Control to Text-to-Image Diffusion Models”), Religa (US 20230315969 A1) and Amit (US 20060103667 A1), as applied to claims 2 and 10 above, and further in view of Maruo; Akito et al. (hereinafter Maruo – US 20220180210 A1). Claim 3: The rejection of claim 2 is incorporated. Gourley-Zhang-Amit does not appear to expressly teach, but Maruo teaches: wherein the genetic algorithm minimizes an energy function […]. (setting values of parameters in a genetic algorithm to appropriate values to that provides the lowest energy value and represent the result of the optimization, ¶¶ 135 and 267) Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, to further modify the method of Gourley to include wherein the genetic algorithm minimizes an energy function, as taught by Maruo. One would have been motivated to make such a combination in order to improve the efficiency of the method by optimizing selection of the layout, Maruo ¶ 8 and Amit ¶¶ 36, 154 and 170. Gourley-Zhang-Amit-Maruo does not appear to expressly teach, but Official Notice teaches: that the parameters optimized are a visual saliency loss, an alignment loss, an overlap loss, and a reading order loss (the examiner takes Official Notice that each of these parameters of visual saliency, alignment, overlap and reading order are well-known parameters of document design. It was well within the capabilities of a person having ordinary skill in the art to have realized that the optimization discussed above could be used in combination with any known design parameter that a designer desires to optimize). Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, to further modify the method of Gourley to include that the parameters optimized are a visual saliency loss, an alignment loss, an overlap loss, and a reading order loss, as taught by Official Notice. One would have been motivated to make such a combination in order to improve the practicality and flexibility of the method to apply layout optimization/selection based on any known parameter(s) or combinations thereof. Claim 11: The rejection of claim 10 is incorporated. Claim(s) 11 is directed to a medium for accomplishing the steps of the method in claim 3, and is rejected using similar rationale(s). Claim(s) 4 and 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gourley (US 20190130031 A1) in view of Zhang (NPL, “Adding Conditional Control to Text-to-Image Diffusion Models”), Religa (US 20230315969 A1), Amit (US 20060103667 A1) and Maruo (US 20220180210 A1), as applied to claims 3 and 11 above, and further in view of Dontcheva; Lubomira A. et al. (hereinafter Dontcheva – US 20130073952 A1). Claim 4: The rejection of claim 3 is incorporated. Gourley-Zhang-Amit-Maruo does not appear to expressly teach, but Dontcheva teaches: wherein the reading order loss is based on an order of the text elements included in the structured representation. (a method includes applying a reading order algorithm to an arrangement of textual elements in order to provide an optimized path/order of textual elements, Dontcheva Claims 4 and 6) Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, to further modify the method of Gourley to include wherein the reading order loss is based on an order of the text elements included in the structured representation, as taught by Dontcheva. One would have been motivated to make such a combination in order to further improve the usability of the method by improving the readability of the textual elements, Dontcheva Claim 4. Claim 12: The rejection of claim 11 is incorporated. Claim(s) 12 is directed to a medium for accomplishing the steps of the method in claim 4, and is rejected using similar rationale(s). Claim(s) 6 and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gourley (US 20190130031 A1) in view of Zhang (NPL, “Adding Conditional Control to Text-to-Image Diffusion Models”) and Religa (US 20230315969 A1), as applied to claims 1 and 9 above, and further in view of Kneubuehler; Dario et al. (hereinafter Kneubuehler – US 20250053619 A1). Claim 6: The rejection of claim 5 is incorporated. Gourley-Zhang does not appear to expressly teach, but Kneubuehler teaches: wherein the diffusion model is trained using a triplet dataset, (An optimization method, ¶¶ 19 and 144, may include a combine different types of classifiers, such as diffusion and triplet loss based neural network classifiers, ¶ 54. Wherein the triplets are used to train the neural network, ¶ 145. The entire process described is called "triplet loss" and is a supervised method for training models to learn embeddings. A triplet loss is a loss function for machine learning algorithms where a reference input (called anchor) is compared to a matching input (called positive) and a non-matching input (called negative). The distance from the anchor to the positive is minimized, and the distance from the anchor to the negative input is maximized. Discussion of the anchor, positive and negative is reflected in, e.g., ¶¶ 145-146 and 150-151) wherein the triplet dataset includes a training canvas, a training prompt, and a training summary page. (it was well within the capabilities of a person having ordinary skill in the art to have realized that in applying Kneubuehler to Gourley-Zhang, Kneubuehler ¶ 145, an algorithm may be trained based on the anchor images [training canvas], corresponding training prompt [text/keywords] or summary page [text/keywords] that are semantically related [positive] to the canvas [images], and different, unrelated [negative] training prompt [text/keywords] or summary page [text/keywords] that is not associated with the anchor canvas [images].) Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, to further modify the method of Gourley to include wherein the diffusion model is trained using a triplet dataset, wherein the triplet dataset includes a training canvas, a training prompt, and a training summary page, as taught by Kneubuehler. One would have been motivated to make such a combination in order to improve the effectiveness of the method because by training on triplet loss, the neural network learns robust and discriminative features of meshes of objects and are effective at generating embeddings that can be utilized for similarity measurements, Kneubuehler ¶ 152. Claim 14: The rejection of claim 13 is incorporated. Claim(s) 14 is directed to a medium for accomplishing the steps of the method in claim 6, and is rejected using similar rationale(s). Response to Arguments The applicant’s 103 arguments, see Remarks Pg(s) 7-11, are fully considered but are moot in view of the new grounds of rejection presented above. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Below is a list of these references, including why they are pertinent: Sewak; Mohit et al. US 20200097569 A1, is pertinent to claim 1 for disclosing methods and systems for generating cognitive real-time pictorial summary scenes, Abstract. Amid; David et al. US 20140136460 A1, is pertinent to claims 2 and 6 for disclosing a method of a visualizing multi objective designs, such as Pareto optimal solutions, which comply with a plurality of objectives in an objective space in a single presentation, according to some embodiments of the present invention, ¶ 8 and fig. 1. Any inquiry concerning this communication or earlier communications from the examiner should be directed to GABRIEL S MERCADO whose telephone number is (408)918-7537. The examiner can normally be reached Mon-Fri 8am-5pm (Eastern Time). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kieu Vu can be reached at (571) 272-4057. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Gabriel Mercado/Primary Examiner, Art Unit 2171
Read full office action

Prosecution Timeline

Mar 01, 2024
Application Filed
Oct 31, 2025
Non-Final Rejection mailed — §103
Jan 27, 2026
Examiner Interview Summary
Jan 27, 2026
Applicant Interview (Telephonic)
Jan 28, 2026
Response Filed
May 26, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705260
USER-DEFINED GRAPHICAL HIERARCHIES
2y 11m to grant Granted Aug 11, 2026
Patent 12656927
POSITION INPUT TERMINAL WITH Z-POSITION DURATION-BASED AND Z-POSITION RANGE-BASED MODE SWITCHING
2y 9m to grant Granted Jun 16, 2026
Patent 12543983
SYSTEMS AND METHODS FOR EMOTION PREDICTION
3y 1m to grant Granted Feb 10, 2026
Patent 12535942
BLOWOUT PREVENTER SYSTEM WITH DATA PLAYBACK
5y 6m to grant Granted Jan 27, 2026
Patent 12511024
Multi-Application Interaction Method
2y 9m to grant Granted Dec 30, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
42%
Grant Probability
69%
With Interview (+26.6%)
3y 5m (~1y 0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 206 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month