Prosecution Insights
Last updated: October 01, 2026
Application No. 18/619,667

GENERATING A DIGITAL POSTER INCLUDING MULTIMODAL CONTENT EXTRACTED FROM A SOURCE DOCUMENT

Non-Final OA §103§112
Filed
Mar 28, 2024
Examiner
PARCHER, DANIEL W
Art Unit
Tech Center
Assignee
Adobe Inc.
OA Round
1 (Non-Final)
61%
Grant Probability
Moderate
1-2
OA Rounds
6m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 61% of resolved cases
61%
Career Allowance Rate
170 granted / 278 resolved
+1.2% vs TC avg
Strong +58% interview lift
Without
With
+57.5%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
28 currently pending
Career history
308
Total Applications
across all art units

Statute-Specific Performance

§101
5.3%
-34.7% vs TC avg
§103
58.2%
+18.2% vs TC avg
§102
15.0%
-25.0% vs TC avg
§112
18.3%
-21.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 278 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Allowable Subject Matter Claims 11-12 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 6 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 6 references “the weights of the submodular function” when no such function has been previously introduced. This antecedent ambiguity renders the scope of the claim indefinite. Prior Art Listed herein below are the prior art references relied upon in this Office Action: Kirke (US Patent Number 11,995,120), referred to as Kirke herein. Dolhansky et al., “Deep Submodular Functions: Definitions & Learning” https://proceedings.neurips.cc/paper_files/paper/2016/file/7fea637fd6d02b8f0adf6f7dc36aed93-Paper.pdf, referred to as Dolhansky herein. Modani et al. (US Patent Application Publication 2017/0147544), referred to as Modani herein. Nacson et al., “Stochastic Gradient Descent on Separable Data: Exact Convergence with a Fixed Learning Rate”, referred to as Nacson herein. OH (US Patent Application Publication 2025/0307832), referred to as OH herein. Gil et al. (US Patent Application Publication 2023/0351091), referred to as Gil herein. Examiner’s Note Strikethrough notation in the pending claims has been added by the Examiner. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-2, 4-6, 8-9, 13, 15, and 17-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kirke in view of Dolhansky. Regarding claim 1, Kirke discloses a computer-implemented method comprising: generating, by at least one processor utilizing an encoder neural network (Kirke, Abstract and Fig. 1 with 1:22-35 – hardware memory containing instructions executed by computer processor. 13:55-14:34 – neural network), embedding vectors representing multimodal content of a digital document comprising text and images (Kirke, 29:61-30:11, 31:9-50 – embedding vectors representing input. 4:49-6:4, 29:30-60, 30:12-50, 40:50-41:22 – text and image input. Content data is a written format work that may be used for drafting a paper or literary work for example); determining, by the at least one processor generating, utilizing a large language model, a summary of the multimodal content of the digital document from a prompt based on the content subset (Kirke, 39:26-40:49 – generated summary or review. 27:20-29:29 – LLM, summarization); generating, by the at least one processor and for display at a client device, a digital poster comprising the summary of the multimodal content generated via the large language model (Kirke, 39:26-40:49 – generated poster). However, Kirke appears not to expressly disclose the limitations shown in strikethrough above. However, in the same field of endeavor, Dolhansky discloses document summarization and image summarizations (Dolhansky, Page 7-8), including utilizing a deep submodular function (Dolhansky, Abstract and Pages 3-5). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the machine learning of Kirke to include deep submodular functions based on the teachings of Dolhansky. The motivation for doing so would have been to more effectively capture feature interaction and redundancy (Dolhansky, Page 3), improving summary accuracy. Regarding claim 2, Kirke as modified discloses the elements of claim 1 above and further discloses wherein generating the embedding vectors representing the multimodal content of the digital document comprises: extracting the text and the images of the digital document (Kirke, 4:49-6:4, 29:32-60, 30:12-50, 40:50-41:22 – text and image input. Content data is a written format work that may be used for drafting a paper or literary work for example); determining text segments from the extracted text of the digital document (Kirke, 39:26-40:49 – the user can constrain the data via a parameter such as outreach data. 30:13-50 – attention is limited to a relevant subset of the input data including text and images); and generating, utilizing the encoder neural network, the embedding vectors representing the text segments and the images in a single embedding space (Kirke, 29:61-30:11, 31:9-50 – embedding vectors representing input. Dolhansky, Page 7 – same feature space for cross-modal learning). Regarding claim 4, Kirke as modified discloses the elements of claim 1 above and further discloses wherein determining the content subset comprises determining, utilizing the deep submodular function, one or more embedding vectors that provide diversity of meaning across the content subset by minimizing repetition of meaning across the one or more embedding vectors according to a diversity component of the deep submodular function (Dolhansky, Page 3 - discounted value assigned to redundant features results in less contribution from redundant features. Page 1 – model achieves information diversity). Regarding claim 5, Kirke as modified discloses the elements of claim 1 above and further discloses wherein determining the content subset comprises determining, utilizing the deep submodular function, one or more text segment vectors that align with one or more image vectors according to an alignment component of the deep submodular function (Kirke, 30:12-50 – alignment scores. Dolhansky, Abstract and Pages 3-5 – Deep submodular function). Regarding claim 6, Kirke as modified discloses the elements of claim 1 above and further discloses wherein determining the content subset by utilizing the deep submodular function on the embedding vectors comprises determining at least one of a coverage component, a diversity component, or an alignment component of the deep submodular function by iteratively optimizing a chosen embedding vector subset and weights of the submodular function to maximize the deep submodular function (Kirke, 30:12-31:8 – iterative vector attention weights determination. Dolhansky, Page 7 – deep submodular function selection optimization). Regarding claim 8, Kirke discloses a system comprising: one or more memory devices; and one or more processors configured to cause the system to (Kirke, Abstract and 1:22-35 – hardware memory containing instructions executed by computer processor. 13:55-14:34 – neural network): determine, from embedding vectors representing multimodal content of a digital document can constrain the data via a parameter such as outreach data. 30:13-50 – attention is limited to a relevant subset of the input data including text and images); generate, utilizing a large language model, a summary of the multimodal content of the digital document from a prompt based on the content subset (Kirke, 39:26-40:49 – generated summary or review. 27:20-29:29 – LLM, summarization); and generate, for display at a client device, a digital poster comprising the summary of the multimodal content generated via the large language model by: determining one or more summary elements of the digital poster based on the summary of the multimodal content (Kirke, 39:26-40:49 – generated poster); and determining a layout of the one or more summary elements in the digital poster according to attributes of the one or more summary elements (Kirke, 20:61-21:17, 32:48-33:56, 37:17-39 – layout determination). However, Kirke appears not to expressly disclose the limitations shown in strikethrough above. However, in the same field of endeavor, Dolhansky discloses document summarization and image summarizations (Dolhansky, Page 7-8), including utilizing a deep submodular function (Dolhansky, Abstract and Pages 3-5). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the machine learning of Kirke to include deep submodular functions based on the teachings of Dolhansky. The motivation for doing so would have been to more effectively capture feature interaction and redundancy (Dolhansky, Page 3), improving summary accuracy. Regarding claim 9, Kirke as modified discloses the elements of claim 8 above and further discloses wherein the one or more processors are further configured to determine, utilizing one or more machine learning models, one or more design elements comprising one or more fonts or one or more colors of the digital poster based on the summary of the multimodal content (Kirke, 20:61-21:17, 32:48-33:56, 37:17-39, 38:5-39:25 – font, layout determination). Regarding claim 13, Kirke as modified discloses the elements of claim 8 above and further discloses wherein the one or more processors are further configured to determine the layout of the one or more summary elements by: determining, from the summary of the multimodal content, a number of the one or more summary elements (Kirke, 39:26-40:49 – populating an editable form with outreach data including the generated summary to generate a poster or flyer. The Examiner is including the phrase “a number of” to include one); and determining, based on the number of the one or more summary elements and the attributes of the one or more summary elements, a spatial arrangement of the one or more summary elements (Kirke, 20:61-21:17, 32:48-33:56, 37:17-39 – layout determination). Regarding claim 15, Kirke discloses a non-transitory computer readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising (Kirke, Abstract and 1:22-35 – hardware memory containing instructions executed by computer processor. 13:55-14:34 – neural network): generate, by at least one processor utilizing an encoder neural network, embedding vectors representing multimodal content of a digital document comprising text and images; determine, by the at least one processor outreach data. 30:13-50 – attention is limited to a relevant subset of the input data including text and images); generate, utilizing a large language model, a summary of the multimodal content of the digital document from a prompt based on the content subset (Kirke, 39:26-40:49 – generated summary or review. 27:20-29:29 – LLM, summarization); and generate, for display at a client device, a digital poster comprising the summary of the multimodal content generated via the large language model by (Kirke, 39:26-40:49 – generated poster) determining a layout and a formatting of one or more summary elements of the digital poster based on the summary of the multimodal content (Kirke, 20:61-21:17, 32:48-33:56, 37:17-39 – layout determination. 20:61-21:17, 32:48-33:56, 37:17-39 – font, layout determination). However, Kirke appears not to expressly disclose the limitations shown in strikethrough above. However, in the same field of endeavor, Dolhansky discloses document summarization and image summarizations (Dolhansky, Page 7-8), including utilizing a deep submodular function (Dolhansky, Abstract and Pages 3-5). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the machine learning of Kirke to include deep submodular functions based on the teachings of Dolhansky. The motivation for doing so would have been to more effectively capture feature interaction and redundancy (Dolhansky, Page 3), improving summary accuracy. Regarding claim 17, Kirke as modified discloses the elements of claim 15 above and further discloses wherein determining the content subset comprises determining, utilizing the deep submodular function, one or more embedding vectors that provide diversity across the content subset by minimizing repetition of meaning across the one or more embedding vectors according to a diversity component of the deep submodular function (Dolhansky, Page 3 - discounted value assigned to redundant features results in less contribution from redundant features. Page 1 – model achieves information diversity). Regarding claim 18, Kirke as modified discloses the elements of claim 15 above and further discloses wherein determining the content subset comprises determining, utilizing the deep submodular function, one or more text segment vectors that align with one or more image vectors according to an alignment component of the deep submodular function (Kirke, 30:12-50 – alignment scores. Dolhansky, Abstract and Pages 3-5 – Deep submodular function). Regarding claim 19, Kirke as modified discloses the elements of claim 15 above and further discloses wherein the operations further comprise determining the layout by: determining, from the summary of the multimodal content, a number of the one or more summary elements (Kirke, 39:26-40:49 – populating an editable form with outreach data including the generated summary to generate a poster or flyer. The Examiner is including the phrase “a number of” to include one); and determining, by a layout determination model and based on the number of the one or more summary elements and attributes of the one or more summary elements, a spatial arrangement of the one or more summary elements (Kirke, 20:61-21:17, 32:48-33:56, 37:17-39 – layout determination). Claim(s) 3, 14, and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kirke in view of Dolhansky in further view of Modani. Regarding claim 3, Kirke as modified discloses the elements of claim 1 above and further discloses wherein determining the content subset comprises determining, utilizing the deep submodular function, one or more embedding vectors that collectively summarize the digital document according to a However, Kirke as modified appears not to expressly disclose the limitations shown in strikethrough above. However, in the same field of endeavor, Modani discloses multi-modal document summarization for text and images (Modani, Abstract), including a coverage component (Modani, ¶0035-¶0036, ¶0044-¶0053). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the deep submodular function of Kirke as modified to include a coverage component based on the teachings of Modani. The motivation for doing so would have been to more effectively summarize the document by accounting for a larger portion of the document (Modani, ¶0052). Regarding claim 14, Kirke as modified discloses the elements of claim 8 above and further discloses wherein determining the content subset of the digital document comprises utilizing the deep submodular function to determine one or more embedding vectors that: collectively summarize the digital document provide diversity of the content subset by minimizing repetition of meaning across the one or more embedding vectors according to a diversity component of the deep submodular function (Dolhansky, Page 3 - discounted value assigned to redundant features results in less contribution from redundant features. Page 1 – model achieves information diversity); and align one or more text segment vectors with one or more image vectors according to an alignment component of the deep submodular function (Kirke, 30:12-50 – alignment scores. Dolhansky, Abstract and Pages 3-5 – Deep submodular function). However, Kirke as modified appears not to expressly disclose the limitations shown in strikethrough above. However, in the same field of endeavor, Modani discloses multi-modal document summarization for text and images (Modani, Abstract), including a coverage component (Modani, ¶0035-¶0036, ¶0044-¶0053). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the deep submodular function of Kirke as modified to include a coverage component based on the teachings of Modani. The motivation for doing so would have been to more effectively summarize the document by accounting for a larger portion of the document (Modani, ¶0052). Regarding claim 16, Kirke as modified discloses the elements of claim 15 above and further discloses wherein determining the content subset comprises determining, utilizing the deep submodular function, one or more embedding vectors that collectively summarize the digital document according to a However, Kirke as modified appears not to expressly disclose the limitations shown in strikethrough above. However, in the same field of endeavor, Modani discloses multi-modal document summarization for text and images (Modani, Abstract), including a coverage component (Modani, ¶0035-¶0036, ¶0044-¶0053). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the deep submodular function of Kirke as modified to include a coverage component based on the teachings of Modani. The motivation for doing so would have been to more effectively summarize the document by accounting for a larger portion of the document (Modani, ¶0052). Claim(s) 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kirke in view of Dolhansky in further view of Nacson. Regarding claim 7, Kirke as modified discloses the elements of claim 1 above and further discloses adjusting parameters of the deep submodular function in a framework of a neural network by reducing an output of a loss function utilizing a projected gradient descent algorithm However, Kirke as modified appears not to expressly disclose the limitations in strikethrough above. However, in the same field of endeavor, Dolhansky discloses a machine learning gradient descent algorithm (Nacson, Abstract), including a fixed learning rate (Nacson, Abstract with Pages 1-2). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the deep submodular function of Kirke as modified to include a coverage component based on the teachings of Modani. The motivation for doing so would have been to improve convergence speed (Nacson, Pages 1-2 and 7). Claim(s) 10 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kirke in view of Dolhansky in further view of OH in further view of Gil. Regarding claim 10, Kirke as modified discloses the elements of claim 9 above and further discloses wherein determining, utilizing the one or more machine learning models, the one or more design elements of the digital poster comprises: determining a title of the digital poster However, Kirke as modified appears not to expressly disclose the limitations in strikethrough above. However, in the same field of endeavor, OH discloses AI summarization of text (OH, Abstract with ¶0072-¶0074), including determining a title from the summary (OH, ¶0074). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the title determination of Kirke as modified to include determination from the summary based on the teachings of OH. The motivation for doing so would have been to more include additional document information to improve the accuracy and fit of the determined title. However, Kirke as modified appears not to expressly disclose determination of the font from the title. However, in the same field of endeavor, Gil discloses machine learning document generation (Gil, Abstract with ¶0020-¶0021), including determining font from the document tile (Gil, ¶0059). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the font determination of Kirke as modified to include determination from the title based on the teachings of Gil. The motivation for doing so would have been to more include additional document information to improve the accuracy and fit of the determined font. Regarding claim 20, Kirke as modified discloses the elements of claim 15 above and further discloses wherein the operations further comprise determining, utilizing one or more machine learning models, one or more design elements of the digital poster based on the summary of the multimodal content by: determining a title of the digital poster However, Kirke as modified appears not to expressly disclose the limitations in strikethrough above. However, in the same field of endeavor, OH discloses AI summarization of text (OH, Abstract with ¶0072-¶0074), including determining a title from the summary (OH, ¶0074). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the title determination of Kirke as modified to include determination from the summary based on the teachings of OH. The motivation for doing so would have been to more include additional document information to improve the accuracy and fit of the determined title. However, Kirke as modified appears not to expressly disclose determination of the font from the title. However, in the same field of endeavor, Gil discloses machine learning document generation (Gil, Abstract with ¶0020-¶0021), including determining font from the document tile (Gil, ¶0059). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the font determination of Kirke as modified to include determination from the title based on the teachings of Gil. The motivation for doing so would have been to more include additional document information to improve the accuracy and fit of the determined font. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. References are at least relevant as indicated in the corresponding summary. Dibia (US Patent Application Publication 2024/0185490) – determining color palette in generation of summarized output. Xuan etal. (US Patent Application Publication 2025/005822) – extracting color palettes and dominant colors. Kamarei (US Patent Application Publication 2025/0356447) – machine learning determination of title. Nahum et al. (US Patent Application Publication 2024/0403545) – machine learning determination of title. Manjuanth et al. (US Patent Application Publication 2021/0151038)– machine learning determination of title. Deutch et al. (US Patent Application Publication 2025/0139057) – machine learning determination of title. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIEL W PARCHER whose telephone number is (303)297-4281. The examiner can normally be reached Monday - Friday, 9:00am - 5:00pm, Mountain Time. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, William Bashore can be reached at (571)272-4088 (Eastern Time). The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DANIEL W PARCHER/Primary Examiner, Art Unit 2174
Read full office action

Prosecution Timeline

Mar 28, 2024
Application Filed
Aug 04, 2026
Examiner Interview (Telephonic)
Sep 01, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12675300
SCHEMA DRIVEN USER INTERFACE CREATION TO DEVELOP AUTONOMOUS DRIVING APPLICATIONS
4y 2m to grant Granted Jul 07, 2026
Patent 12656941
METHOD, APPARATUS, DEVICE AND STORAGE MEDIUM FOR DISPLAY MODE SWITCHING
2y 2m to grant Granted Jun 16, 2026
Patent 12632155
EDITING TECHNIQUES FOR INTERACTIVE VIDEOS
4y 11m to grant Granted May 19, 2026
Patent 12632905
COMPUTING SYSTEM FOR CLASSIFYING TAX EFFECTIVE DATE
3y 6m to grant Granted May 19, 2026
Patent 12621534
REFRESHING METHOD AND DISPLAY APPARATUS
2y 8m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
61%
Grant Probability
99%
With Interview (+57.5%)
3y 0m (~6m remaining)
Median Time to Grant
Low
PTA Risk
Based on 278 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month