Prosecution Insights
Last updated: August 17, 2026
Application No. 18/678,134

AI-BASED CONTENT TRANSFORMATION INTO DIAGRAMS

Non-Final OA §101§103§112
Filed
May 30, 2024
Examiner
PHAKOUSONH, DARAVANH
Art Unit
Tech Center
Assignee
Microsoft Technology Licensing, LLC
OA Round
1 (Non-Final)
33%
Grant Probability
At Risk
1-2
OA Rounds
1y 5m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants only 33% of cases
33%
Career Allowance Rate
1 granted / 3 resolved
-26.7% vs TC avg
Strong +100% interview lift
Without
With
+100.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 8m
Avg Prosecution
24 currently pending
Career history
39
Total Applications
across all art units

Statute-Specific Performance

§101
48.7%
+8.7% vs TC avg
§103
14.5%
-25.5% vs TC avg
§102
21.4%
-18.6% vs TC avg
§112
13.7%
-26.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 3 resolved cases

Office Action

§101 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 4, 16, and 20 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The claim recites that the first instruction string includes instructions to iteratively extract diagram data and generate the diagram “until the diagram meets a threshold of representing the user intent.” However, it is unclear how the degree to which the generated diagram represents the user intent is determined, what degree of representation constitutes the claimed threshold, or how it is determined that the threshold has been met such that the iterative extraction and diagram-generation process terminates. Although paragraph [0021] and [0041] state that the diagram may be refined until it “meets the expected standards and accurately represents the intended information,” the Specification does not identify the expected standards, define the required degree of representation or accuracy, or disclose any evaluation or comparison used to determine that the diagram has reached the claimed threshold. Accordingly, a person of ordinary skill in the art would not be able to determine with reasonable certainty when the diagram meets the claimed threshold of representing user intent. Claim 6 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 6 recites the limitation "the predetermined prompt"; however, claim 1 does not previously recite or otherwise provide antecedent basis for “the predetermined prompt.” There is insufficient antecedent basis for this limitation in the claim. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more. 101 Subject Matter Eligibility Analysis Step 1: Claims 1-20 are within the four statutory categories (a process, machine, manufacture or composition of matter). Step 2A Prong One, Step 2A Prong Two, and Step 2B Analysis: Step 2A Prong One asks if the claim recites a judicial exception (abstract idea, law of nature, or natural phenomenon). If the claim recites a judicial exception, analysis proceeds to Step 2A Prong Two, which asks if the claim recites additional elements that integrate the abstract idea into a practical application. If the claim does not integrate the judicial exception, analysis proceeds to Step 2B, which asks if the claim amounts to significantly more than the judicial exception. If the claim does not amount to significantly more than the judicial exception, the claim is not eligible subject matter under 35 U.S.C. 101. None of the claims represent an improvement to technology. Claims 1-12 and 17-20 are directed to storage mediums and processors which are machines. Claims 13-16 are directed to a method consisting of a series of steps, meaning that it is directed to the statutory category of process. Regarding claim 1, the following claim elements are abstract ideas: constructing… a first prompt by appending the user prompt and the digital content to a first instruction string, the first instruction string including instructions to a generative model to identify semantic context of the digital content based on metadata of the digital content, to identify at least one of a text data item, an audio data item, a video data item, or a structured file item, embedded in the digital content to generate at least one of a text transcript of the audio data item, a text transcript of the video data item, a text transcript of the structured file item, a text description of the audio data item, a textual description of the video data item, or a text description of the structured file item (This is an abstract idea of a mental process. The limitation involves combining instructions with a user request and content, reviewing metadata to determine the semantic context of the content, identifying different types of embedded information, and describing or transcribing that information. A person could review a user request and digital content, append the request and content to written instructions, examine metadata such as title, source, author, date, or file type to determine context, identify whether the content contains text, audio, video, or structured information, and manually transcribe or describe the identified information using observation and judgement. These steps can be practically performed in the human mind with the aid of pen and paper and therefore falls within the mental process grouping of abstract ideas.), to semantically analyze and extract diagram data from at least one of the text data item, the text transcripts, or the textual descriptions based on the semantic context, and to generate the diagram of the digital content based on the diagram data (This is an abstract idea of a mental process. The limitation involves reviewing textual information in view of its context, identifying and extracting information relevant to a diagram, and organizing that information into a diagram. A person could read the text, transcripts, or descriptions, determine their meaning based on the surrounding context, identify relevant concepts and relationships, record those concepts and relationships as diagram data, and manually draw a diagram representing the digital content. These steps can be practically performed in the human mind with the aid of pen and paper or basic computational tools such as a spreadsheet and therefore falls within the mental process grouping of abstract ideas.); The following claim elements are additional elements which, taken alone or in combination with the other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: processor; and a machine-readable storage medium (This a high-level recitation of generic computer components for performing the abstract idea. See MPEP 2106.05(f).) receiving, via a client device, a user prompt requesting a diagram representing digital content, wherein the digital content includes any of text, audio, video, or structured file (The step of “receiving” the user prompt and digital content is merely generic data transmission operation that has been recognized by the courts as well-understood, routine, and conventional activity. See MPEP 2106.05(d)(II).); providing, via the prompt construction unit, as an input the first prompt to the generative model (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).) receiving as an output the diagram from the generative model (The step of “receiving” the diagram is merely a generic data gathering operation that has been recognized by the courts as well-understood, routine, and conventional activity. See MPEP 2106.05(d)(II)(i).); providing the diagram to the client device to be presented on a user interface of the client device (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).). Regarding claim 2, the rejection of claim 1 is incorporated herein. Further, claim 2 recites the following abstract ideas: to determine a diagram type of the diagram based on at least one of the semantic context, the diagram data, a user intent, or a level of detail (This is an abstract idea of a mental process. The limitation involves evaluating the information and deciding what visual format would be communicate it, such as a flowchart for sequential steps or an organizational chart for hierarchical relationships. This type of selection relies on observation, evaluation, and judgement and can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas. See MPEP 2106.04(a)(2)(III).), and The following claim elements are additional elements which, taken alone or in combination with the other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: wherein the diagram type is a timeline, flowchart, decision tree, mind map, organization chart, fishbone, bar chart, scatter plot, pie chart, histogram, or heat map (This limitation merely specifies the type or format of the resulting diagram and adds insignificant extra-solution activity to the abstract idea.). Regarding claim 3, the rejection of claim 2 is incorporated herein. Further, claim 3 recites the following abstract ideas: to extract the user intent or the level of detail from the user prompt, or to infer the user intent or the level of detail from at least one of the semantic context or the diagram data (This is an abstract idea of a mental process. The limitation involves interpreting a request or surrounding information to determine what the user wants and how much information should be included. This type of information relies on observation, evaluation, and judgement and can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.). Regarding claim 4, the rejection of claim 2 is incorporated herein. Further, claim 4 recites the following abstract ideas: to iteratively extract the diagram data from the at least one of the text data item, the text transcripts, or the textual descriptions based on the semantic context and the user intent (This is an abstract idea of a mental process. The limitation involves repeatedly reviewing textual information, interpreting its meaning and the user’s purpose, and identifying information relevant to the diagram. This iterative analysis relies on observation, evaluation, and judgement and can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.), and to generate the diagram of the digital content based on the diagram data and the user intent, until the diagram meets a threshold of representing the user intent (This is an abstract idea of a mental process. The limitation involves organizing information into a visual representation, evaluating whether the representation sufficiently reflects the intended purpose, and revising it until an acceptable result is reached. This iterative evaluation and revision relies on observation and judgement and can be practically performed in the human mind with the aid of pen and paper or basic computational tools and therefore falls within the mental process grouping of abstract ideas.). Regarding claim 5, the rejection of claim 1 is incorporated herein. Further, claim 5 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: wherein the user prompt is a predetermined prompt selected at the client device for the digital content (This limitation merely specifies the source and manner of selecting the prompt and adds insignificant extra-solution activity to the abstract idea. See MPEP 2106.05(g).). Regarding claim 6, the rejection of claim 1 is incorporated herein. Further, claim 6 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: wherein the predetermined prompt is expending ideas, extracting action items, finding pros and cons, generating a decision making flowchart, generating a SWOT analysis, or summarizing ideas (This limitation merely specifies the subject matter or purpose of the predetermined prompt and adds insignificant extra-solution activity to the abstract idea.). Regarding claim 7, the rejection of claim 1 is incorporated herein. Further, claim 7 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: receiving at least one user feedback on the diagram via the user interface of the client device (The step of “receiving” user feedback through the user interface is merely a generic data transmission operation that has been recognized as well-understood, routine, and conventional activity.). Regarding claim 8, the rejection of claim 7 is incorporated herein. Further, claim 8 recites the following abstract ideas: constructing… a second prompt by appending the feedback and the diagram to a second instruction string, the second instruction string including instructions to the generative model to generate at least another diagram based on the feedback and the diagram, by adjusting one or more attributes of the diagram based on the feedback (This is an abstract idea of a mental process. The limitation involves combining existing information with feedback, interpreting the requested changes, and revising the visual representation by adjusting its attributes. The type of review and revision relies on observation, evaluation, and judgement and can be practically performed in the human mind with the aid of pen and paper or basic computational tool such as a spreadsheet and therefore falls within the mental process grouping of abstract ideas.); providing, via the prompt construction unit, as an input the second prompt to the generative model (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).) receiving as an output the other diagram of the digital content from the generative model (The step of “receiving” the diagram is merely a generic data gathering operation that has been recognized by the courts as well-understood, routine, and conventional activity. See MPEP 2106.05(d)(II)(i).); providing the other diagram to the client device to be presented on the user interface of the client device (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).). Regarding claim 9, the rejection of claim 7 is incorporated herein. Further, claim 9 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: wherein the user feedback is collected via a user selection of at least one of a thumbs-up tab, a thumbs-down tab, a neutral tab, or a generating-more-image tab, a textual input, or a combination thereof (This limitation merely specifies the manner in which the feedback is selected or entered and adds insignificant extra-solution activity to the abstract idea.). Regarding claim 10, the rejection of claim 1 is incorporated herein. Further, claim 10 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: causing the user interface to receive a user confirmation of the diagram; and causing a publication of the diagram (This limitation merely specify receiving approval and publishing the resulting diagram and add insignificant extra-solution activity to the abstract idea.). Regarding claim 11, the rejection of claim 1 is incorporated herein. Further, claim 11 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: wherein the generative model is a language model or a multimodal model (This limitation merely identifies the generative model as a language model or a multimodal model, which is a high-level recitation of a generic computer component for performing the abstract idea. See MPEP 2106.05.). Regarding claim 12, the rejection of claim 1 is incorporated herein. Further, claim 12 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: wherein the digital content and the user prompt are received via a software application, and wherein the software application is a virtual meeting and collaboration application, a digital whiteboard application, an employee experience application, an online collaboration application, a calendar application, an email application, a task management application, a team-work planning application, a software development application, an enterprise accounting and sales application, a social media application, or an online encyclopedia (This limitation merely specifies the type of software application through which the digital content and user prompt are received and adds insignificant extra-solution activity to the abstract idea.). Regarding claim 13, the following claim elements are abstract ideas: constructing…a first prompt by appending the user prompt and the digital content to a first instruction string, the first instruction string including instructions to a generative model to identify semantic context of the digital content based on metadata of the digital content, to identify at least one of a text data item, an audio data item, a video data item, or a structured file item, embedded in the digital content to generate at least one of a text transcript of the audio data item, a text transcript of the video data item, a text transcript of the structured file item, a text description of the audio data item, a textual description of the video data item, or a text description of the structured file item (This is an abstract idea of a mental process. The limitation involves combining instructions with a user request and content, reviewing metadata to determine the semantic context of the content, identifying different types of embedded information, and describing or transcribing that information. A person could review a user request and digital content, append the request and content to written instructions, examine metadata such as title, source, author, date, or file type to determine context, identify whether the content contains text, audio, video, or structured information, and manually transcribe or describe the identified information using observation and judgement. These steps can be practically performed in the human mind with the aid of pen and paper and therefore falls within the mental process grouping of abstract ideas.), to semantically analyze and extract diagram data from at least one of the text data item, the text transcripts, or the textual descriptions based on the semantic context, and to generate the diagram of the digital content based on the diagram data (This is an abstract idea of a mental process. The limitation involves reviewing textual information in view of its context, identifying and extracting information relevant to a diagram, and organizing that information into a diagram. A person could read the text, transcripts, or descriptions, determine their meaning based on the surrounding context, identify relevant concepts and relationships, record those concepts and relationships as diagram data, and manually draw a diagram representing the digital content. These steps can be practically performed in the human mind with the aid of pen and paper or basic computational tools such as a spreadsheet and therefore falls within the mental process grouping of abstract ideas.); The following claim elements are additional elements which, taken alone or in combination with the other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: receiving, via a client device, a user prompt requesting a diagram representing digital content, wherein the digital content includes any of text, audio, video, or structured file (The step of “receiving” the user prompt and digital content is merely generic data transmission operation that has been recognized by the courts as well-understood, routine, and conventional activity. See MPEP 2106.05(d)(II).); providing, via the prompt construction unit, as an input the first prompt to the generative model (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).) receiving as an output the diagram from the generative model (The step of “receiving” the diagram is merely a generic data gathering operation that has been recognized by the courts as well-understood, routine, and conventional activity. See MPEP 2106.05(d)(II)(i).); providing the diagram to the client device to be presented on a user interface of the client device (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).). Regarding claim 14, the rejection of claim 13 is incorporated herein. The claim recites similar limitations corresponding to claim 2. Therefore, the same subject matter analysis that was utilized for claim 2, as described above, is equally applicable to claim 14. Therefore, claim 14 is ineligible. Regarding claim 15, the rejection of claim 14 is incorporated herein. The claim recites similar limitations corresponding to claim 3. Therefore, the same subject matter analysis that was utilized for claim 3, as described above, is equally applicable to claim 15. Therefore, claim 15 is ineligible. Regarding claim 16, the rejection of claim 14 is incorporated herein. The claim recites similar limitations corresponding to claim 4. Therefore, the same subject matter analysis that was utilized for claim 4, as described above, is equally applicable to claim 16. Therefore, claim 16 is ineligible. Regarding claim 17, the following claim elements are abstract ideas: constructing…a first prompt by appending the user prompt and the digital content to a first instruction string, the first instruction string including instructions to a generative model to identify semantic context of the digital content based on metadata of the digital content, to identify at least one of a text data item, an audio data item, a video data item, or a structured file item, embedded in the digital content to generate at least one of a text transcript of the audio data item, a text transcript of the video data item, a text transcript of the structured file item, a text description of the audio data item, a textual description of the video data item, or a text description of the structured file item (This is an abstract idea of a mental process. The limitation involves combining instructions with a user request and content, reviewing metadata to determine the semantic context of the content, identifying different types of embedded information, and describing or transcribing that information. A person could review a user request and digital content, append the request and content to written instructions, examine metadata such as title, source, author, date, or file type to determine context, identify whether the content contains text, audio, video, or structured information, and manually transcribe or describe the identified information using observation and judgement. These steps can be practically performed in the human mind with the aid of pen and paper and therefore falls within the mental process grouping of abstract ideas.), to semantically analyze and extract diagram data from at least one of the text data item, the text transcripts, or the textual descriptions based on the semantic context, and to generate the diagram of the digital content based on the diagram data (This is an abstract idea of a mental process. The limitation involves reviewing textual information in view of its context, identifying and extracting information relevant to a diagram, and organizing that information into a diagram. A person could read the text, transcripts, or descriptions, determine their meaning based on the surrounding context, identify relevant concepts and relationships, record those concepts and relationships as diagram data, and manually draw a diagram representing the digital content. These steps can be practically performed in the human mind with the aid of pen and paper or basic computational tools such as a spreadsheet and therefore falls within the mental process grouping of abstract ideas.); The following claim elements are additional elements which, taken alone or in combination with the other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: A non-transitory computer readable medium (This a high-level recitation of generic computer components for performing the abstract idea. See MPEP 2106.05(f).) receiving, via a client device, a user prompt requesting a diagram representing digital content, wherein the digital content includes any of text, audio, video, or structured file (The step of “receiving” the user prompt and digital content is merely generic data transmission operation that has been recognized by the courts as well-understood, routine, and conventional activity. See MPEP 2106.05(d)(II).); providing, via the prompt construction unit, as an input the first prompt to the generative model (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).) receiving as an output the diagram from the generative model (The step of “receiving” the diagram is merely a generic data gathering operation that has been recognized by the courts as well-understood, routine, and conventional activity. See MPEP 2106.05(d)(II)(i).); providing the diagram to the client device to be presented on a user interface of the client device (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).). Regarding claim 18, the rejection of claim 17 is incorporated herein. The claim recites similar limitations corresponding to claim 2. Therefore, the same subject matter analysis that was utilized for claim 2, as described above, is equally applicable to claim 18. Therefore, claim 18 is ineligible. Regarding claim 19, the rejection of claim 18 is incorporated herein. The claim recites similar limitations corresponding to claim 3. Therefore, the same subject matter analysis that was utilized for claim 3, as described above, is equally applicable to claim 19. Therefore, claim 19 is ineligible. Regarding claim 20, the rejection of claim 18 is incorporated herein. The claim recites similar limitations corresponding to claim 4. Therefore, the same subject matter analysis that was utilized for claim 4, as described above, is equally applicable to claim 20. Therefore, claim 20 is ineligible. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-9, and 11, 13-20 are rejected under the 35 U.S.C. 103 as being unpatentable over Jiang et al., (NPL: “Graphologue: Exploring Large Language Model Responses with Interactive Diagrams” (Published: 2023)) In view of Orozco et al., (Pub. No.: US 20240394945 A1 (Filed: May 22, 2024)). Regarding claim 1, Jiang teaches the following limitations: receiving, via a client device, a user prompt requesting a diagram representing digital content, wherein the digital content includes any of text, audio, video, or structured file (Jiang, [Abstract] “Large language models (LLMs) have recently soared in popularity due to their ease of access and the unprecedented ability to synthesize text responses to diverse user questions… We present Graphologue, an interactive system that converts text-based responses from LLMs into graphical diagrams to facilitate information-seeking and question-answering tasks. Graphologue employs novel prompting strategies and interface designs to extract entities and relationships from LLM responses and constructs node-link diagrams in real-time.” [Fig. 1] “Graphologue constructs an interactive diagram in real-time as GPT-4 text responses are streamed in.” – Jiang teaches an interactive system that receives user questions through an interface and converts the resulting text-based LLM responses into graphical node-link diagrams. Under the broadest reasonable interpretation, the user question corresponds to the claimed user prompt, and the LLM text response corresponds to digital content including text that is represented by the generated diagram. Accordingly, Jiang teaches receiving, via a client device, a user prompt for a diagram representing text content.); providing, via the prompt construction unit, as an input the first prompt to the generative model and receiving as an output the diagram from the generative model (Jiang, [section 6] “We develop and test the following prompting strategies with OpenAI’s GPT-4, the most advanced and publicly available LLM to date.” [section 6.1] “We outline our key prompt components, which work together to instruct GPT-4 to generate an initial response that facilitates dynamic diagram construction (D1) and enables easy control over the complexity of presented information (D2) through interactive diagrams (Appendix A.1).” [Figure 2, page 6] “As GPT-4 responses stream in, Graphologue parses them in real-time, removes inline annotations for the interface, extracts entities and relationships, and constructs the corresponding diagrams.” – Jiang teaches providing the constructed prompt components as input to GPT-4, which corresponds to the claimed generative model, to generate an annotated response facilitating diagram construction. Jiang further teaches receiving the streamed GPT-4 response, extracting entities and relationships, and constructing the corresponding rendered diagram. Under BRI, the Graphologue prompt corresponds to the first prompt provided as input to the generative model, and the corresponding rendered diagram constitutes the diagram output.); and providing the diagram to the client device to be presented on a user interface of the client device (Jiang, [section 4, D3] “By utilizing diagrams as the main interface with LLMs, typical information tasks should be supported through interaction with the nodes and links in these diagrams” [section 5] “In Graphologue, she started by typing ‘What is an earthquake?’. As response text streamed into the interface from the LLM, she noticed a node-link diagram was being constructed piece by piece on the side simultaneously” [Fig. 2, page 6] “As GPT-4 responses stream in, Graphologue parses them in real-time, removes inline annotations for the interface, extracts entities and relationships, and constructs the corresponding diagrams.” – Jiang teaches presenting the constructed node-link diagram through Graphologue’s user interface as the LLM response streams into the interface. Under BRI, the device executing and displaying the Graphologue interface corresponds to the client device, and displaying the diagram within that interface corresponds to providing the diagram to be presented on the user interface of the client device.). However, Jiang does not teach but Jiang in view of Orozco teaches the following limitations: A data processing system comprising: a processor; and a machine-readable storage medium storing executable instructions that, when executed, cause the processor alone or in combination with other processors to perform operations of (Orozco, paragraph [0151] “computing device 515 can include one or more processors 513 and memory 514. The memory 514 can store information accessible by the processor(s) 513, including instructions 521 that can be executed by the processor(s) 513. The memory 514 can also include data 523 that can be retrieved, manipulated, or stored by the processor(s) 513. The memory 514 can be a type of non-transitory computer readable medium capable of storing information accessible by the processor(s) 513, such as volatile and non-volatile memory. The processor(s) 513 can include one or more central processing units (CPUs), graphic processing units (GPUs), field-programmable gate arrays (FPGAs), and/or application-specific integrated circuits (ASICs), such as tensor processing units (TPUs).”). constructing, via a prompt construction unit, a first prompt by appending the user prompt and the digital content to a first instruction string, the first instruction string including instructions to a generative model to identify semantic context of the digital content based on metadata of the digital content, to identify at least one of a text data item, an audio data item, a video data item, or a structured file item, embedded in the digital content to generate at least one of a text transcript of the audio data item, a text transcript of the video data item, a text transcript of the structured file item, a text description of the audio data item, a textual description of the video data item, or a text description of the structured file item (Jiang, [section 6] “To enable the diagrams to be constructed simultaneously as the response streams in, we iteratively develop our prompts to have the LLM annotate the entities and relationships inline with the tokens… We develop and test the following prompting strategies with OpenAI’s GPT-4, the most advanced and publicly available LLM to date. A full list of original prompts can be found in Appendix A.” [section 6.1] “We outline our key prompt components, which work together to instruct GPT-4 to generate an initial response that facilitates dynamic diagram construction (D1) and enables easy control over the complexity of presented information (D2) through interactive diagrams (Appendix A.1).” [section 10.2] “Moreover, a text input box can be provided to allow customized requests. As the user explores the knowledge space through the graphical interface, their prior actions and the current diagram can be leveraged as context to construct prompts for LLMs to get responses that are better aligned with the user’s needs.” [Appendix A.1, page 17] “System Please provide a well-structured response to the user’s question in multiple paragraphs… The user’s goal is to construct a concept map to visually explain your response. To achieve this, annotate the key entities and relationships inline for each sentence in the paragraphs.” Orozco, paragraph [0034] “ As an example, the modifications may be related to an upcoming event, and the modification inputs 130 can include the name of the event, a description of the event, keywords related to the event, and so on. The modification inputs 130 can include natural language, tags, titles, etc.” [0036] “The modification inputs 130 can include data of different modalities, such as, images, video, computer drawings, audio, text, and so on.” [0037] “For example, the seed generation engine 195 can generate, from the modification inputs 130, a natural language prompt describing elements of each of the modification inputs 130. The prompt can include the textual information included as part of the inputs 130, as well as textual descriptions of inputs of other modalities provided, e.g., text descriptions of images, video, or audio transcripts.” [0049] “ An embedding at least partially encodes some semantic meaning for the text, image, video, etc., represented by the embedding.” [0080] “ The generative model 310 provides, as output, the generated features associated with the existing digital content. The generated features may be, for example embeddings associated with the existing digital content and/or a natural language description of the existing digital content.” – Jiang teaches developing and constructing prompts using customized user requests and contextual digital content and supplying GPT-4 with predefined textual instructions. Under BRI, Jiang’s System instruction block corresponds to the first instruction string. Orozco teaches generating a prompt from multimodal digital content that includes textual information, textual descriptions, and transcripts. Orozco further teaches keywords, tags, and titles associated with the content and embeddings that encode semantic meaning. Under BRI, the keywords, tags, and titles correspond to metadata, and the encoded semantic meaning corresponds to semantic context. Thus, Jiang in view of Orozco teaches the claimed prompt construction, context identification, and content-item identification, and generation of a corresponding transcript or text description.), to semantically analyze and extract diagram data from at least one of the text data item, the text transcripts, or the textual descriptions based on the semantic context, and to generate the diagram of the digital content based on the diagram data (Jiang, [Abstract] “Graphologue employs novel prompting strategies and interface designs to extract entities and relationships from LLM responses and constructs node-link diagrams in real-time.” [section 6.1.2] “GPT-4 is instructed to annotate entities in the text to serve as nodes in the diagrams.” [section 6.1.3] “In addition to entities, we prompt GPT-4 to identify and annotate relationships between these entities inline (Figure 2). These relationships serve as the links in the node-link diagram…. GPT-4 is instructed to include them in a single annotation, which we utilize to organize the connections together when rendering them on the canvas” [Appendix, A.1, page 17] “You should try to annotate at least one relationship for each entity. Relationships should only connect entities that appear in the response. You can arrange the sentences in a way that facilitates the annotation of entities and relationships, but the arrangement should not alter their meaning, and they should still flow naturally in language.” [section 9.4.3] “They selected the node ‘particles’ in the diagram and clicked the ‘Explain’ button. This action extended the paragraph and constructed a new part of the diagram stemming from the node, serving to clarify the context-specific meaning of ‘particles.’” Orozco, paragraph [0037] “The prompt can include the textual information included as part of the inputs 130, as well as textual descriptions of inputs of other modalities provided, e.g., text descriptions of images, video, or audio transcripts.” [0049] “An embedding at least partially encodes some semantic meaning for the text, image, video, etc., represented by the embedding.” – Jiang teaches analyzing the meaning of text to identify entities and relationships and clarifying the context-specific meaning of terms. Under BRI, this corresponds to semantically analyzing the text based on semantic context. Jiang further represents the identified entities and relationships as nodes, links, and annotations, which correspond to diagram data, and renders the node-link diagram from the data. Orozco teaches that the analyzed input may include textual content, transcripts, or textual descriptions derived from multimodal digital content and the embeddings encode the semantic meaning of that content.); Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, having a combination of Jiang and Orozco before them, to incorporate Orozco’s processing of multimodal digital content, including the use of keywords, tags, and titles to identify semantic context and the generation of textual descriptions or transcripts of non-text content, into the prompt-based diagram generation system of Jiang. One would have been motivated to make such a combination in order to allow Jiang’s system to generate diagrams from digital content provided in different modalities rather than being limited to existing text responses. This would predictably enable Jiang’s generative model to semantically analyze textual descriptions or transcripts representing audio, video, or other digital content and extract the entities and relationships used to construct the corresponding diagram. Regarding claim 2, Jiang in view of Orozco teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. Jiang in view of Orozco further teaches: wherein the first instruction string further includes instructions to determine a diagram type of the diagram based on at least one of the semantic context, the diagram data, a user intent, or a level of detail, and wherein the diagram type is a timeline, flowchart, decision tree, mind map, organization chart, fishbone, bar chart, scatter plot, pie chart, histogram, or heat map (Jiang, [section 2.1] “Natural language interfaces offer the key benefit of enabling users to directly articulate their intended actions and goals without learning and utilizing complex manual user interfaces.” [section 9.4.1] “Participants found that the node-link diagrams enhanced their understanding of the diverse relationships inherent in the topic they explored (Figure 9.1). For instance, P5 stated that diagrams helped “visualize the connections,” and P4 suggested they aid in comprehending how different aspects connect. Participants particularly praised the diagrams for providing an “overall view of a topic” and for “understanding a set of instructions” (P8), likening it to a “mind map” (P8), and stating that it provides an understanding of the organization of information (P5)… The bidirectional mapping between the diagram and paragraph through highlighting “is a good visual cue to get the users’ attention to understand the various relationships” (P5) and helps users locate and comprehend various terms and concepts from the diagram with explanations easily” [section 9.4.2] “Most participants found the amount of information presented in the responses to be concise (Figure 9.3), and they appreciated the ability to control the level of detail they wished to see with Graphologue (Figure 9.4)” – Jiang teaches receiving natural-language input that expresses the user’s intended actions and goals and generating a node-link diagram that represents concepts, relationships, connections, and organization of the underlying information. Jiang describes the resulting diagram as being like a mind map and further permits control over the level of detail presented. Under BRI, the expressed actions and goals correspond to user intent, the meaning conveyed by the represented concepts and relationships corresponds to semantic context, and the node-link representation corresponds to the claimed mind map type. Thus, Jiang teaches determining a mind map diagram type based on user intent, semantic context or level of detail.). Regarding claim 3, Jiang in view of Orozco teaches all the elements of claim 2, therefore is rejected for the same reasons as those presented for claim 2. Jiang in view of Orozco further teaches: wherein the first instruction string further includes instructions to extract the user intent or the level of detail from the user prompt, or to infer the user intent or the level of detail from at least one of the semantic context or the diagram data (Jiang, [section 2.1] “Natural language interfaces offer the key benefit of enabling users to directly articulate their intended actions and goals without learning and utilizing complex manual user interfaces… Another approach to natural language interfaces has been extracting user intents from their natural expressions to enable less rigid communication between humans and computers by leveraging advanced natural language understanding and domain-specific knowledge. Iris, for example, enables users to describe data analysis goals and disambiguate system interpretations using natural expressions [32]. CrossData infers the desired data values and operations from text to report in the data insights without instructing the system [22]. Crosspower employs a human-in-the-loop approach by enabling users to interact with linguistic structures in a video script to convey high-level design goals regarding the graphical content and structures… Many tasks require users to go through arduous and time-consuming prompt engineering to produce well-crafted prompts, thereby ensuring results that align with their intents” – Jiang teaches that natural language expressions directly communicate a user’s intended actions and goals and expressly teaches extracting user intents from such natural expressions. Jiang further teaches inferring desired data values, operations, and high-level graphical design goals from textual input. Under BRI, the natural language expression or textual input corresponds to the user prompt, and the intended actions, goals, desired values, and desired operations correspond to user intent. Thus, Jiang teaches extracting or inferring user intent from the user prompt.). Regarding claim 4, Jiang in view of Orozco teaches all the elements of claim 2, therefore is rejected for the same reasons as those presented for claim 2. Jiang in view of Orozco further teaches: wherein the first instruction string further includes instructions to iteratively extract the diagram data from the at least one of the text data item, the text transcripts, or the textual descriptions based on the semantic context and the user intent, and to generate the diagram of the digital content based on the diagram data and the user intent, until the diagram meets a threshold of representing the user intent (Jiang, [section 4, D2] “To avoid overwhelming the users with complexity, the users should be able to flexibly control the amount of information to be visualized in the diagram and how the available information should be revealed.” [section 4, D3] “Users should be able to interact with the diagram to acquire more information, such as further exploring an unfamiliar concept by requesting more explanations or examples. Similarly, users should be able to collapse or trim parts of the diagrams if they are irrelevant to their goal. Explorations beyond the initial prompt and response should be organized through expanding and trimming of the diagrams.” [section 6] “To enable the diagrams to be constructed simultaneously as the response streams in, we iteratively develop our prompts to have the LLM annotate the entities and relationships inline with the tokens. This enables Graphologue to provide both the text responses and the diagrams at the same time.” [section 6.1.2] “GPT-4 is instructed to annotate entities in the text to serve as nodes in the diagrams.” [section 6.1.3] “In addition to entities, we prompt GPT-4 to identify and annotate relationships between these entities inline (Figure 2). These relationships serve as the links in the node-link diagram.” [Fig. 2, page 6] “As GPT-4 responses stream in, Graphologue parses them in real-time, removes inline annotations for the interface, extracts entities and relationships, and constructs the corresponding diagrams.” – Jiang teaches that GPT-4 annotates entities and relationships inline as the textual response is generated and that Graphologue parse the streamed text to repeatedly extract those entities and relationships. Under BRI, the streamed LLM response corresponds to the claimed text data item, and the extracted entities and relationships correspond to the claimed diagram data, since they are used as the nodes and links of the diagram. Jiang therefore teaches iteratively extracting diagram data from text and using the extracted diagram data to construct the corresponding diagram. Jiang further teaches that the user may iteratively expand the diagram to obtain additional explanations or examples and trim portions that are irrelevant to the user’s goal, while controlling the amount of information represented. Under BRI, the user’s goal and requests for additional or reduced information correspond to user intent, and continuing the interactive extraction and diagram generation until the diagram contains the contextual information and level of detail sought by the user corresponds to generating the diagram until it meets a qualitative threshold of representing the user intent.). Regarding claim 5, Jiang in view of Orozco teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. Jiang in view of Orozco further teaches: wherein the user prompt is a predetermined prompt selected at the client device for the digital content (Jiang, [section 9.4.3] “Graphologue made it particularly easy to construct prompts for creating examples and explanations during learning activities, which can often be monotonous and time-consuming… However, with Graphologue, acquiring context-specific examples is straightforward: “clicking the ‘Examples’ button just does it, and (you) don’t have to think of another prompt” (P5)… They selected the node ‘particles’ in the diagram and clicked the ‘Explain’ button. This action extended the paragraph and constructed a new part of the diagram stemming from the node, serving to clarify the context-specific meaning of ‘particles.’” – Jiang teaches predefined “Examples” and “Explain” controls associated with preconstructed prompts for selected digital content. The user selects one of the controls through the Graphologue user interface, and the corresponding prompt is invoked for the selected diagram content without requiring the user to formulate another prompt. Under BRI, the preconstructed prompt associated with the selected control corresponds to a predetermined user prompt selected at the client device for the digital content.) Regarding claim 6, Jiang in view of Orozco teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. Jiang in view of Orozco further teaches: wherein the predetermined prompt is expending ideas, extracting action items, finding pros and cons, generating a decision making flowchart, generating a SWOT analysis, or summarizing ideas (Jiang, [section 9.4.3] “Graphologue made it particularly easy to construct prompts for creating examples and explanations during learning activities, which can often be monotonous and time-consuming… However, with Graphologue, acquiring context-specific examples is straightforward: “clicking the ‘Examples’ button just does it, and (you) don’t have to think of another prompt” (P5)… They selected the node ‘particles’ in the diagram and clicked the ‘Explain’ button. This action extended the paragraph and constructed a new part of the diagram stemming from the node, serving to clarify the context-specific meaning of ‘particles.’” [section 9.4.1] “Participants particularly praised the diagrams for providing an “overall view of a topic” and for “understanding a set of instructions” (P8), likening it to a “mind map” (P8), and stating that it provides an understanding of the organization of information (P5).” – The Specification explains that “expend ideas” is a selectable prompt suggestion that is incorporated into the user prompt and used to generate a draft mind map. The Specification further describes the corresponding operation as transferring content into a mind map and expanding each idea to multiple levels of detail. Jiang similarly teaches predefined “Examples” and “Explain” controls that construct prompts for generating additional examples and explanations relating to selected diagram content. Selection of the “Explain” control extends the existing content and constructs a new portion of the diagram stemming from the selected idea. Jiang further characterizes its diagrams as mind-map-like representations that organize related information. Under BRI, and consistent with the Specification’s description of “expend ideas,” Jiang’s generation of additional explanatory content and extension of a mind-map-like diagram from a selected idea corresponds to the claimed alternative of expending ideas.) Regarding claim 7, Jiang in view of Orozco teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. Jiang in view of Orozco further teaches: receiving at least one user feedback on the diagram via the user interface of the client device (Jiang, [section 9.4.3] “P4, for example, intended to understand the meaning of ‘particles’ in “charged particles from the sun.” They selected the node ‘particles’ in the diagram and clicked the ‘Explain’ button. This action extended the paragraph and constructed a new part of the diagram stemming from the node, serving to clarify the context-specific meaning of ‘particles.’ P4” – The Specification describes user feedback as input received through the diagram user interface, including interface selections and indications that the diagram should be improved, and teaches using that feedback to imp, rove or regenerate the diagram (Spec, paragraphs [0096]-[0097]). Jiang similarly teaches receiving user input through the Graphologue interface in which a user selects a particular node in the displayed diagram and activates the “Explain” control. The selected node identifies the diagram content to which the input relates, while activation of the control communicates that the selected content should be clarified or expanded. Under BRI, this user input concerning the displayed diagram corresponds to receiving user feedback on the diagram through the user interface of the client device.). Regarding claim 8, Jiang in view of Orozco teaches all the elements of claim 7, therefore is rejected for the same reasons as those presented for claim 7. Jiang in view of Orozco further teaches: constructing, via the prompt construction unit, a second prompt by appending the feedback and the diagram to a second instruction string, the second instruction string including instructions to the generative model to generate at least another diagram based on the feedback and the diagram, by adjusting one or more attributes of the diagram based on the feedback (Jiang, [section 9.4.3] “P4, for example, intended to understand the meaning of ‘particles’ in “charged particles from the sun.” They selected the node ‘particles’ in the diagram and clicked the ‘Explain’ button. This action extended the paragraph and constructed a new part of the diagram stemming from the node, serving to clarify the context-specific meaning of ‘particles.’ P4” [section 6.2.2] “In the self-correction prompt, the previously generated paragraph serves as context. We pinpoint the specific sentence requiring correction and dynamically describe the issue, e.g., “entities labeled $N11 and $N12 are mentioned but lack connecting relationships.” We instruct GPT-4 to re-annotate the sentence or slightly rewrite it if needed for better annotation (Appendix A.2)… Once an updated annotation is complete, the diagram is adjusted and animated to reflect the changes.” – Jiang teaches receiving user feedback by allowing a user to select an existing diagram content and invoke “Explain” operation, thereby identifying the portion of the diagram and the requested modification. Jiang further teaches constructing a subsequent prompt in which previously generated content serves as context, the requested issue or modification is dynamically included, and instructions direct GPT-4 to reannotate or rewrite the affected content. Under BRI, combining user feedback and existing diagram content with the follow-up instructions corresponds to appending the feedback and the diagram to a second instruction string. The resulting prompt causes new or corrected annotated content to be generated and adjusts attributes of the diagram, including nodes, labels, relationships, content, and structure, based on feedback.); providing, via the prompt construction unit, as an input the second prompt to the generative model and receiving as an output the other diagram of the digital content from the generative model; and providing the other diagram to the client device to be presented on the user interface of the client device (Jiang, [section 6.2.2] “We instruct GPT-4 to re-annotate the sentence or slightly rewrite it if needed for better annotation (Appendix A.2)… Once an updated annotation is complete, the diagram is adjusted and animated to reflect the changes.” [section 7.1] “To accomplish this, we parse the text streamed in from GPT-4 and immediately add new entities and relationships to diagrams as they are generated. This ensures that users have a responsive and up-to-date visual representation of the LLM-generated information.” [Fig. 3, page 7] “The Graphologue interface, including the question input box (a), text response blocks (b), and the diagrams (c).” – Jiang teaches providing the subsequent corrective prompt to GPT-4 and receiving updated annotated content form GPT-4 in response. Graphologue uses the updated output to adjust the diagram and immediately adds the newly generated entities and relationships to the displayed diagram. The updated diagram is presented to the user through the Graphologue interface. Under BRI, this corresponds to providing the second prompt to the generative model, receiving the resulting other diagram, and presenting the diagram through the client-device user interface.). Regarding claim 9, Jiang in view of Orozco teaches all the elements of claim 7, therefore is rejected for the same reasons as those presented for claim 7. Jiang in view of Orozco further teaches: wherein the user feedback is collected via a user selection of at least one of a thumbs-up tab, a thumbs-down tab, a neutral tab, or a generating-more-image tab, a textual input, or a combination thereof (Jiang, [Abstract] “users can interact with the diagrams to flexibly adjust the graphical presentation and to submit context-specific prompts to obtain more information.” [section 5] “In Graphologue, she started by typing ‘What is an earthquake?’.” [section 10.2] “The graphical user interface of Graphologue enables users to employ direct manipulation with the diagram to request explanations and examples from LLMs, saving users’ efforts to manually craft textual prompts.” – Jiang teaches collecting user feedback through the Graphologue user interface by receiving textual input in the form of typed prompts and by receiving user selections requesting additional explanations or examples concerning the displayed diagram. Under BRI, the typed content-specific prompts correspond to the claimed textual input, while the interface selection requesting additional generated information corresponds to a generating-more-image selection.). Regarding claim 11, Jiang in view of Orozco teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. Jiang in view of Orozco further teaches: wherein the generative model is a language model or a multimodal model (Jiang, [Abstract] “We present Graphologue, an interactive system that converts text-based responses from LLMs into graphical diagrams.” [section 6.2.2] “We instruct GPT-4 to re-annotate the sentence or slightly rewrite it if needed for better annotation”). Regarding claim 13, Jiang teaches the following limitations: receiving, via a client device, a user prompt requesting a diagram representing digital content, wherein the digital content includes any of text, audio, video, or structured file (Jiang, [Abstract] “Large language models (LLMs) have recently soared in popularity due to their ease of access and the unprecedented ability to synthesize text responses to diverse user questions… We present Graphologue, an interactive system that converts text-based responses from LLMs into graphical diagrams to facilitate information-seeking and question-answering tasks. Graphologue employs novel prompting strategies and interface designs to extract entities and relationships from LLM responses and constructs node-link diagrams in real-time.” [Fig. 1] “Graphologue constructs an interactive diagram in real-time as GPT-4 text responses are streamed in.” – Jiang teaches an interactive system that receives user questions through an interface and converts the resulting text-based LLM responses into graphical node-link diagrams. Under the broadest reasonable interpretation, the user question corresponds to the claimed user prompt, and the LLM text response corresponds to digital content including text that is represented by the generated diagram. Accordingly, Jiang teaches receiving, via a client device, a user prompt for a diagram representing text content.); providing, via the prompt construction unit, as an input the first prompt to the generative model and receiving as an output the diagram from the generative model (Jiang, [section 6] “We develop and test the following prompting strategies with OpenAI’s GPT-4, the most advanced and publicly available LLM to date.” [section 6.1] “We outline our key prompt components, which work together to instruct GPT-4 to generate an initial response that facilitates dynamic diagram construction (D1) and enables easy control over the complexity of presented information (D2) through interactive diagrams (Appendix A.1).” [Figure 2, page 6] “As GPT-4 responses stream in, Graphologue parses them in real-time, removes inline annotations for the interface, extracts entities and relationships, and constructs the corresponding diagrams.” – Jiang teaches providing the constructed prompt components as input to GPT-4, which corresponds to the claimed generative model, to generate an annotated response facilitating diagram construction. Jiang further teaches receiving the streamed GPT-4 response, extracting entities and relationships, and constructing the corresponding rendered diagram. Under BRI, the Graphologue prompt corresponds to the first prompt provided as input to the generative model, and the corresponding rendered diagram constitutes the diagram output.); and providing the diagram to the client device to be presented on a user interface of the client device (Jiang, [section 4, D3] “By utilizing diagrams as the main interface with LLMs, typical information tasks should be supported through interaction with the nodes and links in these diagrams” [section 5] “In Graphologue, she started by typing ‘What is an earthquake?’. As response text streamed into the interface from the LLM, she noticed a node-link diagram was being constructed piece by piece on the side simultaneously” [Fig. 2, page 6] “As GPT-4 responses stream in, Graphologue parses them in real-time, removes inline annotations for the interface, extracts entities and relationships, and constructs the corresponding diagrams.” – Jiang teaches presenting the constructed node-link diagram through Graphologue’s user interface as the LLM response streams into the interface. Under BRI, the device executing and displaying the Graphologue interface corresponds to the client device, and displaying the diagram within that interface corresponds to providing the diagram to be presented on the user interface of the client device.). However, Jiang does not teach but Jiang in view of Orozco teaches the following limitations: constructing, via a prompt construction unit, a first prompt by appending the user prompt and the digital content to a first instruction string, the first instruction string including instructions to a generative model to identify semantic context of the digital content based on metadata of the digital content, to identify at least one of a text data item, an audio data item, a video data item, or a structured file item, embedded in the digital content to generate at least one of a text transcript of the audio data item, a text transcript of the video data item, a text transcript of the structured file item, a text description of the audio data item, a textual description of the video data item, or a text description of the structured file item (Jiang, [section 6] “To enable the diagrams to be constructed simultaneously as the response streams in, we iteratively develop our prompts to have the LLM annotate the entities and relationships inline with the tokens… We develop and test the following prompting strategies with OpenAI’s GPT-4, the most advanced and publicly available LLM to date. A full list of original prompts can be found in Appendix A.” [section 6.1] “We outline our key prompt components, which work together to instruct GPT-4 to generate an initial response that facilitates dynamic diagram construction (D1) and enables easy control over the complexity of presented information (D2) through interactive diagrams (Appendix A.1).” [section 10.2] “Moreover, a text input box can be provided to allow customized requests. As the user explores the knowledge space through the graphical interface, their prior actions and the current diagram can be leveraged as context to construct prompts for LLMs to get responses that are better aligned with the user’s needs.” [Appendix A.1, page 17] “System Please provide a well-structured response to the user’s question in multiple paragraphs… The user’s goal is to construct a concept map to visually explain your response. To achieve this, annotate the key entities and relationships inline for each sentence in the paragraphs.” Orozco, paragraph [0034] “ As an example, the modifications may be related to an upcoming event, and the modification inputs 130 can include the name of the event, a description of the event, keywords related to the event, and so on. The modification inputs 130 can include natural language, tags, titles, etc.” [0036] “The modification inputs 130 can include data of different modalities, such as, images, video, computer drawings, audio, text, and so on.” [0037] “For example, the seed generation engine 195 can generate, from the modification inputs 130, a natural language prompt describing elements of each of the modification inputs 130. The prompt can include the textual information included as part of the inputs 130, as well as textual descriptions of inputs of other modalities provided, e.g., text descriptions of images, video, or audio transcripts.” [0049] “ An embedding at least partially encodes some semantic meaning for the text, image, video, etc., represented by the embedding.” [0080] “ The generative model 310 provides, as output, the generated features associated with the existing digital content. The generated features may be, for example embeddings associated with the existing digital content and/or a natural language description of the existing digital content.” – Jiang teaches developing and constructing prompts using customized user requests and contextual digital content and supplying GPT-4 with predefined textual instructions. Under BRI, Jiang’s System instruction block corresponds to the first instruction string. Orozco teaches generating a prompt from multimodal digital content that includes textual information, textual descriptions, and transcripts. Orozco further teaches keywords, tags, and titles associated with the content and embeddings that encode semantic meaning. Under BRI, the keywords, tags, and titles correspond to metadata, and the encoded semantic meaning corresponds to semantic context. Thus, Jiang in view of Orozco teaches the claimed prompt construction, context identification, and content-item identification, and generation of a corresponding transcript or text description.), to semantically analyze and extract diagram data from at least one of the text data item, the text transcripts, or the textual descriptions based on the semantic context, and to generate the diagram of the digital content based on the diagram data (Jiang, [Abstract] “Graphologue employs novel prompting strategies and interface designs to extract entities and relationships from LLM responses and constructs node-link diagrams in real-time.” [section 6.1.2] “GPT-4 is instructed to annotate entities in the text to serve as nodes in the diagrams.” [section 6.1.3] “In addition to entities, we prompt GPT-4 to identify and annotate relationships between these entities inline (Figure 2). These relationships serve as the links in the node-link diagram…. GPT-4 is instructed to include them in a single annotation, which we utilize to organize the connections together when rendering them on the canvas” [Appendix, A.1, page 17] “You should try to annotate at least one relationship for each entity. Relationships should only connect entities that appear in the response. You can arrange the sentences in a way that facilitates the annotation of entities and relationships, but the arrangement should not alter their meaning, and they should still flow naturally in language.” [section 9.4.3] “They selected the node ‘particles’ in the diagram and clicked the ‘Explain’ button. This action extended the paragraph and constructed a new part of the diagram stemming from the node, serving to clarify the context-specific meaning of ‘particles.’” Orozco, paragraph [0037] “The prompt can include the textual information included as part of the inputs 130, as well as textual descriptions of inputs of other modalities provided, e.g., text descriptions of images, video, or audio transcripts.” [0049] “An embedding at least partially encodes some semantic meaning for the text, image, video, etc., represented by the embedding.” – Jiang teaches analyzing the meaning of text to identify entities and relationships and clarifying the context-specific meaning of terms. Under BRI, this corresponds to semantically analyzing the text based on semantic context. Jiang further represents the identified entities and relationships as nodes, links, and annotations, which correspond to diagram data, and renders the node-link diagram from the data. Orozco teaches that the analyzed input may include textual content, transcripts, or textual descriptions derived from multimodal digital content and the embeddings encode the semantic meaning of that content.); Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, having a combination of Jiang and Orozco before them, to incorporate Orozco’s processing of multimodal digital content, including the use of keywords, tags, and titles to identify semantic context and the generation of textual descriptions or transcripts of non-text content, into the prompt-based diagram generation system of Jiang. One would have been motivated to make such a combination in order to allow Jiang’s system to generate diagrams from digital content provided in different modalities rather than being limited to existing text responses. This would predictably enable Jiang’s generative model to semantically analyze textual descriptions or transcripts representing audio, video, or other digital content and extract the entities and relationships used to construct the corresponding diagram. Regarding claim 14, Jiang in view of Orozco teaches all the elements of claim 13, therefore is rejected for the same reasons as those presented for claim 13. The claim recites similar limitations corresponding to claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale. Regarding claim 15, Jiang in view of Orozco teaches all the elements of claim 14, therefore is rejected for the same reasons as those presented for claim 14. The claim recites similar limitations corresponding to claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rationale. Regarding claim 16, Jiang in view of Orozco teaches all the elements of claim 14, therefore is rejected for the same reasons as those presented for claim 14. The claim recites similar limitations corresponding to claim 4 and is rejected for similar reasons as claim 4 using similar teachings and rationale. Regarding claim 17, Jiang teaches the following limitations: receiving, via a client device, a user prompt requesting a diagram representing digital content, wherein the digital content includes any of text, audio, video, or structured file (Jiang, [Abstract] “Large language models (LLMs) have recently soared in popularity due to their ease of access and the unprecedented ability to synthesize text responses to diverse user questions… We present Graphologue, an interactive system that converts text-based responses from LLMs into graphical diagrams to facilitate information-seeking and question-answering tasks. Graphologue employs novel prompting strategies and interface designs to extract entities and relationships from LLM responses and constructs node-link diagrams in real-time.” [Fig. 1] “Graphologue constructs an interactive diagram in real-time as GPT-4 text responses are streamed in.” – Jiang teaches an interactive system that receives user questions through an interface and converts the resulting text-based LLM responses into graphical node-link diagrams. Under the broadest reasonable interpretation, the user question corresponds to the claimed user prompt, and the LLM text response corresponds to digital content including text that is represented by the generated diagram. Accordingly, Jiang teaches receiving, via a client device, a user prompt for a diagram representing text content.); providing, via the prompt construction unit, as an input the first prompt to the generative model and receiving as an output the diagram from the generative model (Jiang, [section 6] “We develop and test the following prompting strategies with OpenAI’s GPT-4, the most advanced and publicly available LLM to date.” [section 6.1] “We outline our key prompt components, which work together to instruct GPT-4 to generate an initial response that facilitates dynamic diagram construction (D1) and enables easy control over the complexity of presented information (D2) through interactive diagrams (Appendix A.1).” [Figure 2, page 6] “As GPT-4 responses stream in, Graphologue parses them in real-time, removes inline annotations for the interface, extracts entities and relationships, and constructs the corresponding diagrams.” – Jiang teaches providing the constructed prompt components as input to GPT-4, which corresponds to the claimed generative model, to generate an annotated response facilitating diagram construction. Jiang further teaches receiving the streamed GPT-4 response, extracting entities and relationships, and constructing the corresponding rendered diagram. Under BRI, the Graphologue prompt corresponds to the first prompt provided as input to the generative model, and the corresponding rendered diagram constitutes the diagram output.); and providing the diagram to the client device to be presented on a user interface of the client device (Jiang, [section 4, D3] “By utilizing diagrams as the main interface with LLMs, typical information tasks should be supported through interaction with the nodes and links in these diagrams” [section 5] “In Graphologue, she started by typing ‘What is an earthquake?’. As response text streamed into the interface from the LLM, she noticed a node-link diagram was being constructed piece by piece on the side simultaneously” [Fig. 2, page 6] “As GPT-4 responses stream in, Graphologue parses them in real-time, removes inline annotations for the interface, extracts entities and relationships, and constructs the corresponding diagrams.” – Jiang teaches presenting the constructed node-link diagram through Graphologue’s user interface as the LLM response streams into the interface. Under BRI, the device executing and displaying the Graphologue interface corresponds to the client device, and displaying the diagram within that interface corresponds to providing the diagram to be presented on the user interface of the client device.). However, Jiang does not teach but Jiang in view of Orozco teaches the following limitations: A non-transitory computer readable medium on which are stored instructions that, when executed, cause a programmable device to perform functions of (Orozco, paragraph [0163] “ Aspects of this disclosure can further be implemented as one or more computer programs, such as one or more engines or modules of computer program instructions encoded on one or more tangible non-transitory computer storage media for execution by, or to control the operation of, one or more data processing apparatus. “): constructing, via a prompt construction unit, a first prompt by appending the user prompt and the digital content to a first instruction string, the first instruction string including instructions to a generative model to identify semantic context of the digital content based on metadata of the digital content, to identify at least one of a text data item, an audio data item, a video data item, or a structured file item, embedded in the digital content to generate at least one of a text transcript of the audio data item, a text transcript of the video data item, a text transcript of the structured file item, a text description of the audio data item, a textual description of the video data item, or a text description of the structured file item (Jiang, [section 6] “To enable the diagrams to be constructed simultaneously as the response streams in, we iteratively develop our prompts to have the LLM annotate the entities and relationships inline with the tokens… We develop and test the following prompting strategies with OpenAI’s GPT-4, the most advanced and publicly available LLM to date. A full list of original prompts can be found in Appendix A.” [section 6.1] “We outline our key prompt components, which work together to instruct GPT-4 to generate an initial response that facilitates dynamic diagram construction (D1) and enables easy control over the complexity of presented information (D2) through interactive diagrams (Appendix A.1).” [section 10.2] “Moreover, a text input box can be provided to allow customized requests. As the user explores the knowledge space through the graphical interface, their prior actions and the current diagram can be leveraged as context to construct prompts for LLMs to get responses that are better aligned with the user’s needs.” [Appendix A.1, page 17] “System Please provide a well-structured response to the user’s question in multiple paragraphs… The user’s goal is to construct a concept map to visually explain your response. To achieve this, annotate the key entities and relationships inline for each sentence in the paragraphs.” Orozco, paragraph [0034] “ As an example, the modifications may be related to an upcoming event, and the modification inputs 130 can include the name of the event, a description of the event, keywords related to the event, and so on. The modification inputs 130 can include natural language, tags, titles, etc.” [0036] “The modification inputs 130 can include data of different modalities, such as, images, video, computer drawings, audio, text, and so on.” [0037] “For example, the seed generation engine 195 can generate, from the modification inputs 130, a natural language prompt describing elements of each of the modification inputs 130. The prompt can include the textual information included as part of the inputs 130, as well as textual descriptions of inputs of other modalities provided, e.g., text descriptions of images, video, or audio transcripts.” [0049] “ An embedding at least partially encodes some semantic meaning for the text, image, video, etc., represented by the embedding.” [0080] “ The generative model 310 provides, as output, the generated features associated with the existing digital content. The generated features may be, for example embeddings associated with the existing digital content and/or a natural language description of the existing digital content.” – Jiang teaches developing and constructing prompts using customized user requests and contextual digital content and supplying GPT-4 with predefined textual instructions. Under BRI, Jiang’s System instruction block corresponds to the first instruction string. Orozco teaches generating a prompt from multimodal digital content that includes textual information, textual descriptions, and transcripts. Orozco further teaches keywords, tags, and titles associated with the content and embeddings that encode semantic meaning. Under BRI, the keywords, tags, and titles correspond to metadata, and the encoded semantic meaning corresponds to semantic context. Thus, Jiang in view of Orozco teaches the claimed prompt construction, context identification, and content-item identification, and generation of a corresponding transcript or text description.), to semantically analyze and extract diagram data from at least one of the text data item, the text transcripts, or the textual descriptions based on the semantic context, and to generate the diagram of the digital content based on the diagram data (Jiang, [Abstract] “Graphologue employs novel prompting strategies and interface designs to extract entities and relationships from LLM responses and constructs node-link diagrams in real-time.” [section 6.1.2] “GPT-4 is instructed to annotate entities in the text to serve as nodes in the diagrams.” [section 6.1.3] “In addition to entities, we prompt GPT-4 to identify and annotate relationships between these entities inline (Figure 2). These relationships serve as the links in the node-link diagram…. GPT-4 is instructed to include them in a single annotation, which we utilize to organize the connections together when rendering them on the canvas” [Appendix, A.1, page 17] “You should try to annotate at least one relationship for each entity. Relationships should only connect entities that appear in the response. You can arrange the sentences in a way that facilitates the annotation of entities and relationships, but the arrangement should not alter their meaning, and they should still flow naturally in language.” [section 9.4.3] “They selected the node ‘particles’ in the diagram and clicked the ‘Explain’ button. This action extended the paragraph and constructed a new part of the diagram stemming from the node, serving to clarify the context-specific meaning of ‘particles.’” Orozco, paragraph [0037] “The prompt can include the textual information included as part of the inputs 130, as well as textual descriptions of inputs of other modalities provided, e.g., text descriptions of images, video, or audio transcripts.” [0049] “An embedding at least partially encodes some semantic meaning for the text, image, video, etc., represented by the embedding.” – Jiang teaches analyzing the meaning of text to identify entities and relationships and clarifying the context-specific meaning of terms. Under BRI, this corresponds to semantically analyzing the text based on semantic context. Jiang further represents the identified entities and relationships as nodes, links, and annotations, which correspond to diagram data, and renders the node-link diagram from the data. Orozco teaches that the analyzed input may include textual content, transcripts, or textual descriptions derived from multimodal digital content and the embeddings encode the semantic meaning of that content.); Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, having a combination of Jiang and Orozco before them, to incorporate Orozco’s processing of multimodal digital content, including the use of keywords, tags, and titles to identify semantic context and the generation of textual descriptions or transcripts of non-text content, into the prompt-based diagram generation system of Jiang. One would have been motivated to make such a combination in order to allow Jiang’s system to generate diagrams from digital content provided in different modalities rather than being limited to existing text responses. This would predictably enable Jiang’s generative model to semantically analyze textual descriptions or transcripts representing audio, video, or other digital content and extract the entities and relationships used to construct the corresponding diagram. Regarding claim 18, Jiang in view of Orozco teaches all the elements of claim 17, therefore is rejected for the same reasons as those presented for claim 17. The claim recites similar limitations corresponding to claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale. Regarding claim 19, Jiang in view of Orozco teaches all the elements of claim 18, therefore is rejected for the same reasons as those presented for claim 18. The claim recites similar limitations corresponding to claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rationale. Regarding claim 20, Jiang in view of Orozco teaches all the elements of claim 18, therefore is rejected for the same reasons as those presented for claim 18. The claim recites similar limitations corresponding to claim 4 and is rejected for similar reasons as claim 4 using similar teachings and rationale. Claim 10 is rejected under the 35 U.S.C. 103 as being unpatentable over Jiang et al., (NPL: “Graphologue: Exploring Large Language Model Responses with Interactive Diagrams” (Published: 2023)) In view of Orozco et al., (Pub. No.: US 20240394945 A1 (Filed: May 22, 2024)) further in view of Acuff et al., (Pub. No.: US 20220292427 A1 (Filed: 2022)). Regarding claim 10, Jiang in view of Orozco teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. However, Jiang in view of Orozco does not teach but Jiang in view of Orozco further in view of Acuff teaches: causing the user interface to receive a user confirmation of the diagram; and causing a publication of the diagram (Jiang, [Abstract] “Further, users can interact with the diagrams to flexibly adjust the graphical presentation and to submit context-specific prompts to obtain more information.” Acuff, paragraph [0107] “ a user can review the data and perform an interaction using a user interface…the system can provide the alert to the user through the user interface, and then the user can confirm or deny the accuracy of the alert using the user interface.” [0138] “ Following the steps collectively labeled under “Cognition Studio”, a user such as a business analyst publishes the scenario(s) to a data repository labeled in the diagram of FIG. 13 as “Cognition Repository”.” – Jiang teaches presenting the resulting diagram through an interactive user interface for user review and manipulation. Acuff teaches receiving, through a user interface, a user decision confirming the accuracy of the displayed content and publishing completed content into a repository. Applied to Jiang’s resulting diagram, these teachings correspond to causing the user interface to receive a user confirmation that the diagram is satisfactory and causing the confirmed diagram to be published.). Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, having a combination of Jiang and Acuff before them, to incorporate the user confirmation and publication workflow taught by Acuff into the interactive diagram generation system of Jiang. One would have been motivated to make such a combination in order to allow a user to confirm that a generated or refined diagram is satisfactory before the diagram is published. This would help ensure that the published diagram accurately reflects the user’s intended content and presentation and reduce the likelihood that an incomplete, inaccurate, or unintended diagram is distributed. Claim 11 is rejected under the 35 U.S.C. 103 as being unpatentable over Jiang et al., (NPL: “Graphologue: Exploring Large Language Model Responses with Interactive Diagrams” (Published: 2023)) In view of Orozco et al., (Pub. No.: US 20240394945 A1 (Filed: May 22, 2024)) further in view of Hu et al., (NPL: “Visualizing Social Media Content with SentenTree” (Filed: 2017)). Regarding claim 12, Jiang in view of Orozco teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. However, Jiang in view of Orozco does not teach but Jiang in view of Orozco further in view of Hu teaches: wherein the digital content and the user prompt are received via a software application, and wherein the software application is a virtual meeting and collaboration application, a digital whiteboard application, an employee experience application, an online collaboration application, a calendar application, an email application, a task management application, a team-work planning application, a software development application, an enterprise accounting and sales application, a social media application, or an online encyclopedia (Jiang, [section 5] “In Graphologue, she started by typing ‘What is an earthquake?’” [Fig. 3, page 7] “The Graphologue interface, including the question input box (a), text response blocks (b), and the diagrams (c).” Hu, [Abstract] “ We introduce SentenTree, a novel technique for visualizing the content of unstructured social media text…It is implemented as a lightweight application that runs in the browser.” [section 3.5] “In this implementation, a person provides raw text (e.g., tweets) to the application.” – Jiang teaches receiving a user prompt through a software application by providing a question input box through which the user types a request. Hu teaches a browser-based software application that receives digital content in the form of social-media text, including tweets, and generates a node-link diagram from that content. Under BRI, the browser application receiving tweets corresponds to a social media application receiving digital content, while Jiang’s question-input interface corresponds to receiving the user prompt through a software application.) Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, having a combination of Jiang and Hu before them, to configure Jiang’s diagram generation system to receive social media text through a browser-based software application, as taught by Hu. One would have been motivated to make such a combination because Hu teaches that social media collections contain large volumes of unstructured text that are difficult to understand and that graphical representations provide a rapid overview of their key concepts and opinions. This would allow Jiang’s interactive diagram system to more efficiently present and organize information obtained from social-media content. Conclusion The prior art of record and not relied upon is consider pertinent to Applicant’s disclosure: 1. Tang, C. L., Liao, J., Wang, H. C., Sung, C. Y., & Lin, W. C. (2021, April). Conceptguide: Supporting online video learning with concept map-based recommendation of learning path. In Proceedings of the Web Conference 2021 (pp. 2757-2768). – teaches compiling of transcripts from multiple videos and generating a concept map representing relationships among concepts identified in video content. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Daravanh Phakousonh whose telephone number is (571)272-6324. The examiner can normally be reached Mon - Thurs 7 AM - 5 PM, Every other Friday 7 AM - 4PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li B Zhen can be reached at 571-272-3768. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Daravanh Phakousonh/Examiner, Art Unit 2121 /Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121
Read full office action

Prosecution Timeline

May 30, 2024
Application Filed
Jul 30, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12572821
ACCURACY PRIOR AND DIVERSITY PRIOR BASED FUTURE PREDICTION
4y 0m to grant Granted Mar 10, 2026
Study what changed to get past this examiner. Based on 1 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
33%
Grant Probability
99%
With Interview (+100.0%)
3y 8m (~1y 5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 3 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month