Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1-20 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Saleh et al. (Pub. No. US 2025/0165697).
Regarding claim 1, Saleh discloses a method for model instruction processing comprising:
acquiring an initial instruction input by a user (Specifically, Saleh discloses that a user selects a modifier key component (e.g., a “Topic” button), causing a corresponding modifier key string (e.g., “Suggest ideas for”) to be added to a prompt displayed in a textbox, and that the user thereafter enters a natural language string customizing the instruction (e.g., “a summer vacation”), the combination of which constitutes the initial instruction; see FIG. 3, par. 60, pane 301 and prompt 335);
generating a first display interface based on the initial instruction, with initial content corresponding to at least one field in the first display interface being determined based on the initial instruction (In Fig. 3, pane 302, prompt 335 is based on the initial instruction. The computing device sends the prompt in its current state, including the initial instruction, to a foundation model, which returns suggested modifier values that are displayed as selectable modifier value components corresponding to the field; see modifier value components 333, e.g. “Exciting,” “Relaxing,” “Budget-friendly,” and “Romantic” generated for the “Tone” field based on the prompt “Suggest ideas for a summer vacation”)); and
determining a first instruction (“Suggest ideas for a summer vacation. Make it sound relaxing. Include keywords beach, golf”) for inputting into a first target model (the foundation model) based on a field currently displayed in the first display interface and content corresponding to the field (upon selection of the “Generate” button, the application submits the entire prompt — built from the modifier keys (fields) and modifier values (content) currently displayed — to the foundation model for content generation; see FIG. 3, pane 304, “Generate” button 337, and par. 64).
Saleh therefore discloses each and every limitation of claim 1.
Regarding claim 2, Saleh discloses that generating the first display interface based on the initial instruction comprises determining a second instruction for inputting into a second target model based on the initial instruction, inputting the second instruction into the second target model to obtain a second output, and determining the initial content corresponding to the at least one field based on the second output (Specifically, Saleh discloses that, to obtain suggested modifier values for a selected or newly added modifier key, the computing device sends the entire prompt in its current, incomplete state (a “second instruction” built at least in part from the initial instruction) to the foundation model (the “second target model”) and instructs the foundation model to generate values (a “second output”) for completing the instruction; the computing device then displays those values as the initial content of the field; see pars. 61-62).
Regarding claim 3, Saleh discloses that the initial instruction includes an instruction for instructing the first target model to generate a text, and that the at least one field comprises one ore more of the following: a theme, a literary form, audience, a style, a rhythm, or content points (Specifically, Saleh discloses that, for a word processing document, the modifier key options may add instructions directed to “topic, audience, tone, platform, perspective, keywords to use, words to avoid, ideas with which to start or end the to-be-generated content, length, and so on,” (see par. 19) and that modifier keys may be generally ordered from “Topic,” “Audience,” and “Tone” to more specific keys. The disclosed “Audience” field (see Fig. 1) corresponds directly to the claimed “audience” field, and the disclosed “Tone” field (See Figs. 1 and 3) corresponds to the claimed “style” field. Because claim 3 requires only “one or more” of the recited fields, Saleh’s disclosure of an “Audience” field (or, alternatively, a “Tone” field) is sufficient to meet this limitation (see FIG. 7, table 700; description of modifier key components for word processing documents).
Regarding claim 4, Saleh discloses that the initial instruction includes an instruction for instructing the first target model to generate an image, and that the at least one field comprises one or more of the followings: a theme, a form, a painting type, an environment, light, a style, a color, an emotion, or a composition (Specifically, Saleh discloses that, for generating an image, the modifier key components may add modifier keys directed to “subject matter, style (e.g., artist or art school, photorealistic, cartoon, etc.), camera lens, dimensions or aspect ratio, background, depth of field, color palette, etc.” (see par. 19). The disclosed “style” and “color palette” fields correspond directly to the claimed “style” and “color” fields. Because claim 4 requires only “one or more” of the recited fields, this disclosure is sufficient to meet this limitation; see description of modifier key components for image generation).
Regarding claim 5, Saleh discloses that the first target model comprises a language model, and that the method further comprises inputting the first instruction into the first target model to obtain a first target output (Specifically, Saleh discloses, e.g. in par. 27, that the foundation model is representative of a deep learning model, such as BERT, ERNIE, T5, XLNet, or a generative pretrained transformer computing architecture such as GPT-3, GPT-3.5, ChatGPT, or GPT-4 — each a “language model” — and that, upon receiving the completed prompt (the first instruction), the application submits it to the foundation model, which generates and returns a reply (a first target output) that is displayed in the user interface (see description of foundation model 150/450; step 207 and step 209; FIG. 3, pane 304, “Generate” button 337).
Regarding claim 6, Saleh discloses, in response to an edit operation of the user on the field in the first display interface, updating the field displayed in the first display interface or the content corresponding to the field, wherein the edit operation includes modifying the content corresponding to the field (see disclosure that “the user can edit the modifier key in the textbox, then request new suggestions for modifier values,” and that the user may enter a custom modifier value (see par. 16), and adding a field and content corresponding to the field (see disclosure that the user sequentially selects additional modifier key components — e.g., first “Topic,” then “Tone,” then “Keywords” — to add new modifier keys and corresponding modifier values to the prompt; FIG. 3, panes 301–304; FIG. 6, workflow 600).
Regarding claim 7, Saleh discloses that the second target model comprises a language model, for the same reasons discussed above with respect to claim 5: the foundation model used to generate the suggested modifier values (the second output) is the same foundation model — a language model such as GPT-3, GPT-3.5, ChatGPT, or GPT-4 — used to generate the first target output (see description of foundation model 150/450).
Regarding claim 8, Saleh discloses an electronic device comprising at least one memory and at least one processor configured to perform the acquiring, generating, and determining steps discussed above with respect to claim 1. Specifically, Saleh discloses computing device 801, comprising processing system 802 operatively coupled with storage system 803, which stores software 805 including program instructions (prompt process 806) that, when executed by processing system 802, direct processing system 802 to perform the operations described with respect to process 200 and workflow 600 — i.e., to acquire an initial instruction, generate a first display interface with initial field content based on that instruction, and determine a first instruction based on the field currently displayed and its content, for the same reasons set forth above with respect to claim 1 (see FIG. 8, elements 801–805).
Regarding claim 9, Saleh discloses the limitations of claim 9, which parallel those of claim 2, for the same reasons discussed above with respect to claim 2, as implemented by the program code of software 805 executed by processing system 802 (see FIG. 8; step 203, step 205; FIG. 6, workflow 600).
Regarding claim 10, Saleh discloses the limitations of claim 10, which parallel those of claim 3, for the same reasons discussed above with respect to claim 3 (see FIG. 7, table 700; description of modifier key components for word processing documents).
Regarding claim 11, Saleh discloses the limitations of claim 11, which parallel those of claim 4, for the same reasons discussed above with respect to claim 4 (see description of modifier key components for image generation).
Regarding claim 12, Saleh discloses the limitations of claim 12, which parallel those of claim 5, for the same reasons discussed above with respect to claim 5 (see description of foundation model 150/450; step 207, step 209).
Regarding claim 13, Saleh discloses the limitations of claim 13, which parallel those of claim 6, for the same reasons discussed above with respect to claim 6, as implemented by the program code of software 805 executed by processing system 802.
Regarding claim 14, Saleh discloses the limitations of claim 14, which parallel those of claim 7, for the same reasons discussed above with respect to claim 7 (see description of foundation model 150/450).
Regarding claim 15, Saleh discloses a non-transitory computer storage medium storing program code that, when executed by a computer device, causes the computer device to perform the acquiring, generating, and determining steps discussed above with respect to claim 1. Specifically, Saleh discloses storage system 803, a computer-readable storage medium storing software 805 (program code), which, when executed by processing system 802, directs the computing device to perform the operations described with respect to process 200 and workflow 600, for the same reasons set forth above with respect to claim 1 (see FIG. 8, elements 802–805).
Regarding claim 16, Saleh discloses the limitations of claim 16, which parallel those of claim 2, for the same reasons discussed above with respect to claim 2 (see step 203, step 205; FIG. 6, workflow 600).
Regarding claim 17, Saleh discloses the limitations of claim 17, which parallel those of claim 3, for the same reasons discussed above with respect to claim 3 (see FIG. 7, table 700; description of modifier key components for word processing documents).
Regarding claim 18, Saleh discloses the limitations of claim 18, which parallel those of claim 4, for the same reasons discussed above with respect to claim 4 (see description of modifier key components for image generation).
Regarding claim 19, Saleh discloses the limitations of claim 19, which parallel those of claim 5, for the same reasons discussed above with respect to claim 5 (see description of foundation model 150/450; step 207, step 209).
Regarding claim 20, Saleh discloses the limitations of claim 20, which parallel those of claim 7, for the same reasons discussed above with respect to claim 7 (see description of foundation model 150/450).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHONG X NGUYEN whose telephone number is (571)270-1591. The examiner can normally be reached Mon-Fri 8am - 5pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Poon can be reached at (571)272-7440. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PHONG X NGUYEN/ Primary Patent Examiner, Art Unit 2617