DETAILED ACTION
Introduction
This office action is in response to applicant’s amendment filed 4/27/2026. Claims 21-40 and 42 are currently pending and have been examined. Applicant’s IDS have been considered. There is no claim to foreign priority.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments, see remarks, filed 4/27/2026, with respect to claims 39 and 40 have been fully considered and are not persuasive.
Applicant argues, “identify an inconsistency between one or more terms in the first user input and one or more terms in the second user input, wherein the inconsistency comprises a contradiction or mismatch between attributes associated with the first user input and the second user input;
generate a recommendation to modify at least one of the first user input or the second user input to resolve the identified inconsistency;
“As claimed, the claim facilitates correction of inconsistencies between first and second user inputs during iterative content generation. An inconsistency "comprises a contradiction or mismatch between attributes associated with the first user input and the second user input." By correcting inputs during iterative content generation, content resulting from the first user input and the second user input may be made more consistent so that the contradiction or mismatch is resolved.”
“The Office Action acknowledges that Gupta does not teach or suggest these features as recited prior to the instant amendment to the claims, but alleges that Manjunath at paras. 0292- 295, 409, and 253 cures this deficiency. However, Manjunath generally relates to context-aware digital assistant interactions in a multi-participant communication session, in which prior user inputs are shared to enable interpretation of subsequent user inputs. Manjunath in the relied upon passages at best describe using context information from a first user input to disambiguate or interpret a second user input and to continue a conversational interaction. This does not teach or suggest identifying an inconsistency comprising a contradiction or mismatch between attributes associated with the first user input and the second user input for which a recommendation to resolve the inconsistency is generated as claimed. For at least this reason, independent claim 39 and its dependent claim 40 re allowable over Gupta and Manjunath.”
The Examiner notes, the newly claimed limitation includes, “identify an inconsistency between one or more terms in the first user input and one or more terms in the second user input, wherein the inconsistency comprises a contradiction or mismatch between attributes associated with the first user input and the second user input;”
The Examiner notes in at least part of the cited portion of Manjunath, it is disclosed,
“[0292] In some examples, user device 804 obtains the second digital assistant response further based on context information associated with the first user input (and received with the first digital assistant response). Specifically, in some examples, user device 804 (or one or more servers) uses context information associated with the first user input to disambiguate one or more ambiguous terms or phrases in the second user input (e.g., when determining a user intent based on the second user input) and subsequently determine the second digital assistant response based on the disambiguated second user input. For example, if the first user input is “Hey Siri, what's the weather like in Palo Alto?” and the second user input is “Hey Siri, how long will it take me to drive there?”, user device 804 (or one or more servers) can use context information associated with the first user input (e.g., a dialog history received from user device 802) to disambiguate the term “there” and determine that the second user input represents a user request for navigation information (i.e., travel time) from a current location of user device 804 to Palo Alto (e.g., because “there” refers to the location mentioned in the first user input (i.e., Palo Alto)). As another example, if the first user input is “Hey Siri, what's the weather like in Palo Alto?” and the second user input is “Hey Siri, how about in New York?”, user device 804 (or one or more servers) can use context information associated with the first user input (e.g., a dialog history received from user device 802) to disambiguate the phrase “how about in New York” and determine that, the second user input represents a user request for weather information with respect to New York (e.g., because “how about in New York” refers to the task requested in the first user input).”
The Examiner notes, the concept of identifying an inconsistency between one or ore terms in the second user input, wherein the inconsistency comprises a contradiction or mismatch between attributes associated with the first user input and the second user input, does not include enough detail to overcome the current rejection. Wherein, the newly added limitation is considered extremely broad, and a reasonable interpretation of the claimed limitation includes:
The terms in the claims as the words as described by Manjunath.
Inconsistency as a contradiction or mismatch between the attributes associated with the first user input and the second user input, as the location information in between the first and second input or contradictory or do not match, i.e. Palo Alto and New York, the attributes comprising contextual word meanings, location information, intent, etc.
Manjunath used both inputs to disambiguate information and iteratively refines the user intent, as a modified recommendation to either input, and generate a response based thereon.
The Examiner notes, that in combination with Gupta, there is clearly a motivation to combine a system that receives multiple inputs in an iterative manner and disambiguate the input when an inconsistence or mismatch occurs between successive inputs. The Examiner notes should the applicant decide to define what type of inconsistency, contraction and/or mismatch, or the manner of similarity calculation, difference, or threshold between the differences, with respect to the inconsistency, then there may be a distinction between the current prior art and newly submitted claims which include these features. However, as it stands, all that is required is met with respect to the prior art. A mismatch or contradiction is immensely broad, and is covered by the current prior art. Gupta still includes second input operations, which modifies, by the language model, the content based on the second user input, but would be enhanced and/or improved by being able to disambiguate between “terms” in the first and second user input, using attributes associated therewith, in order to define a clear intent (as described below). Therefore, the applicant’s corresponding arguments with respect to the rejection of claims 39 and 40 are deemed non-persuasive.
Allowable Subject Matter
Claims 21-38 and 42 are allowed.
The following is an examiner’s statement of reasons for allowance:
The instant application is deemed to be directed to a non-obvious improvement over the disclosure of the previously cited prior art (See the previous office action, and pertinent prior art as cited in the PTO-892).
None of the above references teach alone or in obvious combination:
Regarding claim 21, A system, comprising:
a processor programmed to:
receive, during iterative content generation, a first user input;
identify a context associated with the first user input;
execute, based on the user input and the identified context, a language model trained to generate text and/or an image model trained to generate images;
generate, as an output of the language model and/or the image model, content based on the user input and the identified context;
(i) after the content is generated, automatically generate a prompt to verify that the content satisfies a request from the first user input;
(ii) execute the image model based on the automatically generated prompt;
(iii) generate a response that indicates whether or not the content satisfies the request;
receive, during the iterative content generation, a second user input; and
modify, by the language model and/or the image model, the content based on the second user input.”
Independent claim 35 sets forth similar limitations as independent claim 21, and is thus allowed based on similar reasons and rationale.
Dependent claim 22-34 and 36-38 and 42 are allowed, as they depend from their respective allowed parent claims.
Any comments considered necessary by applicant must be submitted no later than the payment of the issue fee and, to avoid processing delays, should preferably accompany the issue fee. Such submissions should be clearly labeled “Comments on Statement of Reasons for Allowance.”
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 39 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gupta et al. (Gupta, US 2025/0131020) in view of Manjunath et al. (Manjunath, US 2021/0249009).
As per claim 39, Gupta teaches a computer a non-transitory computer readable medium storing instructions that, when executed by a processor, programs the processor to (paragraph [0011]-see his non-transitory processor readable storage medium discussion):
receive, during iterative content generation, a first user input (paragraph [0085, 0002, 0037]-including an initial prompt for iterative content generation, his generative language models, Fig. 3);
identify a context associated with the first user input (ibid-paragraphs [0084, 0085, 0006]-see his contextual information associated with a user);
execute, based on the user input and the identified context, a language model trained to generate text and/or an image model trained to generate images (ibid, paragraphs [0085, 0037-0040, 0002]-his LXM, generative AI large language model, for text summarization and image generation, see his input and context, used to generate the output, as image and/or text);
generate, as an output of the language model and/or the image model, content based on the user input and the identified context (ibid);
receive, during the iterative content generation, a second user input (ibid-paragraph [0085]);
[identify an inconsistency between one or more terms in the first user input and one or more terms in the second user input, wherein the inconsistency comprises a contradiction or mismatch between attributes associated with the first user input and the second user input];
[generate a recommendation to modify at least one of the first user input or the second user input to resolve the identified inconsistency]; and
modify, by the language model and/or the image model, the content based on the second user input (ibid, paragraph [0085, 0113]-as including his user second input operations, during the iterative content generation, and subsequent content generation as the modified content result, see paragraphs [0061-0065]-which further detail the above cited sections and Fig. 3, updating the content generated by the LXM discussion, see above response to arguments for expatiated content).
Gupta lacks explicitly teaching that which Manjunath teaches, identify an inconsistency between one or more terms in the first user input and one or more terms in the second user input, wherein the inconsistency comprises a contradiction or mismatch between attributes associated with the first user input and the second user input (paragraph [0292-0295, 0409, 0253]-his first user input and second user input, and ambiguity found between the input terms within the inputs, the terms in the claims as the words as described by Manjunath, inconsistency as a contradiction or mismatch between the attributes associated with the first user input and the second user input, as the location information in between the first and second input or contradictory or do not match, i.e. Palo Alto and New York, the attributes comprising contextual word meanings, location information, intent, etc. to disambiguate information and iteratively refines the user intent, as a modified recommendation to either input, and generate a response based thereon);
generate a recommendation to modify at least one of the first user input or the second user input to resolve the identified inconsistency (ibid-his generated modification to the input, wherein the generated recommendation to modify the input, to resolve the inconsistency/ambiguity is selected and used to modify the parameters of the input, and further executed, in his content generation environment).
Thus, it would have been obvious to one of ordinary skill in the linguistics art, before the effective filing date of the invention, as all the claimed elements were known in the prior art and one skilled in the art could have combined the elements as claimed by known methods (computer implemented techniques and algorithms combining processes and steps in natural language processing), in view of the teachings of Gupta and Manjunath to combine the prior art element of generation and modification of content as taught by Gupta with disambiguating inconsistencies found between a first input and a second input, and generating a disambiguation modification, which is used to resolve the ambiguity as taught by Manjunath as each element performs the same function as it does separately, as the combination would yield predictable results, KSR International Co. v. Teleflex Inc., 550 US. -- 82 USPQ2nd 1385 (2007), wherein the predictable result would be disambiguating an input using parameters, as recommendations, from another input in order to resolve any inconsistencies and execute a task (ibid-Manjunath).
Claim(s) 40 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gupta et al. (Gupta, US 2025/0131020) in view of Manjunath as applied to claim 39 above, and further in view of Gusarov et al. (Gusarov, US 2025/0148218).
As per claim 40, Gupta with Manjunath make obvious the non-transitory computer readable medium of claim 39, but lack explicitly teaching that which Gusarov teaches, wherein the first user input comprises an image and the content to be changed comprises text that describes the image (paragraphs [0084-0086]-his request to describe an image), and wherein the instructions, when executed by the processor, further program the processor to:
generate a first instruction, for input to an image model, to describe the image (ibid-paragraphs [0084-0086]);
execute the image model with the image and the instruction to generate the text that describes the image (ibid, paragraph [0084-0094]-see his AI, machine generation of description of image discussion);
generate a second instruction, for input to the language model, to generate text based on the text that describes the image, wherein the content includes the text generated by the language model (ibid-is AI, machine learning model, which generates content, which includes the text generated by the model, the cited section detailing the model generating text from an image, based on instructions and then generating from the generated text, new text based…such as a translation).
Thus, it would have been obvious to one of ordinary skill in the linguistics art, before the effective filing date of the invention, as all the claimed elements were known in the prior art and one skilled in the art could have combined the elements as claimed by known methods (computer implemented techniques and algorithms combining processes and steps in natural language processing), in view of the teachings of Gupta, Manjunath and Gusarov to combine the prior art element of generation and modification of content as taught by Gupta with generating a description as the generated content as taught by Gusarov as each element performs the same function as it does separately, as the combination would yield predictable results, KSR International Co. v. Teleflex Inc., 550 US. -- 82 USPQ2nd 1385 (2007), wherein the predictable result would be using a visual language model to perform visual question answering (ibid-Gusarov).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure (See PTO-892).
Brush et al. (Brush, US 2023/0368151) describes inconsistent textual first input with respect to a second input, wherein a recommendation is provided to resolve the inconsistency.
Sharifi et al. (Sharifi, US 2024/0119088) teaches identifying conflict between one or more terms in a first user input and one or more terns in a second user input, and resolving the inconsistency.
Applicant's amendment necessitated the ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LAMONT M SPOONER whose telephone number is (571)272-7613. The examiner can normally be reached 8:00 AM -5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571)272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LAMONT M SPOONER/ Primary Examiner, Art Unit 2657
7/6/2026