Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments with respect to the 35 U.S.C 101 rejections have been fully considered and are persuasive. The rejection of claims 1-20 under 35 U.S.C. 101 has been withdrawn.
Applicant’s arguments with respect to the 35 U.S.C 112 rejections have been fully considered and are persuasive. The rejection of claims 19 under 35 U.S.C. 112 has been withdrawn.
Applicant’s arguments with respect to the rejection(s) of claim(s) 1-8 and 10-20 under 35 U.S.C. 102(a)(2) have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Yushkina et al.
Siegenthaler discloses the system receives selection of an entire image, input the entire image to the image model to generate descriptors, and generates the final input prompt using the input query and the descriptors for the entire image. The image model splits a single image into sub-images that form semantic units to identify objects in the image (paragraph [0054]). The additional context therefore “comprises a first portion of an image” in the sense that one or more identities of objects/entities in the image are generated using the image model, where each object/entity comprises a portion of the image (paragraphs [0079-0080]). However, Siegenthaler does not expressly disclose the newly added limitation of receiving a selection of a first sub-portion of an image as input to the image model.
Yushkina et al. disclose a method for generating a context-enriched response comprising receiving a selection of a first sub-portion of an image as context input. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the method of generating additional context using an image-based ML model to generate a description of a first sub-portion of an image as disclosed by Siegenthaler to a selection of a first sub-portion of an image for the reasons provided in the 103 rejections below.
Applicant’s amendments necessitated the new grounds of rejection. Accordingly, this action is FINAL.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 4-11 and 14-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Siegenthaler et al. (U.S. Patent Application Pub. No. 2025/0061146, hereinafter “Siegenthaler”), in view of Yushkina et al. (U.S. Patent Application Pub. No. 2024/0281481, hereinafter “Yushkina”).
In regard to claim 1, Siegenthaler discloses a computer-implemented method for generating a context-enriched response (Fig. 4, 400), the method comprising:
generating additional context for a prompt input based on a context input comprising an image (a query comprising an input text and an input image is received, paragraph [0076]), wherein generating the additional context comprises causing an image-based machine learning (ML) model to generate a description of a sub-portion of the image (one or more image models split a single image into sub-images that form semantic units to identify objects, paragraph [0054]; for each object comprising a portion of the image, a description of the object is generated, paragraphs [0079-0080]);
combining the additional context with the prompt input to generate a context-enriched prompt (the system generates an input prompt for an LLM from the text query and natural language descriptors of the input image, paragraph [0086]); and
executing one or more generative machine learning (ML) models on the context-enriched prompt to generate the context-enriched response (the system generates a response to the query by providing the input prompt to a large language model (LLM), paragraph [0087]).
While Siegenthaler discloses an ML model generates a description of sub-portions of an image, Siegenthaler do not expressly disclose receiving a selection of a first sub-portion of an image as a context input.
Yushkina discloses a method for generating additional context comprising receiving a selection of a first sub-portion of an image as a context input (a region search control allows a user to select a first sub-portion of an image as context input for generating a context-enriched response, paragraphs [0050] and [0057-0058]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to receive a selection of a first sub-portion of an image as context input to the method of Siegenthaler, because it would help the system provide a more relevant generated response by allowing the user to specify a particular object in the image, as taught by Yushkina (paragraph [0074]).
In regard to claim 4, Siegenthaler discloses generating the additional context comprises determining a first set of annotations corresponding to the first portion of the image (one or more texts extracted from the image, paragraph [0079]).
As noted with respect to claim 1, Siegenthaler does not disclose the context input is a sub-portion of the image. However, for the same reasons as claim 1, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to determine a first set of annotations corresponding to the first sub-portion of the image in view of Yushkina.
In regard to claim 5, Siegenthaler discloses generating the additional context comprises:
identifying a first object within the first portion of the image (objects/entities within the image, paragraph [0079]); and
generating a first set of data corresponding to the first object (the contextual information is provided as structured data, see Abstract and paragraph [0055]).
As noted with respect to claim 1, Siegenthaler does not disclose the context input is a sub-portion of the image. However, for the same reasons as claim 1, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to identify a first object within the first sub-portion of the image in view of Yushkina.
In regard to claim 6, Siegenthaler discloses the additional context comprises a first portion of text (natural language descriptions of the image, paragraph [0045]), the prompt input comprises a second portion of text (input text query, paragraph [0045]), and combining the additional context with the prompt input comprises concatenating the first portion of text and the second portion of text (a natural language prompt comprising the contextual information and query text, paragraph [0046]).
In regard to claim 7, Siegenthaler discloses receiving a compound prompt that includes the prompt input and the context input (a combined text and image query, paragraph [0038]).
In regard to claim 8, Siegenthaler discloses the compound prompt comprises a multimodal prompt (a combined text and image query, paragraph [0038]).
In regard to claim 10, Siegenthaler discloses at least a portion of the additional context comprises a prompt history associated with the one or more generative ML models (the prompt is enriched with conversation history, paragraph [0046]).
In regard to claim 11, Siegenthaler discloses one or more non-transitory computer-readable media (paragraph [0121]) including instructions that, when executed by one or more processors, cause the one or more processors to generate a context-enriched response by performing the steps of:
generating additional context for a prompt input based on a context input comprising an image (a query comprising an input text and an input image is received, paragraph [0076]), wherein generating the additional context comprises causing an image-based machine learning (ML) model to generate a description of a sub-portion of the image (one or more image models split a single image into sub-images that form semantic units to identify objects, paragraph [0054]; for each object comprising a portion of the image, a description of the object is generated, paragraphs [0079-0080]);
combining the additional context with the prompt input to generate a context-enriched prompt (the system generates an input prompt for an LLM from the text query and natural language descriptors of the input image, paragraph [0086]); and
executing one or more generative machine learning (ML) models on the context-enriched prompt to generate the context-enriched response (the system generates a response to the query by providing the input prompt to a large language model (LLM), paragraph [0087]).
While Siegenthaler discloses an ML model generates a description of sub-portions of an image, Siegenthaler do not expressly disclose receiving a selection of a first sub-portion of an image as a context input.
Yushkina discloses a method for generating additional context comprising receiving a selection of a first sub-portion of an image as a context input (a region search control allows a user to select a first sub-portion of an image as context input for generating a context-enriched response, paragraphs [0050] and [0057-0058]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to receive a selection of a first sub-portion of an image as context input to the method of Siegenthaler, because it would help the system provide a more relevant generated response by allowing the user to specify a particular object in the image, as taught by Yushkina (paragraph [0074]).
In regard to claim 14, Siegenthaler discloses the step of generating the additional context comprises determining a first set of annotations corresponding to the first portion of the image (one or more texts extracted from the image, paragraph [0079]).
As noted with respect to claim 1, Siegenthaler does not disclose the context input is a sub-portion of the image. However, for the same reasons as claim 1, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to determine a first set of annotations corresponding to the first sub-portion of the image in view of Yushkina.
In regard to claim 15, Siegenthaler discloses the step of generating the additional context comprises:
identifying a first object within the first portion of the image (objects/entities within the image, paragraph [0079]); and
generating a first set of data corresponding to the first object (the contextual information is provided as structured data, see Abstract and paragraph [0055]).
As noted with respect to claim 1, Siegenthaler does not disclose the context input is a sub-portion of the image. However, for the same reasons as claim 1, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to identify a first object within the first sub-portion of the image in view of Yushkina.
In regard to claim 16, Siegenthaler discloses the additional context comprises a first portion of text (natural language descriptions of the image, paragraph [0045]), the prompt input comprises a second portion of text (input text query, paragraph [0045]), and combining the additional context with the prompt input comprises concatenating the first portion of text and the second portion of text (a natural language prompt comprising the contextual information and query text, paragraph [0046]).
In regard to claim 17, Siegenthaler discloses the step of receiving a multimodal prompt that includes the prompt input and the context input, wherein the multimodal prompt includes data from at least two different modalities (a combined text and image query, paragraph [0038]).
In regard to claim 18, Siegenthaler discloses the context input comprises a portion of domain data corresponding to a first domain of knowledge (the context comprises knowledge domains associated with a user profile, paragraph [0077]).
In regard to claim 19, Siegenthaler discloses at least a portion of the additional context comprises a prompt history associated with a first domain of knowledge (the prompt is enriched with conversation history, paragraph [0046]).
In regard to claim 20, Siegenthaler discloses a system (Fig. 6, 610) comprising:
one or more memories storing instructions (memory subsystem 625); and
one or more processors coupled to the one or more memories (processors 614) that, when executing the instructions, perform the steps of:
generating additional context for a prompt input based on a context input comprising an image (a query comprising an input text and an input image is received, paragraph [0076]), wherein generating the additional context comprises causing an image-based machine learning (ML) model to generate a description of a sub-portion of the image (one or more image models split a single image into sub-images that form semantic units to identify objects, paragraph [0054]; for each object comprising a portion of the image, a description of the object is generated, paragraphs [0079-0080]);
combining the additional context with the prompt input to generate a context-enriched prompt (the system generates an input prompt for an LLM from the text query and natural language descriptors of the input image, paragraph [0086]); and
executing one or more generative machine learning (ML) models on the context-enriched prompt to generate the context-enriched response (the system generates a response to the query by providing the input prompt to a large language model (LLM), paragraph [0087]).
While Siegenthaler discloses an ML model generates a description of sub-portions of an image, Siegenthaler do not expressly disclose receiving a selection of a first sub-portion of an image as a context input.
Yushkina discloses a method for generating additional context comprising receiving a selection of a first sub-portion of an image as a context input (a region search control allows a user to select a first sub-portion of an image as context input for generating a context-enriched response, paragraphs [0050] and [0057-0058]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to receive a selection of a first sub-portion of an image as context input to the method of Siegenthaler, because it would help the system provide a more relevant generated response by allowing the user to specify a particular object in the image, as taught by Yushkina (paragraph [0074]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Abdishektaei et al. disclose a method for generating a textual description of a sub-portion of an image.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRIAN LOUIS ALBERTALLI whose telephone number is (571)272-7616. The examiner can normally be reached M-F 8AM-3PM, 4PM-5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached at 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
BLA 7/22/26
/BRIAN L ALBERTALLI/Primary Examiner, Art Unit 2656