DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 05/06/2025 has/have been considered by the examiner.
Specification
Applicant is reminded of the proper language and format for an abstract of the disclosure.
The abstract should be in narrative form and generally limited to a single paragraph on a separate sheet within the range of 50 to 150 words in length. The abstract should describe the disclosure sufficiently to assist readers in deciding whether there is a need for consulting the full patent text for details.
The language should be clear and concise and should not repeat information given in the title. It should avoid using phrases which can be implied, such as, “The disclosure concerns,” “The disclosure defined by this invention,” “The disclosure describes,” etc. In addition, the form and legal phraseology often used in patent claims, such as “means” and “said,” should be avoided.
The abstract of the disclosure is objected to because it has phrases “Other embodiments are disclosed”. A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-5, 7-8, 11-14, 16 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Phung et al (CA 3119236), hereinafter Phung in view of Srivatsan et al (arXiv:2305.14779v2 3 Oct 2023), hereinafter Srivatsan.
-Regarding claim 1, Phung discloses a system comprising (FIG. 14; [0012]; [0057]): a processor; and a non-transitory computer-readable media storing computing instructions that, when executed on the processor, cause the processor to perform (Abstract; FIGS. 1-14): receiving, from a user, an image of a product (FIGS. 1-2, 4, 8-9, 11); receiving, from the user, user-submitted logo alt text describing a brand of the product in the image (FIG. 5, product attributes 320; FIG. 8, user 10; FIGS. 2, 5-7, 9A-13; [0036], “ image of the product”; [0039], “… a plurality of attributes … brand, color, style … brand, model, …”; [0042]; [0049], “user 10 enters … or specifies product attributes 320”; [0050], “a product description based on the product attributes”); receiving, from the user, user-submitted image alt text describing the image (FIG. 5, product title 310; FIG. 8, user 10; FIGS. 2, 5-7, 9A-13; [0003], [0036], “a title describing the product …”; [0042], “product description … from product title … or a product’s attributes, or both …”; [0053], “product images … along with the alt-text” ); extracting brand information from the user-submitted logo alt text (FIG. 8, product attributes 320; [0045], “extracts … product attributes …”); extracting product information from the user-submitted image alt text (FIG. 8, product title 310; [0045], “extracts the product title …”); and generating a recommended image alt text describing the image and including the brand information, as extracted, and the product information, as extracted, by using an AI model (FIGS. 7-8, 10-11, 13; [0053]-[0054]).
Phung does not disclose generating an instruction prompt that includes the image and the text information associated with the image and generating a recommended image alt text describing the image and the associated text information by querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt.
In the same field of endeavor, Srivatsan teaches a method for generating alternative text (alt-text) description for images shared on social media (Srivatsan: Abstract; FIGS. 1-3). Srivatsan further teaches generating an instruction prompt that includes the image and the text information associated with the image and generating a recommended image alt text describing the image and the associated text information by querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt (Srivatsan: FIG.2
PNG
media_image1.png
384
773
media_image1.png
Greyscale
; Page 5, 4th paragraph, “… obtain the complete prefix p ⊕ t. We can condition on this, treating it as … a multimodal prompt to our language model …”).
Therefore, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to combine the teaching of Phung with the teaching of Srivatsan by using generating an instruction prompt and querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt in order to provide more literally descriptive and context-specific alternative text description for the image (Srivatsan: Abstract).
-Regarding claim 11, Phung discloses a method implemented via execution of computing instructions configured to run at a processor, the method comprising (Abstract; FIGS. 1-14): receiving, from a user, an image of a product (FIGS. 1-2, 4, 8-9, 11); receiving, from the user, user-submitted logo alt text describing a brand of the product in the image (FIG. 5, product attributes 320; FIG. 8, user 10; FIGS. 2, 5-7, 9A-13; [0036], “ image of the product”; [0039], “… a plurality of attributes … brand, color, style … brand, model, …”; [0042]; [0049], “user 10 enters … or specifies product attributes 320”; [0050], “a product description based on the product attributes”); receiving, from the user, user-submitted image alt text describing the image (FIG. 5, product title 310; FIG. 8, user 10; FIGS. 2, 5-7, 9A-13; [0003], [0036], “a title describing the product …”; [0042], “product description … from product title … or a product’s attributes, or both …”; [0053], “product images … along with the alt-text” ); extracting brand information from the user-submitted logo alt text (FIG. 8, product attributes 320; [0045], “extracts … product attributes …”); extracting product information from the user-submitted image alt text (FIG. 8, product title 310; [0045], “extracts the product title …”); and generating a recommended image alt text describing the image and including the brand information, as extracted, and the product information, as extracted, by using an AI model (FIGS. 7-8, 10-11, 13; [0053]-[0054]).
Phung does not disclose generating an instruction prompt that includes the image and the text information associated with the image, generating a recommended image alt text describing the image and the associated text information by querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt, and validating the recommended image alt text generated by the multimodal GenAI model.
In the same field of endeavor, Srivatsan teaches a method for generating alternative text (alt-text) description for images shared on social media (Srivatsan: Abstract; FIGS. 1-3). Srivatsan further teaches generating an instruction prompt that includes the image and the text information associated with the image and generating a recommended image alt text describing the image and the associated text information by querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt (Srivatsan: FIG.2; Page 5, 4th paragraph, “… obtain the complete prefix p ⊕ t. We can condition on this, treating it as … a multimodal prompt to our language model …”), and validating the recommended image alt text generated by the multimodal GenAI model (Srivatsan: Table 1; Secs. 5-6).
Therefore, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to combine the teaching of Phung with the teaching of Srivatsan by using generating an instruction prompt and querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt in order to provide more literally descriptive and context-specific alternative text description for the image (Srivatsan: Abstract).
-Regarding claim 19, Phung discloses a non-transitory computer readable storage medium storing computing instructions, the computing instructions, when run on a processor, causing the processor to perform operations comprising (Abstract; FIGS. 1-14): receiving, from a user, an image of a product (FIGS. 1-2, 4, 8-9, 11); receiving, from the user, user-submitted logo alt text describing a brand of the product in the image (FIG. 5, product attributes 320; FIG. 8, user 10; FIGS. 2, 5-7, 9A-13; [0036], “ image of the product”; [0039], “… a plurality of attributes … brand, color, style … brand, model, …”; [0042]; [0049], “user 10 enters … or specifies product attributes 320”; [0050], “a product description based on the product attributes”); receiving, from the user, user-submitted image alt text describing the image (FIG. 5, product title 310; FIG. 8, user 10; FIGS. 2, 5-7, 9A-13; [0003], [0036], “a title describing the product …”; [0042], “product description … from product title … or a product’s attributes, or both …”; [0053], “product images … along with the alt-text” ); extracting brand information from the user-submitted logo alt text (FIG. 8, product attributes 320; [0045], “extracts … product attributes …”); extracting product information from the user-submitted image alt text (FIG. 8, product title 310; [0045], “extracts the product title …”); and generating a recommended image alt text describing the image and including the brand information, as extracted, and the product information, as extracted, by using an AI model (FIGS. 7-8, 10-11, 13; [0053]-[0054]).
Phung does not disclose generating an instruction prompt that includes the image and the text information associated with the image, generating a recommended image alt text describing the image and the associated text information by querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt, and validating the recommended image alt text generated by the multimodal GenAI model by: comparing the recommended image alt text generated by the multimodal GenAI model to the user-submitted image alt text to generate a comparison result; and selecting one of the recommended image alt text or the user-submitted image alt text, based on the comparison result.
In the same field of endeavor, Srivatsan teaches a method for generating alternative text (alt-text) description for images shared on social media (Srivatsan: Abstract; FIGS. 1-3). Srivatsan further teaches generating an instruction prompt that includes the image and the text information associated with the image and generating a recommended image alt text describing the image and the associated text information by querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt (Srivatsan: FIG.2; Page 5, 4th paragraph, “… obtain the complete prefix p ⊕ t. We can condition on this, treating it as … a multimodal prompt to our language model …”), and validating the recommended image alt text generated by the multimodal GenAI model (Srivatsan: Tables 1-2; Secs. 5-6) by: comparing the recommended image alt text generated by the multimodal GenAI model to the user-submitted image alt text to generate a comparison result; and selecting one of the recommended image alt text or the user-submitted image alt text, based on the comparison result (Srivatsan: FIG. 3; Tables 1-2; Secs. 5-6; Page 2, last paragraph; Page 3, last paragraph, “ranking or selecting …”).
Therefore, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to combine the teaching of Phung with the teaching of Srivatsan by using generating an instruction prompt and querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt in order to provide more literally descriptive and context-specific alternative text description for the image (Srivatsan: Abstract).
-Regarding claims 2 and 12, Phung in view of Srivatsan teaches the system of claim 1 and the method of claim 11. The combination further teaches receiving the user-submitted image alt text including any one or more of product type, product category, product size, product style, product quantity, product cost, product weight, product color, product shape, product specifications, product description, related product information, and product promotional information (Phung: FIGS. 9A-9B).
-Regarding claims 3 and 13, Phung in view of Srivatsan teaches the system of claim 1 and the method of claim 11. The combination further teaches receiving an advertising image for sale of the product on a website, the advertising image optionally including additional image features in addition to the product (Phung: FIGS. 4, 6-7; [0037]).
-Regarding claim 4, Phung in view of Srivatsan teaches the system of claim 1. The combination further teaches receiving the image of the product without any visible brand information in the image of the product (Phung: FIG. 4; [0037]).
-Regarding claims 5 and 14, Phung in view of Srivatsan teaches the system of claim 1 and the method of claim 11.
Phung does not disclose validating the recommended image alt text generated by the multimodal GenAI model by: comparing the recommended image alt text generated by the multimodal GenAI model to the user-submitted image alt text to generate a comparison result; and selecting one of the recommended image alt text or the user-submitted image alt text, based on the comparison result.
In the same field of endeavor, Srivatsan teaches a method for generating alternative text (alt-text) description for images shared on social media (Srivatsan: Abstract; FIGS. 1-3). Srivatsan further teaches validating the recommended image alt text generated by the multimodal GenAI model (Srivatsan: Tables 1-2; Secs. 5-6) by: comparing the recommended image alt text generated by the multimodal GenAI model to the user-submitted image alt text to generate a comparison result; and selecting one of the recommended image alt text or the user-submitted image alt text, based on the comparison result (Srivatsan: FIG. 3; Tables 1-2; Secs. 5-6; Page 2, last paragraph; Page 3, last paragraph, “ranking or selecting …”).
Therefore, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to combine the teaching of Phung with the teaching of Srivatsan by using generating an instruction prompt and querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt in order to provide more literally descriptive and context-specific alternative text description for the image (Srivatsan: Abstract).
-Regarding claims 7 and 16, Phung in view of Srivatsan teaches the system of claim 1 and the method of claim 11. The combination further teaches post-processing of the recommended image alt text to improve any one or more of readability, searchability, and accuracy of the recommended image alt text (Phung: FIGS. 8, 11, post processor 220; [0049], page 11, lines 3-10; [0054], page 13, lines 3-10, “post-processed … allowing the product to be located … via a search engine …”).
-Regarding claim 8, Phung in view of Srivatsan teaches the system of claim 1.
Phung does not disclose training the multimodal GenAI model on pairs of model images and pre-approved image alt text associated therewith, before querying of the multimodal GenAI model with the instruction prompt.
In the same field of endeavor, Srivatsan teaches a method for generating alternative text (alt-text) description for images shared on social media (Srivatsan: Abstract; FIGS. 1-3). Srivatsan further teaches training the multimodal GenAI model on pairs of model images and pre-approved image alt text associated therewith, before querying of the multimodal GenAI model with the instruction prompt (Srivatsan: Page 4, Sec. 3.).
Therefore, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to combine the teaching of Phung with the teaching of Srivatsan by using generating an instruction prompt and querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt in order to provide more literally descriptive and context-specific alternative text description for the image (Srivatsan: Abstract).
Claim(s) 6, 15 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Phung et al (CA 3119236), hereinafter Phung in view of Srivatsan et al (arXiv:2305.14779v2 3 Oct 2023), hereinafter Srivatsan, and further in view of Musundi et al (US 20250094720 A1), hereinafter Musundi.
-Regarding claims 6, 15 and 20, Phung in view of Srivatsan teaches the system of claim 5, the method of claim 14, and non-transitory computer readable storage medium of claim 19.
Phung in view of Srivatsan does not teach selecting the recommended image alt text for recommendation to the user when the comparison result indicates the number of differences between the recommended image alt text and the user-submitted image alt text exceeds a threshold value; or selecting the user-submitted image alt text for recommendation to the user when the comparison result indicates the number of differences between the recommended image alt text and the user-submitted image alt text does not exceed the threshold value.
However, Musundi is an analogous art pertinent to the problem to be solved in this application and teaches a method for validation of alt text for images in web pages includes extracting image data from the web pages, the image data including source data and alt text data for a plurality of image elements in the web pages (Musundi: Abstract; FIGS. 1-6). Musundi further teaches selecting the recommended image alt text for recommendation to the user when the comparison result indicates the number of differences between the recommended image alt text and the user-submitted image alt text exceeds a threshold value; or selecting the user- submitted image alt text for recommendation to the user when the comparison result indicates the number of differences between the recommended image alt text and the user-submitted image alt text does not exceed the threshold value (Musundi: FIG. 1, alt text validation application 128; FIGS. 2-4; [0025]; [0019], “Depending on the assessed accuracy, the system may also be configured to generate an alt text suggestion for the image element, for example, if the accuracy of the current alt text is below a threshold accuracy level”; [0026]-[0027]; [0036], “when the similarity score for the alt text of an image element is below a predefined threshold, the system generates new alt text which can be presented to a user as a suggestion for modifying or replacing the inaccurate alt text”; [0037]).
Therefore, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify the teaching of Phung in view of Srivatsan with the teaching of Musundi by comparing differences between the recommended image alt text and the user-submitted image alt text with a predetermined threshold in order to generate an alt text with predetermined accuracy.
Allowable Subject Matter
Claims 9-10 and 17-18 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to XIAO LIU whose telephone number is (571)272-4539. The examiner can normally be reached Monday-Thursday and Alternate Fridays 8:30-4:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Mehmood can be reached at (571) 272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/XIAO LIU/Primary Examiner, Art Unit 2664