Prosecution Insights
Last updated: October 01, 2026
Application No. 19/041,244

PRODUCT-INCLUSIVE IMAGE ALT TEXT GENERATION

Non-Final OA §103
Filed
Jan 30, 2025
Priority
Jan 30, 2024 — provisional 63/627,046
Examiner
LIU, XIAO
Art Unit
Tech Center
Assignee
Walmart Apollo LLC
OA Round
1 (Non-Final)
88%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 88% — above average
88%
Career Allowance Rate
279 granted / 318 resolved
+27.7% vs TC avg
Moderate +12% lift
Without
With
+12.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
30 currently pending
Career history
349
Total Applications
across all art units

Statute-Specific Performance

§101
7.6%
-32.4% vs TC avg
§103
53.1%
+13.1% vs TC avg
§102
17.4%
-22.6% vs TC avg
§112
16.3%
-23.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 318 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 05/06/2025 has/have been considered by the examiner. Specification Applicant is reminded of the proper language and format for an abstract of the disclosure. The abstract should be in narrative form and generally limited to a single paragraph on a separate sheet within the range of 50 to 150 words in length. The abstract should describe the disclosure sufficiently to assist readers in deciding whether there is a need for consulting the full patent text for details. The language should be clear and concise and should not repeat information given in the title. It should avoid using phrases which can be implied, such as, “The disclosure concerns,” “The disclosure defined by this invention,” “The disclosure describes,” etc. In addition, the form and legal phraseology often used in patent claims, such as “means” and “said,” should be avoided. The abstract of the disclosure is objected to because it has phrases “Other embodiments are disclosed”. A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-5, 7-8, 11-14, 16 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Phung et al (CA 3119236), hereinafter Phung in view of Srivatsan et al (arXiv:2305.14779v2 3 Oct 2023), hereinafter Srivatsan. -Regarding claim 1, Phung discloses a system comprising (FIG. 14; [0012]; [0057]): a processor; and a non-transitory computer-readable media storing computing instructions that, when executed on the processor, cause the processor to perform (Abstract; FIGS. 1-14): receiving, from a user, an image of a product (FIGS. 1-2, 4, 8-9, 11); receiving, from the user, user-submitted logo alt text describing a brand of the product in the image (FIG. 5, product attributes 320; FIG. 8, user 10; FIGS. 2, 5-7, 9A-13; [0036], “ image of the product”; [0039], “… a plurality of attributes … brand, color, style … brand, model, …”; [0042]; [0049], “user 10 enters … or specifies product attributes 320”; [0050], “a product description based on the product attributes”); receiving, from the user, user-submitted image alt text describing the image (FIG. 5, product title 310; FIG. 8, user 10; FIGS. 2, 5-7, 9A-13; [0003], [0036], “a title describing the product …”; [0042], “product description … from product title … or a product’s attributes, or both …”; [0053], “product images … along with the alt-text” ); extracting brand information from the user-submitted logo alt text (FIG. 8, product attributes 320; [0045], “extracts … product attributes …”); extracting product information from the user-submitted image alt text (FIG. 8, product title 310; [0045], “extracts the product title …”); and generating a recommended image alt text describing the image and including the brand information, as extracted, and the product information, as extracted, by using an AI model (FIGS. 7-8, 10-11, 13; [0053]-[0054]). Phung does not disclose generating an instruction prompt that includes the image and the text information associated with the image and generating a recommended image alt text describing the image and the associated text information by querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt. In the same field of endeavor, Srivatsan teaches a method for generating alternative text (alt-text) description for images shared on social media (Srivatsan: Abstract; FIGS. 1-3). Srivatsan further teaches generating an instruction prompt that includes the image and the text information associated with the image and generating a recommended image alt text describing the image and the associated text information by querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt (Srivatsan: FIG.2 PNG media_image1.png 384 773 media_image1.png Greyscale ; Page 5, 4th paragraph, “… obtain the complete prefix p ⊕ t. We can condition on this, treating it as … a multimodal prompt to our language model …”). Therefore, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to combine the teaching of Phung with the teaching of Srivatsan by using generating an instruction prompt and querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt in order to provide more literally descriptive and context-specific alternative text description for the image (Srivatsan: Abstract). -Regarding claim 11, Phung discloses a method implemented via execution of computing instructions configured to run at a processor, the method comprising (Abstract; FIGS. 1-14): receiving, from a user, an image of a product (FIGS. 1-2, 4, 8-9, 11); receiving, from the user, user-submitted logo alt text describing a brand of the product in the image (FIG. 5, product attributes 320; FIG. 8, user 10; FIGS. 2, 5-7, 9A-13; [0036], “ image of the product”; [0039], “… a plurality of attributes … brand, color, style … brand, model, …”; [0042]; [0049], “user 10 enters … or specifies product attributes 320”; [0050], “a product description based on the product attributes”); receiving, from the user, user-submitted image alt text describing the image (FIG. 5, product title 310; FIG. 8, user 10; FIGS. 2, 5-7, 9A-13; [0003], [0036], “a title describing the product …”; [0042], “product description … from product title … or a product’s attributes, or both …”; [0053], “product images … along with the alt-text” ); extracting brand information from the user-submitted logo alt text (FIG. 8, product attributes 320; [0045], “extracts … product attributes …”); extracting product information from the user-submitted image alt text (FIG. 8, product title 310; [0045], “extracts the product title …”); and generating a recommended image alt text describing the image and including the brand information, as extracted, and the product information, as extracted, by using an AI model (FIGS. 7-8, 10-11, 13; [0053]-[0054]). Phung does not disclose generating an instruction prompt that includes the image and the text information associated with the image, generating a recommended image alt text describing the image and the associated text information by querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt, and validating the recommended image alt text generated by the multimodal GenAI model. In the same field of endeavor, Srivatsan teaches a method for generating alternative text (alt-text) description for images shared on social media (Srivatsan: Abstract; FIGS. 1-3). Srivatsan further teaches generating an instruction prompt that includes the image and the text information associated with the image and generating a recommended image alt text describing the image and the associated text information by querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt (Srivatsan: FIG.2; Page 5, 4th paragraph, “… obtain the complete prefix p ⊕ t. We can condition on this, treating it as … a multimodal prompt to our language model …”), and validating the recommended image alt text generated by the multimodal GenAI model (Srivatsan: Table 1; Secs. 5-6). Therefore, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to combine the teaching of Phung with the teaching of Srivatsan by using generating an instruction prompt and querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt in order to provide more literally descriptive and context-specific alternative text description for the image (Srivatsan: Abstract). -Regarding claim 19, Phung discloses a non-transitory computer readable storage medium storing computing instructions, the computing instructions, when run on a processor, causing the processor to perform operations comprising (Abstract; FIGS. 1-14): receiving, from a user, an image of a product (FIGS. 1-2, 4, 8-9, 11); receiving, from the user, user-submitted logo alt text describing a brand of the product in the image (FIG. 5, product attributes 320; FIG. 8, user 10; FIGS. 2, 5-7, 9A-13; [0036], “ image of the product”; [0039], “… a plurality of attributes … brand, color, style … brand, model, …”; [0042]; [0049], “user 10 enters … or specifies product attributes 320”; [0050], “a product description based on the product attributes”); receiving, from the user, user-submitted image alt text describing the image (FIG. 5, product title 310; FIG. 8, user 10; FIGS. 2, 5-7, 9A-13; [0003], [0036], “a title describing the product …”; [0042], “product description … from product title … or a product’s attributes, or both …”; [0053], “product images … along with the alt-text” ); extracting brand information from the user-submitted logo alt text (FIG. 8, product attributes 320; [0045], “extracts … product attributes …”); extracting product information from the user-submitted image alt text (FIG. 8, product title 310; [0045], “extracts the product title …”); and generating a recommended image alt text describing the image and including the brand information, as extracted, and the product information, as extracted, by using an AI model (FIGS. 7-8, 10-11, 13; [0053]-[0054]). Phung does not disclose generating an instruction prompt that includes the image and the text information associated with the image, generating a recommended image alt text describing the image and the associated text information by querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt, and validating the recommended image alt text generated by the multimodal GenAI model by: comparing the recommended image alt text generated by the multimodal GenAI model to the user-submitted image alt text to generate a comparison result; and selecting one of the recommended image alt text or the user-submitted image alt text, based on the comparison result. In the same field of endeavor, Srivatsan teaches a method for generating alternative text (alt-text) description for images shared on social media (Srivatsan: Abstract; FIGS. 1-3). Srivatsan further teaches generating an instruction prompt that includes the image and the text information associated with the image and generating a recommended image alt text describing the image and the associated text information by querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt (Srivatsan: FIG.2; Page 5, 4th paragraph, “… obtain the complete prefix p ⊕ t. We can condition on this, treating it as … a multimodal prompt to our language model …”), and validating the recommended image alt text generated by the multimodal GenAI model (Srivatsan: Tables 1-2; Secs. 5-6) by: comparing the recommended image alt text generated by the multimodal GenAI model to the user-submitted image alt text to generate a comparison result; and selecting one of the recommended image alt text or the user-submitted image alt text, based on the comparison result (Srivatsan: FIG. 3; Tables 1-2; Secs. 5-6; Page 2, last paragraph; Page 3, last paragraph, “ranking or selecting …”). Therefore, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to combine the teaching of Phung with the teaching of Srivatsan by using generating an instruction prompt and querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt in order to provide more literally descriptive and context-specific alternative text description for the image (Srivatsan: Abstract). -Regarding claims 2 and 12, Phung in view of Srivatsan teaches the system of claim 1 and the method of claim 11. The combination further teaches receiving the user-submitted image alt text including any one or more of product type, product category, product size, product style, product quantity, product cost, product weight, product color, product shape, product specifications, product description, related product information, and product promotional information (Phung: FIGS. 9A-9B). -Regarding claims 3 and 13, Phung in view of Srivatsan teaches the system of claim 1 and the method of claim 11. The combination further teaches receiving an advertising image for sale of the product on a website, the advertising image optionally including additional image features in addition to the product (Phung: FIGS. 4, 6-7; [0037]). -Regarding claim 4, Phung in view of Srivatsan teaches the system of claim 1. The combination further teaches receiving the image of the product without any visible brand information in the image of the product (Phung: FIG. 4; [0037]). -Regarding claims 5 and 14, Phung in view of Srivatsan teaches the system of claim 1 and the method of claim 11. Phung does not disclose validating the recommended image alt text generated by the multimodal GenAI model by: comparing the recommended image alt text generated by the multimodal GenAI model to the user-submitted image alt text to generate a comparison result; and selecting one of the recommended image alt text or the user-submitted image alt text, based on the comparison result. In the same field of endeavor, Srivatsan teaches a method for generating alternative text (alt-text) description for images shared on social media (Srivatsan: Abstract; FIGS. 1-3). Srivatsan further teaches validating the recommended image alt text generated by the multimodal GenAI model (Srivatsan: Tables 1-2; Secs. 5-6) by: comparing the recommended image alt text generated by the multimodal GenAI model to the user-submitted image alt text to generate a comparison result; and selecting one of the recommended image alt text or the user-submitted image alt text, based on the comparison result (Srivatsan: FIG. 3; Tables 1-2; Secs. 5-6; Page 2, last paragraph; Page 3, last paragraph, “ranking or selecting …”). Therefore, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to combine the teaching of Phung with the teaching of Srivatsan by using generating an instruction prompt and querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt in order to provide more literally descriptive and context-specific alternative text description for the image (Srivatsan: Abstract). -Regarding claims 7 and 16, Phung in view of Srivatsan teaches the system of claim 1 and the method of claim 11. The combination further teaches post-processing of the recommended image alt text to improve any one or more of readability, searchability, and accuracy of the recommended image alt text (Phung: FIGS. 8, 11, post processor 220; [0049], page 11, lines 3-10; [0054], page 13, lines 3-10, “post-processed … allowing the product to be located … via a search engine …”). -Regarding claim 8, Phung in view of Srivatsan teaches the system of claim 1. Phung does not disclose training the multimodal GenAI model on pairs of model images and pre-approved image alt text associated therewith, before querying of the multimodal GenAI model with the instruction prompt. In the same field of endeavor, Srivatsan teaches a method for generating alternative text (alt-text) description for images shared on social media (Srivatsan: Abstract; FIGS. 1-3). Srivatsan further teaches training the multimodal GenAI model on pairs of model images and pre-approved image alt text associated therewith, before querying of the multimodal GenAI model with the instruction prompt (Srivatsan: Page 4, Sec. 3.). Therefore, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to combine the teaching of Phung with the teaching of Srivatsan by using generating an instruction prompt and querying a multimodal generative artificial intelligence (multimodal GenAI) model with the instruction prompt in order to provide more literally descriptive and context-specific alternative text description for the image (Srivatsan: Abstract). Claim(s) 6, 15 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Phung et al (CA 3119236), hereinafter Phung in view of Srivatsan et al (arXiv:2305.14779v2 3 Oct 2023), hereinafter Srivatsan, and further in view of Musundi et al (US 20250094720 A1), hereinafter Musundi. -Regarding claims 6, 15 and 20, Phung in view of Srivatsan teaches the system of claim 5, the method of claim 14, and non-transitory computer readable storage medium of claim 19. Phung in view of Srivatsan does not teach selecting the recommended image alt text for recommendation to the user when the comparison result indicates the number of differences between the recommended image alt text and the user-submitted image alt text exceeds a threshold value; or selecting the user-submitted image alt text for recommendation to the user when the comparison result indicates the number of differences between the recommended image alt text and the user-submitted image alt text does not exceed the threshold value. However, Musundi is an analogous art pertinent to the problem to be solved in this application and teaches a method for validation of alt text for images in web pages includes extracting image data from the web pages, the image data including source data and alt text data for a plurality of image elements in the web pages (Musundi: Abstract; FIGS. 1-6). Musundi further teaches selecting the recommended image alt text for recommendation to the user when the comparison result indicates the number of differences between the recommended image alt text and the user-submitted image alt text exceeds a threshold value; or selecting the user- submitted image alt text for recommendation to the user when the comparison result indicates the number of differences between the recommended image alt text and the user-submitted image alt text does not exceed the threshold value (Musundi: FIG. 1, alt text validation application 128; FIGS. 2-4; [0025]; [0019], “Depending on the assessed accuracy, the system may also be configured to generate an alt text suggestion for the image element, for example, if the accuracy of the current alt text is below a threshold accuracy level”; [0026]-[0027]; [0036], “when the similarity score for the alt text of an image element is below a predefined threshold, the system generates new alt text which can be presented to a user as a suggestion for modifying or replacing the inaccurate alt text”; [0037]). Therefore, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify the teaching of Phung in view of Srivatsan with the teaching of Musundi by comparing differences between the recommended image alt text and the user-submitted image alt text with a predetermined threshold in order to generate an alt text with predetermined accuracy. Allowable Subject Matter Claims 9-10 and 17-18 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to XIAO LIU whose telephone number is (571)272-4539. The examiner can normally be reached Monday-Thursday and Alternate Fridays 8:30-4:30. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Mehmood can be reached at (571) 272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /XIAO LIU/Primary Examiner, Art Unit 2664
Read full office action

Prosecution Timeline

Jan 30, 2025
Application Filed
Sep 24, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731389
VIDEO-BASED SURGICAL SKILL ASSESSMENT USING TOOL TRACKING
3y 5m to grant Granted Sep 08, 2026
Patent 12731415
SYSTEMS AND METHODS FOR DETECTING A SOFT POINT ON A ROAD USING A HARD POINT
2y 7m to grant Granted Sep 08, 2026
Patent 12730190
OBJECT DETECTION AND CLASSIFICATION USING LIDAR RANGE IMAGES FOR AUTONOMOUS MACHINE APPLICATIONS
2y 9m to grant Granted Sep 08, 2026
Patent 12726581
METHOD FOR REPRESENTING A HARMONIZED OBSCURED AREA OF AN ENVIRONMENT OF A MOBILE PLATFORM
4y 9m to grant Granted Sep 01, 2026
Patent 12725409
Deep Learning for Electromagnetic Imaging of Stored Commodities
2y 11m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
88%
Grant Probability
99%
With Interview (+12.0%)
2y 6m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 318 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month