Prosecution Insights
Last updated: October 02, 2026
Application No. 18/505,916

AUTOMATED IMAGE CAPTIONING BASED ON COMPUTER VISION AND NATURAL LANGUAGE PROCESSING

Final Rejection §101§102
Filed
Nov 09, 2023
Examiner
GORADIA, SHEFALI DINESH
Art Unit
2676
Tech Center
2600 — Communications
Assignee
Snap Inc.
OA Round
2 (Final)
90%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 90% — above average
90%
Career Allowance Rate
558 granted / 618 resolved
+28.3% vs TC avg
Moderate +11% lift
Without
With
+11.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
22 currently pending
Career history
637
Total Applications
across all art units

Statute-Specific Performance

§101
17.1%
-22.9% vs TC avg
§103
36.3%
-3.7% vs TC avg
§102
24.8%
-15.2% vs TC avg
§112
12.5%
-27.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 618 resolved cases

Office Action

§101 §102
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The amendment was filed on 6/30/2026. Claims 1-6, 8-12, and 15-19 are pending. Claims 7, 13, and 20 are canceled. Claim 14 is withdrawn. Response to Arguments Applicants’ arguments filed under Remarks on pages 7-11 with respect to rejections under 35 USC 101 on 6/30/2026 have been fully considered but they are not persuasive. Applicants state on page 8 that: PNG media_image1.png 444 690 media_image1.png Greyscale The Examiner respectfully disagrees. The claim recites “providing the prompt to a…model”. This step simply provides the prompt. There is no processing or any other detail. Then the step of “generating one or more caption options based on the…model and the prompt”, is interpreted to have been possibly be performed by a human with an aid of pen and paper for example as the apparatus doing the ‘generating’. The claim nowhere recites that this so called model is anything special or performing any special steps to achieve the ‘generating’. Once a prompt of provided to a model, such as, “draw me a picture of a dog”, caption can be ‘generated’ based on a model (person using the natural language) and the prompt itself. Paragraph [0019] of the pending specification states “[0019] In some embodiments, the user can provide an input to select a desired tone or style for the captions, such as funny, serious, or sarcastic. The prompt can reflect this selected tone so the language model generates caption options in an appropriate style.” Per this, once a prompt is provided, one can generate a picture of a dog or a caption option related to that as either funny or serious, for example. On page 9 of the Remarks, the applicant argues by stating: PNG media_image2.png 662 661 media_image2.png Greyscale PNG media_image3.png 51 632 media_image3.png Greyscale The Examiner respectfully disagrees. Example 37, ‘determining’ step is nothing like ‘generating’ steps in the pending claims. Claim 1 of Example 37 was found to have been integrated into a practical application under step 2A prong 2 because “the additional elements recite a specific manner of automatically displaying icons to the user based on usage which provides a specific improvement over prior systems, resulting in an improved user interface for electronic devices”. Graphical User Interface (GUI) is a specific structure, the receiving was done via a GUI and moving of the most used icons to a position were performed on the GUI. There is no structure recited in pending claims that aligns with GUI or alike. The receiving request simply comprises a selection of a graphical icon, and the prompt generated is in response to the request that comprises the selection of the icon. There is no arrangement being done, like in Example 37. There is no GUI being claimed. In the pending system claim 8, there is simply a generic processor and a memory. Hence, the pending claims are not similar to the Example 37 of the 2019 PEG subject mattery eligibility example: abstract ideas. On page 10 of the Remarks, the applicant argues by stating: PNG media_image4.png 306 659 media_image4.png Greyscale The Examiner respectfully disagrees. As discussed with respect to Step 2A Prong Two, and Example 37 of the 2019 PEG, the additional element in the claim amounts to no more than mere instructions to apply the exception using a generic computer component. The same analysis applies here in Step 2B, i.e., mere instructions to apply an exception using a generic computer component cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. The recitation of additional elements in the pending claims of “receiving” and “generating” steps, does not provide any indication that the client device, memory, processor, etc. is anything other than a generic, off the-shelf computer components. Thus, even when viewed as a whole, nothing in the claim adds significantly more (i.e., an inventive concept) to the abstract idea. On page 12 of the Remarks, regarding the rejections under 35 USC 102, the Applicant argues by stating: PNG media_image5.png 800 823 media_image5.png Greyscale The Examiner respectfully disagrees. Paragraph [0090] of Valliani states: [0092] In some embodiments, presentation component 218 generates user interface features associated with a caption. Such features can include interface elements (such as graphics buttons, sliders, menus, audio prompts, alerts, alarms, vibrations, pop-up windows, notification-bar or status-bar items, in-app notifications, or other similar features for interfacing with a user), queries, and prompts. For example, presentation component 218 may query the user regarding user preferences for captions, such as asking the user “Keep showing you similar captions in the future?” or “Please rate the accuracy of this caption from 1-5 . . . .” Some embodiments of presentation component 218 capture user responses (e.g., modifications) to captions or user activity associated with captions (e.g., sharing, saving, dismissing, deleting). [Emphasis added]. This paragraph of Valliani shows that the request is received from a user by component 218 to re-generate/modify/update caption and said request is received by user of a graphic button from the interface, as claimed in the amended pending claim. This prompt generated eventually is in response to the user selecting graphic button, as disclosed in this paragraph by Valliani. Therefore, the rejection stands. Information Disclosure Statement The information disclosure statement (IDS) submitted on 6/30/2026 has been considered by the examiner. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-6, 8-12 and 15-19 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., abstract idea – mental process) without significantly more. Claim 1 is used as an example. Claims 8 and 15 recite a system and non-transitory computer-readable medium, respectively, having a memory and a processor. The two-part test to identify claims that are directed to a judicial exception (Step 2A) and to then evaluate if additional elements of the claim provide an inventive concept (Step 2B) are: (1) Are the claims directed to a process, machine, manufacture or composition of matter; (2A) Prong One: Are the claims directed to a judicially recognized exception, i.e., a law of nature, a natural phenomenon, or an abstract idea; Prong Two: If the claims are directed to a judicial exception under Prong One, then is the judicial exception integrated into a practical application; (2B) If the claims are directed to a judicial exception and do not integrate the judicial exception, do the claims provide an inventive concept. Claim 1. A method comprising: (a) receiving, from a client device, image data; (b) detecting one or more objects depicted by the image data; (c) receiving a request to generate a caption, the request comprising a selection of a graphic icon presented among a set of graphical icons; (d) generating a prompt based on the one or more objects detected within the image data, the prompt generated responsive to the request that comprises the election of the graphical icon; (e) providing the prompt to a natural language processing model; (f) generating one or more caption options based on the natural language processing model and the prompt; and (g) causing display of a presentation of the one or more caption options at the client device. [emphasis added]. With regard to (1), the instant claims recite an apparatus and a method, therefore the answer is "yes". With regard to (2A), Prong One: Yes. When viewed under the broadest most reasonable interpretation, the instant claims are directed to a Judicial Exception – an abstract idea belonging to the group of mental process – concepts that are practicably performed in the human mind (including an observation, evaluation, judgement, opinion). The steps of (b), (d), and (f) (above in emphasized claim 1) are generically recited and nothing in these steps precludes the steps from practically being performed by a human equipped with an appropriate apparatus. It can be interpreted as merely looking at the image/data and determining a region/object in the image and writing out a ‘prompt’ related to or describes that region/object. There is nothing in the claim that requires more than an operation that a human, armed with the appropriate apparatus, pen and a paper, cannot perform. The detecting and generating, under its broadest reasonable interpretation, covers performance of the limitation in the mind. The claim encompasses the user looking at data/image once the image is received, region/object such as data with location/time/user information, etc. of a section of the image can be determined. This way, essentially one can present/output information about the section of an image that represents that context. Thus, these limitations are a mental process. With regard to (2A), Prong Two: No. The instant claims do not apply, rely on, or use the judicial exception in a manner that imposes a meaningful limit on the judicial exception of (a)/(c) “receiving”, (e) “providing”, and (g) “causing display”, and therefore does not integrate the judicial exception into a practical application. The use of a system/memory/processor to receive an image (i.e., “data”) at a high level of generality such that said “data” can be used in the operation of the recited judicial exception (the mental step of “receiving”). Providing/supplying/receiving “data” does not provide for “integration” of the abstract idea into a practical application, as said data do not change the way in which said system operates. Receiving image data and a request to generate a caption is nothing more than a user selecting a button the screen to do what the button is designed to do. There are no specifics on how the data is received or how the request is received, other than a user selecting a button on a screen, for example. This can be interpreted as “visualization”. Even if this step is by a “processor” that may be, for example, a camera, or, by a screen having button/icon. A camera/sensor/screen having options is well known in the field, and receiving data from a camera/sensor is also well known. Even if this step of “providing” the prompt to a natural language processing model, that may be, for example, a human/person. Therefore, once a data/image is provided to a person, a person is able to generate a caption describing what’s in the image or something about said object depicted by the image. This generic processor limitation is no more than mere instructions to apply the exception using a generic computer component. The ‘causing display of presentation’ is simply outputting/writing out a caption/description about the image or object depicted by the image. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. In conclusion, the claim as a whole does not provide for “integration” of the abstract idea into a practical application. The claim is directed to the abstract idea. With regard to (2B), as discussed with respect to Step 2A Prong Two, the additional element in the claim amounts to no more than mere instructions to apply the exception using a generic computer component. The same analysis applies here, i.e., mere instructions to apply an exception using a generic computer component cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. The pending claims do not show what is more than a routine in the art presented in the claims, i.e., the additional elements are nothing more than routine and well-known steps. There is no improvement to technology here. There are only steps of (b), (d), and (f) with additional elements of (a)/(b) “receiving”, (e) “providing”, and (g) “causing display” and it has not been shown that the mental process allows the “technology” to do something that it previously was not able to do. Therefore, the claims 1, 8, and 15 are ineligible under 35 USC 101. With regard to dependent claims 2-6, 9-12, and 16-19, similar analysis is applied and therefore does not integrate the judicial exception into a practical application – does not provide significant more than the judicial exception. These claims are similarly rejected for the same reasons discussed in view of steps recited in claim 1 and not repeated herewith. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-6, 8-12 and 15-19 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by US 2017/0132821 to Valliani et al. (hereafter, “Valliani”). With regard to claim 1 Valliani discloses a method comprising: receiving, from a client device, image data (paragraphs [0068, 0070] where image data is received); detecting one or more objects depicted by the image data (detecting context by extractor 264, paragraphs [0069, 0076-0078); receiving a request to generate a caption, the request comprising a selection of a graphic icon presented among a set of graphical icons (paragraph [0092]); generating a prompt based on the one or more objects detected within the image data, the prompt generated responsive to the request that comprises the election of the graphical icon (paragraphs [0079, 0092]); providing the prompt to a natural language processing model (paragraph [0085]); generating one or more caption options based on the natural language processing model and the prompt (paragraph [0085]); and causing display of a presentation of the one or more caption options at the client device (paragraph [0092]). With regard to claim 2 Valliani discloses wherein the generating the prompt further comprises: accessing contextual data at the client device; and generating the prompt based on the one or more objects detected within the image data and the contextual data (paragraphs [0022, 0040, 0100] where “using user data from other users (i.e., crowdsourcing data) for determining typical user media sharing and caption patterns for events of similar types”). With regard to claim 3 Valliani discloses wherein the contextual data includes one or more of: location data; temporal data that indicates a time of day; and user profile data (paragraph [0040] where “if many people in a particular location on a particular day are sharing images, then a media-sharing event may be detected and captions automatically generated when a user takes a picture at the location on the particular day.”). With regard to claim 4 Valliani discloses wherein the generating the prompt further comprises: receiving an input that defines a tone; and generating the prompt based on the one or more objects detected within the image data and the tone defined by the input (paragraphs [0032-0035] where a tone can be interpreted as a demographic information or a particular scenario; also, paragraphs [0056-0059, 0080-0081]). With regard to claim 5 Valliani discloses wherein the causing display of the presentation of the one or more caption options at the client device further comprises: determining a ranking of the one or more caption options; and causing display of the presentation of the one or more caption options based on the ranking (paragraph [0056] where “first event may be correlated with variables of a second event to identify in-common variables for determining a likely pattern. For example, where a first event comprises, a user posting a digital image of food with a caption from a restaurant on a first Saturday and a second event comprises user posting a digital image with a caption from a different restaurant on the following Saturday, a pattern may be determined that the user posts pictures taken in a restaurant on Saturday.”). With regard to claim 6 Valliani discloses wherein the generating the prompt based on the one or more objects detected within the image data further comprises: receiving an input that selects an object from among the one or more objects detected within the image data; and generating the prompt based on the object selected by the input (paragraph [0028], “selects a portion of the image that is associated with a recognizable object. The portion of the image may be selected prior to recognition of an object in the image”; paragraph [0071]). With regard to claims 8 and 15, claims 8 and 15 are rejected same as claim 1 and the arguments similar to that presented above for claim 1 are equally applicable to claims 8 and 15 A. Valliani discloses a system comprising a memory and a processor as shown in Figures 1-2 and 6, and all of the other limitations similar to claim 1 are not repeated herein, but incorporated by reference. With regard to claims 9 and 16, claims 9 and 16 are rejected same as claim 2 and the arguments similar to that presented above for claim 2 are equally applicable to claims 9 and 16, and all of the other limitations similar to claim 2 are not repeated herein, but incorporated by reference. With regard to claims 10 and 17, claims 10 and 17 are rejected same as claim 3 and the arguments similar to that presented above for claim 3 are equally applicable to claims 10 and 17, and all of the other limitations similar to claim 3 are not repeated herein, but incorporated by reference. With regard to claims 11 and 18, claims 11 and 18 are rejected same as claim 4 and the arguments similar to that presented above for claim 4 are equally applicable to claims 11 and 18, and all of the other limitations similar to claim 4 are not repeated herein, but incorporated by reference. With regard to claims 12 and 19, claims 12 and 19 are rejected same as claim 5 and the arguments similar to that presented above for claim 5 are equally applicable to claims 12 and 19, and all of the other limitations similar to claim 5 are not repeated herein, but incorporated by reference. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHEFALI D. GORADIA whose telephone number is (571)272-8958. The examiner can normally be reached on Monday-Thursday 8AM-6PM, Friday 8AM-12PM. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Henok Shiferaw can be reached on 571-272-4637. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. SHEFALI D. GORADIA Primary Patent Examiner Art Unit 2676 /SHEFALI D GORADIA/Primary Patent Examiner, Art Unit 2676
Read full office action

Prosecution Timeline

Nov 09, 2023
Application Filed
Apr 01, 2026
Non-Final Rejection mailed — §101, §102
Jun 30, 2026
Response Filed
Sep 15, 2026
Final Rejection mailed — §101, §102 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749191
MULTIMODAL PREDICTION OF VISUAL ACUITY RESPONSE
3y 4m to grant Granted Sep 29, 2026
Patent 12744870
COLOR ANGLE SCANNING IN CROP ROW DETECTION
2y 8m to grant Granted Sep 22, 2026
Patent 12737892
SYSTEM AND METHODS FOR AUTOMATED PHOTOSENSITIVITY DETECTION
2y 5m to grant Granted Sep 15, 2026
Patent 12738003
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND A COMPUTER-READABLE STORAGE MEDIUM STORING AN INFORMATION PROCESSING PROGRAM
2y 6m to grant Granted Sep 15, 2026
Patent 12734616
METHOD FOR QUALITY CONTROL OF A WELDING JOINT BETWEEN A PAIR OF ENDS OF CONDUCTING ELEMENTS OF AN INDUCTIVE WINDING OF A STATOR
2y 0m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
90%
Grant Probability
99%
With Interview (+11.4%)
2y 5m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 618 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month