Prosecution Insights
Last updated: October 02, 2026
Application No. 18/762,810

CHALLENGE-RESPONSE AUTHENTICATION USING GENERATIVE ARTIFICIAL INTELLIGENCE

Non-Final OA §103
Filed
Jul 03, 2024
Examiner
HABASHI, DANIEL MONIS S
Art Unit
2407
Tech Center
2400 — Computer Networks
Assignee
Kyndryl Inc.
OA Round
3 (Non-Final)
Grant Probability
Favorable
3-4
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-58.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
9 currently pending
Career history
11
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on May 27, 2026 has been entered. Response to Amendment The Amendment filed May 12, 2026 has been entered. Claims 5, 12, and 19 are cancelled. Claims 1-4, 6-11, 13, 15-18 and 20, as well as the newly added claim 21, remain pending in the application. Response to Arguments Regarding the objection of claim 7 as previously set forth in the Office Action dated March 20, 2026, the amendments to the claim overcomes the objection. Accordingly, the objection is withdrawn. Regarding the rejection of claims 1-6, 8-13, and 15-20 under 35 U.S.C. §103 over US 20240320310 by Callegari et al. in view of US 10097360 to Hachey and US 20170366564 by Ping et al. (hereinafter “Callegari”, “Hachey”, and “Ping” respectively), Applicant's arguments filed May 12, 2026 (hereinafter Remarks) have been fully considered but they are not persuasive. Regarding claim 1, Applicant alleges that Examiner previously took Official Notice regarding the output format of the CAPTCHAs (Remarks, p. 9: “The Office Action appears to take Official Notice regarding the deficiencies. Concerning the rejections of the dependent claims under 35 U.S.C. 103 over the above references in view of “routine skill in the art”, “obvious matter of design choice,” and “Official Notice” pursuant to at least section 2144 of the MPEP, Applicant respectfully traverse the same in view of the following comments found in section 2144.”) Examiner responds by noting that no language indicating Official Notice (including the terms “routine skill in the art”, “obvious matter of design choice”, and “Official Notice”) was used. Rather, Examiner interpreted the broad language of Ping and made such interpretation of record. MPEP §2144.03(C) discusses traversal of rejections based on Official Notice: “If applicant adequately traverses the examiner’s assertion of official notice, the examiner must provide documentary evidence in the next Office action if the rejection is to be maintained.” Upon reconsideration of the application and re-examination of the prior art, the use of audio and video CAPTCHAs is disclosed within Callegari (Callegari [0112]: “It will be appreciated that input 902 and generative model output 906 may each include any of a variety of content types, including, but not limited to, text output, image output, audio output, video output, programmatic output, and/or binary output, among other examples.”) Examiner maintains that no Official Notice was taken, but the cited quotation is documentary evidence to support a rejection under 35 U.S.C. §103 or, if Official Notice was inadvertently taken, to fulfill the obligation imposed by MPEP §2144.03(C). In either case, Applicant’s arguments are not persuasive. Examiner notes that Callegari still does not disclose randomly selecting the output type, only that multiple (including audio and video) may be used. Regarding the now-cancelled claim 5, Applicant argues that Callegari in view of Hachey and Ping does not render obvious “generating a modified prompt: by removing an entity from the prompt such that removing the entity causes an attribute of the audio output type or the video output type in the solution object to be missing from the candidate objects, or by inserting the entity that is not present in the prompt such that inserting the entity causes the attribute of the audio output type or the video output type to be present in the candidate objects but not present in the solution object” as now recited in claim 1 (Remarks, p. 11). For ease of reference, the claim can be broken down into its 2 alternative cases, summarized as follows: The prompt has an entity removed such that the candidates do not have the entity, but the solution does. The question posed to the user is “Which audio/video contains the entity?” The prompt has an entity added such that the candidates have the entity, but the solution does not. The question posed to the user is “Which audio/video does NOT contain the entity?” Examiner responds that Callegari’s description of how the prompts are modified renders obvious both cases. Callegari describes its procedure in terms of “a plurality of categories of variables (e.g., including a subject, a verb, a setting, a style, etc.)” (Callegari, [0005]). These variables are substituted into prompt templates. Consider the example of Callegari FIG. 2. [0040] describes that “the first image 204, the second image 206, and the third image 208 may all be generated using the same first prompt (e.g., a Picasso image of a horse jumping over a fence in space)” while “the fourth image 210 may be generated using a second prompt that is different than the first prompt (e.g., a Picasso image of a lion jumping over a fence in space).” The difference between the first prompt and the second prompt is the exchange of subject (a horse was removed and a lion was added). The instruction 202 directs the user to select the images that contain the aspect removed from the prompt (horses). Bearing in mind that Callegari teaches audio or video format, this directly maps to the first case above (where the prompt may read “a Picasso-style video” or “a cartoon” rather than “a Picasso image”). Therefore, Callegari at least teaches “generating a modified prompt: by removing an entity from the prompt such that removing the entity causes an attribute of the audio output type or the video output type in the solution object to be missing from the candidate objects…”. The second case is readily obvious from the first- rather than removing or replacing an object from the prompt, add an object. Now the unmodified prompt generates the solution while the new prompt generates the candidates. Therefore, Callegari teaches “generating a modified prompt: by removing an entity from the prompt such that removing the entity causes an attribute of the audio output type or the video output type in the solution object to be missing from the candidate objects, or by inserting the entity that is not present in the prompt such that inserting the entity causes the attribute of the audio output type or the video output type to be present in the candidate objects but not present in the solution object”. Applicant further argues that Callegari does not disclose “audio output type or the video output type” (Remarks, p. 11: “Callegari combined with the cited art does not and would not modify any alleged prompt to the alleged modified prompt “by removing an entity… caus[ing] an attribute of the audio output type or the video output type in the solution object to be missing from the candidate objects” or “by inserting the entity… not present in the prompt… caus[ing] the attribute of the audio output type or the video output type to be present in the candidate objects but not present in the solution object”). Callegari does disclose audio and video output types (Callegari [0112]: “It will be appreciated that input 902 and generative model output 906 may each include any of a variety of content types, including, but not limited to, text output, image output, audio output, video output, programmatic output, and/or binary output, among other examples.” See also [0111]: “With reference first to FIG. 9A, conceptual diagram 900 depicts an overview of pre-trained generative model package 904 that processes an input 902 to generate output for CAPTCHA images 906 according to aspects described herein. Examples of pre-trained generative model package 904 includes, but is not limited to, Megatron-Turing Natural Language Generation model (MT-NLG), Generative Pre-trained Transformer 3 (GPT-3), Generative Pre-trained Transformer 4 (GPT-4), BigScience BLOOM (Large Open-science Open-access Multilingual Language Model), DALL-E, DALL-E 2, Stable Diffusion, or Jukebox.” Examiner has emphasized models with video- or audio-generation capabilities), but not their random selection. Therefore, Applicant’s arguments are not convincing. Claim Rejections - 35 USC § 103 The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. Claims 1-4, 6, 8-11, 13, 15-18 and 20-21 are rejected under 35 U.S.C. 103 as being unpatentable over US 20240320310 to Callegari et al. (hereinafter “Callegari”), in view of US 20170366564 to Ping et al. (hereinafter “Ping”) and further in view of US 10097360 to Hachey (hereinafter “Hachey”). Regarding claim 1, Callegari discloses: A computer-implemented method, comprising: [] selecting from an audio output type and a video output type a randomly selected output type (Callegari [0112]: “It will be appreciated that input 902 and generative model output 906 may each include any of a variety of content types, including, but not limited to, text output, image output, audio output, video output, programmatic output, and/or binary output, among other examples.”); generating a prompt (Callegari [0040]: “…developed to generate images from natural language descriptions (e.g., prompts)…”) for an object of the [] selected output type (Callegari [0112]: “It will be appreciated that… generative model output 906 may… include any of a variety of content types, including, but not limited to… audio output, video output… among other examples.”); generating a solution object of the [] selected output type based on the prompt using a generative artificial intelligence (AI) engine (Callegari [0040]: “…developed to generate images from natural language descriptions (e.g., prompts)…”)[]; generating a modified prompt: by removing an entity (Callegari FIG. 2: The horse) from the prompt (Callegari [0040]: “For example, the first image 204, the second image 206, and the third image 208 may all be generated using the same first prompt (e.g., a Picasso image of a horse jumping over a fence in space).”) such that removing the entity causes an attribute (Callegari FIG. 2: The presence of the horse) of the audio output type or the video output type in the solution object (Callegari FIG 2: 204, 206, 208) to be missing from the candidate objects (Callegari FIG 2: 210), or by inserting the entity that is not present in the prompt such that inserting the entity causes the attribute of the audio output type or the video output type to be present in the candidate objects but not present in the solution object; generating the candidate objects of the [] selected output type based on the modified prompt using the generative AI engine (Callegari [0040]: “…the fourth image 210 may be generated by using a second prompt that is different from the first…”); presenting a challenge-response test comprising a question (Callegari FIG. 2, 202: “question” is the text that prompts the user to select the appropriate objects) based on the prompt, the solution object, and the number of the candidate objects [] to a user device (Callegari Fig. 3, 304); and in response to receiving a response to the challenge-response test from the user device comprising an object selection (Callegari [0054]: “At operation 308, it is determined if the selection [by the user] is correct based on the description provided at operation 304…”), performing a responsive action (Callegari [0055]: “…If the selection is not correct based on the provided description, flow branches “NO” to operation 310…”; [0057]: “If the selection is correct based on the provided description, flow branches “YES” to operation 312…”, where the responsive action is either 310 or 312). Callegari does not disclose randomly selecting the CAPTCHA media type. However, Ping discloses: randomly selecting a model output type from among various model output types (Ping [0085]: “A CAPTCHA of a type is randomly selected from CAPTCHAs of different types…”). Callegari and Ping are art analogous to the claimed invention because all are directed towards AI-generated CAPTCHA technology. It would have been obvious to a person having ordinary skill in the art, prior to the effective filing date of the claimed invention, to randomly select the output type as taught by Ping from the types taught by Callegari because this adds another layer of randomness that could thwart automated systems attempting to bypass the security measures. Neither Callegari nor Ping disclose a maximum number of objects generated. However, Hachey discloses: a maximum value is set for a number of candidate objects (Hachey 6:7-15: “In some embodiments, the number of images included in the set of selected images may vary. For example, in one embodiment a set of six images may be selected, while in other embodiments the number of images selected can vary between two and twelve. Any number of selected images are contemplated...”); randomly generating a random value for the number of the candidate objects that is less than the maximum value (Hachey 6:7-15: “…the number of images can vary between two and twelve. Any number of selected images are contemplated. Furthermore, in one embodiment, the number of images that are selected varies upon each invocation of the visual CAPTCHA. For example, upon the first invocation, the system may select a set of six images, while on a subsequent invocation, eight images are selected.”) Hachey is art analogous to the claimed invention because both are directed towards CAPTCHA technology. It would have been obvious to a person having ordinary skill in the art, prior to the effective filing date of the claimed invention, to randomly generate the number of objects presented, subject to a maximum value, as taught by Hachey because randomizing the number of images presented in a challenge increases the difficulty of the challenge for automated systems (Hachey 6:17-21: “By varying the number of images that are presented per invocation of the visual CAPTCHA, embodiments of the present invention may advantageously thwart automated systems that may attempt to use probabilistic analysis to defeat the visual CAPTCHA.”) Regarding claim 2, Callegari in view of Ping and further view of Hachey discloses: The computer-implemented method of claim 1, wherein the generative AI engine is a text-to-image generative AI engine, a text-to-video generative AI engine, a text-to-audio generative AI engine, or a text-to-text generative AI engine (Callegari [0112]: “It will be appreciated that input 902 and generative model output 906 may each include any of a variety of content types, including, but not limited to, text output, image output, audio output, video output, programmatic output, and/or binary output, among other examples.”) Regarding claim 3, Callegari in view of Ping and further view of Hachey discloses: The computer-implemented method of claim 1, wherein the responsive action comprises: determining that the object selection of the response matches the solution object (Callegari [0054]: “At operation 308, it is determined if the selection [by the user] is correct based on the description provided at operation 304…”; see also [0057]: “If the selection is correct based on the provided description, flow branches “YES” to operation 312…”); and granting the user device access to a protected resource (Callegari [0058]: “…the indication that the selection is correct may be the execution of a process, such as granting access to a system protected by the CAPTCHA generated via method 300.”) Regarding claim 4, Callegari in view of Ping and further view of Hachey discloses: The computer-implemented method of claim 1, wherein the responsive action comprises: determining that the object selection of the response does not match the solution object (Callegari [0055]: “…If the selection is not correct based on the provided description, flow branches “NO” to operation 310…”); and performing a security action comprising preventing access to a protected resource for the user device (Callegari [0056]: “…the indication that the selection is incorrect may be the execution of a process, such as locking a user out of a system protected by the CAPTCHA generated via method 300…”) or presenting a second question, a second solution object, and a second candidate object to the user of the user device (Callegari [0056]: “…when the method 300 reaches operation 310 [incorrect selection], the method 300 may return to operation 302 and generate a second plurality of images using the generative model…”). Regarding claim 6, Callegari in view of Ping and further view of Hachey discloses: The computer-implemented method of claim 1, wherein the operations further comprise: generating, by a text-to-text generative AI engine, the question based on the prompt (Callegari [0052]: “… the description may be generated based on one or more of the variables used to generate the plurality of images… the description may instruct a user to select images based on a similarity or difference…”). Claim 8 recites: A system comprising: a memory (Callegari Fig. 11, 1162) having computer readable instructions (Callegari Fig. 11, 1166); and one or more processors (Callegari Fig. 11, 1160-1161) for executing the computer readable instructions, the computer readable instructions controlling the one or more processors to perform operations comprising: the method of claim 1. Therefore, claim 8 recites essentially the same material as claim 1, and is rejected for the same reasons. Claim 9 recites essentially the same material as claim 2, and is rejected for the same reasons. Claim 10 recites essentially the same material as claim 3, and is rejected for the same reasons. Claim 11 recites essentially the same material as claim 4, and is rejected for the same reasons. Claim 13 recites essentially the same material as claim 6, and is rejected for the same reasons. Claim 15 recites: A computer program product comprising a computer readable storage medium (Callegari Fig. 11, 1168) having program instructions embodied therewith (Callegari Fig. 11, 1166), the program instructions executable by one or more processors (Callegari Fig. 11, 1160-1161) to cause the one or more processors to perform operations comprising: the method of claim 1. Therefore, claim 15 recites essentially the same content as claims 1 and 8, and is rejected for the same reasons. Claim 16 recites essentially the same content as claims 2 and 9, and is rejected for the same reasons. Claim 17 recites essentially the same content as claims 3 and 10, and is rejected for the same reasons. Claim 18 recites essentially the same content as claims 4 and 11, and is rejected for the same reasons. Claim 20 recites essentially the same content as claims 6 and 13, and is rejected for the same reasons. Regarding claim 21, Callegari in view of Ping and further view of Hachey discloses: The computer-implemented method of claim 1, wherein the prompt is used to generate (Callegari [0039]: “The instruction 202 may correspond to one of a similarity or difference between the plurality of images. For example, the instruction 202 illustrated in FIG. 2 instructs a user to “select all of the images that show a horse.” Therefore, the illustrated instruction 202 corresponds to a similarity between each of the plurality of images 204-210. In some examples, the instruction 202 corresponds to a difference between each of the plurality of images 204-210, such as by stating “select the images that do not show a horse.”) the question (Callegari Fig. 2, 202. Examiner notes that “Select all images that show a horse” is obviously semantically equivalent to “Which images show a horse?”) that inquires about either the attribute of the audio output type or the video output type being present or missing (Callegari [0039]: “For example, the instruction 202 illustrated in FIG. 2 instructs a user to “select all of the images that show a horse.” Therefore, the illustrated instruction 202 corresponds to a similarity between each of the plurality of images 204-210. In some examples, the instruction 202 corresponds to a difference between each of the plurality of images 204-210, such as by stating “select the images that do not show a horse.”). Callegari in view of Ping and further view of Hachey also discloses the use of text-to-text AI engines (Callegari [0112]: “It will be appreciated that input 902 and generative model output 906 may each include any of a variety of content types, including, but not limited to, text output…”). Callegari in view of Ping and further view of Hachey does not explicitly recite the prompt is input to a text-to-text AI engine to transform the prompt into the question. However, it would have been obvious to a person having ordinary skill in the art, prior to the effective filing date of the claimed invention, to use the described text-to-text AI models (e.g., MT-NLG, GPT-3, or GPT-4. See Callegari [0111]) to generate the question from the prompt, particularly at the same time as generating the alternate prompt, in order to consolidate AI text processing into one step to produce higher-quality outputs from the model as well as reduce redundant usage of the model (e.g., when the model is provided a service which charges a user or company per query/generation). Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Callegari in view of Ping and further view of Hachey as applied to claim 1 above, and further in view of “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding” by J. Devlin et al., published May 2019 and retrieved from https://arxiv.org/abs/1810.04805 (hereinafter “Devlin”). Regarding claim 7, Callegari in view of Ping and further view of Hachey discloses: The computer-implemented method of claim 1, wherein the generating the prompt for the object of the randomly selected output type involves “generating” (Callegari [0043]: “To generate images according to aspects provided herein, prompts may be created by fixing a variable for one or more categories of the plurality of categories and altering (e.g., randomizing) a variable for one or more other categories of the plurality of categories”; Callegari [0042]: “In some examples, the prompts may be generated based on interests specific to a user (e.g., from a database of personal data that is collected with a user's permission)… [or] demographic features of a user… [or] the prompts may be generated based on geographic boundaries corresponding to where a user is located and/or cultural norms associated with the geographic boundaries…”). Callegari in view of Ping and further view of Hachey is silent as to whether the system uses the generative AI engine to create the prompt. However, Devlin teaches Bidirectional Encoder Representations from Transformers (BERT), a language representation artificial intelligence model (Devlin p. 1, Abstract) that is trained on a task that is essentially identical to prompt generation. BERT is trained, in part, on Masked Language Model (Masked LM) training task, where a token (conceptually, a word) is masked out. The model’s goal is then to successfully predict words that fill in the blanks left by the masked tokens. Adapting the example provided by Devlin (p. 12, Appendix A, A.1 “Masked LM and the Masking Procedure”), the sentence “My dog is hairy” may be masked as “My dog is [MASK]”. BERT predict words that are likely to complete the sentence, (e.g., “hairy”, “friendly”) and avoids unlikely ones (e.g., “human”). This encompasses the prompt generation process of fixing some variables and randomizing others as taught by Callegari. The masked token is the variable being randomized, with the likelihood of various terms determined by BERT’s training and the model confidence for each of the terms. Examiner notes that one of ordinary skill in the art would recognize that while BERT’s technology is not especially unique, and that other similar generative models of varying architectures could accomplish the same task, its use of Masked LM provides additional motivation for use in the prompt generation task described by Callegari. Devlin is art analogous to the claimed invention because both are directed towards advances in and uses of generative AI. It would have been obvious to a person having ordinary skill in the art, prior to the effective filing date of the claimed invention, to use the model and technology of BERT, or a similar generative AI model, to generate prompts in a fashion similar to that described by Callegari [0043] to generate prompts faster and more consistently compared to manual prompt generation. Additionally, BERT or a similar model can be trained on “a database of personal data that is collected with a user's permission”, “which may make corresponding CATPCHAs relatively more effective for and/or enjoyable to a user” (Callegari [0042]). Conclusion The following prior art made of record and not relied upon is considered pertinent to applicant’s disclosure: “TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering” by Hu et al. discloses an AI evaluation metric that, as part of its process, uses one a text-to-image model to transform a prompt into an image and a text-to-text language model to generate a question from the prompt (see Figure 2(a)). Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIEL HABASHI whose telephone number is (571)272-2245. The examiner can normally be reached M-F: 9 AM-6 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Catherine Thiaw can be reached at (571)270-1138. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DH/ Examiner Art Unit 2407 /Catherine Thiaw/Supervisory Patent Examiner, Art Unit 2407 8/20/2026
Read full office action

Prosecution Timeline

Show 6 earlier events
Mar 20, 2026
Final Rejection mailed — §103
May 01, 2026
Interview Requested
May 07, 2026
Applicant Interview (Telephonic)
May 07, 2026
Examiner Interview Summary
May 12, 2026
Response after Non-Final Action
May 27, 2026
Request for Continued Examination
Jun 03, 2026
Response after Non-Final Action
Aug 24, 2026
Non-Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
Grant Probability
High
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month