Prosecution Insights
Last updated: October 02, 2026
Application No. 18/789,533

ADJUSTING PROBABILITY OF AN END-OF-SENTENCE TOKEN IN A GENERATIVE ARTIFICIAL INTELLIGENCE MODEL

Final Rejection §101§102§103§Other
Filed
Jul 30, 2024
Examiner
MCCORD, PAUL C
Art Unit
2692
Tech Center
2600 — Communications
Assignee
Microsoft Technology Licensing, LLC
OA Round
2 (Final)
69%
Grant Probability
Favorable
3-4
OA Rounds
1y 3m
Est. Remaining
95%
With Interview

Examiner Intelligence

Grants 69% — above average
69%
Career Allowance Rate
405 granted / 585 resolved
+7.2% vs TC avg
Strong +26% interview lift
Without
With
+25.9%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
37 currently pending
Career history
621
Total Applications
across all art units

Statute-Specific Performance

§101
5.4%
-34.6% vs TC avg
§103
60.9%
+20.9% vs TC avg
§102
8.8%
-31.2% vs TC avg
§112
19.1%
-20.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 585 resolved cases

Office Action

§101 §102 §103 §Other
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 17-19, 21, 22 rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter because the broadest reasonable interpretation of the claimed one or more storage devices includes a signal embodiment; paragraph 100 of the instant specification recites that the storage may comprise “any other media which can be used to store information, and which can be accessed within the computer system,” and paragraph 105 points out that computer readable storage such as computer readable media may comprise “volatile media,” and as such the claims are considered to comprise a signal per se. Applicants amendments to claims 1-20 suffice to obviate the instant 35 U.S.C. 101 rejection with regard to the claimed invention being directed to a judicial exception without significantly more. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1, 3-5, 7, 8, 12-15, 17-19, 21-24 is/are rejected under 35 U.S.C. 102a1 as being anticipated by Gao: “Cross Modal Compression With Variable Rate Prompt,” (copy provided by Examiner; copyright 2023). Regarding claim 1 Gao teaches: A computing device comprising a processor system and memory, wherein the computing device implements a generative artificial intelligence ("Al") model (Gao: Abstract; § III.B.1: an AI model operable to manage EOS complexity by determining the likelihood of an EOS token such as using cross modal compression (CMC) to encode an image as text) configured to perform operations comprising: accepting, at the generative Al model, input (Gao: Abstract; § I., III.B.1, III.B.3; Fig 2, 3: an encoder receives image data as input); producing, with the generative Al model, input tokens that encode the input (Gao: Abstract; § I., III.B.1, III.B.3; Fig 2, 3: system generates tokens, text, etc. of a desired length DL for encoding the input image); and producing, with the generative Al model, a text response using the input tokens that encode the input (Gao: Abstract; § I., III.B.1, III.B.3; Fig 2, 3: system encodes a variable rate prompt to generate a text response from the tokens, text, etc. such as through multiple successive decoding steps), wherein the producing the text response includes multiple iterations of output token generation (id.), wherein a probability of an end-of-sentence ("EOS") token increases in successive iterations among the multiple iterations of output token generation (Gao: § I., III.B.2, III.B.3; Fig 2, 3: an exponential decay probability of the likelihood of an EOS token (DEP) such as to manage overall length with respect to keeping grammatically correct), and wherein the producing the text response includes, in a current iteration of the multiple iterations of output token generation (Gao: § I., III.B.2, III.B.3; Fig 2, 3; Eqn 2, 3: DEP adjustment applied at each iteration decoding step in an autoregressive manner): determining the probability of the EOS token in a previous iteration of the multiple iterations of output token generation, the previous iteration immediately preceding the current iteration (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: EOS probability determined for each of a plurality of decoding steps, iterations, etc. based on the preceding step(s)); increasing the probability of the EOS token in the current iteration by adjusting, according to an exponential growth factor, the probability of the EOS token in the previous iteration (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: EOS probability increases exponentially; “As i increases, EOS probability increases,”); producing one or more output tokens, wherein each of the one or more output tokens is the EOS token or a text token representing one or more words (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: system selects from a vocabulary comprising text tokens and EOS token at each step); and based at least in part on whether the EOS token has been produced in the current iteration, determining whether or not to complete the producing the text response (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: system continues generation of text until an EOS token indicates completion). Regarding claim 3 Gao teaches: The computing device of claim 1, wherein the generative Al model is a vision language model, and wherein a visual encoder and/or text encoder of the generative Al model accept the input and produce the input tokens that encode the input (Gao: Abstract; § I., III.B.1, III.B.3; Fig 2, 3: a vison to language model wherein an image encoder generates text, tokens, etc. which encode the input for use by a decoder back end). Regarding claim 4 Gao teaches: The computing device of claim 1, wherein a text decoder of the generative Al model produces the text response (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: system selects from a vocabulary comprising text tokens and EOS token at each successive decoding step). Regarding claim 5 Gao teaches: The computing device of claim 1, wherein the producing the text response further includes, in the current iteration if the EOS token was produced, completing the producing the text response; and otherwise, the EOS token not having been produced, continuing in a next iteration among the multiple iterations of output token generation ach of the multiple iterations of output token generation: increasing the probability of the EOS token; producing one or more output tokens, wherein each of the one or more output tokens is the EOS token or a text token representing one or more words (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: system generates an output text by selecting from a vocabulary comprising text tokens and EOS token at each successive decoding step upon arriving at an EOS token ). Regarding claim 7 Gao teaches: The computing device of claim 1, wherein the probability of the EOS token increases in the successive iterations according to: PEoSs1,…,sk-t=1-1-PEoSs1,…,sk-t-11+∝, wherein k represents a target limit on count of output tokens, t represents a counter that decreases in the successive iterations, PEoSs1,…,sk-t represents the probability of the EOS token in the current iteration, PEoSs1,…,sk-t-1represents the probability of the EOS token in the previous iteration, and ∝ represents a hyper parameter that controls a rate of the exponential growth factor (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: system selects from a vocabulary comprising text tokens and EOS token at each step and manages a dynamic EOS probability by managing the rate of exponential decay over successive decoding steps in a manner considered mathematically equivalent to the recited formula). Regarding claim 8 Gao teaches: The computing device of claim 7, wherein the hyper parameter that controls the rate of the exponential growth factor is in a range of (0, …, 1) (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: 0 to 1 set as the exponential decay factor as the 0 to 1 range of the hyperparameter directs the exponential decay of the EOS probability over 0:1). Regarding claim 12 Gao teaches: The computing device of claim 1, wherein the generative AI model has been trained in a training process comprising: receiving an initial training set comprising images and initial text captions, each of the initial text captions being associated with an image among the images (Gao: § I., III.B.2, III.B.3, III.B.4, IV.A, IV.C, IV.D; Fig 2, 3; Eqn 2, 3: system trained, tested on MSCOCO, which comprises images and plurality of captions potentially relevant thereto); updating the initial training set, including distilling the initial text captions into final text captions (Gao: § III.B.2, III.B.3, III.B.4, IV.A, IV.C, IV.D; Fig 2, 3; Eqn 2, 3: system generates coherent text by random truncation which is used to train the model), wherein a given final text caption among the final text captions is: generated using a corresponding initial text caption among the initial text captions; more concise than the corresponding initial text caption; and associated with the image that is associated with the corresponding initial text caption (Gao: § III.B.2, III.B.3, III.B.4, IV.A, IV.C, IV.D; Fig 2, 3; Eqn 2, 3: system trained to generate short text and use same to improve the generation of more concise and/or accurate captions); and adjusting the generative AI model using the updated training set (Gao: § III.B.2, III.B.3, III.B.4, IV.A, IV.C, IV.D; Fig 2, 3; Eqn 2, 3: system trained and adapted to generate short text and use same to improve the generation of more concise and/or accurate captions such as by improving the model). Regarding claims 13, 17—the claims are considered to recite substantially similar subject matter to that of claim 1 and are similarly rejected. Regarding claims 14, 18—the claims are considered to recite substantially similar subject matter to that of claim 3 and are similarly rejected. Regarding claims 15, 19—the claims are considered to recite substantially similar subject matter to that of claim 5 and are similarly rejected. Regarding claims 21, 23—the claims are considered to recite substantially similar subject matter to that of claim 7 and are similarly rejected. Regarding claims 22, 24—the claims are considered to recite substantially similar subject matter to that of claim 8 and are similarly rejected. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-5, 7-9, 11-15, 17-19, 21-24 is/are rejected under 35 U.S.C. 102a1 as being anticipated by Gao: “Cross Modal Compression With Variable Rate Prompt,” (copy provided by Examiner; copyright 2023) further in view of Song: 20210286951. Regarding claim 1 Gao teaches: A computing device comprising a processor system and memory, wherein the computing device implements a generative artificial intelligence ("Al") model (Gao: Abstract; § III.B.1: an AI model operable to manage EOS complexity by determining the likelihood of an EOS token such as using cross modal compression (CMC) to encode an image as text) configured to perform operations comprising: accepting, at the generative Al model, input (Gao: Abstract; § I., III.B.1, III.B.3; Fig 2, 3: an encoder receives image data as input); producing, with the generative Al model, input tokens that encode the input (Gao: Abstract; § I., III.B.1, III.B.3; Fig 2, 3: system generates tokens, text, etc. of a desired length DL for encoding the input image); and producing, with the generative Al model, a text response using the input tokens that encode the input (Gao: Abstract; § I., III.B.1, III.B.3; Fig 2, 3: system encodes a variable rate prompt to generate a text response from the tokens, text, etc. such as through multiple successive decoding steps), wherein the producing the text response includes multiple iterations of output token generation (id.), wherein a probability of an end-of-sentence ("EOS") token increases in successive iterations among the multiple iterations of output token generation (Gao: § I., III.B.2, III.B.3; Fig 2, 3: an exponential decay probability of the likelihood of an EOS token (DEP) such as to manage overall length with respect to keeping grammatically correct), and wherein the producing the text response includes, in a current iteration of the multiple iterations of output token generation (Gao: § I., III.B.2, III.B.3; Fig 2, 3; Eqn 2, 3: DEP adjustment applied at each iteration decoding step in an autoregressive manner): determining the probability of the EOS token in a previous iteration of the multiple iterations of output token generation, the previous iteration immediately preceding the current iteration (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: EOS probability determined for each of a plurality of decoding steps, iterations, etc. based on the preceding step(s)); increasing the probability of the EOS token in the current iteration by adjusting, according to an exponential growth factor, the probability of the EOS token in the previous iteration (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: EOS probability increases exponentially; “As i increases, EOS probability increases,”); producing one or more output tokens, wherein each of the one or more output tokens is the EOS token or a text token representing one or more words (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: system selects from a vocabulary comprising text tokens and EOS token at each step); and based at least in part on whether the EOS token has been produced in the current iteration, determining whether or not to complete the producing the text response (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: system continues generation of text until an EOS token indicates completion). It may be argued that Gao does not explicitly teach the recited EOS probability adjustment in a step by step manner. In a related field of endeavor Song teaches a transformer model generative of output text (Song: Abstract) by producing, with the generative AI model, input tokens that encode the input (Song: Abstract; ¶ 3, 4, 26-29, 31-36; Fig 4: a transformer neural model receives tokenized text as an input generates target summary text as output stepwise over a plurality of tokens); producing, with the generative AI model, a text response using the input tokens that encode the input (id.), wherein a probability of a token over stepwise successive iterations among the multiple iterations of output token generation to apply a diminishing reward to candidate words, tokens, etc. generated as output and comprising applying a diminishing reward value to candidate words in a sequence of words to thereby maintain the sequence at or about a length threshold (Song: ¶ 3, 4, 26-29, 31-36; Claim 1). It would have been obvious to one of ordinary skill in the art before the effective filing date of the instant application to modify the decaying EOS likelihoods as taught or suggested by Gao with the per step recurrence score adjustment discussed by Song thereby utilizing an effective length penalty applied as a per-step recurrence as taught or suggested by Song to bound the length of the Gao and/or Song output sequences such that as the length of a caption increases the probability of any particular word or token decreases and the probability of an EOS token increases for at least the purpose of managing output sequences to reliably terminate at or below a target length thereby reducing runtime, compute, etc. of the processing while improving readability, consistency, accuracy, etc. of the output; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 2 Gao in view of Song teaches or suggests: The computing device of claim 1, wherein the generative AI model is a large language model (Song: Abstract: a generative transformer model), and wherein a text encoder of the generative AI model accepts the input and produces the input tokens that encode the input (Gao: Abstract; § I., III.B.1, III.B.3; Fig 2, 3: system generates tokens, text, etc. of a desired length DL for encoding the input image); (Song: Abstract; ¶ 3, 4, 26-29, 31-36; Fig 4: a transformer neural model receives tokenized text as an input generates target summary text as output stepwise over a plurality of tokens). The claim is considered obvious over Gao as modified by Song as addressed in the base claim as it would have been obvious to apply the further teaching of Gao and/or Song to the modified device of Gao and Song; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 3 Gao in view of Song teaches or suggests: The computing device of claim 1, wherein the generative AI model is a vision language model (Gao: Abstract; § I., III.B.1, III.B.3; Fig 2, 3: a vison to language model), and wherein a visual encoder and/or text encoder of the generative Al model accept the input and produce the input tokens that encode the input. (Gao: Abstract; § I., III.B.1, III.B.3; Fig 2, 3: a vison to language model wherein an image encoder generates text, tokens, etc. which encode the input for use by a decoder back end); (Song: Abstract, etc.). Further, Examiner has taken official notice which Applicant has failed to timely and explicitly traverse and it is thus additionally accepted as Admitted Prior Art (APA: please see MPEP 2144.03) that adapting a vision model to comprise a vision language model adapted to encode, tokenize text, etc. would have comprised an obvious inclusion for at least the purpose of integrating modalities such as vision for the purpose of performance of multi-modal predictive tasks. The claim is considered obvious over Gao as modified by Song as addressed in the base claim as it would have been obvious to apply the further teaching of Gao and/or Song to the modified device of Gao and Song; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 4 Gao in view of Song teaches or suggests: The computing device of claim 1, wherein a text decoder of the generative AI model produces the text response (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: system selects from a vocabulary comprising text tokens and EOS token at each successive decoding step); (Song: Abstract; ¶ 36; Fig 4: a transformer neural model receives tokenized text as an input generates target summary text as output). The claim is considered obvious over Gao as modified by Song as addressed in the base claim as it would have been obvious to apply the further teaching of Gao and/or Song to the modified device of Gao and Song; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 5 Gao in view of Song teaches or suggests: The computing device of claim 1, wherein the producing the text response includes, in each of the multiple iterations of output token generation: increasing the probability of the EOS token (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: EOS probability increases exponentially; “As i increases, EOS probability increases,”); (Song: ¶ 3, 4, Claim 1: a length threshold and/or diminishing reward value of words iteratively increases the probability of a shorter sentence and increases the respective probability of the Gao EOS token); producing one or more output tokens, wherein each of the one or more output tokens is the EOS token or a text token representing one or more words; if the EOS token was produced, completing the producing the text response (Lee: ¶ 59); and otherwise, the EOS token not having been produced, continuing in a next iteration among the multiple iterations of output token generation (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: system generates an output text by selecting from a vocabulary comprising text tokens and EOS token at each successive decoding step upon arriving at an EOS token ). The claim is considered obvious over Gao as modified by Song as addressed in the base claim as it would have been obvious to apply the further teaching of Gao and/or Song to the modified device of Gao and Song; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 7 Gao in view of Song teaches or suggests: The computing device of claim 1, wherein the probability of the EOS token increases in the successive iterations according to: PEoSs1,…,sk-t=1-1-PEoSs1,…,sk-t-11+∝, wherein k represents a target limit on count of output tokens, t represents a counter that decreases in the successive iterations, PEoSs1,…,sk-t represents the probability of the EOS token in the current iteration, PEoSs1,…,sk-t-1represents the probability of the EOS token in the previous iteration, and ∝ represents a hyper parameter that controls a rate of the exponential growth factor (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: system selects from a vocabulary comprising text tokens and EOS token at each step and manages a dynamic EOS probability by managing the rate of exponential decay over successive decoding steps in a manner considered mathematically equivalent to the recited formula); (Song: 6, 35, 36: output sequence iterated in a per word, token, etc. fashion and based on probabilities thereof). The claim is considered obvious over Gao as modified by Song as addressed in the base claim as it would have been obvious to apply the further teaching of Gao and/or Song to the modified device of Gao and Song; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 8 Gao in view of Song teaches or suggests: The computing device of claim 7, wherein the hyper parameter that controls the rate of the exponential growth factor is in a range of (0, …, 1) (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: 0 to 1 set as the exponential decay factor as the 0 to 1 range of the hyperparameter directs the exponential decay of the EOS probability over 0:1). The claim is considered obvious over Gao as modified by Song as addressed in the base claim as it would have been obvious to apply the further teaching of Gao and/or Song to the modified device of Gao and Song; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 9 Gao in view of Song teaches or suggests: The computing device of claim 7, wherein the hyper parameter that controls the rate of the exponential growth factor has been set in a training process (Gao: § III.B.2, III.B.3, III.B.4, IV.A, IV.C, IV.D; Fig 2, 3; Eqn 2, 3: system generates coherent text by random truncation which is used to train and improve the model) comprising: receiving a training set comprising images and text captions, each of the text captions being associated with an image among the images (Gao: § I., III.B.2, III.B.3, III.B.4, IV.A, IV.C, IV.D; Fig 2, 3; Eqn 2, 3: system trained, tested on MSCOCO, which comprises images and plurality of captions potentially relevant thereto); and adjusting the generative AI model using the training set, including adjusting the hyper parameter that controls the rate of the exponential growth factor comprising: receiving a training set comprising of text data (Lee: ¶ 3, 4, 59, 67, 68, 92, 104: system for training a neural network to generate particular sequences of tokens); (Song: ¶ 3, 4, 6, 14, 35, 36: system for training a neural network to generate particular sequences of tokens); and adjusting the generative AI model using the training set, including adjusting the hyper parameter that controls the rate of the exponential growth factor (Gao: § III.B.2, III.B.3, III.B.4; Fig 2, 3; Eqn 2, 3: EOS probability determined for each of a plurality of decoding steps, iterations, etc. based on the preceding step(s) such that the EOS probability increases exponentially; “As i increases, EOS probability increases,”); (Song: 3, 35: the algorithm adjusts a reward value to increase and or decrease an overall length in concert with a defined length threshold). The claim is considered obvious over Gao as modified by Song as addressed in the base claim as it would have been obvious to apply the further teaching of Gao and/or Song to the modified device of Gao and Song; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 11 Gao in view of Song teaches or suggests: The computing device of claim 1, wherein the operations further comprise, during inference using the generative AI model ((please see claim 1 supra; such as using the Gao in view of Song method): identifying one or more representative units of video (Gao: III.A: system operative upon video to transform an image therein to a text representation); (please see ¶ 42-44 of the instant specification such as using well known scene detection and keyframe determination functionality of Azure AI video indexer ,or Amazon recognition video, or PySceneDetect, or other well-known methods therefor detailed by Applicant); ranking the one or more representative units (please see claim 1 supra; and Song: ¶ 26-29: ranking of candidate outputs based on a best-first search or based on any of the well-known ranking functionality detailed by Applicant); and based on results of the ranking the one or more representative units, selecting a particular representative of the one or more representative units, wherein the particular representative unit is provided to the generative AI model as the input (please see claim 1 supra; Gao: Abstract: cross modal compression of image data, generation of captions therefrom, etc.; see additionally ¶ 42-44 of the instant specification ). Thus Examiner accepts Applicants admission of the relevant prior art and its suitability for the claimed purpose which would be obvious to combine with the Gao in view of Song system and method to thereby perform keyframe determination in the matter disclosed as well-known and to operate thereupon by the Gao in view of Song system and method cross modal image compression system to provide a summary thereof for the purpose of thereby additionally providing a summary for a video or scene therein; one of ordinary skill in the art would have expected only predictable results therefrom. The claim is thus considered obvious over Gao as modified by Song as addressed in the base claim and in view of the well-known nature of the claimed subject matter as it would have been obvious to apply the further teachings of Gao, Song, and/or Azure, Amazon, etc. to the modified device of Gao and Song; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 12 Gao in view of Song teaches or suggests: The computing device of claim 1, wherein the generative AI model has been trained in a training process comprising: receiving an initial training set comprising images and initial text captions, each of the initial text captions being associated with an image among the images (Gao: § I., III.B.2, III.B.3, III.B.4, IV.A, IV.C, IV.D; Fig 2, 3; Eqn 2, 3: system trained, tested on MSCOCO, which comprises images and plurality of captions potentially relevant thereto): (Song: Abstract; ¶ 38: such as comprising text input); updating the initial training set, including distilling the initial text captions into final text captions (Gao: § III.B.2, III.B.3, III.B.4, IV.A, IV.C, IV.D; Fig 2, 3; Eqn 2, 3: system generates coherent text by random truncation which is used to train the model); (Song: Abstract; ¶ 38: model trained on input and generated output), wherein a given final text caption among the final text captions is: generated using a corresponding initial text caption among the initial text captions; more concise than the corresponding initial text caption; and associated with the image that is associated with the corresponding initial text caption (Gao: § III.B.2, III.B.3, III.B.4, IV.A, IV.C, IV.D; Fig 2, 3; Eqn 2, 3: system trained to generate short text and use same to improve the generation of more concise and/or accurate captions); (Song: ¶ 3, 4, 33-36: such as based on adjusting length values of the word reward algorithm of the length reward); and adjusting the generative AI model using the updated training set (Gao: § III.B.2, III.B.3, III.B.4, IV.A, IV.C, IV.D; Fig 2, 3; Eqn 2, 3: system trained and adapted to generate short text and use same to improve the generation of more concise and/or accurate captions such as by improving the model); (Song: ¶ 3, 4, 33-36, 38: system iteratively improves a word output by iterative adjustment of training data based thereon). Regarding claims 13, 17—the claims are considered to recite substantially similar subject matter to that of claim 1 and are similarly rejected. Regarding claims 14, 18—the claims are considered to recite substantially similar subject matter to that of claim 3 and are similarly rejected. Regarding claims 15, 19—the claims are considered to recite substantially similar subject matter to that of claim 5 and are similarly rejected. Regarding claims 21, 23—the claims are considered to recite substantially similar subject matter to that of claim 7 and are similarly rejected. Regarding claims 22, 24—the claims are considered to recite substantially similar subject matter to that of claim 8 and are similarly rejected. Response to Arguments Applicant’s arguments in concert with claim amendments, see Remarks and Claims, filed 4/24/26, with respect to the rejection(s) of claim(s) 1-20 under 35 USC 103 variously over Lee in view of Song, Lee in view of Song in view of the Hugging Face transformer model specification 2.42; Lee in view of Song in view of the Hugging Face transformer model specification 2.42 in view of Seybold; and Lee in view of Song in view of Seybold, have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Gao and/or Gao in view of Song. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to PAUL C MCCORD whose telephone number is (571)270-3701. The examiner can normally be reached 730-630 M-F. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, CAROLYN EDWARDS can be reached at (571) 270-7136. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PAUL C MCCORD/Primary Examiner, Art Unit 2692
Read full office action

Prosecution Timeline

Jul 30, 2024
Application Filed
Feb 04, 2026
Non-Final Rejection mailed — §101, §102, §103
Apr 24, 2026
Response Filed
Jul 16, 2026
Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12750630
MEDIA PLAYBACK BASED ON SENSOR DATA
2y 7m to grant Granted Sep 29, 2026
Patent 12745041
Playback Device Pairing
3y 0m to grant Granted Sep 22, 2026
Patent 12724972
AUTOMATIC SENTENCE CONDITION MATCHING USING NATURAL LANGUAGE PROCESSING
2y 11m to grant Granted Sep 01, 2026
Patent 12718824
Audio Signal Encoding Method, Decoding Method, Encoding Device, and Decoding Device
3y 10m to grant Granted Aug 25, 2026
Patent 12718804
DOMAIN SPECIALTY INSTRUCTION GENERATION FOR TEXT ANALYSIS TASKS
3y 1m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
69%
Grant Probability
95%
With Interview (+25.9%)
3y 5m (~1y 3m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 585 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month