Prosecution Insights
Last updated: August 16, 2026
Application No. 18/974,344

TEXT GENERATION METHOD AND APPARATUS, ELECTRONIC DEVICE, AND STORAGE MEDIUM

Non-Final OA §101§103
Filed
Dec 09, 2024
Priority
Dec 11, 2023 — CN 202311695859.4
Examiner
BEZUAYEHU, SOLOMON G
Art Unit
Tech Center
Assignee
Lemon Inc.
OA Round
1 (Non-Final)
75%
Grant Probability
Favorable
1-2
OA Rounds
1y 6m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 75% — above average
75%
Career Allowance Rate
473 granted / 627 resolved
+15.4% vs TC avg
Strong +30% interview lift
Without
With
+30.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
37 currently pending
Career history
664
Total Applications
across all art units

Statute-Specific Performance

§101
17.3%
-22.7% vs TC avg
§103
52.1%
+12.1% vs TC avg
§102
12.9%
-27.1% vs TC avg
§112
10.2%
-29.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 627 resolved cases

Office Action

§101 §103
DETAILED ACTION Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. When reviewing independent claim 1, and based upon consideration of all of the relevant factors with respect to the claim as a whole, claims 1-20 are held to claim an abstract idea without reciting elements that amount to significantly more than the abstract idea and is/are therefore rejected as ineligible subject matter under 35 U.S.C. 101. The Examiner will analyze Claim 1, and similar rationale applies to independent Claim 9 and 17. The rationale, under MPEP § 2106, for this finding is explained below. The claimed invention (1) must be directed to one of the four statutory categories, and (2) must not be wholly directed to subject matter encompassing a judicially recognized exception, as defined below. The following two step analysis is used to evaluate these criteria. Step 1: Is the claim directed to one of the four patent-eligible subject matter categories: process, machine, manufacture, or composition of matter? When examining the claim under 35 U.S.C. 101, the Examiner interprets that the claims is related to a process since the claim is directed to a text generation method. Step 2a, Prong 1: Does the claim wholly embrace a judicially recognized exception, which includes laws of nature, physical phenomena, and abstract ideas, or is it a particular practical application of a judicial exception? The Examiner interprets that the judicial exception applies since Claim 1 limitation of extracting events from a video to be processed, and determining target video frames corresponding to the events [mental process. Investigator reviewing an incident footage and taking event notes]; extracting frame features of the target video frames, and determining event features based on the frame features [Also a mental process. The investigator determining a precursor to an event]; concatenating event features the in an extraction order of corresponding events, to generate a prompt text [also mental process. The investigator putting incidents together and asking questions in order to determine what happened]; and generating, by a first language model, a description text of the video to be processed based on the prompt text [Also a mental process. The investigator making a decision (description) based on the answer to the question] are directed to an abstract. If/when the claim recites a judicial exception (i.e., an abstract idea enumerated in MPEP § 2106.04(a), a law of nature, or a natural phenomenon), the claim requires further analysis in Prong Two. Step 2a, Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? No Step 2b: If a judicial exception into a practical application is not recited in the claim, the Examiner must interpret if the claim recites additional elements that amount to significantly more than the judicial exception. The Examiner interprets that the Claims do not amount to significantly more. Furthermore, the generic computer components or machine learning algorithm of the processor/memory recited as performing generic computer or machine learning functions that are well-understood, routine and conventional activities amount to no more than implementing the abstract idea with a computerized system. The Examiner finds that Claims 2-8 does not state significantly more since the claim only recites additional steps for analyzing video using machine learning model in order to generate text. Thus, claims 1-20 recite the same abstract idea and therefore are not drawn to the eligible subject matter as they are directed to the abstract idea without significantly more. Therefore, all claims are rejected under 35 U.S.C. 101. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 2, 4, 6-10, 12, 14-18 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Lee et al. (Pub. No. US 2024/0362272) in view of Hirshberg et al. (Pub. No. US 2024/0370661 hereinafter “Hir”). Regarding claim 1, Lee teaches a text generation method, comprising: extracting events (video clips) from a video to be processed [Para. 119 “The video analysis system 130 obtains the video for the query. The video analysis system 130 may optionally apply segmentation techniques, such as scene boundary detection and/or key frame detection to identify a plurality of video clips from the video file”. The phrase “to be processed” in the claim, is nothing more than an intended use of the video file. Therefore, any video file could be considered “to be processed video file” until is actually processed], and determining target video frames (frame data) corresponding to the events (video clips) [Para. 119 “The video analysis system 130 obtains the video for the query. The video analysis system 130 may optionally apply segmentation techniques, such as scene boundary detection and/or key frame detection to identify a plurality of video clips from the video file.”; para. 113 “as described above, the video encoder 810 is coupled to receive the frame data, audio data, and/or text data of a video or video clip and generate a set of video embeddings 822 numerically representing the video or video clip in a latent space”. Any frame could be considered the target frame. The term “corresponding” is broad]; extracting frame features (visual embeddings) of the target video frames (visual embeddings) [Para. 69 “The visual encoder 212 is coupled to receive a sequence of frames and generate a set of visual embeddings that encode the visual information in the frames”; Para. 113 “As described above, the video encoder 810 is coupled to receive the frame data, audio data, and/or text data of a video or video clip and generate a set of video embeddings 822 numerically representing the video or video clip in a latent space”], and determining event features (video-language-aligned embeddings) based on the frame features (visual embedding) [Para. 71 “The multimodal encoder 220 is coupled to receive the set of visual embeddings, the set of audio embeddings, and the set of text embeddings, and generate the set of video embeddings 222”; Para. 114 “The alignment model 840 is coupled to receive the set of video embeddings 822 and generate a set of video-language-aligned embeddings 828”]. Lee also teaches having event features (video-language-aligned embeddings), corresponding event (video clips) and concatenating prompts embeddings to video-language aligned embeddings [Para. 119 “The video analysis system 130 may optionally apply segmentation techniques, such as scene boundary detection and/or key frame detection to identify a plurality of video clips from the video file”; Para. 115 “In one instance, the set of prompt embeddings 825 are concatenated to the set of video-language-aligned embeddings 828 to generate a concatenated tensor”]. However, Lee doesn’t explicitly teach concatenating the event features in an extraction order (of corresponding events, to generate a prompt text. Hir teaches concatenating the event features in an extraction order (aggregated timelines) of corresponding events, to generate a prompt text (summary prompt) [Para. 12 “The systems also generate an aggregated timeline of the audio insights and the visual insights which temporally aligns the audio insights and the visual insights. This aggregated timeline is segmented into a plurality of coherent segments, wherein each of the coherent segments includes a unique combination of audio insights and visual insights”; Para. 15 “a summary prompt is generated for each chunk in the set of chunks, which includes (i) the audio insights and visual insights of the coherent segments of each chunk and (ii) the selected summary style.”]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Lee’s embedding-concatenation input to the language model by incorporating Hir’s extraction order (aggregated timeline) and prompt text (summary prompt) so that Lee’s video-language-aligned embeddings for video clips are arranged according to temporal alignment and used with a generated prompt for video description. This medication improves Lee by preserving temporal relationships among clip-level event inputs, thereby improving temporal coherence of the generated description text. generating, by a first language model, a description text (video-level text) of the video to be processed based on the prompt text [Para. 131 “the video analysis system 130 applies the LLM 970 to the dense clip descriptions and the set of prompt embeddings to generate a list of video-level text, including at least one of a title, hashtags, topic, summary, chapters, highlights, dense narrations, and the like of the video”]. Regarding claims 2, 10 and 18, Lee teaches wherein determining the event features (video-language-aligned embeddings) based on the frame features (video embeddings) comprises: transforming the frame features into an input space of the first language model (LLM’s 850 text domain), to obtain transformed features (video-language-aligned embeddings) [Para. 114]; and determining the event features based on the transformed features (video-language-aligned embeddings) [Para. 114]. Regarding claims 4, 12 and 20, Lee teaches predefined order prompts and generation instruction prompts (generate a dense description) [Para. 82 and 121]. However, Lee doesn’t explicitly teach the rest of claim limitations. However, Hir teaches wherein concatenating the event features in the extraction order of the corresponding events comprises: concatenating the event features in the extraction order (aggregated timeline) of the corresponding events based on predefined order prompts and generation instruction prompts (summary prompt) [Para. 12, and 15]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Lee’s embedding-concatenation input to the language model by incorporating Hir’s extraction order (aggregated timeline) and prompt text (summary prompt) so that Lee’s video-language-aligned embeddings for video clips are arranged according to temporal alignment and used with a generated prompt for video description. This medication improves Lee by preserving temporal relationships among clip-level event inputs, thereby improving temporal coherence of the generated description text. Claims 9 and 17 are rejected for the same reason as claim 1 above. Furthermore, Lee teaches an electronic device, comprising: one or more processors and non-transitory medium; and a storage apparatus configured to store one or more programs, wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to perform the claim limitations [fig. 2, 4 and related description]. Regarding claims 6 and 14, Lee teaches performing speech recognition on the audio data, to obtain a speech text [Para. 67]. However, Lee doesn’t explicitly teach the rest of claim limitations. Hir teaches wherein in response to the video to be processed containing audio data, the method further comprises: performing speech recognition on the audio data, to obtain a speech text (transcript) [Para. 44]; and accordingly, generating the prompt text (summary prompt) further comprises: concatenating the concatenated event features with the speech text, to generate the prompt text [Para. 44]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Lee’s embedding-concatenation input to the language model by incorporating Hir’s extraction order (aggregated timeline) and prompt text (summary prompt) so that Lee’s video-language-aligned embeddings for video clips are arranged according to temporal alignment and used with a generated prompt for video description. This medication improves Lee by preserving temporal relationships among clip-level event inputs, thereby improving temporal coherence of the generated description text. Regarding claims 7 and 15, Lee teaches performing feature extraction on the audio data, to obtain an audio feature (audio embeddings) [Para. 69]. However, Lee doesn’t explicitly teach the rest of claim limitations. Hir teaches wherein in response to the video to be processed containing audio data, the method further comprises: performing feature extraction on the audio data, to obtain an audio feature (audio insights); and accordingly, generating the prompt text (summary prompt) further comprises: concatenating the concatenated event features with the audio feature, to generate the prompt text [Para. 15]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Lee’s embedding-concatenation input to the language model by incorporating Hir’s extraction order (aggregated timeline) and prompt text (summary prompt) so that Lee’s video-language-aligned embeddings for video clips are arranged according to temporal alignment and used with a generated prompt for video description. This medication improves Lee by preserving temporal relationships among clip-level event inputs, thereby improving temporal coherence of the generated description text. Regarding claims 8 and 16, Lee teaches wherein after the description text (clip descriptions) is generated, the method further comprises: obtaining a question (customer user query) text of the video to be processed; and generating, by a second language model (LLM 970), an answer text (response) based on the description text (clip descriptions) and the question text (custom user query) [Para. 150. Video-language-aligned embeddings is same as the transformed and event features because Lee teaches that video embeddings for reach video clip are transformed/projected into LLM-text-domain video-language-aligned embeddings, making them the transformed feature representation of the mapped event/video clip, see para. 114, and 121]. Claims 3, 5, 11, 13, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Lee et al. (Pub. No. US 2024/0362272) in view of Hirshberg et al. (Pub. No. US 2024/0370661 hereinafter “Hir”) and Panagopoulou et al. (Pub. No. US 2024/0370718 hereinafter “Pana”). Regarding claims 3, 11, and 19, Lee in view of Hir doesn’t explicitly teach the claim limitations. However, Pana teaches wherein determining the event features (instruction-aware input representation) based on the transformed features comprises at least one of: concatenating the transformed features, to obtain the event features [Para. 50]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Lee in view of Hir’s alignment model processing by incorporating Pana’s teaching of concatenating transformed feature for successive video frames to form the event level representation used by Lee’s language model. This medication improves Lee by preserving frame to frame changes in the transformed representation, thereby improving temporal fidelity of event feature generation. Regarding claims 5 and 13, Lee in view of Hir doesn’t explicitly teach the claim limitations. However, Pana teaches wherein concatenating the event features in the extraction order of the corresponding events based on the predefined order prompts and the generation instruction prompts comprises: concatenating the event features after corresponding order prompts according to the extraction order of the events, to generate an event order prompt text; and concatenating the generation instruction prompts after the event order prompt text, to generate the prompt text [Para. 34, fig. 3 and related description]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Lee in view of Hir’s alignment model processing by incorporating Pana’s teaching of concatenating transformed feature for successive video frames to form the event level representation used by Lee’s language model. This medication improves Lee by preserving frame to frame changes in the transformed representation, thereby improving temporal fidelity of event feature generation. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to SOLOMON G BEZUAYEHU whose telephone number is (571)270-7452. The examiner can normally be reached on Monday-Friday 10 AM-7 PM. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, O’Neal Mistry can be reached on 313-446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-0101 (IN USA OR CANADA) or 571-272-1000. /SOLOMON G BEZUAYEHU/ Primary Examiner, Art Unit 2666
Read full office action

Prosecution Timeline

Dec 09, 2024
Application Filed
Jul 15, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12688706
System for Estimating Changes in Lane Marker Geometry
2y 9m to grant Granted Jul 21, 2026
Patent 12682658
WHITE LINE RECOGNITION DEVICE, MOBILE OBJECT CONTROL SYSTEM, AND WHITE LINE RECOGNITION METHOD
3y 3m to grant Granted Jul 14, 2026
Patent 12676012
SYSTEM AND METHOD FOR LANE GRAPH ESTIMATION
2y 9m to grant Granted Jul 07, 2026
Patent 12670758
AUTHENTICATION DEVICE AND VEHICLE HAVING THE SAME
3y 6m to grant Granted Jun 30, 2026
Patent 12671766
SYSTEM AND METHOD FOR ELECTRONIC NOTIFICATION IN INSTITUTIONAL COMMUNICATIONS
2y 6m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
75%
Grant Probability
99%
With Interview (+30.2%)
3y 3m (~1y 6m remaining)
Median Time to Grant
Low
PTA Risk
Based on 627 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month