Prosecution Insights
Last updated: October 02, 2026
Application No. 18/680,004

SYSTEMS, METHODS, AND APPARATUSES FOR EVALUATING CONTENT

Non-Final OA §101§103
Filed
May 31, 2024
Examiner
YANG, JIANXUN
Art Unit
2662
Tech Center
2600 — Communications
Assignee
Comcast Cable Communications LLC
OA Round
1 (Non-Final)
74%
Grant Probability
Favorable
1-2
OA Rounds
3m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
491 granted / 663 resolved
+12.1% vs TC avg
Strong +19% interview lift
Without
With
+19.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
43 currently pending
Career history
700
Total Applications
across all art units

Statute-Specific Performance

§101
4.6%
-35.4% vs TC avg
§103
66.2%
+26.2% vs TC avg
§102
5.9%
-34.1% vs TC avg
§112
17.4%
-22.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 663 resolved cases

Office Action

§101 §103
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are pending. Claims 4 and 12 are withdrawn. Claims 1-3, 5-11 and 13-20 are examined. Response to Arguments (Election Requirement of Species) Applicant's election with traverse of Species 2 (claims 3 and 11) in the reply filed on 6/5/2026 is acknowledged. The traversal is on the ground(s) that the species defined by the Examiner are not mutually exclusive and that no undue burden would result from examining all of the claims. After reviewing the applicant's arguments, the examiner agrees to modify the restriction requirement to withdraw Species 1, 4, 5, and 6. Consequently, per applicant’s election of Species 2, claims 4 and 12 (Species 3) are withdrawn from consideration. Claims 1-3, 5-11 and 13-20 are examined in this office action. Species 2, encompassing claims 3 and 11, is directed specifically to determining a stability level or motion factor by comparing a frame to an immediately preceding video frame. In contrast, Species 3, which includes claims 4 and 12, is directed specifically to making this determination by comparing a frame to an immediately following video frame. Practicing the specific method of Species 2 does not inherently require the specific method of Species 3, and vice versa, rendering the claims drawn to distinct, mutually exclusive operations. Furthermore, the examiner disagrees with the applicant's argument that no serious search burden exists merely because all claims relate to the same core technical field of video frame analysis and summarization. A serious search burden remains between the two distinct species. Species 2 requires searching for prior art that utilizes historical frame buffering and retroactive analysis of preceding frames. Conversely, Species 3 requires searching for prior art that utilizes predictive analysis, look-ahead processing, or future-frame buffering. These differing technical methodologies demand distinct search strategies and queries, thereby establishing a valid search burden. For these reasons, the restriction requirement between Species 2 and Species 3 is proper and is maintained. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 11-3, 5-11 and 13-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim(s) 1, 9 and 16 recite a process, which is a statutory category of invention under 35 U.S.C. § 101. Step 1: The Statutory Categories Claim(s) 1, 8, and 15 recite(s) a "processor," a "system," and a "method," which fall under the statutory categories of a machine and a process. Step 2A: The Judicial Exceptions Prong 1: do the claims recite an exception? Claim(s) 1, 9 and 16 are directed to the abstract idea of a mental process and organizing human activity (specifically, observing video frames, evaluating their stability or motion, and summarizing the content based on the clearest visual information). Prong 2: is the exception integrated into a practical application? The claim(s) do not integrate the abstract idea into a practical application. The claims recite applying the abstract idea using a generic "computing device." Implementing an abstract summarization concept on a generic computer as a tool does not improve the functioning of the computer itself or provide a specific improvement to another technology. Step 2B: The Inventive Concept Do the claims amount to "significantly more" than the exception? The additional elements in the claims (receiving data, determining stability/motion/pixel changes, selecting frames, and generating a summary) amount to no more than well-understood, routine, and conventional data gathering and processing activities. The use of a generic computing device to simply automate the mental process of selecting stable video frames for a summary does not add a significant inventive concept. Conclusion: Claim(s) 1, 9 and 16 are directed to an abstract idea and lack an inventive concept. Claim(s) 1, 9 and 16 are rejected as ineligible subject matter under 35 U.S.C. § 101. Regarding dependent claims 2-3, 5-8, 10-11, 13-15 and 17-20: limitations in these dependent claims have been examined in a similar way as to the above independent claims. It was found that claims 2-3, 5-8, 10-11, 13-15 and 17-20 are ineligible subject matter under 35 U.S.C. § 101: Claims 2 and 10: Ineligible. Routine data organization (grouping by scenes). Claims 3 and 11: Ineligible. Mathematical calculation/mental process (comparing adjacent data frames). Claims 5, 13 and 7: Ineligible. Generic application of a machine-learning model as a tool. Claims 6, 14 and 18: Ineligible. Mental process/abstract idea (transcription and generic LLM summarization). Claims 7, 15 and 20: Ineligible. Mere field of use restriction (advertising). Claims 8 and 19: Ineligible. Organizing human activity/mental process (determining demographics). Claim Rejections - 35 USC § 103 The following is a quotation of pre-AIA 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action: (a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made. Claim(s) 1-3, 7, 9-11, 15-16 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hua et al (US20080192840A1) in view of Dirik et al (US20120106925A1)). Regarding claims 1, 9 and 16, Hua teaches a method comprising: receiving, by a computing device, a content item comprising a plurality of video frames; (Hua, "an exemplary system for implementing embodiments includes a general purpose computing system environment, such as computing system environment 100.", [0020], "The computing system environment 100 may also include a number of audio/video inputs and outputs 118 for receiving and transmitting video content.", [0022], "video clip 230", [0024], "frames of the video clip 230", [0025]; a computing device (computing system environment 100) that receives a content item (video content/video clip) made up of a plurality of video frames) determining a stability level for one or more of the plurality of video frames; (Hua, "analyzing frames of the video clip to determine which frames are stable", [0018]; "The stableness of a frame may be represented by its histogram delta.", [0034]; analyzing the frames to determine their stableness, explicitly noting that this stability level can be calculated and represented numerically by a histogram delta) selecting, based on the determined stability level, a portion of the plurality of video frames, (Hua, "The attribute delta and segment forming portion 2122 is operable to classify the frames of the video clip 230 into two classes (e.g., stable and non-stable) based on their respective attribute deltas ... Once the stable frames have been determined, connecting the stable frames results in a set of stable video segments.", [0026]; classifying and selecting a specific portion of frames (the stable frames/segments) based on their evaluated stability level, represented by the attribute deltas) wherein the portion of the plurality of video frames have a higher stability level relative to other video frames of the plurality of video frames; and (Hua, “In addition, the candidate segments are also relatively more “stable” than other segments.", [0024], "A stable frame is a frame that exhibits a low degree of movement and/or change with respect to the preceding and following frames.", [0025]; this establishes that the selected portion of frames (stable frames forming the candidate segments) possess a higher stability level compared to the non-selected, non-stable video frames) generating, based on the portion of the plurality of video frames, a summary of content in the content item. (Hua, "The representative thumbnail is then selected from among the frames of the candidate segments.", [0018]; Dirik, "The static summary is generated by combining thumbnail images of the selected shots.", [0010]; Hua teaches generating representative thumbnails based on the stable candidate segments to visually represent the video content. However, Hua does not explicitly use the terminology of generating a "summary" of the content. Dirik explicitly teaches generating a static summary of the video content by combining selected thumbnail frames. While Hua teaches extracting and selecting representative stable frames (thumbnails) from a video clip, it lacks the explicit phrasing of aggregating these frames specifically into a "summary" of the content. Dirik explicitly teaches generating a static video summary by combining the selected representative thumbnails) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate Dirik into Hua in order to arrange the extracted stable representative thumbnails into a consolidated static summary, thereby providing the user with a comprehensive and easily scannable overview of the overall video content. The combination of Hua and Dirik also teaches other enhanced capabilities) Regarding claims 2 and 10, the combination of Hua and Dirik also teaches its/their respective base claim(s). The combination further teaches the method of claim 1, wherein selecting the portion of the plurality of video frames comprises: determining a plurality of scenes in the content item; and (Hua, "In addition to selecting a single thumbnail that is representative of an entire program, the technology thus described may also be used to select multiple thumbnails, where each is representative of individual sections, scenes, chapters, etc., of a program.", [0049]; identifying a plurality of segments/divisions in the content item, specifically referring to them as "scenes" within a program) selecting, for each of the plurality of scenes, a video frame, of the plurality of video frames, with a higher stability level relative to other video frames for the respective scene, of the plurality of scenes, for inclusion in the portion of the plurality of video frames. (Hua, "select multiple thumbnails, where each is representative of individual sections, scenes, chapters, etc., of a program.", [0049], "In one embodiment, the thumbnail selector 2200 selects the most stable frame in the highest confidence candidate segment.", [0034]; for the determined scenes (or candidate segments), a video frame is selected as the representative thumbnail based on it being the "most stable frame" (having a higher stability level) relative to the other frames in that segment/scene) Regarding claims 3 and 11, the combination of Hua and Dirik also teaches its/their respective base claim(s). The combination further teaches the method of claim 1, wherein determining the stability level for the one or more of the plurality of video frames comprises: determining a first video frame of the plurality of video frames and a second video frame of the plurality of video frames that immediately precedes the first video frame; and (Hua, "The system parameters tuning portion 2114 calculates, with respect to a particular frame, an attribute delta between that frame and the frame prior to and after it.", [0025]; evaluating a specific frame (first video frame) alongside the frame that comes immediately prior to it (second video frame that immediately precedes it) determining, based on a quantity of changes from the second video frame to the first video frame, the stability level for the first video frame. (Hua, "A stable frame is a frame that exhibits a low degree of movement and/or change with respect to the preceding and following frames ... The system parameters tuning portion 2114 calculates, with respect to a particular frame, an attribute delta between that frame and the frame prior to and after it", [0025]; determining the stability of a frame based on the degree of change (attribute delta/quantity of changes) compared to its preceding frame) Regarding claims 4 and 12, the combination of Hua and Dirik also teaches its/their respective base claim(s). The combination further teaches the method of claim 1, wherein determining the stability level for the one or more of the plurality of video frames comprises: determining a first video frame of the plurality of video frames and a second video frame of the plurality of video frames that immediately follows the first video frame; and (Hua, "The system parameters tuning portion 2114 calculates, with respect to a particular frame, an attribute delta between that frame and the frame prior to and after it", [0025]; evaluating a specific frame (first video frame) alongside the frame that comes after it (second video frame that immediately follows it)) determining, based on a quantity of changes from the first video frame to the second video frame, the stability level for the first video frame. (Hua, "A stable frame is a frame that exhibits a low degree of movement and/or change with respect to the preceding and following frames ... The system parameters tuning portion 2114 calculates, with respect to a particular frame, an attribute delta between that frame and the frame prior to and after it", [0025]; determining the stability level of a frame based on the calculated change (attribute delta) compared to its following frame) Regarding claims 7, 15 and 20, the combination of Hua and Dirik also teaches its/their respective base claim(s). The combination further teaches the method of claim 1, wherein the content item comprises one of an advertisement or an offer of goods or services. (Hua, "Non-program content includes, for instance, commercials, credit sequences, blurry frames, black frames, etc.", [0018]; Hua references commercials only as examples of non-program content to be excluded from consideration when selecting thumbnails, the content item processed by Hua is a television program recording, and commercials are treated as undesirable content. Hua does not teach that the content item itself is an advertisement, however, Dirik teaches, "we introduce a low complexity video summarization method for different types of enterprise videos such as, for example, commercial clips, talk videos, or e-meeting screen and slide share recordings", [0041]; the content item being summarized can be "commercial clips" (an advertisement)) Claim(s) 5-6, 13-14 and 17-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hua et al (US20080192840A1) in view of Dirik et al (US20120106925A1) and further in view of Gardner et al (US12008332B1). Regarding claims 5, 13 and 17, the combination of Hua and Dirik also teaches its/their respective base claim(s). The combination does not expressly disclose but Gardner teaches the method of claim 1, wherein generating the summary of the content in the content item comprises determining, based on the portion of the plurality of video frames and a machine-learning prediction model, the summary of the content. (Gardner, "For image and video inputs, multimedia analysis techniques may be applied to extract salient objects, people, scenes, and text segments from the visual content ... The system's summarization models may be trained to incorporate both textual and visual relevance cues", "After extracting key textual and visual elements from multimedia, the system's summarization engine may condense this content", c12:30-45; extracting visual cues from image and video inputs (video frames) and using summarization models (machine-learning prediction models) trained on these visual relevance cues to condense the content (generate the summary)) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate Dirik into Hua in order to generate highly accurate semantic textual summaries of the video content using the trained system's summarization models for improvement of the baseline video summarization systems. The combination of Hua, Dirik and Gardner also teaches other enhanced capabilities. Regarding claims 6, 14 and 18, the combination of Hua and Dirik also teaches its/their respective base claim(s). The combination of Hua, Dirik and Gardner teaches the method of claim 1, further comprising: determining a plurality of spoken words in the content item; (Gardner, "For audio input such as podcasts, speeches, and phone calls, automated speech recognition may first be used to transcribe the audio to text.", c12:20-30; using speech recognition to transcribe spoken words from the audio inputs of the content item) generating, based on the plurality of spoken words in the content item, a textual representation of the plurality of spoken words; (Gardner, "For audio input such as podcasts, speeches, and phone calls, automated speech recognition may first be used to transcribe the audio to text. The resulting text transcript may then summarized using the system's text summarization capabilities", c12:20-30; the transcribed audio generates a text transcript (textual representation) of the spoken words) determining, one or more words presented in the plurality of video frames of the content item; and (Gardner, "Optical character recognition (OCR) may be used to extract text from image or video files.", c13:1-5; "For image and video inputs, multimedia analysis techniques may be applied to extract salient objects, people, scenes, and text segments from the visual content.", c12:30-35; using OCR to extract text segments (words) directly from the visual content of the video files (video frames)) generating, by a large language model and based on the summary of content in the content item, the textual representation of the plurality of spoken words, and the one or more words presented in the plurality of video frames of the content item, a second summary of the content. (Gardner, "The prompts can include the original text to summarize along with instructions tailored to elicit the target summary characteristics from the LLM.", c6:45-50; "In example embodiments, the system is configured to use such a loop to iteratively re-prompt the LLM API with templates designed for higher abstraction levels each iteration ... The LLM output summary may then be used as the content for the next iteration's increased abstraction prompt.", c51:40-50; utilizing a Large Language Model (LLM) to iteratively generate new summaries. In the iterative loop, the LLM generates a subsequent (second) summary based on the prior summary ("LLM output summary") as well as the original source text, which Gardner previously defines as including both the audio transcripts (spoken words) and OCR text (words presented in video frames)) Claim(s) 8 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hua et al (US20080192840A1) in view of Dirik et al (US20120106925A1) and further in view of Duque et al (US20150081604A1). Regarding claims 8 and 19, the combination of Hua and Dirik also teaches its/their respective base claim(s). The combination does not expressly disclose but Duque teaches the method of claim 1, further comprising determining, based on the summary of the content in the content item, at least one user demographic associated with the content item. (Duque, "A video demographics analysis system produces demographic classifier models that predict the demographic characteristics associated with a video", [0007]; "The features for a video may further include textual metadata, such as the video title and any tag words or phrases assigned to the video.", [0047]; "Another usage scenario is prediction of demographic attribute values for a video, such as newly submitted video.", [0074]; determining user demographics based on analysis/summary of video content (via classifier model predicting demographics from video features)) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate Duque's demographic analysis into Dirik's summarization process in order to enable a system which could intelligently select representative thumbnails that appeal specifically to the predicted target audience (e.g., age, gender) of the newly submitted video, rather than just generating a generic summary. The combination of Hua, Dirik and Duque also teaches other enhanced capabilities. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JIANXUN YANG whose telephone number is (571)272-9874. The examiner can normally be reached on MON-FRI: 8AM-5PM Pacific Time. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached on (571)272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272- 1000. /JIANXUN YANG/ Primary Examiner, Art Unit 2662 8/8/2026
Read full office action

Prosecution Timeline

May 31, 2024
Application Filed
Apr 07, 2026
Interview Requested
May 06, 2026
Interview Requested
May 14, 2026
Applicant Interview (Telephonic)
May 14, 2026
Examiner Interview Summary
Aug 12, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743866
FUNDAMENTAL MATRIX GENERATION APPARATUS, CONTROL METHOD, AND COMPUTER-READABLE MEDIUM
3y 0m to grant Granted Sep 22, 2026
Patent 12743791
SYSTEMS AND METHODS FOR MULTI-BRANCH VIDEO OBJECT DETECTION FRAMEWORK
3y 3m to grant Granted Sep 22, 2026
Patent 12743884
VERSATILE ACTION MODELS (VAMOS) FOR VIDEO UNDERSTANDING
2y 6m to grant Granted Sep 22, 2026
Patent 12738036
IMAGE PROCESSING APPARATUS, IMAGE PROCESSING METHOD, AND IMAGE PROCESSING PROGRAM
2y 5m to grant Granted Sep 15, 2026
Patent 12731362
IMAGE PROCESSING DEVICE, IMAGE PROCESSING METHOD, AND PROGRAM
3y 2m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
74%
Grant Probability
93%
With Interview (+19.3%)
2y 7m (~3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 663 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month