Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This office action is sent in response to Applicant’s communication received on
10/09/2024 for the application number 18911114. The office hereby acknowledges receipt of the
following placed of record in the file: Specification, Abstract, Oath/Declaration and claims.
Status of the claims
Claims 1-20 are presented for examination.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 1/29/2026 was filed before the mailing date of the first office action. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 20 rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter because the claim recites a "computer program product" which in light of the disclosure, is directed to software per se.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because as explained below.
Claim 1 recites a computer-implemented video processing method comprising:
receiving, at a question-answering system, a question input conveying a question about a video, and initial video data comprising frames obtained from the video at a first frame rate;
determining, by the question-answering system, whether or not the question can be answered using the initial video data,
receiving, at the question-answering system, further video data comprising frames obtained from the video at a second frame rate, wherein the second frame rate is higher than the first frame rate;
determining, by the question-answering system, a question-answering output, wherein the question-answering output conveys an answer to the question, wherein determining the question-answering output comprises the question-answering system processing the further video data;
outputting the question-answering output from the question-answering system.
Step (a) comprises a mental process. This step can be performed by a human as a person
can receive a question as well as view a video at a first frame rate.
Step (b) comprises a mental process. This step can be performed by a human as a person can watch a video and determine whether the data answers a question.
Step (c) comprises a mental process. This step can be performed by a human as person can receive a second video in a higher frame rate and view it.
Step (d) comprises a mental process. This step can be performed by a human as person can as a human can determine whether a video has sufficient information to answer a predetermined question.
Step (e) comprises a mental process. This step can be performed by a human as person can provide an answer to an initial question.
Step 1: This part of the eligibility analysis evaluates whether the claim falls within any
statutory category. See MPEP 2106.03. The claim recites at least system. Thus, the claim is a
method, which is one of the statutory categories of invention. (Step 1: YES). Step 2A, Prong
One: This part of the eligibility analysis evaluates whether the claim recites a judicial exception. As explained in MPEP 2106.04, subsection II, a claim “recites” a judicial exception when the
judicial exception is “set forth” or “described” in the claim. As discussed above, the broadest
reasonable interpretation of steps (a)-(e) recites a mental process.
Specifically, step (a) can be performed by a human as a person can receive a question as well as view a video at a first frame rate.
Step (b) can be performed by a human as a person can watch a video and determine whether the data answers a question.
Step (c) can be performed by a human as person can receive a second video in a higher frame rate and view it.
Step (d) can be performed by a human as person can as a human can determine whether a video has sufficient information to answer a predetermined question.
Step (e) can be performed by a human as person can provide an answer to an initial question.
Hence the claim encompasses mental processes practically performed in the human mind
by observation, evaluation, judgement, and opinion. See MPEP 2106.04(a)(2), subsection III.
(Step 2A, Prong One: YES).
Step 2A, Prong Two: This part of the eligibility analysis evaluates whether the claim as a
whole integrates the recited judicial exception into a practical application of the exception or
whether the claim is “directed to” the judicial exception. This evaluation is performed by (1)
identifying whether there are any additional elements recited in the claim beyond the judicial
exception, and (2) evaluating those additional elements individually and in combination to
determine whether the claim as a whole integrates the exception into a practical application. See
MPEP 2106.04(d).
The claim recited additional elements including a question-answering system and video data. However, these elements are recited at a high level of generality and perform generic computer functions, such as receiving data, processing data, generating content, and providing output. The use of these elements to receive a query, and a video, and determine qualities about said video merely automates the mental processes described above using generic computer components. Such implementation does not impose any meaningful limit on the judicial exception. Accordingly, these elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. (Step 2A, Prong Two: NO), and the claim is directed to the judicial exception. (Step 2A: YES)
Step 2B: This part of the eligibility analysis evaluates whether the claim as a whole
amounts to significantly more than the recited exception i.e., whether any additional element, or combination of additional elements, adds an inventive concept to the claim. As discussed with
respect to Step 2A, Prong Two, question-answering system and video data comprise additional elements that perform well-understood, routine, and conventional activities in the field such as
receiving data, processing data, generating content, and providing output. See MPEP 2106.05(g).
As known in the art these elements are well understood, routine, and conventional functions of a
computing device. Even when considered in combination these additional elements merely
implement the abstract idea using generic computer components and perform insignificant extra -
solutional activity, which does not provide an inventive concept. The claim is not patent eligible.
Claim 2 recites a mental process as a human can request additional information to answer a question.
Claim 3 recites a mental process as a human can identify one or more parameters for obtaining additional information.
Claim 4 recites a mental process as a human can identify one or more time segments from which information should be obtained.
Claim 5 recites a mental process as a human can as a human can identify a frame rate for obtaining information.
Claim 6 recites a mental process as a human can as a human can request additional information and receive the requested information.
Claim 7 recites a mental process as a human can utilize a visual language model or other reasoning tool to answer a question.
Claim 8 recites a mental process as a human can determine that sufficient information is available, formulate an answer, and communicate the answer.
Claim 9 recites a mental process as a human can determine an answer based on tokenized video date representing obtained information.
Claim 10 recites a mental process as a human can determine an answer based on tokenized video data representing additional obtained information.
Claim 11 recites a mental process as a human can determine whether available information is sufficient to answer a question.
Claim 12 recites a mental process as a human can determine whether additional information is sufficient to answer a question.
Claim 13 recites a mental process as a human can determine whether additional information is sufficient to answer a question before obtaining further information.
Claim 14 recites a mental process as a human can iteratively obtain additional information and repeatedly evaluate whether a question can be answered.
Claim 15 recites a mental process as a human can determine whether subsequently obtained information is sufficient to answer a question.
Claim 16 recites a mental process as a human can select increasingly detailed information by increasing the frame rate used to obtain information.
Claim 17 recites a mental process as a human can increase the amount of information obtained by utilizing a higher frame rate.
Claim 18 recites a mental process as a human can as a human can identify one or more time segments from which subsequently obtained information should be extracted.
Claim 19 & 20 recite substantially the same limitations as claim 1, but in different
statutory categories (system and computer program product). Accordingly, they are directed to
the same abstract idea as claim 1.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. ("VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos", arXiv:2405.19209v1, May 29, 2024, 20 pages.) in view of Bersak et al. (US 20250131221 A1)
Regarding claim 1, Wang teaches A computer-implemented video processing method, comprising (Fig. 1 teaches the system as a whole): receiving, at a question-answering system, a question input conveying a question about a video (Sec 3.1: “given the video and a query about it”), and initial video data comprising frames obtained from the video at a first frame rate (Abstract, “extracts query-relevant information from the input video through an iterative process”); determining, by the question-answering system, whether or not the question can be answered using the initial video data (Sec 3.1, Relevance Scoring, “If the number of highly relevant clusters is below the requirement, that indicates the information extracted from the current set of keyframes is insufficient for the LLM to answer the video query”), and if it is determined that the question cannot be answered using the initial video data: receiving, at the question-answering system, further video data comprising more detail (Pg. 4 Left col., “We increase the number of clusters k by double the original number and repeat the clustering, captioning, and relevance scoring operations.”), and determining, by the question-answering system, a question-answering output (Fig. 2, shows the output of the system), wherein the question-answering output conveys an answer to the question (Fig. 2, the output is “predicted answer”), wherein determining the question-answering output comprises the question-answering system processing the further video data (Pg. 4 Left col., “repeat the clustering, captioning, and relevance scoring operations.”); and outputting the question-answering output from the question-answering system. (Fig. 2).
Wang does not teach further data comprising frames obtained from the video at a second frame rate, wherein the second frame rate is higher than the first frame rate.
However, Bersak teaches further data comprising frames obtained from the video at a second frame rate, wherein the second frame rate is higher than the first frame rate. (Para 0071, “Then, at the step 400-2, the frame rate is gradually increased to be higher than a predetermined threshold”).
It would have been obvious to a person of ordinary skill in the art to modify Wang before the effective filing date in such a way as to incorporate the teachings of Bersak in order to use the second frame rate acquisition of Bersak to provide additional video information for improved question answering (0071).
Regarding claim 2, Wang teaches wherein the question-answering system comprises a planner configured for requesting the further video data (Sec 3.1, VIDEOTREE iterate these operations until getting enough query-relevant information from long videos in an adaptive manner).
Regarding claim 3, Wang teaches the planner identifying one or more parameters for the further video data. (Sec 3.1 Relevance Scoring, “we first feed all cluster captions from the last operation and the video query Q into the LLM and output a set of relevance scores for each cluster.”).
Regarding claim 4, Wang teaches wherein the one or more parameters identify one or more time segments of the video from which the frames obtained from the video at the second frame rate are to be extracted. (Sec 3.2 Relevance-Guided Depth Expansion, “we expand the depth of the tree by sub-clustering the clusters with higher relevance scores from the first step.”).)
Regarding claim 5, Wang does not teach wherein the one or more parameters identify the second frame rate.
However, Bersak teaches wherein the one or more parameters identify the second frame rate. (Para 0071, “Then, at the step 400-2, the frame rate is gradually increased to be higher than a predetermined threshold” wherein this predetermined threshold comprises a parameter).
It would have been obvious to a person of ordinary skill in the art to modify Wang before the effective filing date in such a way as to incorporate the teachings of Bersak in order to use the second frame rate acquisition of Bersak to provide additional video information for improved question answering (Para 0071).
Regarding claim 6, Wang teaches the planner generating an output to request the further video data from a video content provision system (3.1, “VIDEOTREE iterate these operations
until getting enough query-relevant information from long videos in an adaptive manner.” Wherein the LLM comprises the planner).
Wang does not teach wherein the video content provision system is configured to extract frames from the video at the second frame rate.
However, Bersak teaches wherein the video content provision system is configured to extract frames from the video at the second frame rate. (Para 0071, “Then, at the step 400-2, the frame rate is gradually increased to be higher than a predetermined threshold”).
It would have been obvious to a person of ordinary skill in the art to modify Wang before the effective filing date in such a way as to incorporate the teachings of Bersak in order to use the second frame rate acquisition of Bersak to provide additional video information for improved question answering (0071).
Regarding claim 7, Wang teaches wherein the planner comprises a Visual Language Model (VLM). (Introduction, “For each cluster, we obtain a single keyframe and pass it to a VLM captioner.”).
Regarding claim 8, Wang teaches wherein if it is determined that the question can be answered using the initial video data then the method comprises: determining, by the question answering system, a question-answering output using the initial video data (Fig. 2, If there are enough high relevance clusters it simply continues to eventually outputting an answer), wherein the question answering output conveys an answer to the question (Fig. 2, output is predicted answer); and outputting the question-answering output determined using the initial video data from the question-answering system. (Fig. 2).
Regarding claim 9, Wang teaches wherein determining, by the question-answering system, the question-answering output using the initial video data comprises the question-answering model determining the answer to the question based on tokenized video data derived from the initial video data. (3.1 Visual Clustering, “we first propose a visual clustering operation that groups similar frames together before extracting information from them” where partitioning the video data into clusters for processing comprises tokenization of the video data)
Regarding claim 10 Wang teaches wherein the question-answering system processing the further video data comprises the question-answering model determining the answer to the question based on tokenized video data derived from the further video data. (3.2 Relevance-Guided Depth Expansion, “we expand the depth of the tree by sub-clustering the clusters
with higher relevance scores from the first step.” Where further breaking down portions of the video comprises tokenizing the video data).
Regarding claim 11, Wang teaches wherein determining whether or not the question can be answered using the initial video data comprises the planner determining whether or not the question can be answered using the initial video data. (3.1, Relevance Scoring, “We use the LLM to decide whether the extracted key information from each cluster is sufficient for the LLM to answer the given query”).
Regarding claim 12, Wang teaches wherein the question-answering system processing the further video data comprises determining, by the question-answering system, whether or not the question can be answered using the further video data. (3.1, Relevance Scoring, “If the number of highly relevant clusters is below the requirement, that indicates the information extracted from the current cluster assignment is insufficient”.
Regarding claim 13, Wang teaches wherein the question-answering system processing the further video data comprises the planner determining whether or not the question can be answered using the further video data. (3.1, Relevance Scoring, “We use the LLM to decide whether the extracted key information from each cluster is sufficient for the LLM to answer the given query”).
Regarding claim 14, Wang teaches wherein processing the further video data forms part of an iterative loop (3.1 Relevance Scoring, “we adaptively extract the query-relevant information within the video by iterating the clustering, captioning, and relevance scoring operation”), wherein the iterative loop comprises the planner identifying subsequent parameters for extracting subsequent frames from the video (3.1 Relevance Scoring, “We increase the number of clusters k and repeat the clustering, captioning, and relevance scoring operations.”) and the question-answering system receiving the subsequent frames and determining whether or not the question can be answered using the subsequent frames (3.1 Relevance Scoring, “We use the LLM to decide whether the extracted key information from each cluster is sufficient for the LLM to answer the given query”), wherein the iterative loop is performed until the question-answering system determines that the question can be answered using the subsequent frames. (3.1 Relevance Scoring).
Regarding claim 15, Wang teaches wherein determining whether or not the question can be answered using the subsequent frames comprises the planner determining whether or not the question can be answered using the subsequent frames. (3.1 Relevance Scoring, “We use the LLM to decide whether the extracted key information from each cluster is sufficient for the LLM to answer the given query”)
Regarding claim 16, Wang teaches wherein each time the iterative loop is performed, using greater and greater degree of details.
Wang does not teach wherein the subsequent parameters comprise a subsequent frame rate for extracting the frames from the video, and wherein each time the iterative loop is performed the planner increases the subsequent frame rate.
However, Bersak teaches wherein the subsequent parameters comprise a subsequent frame rate for extracting the frames from the video (Para 0071, “Then, at the step 400-2, the frame rate is gradually increased to be higher than a predetermined threshold” where the predetermined threshold comprises a parameter).
It would have been obvious to a person of ordinary skill in the art to modify Wang before the effective filing date in such a way as to incorporate the teachings of Bersak in order to use the second frame rate acquisition of Bersak to provide additional video information for improved question answering (0071).
Regarding claim 17, Wang does not teach wherein the subsequent frame rate is higher than the second frame rate.
However, Bersak teaches wherein the subsequent frame rate is higher than the second frame rate. (Para 0071, “Then, at the step 400-2, the frame rate is gradually increased to be higher than a predetermined threshold”).
It would have been obvious to a person of ordinary skill in the art to modify Wang before the effective filing date in such a way as to incorporate the teachings of Bersak in order to use the second frame rate acquisition of Bersak to provide additional video information for improved question answering (0071).
Regarding claim 18, Wang teaches wherein each time the iterative loop is performed the planner identifies one or more time segments of the video from which the subsequent frames are to be extracted. (3.1 Relevance-Guided Depth Expansion, “we expand the depth of the tree by sub-clustering the clusters with higher relevance scores from the first step.”)
Claim 19 & 20 are analogous to claim 1 in that they recite substantially the same
limitations. They are therefore rejected for the same reasons set forth above.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL ALAN FOSTER JR. whose telephone number is (571)272-8874. The examiner can normally be reached M - F 8:00am - 6:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at (571) 272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL A FOSTER JR/Examiner, Art Unit 2654
/HAI PHAN/Supervisory Patent Examiner, Art Unit 2654