Prosecution Insights
Last updated: October 02, 2026
Application No. 18/819,643

SMART FRAME SELECTION VIA ACTIVITY-BASED RANKING AND OPTIMIZATION

Non-Final OA §101§103
Filed
Aug 29, 2024
Examiner
ANDERSON, BRODERICK C
Art Unit
2178
Tech Center
2100 — Computer Architecture & Software
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
74%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
198 granted / 266 resolved
+19.4% vs TC avg
Strong +18% interview lift
Without
With
+18.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
18 currently pending
Career history
288
Total Applications
across all art units

Statute-Specific Performance

§101
9.8%
-30.2% vs TC avg
§103
62.7%
+22.7% vs TC avg
§102
18.7%
-21.3% vs TC avg
§112
6.0%
-34.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 266 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Specification The lengthy specification has not been checked to the extent necessary to determine the presence of all possible minor errors. Applicant’s cooperation is requested in correcting any errors of which applicant may become aware in the specification. Applicant is reminded of the proper language and format for an abstract of the disclosure. The abstract should be in narrative form and generally limited to a single paragraph on a separate sheet within the range of 50 to 150 words in length. The abstract should describe the disclosure sufficiently to assist readers in deciding whether there is a need for consulting the full patent text for details. The language should be clear and concise and should not repeat information given in the title. It should avoid using phrases which can be implied, such as, “The disclosure concerns,” “The disclosure defined by this invention,” “The disclosure describes,” etc. In addition, the form and legal phraseology often used in patent claims, such as “means” and “said,” should be avoided. The abstract of the disclosure is objected to because the term “disclosed” is used. A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b). Drawings The drawings filed 8/29/2024 were accepted. Claim Objections Claim 9 is objected to because of the following informalities: “comprises text data of the video stream and of text data of content within the plurality of frames” should be replaced with “comprises text data of the video stream and Appropriate correction is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-4, 6-17, and 19-20 are rejected under 35 U.S.C. 101. Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claims do not fall within at least one of the four categories of patent eligible subject matter because they are directed to an abstract idea without significantly more. The claims recite the abstract idea of generating rankings and determine a plurality of frames. Step 2A, Prong 1 The limitations that describe the generating rankings and determine frames are processes that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. The claims also include elements of processors, circuits, a ranking model, a machine learning model, the buffer storing, and providing the subset of frames, however nothing in the claims precludes the steps from practically being performed in the mind. Step 2A, Prong 2 The judicial exception is not integrated into a practical application because the additional elements regarding processors, circuits, a ranking model, a machine learning model, the buffer storing, and providing the subset of frames are considered insignificant extra-solution activity. These limitations are not considered improvements to the functioning of a technology or technical field. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the extrasolutionary elements are not considered significantly more than just applying the steps of generating rankings and determine frames. Step 2B In addition to the abstract idea, the claims have the processors, circuits, a ranking model, a machine learning model, the buffer storing, and providing the subset of frames, but they represent only well-understood, routine, conventional activity that can be performed on generic computers. The transmitting of data (e.g. providing the subset of frames) has been recognized by the courts as being well-understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. See buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) and MPEP 2106.05(d), subsection II. The storing of data (storing frames in the buffer) has been recognized by the courts as being well-understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. See Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93. The processors, circuits, ranking model, and machine learning model are all considered mere instructions to apply an exemption using generic computer components. Mere instructions to apply an exception using generic computer components cannot provide an inventive concept. The claims are not patent eligible. As per claim 2, this claim has similar determining and buffer storing (the updating of the buffer is just storing data in the buffer again) and is rejected similarly to claim 1. As per claim 3, this claim has similar receiving, generating, and machine-learning model and is rejected similarly to claim 1. As per claim 3, this claim also recites an additional abstract idea of representing (the first subset of frames represent the summarization of the video stream). The representing is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. There are no other additional elements. As per claim 4, this claim has similar determining, machine-learning model, and determination elements and is rejected similarly to claim 1. This claim also has similar generating elements as claim 3 and is rejected similarly to claim 3. This claim also recites an additional abstract idea of detecting actions, objects, or movements. The detecting is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. There are no other additional elements Claim 6 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. As per claim 6, this claim recites additional elements of receiving an encoded bitstream, and decoding the encoded bitstream. (Step 2A, prong 2) The judicial exception is not integrated into a practical application because the additional elements regarding receiving an encoded bitstream, and decoding the encoded bitstream are considered insignificant extra-solution activity. These limitations are not considered improvements to the functioning of a technology or technical field. (Step 2B) The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the receiving an encoded bitstream and decoding the encoded bitstream are not considered significantly more than the judicial exception. The additional elements represent only well-understood, routine, conventional activity that can be performed on generic computer systems. The receiving of data has been recognized by the courts as being well-understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. See buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) and MPEP 2106.05(d), subsection II. McCarthy (US20130195206A1; filed 2012) discloses how well-understood, routine, and conventional the decoding is: McCarthy, paragraphs 1-2: “Video encoding is commonly used to transmit digital video via terrestrial broadcast, via cable TV, or via satellite TV services… A video encoder may be programmed to try to maintain a certain level of video quality so a user viewing the video after decoding is satisfied.” The claims are not patent eligible. As per claim 7, this claim has similar obtaining steps (the obtaining is the result of the decoding of claim 6) and is rejected similarly to claim 6. Claim 8 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. As per claim 8, this claim recites additional abstract ideas of detecting actions or movements, detecting and tracking objects, generating a plurality of rankings, identifying areas, and prioritizing an area, and additional elements of a computer vision model. (Step 2A, Prong 1) The limitations that describe the detecting actions or movements, detecting and tracking objects, generating a plurality of rankings, identifying areas, and prioritizing an area are processes that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. The claims also include elements of a computer vision model, however nothing in the claims precludes the steps from practically being performed in the mind. (Step 2A, prong 2) The judicial exception is not integrated into a practical application because the additional elements regarding a computer vision model are considered insignificant extra-solution activity. These limitations are not considered improvements to the functioning of a technology or technical field. (Step 2B) The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the computer vision model is not considered significantly more than the judicial exception. The computer vision model is considered mere instructions to apply an exemption using generic computer components. Mere instructions to apply an exception using generic computer components cannot provide an inventive concept. The claims are not patent eligible. As per claim 9, this claim has similar metadata (received in claim 1) and is rejected similarly to claim 1. As per claim 10, this claim has similar determining and is rejected similarly to claim 1. As per claim 11, this claim has similar buffer storing (the maintaining of the buffer is just storing data in the buffer again) and is rejected similarly to claim 1. As per claim 12, this claim has similar storing and is rejected similarly to claim 1. Claim 12 also recites an additional element of transferring frames. (Step 2A, prong 2) The judicial exception is not integrated into a practical application because the additional elements regarding transferring frames are considered insignificant extra-solution activity. These limitations are not considered improvements to the functioning of a technology or technical field. (Step 2B) The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the transferring frames are not considered significantly more than the judicial exception. The additional elements represent only well-understood, routine, conventional activity that can be performed on generic computer systems. The sending and receiving of data has been recognized by the courts as being well-understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. See buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) and MPEP 2106.05(d), subsection II. The claims are not patent eligible. Claim 13 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. As per claim 13, this claim recites an additional abstract idea of configuring the buffer. The providing of an configuring is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. There are no other additional elements. Claim 14 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. As per claim 14, this claim recites an additional element of a system. (Step 2A, prong 2) The judicial exception is not integrated into a practical application because the additional elements regarding a system are considered insignificant extra-solution activity. These limitations are not considered improvements to the functioning of a technology or technical field. (Step 2B) The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the system is not considered significantly more than the judicial exception. The system is considered mere instructions to apply an exemption using generic computer components. Mere instructions to apply an exception using generic computer components cannot provide an inventive concept. The claims are not patent eligible. Claim 15 recites substantially similar limitations to claim 1 and is thus rejected along the same rationale. Claim 16 recites substantially similar limitations to claim 2 and is thus rejected along the same rationale. Claim 17 recites substantially similar limitations to claims 3-4 combined, and is thus rejected along the same rationale. Claim 19 recites substantially similar limitations to claims 6-7 combined, and is thus rejected along the same rationale. Claim 20 recites substantially similar limitations to claim 1 and is thus rejected along the same rationale. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-5, 10-11, 13, 15-18, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Saggi et al (US20200372066A1; filed 5/22/2020) in view of Chakraborty et al (US20160070963A1; filed 9/4/2014) and Lee et al (US20250037409A1; filed 7/28/2023). With regards to claim 1, Saggi et al discloses One or more processors comprising: one or more circuits to (Saggi et al, paragraph 27: “at least one processor”): receive a plurality of frames and metadata from a capture device capturing a video stream (Saggi et al, paragraph 9: “the system receives an input, which may include at least one video (e.g., provided as an upload to the system, a link, etc.);” abstract: “Multimedia content may be received and a plurality of frames and audio, visual, and metadata elements associated therewith are extracted from the multimedia content”); generate, using a ranking model (Saggi et al, paragraph 111: “steps 404-408 are a process for training a machine learning model for generating summarization of original content based on optimized aggregation of importance sub-scores into importance scores that are used to define moment candidates for potential inclusion in a final summarization. According to one embodiment, the steps 404-408 are performed by the model service 209 and using one or more training datasets retrieved from training data 216;” the sub-scores are used for the ranking of frames, so the machine learning model is interpreted as a ranking model), a plurality of rankings for the plurality of frames based on a plurality of video parameters of the plurality of frames and the metadata of the video stream, wherein the plurality of rankings correspond to a summarization of the video stream (Saggi et al, paragraph 27: “computing, for each frame, a plurality of sub-scores based on the keyword mapping, the one or more audio elements, the one or more visual elements, and the metadata… merging the one or more top-ranked frames into one or more moments based on a sequential similarity analysis of the determined one or more top-ranked frames, wherein the merging includes aggregating one or more of the one or more of the audio elements, the one or more visual elements, and the metadata of each of the one or more top-ranked frames; and 10) aggregating the one or more moments into a final summarization of the multimedia content”); determine at least one of the plurality of frames to provide… based on the plurality of rankings, wherein the at least one… stores a first subset of frames of the plurality of frames (Saggi et al, paragraph 55: “the system performs a process for identifying and providing key video points, wherein the process includes… 6) ranking the one or more segments based on the importance metric of each; 7) selecting a plurality of the one or more segments that were highly ranked; 8) generating titles and descriptions for the plurality of segments; 9) combining the titles, descriptions, and plurality of segments into a single video or a collection of key moments (e.g., the plurality of segments); and 10) providing the single video or collection of moments to users”). However, Saggi et al does not disclose at least one first buffer… buffer… provide, from the at least one first buffer, the first subset of frames as input to a first machine-learning model. Chakraborty et al teaches at least one first buffer… buffer… provide, from the at least one first buffer (Chakraborty et al, paragraph 32: "As one example, a 20 summary frames may be selected as representative of 1,000,000 or more video data frames streamed over a day on a platform including a CM operating at 30 fps. A circular buffer retaining streamed video data… it may be continuously overwritten in sole reliance of the stored summary images"). It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Saggi et al and Chakraborty et al such that the frames of the summary video are stored in a buffer. This would have reduced the amount of memory used to store the summary data (Chakraborty et al, paragraph 32: “A circular buffer retaining streamed video data may be relatively small, much less than would be required to store all the day's streamed video data frames”). Lee et al teaches the first subset of frames as input to a first machine-learning model (Lee et al, abstract: “a machine-learned adaptive thresholding model configured to receive a query (e.g., text, image, video)… The video analysis system filters video segments for the query that are associated with relevance scores above the predicted threshold generated by the adaptive thresholding model”). It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Saggi et al, Chakraborty et al, and Lee et al such that the video summary could be used as input of a machine learning model. This would have enabled the model to extract desired portions of the video in response to user queries (Lee et al, paragraph 16: “video segments that have predicted relevance scores above the predetermined threshold are determined to be relevant and are provided as a response to the user.”). With regards to claim 2, which depends on claim 1, Saggi et al discloses determine a second subset of frames of the first subset of frames based on metadata of the first subset of frames; and update… based on the second subset of frames (Saggi et al, paragraph 107: “According to one embodiment the metadata file includes a table of timestamps corresponding to the portions of the original content that were extracted and merged to generate the final summarization;” the merging of the metadata of the subsets of frames are interpreted as updating the frames). However, Saggi et al does not disclose update the at least one first buffer. Chakraborty et al teaches update the at least one first buffer (Chakraborty et al, paragraph 32: "A circular buffer retaining streamed video data… it may be continuously overwritten in sole reliance of the stored summary images"). It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Saggi et al, Chakraborty et al, and Lee et al such that the frames of the summary video are stored in a buffer. This would have reduced the amount of memory used to store the summary data (Chakraborty et al, paragraph 32: “A circular buffer retaining streamed video data may be relatively small, much less than would be required to store all the day's streamed video data frames”). With regards to claim 3, which depends on claim 1, Saggi et al discloses wherein the first subset of frames represent the summarization of the video stream (Saggi et al, paragraph 27: “aggregating the one or more moments into a final summarization of the multimedia content”). However, Saggi et al does not disclose receive a query regarding content of the video stream; and generate, using the first machine-learning model, an output based on the first subset of frames, the output comprising a response to the query extracting video content of the video stream. Lee et al teaches receive a query regarding content of the video stream; and generate, using the first machine-learning model, an output based on the first subset of frames, the output comprising a response to the query extracting video content of the video stream (Lee et al, paragraph 14: “In one instance, the video retrieval model is configured to receive a pair of the user query and a respective video segment and generate a relevance score for the pair that indicates how relevant the video segment is to the user query;” paragraph 27: “users of client devices 116 can submit queries to the video analysis system 130 to retrieve video segments or videos that are relevant to the user query”). It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Saggi et al, Chakraborty et al, and Lee et al such that the video summary could be used as input of a machine learning model. This would have enabled the model to extract desired portions of the video in response to user queries (Lee et al, paragraph 16: “video segments that have predicted relevance scores above the predetermined threshold are determined to be relevant and are provided as a response to the user.”). With regards to claim 4, which depends on claim 3, Saggi et al discloses wherein the summarization represented in the plurality of rankings correspond to the determination of the first subset of frames representing one or more temporal or spatial segments of the video stream (Saggi et al, paragraph 111: “steps 404-408 are a process for training a machine learning model for generating summarization of original content based on optimized aggregation of importance sub-scores into importance scores that are used to define moment candidates for potential inclusion in a final summarization. According to one embodiment, the steps 404-408 are performed by the model service 209 and using one or more training datasets retrieved from training data 216;” any frame can be interpreted as representing a temporal segment of a video). However, Saggi et al does not disclose in response to receiving the query, determine a third subset of frames to apply to the first machine-learning model to generate the output based on detecting, using a second machine-learning model, one or more actions, objects, or movements described in the query. Lee et al teaches in response to receiving the query, determine a third subset of frames to apply to the first machine-learning model to generate the output based on detecting, using a second machine-learning model (Lee, paragraph 14: “the video analysis system 130 performs the relevance analysis by applying one or more machine-learned video retrieval models to the query and the videos managed by the video analysis system 130;” one of the other models in the one or more models is interpreted as the second machine-learning model), one or more actions, objects, or movements described in the query (Lee et al, paragraph 15: “For example, a high relevance score indicates that the video retrieval model identified the video segment as being highly relevant to the user query. For example, the video segment may include entities (e.g., objects, persons, scenery) specified in the query, characteristics described in the query, and the like”). It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Saggi et al, Chakraborty et al, and Lee et al such that the user could submit a query to find portions of the video summary with specific objects. This would have enabled the model to extract desired portions of the video in response to user queries (Lee et al, paragraph 16: “video segments that have predicted relevance scores above the predetermined threshold are determined to be relevant and are provided as a response to the user.”). With regards to claim 5, which depends on claim 1, Saggi et al discloses generating, using the ranking model, the plurality of rankings further comprises applying differential weighting to the plurality of video parameters (Saggi et al, paragraph 10: “in various embodiments, the technology described herein may improve on the deficiencies of previous approaches by using weighted outputs of topic transition detection, in combination with weighted outputs of other analyses, to provide novel multimodal summarization generation processes”); and at least one first video parameter is assigned a higher weight according to the ranking model than at least one second video parameter based on the metadata of the plurality of frames (Saggi et al, paragraph 19: “computing, for each frame, a plurality of sub-scores based on the keyword mapping, the one or more audio elements, of the one or more visual elements, and metadata… wherein generating the importance score includes weighting each of the plurality of sub-scores according to predetermined weight values and aggregating the weighted sub-scores”). With regards to claim 10, which depends on claim 1, Saggi et al discloses the first subset of frames is further determined based on a plurality of similarity metrics of the plurality of frames, wherein the plurality of similarity metrics are determined… the first subset of frames is further determined based on a minimum distance metric between the plurality of frames (Saggi et al, paragraph 19: “merging the one or more top-ranked frames into one or more moments based on a sequential similarity analysis of the determined one or more top-ranked frames, wherein the merging includes aggregating one or more of the one or more audio elements, the one or more visual elements, and the metadata of each of the one or more top-ranked frames”). However, Saggi et al does not disclose using at least one of (i) a cosine distance, (ii) a Siamese network, (iii) a structural similarity, or (iv) background subtraction. Chakraborty et al teaches using at least one of (i) a cosine distance, (ii) a Siamese network, (iii) a structural similarity, or (iv) background subtraction (Chakraborty et al, paragraph 49: “the similarity between the feature vectors associated with the pair of frames i and j in the frame set V. While there are many techniques for measuring similarity of two vectors, the inventors have found the cosine similarity metric to work well for HOG feature vectors”). It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Saggi et al, Chakraborty et al, and Lee et al such that the distance between frames is measured using a cosine distance. This measurement was found to “work well” (Chakraborty et al, paragraph 49: “the inventors have found the cosine similarity metric to work well for HOG feature vectors”). With regards to claim 11, which depends on claim 1, Saggi et al does not disclose maintain the at least one first buffer containing a predetermined maximum number of frames based on the plurality of rankings. Chakraborty et al teaches maintain the at least one first buffer containing a predetermined maximum number of frames based on the plurality of rankings (Chakraborty et al, Paragraph 45: "With k slots available, k incumbent stream summary frames summarize any number of prior frame sets 471 that were exposed and processed through the summarization process earlier in time. For example, a snapshot of stream summary 465 includes incumbent frame i from set V most recently processed, incumbent frame i+j from a frame set V−3, etc. Looking forward in time, any number of new frame sets 472 will be exposed and processed through the summarization process later in time (e.g., beginning with set V, and ending with V+m). In response to receiving each new frame set (e.g., V, V+1, etc.) a summarization iteration is performed where the incumbent k stream summary frames and n non-incumbent frames are the batch of candidate frames for selection through application of an objective function 466. With each iteration, one or more incumbent frame may retain a slot within stream summary 465, and one or more incumbent frame may be evicted from stream summary 465 in preference of a non-incumbent frame included in a new set (e.g., set V+1)"). It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Saggi et al, Chakraborty et al, and Lee et al such that the frames of the summary video are stored in a buffer with a fixed maximum size. This would have reduced the amount of memory used to store the summary data (Chakraborty et al, paragraph 32: “A circular buffer retaining streamed video data may be relatively small, much less than would be required to store all the day's streamed video data frames”). With regards to claim 13, which depends on claim 1, Saggi et al discloses wherein the video stream is at least one of a live stream or an offline stream stored in a file (Saggi et al, paragraph 9: “the system may receive an informational video (for example, a webinar recording), as an upload”). However, Saggi et al does not disclose wherein the one or more circuits are to: configure the at least one first buffer for the live stream or the offline stream to perform frame storage, wherein the at least one first buffer is configured to perform at least one of: (i) a circularity process on the first subset of frames stored in the at least one first buffer, (ii) segmenting of the video stream into one or more segments comprising a fourth subset of frames of the plurality of frames based on at least one segmentation parameter, or (iii) storing a fifth subset of frames of the plurality of frames from a previous segment of the one or more segments and updating the fifth subset of frames based on an updating parameter. Chakraborty et al teaches wherein the one or more circuits are to: configure the at least one first buffer for the live stream or the offline stream to perform frame storage, wherein the at least one first buffer is configured to perform at least one of: (i) a circularity process on the first subset of frames stored in the at least one first buffer (Chakraborty et al, paragraph 32: "A circular buffer retaining streamed video data… it may be continuously overwritten in sole reliance of the stored summary images"), (ii) segmenting of the video stream into one or more segments comprising a fourth subset of frames of the plurality of frames based on at least one segmentation parameter, or (iii) storing a fifth subset of frames of the plurality of frames from a previous segment of the one or more segments and updating the fifth subset of frames based on an updating parameter. It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Saggi et al, Chakraborty et al, and Lee et al such that the frames of the summary video are stored in a buffer. This would have reduced the amount of memory used to store the summary data (Chakraborty et al, paragraph 32: “A circular buffer retaining streamed video data may be relatively small, much less than would be required to store all the day's streamed video data frames”). Claim 15 recites substantially similar limitations to claim 1 and is thus rejected along the same rationale. Claim 16 recites substantially similar limitations to claim 2 and is thus rejected along the same rationale. Claim 17 recites substantially similar limitations to claims 3-4 combined, and is thus rejected along the same rationale. Claim 18 recites substantially similar limitations to claim 5 and is thus rejected along the same rationale. Claim 20 recites substantially similar limitations to claim 1 and is thus rejected along the same rationale. Claim(s) 6-7 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Saggi et al in view of Chakraborty et al and Lee et al, and further in view of McCarthy (US20130195206A1; filed 1/31/2012). With regards to claim 6, which depends on claim 1, Saggi et al, Chakraborty et al, and Lee et al do not disclose receive an encoded bitstream of the video stream; and decode the encoded bitstream to extract the plurality of frames, the plurality of video parameters, and the metadata of the video stream. However, McCarthy teaches receive an encoded bitstream of the video stream (McCarthy, paragraph 21: “The video encoding system 100 receives a video sequence 101. The video sequence 101 may be included in a video bitstream and includes frames or pictures which may be stored for encoding.”); and decode the encoded bitstream to extract the plurality of frames (McCarthy, paragraph 26: “The video decoding system 302 includes a receiver buffer 350, a decoding unit 351”), the plurality of video parameters, and the metadata of the video stream (McCarthy, paragraph 23: “The perceptual engine 120 also may generate video quality metadata 103 including video quality metrics according to perceptual representations generated for the encoded pictures. The video quality metadata may be included in or associated as metadata with the compressed video bitstream output by the video encoding system 100.” Paragraph 31: “The receiver buffer 350 of the video decoding system 302 may temporarily store encoded information including motion vectors, residual pictures and video quality metadata from the video encoding system 301”). It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Saggi et al, Chakraborty et al, Lee et al, and McCarthy such that the video stream is encoded for the transmission. This would have reduced bandwidth usage (McCarthy, paragraph 2: “A common goal for video compression is to minimize bandwidth for video transmission while maintaining video quality”). With regards to claim 7, which depends on claim 6, Saggi et al, Chakraborty et al, and Lee et al do not disclose wherein the plurality of video parameters comprise at least one of: one or more motion vectors obtained from the encoded bitstream, the one or more motion vectors corresponding to movement data of one or more objects in the plurality of frames; instantaneous decoder refresh (IDR) frames or scene change indicators obtained from the encoded bitstream, the IDR frames or scene change indicators corresponding to content updates in the plurality of frames; one or more bitrate variations obtained from the encoded bitstream, the one or more bitrate variations corresponding to data rate updates used to encode the video stream; or one or more optical flow motion vectors obtained from the encoded bitstream, the one or more optical flow motion vectors corresponding to movement data of one or more objects in consecutive frames of the plurality of frames. However, McCarthy teaches wherein the plurality of video parameters comprise at least one of: one or more motion vectors obtained from the encoded bitstream, the one or more motion vectors corresponding to movement data of one or more objects in the plurality of frames (McCarthy, paragraph 23: “the encoding unit generates motion vectors and predicted pictures according to a video encoding format;” paragraph 25: “The video encoding system 301 may transmit a compressed video bitstream 305, including motion vectors and other information”); instantaneous decoder refresh (IDR) frames or scene change indicators obtained from the encoded bitstream, the IDR frames or scene change indicators corresponding to content updates in the plurality of frames; one or more bitrate variations obtained from the encoded bitstream, the one or more bitrate variations corresponding to data rate updates used to encode the video stream; or one or more optical flow motion vectors obtained from the encoded bitstream, the one or more optical flow motion vectors corresponding to movement data of one or more objects in consecutive frames of the plurality of frames. It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Saggi et al, Chakraborty et al, Lee et al, and McCarthy such that the video stream is encoded for the transmission. This would have reduced bandwidth usage (McCarthy, paragraph 2: “A common goal for video compression is to minimize bandwidth for video transmission while maintaining video quality;” paragraph 28: “Perceptual representations, motion vectors, predicted pictures and video quality metadata may be derived from the pictures in video sequence 101;” (the predicted pictures don’t require as much transmission, which further reduces bandwidth usage). Claim 19 recites substantially similar limitations to claims 6-7 combined, and is thus rejected along the same rationale. Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Saggi et al in view of Chakraborty et al and Lee et al, and further in view of Takata et al (US20040146275A1; filed 1/16/2004). With regards to claim 9, which depends on claim 1, Saggi et al discloses the metadata of video stream comprises text data of the video stream and of text data of content within the plurality of frames (Saggi et al, paragraph 104: “In various embodiments, metadata of each of moment is augmented with a transcript of the moment and/or a list of top keywords determined from the original content”), the text data of the video stream and of the content comprises at least a type of video (Saggi et al, paragraph 92: “In at least one embodiment, metadata may include, but is not limited to, timestamps, tables of contents, transcripts, titles, descriptions, and other information related to the audio-visual elements or other elements of the original content”). However, Saggi et al, Chakraborty et al, and Lee et al do not disclose an event type being videoed. Takata et al teaches an event type being videoed (Takata et al, Paragraph 79: “FIG. 3 shows an example of video data and metadata attached thereto. The figure shows that metadata is attached to a series of frames included in the video data. The metadata is information representing the contents and features of each piece of data, such as a type of event”). It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Saggi et al, Chakraborty et al, Lee et al, and Takata et al such that the metadata included additional information about the video content. This would have enabled the invention to track additional useful data for multimedia objects (Takata et al, paragraph 46: “A transition effect can be set by using metadata attached to video data. The metadata includes contents of multimedia data and is used in an application such as search, the contents can being described based on a schema standardized by MPEG-7 or the like”). Claim(s) 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Saggi et al in view of Chakraborty et al and Lee et al, and further in view of Panda et al (US20220129679A1; filed 10/27/2020). With regards to claim 14, which depends on claim 1, Saggi et al, Chakraborty et al, and Lee et al do not disclose wherein the plurality of processors are comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system implemented using a robot; an aerial system; a medical system; a boating system; a smart area monitoring system; a system for performing deep learning operations; a system for performing simulation operations; a system for generating or presenting virtual reality (VR) content, augmented reality (AR) content, or mixed reality (MR) content; a system for performing digital twin operations; a system implemented using an edge device; a system incorporating one or more virtual machines (VMs); a system for generating synthetic data; a system implemented at least partially in a data center; a system for performing conversational artificial intelligence (AI) operations; a system for performing generative AI operations; a system implementing language models; a system implementing vision language models (VLMs); a system implementing large language models (LLMs); a system implementing multi-modal language models; a system for hosting one or more real-time streaming applications; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; or a system implemented at least partially using cloud computing resources. However, Panda et al teaches wherein the plurality of processors are comprised in at least one of: … a system implemented at least partially using cloud computing resources (Panda et al, paragraph 34: "one or more elements of the present techniques can optionally be provided as a service in a cloud environment."). It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Saggi et al, Chakraborty et al, Lee et al, and Panda et al such that at least some of the processing is performed using cloud computing. This would have enabled the invention to take advantage of high-powered processing devices (Panda et al, paragraph 34: "the target and related videos segmentation, the feature vector representation, segment scoring, consensus estimation and/or summary generation can be performed on a dedicated cloud server to take advantage of high-powered CPUs and GPUs, after which the result is sent back to the local device"). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Liu (US20160092561A1): Teaches video summarization and metadata usage. Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRODERICK C ANDERSON whose telephone number is (313)446-6566. The examiner can normally be reached Monday-Tuesday, Thursday-Saturday 9-5 PST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Stephen Hong can be reached at 5712724124. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /B.C.A/Examiner, Art Unit 2178 /STEPHEN S HONG/Supervisory Patent Examiner, Art Unit 2178
Read full office action

Prosecution Timeline

Aug 29, 2024
Application Filed
Jul 02, 2026
Non-Final Rejection mailed — §101, §103
Sep 22, 2026
Applicant Interview (Telephonic)
Sep 22, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12718000
Systems and methods for processing designs
1y 6m to grant Granted Aug 25, 2026
Patent 12710971
Digital Character Interactions with Media Items in a Conversational Session
3y 2m to grant Granted Aug 18, 2026
Patent 12705543
LEARNING MACHINE TRAINING BASED ON PLAN TYPES
2y 8m to grant Granted Aug 11, 2026
Patent 12682324
INTELLIGENT SERVICE REQUEST SUBMITTAL SYSTEM
3y 1m to grant Granted Jul 14, 2026
Patent 12681618
REMOTE DEVICE INPUT ENTRY
2y 8m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
74%
Grant Probability
93%
With Interview (+18.3%)
2y 11m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 266 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month