Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status
This instant application No. 19/266,518 has claims 1-20 pending.
Priority
Applicant’s claim for priority of provisional application No. 63/734,067 (filed on December 14, 2024) is acknowledged. The effective filing date is December 14, 2024.
Information Disclosure Statement
As required by M.P.E.P. 609(C), the Applicant’s submission of the Information Disclosure Statement filed on November 5, 2025 is acknowledged by the Examiner and the cited references have been considered in the examination of the claims. As required by M.P.E.P. 609 C(2), a copy of the PTOL-1449 initialed and dated by the Examiner is attached to the instant Office action.
Drawings
The drawings filed on July 11, 2025 are acceptable for examination purposes.
Abstract
The abstract of the disclosure is objected due to the use of implied language. Note that in the abstract, the language should be clear and concise and should not repeat information given in the title. It should avoid using phrases which can be implied, such as, “The disclosure concerns,” “The disclosure defined by this invention,” “The disclosure describes,” etc… See MPEP § 608.01(b).
Note that in the abstract, Applicant cites “…multimodal data processing for content retrieval systems and applications is described herein…” on lines 1-2. This citation clearly repeats the title and invokes implied language. Correction is required (e.g., removal of the entire first sentence of the abstract).
Claim Objections
Claims 18, and 20 are objected for citing …an OS-level virtualization package (e.g., a container)… since the usage of “e.g.” or “for example” renders each claim unclear in scope and the metes and bounds of the element associated with such language cannot be properly assessed and/or determined.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
The claimed invention in claims 1-20 are directed to a judicial exception (i.e., an abstract idea) without significantly more.
Claims 1-20 pass step 1 of the 35 U.S.C. 101 analysis since each claim is either directed to a method, a system comprising one or more processors, or one or more processors comprising processing circuitry (i.e., hardware components such as CPU, GPU per Figure 12).
a. Claims 1, 10, and 19 recite each, in part, steps that are directed to an abstract idea (“Courts have examined claims that required the use of a computer and still found that the underlying, patent-ineligible invention could be performed via pen and paper or in a person’s mind.” Versata Dev. Group v. SAP Am., Inc., 793 F.3d 1306, 1335, 115 USPQ2d 1681, 1702 (Fed. Cir. 2015)) per step 2A – prong 1 since the core of each claim recites steps or functions of generating transcript text from speech using a language model; determining sets of frames and selecting frames from video data based on visual attributes; generating text from selected frames using a vision-language mode; combining speech-derived text and frame-derived text; and generating and storing vector embeddings of the combined text in a database. These limitations fall within the recognized judicial categories of Certain Methods of Organizing Human Activity (i.e., collecting, analyzing, and synthesizing multimodal information) and Mental Processes (i.e., concepts that can be performed in the human mind or with pen/paper such as transcribing speech, reviewing video frames to note visuals, summarizing observations, and filling summaries in an indexed record). The inclusion of automated tools like language models and vision language models merely automates human cognitive functions without altering the fundamental nature of the underlying abstract concepts (Electric Power Group, LLC v. Alstom S.A., 830 F.3d 1350 (Fed. Cir. 2016)).
Per step 2A - prong 2, the claims recite high-level functional concepts without detailing how the underlying processor or computing systems are functioning differently at a physical or hardware level to achieve an improvement in computer functionality. Using generic language models, vision language models, processors, and databases performs routine computer functions (i.e., data processing, pattern recognition, and data storage). Merely invoking off-the-shelf ML tools to carry out generic processing steps does not integrate the abstract idea into a practical application under MPEP 2106.04(d).
Per step 2B, considering the claims elements individually and as an ordered combination, the independent claims fail to recite significantly more than the abstract idea itself. The individual elements processors, video data, audio data, language models, vision language models, and databases are well-understood, routine, and conventional (WURC) components in the art of multi-model data processing. The sequence of steps - processing audio to text, selecting key video frames, describing key video frames with an AI model, concatenating the text, and storing vector embeddings – is a WURC sequence for data indexing and retrieval. The combination adds nothing beyond the abstract idea of organizing and summarizing media content using off-the-shelf software models.
b. Claims 2-3, 11-13, and 15 further recite selecting frame sets based on scene transitions, visual differences, similarity scores, thresholds, chapters, clips, or starting/ending frames.
Per step 2A – prong 2, the additional limitations describe standard mathematical and data manipulation operations (e.g., calculating similarity scores, comparing against threshold, identifying visual differences). These features do not transform the underlying system into a practical application because they recite what logic to apply (i.e., comparing frame differences) using generic mathematical operations rather than a technical improvement to image-processing hardware or low-level signal processing architecture.
Per step 2B, calculating visual differences, comparing similarity scores to a threshold, and dividing video into chapters or clips are WURC techniques in standard digital video editing and video encoding algorithms. Executing conventional scene-cut detection or clip segmentation prior to feeding data into a model does not supply an inventive concept.
c. Claims 4, and 14 specify determining entropy values associated with individual frames, comparing entropy values to a threshold, and selecting frames based on the comparison.
Per step 2A – prong 2, calculating entropy values is a pure mathematical concept (e.g., information entropy / mathematical statistics). Evaluating mathematical entropy values on generic computing hardware without modifying the underlying processor or memory operations fails to integrate the mathematical exception into a practical application.
Per step 2B, mathematical calculation of visual entropy for image feature selection is WURC in image processing and computer vision. Reciting a generic threshold value comparison represents standard algorithmic logic. Taken individually or in combination, using entropy filtering to pare down data sets before model inference fails to provide an inventive concept.
d. Claims 5-6, 8-9, and 16 recite text interleaving via timestamps/associations, chunking text, generating embeddings via encoders, receiving query data, comparing query embeddings to database embeddings, and returning a query response based on RAG systems.
Per step 2A – prong 2, these limitations are directed to additional abstractions such as organizing text documents, matching semantic query vectors, and text synthesis. Applying standard RAG pipelines or timestamp-based text concatenation merely applies generic data indexing and retrieval rules on conventional computers. The ordered combination of running a vector search and generating text output provides no technical improvement to computer functionality or hardware performance and fails to integrate the abstract idea into a practical application.
Per step 2B, matching query vectors to stored embeddings in a database and generating a text response is the standard operational paradigm of search engines and conversational RAG architectures. Chunking text and matching timestamps represent WURC data handling methods which fail to provide an inventive concept.
e. Claims 7, and 17 recite downsampling video data to determine a portion of video frames for subsequent processing.
Per step 2A – prong 2, downsampling is a mathematical data reduction technique. Executing data downsampling on generic video streams does not alter the functional operation of the computer or video processor and remains an unintegrated abstract operation.
Per step 2B, frame rate downsampling (e.g., dropping frames or subsampling) is one of the most WURC preprocessing steps in computer vision. Adding standard downsampling to an abstract multi-modal text generation pipeline does not contribute an inventive concept.
f. Claims 18, and 20 append an extensive list of generic hardware environments (e.g., an autonomous machine, a robot, edge device, cloud computing resources, virtual machines) the system and processor claims.
Per step 2A – prong 2, simply limiting the execution of an abstract idea to a particular technological environment or field of use (e.g., robotics, cloud gaming, autonomous driving, AR/VR) does not integrate the abstract idea into a practical application. See MPEP 2106.05(h).
Per step 2B, broadly invoking generic deployment environments (e.g., data centers, containers, cloud resources, edge devices) amounts to nothing more than telling the practitioner to apply the abstract data processing pipeline on conventional machinery within a specific field of use. This does not supply an inventive concept.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 5-6, 8-11, 15-16, and 18-20 are rejected under AIA 35 U.S.C. 103 as being unpatentable over Mahyar et al. (Pat. No. US 10999566, published on May 4, 2021; hereinafter Mahyar) in view of Rosa (“Video Enriched Retrieval Augmented Generation Using Aligned Video Captions”; published on May 27, 2024).
Regarding claims 1, 10, and 19, Mahyar clearly shows and discloses a method (Abstract); a system comprising: one or more processors to implement the method; and one or more processors comprising: processing circuitry to implement the method (Figure 6) comprising:
generating, using one or more language models and based at least on audio data representing at least speech associated with a video, first text associated with a transcript corresponding to the speech (The audio processing module(s) 330 may be configured to process and/or analyze audio content, such as audible dialog, sound effects, music, and other audio. In some instances, the audio processing module(s) 330 may be configured to convert audio to text and/or perform natural language processing to determine a meaning of certain portions of audio or its corresponding transcription, [Column 10, Lines 29-42]);
determining, based at least on video data representing frames of the video, one or more sets of frames from the frames (the video processing module(s) 320 may be configured to determine frames or sets of frames of video content and may be configured to detect certain features, such as certain objects, as well as actions or events across multiple frames. For example, a video file for a movie may include a first frame, a second frame, and so forth. The video processing module(s) 320 may be configured to detect or analyze frames in video content to determine which frames correspond to the same scene, [Column 10, Lines 4-28]);
determining, based at least on analyzing one or more visual attributes associated with individual frames from the one or more sets of frames, one or more frames from the one or more sets of frames (Using one or more algorithms or modules, the content processing engine 310 may determine the presence of one or more types of objects, faces, and/or scenes in the content, and may output vector data 380. The vector data 380 may include one or more vectors that represent the features of the frames that form a video segment. For example, a video segment of 10 seconds may have 30 frames, and features detected in some or all of the 30 frames may be aggregated and used to generate a vector that may be included in the vector data 380. Video content may include more than one vector for various video segments of the video content, [Column 11, Lines 1-11]);
generating, using one or more vision language models and based at least on the one or more frames, second text associated with the one or more frames (The vector data 380 may be input at a textual description generation engine 390 and/or one or more textual description generation module(s). The textual description generation engine 390 may be configured to generate textual descriptions for video segments using the vector data 380. For example, the textual description generation engine 390 may generate a first textual description using a first vector, and a second textual description using a second vector of the vector data 380, [Column 11, Lines 12-22]);
generating third text associated with the video by at least combining the first text and the second text (The textual description generation engine 120 may generate a vector that combines the individual textual descriptions for the frames and the output of the audio analysis. The second neural network may analyze the individual textual descriptions, aggregate and combine the individual textual descriptions, incorporate analysis of the audio, and output a textual description for the scene and/or segment of video, [Column 4, Lines 36-67]).
Rosa then discloses:
generating third text associated with the video by at least combining the first text and the second text (Figure 2 shows texts of a specific Scene is combined with audio transcript to generate aligned video caption texts, [Page 3]); and
storing, in one or more databases, one or more embeddings associated with the third text (The selected query engine tool vectorizes the query and searches the vector database to retrieve (chunked) aligned video caption text blobs, [Page 4]).
It would have been obvious to an ordinary person skilled in the art at the time of the invention that was effectively filed to incorporate the teachings of Rosa with the teachings of Mahyar for the purpose of enabling efficient, semantic natural language retrieval over video archies without requiring video frames to be constantly reprocessed by computationally expensive ML models at query time thereby reducing context-window overhead and processing resources.
Regarding claims 2, and 11, Mahyar further discloses the determining the one or more sets of frames from the frames comprises:
determining, based at least on the video data, one or more visual differences between one or more consecutive frames from the frames (a video file for a movie may include a first frame, a second frame, and so forth. The video processing module(s) 320 may be configured to detect or analyze frames in video content to determine which frames correspond to the same scene, [Column 10, Lines 4-28]);
determining, based at least on the one or more visual differences, one or more scene transitions within the video; and determining the one or more sets of frames based at least on the one or more scene transitions (a scene may be briefly interrupted by a flashback or cut to a different story, and may resume thereafter. Video processing module(s) 320 may include one or more object recognition algorithms configured to detect at least one of predefined objects, predefined scenery (e.g., certain locations, etc.), and the like, [Column 10, Lines 4-28]).
Regarding claim 5, Rosa further discloses the generating the third text comprises:
determining one or more first timestamps associated with the first text (Figure 2 shows timestamps associated with video scenes #7-#9);
determining one or more second timestamps associated with the second text (Figure 2 shows timestamps associated with audio transcripts of scenes #7-#9); and
generating, based at least on the one or more first timestamps and the one or more second timestamps, the third text by at least inputting one or more portions of the second text into one or more portions of the first text (Figure 2 shows combining texts from a respective scene with audio transcript to create a respective chunk of aligned video caption transcript).
Regarding claim 6, Rosa further discloses the generating the third text comprises:
determining that one or more first portions of the first text are associated with the one or more sets of frames (Figure 2 shows first text portions associated with video scenes #7-#9);
determining that one or more second portions of the second text are associated with the one or more sets of frames (Figure 2 shows second text portions associated with audio transcripts of scenes #7-#9); and
generating the third text by at least combining the one or more first portions of the first text with the one or more second portions of the second text (Figure 2 shows combining texts from a respective scene with texts from audio transcript of the respective scene to create a respective chunk of aligned video caption transcript for that particular scene).
Regarding claim 8, Rosa further discloses:
determining one or more chunks of text that are associated with the third text (Figure 3 shows the aligned video captions are chunked and stored in Vector DB); and
generating, using one or more encoders, the one or more embeddings associated with the one or more chunks of text (The selected query engine tool vectorizes the query and searches the vector database to retrieve (chunked) aligned video caption text blobs, [Page 4, Section 4(2)]).
Regarding claim 9, Rosa further discloses:
receiving input data representing a query associated with the video (In Figure 3 we illustrate the main components in a RAG based AI chat bot application that leverages the aligned video captions to return relevant answers and corresponding video clip source, [Page 3, Section 4]);
generating, based at least on the input data, one or more second embeddings associated with the query (The selected query engine tool vectorizes the query, [Page 4, Section 4(2)]);
determining, based at least on the one or more embeddings stored in the one or more databases and the one or more second embeddings, information associated with the video that is related to the query (The selected query engine tool vectorizes the query and searches the vector database to retrieve (chunked) aligned video caption text blobs, [Page 4, Section 4(2)]); and
generating, based at least on the query and the information, a response associated with the query (The query engine tool interprets the results and summarizes
into a specific pydantic format customized for that answer type; for example, a "how to" response should respond with a bulleted list of steps like in Figure 1, whereas a "place" response would describe a location and why it is notable. Timestamps in retrieved docs help give the application pointers to specific parts of video to enhance user interaction, [Page 4, Section 4(3)]).
Regarding claim 15, Mahyar further discloses the selection of the one or more frames from the one or more sets of frames comprises at least one of: selecting one or more starting frames from the one or more sets of frames; or selecting one or more ending frames from the one or more set of frames (Segments may correspond to events, scenes, and/or other occurrences that may be discrete and/or extractable from the content. In some instances, segments may correspond to certain locations and/or times, certain actors that appear, certain music or sounds, and/or other features of the content. For example, the remote server may determine a first clip or a first segment of a movie using content data associated with the movie, such as video analysis data. The first clip may be a continuous portion of the movie corresponding to a first scene of the movie that occurs from a first timestamp to a second timestamp, [Column 7, Lines 11-21]).
Regarding claim 16, Rosa further discloses generating third text by combining at least a portion of the first text with at least a portion of the second text, wherein the data represents one or more portions of the third text (Figure 2 shows combining texts from a respective scene with texts from audio transcript of the respective scene to create a respective chunk of aligned video caption transcript for that particular scene).
Regarding claims 18, and 20, Mahyar then discloses the system is in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing one or more light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more wireless cellular transmissions using a wireless cellular network; a system that provides one or more cloud gaming applications; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative Al operations; a system for performing one or more conversational Al operations; a system for performing operations using one or more small language models (SLMs); a system for performing operations using one or more large language models (LLMs); a system for performing operations using one or more vision language models (VLMs); a system for performing operations using one or more multi-modal language models (MMLMs); a system for performing operations using one or more vision-language-action (VLA) models; a system for performing one or more conversational Al operations; a system for performing one or more synthetic data generation operations; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; systems using or deploying one or more inference microservices; systems that incorporate deploy one or more machine learning models in a service or microservice along with an OS-level virtualization package (e.g., a container); a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
Claims 3, and 12 are rejected under AIA 35 U.S.C. 103 as being unpatentable over Mahyar in view of Rosa and further in view of Walker et al. (Pat. No. US 6928233, published on August 9, 2005; hereinafter Walker).
Regarding claims 3, and 12, Walker then discloses the determining the one or more sets of frames from the frames comprises:
determining, based at least on the video data, similarity scores indicating visual similarities between consecutive frames from the frames (extracting, from a segment consisting of a sequence of consecutive frames forming together the signal, at least one feature which characterizes the properties of the segment; calculating, using the extracted feature, a criterion for measurement of a similarity between a pair of segments for every extracted feature and measuring a similarity between a pair of segments according to the similarity measurement criterion, [Column 2, Lines 30-47]);
determining one or more threshold scores based at least on the similarity scores; determining that a portion of the similarity scores satisfies the one or more threshold scores (detecting, according to the feature and similarity measurement criterion, two of the segments, whose mutual dissimilarity is less than a predetermined dissimilarity threshold, [Column 2, Lines 30-47]); and
determining the one or more sets of frames based at least on the portion of the similarity scores (grouping the segments into a scene consisting of a sequence of temporally consecutive segments reflecting the semantics of the signal content, [Column 2, Lines 30-47]).
It would have been obvious to an ordinary person skilled in the art at the time of the invention that was effectively filed to incorporate the teachings of Walker with the teachings of Mahyar, as modified by Rosa, for the purpose of detecting two visual segments and/or audio segments whose similarity satisfy a certain threshold and grouping the segments into a scene to enhance search and retrieval of the segments.
Claims 4, and 14 are rejected under AIA 35 U.S.C. 103 as being unpatentable over Mahyar in view of Rosa and further in view of Skinner et al. (Pub. No. US 2020/0042837, published on February 6, 2020; hereinafter Skinner).
Regarding claims 4, and 14, Skinner then discloses the determining the one or more frames from the one or more sets of frames comprises:
determining, based at least on the one or more visual attributes associated with the individual frames from the one or more sets of frames, entropy values associated with the individual frames (expediting the processing of video frames, for instance, by selectively processing a subset of the video frames exhibiting greater than a threshold amount of entropy relative to the preceding frame or frames, [0034]. Segments may be delimited by selected frames. In cases in which subsets of frames are individually selected for processing, some embodiments may form a video segment corresponding to, for example an application window, and there may be multiple overlapping segments in time corresponding to different subsets of the area of the display, with pixels in those different areas, along those different segments, being subject to similar or the same processing, based upon an initially selected frame or region thereof, [0105]);
determining that a portion of the entropy values satisfies a threshold value; and selecting/determining the one or more frames as being associated with the portion of the entropy values (Upon determining that there are more frames to process, some embodiments may select a next frame, as indicated by block 156, and score the inter-frame entropy of that frame relative to earlier frames, as indicated by block 158. Some embodiments may determine whether that score exceeds a threshold, as indicated by block 160. Upon determining that the score does exceed a threshold, some embodiments may designate frames since a previously selected frame as part of the same video segment beginning with that previously selected frame, as indicated by block 162, [0113]).
It would have been obvious to an ordinary person skilled in the art at the time of the invention that was effectively filed to incorporate the teachings of Skinner with the teachings of Mahyar, as modified by Rosa, for the purpose of enhancing pattern matching between different type of visual patterns with a set of frames based on amounts of change in information depicted between the set of frames and an entropy threshold.
Claims 7, and 17 are rejected under AIA 35 U.S.C. 103 as being unpatentable over Mahyar in view of Rosa and further in view of Lu et al. (Pub. No. US 2007/0253594, published on November 1, 2007; hereinafter Lu).
Regarding claims 7, and 17, Lu then discloses determining, based at least on downsampling the video data, a portion of the frames of the video, wherein the determining the one or more sets of frames is based at least on updated video data representing the portion of the frames (Framerate downsampling to the set of common framerates produces multiple groups of frames. The TS is computed over each group of frames, resulting multirate TS. For clarity in this document, the TS computed from a particular group will be labeled by the downsampled framerate of that group. For example, TS6 indicates that it is the TS computed from the group of frames of 6 fps, [0040]).
It would have been obvious to an ordinary person skilled in the art at the time of the invention that was effectively filed to incorporate the teachings of Lu with the teachings of Mahyar, as modified by Rosa, for the purpose of generating a fingerprint for a video object and grouping similar frames based on the generated fingerprints within a sliding window for efficient search and matching of the similar frames.
Claim 13 is rejected under AIA 35 U.S.C. 103 as being unpatentable over Mahyar in view of Rosa and further in view of Lee et al. (Pub. No. US 2024/0362272, filed on April 26, 2024; hereinafter Lee).
Regarding claim 13, Lee then discloses the determination of the one or more sets of frames from the frames comprises:
determining, based at least on the video data, one or more chapters associated with the video (the video analysis system 130 applies the LLM 970 to the dense clip descriptions and the set of prompt embeddings to generate a list of video-level text, including at least one of a title, hashtags, topic, summary, chapters, highlights, dense narrations, and the like of the video, [0131]); and
determining, based at least on the one or more chapters, one or more clips within the one or more chapters, wherein the one or more sets of frames correspond to the one or more clips (FIG. 10 illustrates example screenshots of chapters generated using the clip description model, in accordance with an embodiment. In one example described herein, the video analysis system 130 generates clip descriptions and video-level information for an advertisement video using the process of FIGS. 8 and 9, [0151]-[0165]).
It would have been obvious to an ordinary person skilled in the art at the time of the invention that was effectively filed to incorporate the teachings of Lee with the teachings of Mahyar, as modified by Rosa, for the purpose of generating set of video embeddings by extracting frame data, audio data, or text data from the video content to enhance search and retrieval of matching video embeddings with respect to at least a portion of the query in a latent space.
Pertinent Prior Art
The following references are considered relevant to the claims:
Li et al. (Pub. No. US 2022/0382808) teaches automated product identification within hosted and streamed videos is performed based on video content of a video received at an online video platform and text content associated with the video. A product identification representative of a product featured in the video is determined based on a comparison of the first embeddings against entries of the product candidate index, such as including by a nearest neighbor search responsive to the comparison. An indication of the product identification is then output at the online video platform
Buch et al. (Pub. No. US 2025/0390532) teaches receiving a query relating to a data item that includes multiple data item samples and processing the query and the data item to generate a response to the query. The techniques described include adaptively selecting a subset of the data item samples using a selection neural network conditioned on features of the data item samples and the query. Then processing the subset and query using a downstream task neural network to generate a response to the query.
Contact Information
Any inquiry concerning this communication or earlier communications from the Examiner should be directed to Son Hoang whose telephone number is (571) 270-1752. The Examiner can normally be reached on Monday – Friday (7:00 AM – 4:00 PM).
If attempts to reach the Examiner by telephone are unsuccessful, the Examiner’s supervisor, Sherief Badawi can be reached on (571) 272-9782. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SON T HOANG/ Primary Examiner, Art Unit 2169 July 23, 2026