Prosecution Insights
Last updated: August 17, 2026
Application No. 18/894,746

VIDEO MANAGEMENT IN AN INFORMATION PROCESSING SYSTEM

Non-Final OA §101§103
Filed
Sep 24, 2024
Examiner
ROBERTS, RACHEL L
Art Unit
2674
Tech Center
2600 — Communications
Assignee
Dell Products L.P.
OA Round
1 (Non-Final)
76%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 76% — above average
76%
Career Allowance Rate
25 granted / 33 resolved
+13.8% vs TC avg
Strong +32% interview lift
Without
With
+32.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
26 currently pending
Career history
61
Total Applications
across all art units

Statute-Specific Performance

§101
12.0%
-28.0% vs TC avg
§103
62.5%
+22.5% vs TC avg
§102
7.7%
-32.3% vs TC avg
§112
12.5%
-27.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 33 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The IDS dated 09/24/2024 has been considered and placed in the application file. Claim Interpretation The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification. Under MPEP 2143.03, "All words in a claim must be considered in judging the patentability of that claim against the prior art." In re Wilson, 424 F.2d 1382, 1385, 165 USPQ 494, 496 (CCPA 1970). As a general matter, the grammar and ordinary meaning of terms as understood by one having ordinary skill in the art used in a claim will dictate whether, and to what extent, the language limits the claim scope. Language that suggests or makes a feature or step optional but does not require that feature or step does not limit the scope of a claim under the broadest reasonable claim interpretation. In addition, when a claim requires selection of an element from a list of alternatives, the prior art teaches the element if one of the alternatives is taught by the prior art. See, e.g., Fresenius USA, Inc. v. Baxter Int’l, Inc., 582 F.3d 1288, 1298, 92 USPQ2d 1163, 1171 (Fed. Cir. 2009). Claim 1, Claim 4, Claim 7, Claim 8, Claim 11, Claim 14, Claim 15, Claim 18, and Claim 20 recite “one or more” then listing “contextual attributes” and “ detected contextual attributes”. Since “one or more” is disjunctive, any one of the elements found in the prior art is sufficient to reject the claim. While citations have been provided for completeness and rapid prosecution, only one element is required. Because, on balance, it appears the disjunctive interpretation enjoys the most specification support and for that reason the disjunctive interpretation (one of A, B OR C) is being adopted for the purposes of this Office Action. Applicant’s comments and/or amendments relating to this issue are invited to clarify the claim language and the prosecution history. Claim 4, Claim 11, and Claim 18 recite “one or more” then listing “of text appearing in the video, a face appearing in the video, an object appearing in the video, and a color appearing in the video”. Since “one or more” is disjunctive, any one of the elements found in the prior art is sufficient to reject the claim. While citations have been provided for completeness and rapid prosecution, only one element is required. Because, on balance, it appears the disjunctive interpretation enjoys the most specification support and for that reason the disjunctive interpretation (one of A, B OR C) is being adopted for the purposes of this Office Action. Applicant’s comments and/or amendments relating to this issue are invited to clarify the claim language and the prosecution history. Claim 7, Claim 14, and Claim 20 and recite “one or more” then listing “one or more metadata derived contextual references, one or more video derived contextual references, one or more audio derived contextual references, and one or more video classification references”. Since “one or more” is disjunctive, any one of the elements found in the prior art is sufficient to reject the claim. While citations have been provided for completeness and rapid prosecution, only one element is required. Because, on balance, it appears the disjunctive interpretation enjoys the most specification support and for that reason the disjunctive interpretation (one of A, B OR C) is being adopted for the purposes of this Office Action. Applicant’s comments and/or amendments relating to this issue are invited to clarify the claim language and the prosecution history. Claim Objections Claim 1 is objected to because of the following informalities: Line 3 of Claim 1 currently reads “detecting one or more contextual attributes in the extracted set of frames frames” Examiner believes that the repetition of the word “frames” is a mistake and Line 3 of Claim 1 should read “detecting one or more contextual attributes in the extracted set of frames”. Appropriate correction is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-4, 7, 8-11, 14, 15-18 and 20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. When reviewing independent claims 1, 8, and 15 and based upon consideration of all of the relevant factors with respect to the claim as a whole, claims 1-4, 7, 8-11, 14, 15-18 and 20 are held to claim an abstract idea without reciting elements that amount to significantly more than the abstract idea and is/are therefore rejected as ineligible subject matter under 35 U.S.C. 101. The Examiner will analyze Claim 1. The rationale, under MPEP § 2106, for this finding is explained below: The claimed invention (1) must be directed to one of the four statutory categories, and (2) must not be wholly directed to subject matter encompassing a judicially recognized exception, as defined below. The following two step analysis is used to evaluate these criteria. Step 1: Is the claim directed to one of the four patent-eligible subject matter categories: process, machine, manufacture, or composition of matter? When examining the claim under 35 U.S.C. 101, the Examiner interprets that the claim is related to a process since the claim is directed to a method. Step 2a, Prong 1: Does the claim wholly embrace a judicially recognized exception, which includes laws of nature, physical phenomena, and abstract ideas, or is it a particular practical application of a judicial exception? The Examiner interprets that the judicial exception applies since Claim 1 is directed to the abstract idea of using a mental process to determine, categorize, and rank the contextual elements of frames taken from a video. The claims recite mental processes that include determining contextual elements of a video frame based on the elements present in the claim. Under Berkheimer v. HP, The Federal Circuit determined that these claims were directed to mental processes of parsing and comparing images, because the steps were recited at a high level of generality and merely used computers, specifically a processor coupled to a memory, as a tool to perform the processes. Berkheimer v. HP, Inc., 881 F.3d 1360, 125 USPQ2d 1649 (Fed. Cir. 2018), where the data analysis steps are recited at a high level of generality such that they could practically be performed in the human mind, Electric Power Group v. Alstom, S.A., 830 F.3d 1350, 1353-54, 119 USPQ2d 1739, 1741-42 (Fed. Cir. 2016); If the claim recites a judicial exception (i.e., an abstract idea enumerated in MPEP § 2106.04(a), a law of nature, or a natural phenomenon), the claim requires further analysis in Prong Two. Step 2a, Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? The Examiner interprets that the Claim 1 limitation does not provide additional elements or combination of additional elements to a practical application since the claims are not adding insignificant extra-solution activity to the judicial exception. The abstract idea is not integrated into a practical application because the additional elements fail to provide a technical improvement. Data gathering information that describes the contextual elements of the image not add a practical application and could reasonably be done by a person selecting what contextual elements to focus on in the image still. Selecting frames with the contextual attributes and classifying the frames could be done in a human mind. Generating a context structure based on the contextual attributes and classification is further data manipulation and does not integrate the invention into practical application. The instant specification ¶0004-¶0005 describes the practical application being to add additional contextual elements to help improve search results, this applications fail to provide a technical improvement. Specifically, the analysis method does not integrate a judicial exception into practical application. See Genetic Techs. v. Merial LLC, 818 F.3d 1369, 1376, 118 USPQ2d 1541, 1546 (Fed. Cir. 2016) (eligibility "cannot be furnished by the unpatentable law of nature (or natural phenomenon or abstract idea) itself."). For a claim reciting a judicial exception to be eligible, the additional elements (if any) in the claim must "transform the nature of the claim" into a patent-eligible application of the judicial exception, Alice Corp., 573 U.S. at 217, 110 USPQ2d at 1981, either at Prong Two or in Step 2B. If there are no additional elements in the claim, then it cannot be eligible. In such a case, after making the appropriate rejection, it is a best practice for the examiner to recommend an amendment, if possible, that would resolve eligibility of the claim. Step 2b: If a judicial exception into a practical application is not recited in the claim, the Examiner must interpret if the claim recites additional elements that amount to significantly more than the judicial exception. The Examiner interprets that the Claims do not amount to significantly more since the Claims state analyzing images for well-known characteristics with a high level of generality. The claims lack an inventive concept because the elements, considered individually and as an order combination are an abstract idea, and can easily be performed in the human mind based on data analysis, the data in this case being the context contained in frames from video images, where the data analysis steps are recited at a high level of generality such that they could practically be performed in the human mind, Electric Power Group v. Alstom, S.A., 830 F.3d 1350, 1353-54, 119 USPQ2d 1739, 1741-42 (Fed. Cir. 2016). Claims 2-4, 7, 9-11, 14, 16-18 and 20 depend on the independent claim/s include all the limitations of the independent claim. The Examiner finds that Claims 2, 9, and 16 do not state significantly more since the claim adds limitation responding to a query search based on data gathered, which is further data manipulation activity which is an additional element and under Step 2A prong 2 to be a mere recitation of mental process. It is insignificant extra solution activity of additional data gathering 2106.05(g). The Examiner finds that Claims 3, 10, and 17 do not state significantly more since the claim adds limitation of determining the content of the context groups, which is more data gathering activity, which is an additional element and under Step 2A prong 2 to be a mere recitation of mental process. It is insignificant extra solution activity of additional data gathering 2106.05(g). The Examiner finds that Claims 4, 11, and 18 do not state significantly more since the claim adds limitation determining the content the context groups of is more data gathering, which is more data gathering activity, which is an additional element and under Step 2A prong 2 to be a mere recitation of mental process. It is insignificant extra solution activity of additional data gathering 2106.05(g). The Examiner finds that Claim 7, 14, and 20 do not state significantly more since the claim adds the limitation of creating a context hierarchy ranking system based on the classification and context groups which is mere data manipulation which is an additional element and under Step 2A prong 2 to be a mere recitation of mental process. It is insignificant extra solution activity of additional data manipulation 2106.05(g). Thus, Claims 1-4, 7, 8-11, 14, 15-18 and 20 recite the same abstract idea and therefore are not drawn to the eligible subject matter as they are directed to the abstract idea without significantly more. Therefore, the Examiner interprets that the claims are rejected under 35 U.S.C. 101. For the analogous independent claim 8 and analogous independent claim 15 to the independent claim 1, the analogous limitations can be analyzed in the same way as above for the claim 1 hence rejected under 101. Moreover, the claim of 8 and claim 15 further recites at the same statutory category of “device” and “computer program product” respectively which is a limitation that the examiner interprets that the claim is related to a machine since the claim is directed to a device consisting of a processing platform to implement the method of claim 1, and is consistent with the abstract ideas, the claims recite mental processes that include interpreting data and making decisions. Under Berkheimer v. HP, The Federal Circuit determined that these claims were directed to mental processes of parsing and comparing images, because the steps were recited at a high level of generality and merely used computers and computer components as a tool to perform the processes. Berkheimer v. HP, Inc., 881 F.3d 1360, 125 USPQ2d 1649 (Fed. Cir. 2018). This limitation of a residual component just further implements the abstract ideas to be performed by generic computer or software/hardware components of additional elements of the different types of analyzation methods. Therefore, the Examiner interprets that the claims are rejected under 35 U.S.C. 101. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-20 are rejected under 35 U.S.C. 103 as unpatentable over Patel et al (Patel, Bhagwandas, and Brijmohan Singh. "Content-based video retrieval systems: A review." 2023 3rd International Conference on Innovative Mechanisms for Industry Applications (ICIMIA). IEEE, 2023, hereafter referred to as Patel) in view of Oneill et al (WO Patent Publication WO 2024/161331 A1, hereafter referred to as Oneill). Regarding Claim 1, Patel teaches a method (Patel Abstract discloses methods, approaches, and performance evaluation parameters along with several sub process of video analysis used in different multimedia applications like recognition, segmentation, classification) comprising: extracting a set of frames from a video (Patel Fig 1 and Section II A Keyframe extraction, disclose extracting frames from the video); detecting one or more contextual attributes (Patel Fig 1 and Section II B discloses extracting useful features from the keyframes derived in the first step) in the extracted set of frames frames (Patel Fig 1 and Section II Keyframe extraction, disclose extracting frames from the video), wherein the one or more contextual attributes (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step) correspond to contextual attributes associated with a plurality of contextual groups (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features); selecting sample frames from the extracted set of frames (Patel Fig 1 and Section II A (iv) discloses selecting a subset of key frames) for each of the plurality of contextual groups (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features) for which the one or more detected contextual attributes correspond (Patel Fig 1 and Section II A (iv) discloses selecting a subset of key frames that consider various factors such as scene changes, camera motion, and content redundancy to extract representative frames); generating at least one classification for the video (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features, each of these features corresponding to categories) based on at least a portion of the selected sample frames (Patel Fig 1 and Section II A (iv) discloses selecting a subset of key frames that consider various factors such as scene changes, camera motion, and content redundancy to extract representative frames); and generating a context structure (Patel Section II E and F and Section IV discloses using contextual indexing which incorporates contextual information for enhanced retrieval and enhance the retrieval results by considering the contextual and semantic information present in videos) comprising the one or more detected contextual attributes (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step) and the at least one classification for the video (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features, each of these features corresponding to categories); Patel does not explicitly disclose wherein the above steps are performed in accordance with a processing device comprising a processor operatively coupled to a memory and configured to execute program code. Oneill is in the same field of image analysis of video content. Further, Oneill teaches wherein the above steps are performed in accordance with a processing device (Oneill ¶0012 discloses a computing device having at least one processor or processing system configured with processor-executable instructions to perform various operations corresponding to the methods) comprising a processor operatively coupled to a memory and configured to execute program code (Oneill ¶0012 discloses a non-transitory processor-readable storage medium having stored thereon processor executable instructions configured to cause at least one processor or processing system to perform various operations corresponding to the method operations). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Patel by incorporating the use of a face as a contextual category, integrating a referential context hierarchy, and the execution of the method using a processing device as taught by Oneill; to make an invention that can improve the search results through the use of the hierarchal refence to produce results closer in content to the query, including facial queries and results, automatically using a processing device; thus one of ordinary skilled in the art would be motivated to combine the references since there is a need for information structure that maps various features of media content in a multidimensional space to allow for analysis and interpretation of complex relationships and patterns within the data which further allows for faster and more robust analysis of product presence and context in media discussions as disclosed by Oneill in ¶0010. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Regarding Claim 2, Patel in view of Oneill teaches the method of claim 1, further comprising utilizing the context structure (Patel Section II E and F and Section IV discloses using contextual indexing which incorporates contextual information for enhanced retrieval and enhance the retrieval results by considering the contextual and semantic information present in videos) to respond to a contextual query searching for the video (Patel Section II F discloses the contextual queries that could be used to search for the video). See Claim 1 for rationale, its parent claim. Regarding Claim 3, Patel in view of Oneill teaches the method of claim 1, wherein the plurality of contextual groups (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features) comprise a text-oriented contextual group (Patel Section II, B, 4 discloses a textual group including subtitles and captions), a face-oriented contextual group (Oneill ¶0009, ¶0182, ¶0195, ¶0211, ¶0275, ¶0308 discloses identifying an face in a scene as identifying recurring visual elements), an object-oriented contextual group (Patel Section II, B, 1 discloses a visual motion features group including object boundaries and motion), and a color-oriented contextual group (Patel Section II, B, 2 discloses a visual local features group including color distribution). See Claim 1 for rationale, its parent claim. Regarding Claim 4, Patel in view of Oneill teaches the method of claim 1, wherein the one or more detected contextual attributes (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step) comprise one or more of text appearing in the video (Patel Section II, B, 4 discloses a textual group including subtitles and captions) , a face appearing in the video (Oneill ¶0009, ¶0182, ¶0195, ¶0211, ¶0275, ¶0308 discloses identifying an face in a scene as identifying recurring visual elements), an object appearing in the video (Patel Section II, B, 1 discloses a visual motion features group including object boundaries and motion), and a color appearing in the video (Patel Section II, B, 2 discloses a visual local features group including color distribution). See Claim 1 for rationale, its parent claim. Regarding Claim 5, Patel in view of Oneill teaches the method of claim 1, wherein generating the at least one classification for the video (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features, each of these features corresponding to categories) based on at least a portion of the selected sample frames (Patel Fig 1 and Section II A (iv) discloses selecting a subset of key frames that consider various factors such as scene changes, camera motion, and content redundancy to extract representative frames) further comprises utilizing a long short-term memory architecture (Patel Section II B, 3 and 5 disclose the use of Long Short-Term Memory (LSTM) capturing long-term dependencies in sequential data which is suitable for tasks requiring memory and context modeling such as video summarization) to predict the at least one classification (Patel Section II B, 3 and 5 disclose the use of Long Short-Term Memory (LSTM) to determine the audio and deep features of the video). See Claim 1 for rationale, its parent claim. Regarding Claim 6, Patel in view of Oneill teaches the method of claim 5, wherein utilizing a long short-term memory architecture (Patel Section II B, 3 and 5 disclose the use of Long Short-Term Memory (LSTM) capturing long-term dependencies in sequential data which is suitable for tasks requiring memory and context modeling such as video summarization) to predict the at least one classification (Patel Section II B, 3 and 5 disclose the use of Long Short-Term Memory (LSTM) to determine the audio and deep features of the video) further comprises implementing a temporal attention mechanism (Patel Section II B 3 and 5 discloses using CNNS and DNNS to learns discriminative speech features automatically and captures both spectral and temporal variations) in the long short-term memory architecture (Patel Section II B, 3 and 5 disclose the use of Long Short-Term Memory (LSTM) capturing long-term dependencies in sequential data which is suitable for tasks requiring memory and context modeling such as video summarization) to focus on the most relevant parts of the video for making a classification decision (Patel Section II A (ix) discloses using temporal segmentation techniques consider the temporal continuity of frames to capture significant moments in the extracted video). See Claim 1 for rationale, its parent claim. Regarding Claim 7, Patel in view of Oneill teaches the method of claim 1, wherein generating a context structure (Patel Section II E and F and Section IV discloses using contextual indexing which incorporates contextual information for enhanced retrieval and enhance the retrieval results by considering the contextual and semantic information present in videos) comprising the one or more detected contextual attributes (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step) and the at least one classification for the video (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features, each of these features corresponding to categories) further comprises generating a referential context hierarchy (Oneill ¶0164-¶0165, Fig 2B, ¶0403 discloses using a referential contextual hierarchy using identifiers to categorize the contextual components) comprising one or more metadata derived contextual references (Patel Section II A (viii) discloses extracting features from the video including metadata or user preferences), one or more video derived contextual references (Patel Section II B 1 and 2 disclose extracting visual motion and visual local features from the video frames), one or more audio derived contextual references (Patel Section II B, 3 disclose extracting audio features such as music or speech from the video frames), and one or more video classification references (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features, each of these features corresponding to categories). See Claim 1 for rationale, its parent claim. Regarding Claim 8, Patel teaches extract a set of frames from a video (Patel Fig 1 and Section II A Keyframe extraction, disclose extracting frames from the video); detect one or more contextual attributes (Patel Fig 1 and Section II B discloses extracting useful features from the keyframes derived in the first step) in the extracted set of frames frames (Patel Fig 1 and Section II Keyframe extraction, disclose extracting frames from the video), wherein the one or more contextual attributes (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step) correspond to contextual attributes associated with a plurality of contextual groups (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features); select sample frames from the extracted set of frames (Patel Fig 1 and Section II A (iv) discloses selecting a subset of key frames) for each of the plurality of contextual groups (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features) for which the one or more detected contextual attributes correspond (Patel Fig 1 and Section II A (iv) discloses selecting a subset of key frames that consider various factors such as scene changes, camera motion, and content redundancy to extract representative frames); generate at least one classification for the video (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features, each of these features corresponding to categories) based on at least a portion of the selected sample frames (Patel Fig 1 and Section II A (iv) discloses selecting a subset of key frames that consider various factors such as scene changes, camera motion, and content redundancy to extract representative frames); and generate a context structure (Patel Section II E and F and Section IV discloses using contextual indexing which incorporates contextual information for enhanced retrieval and enhance the retrieval results by considering the contextual and semantic information present in videos) comprising the one or more detected contextual attributes (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step) and the at least one classification for the video (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features, each of these features corresponding to categories). Patel does not explicitly disclose an apparatus comprising: at least one processing platform comprising at least one processor coupled to at least one memory, the at least one processing platform, when executing program code, is configured to. Oneill is in the same field of image analysis of video content. Further, Oneill teaches an apparatus (Oneill ¶0173, ¶0179 discloses a computing device) comprising: at least one processing platform comprising at least one processor coupled to at least one memory (Oneill ¶0012 discloses a non-transitory processor-readable storage medium having stored thereon processor executable instructions configured to cause at least one processor or processing system to perform various operations corresponding to the method operations), the at least one processing platform, when executing program code (Oneill ¶0012 discloses a computing device having at least one processor or processing system configured with processor-executable instructions to perform various operations corresponding to the methods), is configured to. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Patel by incorporating the use of a face as a contextual category, integrating a referential context hierarchy, and the execution of the method using a processing device as taught by Oneill; to make an invention that can improve the search results through the use of the hierarchal refence to produce results closer in content to the query, including facial queries and results, automatically using a processing device; thus one of ordinary skilled in the art would be motivated to combine the references since there is a need for information structure that maps various features of media content in a multidimensional space to allow for analysis and interpretation of complex relationships and patterns within the data which further allows for faster and more robust analysis of product presence and context in media discussions as disclosed by Oneill in ¶0010. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Regarding Claim 9, Patel in view of Oneill teaches the apparatus of claim 8, wherein the at least one processing platform (Oneill ¶0012 discloses a non-transitory processor-readable storage medium having stored thereon processor executable instructions configured to cause at least one processor or processing system to perform various operations corresponding to the method operations) is further configured to utilize the context structure (Patel Section II E and F and Section IV discloses using contextual indexing which incorporates contextual information for enhanced retrieval and enhance the retrieval results by considering the contextual and semantic information present in videos) to respond to a contextual query searching for the video (Patel Section II F discloses the contextual queries that could be used to search for the video). See Claim 8 for rationale, its parent claim. Regarding Claim 10, Patel in view of Oneill teaches the apparatus of claim 8, wherein the plurality of contextual groups (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features) comprise a text-oriented contextual group (Patel Section II, B, 4 discloses a textual group including subtitles and captions), a face-oriented contextual group (Oneill ¶0009, ¶0182, ¶0195, ¶0211, ¶0275, ¶0308 discloses identifying an face in a scene as identifying recurring visual elements), an object-oriented contextual group (Patel Section II, B, 1 discloses a visual motion features group including object boundaries and motion), and a color-oriented contextual group (Patel Section II, B, 2 discloses a visual local features group including color distribution). See Claim 8 for rationale, its parent claim. Regarding Claim 11, Patel in view of Oneill teaches the apparatus of claim 8, wherein the one or more detected contextual attributes (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step) comprise one or more of text appearing in the video(Patel Section II, B, 4 discloses a textual group including subtitles and captions) , a face appearing in the video (Oneill ¶0009, ¶0182, ¶0195, ¶0211, ¶0275, ¶0308 discloses identifying an face in a scene as identifying recurring visual elements) , an object appearing in the video (Patel Section II, B, 1 discloses a visual motion features group including object boundaries and motion), and a color appearing in the video (Patel Section II, B, 2 discloses a visual local features group including color distribution). See Claim 8 for rationale, its parent claim. Regarding Claim 12, Patel in view of Oneill teaches the apparatus of claim 8, wherein generating the at least one classification for the video (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features, each of these features corresponding to categories) based on at least a portion of the selected sample frames (Patel Fig 1 and Section II A (iv) discloses selecting a subset of key frames that consider various factors such as scene changes, camera motion, and content redundancy to extract representative frames) further comprises utilizing a long short-term memory architecture (Patel Section II B, 3 and 5 disclose the use of Long Short-Term Memory (LSTM) capturing long-term dependencies in sequential data which is suitable for tasks requiring memory and context modeling such as video summarization) to predict the at least one classification (Patel Section II B, 3 and 5 disclose the use of Long Short-Term Memory (LSTM) to determine the audio and deep features of the video). See Claim 8 for rationale, its parent claim. Regarding Claim 13, Patel in view of Oneill teaches the apparatus of claim 12, wherein utilizing a long short-term memory architecture (Patel Section II B, 3 and 5 disclose the use of Long Short-Term Memory (LSTM) capturing long-term dependencies in sequential data which is suitable for tasks requiring memory and context modeling such as video summarization) to predict the at least one classification (Patel Section II B, 3 and 5 disclose the use of Long Short-Term Memory (LSTM) to determine the audio and deep features of the video) further comprises implementing a temporal attention mechanism (Patel Section II B 3 and 5 discloses using CNNS and DNNS to learns discriminative speech features automatically and captures both spectral and temporal variations) in the long short-term memory architecture (Patel Section II B, 3 and 5 disclose the use of Long Short-Term Memory (LSTM) capturing long-term dependencies in sequential data which is suitable for tasks requiring memory and context modeling such as video summarization) to focus on the most relevant parts of the video for making a classification decision (Patel Section II A (ix) discloses using temporal segmentation techniques consider the temporal continuity of frames to capture significant moments in the extracted video). See Claim 8 for rationale, its parent claim. Regarding Claim 14, Patel in view of Oneill teaches the apparatus of claim 8, wherein generating a context structure (Patel Section II E and F and Section IV discloses using contextual indexing which incorporates contextual information for enhanced retrieval and enhance the retrieval results by considering the contextual and semantic information present in videos) comprising the one or more detected contextual attributes (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step) and the at least one classification for the video (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features, each of these features corresponding to categories) further comprises generating a referential context hierarchy (O’Neill ¶0164-¶0165, Fig 2B, ¶0403 discloses using a referential contextual hierarchy using identifiers to categorize the contextual components) comprising one or more metadata derived contextual references (Patel Section II A (viii) discloses extracting features from the video including metadata or user preferences), one or more video derived contextual references (Patel Section II B 1 and 2 disclose extracting visual motion and visual local features from the video frames), one or more audio derived contextual references (Patel Section II B, 3 disclose extracting audio features such as music or speech from the video frames), and one or more video classification references (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features, each of these features corresponding to categories). See Claim 8 for rationale, its parent claim. Regarding Claim 15, Patel teaches extract a set of frames from a video (Patel Fig 1 and Section II A Keyframe extraction, disclose extracting frames from the video); detect one or more contextual attributes (Patel Fig 1 and Section II B discloses extracting useful features from the keyframes derived in the first step) in the extracted set of frames frames (Patel Fig 1 and Section II Keyframe extraction, disclose extracting frames from the video), wherein the one or more contextual attributes (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step) correspond to contextual attributes associated with a plurality of contextual groups (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features); select sample frames from the extracted set of frames (Patel Fig 1 and Section II A (iv) discloses selecting a subset of key frames) for each of the plurality of contextual groups (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features) for which the one or more detected contextual attributes correspond (Patel Fig 1 and Section II A (iv) discloses selecting a subset of key frames that consider various factors such as scene changes, camera motion, and content redundancy to extract representative frames); generate at least one classification for the video (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features, each of these features corresponding to categories) based on at least a portion of the selected sample frames (Patel Fig 1 and Section II A (iv) discloses selecting a subset of key frames that consider various factors such as scene changes, camera motion, and content redundancy to extract representative frames); and generate a context structure (Patel Section II E and F and Section IV discloses using contextual indexing which incorporates contextual information for enhanced retrieval and enhance the retrieval results by considering the contextual and semantic information present in videos) comprising the one or more detected contextual attributes (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step) and the at least one classification for the video (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features, each of these features corresponding to categories). Patel does not explicitly disclose a computer program product comprising a non-transitory processor-readable storage medium having stored therein program code of one or more software programs, wherein the program code when executed by at least one processing device causes the at least one processing device to. Oneill is in the same field of image analysis of video content. Further, Oneill teaches a computer program product comprising a non-transitory processor-readable storage medium having stored therein program code of one or more software programs (Oneill ¶0012 discloses a non-transitory processor-readable storage medium having stored thereon processor executable instructions configured to cause at least one processor or processing system to perform various operations corresponding to the method operations) , wherein the program code when executed by at least one processing device causes the at least one processing device (Oneill ¶0012 discloses a computing device having at least one processor or processing system configured with processor-executable instructions to perform various operations corresponding to the methods) to. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Patel by incorporating the use of a face as a contextual category, integrating a referential context hierarchy, and the execution of the method using a processing device as taught by Oneill; to make an invention that can improve the search results through the use of the hierarchal refence to produce results closer in content to the query, including facial queries and results, automatically using a processing device; thus one of ordinary skilled in the art would be motivated to combine the references since there is a need for information structure that maps various features of media content in a multidimensional space to allow for analysis and interpretation of complex relationships and patterns within the data which further allows for faster and more robust analysis of product presence and context in media discussions as disclosed by Oneill in ¶0010. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Regarding Claim 16, Patel in view of Oneill teaches the computer program product of claim 15, further comprising utilizing the context structure (Patel Section II E and F and Section IV discloses using contextual indexing which incorporates contextual information for enhanced retrieval and enhance the retrieval results by considering the contextual and semantic information present in videos) to respond to a contextual query searching for the video (Patel Section II F discloses the contextual queries that could be used to search for the video). See Claim 15 for rationale, its parent claim. Regarding Claim 17, Patel in view of Oneill teaches the computer program product of claim 15, wherein the plurality of contextual groups (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features) comprise a text-oriented contextual group (Patel Section II, B, 4 discloses a textual group including subtitles and captions), a face-oriented contextual group (Oneill ¶0009, ¶0182, ¶0195, ¶0211, ¶0275, ¶0308 discloses identifying an face in a scene as identifying recurring visual elements), an object-oriented contextual group (Patel Section II, B, 1 discloses a visual motion features group including object boundaries and motion), and a color-oriented contextual group (Patel Section II, B, 2 discloses a visual local features group including color distribution). See Claim 15 for rationale, its parent claim. Regarding Claim 18, Patel in view of Oneill teaches the computer program product of claim 17, wherein the one or more detected contextual attributes (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step) comprise one or more of text appearing in the video (Patel Section II, B, 4 discloses a textual group including subtitles and captions), a face appearing in the video (Oneill ¶0009, ¶0182, ¶0195, ¶0211, ¶0275, ¶0308 discloses identifying an face in a scene as identifying recurring visual elements) , an object appearing in the video (Patel Section II, B, 1 discloses a visual motion features group including object boundaries and motion), and a color appearing in the video (Patel Section II, B, 2 discloses a visual local features group including color distribution). See Claim 15 for rationale, its parent claim. Regarding Claim 19, Patel in view of Oneill teaches the computer program product of claim 15, wherein generating the at least one classification for the video (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features, each of these features corresponding to categories) based on at least a portion of the selected sample frames (Patel Fig 1 and Section II A (iv) discloses selecting a subset of key frames that consider various factors such as scene changes, camera motion, and content redundancy to extract representative frames) further comprises utilizing a long short-term memory architecture (Patel Section II B, 3 and 5 disclose the use of Long Short-Term Memory (LSTM) capturing long-term dependencies in sequential data which is suitable for tasks requiring memory and context modeling such as video summarization) to predict the at least one classification (Patel Section II B, 3 and 5 disclose the use of Long Short-Term Memory (LSTM) to determine the audio and deep features of the video), and wherein the long short-term memory architecture implements a temporal attention mechanism (Patel Section II B 3 and 5 discloses using CNNS and DNNS to learns discriminative speech features automatically and captures both spectral and temporal variations) in the long short-term memory architecture (Patel Section II B, 3 and 5 disclose the use of Long Short-Term Memory (LSTM) capturing long-term dependencies in sequential data which is suitable for tasks requiring memory and context modeling such as video summarization) to focus on the most relevant parts of the video for making a classification decision (Patel Section II A (ix) discloses using temporal segmentation techniques consider the temporal continuity of frames to capture significant moments in the extracted video). See Claim 15 for rationale, its parent claim. Regarding Claim 20, Patel in view of Oneill teaches the computer program product of claim 15, wherein generating a context structure (Patel Section II E and F and Section IV discloses using contextual indexing which incorporates contextual information for enhanced retrieval and enhance the retrieval results by considering the contextual and semantic information present in videos) comprising the one or more detected contextual attributes (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step) and the at least one classification for the video (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features, each of these features corresponding to categories) further comprises generating a referential context hierarchy (O’Neill ¶0164-¶0165, Fig 2B, ¶0403 discloses using a referential contextual hierarchy using identifiers to categorize the contextual components) comprising one or more metadata derived contextual references (Patel Section II A (viii) discloses extracting features from the video including metadata or user preferences), one or more video derived contextual references (Patel Section II B 1 and 2 disclose extracting visual motion and visual local features from the video frames), one or more audio derived contextual references (Patel Section II B, 3 disclose extracting audio features such as music or speech from the video frames), and one or more video classification references (Patel Fig 1 and Section II B discloses extracting 5 useful features from the keyframes derived in the first step including visual motion features, visual local features, audio features, textual features, and deep features, each of these features corresponding to categories) See Claim 15 for rationale, its parent claim. Reference Cited The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. US Patent Pub US-20240320952-A1 to Petitpont et al. discloses video frame of the scene may be input into expert machine learning models to output expert machine learning model specific labels associated with the scene, and expert machine learning model-specific markup tags. Hu, Weiming, et al. "A survey on visual content-based video indexing and retrieval." IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews) 41.6 (2011): 797-819. (Year: 2011) discloses general strategies in visual content-based video indexing and retrieval, focusing on methods for video structure analysis. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to RACHEL ROBERTS whose telephone number is (571)272-6413. The examiner can normally be reached Monday- Friday 7:30am- 5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Oneal Mistry can be reached on (313) 446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /RACHEL L ROBERTS/Examiner, Art Unit 2674 /ONEAL R MISTRY/Supervisory Patent Examiner, Art Unit 2674
Read full office action

Prosecution Timeline

Sep 24, 2024
Application Filed
Jul 14, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12702269
IMAGE PROCESSING APPARATUS, IMAGE PROCESSING METHOD, NAVIGATION METHOD AND ENDOSCOPE SYSTEM
3y 10m to grant Granted Aug 11, 2026
Patent 12705689
RGB-IR PIXEL PATTERN CONVERSION VIA ADAPTIVE FILTERING
3y 5m to grant Granted Aug 11, 2026
Patent 12678228
PATIENT-SPECIFIC ANTERIOR PLATE IMPLANTS
4y 0m to grant Granted Jul 14, 2026
Patent 12678230
System and Method for Percutaneous Needle Insertion
3y 4m to grant Granted Jul 14, 2026
Patent 12674780
SYSTEM AND METHOD FOR AUTOMATED DETECTION, CLASSIFICATION, AND REMEDIATION OF DEFECTS USING ULTRASOUND TESTING
3y 8m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
76%
Grant Probability
99%
With Interview (+32.3%)
3y 0m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 33 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month