Prosecution Insights
Last updated: August 17, 2026
Application No. 18/669,706

SYSTEM AND METHOD USING REASONING FOR VIDEO ANOMALY DETECTION WITH LARGE LANGUAGE MODELS

Final Rejection §103
Filed
May 21, 2024
Priority
Feb 14, 2024 — provisional 63/553,550
Examiner
ABOUZAHRA, MAHMOUD KAMAL
Art Unit
2486
Tech Center
2400 — Computer Networks
Assignee
Honda Motor Co., Ltd.
OA Round
2 (Final)
65%
Grant Probability
Favorable
3-4
OA Rounds
5m
Est. Remaining
71%
With Interview

Examiner Intelligence

Grants 65% — above average
65%
Career Allowance Rate
26 granted / 40 resolved
+7.0% vs TC avg
Moderate +6% lift
Without
With
+6.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
21 currently pending
Career history
78
Total Applications
across all art units

Statute-Specific Performance

§103
78.4%
+38.4% vs TC avg
§102
11.0%
-29.0% vs TC avg
§112
3.8%
-36.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 40 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The Amendment filed 04/06/2026 has been entered. Claims 1-20 are pending in this application. Claims 1, 16- 17 and 20 have been amended. Claim 18 is cancelled. Response to Arguments Applicant’s arguments with respect to claims 1, 3, 16, and 20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim 1 is rejected under 35 U.S.C. 103 as being unpatentable over Wesley Kenneth Cobb (US 20110043626 A1) (hereinafter Cobb) further in view of Junaid Iqbal (US 20220262121 A1) (hereinafter Iqbal): Regarding Claim 1, Cobb teaches a video anomaly detection (VAD) system (video anomaly detection [0002];[0008]) comprising: an induction stage receiving a plurality of video frames as a reference (receiving video frames for reference [0006]; [0034]) and deriving a rule for a normal event occurrence based on the plurality of video frames received (deriving the and learning normal behavior from the video data [0023]; [0034]) ; and a deduction stage applying the rule for the normal event occurrence and the corresponding rule for the anomaly event occurrence (checking step that applies what is learned to determine if new activity is abnormal [0070]; [0081]) to determine anomalies in non-reference video frames (determines anomalies in new video frames not used as reference [0023]; [0080]). Cobb does not explicitly teach the following limitations; however, in an analogous art, Iqbal teaches deriving a corresponding rule for an anomaly event occurrence by contrasting the corresponding rule for the anomaly event occurrence to the rule for the normal event occurrence (deriving an abnormality rule by comparing against the normal rule [0046]; [0050]) . It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to improve the detection of abnormal events in video data (Cobb [0026]- [0030]). Claims 2- 13, 15, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Wesley Kenneth Cobb (US 20110043626 A1) (hereinafter Cobb) in view of Junaid Iqbal (US 20220262121 A1) (hereinafter Iqbal) further in view of Kamal Nasrollahi (US 20250252735 A1) (hereinafter Nasrollahi): Regarding Claim 2, Iqbal in view of Cobb teach the VAD system of claim 1; however, do not explicitly teach wherein the induction stage comprises large language models (LLMs) to induce the rule for the normal event occurrence from a representative set of normal scenarios from the plurality of video frames as the reference and which differentiates the normal event occurrence and the anomaly event occurrence. However, in an analogous art, Nasrollahi teaches wherein the induction stage comprises large language models (LLMs) to induce the rule for the normal event occurrence from a representative set of normal scenarios from the plurality of video frames as the reference and which differentiates the normal event occurrence and the anomaly event occurrence ([0106] the user-selected VAD model used by the user according to the disclosure may comprise a MLM or a rules engine configured to perform VAD on the acquired video data corresponding to the said unique identifier or the said subset of the unique identifiers, as need be for the use case. [0088] The said MLM preferably comprises a Large Language Model, LLM, as a first LLM, the first LLM being configured to perform captioning of the video data. [0014]; [0024]; [0097]; [0103]; [0107]). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to improve the context awareness of the anomaly detection (Nasrollahi [0104]-[0105]). Regarding Claim 3, Iqbal in view of Cobb teach the VAD system of claim 1; however, do not explicitly teach wherein the induction stage comprises: a visual perception unit converting visual features into textual descriptions from the plurality of video frames as the reference; a rules generation unit generating the rule for the normal event occurrence and the corresponding rule for the anomaly event occurrence from the textual descriptions; and a rules aggregation unit applying randomize smoothing to generate the rule for the normal event occurrence and the corresponding rule for the anomaly event occurrence. However, in an analogous art, Nasrollahi teaches wherein the induction stage comprises: a visual perception unit converting visual features into textual descriptions from the plurality of video frames as the reference (acquire metadata generated by performing captioning of the video data, wherein the metadata comprises semantic data to represent content in the video data in combination with the unique identifiers; [0079]-[0080]; [0041]); a rules generation unit generating the rule for the normal event occurrence and the corresponding rule for the anomaly event occurrence from the textual descriptions (converting the video description into text in step (S310), then in step (S320) a rule is generated based on the description; [0128]; [0103] [0116]; [0014]; [0107]). a rules aggregation unit applying randomize smoothing to generate the rule for the normal event occurrence and the corresponding rule for the anomaly event occurrence ([0122] According to the disclosure, fact-checking may be done before or after identification of the said user-relevant content. However, it is preferable to perform fact-checking before the said identification, in order to avoid presenting the user with false hits.). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to improve the context awareness of the anomaly detection (Nasrollahi [0104]-[0105]). Regarding Claim 4, Iqbal in view of Cobb and Nasrollahi teach the VAD system of claim 3. Nasrollahi further teaches, wherein the visual perception module uses a Vision Language Model (VLM) to convert visual features into textual descriptions ([0014]: Large Vision Language Models (LVLMs) to detect previously unknown objects). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to improve the context awareness of the anomaly detection (Nasrollahi [0104]-[0105]). Regarding Claim 5, Iqbal in view of Cobb and Nasrollahi teach the VAD system of claim 3. Nasrollahi further teaches wherein the visual perception unit decouples the plurality of video frames into multiple categories ([0085] The step S220′ may be triggered only when a predefined threshold is met (e.g. a threshold for detecting motion or specific objects such as humans or vehicles), for computational efficiency.). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to improve the context awareness of the anomaly detection (Nasrollahi [0104]-[0105]). Regarding Claim 6, Iqbal in view of Cobb and Nasrollahi teach the VAD system of claim 3. Nasrollahi further teaches wherein the two categories are human activities and environmental objects ([0085] The step S220′ may be triggered only when a predefined threshold is met (e.g. a threshold for detecting motion or specific objects such as humans or vehicles), for computational efficiency.). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to improve the context awareness of the anomaly detection (Nasrollahi [0104]-[0105]). Regarding Claim 7, Iqbal in view of Cobb and Nasrollahi teach the VAD system of claim 6. Nasrollahi further teaches wherein the rules generation unit generates rules for human activities and environmental objectives (The metadata may also define what type of object has been detected e.g. person, car, dog, bicycle, and/or characteristics of the object (e.g. colour, speed of movement etc). [0066]). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to improve the context awareness of the anomaly detection (Nasrollahi [0104]-[0105]). Regarding Claim 8, Iqbal in view of Cobb and Nasrollahi teach the VAD system of claim 3. Nasrollahi further teaches wherein the rules generation unit queries the textual descriptions and detects patterns to define the rule for the normal event occurrence ([0041] acquire metadata generated by performing captioning of the video data, wherein the metadata comprises semantic data to represent content in the video data in combination with the unique identifiers). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to improve the context awareness of the anomaly detection (Nasrollahi [0104]-[0105]). Regarding Claim 9, Iqbal in view of Cobb and Nasrollahi teach the VAD system of claim 8. Nasrollahi further teaches wherein the rules generation unit derives the corresponding rule for the anomaly event occurrence based on the rule for the normal event occurrence that has been defined ([0116] Generating the message, instruction, event or additional metadata may be triggered by a rules engine, as a first rules engine, based on the said user-defined semantic query. For instance, the rules engine may comprise a rule such as ‘send a text to person X upon detection of a person falling by camera no. 1’.). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to improve the context awareness of the anomaly detection (Nasrollahi [0104]-[0105]). Regarding Claim 10, Iqbal in view of Cobb and Nasrollahi teach the VAD system of claim 3. Nasrollahi further teaches wherein the rules generation unit uses analogical reasoning (Comparison to evaluate analogously. “[0121] Comparing the said first graph with a graph representing ground truth allows to eliminate fictional triples, and thus to reduce the risk of using erroneous data (e.g. for VAD or statistical purposes).”). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to improve the context awareness of the anomaly detection (Nasrollahi [0104]-[0105]). Regarding Claim 11, Iqbal in view of Cobb and Nasrollahi teach the VAD system of claim 3. Nasrollahi further teaches wherein the rules aggregation unit samples a plurality of batches of the video frames as the reference each containing a predefined number of frames, each of the plurality of batches of the video frames as the reference run independently through the visual perception unit and the rules generation unit to generate the rule for the normal event occurrence ([0070]: clips are a predefined number of frames [0012] object-level analysis is performed independently for each object on each frame,). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to improve the context awareness of the anomaly detection (Nasrollahi [0104]-[0105]). Regarding Claim 12, Iqbal in view of Cobb and Nasrollahi teach the VAD system of claim 11. Nasrollahi further teaches wherein the rules aggregation unit uses a large language model (LLM) with a voting mechanism to generate the rule for the normal event occurrence based on appearance in the plurality of batches of the video frames as the reference ([0005] The identification of whether information is important is typically made by the viewer, although the viewer can be assisted by the alert and/or event identifying that the information could be important. Typically, the viewer is interested to view video data that depicts the motion of objects that are of particular interest, such as people or vehicles.). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to improve the context awareness of the anomaly detection (Nasrollahi [0104]-[0105]). Regarding Claim 13, Iqbal in view of Cobb teach the VAD system of claim 1; however, do not explicitly teach wherein the deduction stage a visual perception unit processing the non-reference video frames and outputting the textual descriptions; a perception smoothing unit using exponential majority smoothing for perception error reduction and temporal consistency; and a robust reasoning module with a double-check system to reduce false negative outputs. However, in an analogous art, Nasrollahi teaches wherein the deduction stage, a visual perception unit processing the non-reference video frames and outputting the textual descriptions ([0079] In step S220′, (content) metadata is generated by performing captioning of the video data. The metadata comprises semantic data (according to a semantic data model, SDM) to represent content in the video data in combination with the unique identifiers (i.e. the respective IDs of the video surveillance cameras).); a perception smoothing unit using exponential majority smoothing for perception error reduction and temporal consistency (Such a graph may serve as a historical record or log of detected subjects, predicates, and objects, and/or serve to fact-check the accuracy of the detection model. [0120]); and a robust reasoning module with a double-check system to reduce false negative outputs (it is preferable to perform fact-checking before the said identification, in order to avoid presenting the user with false hits. [0122]). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to improve the context awareness of the anomaly detection (Nasrollahi [0104]-[0105]). Regarding Claim 15, Iqbal in view of Cobb and Nasrollahi teach the VAD system of claim 13. Nasrollahi further teaches wherein the robust reasoning module uses a large language model (LLM) to take a modified description from each of the nonreference video frames and a dummy answer and checks to confirm if the dummy answer matches a description based on the rule for the normal event occurrence and the corresponding rule for the anomaly event occurrence ([0108] Relying on a LLM to directly detect anomalies from image descriptions as in the prior art is unpredictable. However, the present disclosure offers better control, and both model fine-tuning and embedding training have low computational requirements.). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to improve the context awareness of the anomaly detection (Nasrollahi [0104]-[0105]). Regarding Claim 20, Cobb teaches a method for video anomaly detection (VAD) (video anomaly detection [0002];[0008]) comprising: receiving a plurality of video frames as a reference (receiving video frames for reference [0006]; [0034]); generating a rule for a normal event occurrence based on the detected patterns in the plurality' of video frames received (deriving the and learning normal behavior from the video data [0023]; [0034]); Cobb does not explicitly teach the following limitations; however, in an analogous art, Iqbal teaches deriving a corresponding rule for an anomaly event occurrence based on the rule for the normal event occurrence (deriving an abnormality rule by comparing against the normal rule [0046]; [0050]). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to improve the detection of abnormal events in video data (Cobb [0026]- [0030]). Cobb does not explicitly teach the following limitations; however, in an analogous art, Nasrollahi teaches converting visual features into textual descriptions from the plurality of video frames as the reference (acquire metadata generated by performing captioning of the video data, wherein the metadata comprises semantic data to represent content in the video data in combination with the unique identifiers; [0079]-[0080]; [0041]); querying the textual descriptions to detect patterns ([0041] acquire metadata generated by performing captioning of the video data, wherein the metadata comprises semantic data to represent content in the video data in combination with the unique identifiers). applying randomize smoothing to generate the rule for the normal event occurrence and the corresponding rule for the anomaly event occurrence ([0122] According to the disclosure, fact-checking may be done before or after identification of the said user-relevant content. However, it is preferable to perform fact-checking before the said identification, in order to avoid presenting the user with false hits.); processing non-reference video frames to output textual descriptions from the non-reference video frames ([0079] In step S220′, (content) metadata is generated by performing captioning of the video data. The metadata comprises semantic data (according to a semantic data model, SDM) to represent content in the video data in combination with the unique identifiers (i.e. the respective IDs of the video surveillance cameras).); applying the rule for the normal event occurrence and the corresponding rule for the anomaly event occurrence to determine anomalies in the non-reference video frames ([0129] Next, in a step S310, the method comprises identifying user-relevant content using either one of a user-defined semantic query and a user-selected Video Anomaly Detection—VAD—model, in combination with either one of a unique identifier and a subset of the unique identifiers.); applying exponential majority smoothing for perception error reduction and temporal consistency (Such a graph may serve as a historical record or log of detected subjects, predicates, and objects, and/or serve to fact-check the accuracy of the detection model. [0120]); and applying a double-check system to reduce false negative outputs (it is preferable to perform fact-checking before the said identification, in order to avoid presenting the user with false hits. [0122]). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to improve the context awareness of the anomaly detection (Nasrollahi [0104]-[0105]). Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Wesley Kenneth Cobb (US 20110043626 A1) (hereinafter Cobb) in view of Junaid Iqbal (US 20220262121 A1) (hereinafter Iqbal) in view of Kamal Nasrollahi (US 20250252735 A1) (hereinafter Nasrollahi) further in view of Mia Siemon (US 20240303986 A1) (hereinafter Siemon): Regarding Claim 14, Iqbal in view of Cobb and Nasrollahi teach the VAD system of claim 13; however, do not explicitly teach wherein the perception smoothing unit uses a moving average that places a higher weighted value on more recent data points and focuses on a single category. However, in an analogous art, Siemon teaches wherein the perception smoothing unit uses a moving average that places a higher weighted value on more recent data points and focuses on a single category (Evaluations on macro-level take all test videos concatenated to a single recording into consideration, while those on macro-level report the results obtained through the weighted average after considering each test video individually.[0127]). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to further add the teachings of Siemon as disclosed above, in order to increase the accuracy of the anomaly detection (Siemon, [0132]). Claims 16- 17, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Wesley Kenneth Cobb (US 20110043626 A1) (hereinafter Cobb) in view of Junaid Iqbal (US 20220262121 A1) (hereinafter Iqbal) in view of Kamal Nasrollahi (US 20250252735 A1) (hereinafter Nasrollahi) further in view of Narayanan Ramanathan (US 20210304574 A1) (hereinafter Ramanathan) Regarding Claim 16, Cobb teaches a method for video anomaly detection (VAD) (video anomaly detection [0002];[0008]) comprising: receiving a plurality of video frames as a reference (receiving video frames for reference [0006]; [0034]); deriving a rule for a normal event occurrence based on the video frames received (deriving the and learning normal behavior from the video data [0023]; [0034]), applying the rule for the normal event occurrence and the corresponding rule for the anomaly event occurrence (checking step that applies what is learned to determine if new activity is abnormal [0070]; [0080]- [0081]).to determine anomalies in non-reference video frames (determines anomalies in new video frames not used as reference [0023]; [0080]). Cobb does not explicitly teach the following limitations; however, in an analogous art, Iqbal teaches deriving a corresponding rule for an anomaly event occurrence by contrasting the corresponding rule for the anomaly event occurrence to the rule for the normal event occurrence (deriving an abnormality rule by comparing against the normal rule [0046]; [0050]) . It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to improve the detection of abnormal events in video data (Cobb [0026]- [0030]). Cobb does not explicitly teach the following limitations; however, in an analogous art, Nasrollahi teaches wherein the rule for the normal event occurrence is based on randomized smoothing and aggregating responses from a plurality of randomly selected video frames ([0122] According to the disclosure, fact-checking may be done before or after identification of the said user-relevant content. However, it is preferable to perform fact-checking before the said identification, in order to avoid presenting the user with false hits). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to improve the context awareness of the anomaly detection (Nasrollahi [0104]-[0105]). Nasrollahi does not explicitly teach the following limitations; however, in an analogous art, Ramanathan teaches decoupling the plurality of video frames into multiple categories , wherein the multiple categories comprises at least human activates and environmental objects (splitting the images into categories that are either human activities or non-human activities [0099]; [0109]- [0111]; [0124]- [0125]; [0008]; [0124]). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to add the LLM in video anomaly detection as disclosed above by Nasrollahi to further add the detection and categorization of human activities as disclosed above by Ramanathan to improve video detection accuracy and speed (Ramanathan [0012]). Regarding Claim 17, Iqbal in view of Cobb, Nasrollahi, and Ramanathan teach the method of claim 16. Nasrollahi further teaches wherein deriving the rule for the normal event occurrence and the corresponding rule for the anomaly event comprises: converting visual features in each of the plurality of video frames as the reference into textual descriptions ([0041] acquire metadata generated by performing captioning of the video data, wherein the metadata comprises semantic data to represent content in the video data in combination with the unique identifiers); and generating the rule for the normal event occurrence and the corresponding rule for the anomaly event occurrence from the textual descriptions ([0103] For instance, if the configuration includes ‘detecting a person falling’ in the video data from camera no. 1, a person lying in bed captured by the video camera no. 1 will not trigger a “fall” event (or alarm or alert) in the VMS, whereas a person lying on the floor captured by the video camera no. 1 will.). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to improve the context awareness of the anomaly detection (Nasrollahi [0104]-[0105]). Regarding Claim 19, Iqbal in view of Cobb, Nasrollahi, and Ramanathan teach the method of claim 17. Nasrollahi further teaches processing the non-reference video frames to output textual descriptions of the non-reference video frames ([0079] In step S220′, (content) metadata is generated by performing captioning of the video data. The metadata comprises semantic data (according to a semantic data model, SDM) to represent content in the video data in combination with the unique identifiers (i.e. the respective IDs of the video surveillance cameras).); applying exponential majority smoothing for perception error reduction and temporal consistency (Such a graph may serve as a historical record or log of detected subjects, predicates, and objects, and/or serve to fact-check the accuracy of the detection model. [0120]); and applying a double-check system to reduce false negative outputs (it is preferable to perform fact-checking before the said identification, in order to avoid presenting the user with false hits. [0122]). It would have been obvious to the person having ordinary skill in the art before the effective filling date of the claimed invention to modify the deriving of normal activity in video data as disclosed by Cobb to add the deriving of abnormality rules by comparing such events with the normal activities as disclosed by Cobb above to further add the LLM in video anomaly detection as disclosed above by Nasrollahi to improve the context awareness of the anomaly detection (Nasrollahi [0104]-[0105]). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MAHMOUD KAMAL ABOUZAHRA whose telephone number is (703)756-1694. The examiner can normally be reached M-F 7:00 AM to 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jamie Atala can be reached at (571) 272-7384. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MAHMOUD KAMAL ABOUZAHRA/ Examiner, Art Unit 2486 /JAMIE J ATALA/ Supervisory Patent Examiner, Art Unit 2486
Read full office action

Prosecution Timeline

May 21, 2024
Application Filed
Jan 07, 2026
Non-Final Rejection mailed — §103
Apr 06, 2026
Response Filed
Jul 21, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12676955
ELECTRONIC MIRROR SYSTEM, IMAGING DEVICE, AND ELECTRONIC MIRROR
3y 2m to grant Granted Jul 07, 2026
Patent 12634490
DECODING A VIDEO STREAM ON A CLIENT DEVICE
2y 12m to grant Granted May 19, 2026
Patent 12558845
System and Method for a Three-Dimensional Optical Switch Display Device
5y 3m to grant Granted Feb 24, 2026
Patent 12464148
COMPUTER-IMPLEMENTED MULTI-SCALE MACHINE LEARNING MODEL FOR THE ENHANCEMENT OF COMPRESSED VIDEO
2y 7m to grant Granted Nov 04, 2025
Patent 12422691
VEHICULAR CAMERA ASSEMBLY WITH LENS BARREL WELDED AT IMAGER HOUSING
2y 4m to grant Granted Sep 23, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
65%
Grant Probability
71%
With Interview (+6.3%)
2y 8m (~5m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 40 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month