DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgement is made of Applicant’s claim of priority from Foreign Application No. CN202310900041.5, filed July 20, 2023.
Status of Claims
Claims 1-20 are pending.
Response to Arguments
Applicant’s arguments, see p. 12-13, filed July 24, 2026, with respect to the Claim Objections have been fully considered and are persuasive. The amendments have overcome the previous objection and therefore it has been withdrawn.
Applicant’s arguments, see p. 13, filed July 24, 2026, with respect to the 35 USC 101 abstract idea rejections have been fully considered and are persuasive. The amendments have overcome the previous rejections and therefore they have been withdrawn.
With regards to the 35 USC § 103 rejection of claim 1, applicant argues the claims as amended are not taught by the immediate prior art and the claim should not be rejected under 35 USC § 103. Applicant’s remarks and amendments have been fully considered and are found convincing, however, upon further search and consideration, the newly discovered Zhao prior art document, referenced in the updated rejection below, teaches the limitations as claimed. Please see below for full rejection. Accordingly, applicant’s amendments have necessitated the new grounds of rejection set forth.
Applicant further argues that the references do not teach the limitations that make up dependent claim 3. Examiner respectfully disagrees. As described in the 35 USC 103 rejection below, the Ghanem reference teaches the limitations of claim 3. Applicant argues that Ghanem discloses identifying objects of interest based on a depth or distance of pixels and does not teach the “semantically related patches” as claimed. However, Ghanem teaches associating each object of interest in a current frame with an object of interest in a previous frame. Under the broadest reasonable interpretation of the “semantically related patches”, Examiner asserts that the associated objects of interest in each frame teach these patches and the object of interest in the current frame is extracted by a spatial analysis module as the first spatial semantic feature and the temporal analysis module utilizes the association between the objects of interest in the frames to generate temporal features of the objects of interest (i.e., first temporal semantic features) (see Ghanem, Para. [0050]). Thus, the 35 USC 103 rejection of claim 3 is maintained.
Applicant further argues that the references do not teach the limitations that make up dependent claim 4. Examiner respectfully disagrees. Applicant argues that the Wen reference does not disclose the specific sequence recited in claim 4 of convolution, rearrangement, convolution, rearrangement. However, as described in the 35 USC 103 rejections below the Wen reference is only relied upon to teach the second convolution and rearrangement processes. The Narayanmurthy reference teaches a first convolution process (see Narayanmurthy, Col. 7, lines 13-58) and the Wang reference teaches performing spatial rearranging after a convolution process (see Wang, Para. [0099]). Then, Wen is relied upon to teach performing convolution on the rearranged features of Wang and then performing a channel rearrangement (see Wen, Para. [0035]). Nothing in the Wen reference indicates that the second convolution layer could not perform convolution on features previously rearranged. Thus, the 35 USC 103 rejection of claim 4 is maintained, and consequently, THIS ACTION IS FINAL.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-3, 6, 8, and 16-20 are rejected under 35 U.S.C. 103 as being unpatentable over Narayanmurthy et al. (US 11,328,512 B2) in view of Roto Le (US 2023/0101399 A1) further in view of Zhao et al. (US 2021/0303864 A1), Ghanem et al. (US 2016/0110613 A1) and Qi et al. (US 2023/0154207 A1).
Regarding claim 1, Narayanmurthy teaches a method executed by an electronic device, comprising:
acquiring behavior objects and relevant events associated with the behavior objects in a video to be processed by using an artificial intelligence (AI) network (Narayanmurthy, Col. 8, line 66 – Col. 9, line 23, the behaviour detection module may implement RNN (i.e., AI network) using techniques to detect the behaviour of the subject (i.e., the behaviour object). The behaviour detection may include actions and activities (i.e., relevant events associated with the behaviour objects));
wherein the acquiring the behavior objects and the relevant events associated with the behavior objects comprises extracting semantic features of the video, the semantic features of the video comprising semantic features of each frame (Narayanmurthy, Col. 8, lines 33-65, the first spatial filter is applied on each image frame of the plurality of image frames (i.e., semantic features of each frame). The output of the first spatial filter is a first set of activation maps. The first set of activation maps has reduced dimensions than the plurality of image frames. The one or more spatial features (i.e., semantic features) are extracted from the first set of activation maps. The extracted one or more spatial features may be used to determine low-level spatial features. Further, a first level temporal filter is applied on the first set of activation maps. An output of the first level temporal filter is a first level second set of activation maps which provides low-level temporal features. A second level temporal filer is applied on the output of the first level temporal filter, resulting in a second level second set of activation maps which is used to determine high-level temporal features).
Although Narayanmurthy teaches a user interface (Narayanmurthy, Col. 11, lines 8-15), Narayanmurthy does not explicitly teach “providing a behavior object selection interface based on the acquired behavior objects”, “receiving a behavior object selected through the selection interface by a user” and “providing an event related to the behavior object selected by the user”. However, in an analogous field of endeavor, Le teaches a user interface module identifies objects “A”, “B”, “C”, and “D” in the video data. A user selection of objects “A” and “C” causes the user interface module to generate a user interface with user interface elements displaying the portions of the video data corresponding to objects “A” and “C” but not objects “B” or “D” (Le, Para. [0024]).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the method of Narayanmurthy with the teachings of Le by providing a user a selection interface to select behavior objects and to provide the video data (i.e., event) corresponding to the selected objects. One having ordinary skill in the art would have been motivated to combine these references because doing so would allow for presenting a user with relevant video information for selected objects, as recognized by Le. Thus, the claimed invention would have been obvious to one having ordinary skill in the art before the effective filing date.
Although Narayanmurthy in view of Le teaches displaying portions of the video corresponding to selected objects (Le, Para. [0024]), the references do not explicitly teach “wherein the event related to the behavior object comprises an event clip corresponding to an object with a preset behavior”. However, in an analogous field of endeavor, Zhao teaches based on the target video and the behavior recognition model, whether there is the preset behavior included in the target video is determined. If in the target video there is the preset behavior, it is determined that the output data of the target model meets the condition corresponding to the target model, and a video clip containing the preset behavior is determined as the dynamic cover (i.e., an event clip corresponding to an object with a preset behavior) (Zhao, Para. [0090]). When the user browses the videos in the short video entertainment applications in the terminal device(s) 101, 102, and/or 103, if a video card is loaded, the dynamic cover corresponding to the video is displayed, so that the user may learn video information of the video based on the dynamic cover, which improves an efficiency of information acquisition (Zhao, Para. [0024]).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date to modify the method of Narayanmurthy in view of Le with the teachings of Zhao by including that the event related to the behavior object is an event clip corresponding to an object with a preset behavior. One having ordinary skill in the art would have been motivated to combine these references because doing so would allow for improving efficiency of information acquisition, as recognized by Zhao.
Although Narayanmurthy in view of Le further in view of Zhao teaches determining spatial and temporal features (Narayanmurthy, Col. 8, lines 33-65), the references do not explicitly teach “wherein the extracting semantic features of the video comprises: extracting a first semantic feature of the frame and a second semantic feature of an adjacent frame”, “based on the extracted first semantic feature of the frame, obtaining a first spatial semantic feature”, “based on the extracted second semantic feature of the adjacent frame, obtaining a first temporal semantic feature”. However, in an analogous field of endeavor, Ghanem teaches the spatial analysis module iteratively identifies objects of interest based on local maximum or minimum depth stream values within each frame (i.e., first and second semantic features), removes identified objects of interest, and repeats until all objects of interest have been identified. The temporal analysis module is connected to receive objects of interest identified by the spatial analysis module in a frame, wherein the temporal analysis module associates each object of interest in the current frame with an object of interest identified in a previous frame, wherein the temporal analysis module utilizes the association between current frame objects of interest and previous frame objects of interest to generate temporal features related to each object of interest (i.e., first temporal semantic features of objects in the frame) (Ghanem, Para. [0050]).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date to modify the method of Narayanmurthy in view of Le further in view of Zhao with the teachings of Ghanem by including determining features in the frame and an adjacent frame, determining semantically related patches in the frame (i.e., objects of interest), extracting spatial features of objects and extracting temporal features of objects based on the current and adjacent frame. One having ordinary skill in the art would have been motivated to combine these references because doing so would allow for a video analytic system capable of operating accurately in a variety of applications and conditions, as recognized by Ghanem.
Although Narayanmurthy in view of Le further in view of Zhao and Ghanem teaches determining spatial and temporal features between two frames (Ghanem, Para. [0050]), the references do not explicitly teach “fusing the first spatial semantic feature and the first temporal semantic feature”. However, in an analogous field of endeavor, Qi teaches the spatial and temporal feature information is fused to obtain key features (Qi, Para. [0022]).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date to modify the method of Narayanmurthy in view of Le further in view of Zhao and Ghanem with the teachings of Qi by including fusing the spatial features and temporal features to obtain semantic features of the frame (i.e., key features). One having ordinary skill in the art would have been motivated to combine these references because doing so would allow for detecting behavior objects and actions based on spatial and temporal features, as recognized by Qi. Thus, the claimed invention would have been obvious to one having ordinary skill in the art before the effective filing date.
Regarding claim 2, Narayanmurthy in view of Le further in view of Zhao, Ghanem teaches the method according to claim 1, wherein the acquiring the behavior objects further comprises:
determining, based on the semantic features of the video, the behavior objects and the relevant events in the video (Narayanmurthy, Col. 8, line 66 – Col. 9, line 23, the extracted one or more spatial and temporal features are provided to the FC layer which generates the feature vector. The feature vector is then provided to the behaviour detection module for determining a behaviour of the subject).
Regarding claim 3, Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi teaches the method according to claim 1, wherein the extracting the semantic features further comprises:
for each frame, extracting the first semantic feature of the frame and the second semantic feature of the adjacent frame, based on a convolution module (Ghanem, Para. [0050], the spatial analysis module iteratively identifies objects of interest based on local maximum or minimum depth stream values within each frame (i.e., extracts first semantic feature of frame and second semantic feature of adjacent frame));
determining, based on the first semantic feature and the second semantic feature, first semantically related patches in the frame and second semantically related patches in the adjacent frame (Ghanem, Para. [0050], the temporal analysis module is connected to receive objects of interest identified by the spatial analysis module in a frame, wherein the temporal analysis module associates each object of interest in the current frame with an object of interest identified in a previous frame (i.e., first semantically related patches in the frame and second semantically related patches in the adjacent frame));
extracting, from the first semantically related patches in the frame, the first spatial semantic feature of at least one object in the frame (Ghanem, Para. [0050], the temporal analysis module is connected to receive objects of interest identified by the spatial analysis module in a frame (i.e., extracting the first spatial semantic feature of at least one object in the frame)); and
extracting, from the second semantically related patches in the adjacent frame, the first temporal semantic feature of at least one object in the frame (Ghanem, Para. [0050], the temporal analysis module utilizes the association between current frame objects of interest and previous frame objects of interest to generate temporal features related to each object of interest (i.e., first temporal semantic feature of at least one object in the frame)).
The proposed combination as well as the motivation for combining the Narayanmurthy, Le, Zhao, Ghanem and Qi references presented in the rejection of Claim 1, apply to Claim 3 and are incorporated herein by reference. Thus, the method recited in Claim 3 is met by Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi.
Regarding claim 6, Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi teaches the method according to claim 2, wherein the determining the behavior objects and the relevant events in the video comprises:
determining, based on the semantic features of each frame and using an object mask module, a region where an object in each frame is located (Le, Para. [0020], the object detection model provides, as output, pixel coordinates or bounding boxes identifying the locations of particular objects within the video data);
determining, based on the semantic features of the frame and the region where the object in the frame is located, region features of the region where the object in the frame is located (Le, Para. [0021], classifying an object includes assigning a label or classification to the particular identified object. As an example, assuming video data depicting a presentation where a presenter is speaking to an audience and writing notes on a white board, the object detection model identifies and classifies objects such as “presenter,” “white board,” “audience member,” and the like); and
determining, based on the region features of the region where the object in each frame is located, the behavior objects and the relevant events in the video (Narayanmurthy, Col. 8, line 66 – Col. 9, line 23, the behaviour detection module may implement RNN (i.e., AI network) using techniques to detect the behaviour of the subject (i.e., the behaviour object). The behaviour detection may include actions and activities (i.e., relevant events associated with the behaviour objects)).
The proposed combination as well as the motivation for combining the Narayanmurthy, Le, Zhao, Ghanem and Qi references presented in the rejection of Claim 1, apply to Claim 6 and are incorporated herein by reference. Thus, the method recited in Claim 6 is met by Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi.
Regarding claim 8, Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi teaches the method according to claim 6, wherein the determining the behavior objects comprises:
for the frame, obtaining object features corresponding to the frame based on the region features of the region where the object in the frame is located (Le, Para. [0020], the object detection model provides, as output, pixel coordinates or bounding boxes identifying the locations of particular objects within the video data);
determining, based on the object features and using an object recognition model, whether the behavior objects are contained in the frame (Le, Para. [0021], classifying an object includes assigning a label or classification to the particular identified object. As an example, assuming video data depicting a presentation where a presenter is speaking to an audience and writing notes on a white board, the object detection model identifies and classifies objects such as “presenter,” “white board,” “audience member,” and the like); and
obtaining the behavior objects and the relevant events in the video based on behavior object features, wherein the behavior object features are object features of frames containing the behavior objects (Narayanmurthy, Col. 8, line 66 – Col. 9, line 23, the behaviour detection module may implement RNN (i.e., AI network) using techniques to detect the behaviour of the subject (i.e., the behaviour object). The behaviour detection may include actions and activities (i.e., relevant events associated with the behaviour objects)).
The proposed combination as well as the motivation for combining the Narayanmurthy, Le, Zhao, Ghanem and Qi references presented in the rejection of Claim 1, apply to Claim 8 and are incorporated herein by reference. Thus, the method recited in Claim 8 is met by Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi.
Claims 16-19 recite devices with elements corresponding to the steps recited in Claims 1-3 and 6, respectively. Therefore, the recited elements of these claims are mapped to the proposed combination in the same manner as the corresponding steps in their corresponding method claims. Additionally, the rationale and motivation to combine the Narayanmurthy, Le, Zhao, Ghanem and Qi references, presented in rejection of Claim 1, apply to these claims. Finally, the combination of the Narayanmurthy, Le, Zhao, Ghanem and Qi references discloses a processor (Narayanmurthy, Col. 2, lines 1-22, the computing unit comprising a processor and a memory).
Claim 20 recites a computer-readable storage medium storing a program with instructions corresponding to the steps recited in Claim 1. Therefore, the recited programming instructions of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim. Additionally, the rationale and motivation to combine the Narayanmurthy, Le, Zhao, Ghanem and Qi references, presented in rejection of Claim 1, apply to this claim. Finally, the combination of the Narayanmurthy, Le, Zhao, Ghanem and Qi references discloses a computer readable storage medium (Narayanmurthy, Col. 2, lines 23-42, a non-transitory computer readable medium including instructions stored thereon).
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Narayanmurthy et al. (US 11,328,512 B2) in view of Roto Le (US 2023/0101399 A1) further in view of Zhao et al. (US 2021/0303864 A1), Ghanem et al. (US 2016/0110613 A1) and Qi et al. (US 2023/0154207 A1), as applied to claims 1-3, 6, 8 and 16-20 above, and further in view of Wang et al. (US 2022/0381699 A1) and Wen et al. (CN 115457271 A, machine translation used herein).
Regarding claim 4, Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi teaches the method according to claim 1, wherein, for each frame, the extracting comprises:
performing convolution on the frame by using a first convolution layer (Narayanmurthy, Col. 7, lines 13-58, the plurality of image frames may be represented as an image matrix comprising a plurality of elements. Each element of the matrix represents a pixel value. A convolution kernel is a matrix comprising elements that have to be convoluted with the image matrix/image matrices. The convolution is performed to represent the features of the plurality of image frames).
Although Narayanmurthy in view of Le further in view of Ghanem and Qi teaches performing convolution on the frame (Narayanmurthy, Col. 7, lines 13-58), the references do not explicitly teach “spatially rearranging features of each channel from among features extracted by the first convolution layer”. However, in an analogous field of endeavor, Wang teaches the pixel rearrangement layer may periodically rearrange the low-resolution features of the channels of each pixel obtained by the second convolution to obtain a high-resolution feature image (Wang, Para. [0099]).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the method of Narayanmurthy in view of Le further in view of Ghanem and Qi with the teachings of Wang by including performing spatial rearrangement on the features of the channels of each pixel extracted by the convolution layer. One having ordinary skill in the art would have been motivated to combine these references because doing so would allow for obtaining a high-resolution feature image, as recognized by Wang.
Although Narayanmurthy in view of Le further in view of Ghanem, Qi and Wang teaches rearranging low-resolution features of the channels (Wang, Para. [0099]), the references do not explicitly teach “performing convolution on the rearranged features by a second convolution layer” and “performing channel rearrangement on features of each space in features extracted by the second convolution layer to obtain semantic features of the frame”. However, in an analogous field of endeavor, Wen teaches performing 3x3 convolution, 1x1 convolution, 3x3 dilated convolution and channel rearrangement operations in sequence (Wen, Para. [0035]).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the method of Narayanmurthy in view of Le further in view of Ghanem, Qi and Wang with the teachings of Wen by including performing convolution on the rearranged features of Wang and further performing channel rearrangement on the features to determine features of the frame. One having ordinary skill in the art would have been motivated to combine these references because doing so would allow for improved segmentation performance of a model, as recognized by Wen. Thus, the claimed invention would have been obvious to one having ordinary skill in the art before the effective filing date.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Narayanmurthy et al. (US 11,328,512 B2) in view of Roto Le (US 2023/0101399 A1) further in view of Zhao et al. (US 2021/0303864 A1), Ghanem et al. (US 2016/0110613 A1) and Qi et al. (US 2023/0154207 A1), as applied to claims 1-3, 6, 8 and 16-20 above, and further in view of Wu et al. (US 2021/0383128 A1) and Reda et al. (US 2019/0297326 A1).
Regarding claim 5, Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi teaches the method according to claim 3, as described above.
Although Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi teaches fusing spatial and temporal features (Qi, Para. [0022]), the references do not explicitly teach “fusing the first semantic feature and the second semantic feature to obtain a first fused feature”. However, in an analogous field of endeavor, Wu teaches each of the initial semantic feature blocks is fused, and a fused target semantic feature block is obtained (Wu, Para. [0038]).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the method of Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi with the teachings of Wu by including fusing the first and second semantic features to obtain a first fused feature. One having ordinary skill in the art before the effective filing date of the claimed invention would have been motivated to combine these references because doing so would allow for efficient video recognition, as recognized by Wu.
Although Narayanmurthy in view of Le further in view of Zhao, Ghanem, Qi and Wu teaches determining a first fused feature by fusing first and second semantic features (Wu, Para. [0038]), the references do not explicitly teach “determining, based on the first fused feature and in the frame and the adjacent frame, spatial position offshoot information of other patches semantically related to each patch in the frame relative to the patch, respectively” and “determining, based on the spatial position offshoot information, the first semantically related patches in the frame and the second semantically related patches in the adjacent frame”. However, in an analogous field of endeavor, Reda teaches The sampling of the previous video frame, based on the parameters, is performed to generate pixel values for a predicted video frame, which is the next video frame following the sequence of video frames. The parameters include a displacement vector and one or more convolution kernels for each pixel of the predicted video frame. A spatially-displaced convolution module receives the set of parameters for a pixel of the predicted video frame and samples the previous video frame by performing a convolution operation on a corresponding patch of pixels displaced from a corresponding pixel in the previous video frame. The patch of pixels is identified using the displacement vector from the set of parameters for the pixel and is offset from the corresponding pixel in the previous video frame by a magnitude of the displacement vector (Reda, Para. [0031]).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the method of Narayanmurthy in view of Le further in view of Zhao, Ghanem, Qi and Wu with the teachings of Reda by including determining the semantically related patches in the frame and adjacent frame based off of determining a patch of pixels offset from the corresponding pixel in the previous frame by a magnitude of the displacement vector (i.e., determine semantically related patches based on the spatial position offshoot information). One having ordinary skill in the art would have been motivated to combine these references because doing so would allow for predicting motion of objects and occlusions or disocclusions, as recognized by Reda. Thus, the claimed invention would have been obvious to one having ordinary skill in the art before the effective filing date.
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Narayanmurthy et al. (US 11,328,512 B2) in view of Roto Le (US 2023/0101399 A1) further in view of Zhao et al. (US 2021/0303864 A1), Ghanem et al. (US 2016/0110613 A1) and Qi et al. (US 2023/0154207 A1), as applied to claims 1-3, 6, 8 and 16-20 above, and further in view of Wu et al. (US 2021/0383128 A1).
Regarding claim 7, Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi teaches the method according to claim 6, wherein the determining the region where the object in each frame is located comprises:
performing an object segmentation (Le, Para. [0020], the object detection model provides, as output, pixel coordinates or bounding boxes identifying the locations of particular objects within the video data. As another example, in some implementations, the object detection model provides, as output, an object mask such as an image or pixel mask that maps to particular objects within the video data).
Although Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi teaches performing object segmentation (Le, Para. [0020]), the references do not explicitly teach “for each frame, fusing a first semantic features of the frame and a second semantic features of the adjacent frame to obtain a first fused features”. However, in an analogous field of endeavor, Wu teaches each of the initial semantic feature blocks is fused, and a fused target semantic feature block is obtained (Wu, Para. [0038]).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the method of Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi with the teachings of Wu by including fusing the first and second semantic features to obtain a first fused feature on which to perform the object segmentation taught by Le. One having ordinary skill in the art before the effective filing date of the claimed invention would have been motivated to combine these references because doing so would allow for efficient video recognition, as recognized by Wu. Thus, the claimed invention would have been obvious to one having ordinary skill in the art before the effective filing date.
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Narayanmurthy et al. (US 11,328,512 B2) in view of Roto Le (US 2023/0101399 A1) further in view of Zhao et al. (US 2021/0303864 A1), Ghanem et al. (US 2016/0110613 A1) and Qi et al. (US 2023/0154207 A1), as applied to claims 1-3, 6, 8 and 16-20 above, and further in view of Jiang et al. (US 2024/0046471 A1, with priority to Foreign Application No. CN 202210191770.3, filed February 28, 2022, US PGPub used herein for translation and mapping purposes).
Regarding claim 9, Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi teaches the method according to claim 8, as described above.
Although Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi teaches determining an object from region features (Le, Para. [0020]), the references do not explicitly teach “fusing the region features of the region where the object in the frame is located and the semantic features of the frame to obtain target features of the frame” and “fusing the target features of the frame and the region features of the region where the object in the frame is located to obtain the object features of the frame”. However, in an analogous field of endeavor, Jiang teaches the feature fusion module includes a first fusion unit, configured to perform fusion processing on the image semantic feature and a view feature to obtain a view image semantic feature (i.e., fusing the region features and the semantic features to obtain target features); and a second fusion unit configured to perform feature fusion processing on each view image semantic feature (i.e., fusing the target features of the frame and the region features) (Jiang, Para. [0174]-[0176]).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the method of Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi with the teachings of Jiang by including fusing region features with semantic features to obtain target features and fusing region features with target feature to obtain object features. One having ordinary skill in the art would have been motivated to combine these references because doing so would allow for performing image recognition with reduced computation complexity, as recognized by Jiang. Thus, the claimed invention would have been obvious to one having ordinary skill in the art before the effective filing date.
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Narayanmurthy et al. (US 11,328,512 B2) in view of Roto Le (US 2023/0101399 A1) further in view of Zhao et al. (US 2021/0303864 A1), Ghanem et al. (US 2016/0110613 A1), Qi et al. (US 2023/0154207 A1) and Jiang et al. (US 2024/0046471 A1, with priority to Foreign Application No. CN 202210191770.3, filed February 28, 2022, US PGPub used herein for translation and mapping purposes), as applied to claim 9 above, and further in view of Wu et al. (US 2021/0383128 A1).
Regarding claim 10, Narayanmurthy in view of Le further in view of Zhao, Ghanem, Qi and Jiang teaches the method according to claim 9, wherein the fusing the region features comprises:
extracting, from the region where the object in the frame is located, target region features of the object in the frame (Jiang, Para. [0165]-[0166], the first extraction unit is configured to perform feature extraction processing on the first image window interactive feature to obtain the 2D image window feature (i.e., extracting target region features)); and
fusing the target region features and the semantic features of the frame to obtain the target features of the frame (Jiang, Para. [0175], a first fusion unit, configured to perform fusion processing on the image semantic feature and a view feature to obtain a view image semantic feature).
The proposed combination as well as the motivation for combining the Narayanmurthy, Le, Zhao, Ghanem, Qi and Jiang references presented in the rejection of Claim 9, apply to Claim 10 and are incorporated herein by reference.
Although Narayanmurthy in view of Le further in view of Zhao, Ghanem, Qi and Jiang teaches fusing region features and semantic features in a frame (Jiang, Para. [0175]), the references do not explicitly teach “fusing the region features of the region where the object in the frame is located and the semantic features of the adjacent frame”. However, in an analogous field of endeavor, Wu teaches the feature information of different channels of the each of initial semantic feature blocks is fused, so that the purpose of different initial semantic feature blocks containing part of information of other initial semantic feature blocks adjacent to the initial semantic feature blocks in time sequence is achieved, and thus semantic associations and differences between different video segments can be determined according to each fused target semantic feature block (Wu, Para. [0039]).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the method of Narayanmurthy in view of Le further in view of Zhao, Ghanem, Qi and Jiang with the teachings of Wu by including fusing the region features of the object in the frame with the semantic features of the adjacent frame. One having ordinary skill in the art would have been motivated to combine these references because doing so would allow for efficient video recognition, as recognized by Wu. Thus, the claimed invention would have been obvious to one having ordinary skill in the art before the effective filing date.
Claims 11-12 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Narayanmurthy et al. (US 11,328,512 B2) in view of Roto Le (US 2023/0101399 A1) further in view of Zhao et al. (US 2021/0303864 A1), Ghanem et al. (US 2016/0110613 A1) and Qi et al. (US 2023/0154207 A1), as applied to claims 1-3, 6, 8 and 16-20 above, and further in view of Mwaura et al. (US 2024/0386715 A1, filed May 19, 2023).
Regarding claim 11, Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi teaches the method according to claim 8, as described above.
Although Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi teaches obtaining behavior objects and related events (Narayanmurthy, Col. 8, line 66 – Col. 9), the references do not explicitly teach “aggregating the behavior objects features to obtain at least one aggregation result” and “obtaining, based on each frame corresponding to the at least one aggregation result, the behavior objects and the relevant events corresponding to the at least one aggregation result”. However, in an analogous field of endeavor, Mwaura teaches the processor can aggregate trajectories retrieved in response to a user request and that overlap with at least one specified feature (e.g., a specific object of interest, a specific region of interest, a specific timestamp, a number of objects, etc.) based on trajectory data (e.g., bounding box data, a time window, etc.) to identify at least one event. For example, if the user request specifies a person, the processor can aggregate trajectories captured and stored in the storage device of that person to form the event (Mwaura, Para. [0056]).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date to modify the method of Narayanmurthy in view of Le further in view of Zhao, Ghanem and Qi with the teachings of Mwaura by including aggregating the behavior object features (i.e., trajectories) to obtain an aggregation result and obtain based on the frame and the aggregation result the behavior objects and relevant events. One having ordinary skill in the art before the effective filing date would have been motivated to combine these references because doing so would allow for accurately tracking objects, as recognized by Mwaura. Thus, the claimed invention would have been obvious to one having ordinary skill in the art before the effective filing date.
Regarding claim 12, Narayanmurthy in view of Le further in view of Zhao, Ghanem, Qi and Mwaura teaches the method according to claim 11, wherein the aggregating the behavior object features comprises:
for each behavior object feature, determining at least one similar object feature of the behavior object feature from the object features (Mwaura, Para. [0056], in some cases, some events can overlap with each other or are similar to each other. As such, overlapping events can be combined into a single event to reduce computational overhead, save storage space, and/or reduce redundant and/or duplicate events. For instance, two people moving in a common direction, at a similar velocity, and/or within a vicinity of one another may originally be determined to be two events);
extracting second fused features of the behavior objects based on the behavior object features and the at least one similar object feature (Mwaura, Para. [0056], As both people appear and/or leave a scene at the same time and/or share one or more similar features (e.g., motion vector, speed, location, etc.), the processor 214 can combine the two events into one event); and
aggregating the second fused features corresponding to the behavior object features (Mwaura, Para. [0056], the processor can aggregate trajectories retrieved in response to a user request and that overlap with at least one specified feature (e.g., a specific object of interest, a specific region of interest, a specific timestamp, a number of objects, etc.) based on trajectory data (e.g., bounding box data, a time window, etc.)).
The proposed combination as well as the motivation for combining the Narayanmurthy, Le, Zhao, Ghanem, Qi and Mwaura references presented in the rejection of Claim 11, apply to Claim 12 and are incorporated herein by reference. Thus, the method recited in Claim 12 is met by Narayanmurthy in view of Le further in view of Zhao, Ghanem, Qi and Mwaura.
Regarding claim 14, Narayanmurthy in view of Le further in view of Zhao, Ghanem, Qi and Mwaura teaches the method according to any one of claim 11, wherein the obtaining the behavior objects and the relevant events corresponding to the at least one aggregation result comprises:
for the aggregation result, determining, based on the behavior object feature in the aggregation result, the quality of the behavior object in the frame corresponding to the behavior object feature (Mwaura, Para. [0056], the event can be formed based on an aggregation of detections that satisfies the user request (e.g., specific object of interest, specific region of interest, etc.) (i.e., quality of the behavior object is whether or not it satisfies the user request));
determining the behavior object in the video based on the quality of the behavior object in the frame corresponding to the aggregation result (Mwaura, Para. [0056], the event can be formed based on an aggregation of detections that satisfies the user request (e.g., specific object of interest, specific region of interest, etc.) (i.e., the behavior object is the specific object of interest)); and
determining relevant events of the behavior object based on each frame corresponding to the aggregation result (Mwaura, Para. [0056], the processor can aggregate trajectories retrieved in response to a user request and that overlap with at least one specified feature (e.g., a specific object of interest, a specific region of interest, a specific timestamp, a number of objects, etc.) based on trajectory data (e.g., bounding box data, a time window, etc.) to identify at least one event).
The proposed combination as well as the motivation for combining the Narayanmurthy, Le, Zhao, Ghanem, Qi and Mwaura references presented in the rejection of Claim 11, apply to Claim 14 and are incorporated herein by reference. Thus, the method recited in Claim 14 is met by Narayanmurthy in view of Le further in view of Zhao, Ghanem, Qi and Mwaura.
Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Narayanmurthy et al. (US 11,328,512 B2) in view of Roto Le (US 2023/0101399 A1) further in view of Zhao et al. (US 2021/0303864 A1), Ghanem et al. (US 2016/0110613 A1), Qi et al. (US 2023/0154207 A1) and Mwaura et al. (US 2024/0386715 A1, filed May 19, 2023), as applied to claims 11-13 and 14 above, and further in view of Mei et al. (US 2016/0267179 A1) and Yang et al. (US 2023/0334840 A1, filed June 20, 2023).
Regarding claim 13, Narayanmurthy in view of Le further in view of Zhao, Ghanem, Qi and Mwaura teaches the method according to claim 12, as described above.
Although Narayanmurthy in view of Le further in view of Zhao, Ghanem, Qi and Mwaura teaches determining similar object features (Mwaura, Para. [0056]), the references do not explicitly teach “fusing each similar object feature of the behavior object features to obtain third fused features” and “performing feature extraction on the third fused features in at least two different feature extraction modes to obtain at least two fused object features”. However, in an analogous field of endeavor, Mei teaches fusing the refined similar visual feature points from the audio index and from the visual index together (i.e., fusing each similar object feature to obtain third fused features) and selecting the most (top K) similar visual feature points from them (i.e., performing feature extraction on the third fused features to obtain at least two fused object features) (Mei, Para. [0085]).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date to modify the method of Narayanmurthy in view of Le further in view of Zhao, Ghanem, Qi and Mwaura with the teachings of Mei by including fusing similar features to obtain third fused features and extracting third fused features to obtain at least two fused features. One having ordinary skill in the art would have been motivated to combine these references because doing so would allow for analyzing video content, as recognized by Mei.
Although Narayanmurthy in view of Le further in view of Zhao, Ghanem, Qi, Mwaura and Mei teaches fusing similar object features (Mei, Para. [0085]), the references do not explicitly teach “obtaining a weight corresponding to each fused object features based on a correlation between the behavior object features and each fused object feature” and “performing weighted fusion on the fused object features by using the weight corresponding to each fused object feature to obtain the second fused features of the behavior objects”. However, in an analogous field of endeavor, Yang teaches performing weighted fusion on data features corresponding to data based on the confidences respectively corresponding to the at least two modalities to obtain a fused feature (Yang, Para. [0049]). Significance indicates a weight value of features corresponding to data for the specified classification task. For a specified classification task, in terms of the contributions of the features corresponding to the data included in the same modality to the classification and prediction result, some data features have greater impact on classification and prediction, and some data features include a small amount of information and have less impact on classification and prediction. Therefore, the data features with greater impact are of greater significance (Yang, Para. [0043]).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the method of Narayanmurthy in view of Le further in view of Zhao, Ghanem, Qi, Mwaura and Mei with the teachings of Yang by including assigning weights to the fused object features and performing weighted fusion on the object features based on the assigned weights. One having ordinary skill in the art would have been motivated to combine these references because doing so would allow for accurately performing classification and prediction, as recognized by Yang. Thus, the claimed invention would have been obvious to one having ordinary skill in the art before the effective filing date.
Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Narayanmurthy et al. (US 11,328,512 B2) in view of Roto Le (US 2023/0101399 A1) further in view of Zhao et al. (US 2021/0303864 A1), Ghanem et al. (US 2016/0110613 A1), Qi et al. (US 2023/0154207 A1) and Mwaura et al. (US 2024/0386715 A1, filed May 19, 2023), as applied to claims 11-13 and 14 above, and further in view of Meng et al. (US 2022/0406050 A1).
Regarding claim 15, Narayanmurthy in view of Le further in view of Zhao, Ghanem, Qi and Mwaura teaches the method according to claim 14, as described above.
Although Narayanmurthy in view of Le further in view of Zhao, Ghanem, Qi and Mwaura teaches aggregating results to determine relevant events (Mwaura, Para. [0056]), the references do not explicitly teach “removing a background in each frame corresponding to the aggregation result, and obtaining relevant events based on each frame with the background removed” and “clipping each frame based on an object region in each frame corresponding to the aggregation result, and the obtaining relevant events based on each clipped frame”. However, in an analogous field of endeavor, Meng teaches because the computer vision CNN is enforced to focus on the baseball and the player that are reflected in the aggregated important and finest scale first dimensional receptive field information, the aggregated important and finest scale first dimensional receptive field information is represented by an aggregated important and finest scale first dimensional receptive field information in which the irrelevant background is removed. FIG. 15 is a schematic diagram illustrating a background suppression effect of the aggregated multi-scale first dimensional receptive field information obtaining module in accordance with an embodiment of the present disclosure (Meng, Para. [0088]; Fig. 15).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date to modify the method of Narayanmurthy in view of Le further in view of Zhao, Ghanem, Qi and Mwaura with the teachings of Meng by including removing background based on the aggregation result and obtaining relevant events based on the clipped portion of the aggregation result with the background removed (Fig. 15 of Meng). One having ordinary skill in the art would have been motivated to combine these references because doing so would allow for accurate video classification of human actions or complex events, as recognized by Meng. Thus, the claimed invention would have been obvious to one having ordinary skill in the art before the effective filing date.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Emma Rose Goebel whose telephone number is (703)756-5582. The examiner can normally be reached Monday - Friday 7:30-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached at (571) 272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Emma Rose Goebel/Examiner, Art Unit 2662
/AMANDEEP SAINI/Supervisory Patent Examiner, Art Unit 2662