DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-5, 8 are rejected under 35 U.S.C. 103 as being unpatentable over Kim (US PG Publication 2024/0233385) in view of Bernal (US PG Publication 2018/0063538).
Regarding Claim 1, Kim (US PG Publication 2024/0233385) discloses a method for processing data (video captioning server 120, Fig. 1 [0053]), comprising:
obtaining video data (video caption unit 123 can separate video data into vision data and audio data [0055]; vision server 121 collecting vision data of video data [0054]) and extracting a video feature (creating a vision attention vector [0058]) from the obtained video data (vision data [0058]);
obtaining audio (audio server 122 collecting audio data of video data [0054]) data related to the video data (of video data [0054]; video caption unit 123 can separate video data into vision data and audio data [0055]) and extracting an audio feature (audio attention vector [0058]) from the obtained audio data (audio data [0058]);
detecting a preset event (when a specific dangerous behavior is sensed [0051], automatically detect behavior events [0063]) on the basis of one or more of the video feature and the audio feature (Feature values of an I3D model and a VGGish model can be configured into a multi-modal type in a vanilla transformer architecture [0063]);
upon detecting the preset event (when a specific dangerous behavior is sensed [0051], automatically detect behavior events [0063]) … generating transmission data (reported to a manager and detailed information about a criminal situation is transmitted [0050]) ….
Kim does not disclose, but Bernal (US PG Publication 2018/0063538) teaches upon detecting a preset event (based on the content of the image data, objects classified as vehicles, people, structures, determining the type of activity being carried out [0021]; analyze the raw image data to determine content, step 230, Fig. 2), performing conversion processing (adjusts a compression parameter 340 based on the content of the image data… this becomes a modified compression algorithm [0021]; vector encoding [0038]) and encoding processing (Huffman, arithmetic or Lempel Ziv coding [0038]) on the video feature (by including features descriptive of objects of interest classified as vehicles, features descriptive of objects of interest classified as structures, features descriptive of actions of interest, features that discriminate actions A, B, C and D from each other [0021]) and the audio (frequency and phase descriptors in the case of one-dimensional sequential data such as audio [0030]) feature (the feature space [0038]);
and generating transmission data including (compressed data stream can be stored or transmitted [0031]) the video feature (by including features descriptive of objects of interest classified as vehicles, features descriptive of objects of interest classified as structures, features descriptive of actions of interest, features that discriminate actions A, B, C and D from each other [0021]) and the audio (frequency and phase descriptors in the case of one-dimensional sequential data such as audio [0030]) feature (the compression module may generate a compressed data representation of the feature representation extracted by the feature extraction module [0028]) on which the conversion processing and the encoding processing have been performed (vector quantization encoding and arithmetic/Huffman/Lempel Ziv [0038]).
One of ordinary skill in the art before the application was filed would have been motivated to supplement the event detection of Kim with content-based compression, as in Bernal, because Bernal teaches that compressed image/video data comprises less data than raw image/video data, facilitating better communication of the information via bandwidth constrained communication link and alleviating the bottleneck that communication link 360 could otherwise have posed in the data collection process [0021].
Regarding Claim 2, Kim (US PG Publication 2024/0233385) discloses the method of claim 1, wherein, in the detecting of the preset event (specific dangerous behavior is sensed [0051], automatically detect behavior events [0063]), upon satisfaction of a condition related to the preset event in both the video feature (figure in-video behavior and motion [0061]) and the audio feature (classification of features [0062]), it is determined that the preset event has been detected (automatically detect behavior events and create video caption information in an artificial intelligence model. Accordingly, it is possible to easily figure out the context in each section by automatically setting breakpoints (behavior stop points) using all of vision and audio information [0063]).
Regarding Claim 3, Kim (US PG Publication 2024/0233385) discloses the method of claim 1, wherein, in the detecting of the preset event (specific dangerous behavior is sensed [0051], automatically detect behavior events [0063]), upon satisfaction of a condition related to the preset event in any one of the video feature and the audio feature (classification of features [0062]), it is determined that the preset event has been detected (automatically detect behavior events [0063]).
Regarding Claim 4, Kim (US PG Publication 2024/0233385) discloses the method of claim 1.
Kim does not disclose, but Bernal (US PG Publication 2018/0063538) teaches wherein the performing of the conversion processing and the encoding processing includes:
converting (compressed data stream [0031]) the video feature (by including features descriptive of objects of interest classified as vehicles, features descriptive of objects of interest classified as structures, features descriptive of actions of interest, features that discriminate actions A, B, C and D from each other [0021]) and encoding (Huffman, arithmetic or Lempel Ziv coding [0038]) the converted video feature to generate video feature encoded data (output of Huffman, arithmetic or Lempel Ziv coding [0038]);
and converting (compressed data stream [0031]) the audio feature (frequency and phase descriptors in the case of one-dimensional sequential data such as audio [0030]) and encoding (Huffman, arithmetic or Lempel Ziv coding [0038]) the converted audio feature to generate audio feature encoded data (output of Huffman, arithmetic or Lempel Ziv coding [0038]).
One of ordinary skill in the art before the application was filed would have been motivated to supplement the event detection of Kim with content-based compression, as in Bernal, because Bernal teaches that compressed image/video data comprises less data than raw image/video data, facilitating better communication of the information via bandwidth constrained communication link and alleviating the bottleneck that communication link 360 could otherwise have posed in the data collection process [0021].
Regarding Claim 5, Kim (US PG Publication 2024/0233385) discloses the method of claim 4.
Kim does not disclose, but Bernal (US PG Publication 2018/0063538) teaches wherein the transmission data (output of Huffman, arithmetic or Lempel Ziv coding [0038]) includes the video feature encoded data (features descriptive of objects of interest classified as vehicles, features descriptive of objects of interest classified as structures, features descriptive of actions of interest, features that discriminate actions A, B, C and D from each other [0021]) and the audio feature encoded data (frequency and phase descriptors in the case of one-dimensional sequential data such as audio [0030]).
One of ordinary skill in the art before the application was filed would have been motivated to supplement the event detection of Kim with content-based compression, as in Bernal, because Bernal teaches that compressed image/video data comprises less data than raw image/video data, facilitating better communication of the information via bandwidth constrained communication link and alleviating the bottleneck that communication link 360 could otherwise have posed in the data collection process [0021].
Regarding Claim 8, Kim (US PG Publication 2024/0233385) discloses the method of claim 1, wherein the transmission data includes metadata related to the detected event (subtitle data related to video data may be obtained by a caption unit 242 [0060], Fig. 2).
Claim(s) 6-7, 9 are rejected under 35 U.S.C. 103 as being unpatentable over Kim (US PG Publication 2024/0233385) in view of Bernal (US PG Publication 2018/0063538) and Li (NPL: “Video Description Combining Visual and Audio Features,” SPIE April 2023).
Regarding Claim 6, Kim (US PG Publication 2024/0233385) discloses the method of claim 1.
Kim does not disclose, but Bernal (US PG Publication 2018/0063538) teaches wherein the performing of the conversion processing and the encoding processing includes:
converting (compressed data stream [0031]) the video (by including features descriptive of objects of interest classified as vehicles, features descriptive of objects of interest classified as structures, features descriptive of actions of interest, features that discriminate actions A, B, C and D from each other [0021]) feature (the feature space [0038]);
converting (compressed data stream [0031]) the audio (frequency and phase descriptors in the case of one-dimensional sequential data such as audio [0030]) feature (the feature space [0038]).
Bernal does not teach, but Li (NPL: “Video Description Combining Visual and Audio Features,” SPIE April 2023) teaches fusing the converted video feature and the converted audio feature to generate fusion feature data (By splicing the two features, the spliced features are linearly transformed to unify the features, and then the average pooling operation is used to obtain the global features; features are embedded in time sequence, Section 3.1).
One of ordinary skill in the art before the application was filed would have been motivated to supplement the event detection of Kim with content-based compression, as in Bernal, because Bernal teaches that compressed image/video data comprises less data than raw image/video data, facilitating better communication of the information via bandwidth constrained communication link and alleviating the bottleneck that communication link 360 could otherwise have posed in the data collection process [0021].
One of ordinary skill in the art before the application was filed would have been motivated modify Kim, as supplemented by Bermal, with timing information because Li suggests that fusing the audio-visual data can generate more accurate description text (Section 5) and improve the convergence speed of the model (Section 3.4).
Regarding Claim 7, Kim (US PG Publication 2024/0233385) discloses the method of claim 6.
Kim does not disclose, but Li (NPL: “Video Description Combining Visual and Audio Features,” SPIE April 2023) teaches wherein the transmission data includes fusion-encoded data generated by encoding the fusion feature data (features are embedded in time sequence through time sequence encoding and sent to the encoder of the transformer).
One of ordinary skill in the art before the application was filed would have been motivated modify Kim, as supplemented by Bermal, with Li because Li suggests that fusing the audio-visual data can generate more accurate description text (Section 5) and improve the convergence speed of the model (Section 3.4).
Regarding Claim 9, Kim (US PG Publication 2024/0233385) discloses the method of claim 1.
Kim does not disclose, but Li (NPL: “Video Description Combining Visual and Audio Features,” SPIE April 2023) teaches wherein the transmission data includes time stamp information for synchronization of the video feature and the audio feature (as a sequence file, the video needs to be embedded with timing encoding to enable the model to learn the context information, Section 3.4).
One of ordinary skill in the art before the application was filed would have been motivated modify Kim, as supplemented by Bermal, with Li because Li suggests that fusing the audio-visual data can generate more accurate description text, and embedding timing information enables the encoder to better recognize the sequence information in the video, resulting in a more accurate predictive description of the content (Section 5) and improved convergence speed of the model (Section 3.4).
Claim(s) 10, 11, 12, are rejected under 35 U.S.C. 103 as being unpatentable over Chambers (US PG Publication 2009/0141939) in view of Fujimura (US PG Publication 2023/0153610) and Bernal (US PG Publication 2018/0063538).
Regarding Claim 10, Chambers (US PG Publication 2009/0141939) discloses a method (a process for sending event notices [0025], Fig. 2) comprising:
upon detecting the preset event (event of interest? Yes, Fig. 2, step 120, yes-branch), begin performing (select portion of video associated with event of interest, Step 130, Fig. 2; generate an event notice 140 … the event notice may include a segment of video [0024]) conversion processing (video segment is compressed [0065]) and encoding processing (video segment is encoded [0065]) on the video (video segment [0065]) …;
the conversion processing and encoding processing not being performed prior to detecting the preset event (NO branch at 120 in Fig. 2 does not generate the event notice, which then does not send the video segment);
upon detecting the preset event (event of interest? Yes, Fig. 2, step 120, yes-branch), begin generating transmission data (send event notice to the user, Fig. 2, step 150) including the video (a portion of the video data associated with the event of interest is selected (step 130, FIG. 2) and then the event notice ("ALERT!") is generated (step 140, FIG. 2) and sent to the user [0060]) … on which the conversion processing (video segment is compressed [0065]) and the encoding processing (video segment is encoded [0065]) have been performed (is compressed, encoded [0065]), and
the transmission data not being generated prior to detecting the preset event (NO branch at 120 in Fig. 2 does not generate the event notice, which then does not send the event notice, does not send the video segment).
Chambers does not disclose, but Fujimura (US PG Publication 2023/0153610) teaches obtaining audio data (sound files from the captured video [0066]) and extracting an audio feature (The neural network 140 receives input(s) 120, which is processed by the neurons 145 to produce the output(s) 125 [0056], note, extracting features is an inherent part of neural networks) from the obtained audio data (sound files from the captured video go to the sound processing machine trained network 110 [0066]);
detecting a preset event (sound-processing machine trained network 110 produces from the supplied sound file 120 the impact location 125 of each impact sound of each golf swing [0067]) on the basis of the audio feature (The neural network 140 receives input(s) 120, which is processed by the neurons 145 to produce the output(s) 125 [0056], note, extracting features is an inherent part of neural networks);
upon detecting the preset event (the precise moment of impact from [0067]), begin extracting a video feature from video data related to the detected event (the video-processing machine trained network 310 also receives the output 125 of the sound processing machine trained network 110 for each golf swing [0067]; each portion in the video that is associated with a golf swing is fed to the video processing machine trained network 310 as input to produce the video-processing machine trained network 310 output data 325 [0068]), the video feature not being extracted prior to detecting the preset event (receives the output 125 of the sound processing machine trained network 110 [0067]… each portion in the video that is associated with a golf swing is fed to the video processing machine trained network 310 [0068]; that means that video is sent to the video classifier after the golf swing is detected).
Chambers does not disclose, but Bernal (US PG Publication 2018/0063538) teaches conversion processing (the compression module may generate a compressed data representation of the feature representation extracted by the feature extraction module [0028]; vector encoding [0038]) and encoding processing (Huffman, arithmetic or Lempel Ziv coding [0038]) on the video feature (by including features descriptive of objects of interest classified as vehicles, features descriptive of objects of interest classified as structures, features descriptive of actions of interest, features that discriminate actions A, B, C and D from each other [0021]) and the audio (frequency and phase descriptors in the case of one-dimensional sequential data such as audio [0030]) feature (the feature space [0038]);
transmission data including (compressed data stream can be stored or transmitted [0031]) the video feature (by including features descriptive of objects of interest classified as vehicles, features descriptive of objects of interest classified as structures, features descriptive of actions of interest, features that discriminate actions A, B, C and D from each other [0021]) and the audio (frequency and phase descriptors in the case of one-dimensional sequential data such as audio [0030]) feature (the compression module may generate a compressed data representation of the feature representation extracted by the feature extraction module [0028]) on which the conversion processing and the encoding processing have been performed (vector quantization encoding and arithmetic/Huffman/Lempel Ziv [0038]).
One of ordinary skill in the art before the application was filed would have been motivated to modify event detection of Chambers using Fujimura to identify video segments based on precise moments identified through sound because Fujimura teaches that video does not have sufficient temporal resolution to identify all major events, and relying on sound data can enable the system to interpolate additional data samples in video [0067], improving the system.
One of ordinary skill in the art before the application was filed would have been motivated to replace the compression and encoding method of Chambers with the compression and encoding method of Bernal, because Bernal’s compression and encoding preserves the high fidelity features essential for the decision making task, enabling a computer and a human to analyze information over the network without the loss of data that generally comes from transmitting data over a network, facilitating a better analysis result ([0004], [0005], [0007]) without overburdening the bandwidth constrained communication link [0021].
Regarding Claim 11, Chambers (US PG Publication 2009/0141939) discloses the method of Claim 10.
Chambers does not disclose, but Bernal (US PG Publication 2018/0063538) teaches wherein the performing of the conversion processing and the encoding processing includes:
converting (compressed data stream [0031]) the video feature (by including features descriptive of objects of interest classified as vehicles, features descriptive of objects of interest classified as structures, features descriptive of actions of interest, features that discriminate actions A, B, C and D from each other [0021]) and encoding (Huffman, arithmetic or Lempel Ziv coding [0038]) the converted video feature to generate video feature encoded data (output of Huffman, arithmetic or Lempel Ziv coding [0038]);
and converting (compressed data stream [0031]) the audio feature (frequency and phase descriptors in the case of one-dimensional sequential data such as audio [0030]) and encoding (Huffman, arithmetic or Lempel Ziv coding [0038]) the converted audio feature to generate audio feature encoded data (output of Huffman, arithmetic or Lempel Ziv coding [0038]).
One of ordinary skill in the art before the application was filed would have been motivated to replace the compression and encoding method of Chambers with the compression and encoding method of Bernal, because Bernal’s compression and encoding preserves the high fidelity features essential for the decision making task, enabling a computer and a human to analyze information over the network without the loss of data that generally comes from transmitting data over a network, facilitating a better analysis result ([0004], [0005], [0007]) without overburdening the bandwidth constrained communication link [0021].
Regarding Claim 12, Chambers (US PG Publication 2009/0141939) discloses the method of Claim 11.
Chambers does not disclose, but Bernal (US PG Publication 2018/0063538) teaches wherein the transmission data (output of Huffman, arithmetic or Lempel Ziv coding [0038]) includes the video feature encoded data (features descriptive of objects of interest classified as vehicles, features descriptive of objects of interest classified as structures, features descriptive of actions of interest, features that discriminate actions A, B, C and D from each other [0021]) and the audio feature encoded data (frequency and phase descriptors in the case of one-dimensional sequential data such as audio [0030]).
One of ordinary skill in the art before the application was filed would have been motivated to replace the compression and encoding method of Chambers with the compression and encoding method of Bernal, because Bernal’s compression and encoding preserves the high fidelity features essential for the decision making task, enabling a computer and a human to analyze information over the network without the loss of data that generally comes from transmitting data over a network, facilitating a better analysis result ([0004], [0005], [0007]) without overburdening the bandwidth constrained communication link [0021].
Claim(s) 13-14, 16 are rejected under 35 U.S.C. 103 as being unpatentable over Chambers (US PG Publication 2009/0141939) in view of Fujimura (US PG Publication 2023/0153610), Bernal (US PG Publication 2018/0063538), and Li (NPL: “Video Description Combining Visual and Audio Features,” SPIE April 2023).
Regarding Claim 13, Chambers (US PG Publication 2009/0141939) discloses the method of claim 10, wherein the performing of the conversion processing (video segment is compressed [0065]) and the encoding processing (video segment is encoded [0065]) includes:
converting the video (video segment is compressed [0065]).
Chambers does not disclose, but Bernal (US PG Publication 2018/0063538) teaches
converting (the compression module may generate a compressed data representation of the feature representation extracted by the feature extraction module [0028]; compressed data stream [0031]) the video (by including features descriptive of objects of interest classified as vehicles, features descriptive of objects of interest classified as structures, features descriptive of actions of interest, features that discriminate actions A, B, C and D from each other [0021]) feature (the feature space [0038]);
converting (the compression module may generate a compressed data representation of the feature representation extracted by the feature extraction module [0028]; compressed data stream [0031]) the audio (frequency and phase descriptors in the case of one-dimensional sequential data such as audio [0030]) feature (the feature space [0038]).
Chambers does not disclose, but Li (NPL: “Video Description Combining Visual and Audio Features,” SPIE April 2023) teaches fusing the converted video feature and the converted audio feature to generate fusion feature data (By splicing the two features, the spliced features are linearly transformed to unify the features, and then the average pooling operation is used to obtain the global features; features are embedded in time sequence, Section 3.1).
One of ordinary skill in the art before the application was filed would have been motivated to replace the compression and encoding method of Chambers with the compression and encoding method of Bernal, because Bernal’s compression and encoding preserves the high fidelity features essential for the decision making task, enabling a computer and a human to analyze information over the network without the loss of data that generally comes from transmitting data over a network, facilitating a better analysis result ([0004], [0005], [0007]) without overburdening the bandwidth constrained communication link [0021].
One of ordinary skill in the art before the application was filed would have been motivated modify the compression of Chambers, as modified by Bernal, with fused data and timing information because Li suggests that fusing the audio-visual data can generate more accurate description text (Section 5) and improve the convergence speed of the model (Section 3.4).
Regarding Claim 14, Chambers (US PG Publication 2009/0141939) discloses the method of claim 13.
Chambers does not disclose, but Li (NPL: “Video Description Combining Visual and Audio Features,” SPIE April 2023) teaches wherein the transmission data includes fusion-encoded data obtained by encoding the generated fusion feature data (features are embedded in time sequence through time sequence encoding and sent to the encoder of the transformer).
One of ordinary skill in the art before the application was filed would have been motivated modify the compression of Chambers, as modified by Bernal, with fused data and timing information because Li suggests that fusing the audio-visual data can generate more accurate description text (Section 5) and improve the convergence speed of the model (Section 3.4).
Regarding Claim 16, Chambers (US PG Publication 2009/0141939) discloses the method of claim 10.
Chambers does not disclose, but Li (NPL: “Video Description Combining Visual and Audio Features,” SPIE April 2023) teaches wherein the transmission data includes time stamp information for synchronization of the video feature and the audio feature (as a sequence file, the video needs to be embedded with timing encoding to enable the model to learn the context information, Section 3.4).
One of ordinary skill in the art before the application was filed would have been motivated modify the compression of Chambers, as modified by Bernal, with fused data and timing information because Li suggests that fusing the audio-visual data can generate more accurate description text (Section 5) and improve the convergence speed of the model (Section 3.4).
Claim(s) 15 is rejected under 35 U.S.C. 103 as being unpatentable over Chambers (US PG Publication 2009/0141939) in view of Fujimura (US PG Publication 2023/0153610), Bernal (US PG Publication 2018/0063538), and Kim (US PG Publication 2024/0233385).
Regarding Claim 15, Chambers (US PG Publication 2009/0141939) discloses the method of Claim 10.
Chambers does not disclose, but Kim (US PG Publication 2024/0233385) teaches wherein the transmission data includes metadata related to the detected event (subtitle data related to video data may be obtained by a caption unit 242 [0060], Fig. 2).
One of ordinary skill in the art before the application was filed to supplement the video of Chambers with subtitles, as in Kim, because the subtitles of Kim supply semantic understanding to the video, improving the user’s ability to review alerts in a short time with computer-generated summaries.
Claim(s) 20 is rejected under 35 U.S.C. 103 as being unpatentable over Chambers (US PG Publication 2009/0141939) in view of Kim (US PG Publication 2024/0233385) and Bernal (US PG Publication 2018/0063538).
Regarding Claim 20, Chambers (US PG Publication 2009/0141939) discloses a method for processing data (a process for sending event notices [0025], Fig. 2) comprising:
obtaining video data (video camera 20 is positioned to view the observed location 10 and can provide video data to a video analysis functional block 30 [0023]; camera 20 gives video analysis 30 video of observed location 10, Fig. 1) …;
detecting a preset event (event of interest has occurred [0024]; event of interest? Yes, Fig. 2, step 120, yes-branch) …;
upon detecting the preset event (event of interest? Yes, Fig. 2, step 120, yes-branch), performing (select portion of video associated with event of interest, Step 130, Fig. 2; generate an event notice 140 … the event notice may include a segment of video [0024]) conversion processing (video segment is compressed [0065]) and encoding processing (video segment is encoded [0065]) on the video (video segment [0065]) …;
the conversion processing and encoding processing not being performed prior to detecting the preset event (NO branch at 120 in Fig. 2 does not generate the event notice, which then does not send the video segment);
generating transmission data (send event notice to the user, Fig. 2, step 150) including the video (a portion of the video data associated with the event of interest is selected (step 130, FIG. 2) and then the event notice ("ALERT!") is generated (step 140, FIG. 2) and sent to the user [0060]) … on which the conversion processing (video segment is compressed [0065]) and the encoding processing (video segment is encoded [0065]) have been performed (is compressed, encoded [0065]), and
the transmission data not being generated prior to detecting the preset event (NO branch at 120 in Fig. 2 does not generate the event notice, which then does not send the event notice, does not send the video segment).
Chambers does not disclose, but Kim (US PG Publication 2024/0233385) teaches obtaining video data (video caption unit 123 can separate video data into vision data and audio data [0055]; vision server 121 collecting vision data of video data [0054]) and extracting a video feature (creating a vision attention vector [0058]) from the obtained video data (vision data [0058]);
obtaining audio (audio server 122 collecting audio data of video data [0054]) data related to the video data (of video data [0054]; video caption unit 123 can separate video data into vision data and audio data [0055]) and extracting an audio feature (audio attention vector [0058]) from the obtained audio data (audio data [0058]);
detecting a preset event (when a specific dangerous behavior is sensed [0051], automatically detect behavior events [0063]) on the basis of one or more of the video feature and the audio feature (Feature values of an I3D model and a VGGish model can be configured into a multi-modal type in a vanilla transformer architecture [0063]).
Chambers does not disclose, but Bernal (US PG Publication 2018/0063538) teaches conversion processing (the compression module may generate a compressed data representation of the feature representation extracted by the feature extraction module [0028]; vector encoding [0038]) and encoding processing (Huffman, arithmetic or Lempel Ziv coding [0038]) on the video feature (by including features descriptive of objects of interest classified as vehicles, features descriptive of objects of interest classified as structures, features descriptive of actions of interest, features that discriminate actions A, B, C and D from each other [0021]) and the audio (frequency and phase descriptors in the case of one-dimensional sequential data such as audio [0030]) feature (the feature space [0038]);
transmission data including (compressed data stream can be stored or transmitted [0031]) the video feature (by including features descriptive of objects of interest classified as vehicles, features descriptive of objects of interest classified as structures, features descriptive of actions of interest, features that discriminate actions A, B, C and D from each other [0021]) and the audio (frequency and phase descriptors in the case of one-dimensional sequential data such as audio [0030]) feature (the compression module may generate a compressed data representation of the feature representation extracted by the feature extraction module [0028]) on which the conversion processing and the encoding processing have been performed (vector quantization encoding and arithmetic/Huffman/Lempel Ziv [0038]).
One of ordinary skill in the art before the application was filed would have been motivated replace the event detection of Chambers using the multi-modal event detection of Kim because Kim teaches that detection based on video alone is limited and requires an immense amount of training data [0004], whereas multi-modal detection provides superior results and can automatically provide situation recognition and information [0005], improving the system.
One of ordinary skill in the art before the application was filed would have been motivated to replace the compression and encoding method of Chambers with the compression and encoding method of Bernal, because Bernal’s compression and encoding preserves the high fidelity features essential for the decision making task, enabling a computer and a human to analyze information over the network without the loss of data that generally comes from transmitting data over a network, facilitating a better analysis result ([0004], [0005], [0007]) without overburdening the bandwidth constrained communication link [0021].
Response to Arguments
Applicant’s remarks filed 6/11/2026 have been considered but are unpersuasive.
Applicant argues, regarding Claim 1, error grounds ## 1, 2, and 4, that Examiner has not given weight to “upon detecting the preset event,” Remarks at 11-12; performing conversion 100% of the time is not a reasonable interpretation, Remarks at 12-13; Kim performs conversion regardless of whether the preset event is detected, Remarks at 14. These arguments are not persuasive.
The language of Claim 1 is “upon detecting the preset event, performing conversion.” There is no inverse of this clause in the claim. The broadest reasonable interpretation is, upon detecting, conversion is performed, and upon not-detecting, there are no limitations. Applicant’s interpretation of Claim 1 is predicated on the “fallacy of the inverse:” denying the antecedent. That is, Applicant concludes that because Claim 1 recites that “upon detecting… performing conversion,” then Claim 1 also recites upon not detecting… not performing conversion. This is a false conclusion. Claim 1 is silent on what happens upon not-detecting. Examiner understands Applicant’s intent to not perform conversion, but Applicant must first amend the claim include that limitation. Until then, Applicant is arguing limitations that are not claimed.
Applicant’s assertion that Kim performs conversion 100% of the time is not persuasive. Applicant appears to have mis-mapped “conversion” to steps performed in Kim. Applicant’s definition of “conversion” is “compression.” See, “Both Kim and Bernal performs ‘conversion’ or ‘compression’…,” Remarks at 15; “…[T]he video feature information may be converted by applying one or more methods of … bit reduction, quantization, and filtering”, Spec. at p. 15 ll. 13-15. The office action does not rely on Kim to perform compression. The citations that Applicant relies on to argue that Kim compresses 100% of the time, are instead citations to Kim detecting a preset event, i.e., detecting events using I3D and VGGish, Remarks at 14. As a result, this argument is not persuasive.
Applicant further argues in error #4 that Bernal does not perform “detecting the preset event” during classification; Bernal determines classification 100% of the time; there is no “event detection.” Remarks at 15. This is unclear. The office action maps Bernal’s “classification” to “detecting a preset event.” This mapping is proper. Bernal detects objects of interest classified as people and vehicles, determines a type of activity being carried out, and discriminates actions A, B, C, and D. Bernal at [0021]. This means that Bernal detects preset events such as the presence of humans or vehicles, and actions A, B, C, and D. Applicant’s argument here is not clear.
Applicant’s argument that Bernal performs conversion 100% of the time is unpersuasive. The dispositive question is whether Bernal performs the conversion “upon detecting the preset event,” not whether it is 100% of the time.
While Applicant has proffered the interpretation that “conversion” means “compression,” “conversion” is broader than compression. The as-filed specification states that the converter may convert the video feature into a form suitable for compressing. Spec. at p. 15 ll. 9-11. Bernal converts features into a form suitable for compression upon detecting the presence of humans or vehicles, and actions A, B, C, and D (the preset event). Bernal at [0021]. Bernal selects different features to compress (making suitable for compression) based on the detected content (preset event). Bernal at [0021]. Bernal adjusts a compression parameter based on the content of the image data by including features of descriptive objects. Bernal at [0021]. Because the content of the image data (detecting the preset event) happens before and informs selecting features for compression (making suitable for compression), Bernal teaches “upon detecting the preset event, performing conversion….” Whether Bernal performs compression 100% of the time does not matter as long as Bernal performs conversion upon detecting the preset event.
Applicant argues in #3 that the office action maps “detecting a/the preset event” to two different, mutually exclusive features of Kim: “stop points” and “specific dangerous behavior.” Remarks at 13. “Stop points” and “specific dangerous behavior” are synonymous because stop points are added to video in spaces where dangerous behavior is identified. Kim at [0063]. To avoid confusion, Examiner has mapped detecting “preset events” in Kim to “specific dangerous behaviors.”
Applicant argues, regarding grounds ## 5 and 6 that Li does not perform conversion upon detecting an event. Remarks at 16-17. This argument is unpersuasive because Li is not relied upon to teach this feature.
On page 18 of remarks, Applicant has not presented arguments beyond what has been addressed for Claim 1.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
NPL: Liu, “Visually-aware audio captioning with adaptive audio-visual attention,” arXiv:2210.16428v3, May 2023. – an encoder-decoder architecture for combining visual and audio features for captioning video, using an adaptive self-attention block.
NPL: Lin, “TAVT: Towards Transferable Audio-Visual Text Generation,” Association for Computational Linguistics, July 2023. – an encoder-decoder architecture for adding text to videos using a meta-mapper to identify semantic audio features.
US-20230020834-A1 – encoder-decoder network for captioning video
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHADAN E HAGHANI whose telephone number is (571)270-5631. The examiner can normally be reached M-F 9AM - 5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jay Patel can be reached at 571-272-2988. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHADAN E HAGHANI/ Examiner, Art Unit 2485