DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation
Claims 1-14 are not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because they are all method claims.
Claims 15-19 are not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the recitations of “memory”, “processor” and “instructions” provide sufficient structure to perform all claimed limitations.
Claim 20is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because it is an article of manufacture claim.
Specification
The disclosure is objected to because of the following informalities:
On page 8 paragraph [0030] line 7, the “stream is 224x22x3x10” ought to be changed to “stream is 224x224x3x10”.
Appropriate correction is required.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of pre-AIA 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1, 7, 10 and 12 is/are rejected under 35 U.S.C. 102(a/b) as being anticipated by Ye (“Campus Violence Detection Based on Artificial Intelligent Interpretation of Surveillance Video Sequences”, Remote Sensing, 2021, pages 1-17, Art of record IDS filed on 12/19/2024, referred as Ye hereinafter).
Regarding claim 1 as a representative claim, Ye teaches a method for detecting bullying, the method comprising (see Abstract, pages 1-2 (detecting school bullying by using a convolutional 3d neural network)):
acquiring, from a video camera by at least one processor, a live video stream of a monitored area (see Abstract (sensors and surveillance cameras are used to detect campus violence; campus violence data is extracted from received video images); page 2 (gathering campus violence video and daily life activity video with surveillance cameras; real-time video stream); pages 3-4 (section “Materials and Methods” and figure 1 (obtaining a video stream wherein the camera is surveying an area); This video stream is to be processed by pre-processing and by a neural network; this implies that at least one processor is inherently included to acquire such video);
preprocessing, by the at least one processor, the live video stream into a normalized low resolution video stream (see pages 3-4, section “Materials and Methods” (note paragraphs after figure 1; “112 pixels x 112 pixels” is a normalized low resolution video stream));
applying, by the at least one processor, 3 dimensional enhanced convolution neural network to the normalized low resolution video stream to detect bullying in the normalized low resolution video stream (see Abstract (C3D is a convolution 3D neural network); pages 4-6, section “Materials and Methods”, it discloses to use a neural network on the normalized low resolution video to determine if bullying is present. The neural network is a three dimensional neural network wherein the third dimension for the video is the inserted frames and the time of each image in the video sequence based)); and
transmitting, by a transceiver communicatively coupled with the at least one processor, a notification in response to detecting bullying (see pages 1-2 (“As smartphones became popular,…take out his/her smartphone and operate it to send an alarm to his/her parents or teacher…necessary.” (note last two lines of page 1 and first three lines of page 2, for example); thus, smartphone inherently includes “a transceiver communicatively coupled with the at least one processor”); and page 12 (section “Results” discloses that an alarm is provided for the detection which is regarded as a notification in response to detect bullying as well),
wherein the 3D enhanced CNN includes 2 dimensional video and a third dimension in time (see Abstract (C3D is a convolution 3D neural network; in “video sequence-based”, it comprises 2D video and time which is the third dimension); pages 4-6, section “Materials and Methods”, it discloses to use a three dimensional neural network for detection. The first two dimensions in video processing are the height and the width of the input images wherein the third dimension is the images at the different times wherein the different times are reflected by a different inserted frame or subsequent frames)).
Regarding claim 7, Ye further teaches wherein applying the 3D enhanced CNN to the normalized low resolution video stream comprises normalizing 3 red green blue (RGB) channels (see figure 2 and page 4 (“Every 16 frames…dimensions of one tensor are 2x16x112x112 (2 stands for three color channels, i.e., red, green, and blue. Figure 2 shows the structure of one tensor…”)).
Regarding claim 10, Ye does not disclose claim limitation “wherein the monitored area is a school” (see Abstract and section 1 (school bullying events; school bullying)).
Regarding claim 12, Ye does not further disclose claim limitation “wherein a training dataset of the 3D enhanced CNN comprises a plurality of video clips depicting labelled bullying and non-bully events” (see pages 5-6, section 3.1.3 (Classifier Design) (corresponding labels; predicted labels (non-bully events); real labels).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 3-6, 11, 15 and 17-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ye.
The advanced statements as applied to claims 1, 7, 10 and 12 above are incorporated hereafter.
Regarding claim 3, Ye does not disclose claim limitation “wherein a video frame rate of the live video stream is 5 frames per second”.
However, such claim limitation is well known in the art (Official Notice).
The motivation for doing so is to reduce video stream size.
Therefore, before the effective filing date of the instant claim invention, it would have been obvious to one of ordinary skill in the art to incorporate such claim limitation in combination with Ye for that reasons.
Regarding claim 4, Ye does not disclose claim limitations “wherein raw video resolution of the live video stream is 1920 x 1080 pixels” and “resolution of the normalized low resolution is 224 x 224 pixels”.
However, such claim limitations are well known in the art (Official Notice).
The motivation for doing so is to improve the detection because of higher quality of the 1920x1080 pixels video stream and to reduce computational cost with the low resolution 224x224 pixels. It also retains higher quality on the low resolution that is obtained from the higher resolution.
Therefore, before the effective filing date of the instant claim invention, it would have been obvious to one of ordinary skill in the art to incorporate such claim limitation in combination with Ye for that reasons.
Regarding claim 5, Ye does not disclose claim limitation “wherein the live video stream is sampled at 2 second increments comprising 10 frames”.
However, such claim limitation is well known in the art (Official Notice).
The motivation for doing so is to reduce video stream size and computational cost.
Therefore, before the effective filing date of the instant claim invention, it would have been obvious to one of ordinary skill in the art to incorporate such claim limitation in combination with Ye for that reasons.
Regarding claim 6, Ye does not disclose claim limitation “wherein the live video stream is sampled using a moving window of 5 frames”.
However, such claim limitation is well known in the art (Official Notice).
The motivation for doing so is to reduce computational cost and speed up the detection so that it could notify authority/parents sooner/faster in case bully is detected.
Therefore, before the effective filing date of the instant claim invention, it would have been obvious to one of ordinary skill in the art to incorporate such claim limitation in combination with Ye for that reasons.
Regarding claim 11, Ye does not further disclose claim limitation “wherein the 3D enhanced CNN is a generative adversarial network comprising a first sub-model used to train a second sub-model”.
However, such claim limitation is well known in the art (Official Notice).
The motivation for doing so is to improve network performance.
Therefore, before the effective filing date of the instant claim invention, it would have been obvious to one of ordinary skill in the art to incorporate such claim limitation in combination with Ye for that reasons.
Regarding claim 15, it is noted that claim recites similar claim limitations called for in the counterpart claim 1. Thus, the advanced statements as applied to claim 1 above are incorporated hereinafter. Ye does not further disclose claim limitations “a memory storing computer-executable instructions” and “at least one processor coupled with the memory and configured to execute the computer-executable instructions to”.
However, such claim limitation is well known in the art (Official Notice).
The motivation for doing so is to improve network performance.
Therefore, before the effective filing date of the instant claim invention, it would have been obvious to one of ordinary skill in the art to incorporate such claim limitation in combination with Ye for that reasons.
Regarding claim 17, it is noted that claim recites similar claim limitations called for in the counterpart claim 3 and thus is rejected for the same reasons as well.
Regarding claim 18, it is noted that claim recites similar claim limitations called for in the counterpart claim 4 and thus is rejected for the same reasons as well.
Regarding claim 19, it is noted that claim recites similar claim limitations called for in the counterpart claim 5 and thus is rejected for the same reasons as well.
Regarding claim 20, it is noted that claim recites similar claim limitations called for in the counterpart claim 15 and thus is rejected for the same reasons as well.
Claim(s) 2, 8-9 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ye in view of Sander et al. (“MobileNetV2: Inverted Residuals and Linear Bottlenecks”, arXiv:1801.04381v4 [cs.CV] 21 Mar 2019, pages 1-14, referred as Sander hereinafter).
The advanced statements as applied to claims 1, 7, 10 and 12 above are incorporated hereinafter.
Regarding claim 2, Ye does not disclose claim limitation “wherein the 3D enhanced CNN is an enhanced MobileNet-V2 network”.
However, such claim limitation is well known in the art as evidenced by Sander (see abstract (MobileNetV2); sections 3.1 and 4).
The motivation for doing so is to reduce computational cost as suggested by Sander (see section 3.1).
Therefore, before the effective filing date of the instant claim invention, it would have been obvious to one of ordinary skill in the art to incorporate such claim limitation as taught by Sander in combination with Ye for that reasons. Also, it would have been obvious to incorporate such claim limitation as taught by Sander into Ye because doing so would merely combine prior art elements according known method to yield predictable results.
Regarding claim 8, the combination of Ye and Sander discloses claim limitation “wherein applying the 3D enhanced CNN to the normalized low resolution video stream comprises applying 15 bottlenecks” (see section 4, Table 2 (note the repeated n times of the bottlenecks)).
Regarding claim 9, the combination of Ye and Sander discloses claim limitation “wherein 2 bottlenecks are applied to a bottleneck operation having 28x28x32xl0 parameters and 3 bottlenecks are applied to a bottleneck operation having l4xl4x64x10 parameters” (see section 4, Table 2).
Regarding claim 16, it is noted that claim recites similar claim limitations called for in the counterpart claim 2 and thus is rejected for the same reasons as well.
Allowable Subject Matter
Claims 13-14 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter:
Regarding claim 13, the cited prior art does not teach or suggest claim limitations “wherein each of the plurality of video clips comprise an audio portion and a visual portion, and wherein the training dataset links the audio portion to the visual portion using timestamps, further comprising: detecting a plurality of keywords in the audio portion; and classifying an action in the visual portion”.
Claim 14 depends on claim 13 and thus is allowable for the reasons as well,
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Fu et al. (U.S. Pat. App. Pub. No. 2022/0156554 A1) teaches CNN comprising MobileNet V2 (see figure 14) and bottlenecks (see table 8).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DUY M DANG whose telephone number is (571)272-7389. The examiner can normally be reached Monday to Friday from 7:00AM to 3:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached at 571-272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
DMD
7/2026
/DUY M DANG/Primary Examiner, Art Unit 2662