DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 1 is rejected under 35 U.S.C. 103 as being unpatentable over Sawaya et al (11,830,252) in view of Keel (9,185,147) and Banfield (12,367,233).
Regarding claim 1 Sawaya discloses
Shoplifting detection from surveillance camera using artificial intelligence technology includes the following steps:
step 1: data preprocesssing (note security control device, col. 6 lines 34-35); an input of this step is a video stream transmitted from the surveillance camera (note fig. 3 block s100 and col. 6 lines 35-36, to receive video surveillance data), Sawaya discloses video stream from the surveillance data. Sawaya does not clearly discloses video cutting into 1-second clips with sliding window step of 0.5 seconds; Keel disclose video production including video cutting clips (note Keel, col. 16 lines 26-38, feed input interface for cutting videos). Sawaya and Keel are combinable because they are from the same field of endeavor. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include video cutting in the system of Sawaya as evidenced by Keel. The suggestion/motivation for doing so provides beneficial for the simultaneous monitoring of large numbers of security cameras. For example, a security guard might need to monitor fifty different security cameras. The Feed Input Manager would automatically import short video sequences from cameras whenever motion is detected and package those video clips as Cards (note col. 16 lines 41-47)
processed to be an input of a model (note Sawaya, col. 6 lines 46-47, control device configured to identify using machine learning model),
step 2: Sawaya discloses person feature extraction and shoplifting behavior probability calculation (note Sawaya, fig. 3 block s104, and col. 6 lines 40-45, video surveillance data, person associated with shoplifting activity); Sawaya discloses a deep learning model. Sawaya and Keel do not clearly disclose a hybid deep learning model which combines a three-dimensional convolutional neural network (3D-CNN) and a two-dimensional convolutional neural network (2D-CNN) is used to learn spatial-temporal features and positions of people in a keyframe of the 1-second clip processed in step 1, then a probability of shoplifting behavior of each person in the keyframe is computed and a position of them is determined; Banfield discloses a hybrid deep learning model (note Banfield, col. 8 lines 13-50). Sawaya, Keel and Banfield are combinable because they are from the same field of endeavor. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to substitute Sawaya’s deep learning model with hybrid deep learning model of Banfield. The suggestion/motivation for doing so hybrid deep learning model searching yielding adequate results (note Banfield, col. 1 lines 25-31). It would have been obvious to combine Banfield with Sawaya and Keel to obtain the invention as specified by claim 1.
step 3: Sawaya discloses post-processing and warning (note Sawaya, fig. 3, block s110 and col. 6 lines 60-67, generate audio message associated with shoplifting activity instructs person to leave the premises) in this step, the probability of shoplifting behavior in entire video is aggregated based on the probability of 1-second clips computed in step 2, thereby giving a warning of the shoplifter if any (note Sawaya col. 6 lines 60-67, audio instructing person to leave the premises).
Allowable Subject Matter
Claims 2-4 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter for dependent claim 2-4.
Regarding claim 2, prior art could not be found for the features wherein: in step 1, the video stream from surveillance camera is cut into 1-second clips, then these 1-second clips are presented at the same playback rate of 25 frames per second (FPS); therefore, each video segment includes 25 frames with the sliding window is 0.5 seconds corresponding to 12 frames; in addition, the video stream from surveillance camera is often high resolution (HD, Full HD, 2K, etc.), so it is necessary to downsample and resize the frames to reduce inference costs, increase computational speed, and reduce memory storage; specifically, 12 frames are sampled from 25 frames according to a uniform distribution; after that, the frames are resized to 640×640×3 and normalized to a normal distribution Ν(0,1) to be the input of the deep learning model in step 2. These features in combination with other features could not be found in the prior art.
Regarding claim 3, prior art could not be found for the features wherein: in step 2, the input of this step is the sequence of frames processed in step 1, which is fed into a hybrid deep learning model consisting of two parallel main branches; the first branch is a 3D-CNN model, which is SlowFast-R50; its input is entire 1-second clip segment, and it extracts spatiotemporal features of all frames; the second branch is a 2D-CNN model, which is Yolov5; it returns the position of each person in the frame; next, the spatiotemporal (3-dimensional) features and position of each person from these two branches are combined using the region-of-interest (ROI) align algorithm, and the output is spatiotemporal features of all people; after that, these features are fed into a fully connected layer to classify the actions of each person in the keyframe, and the output is the probability of shoplifting behavior for each person in keyframe; finally, the shoplifting behavior probability in entire video segment is the highest probability among all the people in the keyframe; the bounding boxes of all people are also realigned to match each other in the output video. These features in combination with other features could not be found in the prior art.
Regarding claim 4, the prior art could not be found for the feature wherein: in step 3, with the video stream transmitted from the surveillance camera, the post-processing and warning (if any) are performed every 5 seconds; the input of this step is the probability of shoplifting behavior occuring in each 1-second clip from step 2, denoting this probability as pclip; if pclip≥θ, θ is a threshold, then the 1-second clip is predicted to exist shoplifting behavior; based on experiments, the threshold of θ=0.8 achieves the best prediction accuracy; because the sliding window of 1-second clip segment is 0.5 seconds, to be sure about the prediction result, the number of 1-second clips with probability pclip≥θ is counted and aggregated (ntimes); when n≥k (k is another threshold), the system will alert that shoplifting behavior is occurring; based on experiments, k= 3 gives optimal result; however, k can be considered as a threshold parameter that can be adjusted by the system’s users. These features in combination with other features could not be found in the prior art.
Related Prior Arts
Zhang (12,008,794) video stream transmitted from the surveillance camera (note fig. 2, block 202, and col. 11 lines 27-65, cites video surveillance camera).
Crisfalusi (12,406,503) person feature extraction and shoplifting behavior probability calculation (note col. 10 lines 20-44, person extraction tracking person and behavioral detection unit).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GREGORY M DESIRE whose telephone number is (571)272-7449. The examiner can normally be reached Monday-Friday 6:30am-3:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Henok Shiferaw can be reached at 571-272-4637. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
G.D.
August 14, 2026
/GREGORY M DESIRE/Primary Examiner, Art Unit 2676