DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-3 and 8-14 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by USPN 2022/0222940 to Rawat et al.
With regard to claim 1, Rawat discloses a computer-implemented method for training a video understanding model for classifying video data with improved accuracy, the method comprising:
obtaining, by a computing system comprising a plurality of computing devices, pretrained model data descriptive at least in part of a video understanding model, the video understanding model comprising at least one parameter of a video transformer encoder model (Fig. 2A, and paragraphs [0005]-[0006] and [0046]-[0047], Localization network shown I Fig. 2A is a pretrained model that is used to recognize video objects/actions. The localization network is used to identify and label the actions with bounding box annotations that are both spatial and temporal, and these are interpreted as parameters. The localization network uses a video encoder as shown.);
training, by the computing system, the pretrained model data based at least in part on a first dataset, the first dataset comprising image data, to determine first updated model data (paragraphs [0055]-[0058], Rawat discloses two sets of data sets are used for training the system and then another training set for evaluating the system. The first datasets are used to train the model of the classifier network); and
training, by the computing system, the first updated model data based at least in part on a second dataset, the second dataset comprising video data, to determine trained model data descriptive of a trained version of the video understanding model (paragraphs [0055]-[0058], The second datasets are considered to be the evaluating datasets that are used on the model once it is trained using the first datasets. All the data sets are annotated video data and are used to determine the accuracy or loss of the video action recognition for the video frame sets determined by the classification network).
With regard to claim 2, Rawat discloses the computer-implemented method of claim 1, wherein the video transformer encoder comprises a factorized encoder, the factorized encoder comprising a spatial transformer encoder and a temporal transformer encoder (paragraph [0044] "To perform the localization of each video clip, the method uses an encoder-decoder structure (such as a 3D convolution-based encoder, e.g. I3D [3]) which extracts class-agnostic action features (such as spatio-temporal features that are required for activity localization) and generates segmentation masks for each clip. The decoder uses these features to segment regions from the original input which contain activities." The encoder is sued to generate spatiotemporal annotations for the video segments).
With regard to claim 3, Rawat discloses the computer-implemented method of claim 2, wherein the spatial transformer encoder is configured to receive the plurality of video tokens and produce, in response to receipt of the plurality of video tokens, a plurality of temporal representations (paragraphs [0013]-[0015] and [0041], spatio-temporal action tubelets are identified using bounding boxes over a series of video frames to identify objects/actions. The video tokens are interpreted as the annotations in the form of bounding boxes indexed to specific image frames. The annotations and identified action frames are used to create tubelets that define the action in a spatiotemporal format); and
wherein the temporal transformer encoder is configured to receive the plurality of temporal representations and produce, in response to receipt of the plurality of temporal representations, a spatiotemporal representation of the video data, wherein the spatiotemporal representation is classified to produce the classification output (paragraphs [0013]-[0015] and [0041], spatio-temporal action tubelets are identified using bounding boxes over a series of video frames to identify objects/actions. The annotations and identified action frames are used to create tubelets that define the action in a spatiotemporal format).
With regard to claim 8, Rawat discloses the computer-implemented method of claim 1, wherein training, by the computing system, the first updated model data based at least in part on a second dataset comprises extracting a plurality of video tokens from the video data (paragraphs [0013]-[0015] and [0041], spatio-temporal action tubelets are identified using bounding boxes over a series of video frames to identify objects/actions. The annotations and identified action frames are used to create tubelets that define the action in a spatiotemporal format. The identified tubelets are used and compared with known annotated video frames in order to train the classification network as shown in Fig. 2A).
With regard to claim 9, Rawat discloses the computer-implemented method of claim 8, wherein extracting the plurality of video tokens comprises:
extracting, by the computing system, a plurality of video tubelets from the video data (paragraphs [0013]-[0015] and [0041], spatio-temporal action tubelets are identified using bounding boxes over a series of video frames to identify objects/actions.);
projecting, by the computing system, the plurality of video tubelets to a plurality of tensor representations of the plurality of video tubelets (paragraphs [0050]-[0051], Rawat discloses a tensor representation in the vector set as described); and
merging, by the computing system, the plurality of tensor representations along at least one dimension to produce the plurality of video tokens (paragraphs [0050]-[0052], As series of tubelets are recognized and categorized, they can be merged to represent as action defined by multiple sequential tubelets, thus merging the representation defining the recognized action).
With regard to claim 10, Rawat discloses the computer-implemented method of claim 9, wherein each of the plurality of video tubelets spans one of the plurality of video frames (paragraphs [0013]-[0015] and [0041], spatio-temporal action tubelets are identified using bounding boxes over a series of video frames to identify objects/actions. A tubelet spanning one frame would just be a single bounding box identification for a single frame).
With regard to claim 11, Rawat discloses the computer-implemented method of claim 9, wherein each of the plurality of video tubelets spans two or more of the plurality of video frames (paragraphs [0013]-[0015] and [0041], spatio-temporal action tubelets are identified using bounding boxes over a series of video frames to identify objects/actions).
With regard to claim 12, Rawat discloses the computer-implemented method of claim 9, wherein the plurality of video tokens are single-dimensional (paragraphs [0013]-[0015] and [0041]-[0043], spatio-temporal action tubelets are identified using bounding boxes over a series of video frames to identify objects/actions. A tubelet spanning one frame would just be a single bounding box identification for a single frame which is interpreted as single-dimensional. Video tokens identify the frames as well as the space within the frame that an object/action has been identified. The frame number would be an example of a single-dimensional token).
With regard to claim 13, Rawat discloses the computer-implemented method of claim 9, wherein the plurality of video tubelets are nonoverlapping (paragraph [0041], the plurality of tubelets can be merged. Since the tubelets occur in series in video frames, it is interpreted that they are not overlapping).
With regard to claim 14, Rawat discloses the computer-implemented method of claim 8, wherein positional embeddings are added to the plurality of video tokens and input to the video understanding model (paragraphs [0041], the tubelets are passed to a classification network where the tubelets are recognized and classified in terms of their respective depicted actions. The localization network determines which pixels and their positions correspond to the recognized actions).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of USPN 2022/0222940 to Rawat et al and publication titled “Social Fabric: Tunelet Compositions for Video Relation Detection” to Chen et al.
With regard to claim 15, Rawat discloses the computer-implemented method of claim 1, but does not explicitly disclose wherein the transformer encoder model comprises at least one normalization layer.
Chen discloses a tubelet composition determination similar to that of Rawat and further teaches a normalization layer (Section 3, third paragraph: “On top of the features, we apply layer normalization [3], followed by a linear layer to obtain embedded representation Ri ⊂ R ∈ RN× D. In this D-dimensional embedding space, we learn a set C ∈ RK×D consisting of K primitives. The idea behind our encoding is to describe a tubelet pair entirely as a weighted combination of these primitives.”). Therefore it would have been obvious to one of ordinary skill in the art before time of filing to use the normalization taught by Chen in order to represent the tubelet accurately with weighted factors.
Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of USPN 2022/0222940 to Rawat et al and publication titled “Spatio-temporal Tubelet Feature Aggregation and Object Linkin in Videos” to Cores at al.
With regard to claim 16, Rawat discloses the computer-implemented method of claim 1, but does not explicitly disclose wherein the transformer encoder model comprises at least one multi-layer perceptron layer.
Cores teaches a spatio-temporal tubelet generating system similar to that of Rawat and further teaches a multi-layer perceptron with two fully connected layers (Fig. 1). Therefore it would have been obvious to one of ordinary skill in the art before time of filing to use a perceptron with multiple layers as taught by Cores in order to aid in the classification output of the classifier network of Rawat.
Allowable Subject Matter
Claims 4-7 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claims 4 and 6 contain allowable subject matter. Claims 5 and 7 depend from claims 4 and 6 respectively.
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WESLEY J TUCKER whose telephone number is (571)272-7427. The examiner can normally be reached 9AM-5PM Monday-Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JOHN VILLECCO can be reached at 571-272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WESLEY J TUCKER/Primary Examiner, Art Unit 2661