DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statements (IDSs) submitted on September 1, 2023, February 27, 2025 and August 19, 2026 comply with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Citations which have not been considered, have not been considered because they do not comply with 37 CFR 1.98(b) which states “The date of publication supplied must include at least the month and year of publication, except that the year of publication (without the month) will be accepted if the applicant points out in the information disclosure statement that the year of publication is sufficiently earlier than the effective U.S. filing date and any foreign priority date so that the particular month of publication is not in issue”.
35 USC § 101 Statutory Analysis
The claims do not recite any of the judicial exceptions enumerated in the 2019 Revised Patent Subject Matter Eligibility Guidance. Further, the claims do not recite any method of organizing human activity, such as a fundamental economic concept or managing interactions between people. Finally, the claims do not recite a mathematical relationship, formula, or calculation. Thus, the claims are eligible because they do not recite a judicial exception.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. §102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-6, 9-11 and 15-19 are rejected under 35 U.S.C. §102(a)(1) as being anticipated by Kadkhodamohammadi et al. (WO 2023/021144 A1) (hereafter referred to as “Kadkhodamohammadi”).
With regard to claim 1, Kadkhodamohammadi describes inputting one or more frame-wise inputs associated with a sequence of video frames into a temporal convolutional network (see Figure 2, and Figure 3, element 201, and refer for example to paragraphs [0061] and [0075]); generating, using the temporal convolutional network, one or more frame-wise features based on the one or more frame-wise inputs (see Figure 2, and Figure 3, element 201, and refer for example to paragraphs [0061] and [0075]); generating a graph comprising one or more nodes and one or more edges based on the one or more frame-wise features, wherein a node corresponds to a video frame, and an edge connecting two nodes represents a connection between frame-wise features of two video frames (see Figure 3, element 206, and Figure 4, and refer for example to paragraphs [0077] through [0081]); inputting the graph into a graph neural network (see Figure 3, element 206, and Figure 4, and refer for example to paragraphs [0077] through [0081]); and generating, using the graph neural network, one or more predictions for the one or more nodes of the graph (refer for example to paragraphs [0066] and [0069] through [0071]).
As to claim 2, Kadkhodamohammadi describes wherein the one or more frame-wise inputs associated with a sequence of video frames comprises a first frame-wise input comprising a first vector of features extracted from a first frame in the sequence of video frames and a second frame-wise input comprising a second vector of features extracted from a second frame in the sequence of video frames (refer for example to paragraphs [0075], [0083] and [0089]).
In regard to claim 3, Kadkhodamohammadi describes wherein the one or more frame-wise inputs associated with a sequence of frames comprises a first frame-wise input comprising a first frame in the sequence of video frames and a second frame-wise input comprising a second frame in the sequence of video frames (refer for example to paragraphs [0067], [0075] and [0076]).
With regard to claim 4, Kadkhodamohammadi describes inputting the sequence of video frames into a three-dimensional convolutional neural network and generating, using the three-dimensional convolution neural network, the one or more frame-wise inputs (refer for example to paragraph [0075]).
As to claim 5, Kadkhodamohammadi describes wherein the one or more frame-wise features are generated by a second-to-the-last layer of the temporal convolutional network (as illustrated for example in Figure 3).
In regard to claim 6, Kadkhodamohammadi describes wherein generating the graph further comprises generating the graph further based on the one or more frame-wise inputs (as illustrated for example in Figure 3).
With regard to claim 9, Kadkhodamohammadi describes one or more processors and one or more storage devices storing a machine learning model having processing operations that are performed by the one or more processors (see Figure 1, element 102, and refer for example to paragraph [0051]), the machine learning model comprising a temporal convolutional network to receive one or more frame-wise inputs associated with a sequence of video frames, and output one or more frame-wise features (see Figure 2, and Figure 3, element 201, and refer for example to paragraphs [0061] and [0075]); a converter to generate a graph comprising nodes and one or more edges based on the one or more frame-wise features, wherein a node corresponds to a video frame, and an edge connecting two nodes represents a connection between frame-wise features of two video frames (see Figure 3, element 206, and Figure 4, and refer for example to paragraphs [0077] through [0081]); and a graph neural network to receive the graph, and output one or more predictions for the nodes (refer for example to paragraphs [0066] and [0069] through [0071]).
As to claim 10, Kadkhodamohammadi describes wherein the machine learning model further comprises a three-dimensional convolutional neural network to receive the sequence of video frames, and output the one or more frame-wise inputs (as illustrated for example in Figure 3).
In regard to claim 11, Kadkhodamohammadi describes wherein the temporal convolutional network comprises one or more convolution operators to receive one or more of the frame-wise inputs, a convolutional operator having a kernel size of 1x1 (as illustrated for example in Figure 3).
With regard to claim 15, Kadkhodamohammadi describes wherein the temporal convolutional network comprises a plurality of layers, and the one or more frame-wise features are generated by a second-to-the-last layer of the temporal convolutional network (as illustrated for example in Figure 3).
As to claim 16, Kadkhodamohammadi describes wherein the machine learning model further comprises a fusing block to receive and fuse the one or more frame-wise inputs and the one or more frame-wise features, and the converter is further to receive an output of the fusing block (as illustrated for example in Figure 3).
In regard to claim 17, Kadkhodamohammadi describes wherein the graph has one or more forward edges, one or more backward edges, and one or more un-directed edges (as illustrated for example in Figure 3).
With regard to claim 18, Kadkhodamohammadi describes the graph includes one or more temporal skip connections, wherein a temporal skip connection connects two nodes separated by at least one timestamp (as illustrated for example in Figure 3).
As to claim 19, Kadkhodamohammadi describes one or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors (see Figure 1, element 102, and refer for example to paragraph [0051]) to process, by a temporal convolutional network, one or more frame-wise inputs associated with a sequence of video frames (see Figure 2, and Figure 3, element 201, and refer for example to paragraphs [0061] and [0075]); extract, by the temporal convolutional network, one or more frame-wise features (see Figure 2, and Figure 3, element 201, and refer for example to paragraphs [0061] and [0075]); generate a graph comprising one or more nodes and one or more edges based on the one or more frame-wise features, wherein a node corresponds to a video frame, and an edge connecting two nodes represents a connection between frame-wise features of two video frames (see Figure 2, and Figure 3, element 201, and refer for example to paragraphs [0061] and [0075]); process, by a graph neural network, the graph (see Figure 2, and Figure 3, element 201, and refer for example to paragraphs [0061] and [0075]); and generate, by the graph neural network, one or more predictions for the nodes (refer for example to paragraphs [0066] and [0069] through [0071]).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. §103(a) which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 12 and 20 are rejected under 35 U.S.C. §103(a) as being unpatentable over Kadkhodamohammadi et al. (WO2023/021144 A1) (hereafter referred to as “Kadkhodamohammadi”) in view of Yi et al. (U.S. Patent No. 12,051,275 B2) (hereafter referred to as “Yi”), Chidlovskii et al. (U.S. Patent Application Publication No. US2021/0174513 A1) (hereafter referred to as “Chidlovskii”) or Kalchbrenner et al. (U.S. Patent Application Publication No. US 2021/0019555 A1) (hereafter referred to as “Kalchbrenner”).
The arguments advanced in section 6 above, as to the applicability of , are incorporated herein.
With regard to claims 12 and 20, although Kadkhodamohammadi does not expressly describe that the temporal convolutional network comprises a plurality of dilated convolution layers and applying one or more dilated convolutions with one or more dilation rates, such a technique is well known and widely utilized in the prior art.
Yi discloses a video processing system for recognizing features in video frame sequences (refer for example to the abstract) which utilizes a neural network model to process a video frame sequence (refer for example to column 10, lines 1-20) which provides for the temporal convolutional network comprises a plurality of dilated convolution layers and applying one or more dilated convolutions (refer for example to column 12, lines 17-21).
Chidlovskii discloses a video processing system for recognizing features in video frame sequences which utilizes a convolutional neural network model to process a video frame sequence (see Figure 2A and refer for example to paragraphs [0007], [0008] and [0009]) which provides for the temporal convolutional network comprises a plurality of dilated convolution layers and applying one or more dilated convolutions (refer for example to paragraphs [0047] and [0048]).
Kalchbrenner discloses a system for processing video frames sequences using neural networks (refer for example to the abstract) which utilizes a neural network model to process the video frame sequences to extract features (refer for example to paragraph [0053]) which provides for the temporal convolutional network comprises a plurality of dilated convolution layers and applying one or more dilated convolutions (refer for example to paragraphs [0063] through [0065]).
Given the teachings of the references and the same environment of operation, namely that of recognizing features in video frame sequences using neural network models, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the Kadkhodamohammadi system to provide a plurality of dilated convolution layers in the manner described by Yi, Chidlovskii or Kalchbrenner according to known methods to yield predictable results and would have been motivated to do so with a reasonable expectation of success in order to provide for increased processing efficiency and higher accuracy as suggested by Yi (refer for example to column 1, line 58 through column 2, line 54), Chidlovskii (refer for example to paragraph [0013]), or Kalchbrenner (refer for example to paragraph [0010]), which fails to patentably distinguish over the prior art absent some novel and unexpected result.
Allowable Subject Matter
Claims 7-8 and 13-14 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Relevant Prior Art
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Joseph, Dalli, O’Donncha, Ben and Shi all disclose systems similar to applicant’s claimed invention.
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jose L. Couso whose telephone number is (571) 272-7388. The examiner can normally be reached on Monday through Friday from 5:30am to 1:30pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matthew Bella, can be reached on 571-272-7778. The fax phone number for the organization where this application or proceeding is assigned is (571) 273-8300.
Information regarding the status of an application may be obtained from the Patent Center information webpage on the USPTO website. For more information about the Patent Center, see https://www.uspto.gov/patents/apply/patent-center. Should you have questions about access to the Patent Center, contact the Patent Electronic Business Center (EBC) at 571-272-4100 or via email at: ebc@uspto.gov .
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
/JOSE L COUSO/Primary Examiner, Art Unit 2667
September 15, 2026