Prosecution Insights
Last updated: October 01, 2026
Application No. 18/857,236

VIDEO PROCESSING SYSTEM, VIDEO PROCESSING APPARATUS, AND VIDEO PROCESSING METHOD

Non-Final OA §103§112
Filed
Oct 16, 2024
Priority
Sep 15, 2022 — nonprovisional of PCTJP2022034510
Examiner
BLACKSTEN, SYDNEY LYNN
Art Unit
Tech Center
Assignee
NEC Corporation
OA Round
1 (Non-Final)
100%
Grant Probability
Favorable
1-2
OA Rounds
6m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 100% — above average
100%
Career Allowance Rate
4 granted / 4 resolved
+40.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
22 currently pending
Career history
23
Total Applications
across all art units

Statute-Specific Performance

§101
8.9%
-31.1% vs TC avg
§103
67.7%
+27.7% vs TC avg
§102
4.0%
-36.0% vs TC avg
§112
12.9%
-27.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 4 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION The United States Patent & Trademark Office appreciates the application that is submitted by the inventor/assignee. The United States Patent & Trademark Office reviewed the following application and has made the following comments below. Priority This application claims benefit of foreign priority under 35 U.S.C. 119(a)-(d) of PCT/JP2022/034510, filed on 09/15/2022. Preliminary Amendment Applicant submitted a preliminary amendment on 10/16/2024. The Examiner acknowledges the amendment and has reviewed the claims accordingly. Information Disclosure Statement The information disclosure statement (IDS) submitted on 10/16/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Specification The disclosure is objected to because of the following informalities: In paragraph [0025], line 4 (page 8), the Examiner suggests replacing “termina” with “terminal.” Appropriate correction is required. Drawings The drawings are objected to because in Fig. 5 (reference character 110) denotes “video acquisition uni.” The Examiner recommends the text be changed to “video acquisition unit.” Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that use the word “means,” and are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “recognition means” in claims 6 and 12. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. For the sake of further prosecution, the Examiner will treat “the recognition means” recited in claims 6 and 12 as hardware or software configured to perform their respective functions/operations. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (B) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 6 and 12 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention. The Examiner strongly suggested that appropriate corrections be made to clarify the claim scope. With respect to Claim 6, the claim recites the following, each of which renders the claim indefinite: “ the recognition means ” on lines 1-2 (unclear antecedent basis); the claim contains no earlier recitation or limitation of a recognition means. It is unclear as to what element the limitation “the recognition means” is making reference to. With respect to Claim 12, the claim recites the following, each of which renders the claim indefinite: “ the recognition means ” on lines 1-2 (unclear antecedent basis); the claim contains no earlier recitation or limitation of a recognition means. It is unclear as to what element the limitation “the recognition means” is making reference to. With respect to Claims 6 and 12, the claims recite the following, each of which renders the claim indefinite: Claim limitation: “ the recognition means ” (Claims 6 and 12) invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. The recognition means/unit is mentioned in multiple paragraphs, for example, paragraphs [0007-8], [0014-15], [0035], [0039], [0047], [0049], [0058-59], [0062], [0065-67] and [0085], of the specification, but the specification only discloses the claimed function of the recognition means in the same language as the claim language. This disclosure is not sufficient because it fails to disclose the structure of the above element. Therefore, the claim is indefinite and is rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph. Applicant may: Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph; Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)). If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either: Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action: (a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103(a) are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-5, 7-11, and 13-17 are rejected under 35 U.S.C. 103(a) as being unpatentable over Angelova et al. (U.S. Patent No. 10,013,640, hereafter referred to as Angelova) in view of Xie et al. (U.S. Patent Pub No. 2021/0051368, hereafter referred to as Xie) in further view of Jenni et al. (U.S. Patent Pub. No. 2023/0276084, hereafter referred to as Jenni). Regarding Claim 1, Angelova teaches a video processing system (Col. 9, lines 40-52, Abstract, Fig. 5, Angelova teaches a computing device 500 used for identifying objects from a video.) PNG media_image1.png 399 721 media_image1.png Greyscale comprising: at least one memory storing instructions (Col. 9, lines 53-67, Col. 10, lines 8-11, Fig. 5, Angelova teaches the computing device 500 includes a memory 504 storing instructions.), and at least one processor configured to execute the instructions to (Col. 9, lines 63-67, Fig. 5, Angelova teaches a processor 502 can process instructions for execution within the computing device 500, including instructions stored in the memory 504.): acquire an input video (Col. 1, lines 32-38, Angelova teaches obtaining multiple frames from a video, where each frame of the multiple frames depicts an object to be recognized.); acquire first time difference information between frames of the input video (Col. 8, lines 4-14, Angelova teaches the system may select the multiple frames from the video based on a predetermined time interval. For example, the video may be 5 seconds long, and five video frames may be selected, where each frame is 1 second apart. The Examiner interprets a time interval, for example, 1 second between frames to be “time difference information” between frames. The Examiner interprets the time difference information is acquired since the system selects/chooses the frames based on the amount of time (time interval) between frames, and therefore, the system has acquired the time difference between frames.); and input the input video (Col. 4, lines 6-9, Angelova teaches the input 102 to the feature extraction layers includes a sequence of three video frames ft-1, ft, and ft+1, taken at time t-1, t, and t+1, respectively.) (Col. 7, lines 25-30, Angelova teaches the object recognition model may be trained.) (Col. 1, lines 32-38, Angelova teaches processing, using an object recognition model, the multiple frames from a video to generate data that represents a classification of the object to be recognized.). Angelova does not explicitly disclose input(ting) the first time difference information between the frames of the input video to a trained recognition model trained using a training video and second time difference information between frames of the training video. Xie is in the same field of art of using a long-time short-term memory (LSTM) network to process video data. Further, Xie teaches input(ting) the first time difference information between the frames of the input video to a trained recognition model (Paragraphs [0051-53], Fig. 7, Xie teaches prediction network 306 may learn a model that predicts the dropped-frame ratio based on an input of dropped-frame ratios and timestamps. In some embodiments a long-time short-term memory (time-LSTM) network is used. A time-LSTM may be a variant of an LSTM network that uses time in the prediction. The time-LSTM network uses inputs for the time to model time intervals. Prediction network 306 may include multiple units 702-1 to 702-3 that can each generate a prediction. Each unit 702-1 to 702-3 include time inputs 706-1 to 706-3 that receive the time difference associated. See Fig. 7 below.) PNG media_image2.png 590 820 media_image2.png Greyscale Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Angelova by inputting the time difference information between frames of the video into the time-LSTM that is taught by Xie, to make the invention that uses time difference information inputs to model time intervals; thus, one of ordinary skilled in the art would be motivated to combine the references since videos may be encoded in multiple representations that include different characteristics, such as different frame rates due to dropping frames, etc. The dropping of frames in a video may lead to discontinuity and decrease the quality of the video. For example, when a frame is dropped, the content may be choppy since some frames are not displayed in the video (Xie, Paragraphs [0001-2]). Therefore, by providing context information such as the amount of time between frames (timestamp difference) in the time series to the LSTM network, the LSTM may be able to recognize objects more accurately in videos with varying (non-constant/inconsistent) frame rates between frames. Angelova in view of Xie does not explicitly disclose (a trained recognition model) trained using a training video and second time difference information between frames of the training video. Jenni is in the same field of art of recognizing temporally varying changes such as frame rate changes in videos and determining where frame skippings occur in a video. Further, Jenni teaches (a trained recognition model) trained using a training video (Paragraphs [0057], [0061-62], Fig. 3, Jenni teaches training a machine-learning model with sample sequences of frame skippings from digital videos as training data. See sub-sampled digital video 302 in Fig. 3 below.) [AltContent: arrow] PNG media_image3.png 499 798 media_image3.png Greyscale and second time difference information between frames of the training video (Paragraphs [0057], [0059], [0064], Fig. 3, Jenni teaches the using the sequence of frame skippings 308 of the digital video as ground truth data for the predicted playback speeds. For example, a sequence of frame skippings of 2, 1, 1, and 2 would translate to ground truth playback speed classifications of 1, 0, 0, and 1 in frame order. The Examiner interprets the ground truth playback speed classifications 308 indicate when a frame has been skipped in sub-sampled digital video 302 (i.e., “1” corresponds to a frame skipped in the sub-sampled video, i.e., there is a time difference between (two) frames in the sub-sampled video. “0” corresponds to no frames skipped in the sub-sampled digital video 302, i.e., there is no time difference between the (two) frames in the sub-sampled video.). The Examiner interprets the sequence of frame skippings shown along with sub-sampled digital video (302) is consistent with Applicant’s specification, which states “the time difference information ΔT is 1 when no frame is skipped between the corresponding predetermined frame and the previous frame. The time difference information ΔT is 1+n when n frames are skipped between the corresponding predetermined frame and the previous frame (Applicant’s specification, page 13, paragraph [0048], lines 28-31).” ). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Angelova in view of Xie by training the model using a training video with predetermined frame skips and frame skip information that is taught by Jenni, to make the invention that trains the machine-learning model to recognize and localize temporally varying changes (time changes between frames) in a digital video; thus, one of ordinary skilled in the art would be motivated to combine the references since by iteratively generating predicted playback speeds for the sub-sampled digital video 302 (training video) using the model to determine a classification loss (between the output and ground truth) iteratively trains the model (Jenni, Paragraphs [0060-61]). After training, the model is able to determine per-frame slowness predictions (frame rate changes) with improved accuracy compared to conventional methods (Jenni, Paragraph [0028]). Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. In regards to Claim 2, Angelova in view of Xie in further view of Jenni discloses the video processing system according to claim 1, wherein the trained recognition model (Col. 7, lines 25-31, Angelova teaches the object recognition model may be trained.) is a model including a plurality of cells (Col. 5, lines 45-61, Angelova teaches the LSTM layer 300 includes one or more LSTM memory blocks, including a memory block 320. The LSTM memory block 320 includes an LSTM memory cell 322 that receives an input xt and generates an output mt from the input and from a previous recurrent projected output rt-1.) PNG media_image4.png 385 617 media_image4.png Greyscale of a recurrent neural network (RNN) (Col. 8, lines 27-32, Fig. 1, Angelova teaches the object recognition model is a recurrent neural network that includes a long short-term memory (LSTM) layer. Fig. 1 below is a block diagram of an object recognition model.) PNG media_image5.png 674 578 media_image5.png Greyscale that inputs time-series frames included in the input video (Col.1, lines 54-60, Angelova teaches the multiple frames may be arranged in an order according to their time of occurrence in the video. To process the multiple frames, each frame of the multiple frames may be processed using the LSTM layer in the order according to their time of occurrence in the video.), and the plurality of cells input a parameter corresponding to first time difference information between the frames of the input video (Col. 5, lines 45-61, Col. 6, lines 26-33, Fig. 3A, Angelova teaches the LSTM memory cell 322 generates an output mt from the input xt and the previous recurrent projected output rt-1. For example, the input xt may be the feature output for a video frame at time step t in a video frame sequence. The previous recurrent projected output rt-1 is the projected output generated by the recurrent projection layer from an output rt-1 generated by the cell at the preceding time step in the video frame sequence. Once the output mt has been computed, the recurrent projection layer may compute a recurrent projected output rt for the current time step using output mt. The recurrent projected output rt can then be fed back to memory block for use in computing output mt+1 at the next time step in the video frame sequence.). In regards to Claim 3, Angelova in view of Xie in further view of Jenni discloses the video processing system according to claim 1, wherein the trained recognition model (Col. 7, lines 25-31, Angelova teaches the object recognition model may be trained.) includes a plurality of cells (Col. 5, lines 45-61, Angelova teaches the LSTM layer 300 includes one or more LSTM memory blocks, including a memory block 320. The LSTM memory block 320 includes an LSTM memory cell 322 that receives an input xt and generates an output mt from the input and from a previous recurrent projected output rt-1.) of a recurrent neural network (Col. 8, lines 27-32, Fig. 1, Angelova teaches the object recognition model is a recurrent neural network that includes a long short-term memory (LSTM) layer.) that input time-series frames included in the input video (Col.1, lines 54-60, Angelova teaches the multiple frames may be arranged in an order according to their time of occurrence in the video. To process the multiple frames, each frame of the multiple frames may be processed using the LSTM layer in the order according to their time of occurrence in the video.), and input and output state vectors chronologically (Col. 5, lines 45-61, Angelova teaches the LSTM memory block 320 includes an LSTM memory cell 322 that receives an input xt and generates an output mt from the input and from previous recurrent projected output rt-1. For example, the input xt may be the feature output for a video frame at time step t in a video frame sequence. The previous recurrent projected output rt-1 is the projected output generated by the recurrent projection layer from an output rt-1 generated by the cell at the preceding time step in the video frame sequence. In other words, the previous recurrent projected output rt-1 is fed back to the cell. The Examiner interprets the “recurrent projected output” refers to the projection of the hidden state, which is updated.), PNG media_image6.png 761 659 media_image6.png Greyscale and a state predictor (Col. 8, lines 58-64, Angelova teaches an LSTM layer, which processes the feature data to generate an LSTM output and to update an internal state of the LSTM layer. The Examiner interprets the LSTM layer to be a “state predictor” since it updates the internal state of the LSTM layer and the claim is silent to the meaning of “state predictor.”) that predicts the state vectors based on the first time difference information between the frames of the input video (Col. 8, lines 65-67 through Col. 9, lines 1-8, Col. 3, lines 57-61, Angelova teaches the system may process each frame of the multiple frames using the LSTM layer in the order according to their time occurrence in the video to generate the LSTM output and to update the internal state of the LSTM layer. For example, the forward LSTM layer is configured to process the feature output in forward time steps to generate a forward LSTM output. The Examiner interprets the state vectors are predicted based on time difference information since the frames may be selected based on a predetermined time interval, such as 1 second, and therefore, the system knows the frames are separated by a constant time interval of 1 second.) is inserted between predetermined cells (Col. 5, lines 45-61, Fig. 1, Angelova teaches the previous recurrent output is the projected output generated by the recurrent projection layer from an output rt-1 generated by the cell at the preceding time step in the video frame sequence. The previous recurrent projected output rt-1 is fed back to the cell.). In regards to Claim 4, Angelova in view of Xie in further view of Jenni discloses the video processing system according to claim 3, wherein the trained recognition model into which the state predictor is inserted (Col. 1, lines 51-60, Angelova teaches the feature data may be processed using an LSTM layer to generate an LSTM output and to update an internal state of the LSTM layer. The Examiner interprets the LSTM layer to be the “state predictor” since it updates the internal state of the LSTM layer. In addition, the Examiner interprets the state predictor is “inserted” in the recognition model since the LSTM layer is included in the recognition model.) is trained (Paragraph [0057], Fig. 3, Jenni teaches training a machine-learning model.) using time- series frames (Col. 1, lines 54-55, Angelova teaches the multiple frames may be arranged in an order according to their time of occurrence in the video.) in which frame skipping incurs in a predetermined pattern included in the training video (Paragraphs [0057], [0061-62], [0066], Fig. 3, Jenni teaches training the machine-learning model utilizing a self-supervised learning approach with sample sequences of frame skipping from digital videos as training data.), PNG media_image7.png 138 294 media_image7.png Greyscale the second time difference information between the frames of the training video (Paragraphs [0058-59], Fig. 3, Jenni teaches generating a subsampled video by sampling a sequence of frame skippings for a training digital video. The system utilizes the playback speed prediction machine-learning model to generate a playback speed prediction vector that classifies a playback speed per frame (e.g., the first frame transition represented in the first row of the vector is classified to a 2x playback speed or a single frame skip, the second frame transition represented in the second row of the vector is classified to a 1x playback speed or no frame skip, and so forth. The temporally varying video re-timing system 106 compares the predicted playback speed classifications from the varying predicted playback speeds 306 to ground truth playback speed classifications 308 that correspond to the frame skips from the sub-sampled digital video 302.), PNG media_image8.png 380 229 media_image8.png Greyscale and correct data (Paragraphs [0059], [0064], Fig. 3, Jenni teaches ground truth playback speed classifications 308 that correspond to the frame skips from the sub-sampled digital video. The Examiner interprets ground truth data to be “correct data.”). In regards to Claim 5, Angelova in view of Xie in further view of Jenni discloses the video processing system according to claim 3, wherein the plurality of cells (Col. 5, lines 45-61, Angelova teaches the LSTM layer 300 includes one or more LSTM memory blocks, including a memory block 320. The LSTM memory block 320 includes an LSTM memory cell 322 that receives an input xt and generates an output mt from the input and from a previous recurrent projected output rt-1.) of the trained recognition model (Col. 7, lines 25-29, Angelova teaches the object recognition model may be trained.) using time-series frames (Col. 1, lines 54-55, Angelova teaches the multiple frames may be arranged in an order according to their time of occurrence in the video.) which are included in the training video (Paragraph [0058], Fig. 3, Jenni teaches a sub-sampled digital video 302 created by sampling a sequence of frame skipping’s for a training digital video.) and in which no frame skipping incurs (Col. 4, lines 7-9, Angelova teaches the input 102 (video) includes a sequence of three video frames ft-1, ft, and ft+1, taken at time t-1, t, and t+1, respectively. The Examiner interprets since the video frames are each taken within one second/time interval of each other, i.e. no “frame skipping” has occurred.) and correct data (Paragraphs [0059], [0064], Fig. 3, Jenni teaches ground truth playback speed classifications.), and the state predictor (Col. 4, lines 53-67 through Col. 5, lines 1-17, Fig. 1, Angelova teaches forward LSTM layer 106 processes the internal LSTM state from the preceding state and the feature output to generate a forward LSTM output and update the internal state of the forward LSTM layer 106.) inserted into the trained recognition model (Col. 3, lines 41-52, Fig. 1, Angelova teaches an object recognition model 100 having a convolutional LSTM layer. The Examiner interprets the “LSTM layer” to be a “state predictor” since the LSTM layer updates the internal state of the forward LSTM layer.) is trained (Col. 7, lines 25-30, Angelova teaches the object recognition model may be trained such that the classification layers store received LSTM outputs until the forward LSTM layer has processed all of the frames in the sequence and has generated all the LSTM outputs, before generating the output representing a set of scores.) using a state vector output at time t (where t is natural number) (Col. 6, lines 26-29, Angelova teaches computing the output mt.) and a state vector output at time t+N (where N is a natural number) (Col. 6, lines 50-53, Angelova teaches computing the output mt+1 at the next time step in the video frame sequence.) by the plurality of cells (Col. 5, lines 45-61, Angelova teaches the LSTM memory block which includes an LSTM memory cell receives an input xt and generates output mt from the input and from a previous recurrent projected output rt-1.) when time-series frames (Col. 4, lines 7-9, Angelova teaches the input includes a sequence of three video frames ft-1, ft, and ft+1, taken at time t-1, t, and t+1, respectively.) which are included in the training video (Paragraph [0058], Fig. 3, Jenni teaches a sub-sampled digital video 302.) and in which no frame skipping incurs (Col. 4, lines 7-9, Angelova teaches the input 102 (video) includes a sequence of three video frames ft-1, ft, and ft+1, taken at time t-1, t, and t+1, respectively. The Examiner interprets since the video frames are each taken within one second of each other, no “frame skipping” has occurred.) are input to the plurality of trained cells (Col. 9, lines 2-6, Col. 5, lines 47-50, Angelova teaches the system 100 may process each frame of the multiple frames using the LSTM layer in the order according to their time of occurrence in the video to generate LSTM output and to update the internal state of the LSTM layer. The LSTM layer 300 includes one or more memory blocks, including an LSTM memory block 320. The LSTM memory block 320 includes an LSTM memory cell 322. The Examiner interprets since there are one or more memory blocks, each of which have a memory cell, there are a “plurality of cells”.). In regards to Claim 7, Angelova discloses a video processing apparatus (Abstract, Angelova teaches an apparatus for identifying an object from a video.) comprising: at least one memory storing instructions (Col. 9, lines 53-67, Col. 10, lines 8-11, Fig. 5, Angelova teaches the computing device 500 includes a memory 504 storing instructions.), and at least one processor configured to execute the instructions to (Col. 9, lines 63-67, Angelova teaches a processor 502 can process instructions for execution within the computing device 500, including instructions stored in the memory 504.): acquire an input video (Col. 1, lines 32-38, Angelova teaches obtaining multiple frames from a video.); acquire first time difference information between frames of the input video (Col. 8, lines 4-14, Angelova teaches the system may select the multiple frames from the video based on a predetermined time interval. For example, the video may be 5 seconds long, and five video frames may be selected, where each frame is 1 second apart.); and input the input video (Col. 4, lines 6-9, Angelova teaches the input 102 to the feature extraction layers includes a sequence of three video frames ft-1, ft, and ft+1, taken at time t-1, t, and t+1, respectively.) (Col. 1, lines 32-38, Col. 8, lines 28-29 and lines 40-48, Angelova teaches processing, using an object recognition model the multiple frames from a video to generate data that represents a classification of the object to be recognized.). Angelova does not explicitly disclose input(ing) the first time difference information between the frames of the input video to a trained recognition model trained using a training video and second time difference information between frames of the training video. Xie is in the same field of art of using a long-time short-term memory (LSTM) network to process videos. Further, Xie teaches input(ting) the first time difference information between the frames of the input video to a trained recognition model (Paragraphs [0051-53], Fig. 7, Xie teaches prediction network 306 may learn a model that predicts the dropped-frame ratio based on an input of dropped-frame ratios and timestamps. In some embodiments a long-time short-term memory (time-LSTM) network is used. A time-LSTM may be a variant of an LSTM network that uses time in the prediction. The time-LSTM network uses inputs for the time to model time intervals. Prediction network 306 may include multiple units 702-1 to 702-3 that can each generate a prediction. Each unit 702-1 to 702-3 include time inputs 706-1 to 706-3 that receive the time difference associated.) Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Angelova by inputting the time difference information between frames of the video into the time-LSTM that is taught by Xie, to make the invention that uses time difference information inputs to model time intervals; thus, one of ordinary skilled in the art would be motivated to combine the references since videos may be encoded in multiple representations that include different characteristics, such as different frame rates due to dropping frames, etc. The dropping of frames in a video may lead to discontinuity and decrease the quality of the video. For example, when a frame is dropped, the content may be choppy since some frames are not displayed in the video (Xie, Paragraphs [0001-2]). Therefore, by providing context information such as the amount of time between frames (timestamp difference) in the time series to the LSTM network, the LSTM may be able to recognize objects more accurately in videos with varying (non-constant) frame rates between frames. Angelova in view of Xie does not explicitly disclose (a trained recognition model) trained using a training video and second time difference information between frames of the training video. Jenni is in the same field of art of recognizing temporally varying changes such as frame rate changes in videos and determining where frame skippings occur in a video. Further, Jenni teaches (a trained recognition model) trained using a training video (Paragraphs [0057], [0061-62], Fig. 3, Jenni teaches training a machine-learning model with sample sequences of frame skippings from digital videos as training data. See sub-sampled digital video 302 in Fig. 3 below.) and second time difference information between frames of the training video (Paragraphs [0057], [0059], [0064], Fig. 3, Jenni teaches the using the sequence of frame skippings 308 of the digital video as ground truth data for the predicted playback speeds. For example, a sequence of frame skippings of 2, 1, 1, and 2 would translate to ground truth playback speed classifications of 1, 0, 0, and 1 in frame order. The Examiner interprets the ground truth playback speed classifications 308 indicate when a frame has been skipped in sub-sampled digital video 302 (i.e., “1” corresponds to a frame skipped in the sub-sampled video, i.e., there is a time difference between (two) frames in the sub-sampled video. “0” corresponds to no frames skipped in the sub-sampled digital video 302, i.e., there is no time difference between the (two) frames in the sub-sampled video.). The Examiner interprets the sequence of frame skippings shown along with sub-sampled digital video (302) is consistent with Applicant’s specification, which states “the time difference information ΔT is 1 when no frame is skipped between the corresponding predetermined frame and the previous frame. The time difference information ΔT is 1+n when n frames are skipped between the corresponding predetermined frame and the previous frame (Applicant’s specification, page 13, paragraph [0048], lines 28-31).” ). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Angelova in view of Xie by training the model using a training video with predetermined frame skips and frame skip information that is taught by Jenni, to make the invention that trains the machine-learning model to recognize and localize temporally varying changes (time changes between frames) in a digital video; thus, one of ordinary skilled in the art would be motivated to combine the references since by iteratively generating predicted playback speeds for the sub-sampled digital video 302 (training video) using the model to determine a classification loss (between the output and ground truth) iteratively trains the model (Jenni, Paragraphs [0060-61]). After training, the model is able to determine per-frame slowness predictions with improved accuracy compared to conventional methods (Jenni, Paragraph [0028]). Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. In regards to Claim 8, Angelova in view of Xie in further view of Jenni discloses the video processing apparatus according to claim 7, wherein the trained recognition model is a model including a plurality of cells (Col. 5, lines 45-61, Angelova teaches the LSTM layer 300 includes one or more LSTM memory blocks, including a memory block 320. The LSTM memory block 320 includes an LSTM memory cell 322 that receives an input xt and generates an output mt from the input and from a previous recurrent projected output rt-1.) of a recurrent neural network (RNN) (Col. 8, lines 27-32, Fig. 1, Angelova teaches the object recognition model is a recurrent neural network that includes a long short-term memory (LSTM) layer.) that inputs time-series frames included in the input video (Col.1, lines 54-60, Angelova teaches the multiple frames may be arranged in an order according to their time of occurrence in the video. To process the multiple frames, each frame of the multiple frames may be processed using the LSTM layer in the order according to their time of occurrence in the video.), and the plurality of cells input a parameter corresponding to first time difference information between the frames of the input video (Col. 5, lines 45-61, Col. 6, lines 26-33, Fig. 3A, Angelova teaches the LSTM memory cell 322 generates an output mt from the input xt and the previous recurrent projected output rt-1. For example, the input xt may be the feature output for a video frame at time step t in a video frame sequence. The previous recurrent projected output rt-1 is the projected output generated by the recurrent projection layer from an output rt-1 generated by the cell at the preceding time step in the video frame sequence. Once the output mt has been computed, the recurrent projection layer may compute a recurrent projected output rt for the current time step using output mt. The recurrent projected output rt can then be fed back to memory block for use in computing output mt+1 at the next time step in the video frame sequence.). In regards to Claim 9, Angelova in view of Xie in further view of Jenni discloses the video processing apparatus according to claim 7, wherein the trained recognition model includes a plurality of cells (Col. 5, lines 45-61, Angelova teaches the LSTM layer 300 includes one or more LSTM memory blocks, including a memory block 320. The LSTM memory block 320 includes an LSTM memory cell 322 that receives an input xt and generates an output mt from the input and from a previous recurrent projected output rt-1.) of a recurrent neural network (Col. 8, lines 27-32, Fig. 1, Angelova teaches the object recognition model is a recurrent neural network that includes a long short-term memory (LSTM) layer.) that input time-series frames included in the input video (Col.1, lines 54-60, Angelova teaches the multiple frames may be arranged in an order according to their time of occurrence in the video. To process the multiple frames, each frame of the multiple frames may be processed using the LSTM layer in the order according to their time of occurrence in the video.), and input and output state vectors chronologically (Col. 5, lines 45-61, Angelova teaches the LSTM memory block 320 includes an LSTM memory cell 322 that receives an input xt and generates an output mt from the input and from previous recurrent projected output rt-1. For example, the input xt may be the feature output for a video frame at time step t in a video frame sequence. The previous recurrent projected output rt-1 is the projected output generated by the recurrent projection layer from an output rt-1 generated by the cell at the preceding time step in the video frame sequence. In other words, the previous recurrent projected output rt-1 is fed back to the cell. The Examiner interprets the “recurrent projected output” refers to the projection of the hidden state (mt), which is updated.), and a state predictor (Col. 8, lines 58-64, Angelova teaches an LSTM layer, which processes the feature data to generate an LSTM output and to update an internal state of the LSTM layer.) that predicts the state vectors based on the first time difference information between the frames of the input video (Col. 8, lines 65-67 through Col. 9, lines 1-8, Col. 3, lines 57-61, Angelova teaches the system may process each frame of the multiple frames using the LSTM layer in the order according to their time occurrence in the video to generate the LSTM output and to update the internal state of the LSTM layer. For example, the forward LSTM layer is configured to process the feature output in forward time steps to generate a forward LSTM output. The Examiner interprets the state vectors are predicted based on time difference information since the frames may be selected based on a predetermined time interval, such as 1 second, and therefore, the system knows the frames are separated by a constant time interval of 1 second.) is inserted between predetermined cells (Col. 5, lines 45-61, Fig. 1, Angelova teaches the previous recurrent output is the projected output generated by the recurrent projection layer from an output rt-1 generated by the cell at the preceding time step in the video frame sequence. The previous recurrent projected output rt-1 is fed back to the cell.). In regards to Claim 10, Angelova in view of Xie in further view of Jenni discloses the video processing apparatus according to claim 9, wherein the trained recognition model into which the state predictor is inserted (Col. 1, lines 51-60, Angelova teaches the feature data may be processed using an LSTM layer to generate an LSTM output and to update an internal state of the LSTM layer.) is trained (Paragraph [0057], Fig. 3, Jenni teaches training a machine-learning model.) using time- series frames (Col. 1, lines 54-55, Angelova teaches the multiple frames may be arranged in an order according to their time of occurrence in the video.) in which frame skipping incurs in a predetermined pattern included in the training video (Paragraphs [0057], [0061-62], [0066], Jenni teaches training the machine-learning model utilizing a self-supervised learning approach with sample sequences of frame skipping from digital videos as training data.), the second time difference information between the frames of the training video, and correct data (Paragraph [0058], Fig. 3, Jenni teaches generating a subsampled video by sampling a sequence of frame skipping’s for a training digital video. The system utilizes the playback speed prediction machine-learning model to generate a playback speed prediction vector that classifies a playback speed per frame (e.g., the first frame transition represented in the first row of the vector is classified to a 2x playback speed or a single frame skip, the second frame transition represented in the second row of the vector is classified to a 1x playback speed or no frame skip, and so forth.) and correct data (Paragraphs [0059], [0064], Fig. 3, Jenni teaches ground truth playback speed classifications that correspond to the frame skips from the sub-sampled digital video.). In regards to Claim 11, Angelova in view of Xie in further view of Jenni discloses the video processing apparatus according to claim 9, wherein the plurality of cells of the trained recognition model (Col. 5, lines 45-61, Angelova teaches the LSTM layer 300 includes one or more LSTM memory blocks, including a memory block 320. The LSTM memory block 320 includes an LSTM memory cell 322 that receives an input xt and generates an output mt from the input and from a previous recurrent projected output rt-1.) are trained (Paragraphs [0051-53], Xie teaches prediction network 306 may learn a model that predicts the dropped-frame ratio based on an input of dropped-frame ratios and timestamps. The prediction network 306 may include multiple units 702-1 to 702-3 that can each generate a prediction. The Examiner interprets “units” and “cells” to be synonymous in the context of LSTM models.) using time-series frames (Col. 1, lines 54-55, Angelova teaches the multiple frames may be arranged in an order according to their time of occurrence in the video.) which are included in the training video (Paragraph [0058], Fig. 3, Jenni teaches a sub-sampled digital video 302 created by sampling a sequence of frame skipping’s for a training digital video.) and in which no frame skipping incurs (Col. 4, lines 7-9, Angelova teaches the input 102 (video) includes a sequence of three video frames ft-1, ft, and ft+1, taken at time t-1, t, and t+1, respectively. The Examiner interprets since the video frames are each taken within one second of each other, no “frame skipping” has occurred.) and correct data (Paragraphs [0059], [0064], Fig. 3, Jenni teaches ground truth playback speed classifications.), and the state predictor (Col. 4, lines 53-67 through Col. 5, lines 1-17, Fig. 1, Angelova teaches forward LSTM layer 106 processes the internal LSTM state from the preceding state and the feature output to generate a forward LSTM output and update the internal state of the forward LSTM layer 106.) inserted into the trained recognition model (Col. 3, lines 41-52, Fig. 1, Angelova teaches an object recognition model 100 having a convolutional LSTM layer. The Examiner interprets the “LSTM layer” to be a “state predictor” since the LSTM layer updates the internal state of the forward LSTM layer.) is trained (Col. 7, lines 25-30, Angelova teaches the object recognition model may be trained such that the classification layers store received LSTM outputs until the forward LSTM layer has processed all of the frames in the sequence and has generated all the LSTM outputs, before generating the output representing a set of scores.) using a state vector output at time t (where t is natural number) (Col. 6, lines 26-29, Angelova teaches computing the output mt.) and a state vector output at time t+N (where N is a natural number) (Col. 6, lines 50-53, Angelova teaches computing the output mt+1 at the next time step in the video frame sequence.) by the plurality of cells (Col. 5, lines 45-61, Angelova teaches the LSTM memory block which includes an LSTM memory cell receives an input xt and generates output mt from the input and from a previous recurrent projected output rt-1.) when time-series frames (Col. 4, lines 7-9, Angelova teaches the input includes a sequence of three video frames ft-1, ft, and ft+1, taken at time t-1, t, and t+1, respectively.) which are included in the training video (Paragraph [0058], Fig. 3, Jenni teaches a sub-sampled digital video 302.) and in which no frame skipping incurs (Col. 4, lines 7-9, Angelova teaches the input includes a sequence of three video frames ft-1, ft, and ft+1, taken at time t-1, t, and t+1, respectively. The Examiner interprets no frame skipping occurs since the frames are each taken 1 second apart.) are input to the plurality (Col. 9, lines 2-6, Col. 5, lines 47-50, Angelova teaches the system 100 may process each frame of the multiple frames using the LSTM layer in the order according to their time of occurrence in the video to generate LSTM output and to update the internal state of the LSTM layer. The LSTM layer 300 includes one or more memory blocks, including an LSTM memory block 320. The LSTM memory block 320 includes an LSTM memory cell 322. The Examiner interprets since there are one or more memory blocks, each of which have a memory cell, there are a “plurality of cells”.) of trained cells (Col. 7, lines 25-29, Angelova teaches the object recognition model may be trained. The Examiner interprets the memory cells are trained since the overall model is trained.). In regards to Claim 13, Angelova teaches a video processing method (Abstract, Angelova teaches a method for identifying an object from a video.) comprising: by a computer (Col. 2, lines 46-56, Angelova teaches the computer program configured to perform the actions and methods, is encoded on computer storage devices. One or more computer programs can be so configured by virtue of having instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.), acquiring an input video (Col. 1, lines 32-38, Angelova teaches obtaining multiple frames from a video, where each frame of the multiple frames depicts an object to be recognized.); acquiring first time difference information between frames of the input video (Col. 8, lines 4-14, Angelova teaches the system may select the multiple frames from the video based on a predetermined time interval. For example, the video may be 5 seconds long, and five video frames may be selected, where each frame is 1 second apart. The Examiner interprets a time interval, for example, 1 second between frames to be “time difference information” between frames. The Examiner interprets the time difference information is acquired since the system selects/chooses the frames based on the amount of time (time interval) between frames, and therefore, the system has acquired the time difference between frames.); and inputting the input video (Col. 4, lines 6-9, Angelova teaches the input 102 to the feature extraction layers includes a sequence of three video frames ft-1, ft, and ft+1, taken at time t-1, t, and t+1, respectively.) (Col. 7, lines 25-30, Angelova teaches the object recognition model may be trained.) (Col. 1, lines 32-38, Angelova teaches processing, using an object recognition model, the multiple frames from a video to generate data that represents a classification of the object to be recognized.). Angelova does not explicitly disclose (inputting) the first time difference information between the frames of the input video to a trained recognition model trained using a training video and second time difference information between frames of the training video Xie is in the same field of art of using a long-time short-term memory (LSTM) network to process videos. Further, Xie teaches (inputting) the first time difference information between the frames of the input video to a trained recognition model (Paragraphs [0051-53], Fig. 7, Xie teaches prediction network 306 may learn a model that predicts the dropped-frame ratio based on an input of dropped-frame ratios and timestamps. In some embodiments a long-time short-term memory (time-LSTM) network is used. A time-LSTM may be a variant of an LSTM network that uses time in the prediction. The time-LSTM network uses inputs for the time to model time intervals. Prediction network 306 may include multiple units 702-1 to 702-3 that can each generate a prediction. Each unit 702-1 to 702-3 include time inputs 706-1 to 706-3 that receive the time difference associated.) Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Angelova by inputting the time difference information between frames of the video into the time-LSTM that is taught by Xie, to make the invention that uses time difference information inputs to model time intervals; thus, one of ordinary skilled in the art would be motivated to combine the references since videos may be encoded in multiple representations that include different characteristics, such as different frame rates due to dropping frames, etc. The dropping of frames in a video may lead to discontinuity and decrease the quality of the video. For example, when a frame is dropped, the content may be choppy since some frames are not displayed in the video (Xie, Paragraphs [0001-2]). Therefore, by providing context information such as the amount of time between frames (timestamp difference) in the time series to the LSTM network, the LSTM may be able to recognize objects more accurately in videos with varying (non-constant/inconsistent) frame rates between frames. Angelova in view of Xie does not explicitly disclose (a trained recognition model) trained using a training video and second time difference information between frames of the training video. Jenni is in the same field of art of recognizing temporally varying changes such as frame rate changes in videos and determining where frame skippings occur in a video. Further, Jenni teaches (a trained recognition model) trained using a training video (Paragraphs [0057], [0061-62], Fig. 3, Jenni teaches training a machine-learning model with sample sequences of frame skippings from digital videos as training data. See sub-sampled digital video 302 in Fig. 3 below.) and second time difference information between frames of the training video (Paragraphs [0057], [0059], [0064], Fig. 3, Jenni teaches the using the sequence of frame skippings 308 of the digital video as ground truth data for the predicted playback speeds. For example, a sequence of frame skippings of 2, 1, 1, and 2 would translate to ground truth playback speed classifications of 1, 0, 0, and 1 in frame order. The Examiner interprets the ground truth playback speed classifications 308 indicate when a frame has been skipped in sub-sampled digital video 302 (i.e., “1” corresponds to a frame skipped in the sub-sampled video, i.e., there is a time difference between (two) frames in the sub-sampled video. “0” corresponds to no frames skipped in the sub-sampled digital video 302, i.e., there is no time difference between the (two) frames in the sub-sampled video.). The Examiner interprets the sequence of frame skippings shown along with sub-sampled digital video (302) is consistent with Applicant’s specification, which states “the time difference information ΔT is 1 when no frame is skipped between the corresponding predetermined frame and the previous frame. The time difference information ΔT is 1+n when n frames are skipped between the corresponding predetermined frame and the previous frame (Applicant’s specification, page 13, paragraph [0048], lines 28-31).” ). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Angelova in view of Xie by training the model using a training video with predetermined frame skips and frame skip information that is taught by Jenni, to make the invention that trains the machine-learning model to recognize and localize temporally varying changes (time changes between frames) in a digital video; thus, one of ordinary skilled in the art would be motivated to combine the references since by iteratively generating predicted playback speeds for the sub-sampled digital video 302 (training video) using the model to determine a classification loss (between the output and ground truth) iteratively trains the model (Jenni, Paragraphs [0060-61]). After training, the model is able to determine per-frame slowness predictions with improved accuracy compared to conventional methods (Jenni, Paragraph [0028]). Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. In regards to Claim 14, Angelova in view of Xie in further view of Jenni discloses the video processing method according to claim 13, wherein the trained recognition model (Col. 7, lines 25-31, Angelova teaches the object recognition model may be trained.) is a model including a plurality of cells (Col. 5, lines 45-61, Angelova teaches the LSTM layer 300 includes one or more LSTM memory blocks, including a memory block 320. The LSTM memory block 320 includes an LSTM memory cell 322 that receives an input xt and generates an output mt from the input and from a previous recurrent projected output rt-1.) of a recurrent neural network (RNN) (Col. 8, lines 27-32, Fig. 1, Angelova teaches the object recognition model is a recurrent neural network that includes a long short-term memory (LSTM) layer. Fig. 1 below is a block diagram of an object recognition model.) that inputs time-series frames included in the input video (Col.1, lines 54-60, Angelova teaches the multiple frames may be arranged in an order according to their time of occurrence in the video. To process the multiple frames, each frame of the multiple frames may be processed using the LSTM layer in the order according to their time of occurrence in the video.), and the plurality of cells input a parameter corresponding to first time difference information between the frames of the input video (Col. 5, lines 45-61, Col. 6, lines 26-33, Fig. 3A, Angelova teaches the LSTM memory cell 322 generates an output mt from the input xt and the previous recurrent projected output rt-1. For example, the input xt may be the feature output for a video frame at time step t in a video frame sequence. The previous recurrent projected output rt-1 is the projected output generated by the recurrent projection layer from an output rt-1 generated by the cell at the preceding time step in the video frame sequence. Once the output mt has been computed, the recurrent projection layer may compute a recurrent projected output rt for the current time step using output mt. The recurrent projected output rt can then be fed back to memory block for use in computing output mt+1 at the next time step in the video frame sequence.). In regards to Claim 15, Angelova in view of Xie in further view of Jenni discloses the video processing method according to claim 13, wherein the trained recognition model (Col. 7, lines 25-31, Angelova teaches the object recognition model may be trained.) includes a plurality of cells (Col. 5, lines 45-61, Angelova teaches the LSTM layer 300 includes one or more LSTM memory blocks, including a memory block 320. The LSTM memory block 320 includes an LSTM memory cell 322 that receives an input xt and generates an output mt from the input and from a previous recurrent projected output rt-1.) of a recurrent neural network (Col. 8, lines 27-32, Fig. 1, Angelova teaches the object recognition model is a recurrent neural network that includes a long short-term memory (LSTM) layer.) that input time-series frames included in the input video (Col.1, lines 54-60, Angelova teaches the multiple frames may be arranged in an order according to their time of occurrence in the video. To process the multiple frames, each frame of the multiple frames may be processed using the LSTM layer in the order according to their time of occurrence in the video.), and input and output state vectors chronologically (Col. 5, lines 45-61, Angelova teaches the LSTM memory block 320 includes an LSTM memory cell 322 that receives an input xt and generates an output mt from the input and from previous recurrent projected output rt-1. For example, the input xt may be the feature output for a video frame at time step t in a video frame sequence. The previous recurrent projected output rt-1 is the projected output generated by the recurrent projection layer from an output rt-1 generated by the cell at the preceding time step in the video frame sequence. In other words, the previous recurrent projected output rt-1 is fed back to the cell. The Examiner interprets the “recurrent projected output” refers to the projection of the hidden state, which is updated.), and a state predictor (Col. 8, lines 58-64, Angelova teaches an LSTM layer, which processes the feature data to generate an LSTM output and to update an internal state of the LSTM layer. The Examiner interprets the LSTM layer to be a “state predictor” since it updates the internal state of the LSTM layer.) that predicts the state vectors based on the first time difference information between the frames of the input video (Col. 8, lines 65-67 through Col. 9, lines 1-8, Col. 3, lines 57-61, Angelova teaches the system may process each frame of the multiple frames using the LSTM layer in the order according to their time occurrence in the video to generate the LSTM output and to update the internal state of the LSTM layer. For example, the forward LSTM layer is configured to process the feature output in forward time steps to generate a forward LSTM output. The Examiner interprets the state vectors are predicted based on time difference information since the frames may be selected based on a predetermined time interval, such as 1 second, and therefore, the system knows the frames are separated by a constant time interval of 1 second.) is inserted between predetermined cells (Col. 5, lines 45-61, Fig. 1, Angelova teaches the previous recurrent output is the projected output generated by the recurrent projection layer from an output rt-1 generated by the cell at the preceding time step in the video frame sequence. The previous recurrent projected output rt-1 is fed back to the cell.). In regards to Claim 16, Angelova in view of Xie in further view of Jenni discloses the video processing method according to claim 15, wherein the trained recognition model into which the state predictor is inserted (Col. 1, lines 51-60, Angelova teaches the feature data may be processed using an LSTM layer to generate an LSTM output and to update an internal state of the LSTM layer.) is trained (Paragraph [0057], Fig. 3, Jenni teaches training a machine-learning model.) using time- series frames (Col. 1, lines 54-55, Angelova teaches the multiple frames may be arranged in an order according to their time of occurrence in the video.) in which frame skipping incurs in a predetermined pattern included in the training video (Paragraphs [0057], [0061-62], [0066], Jenni teaches training the machine-learning model utilizing a self-supervised learning approach with sample sequences of frame skipping from digital videos as training data.), the second time difference information between the frames of the training video, and correct data (Paragraph [0058], Fig. 3, Jenni teaches generating a subsampled video by sampling a sequence of frame skipping’s for a training digital video. The system utilizes the playback speed prediction machine-learning model to generate a playback speed prediction vector that classifies a playback speed per frame (e.g., the first frame transition represented in the first row of the vector is classified to a 2x playback speed or a single frame skip, the second frame transition represented in the second row of the vector is classified to a 1x playback speed or no frame skip, and so forth.) and correct data (Paragraphs [0059], [0064], Fig. 3, Jenni teaches ground truth playback speed classifications that correspond to the frame skips from the sub-sampled digital video.). In regards to Claim 17, Angelova in view of Xie in further view of Jenni discloses the video processing method according to claim 15, wherein the plurality of cells of the trained recognition model (Col. 5, lines 45-61, Angelova teaches the LSTM layer 300 includes one or more LSTM memory blocks, including a memory block 320. The LSTM memory block 320 includes an LSTM memory cell 322 that receives an input xt and generates an output mt from the input and from a previous recurrent projected output rt-1.) are trained (Paragraphs [0051-53], Xie teaches prediction network 306 may learn a model that predicts the dropped-frame ratio based on an input of dropped-frame ratios and timestamps. The prediction network 306 may include multiple units 702-1 to 702-3 that can each generate a prediction. The Examiner interprets “units” and “cells” to be synonymous in the context of LSTM models.) using time-series frames (Col. 1, lines 54-55, Angelova teaches the multiple frames may be arranged in an order according to their time of occurrence in the video.) which are included in the training video (Paragraph [0058], Fig. 3, Jenni teaches a sub-sampled digital video 302 created by sampling a sequence of frame skipping’s for a training digital video.) and in which no frame skipping incurs (Col. 4, lines 7-9, Angelova teaches the input 102 (video) includes a sequence of three video frames ft-1, ft, and ft+1, taken at time t-1, t, and t+1, respectively. The Examiner interprets since the video frames are each taken within one second of each other, no “frame skipping” has occurred.) and correct data (Paragraphs [0059], [0064], Fig. 3, Jenni teaches ground truth playback speed classifications.), and the state predictor (Col. 4, lines 53-67 through Col. 5, lines 1-17, Fig. 1, Angelova teaches forward LSTM layer 106 processes the internal LSTM state from the preceding state and the feature output to generate a forward LSTM output and update the internal state of the forward LSTM layer 106.) inserted into the trained recognition model (Col. 3, lines 41-52, Fig. 1, Angelova teaches an object recognition model 100 having a convolutional LSTM layer. The Examiner interprets the “LSTM layer” to be a “state predictor” since the LSTM layer updates the internal state of the forward LSTM layer.) is trained (Col. 7, lines 25-30, Angelova teaches the object recognition model may be trained such that the classification layers store received LSTM outputs until the forward LSTM layer has processed all of the frames in the sequence and has generated all the LSTM outputs, before generating the output representing a set of scores.) using a state vector output at time t (where t is natural number) (Col. 6, lines 26-29, Angelova teaches computing the output mt.) and a state vector output at time t+N (where N is a natural number) (Col. 6, lines 50-53, Angelova teaches computing the output mt+1 at the next time step in the video frame sequence.) by the plurality of cells (Col. 5, lines 45-61, Angelova teaches the LSTM memory block which includes an LSTM memory cell receives an input xt and generates output mt from the input and from a previous recurrent projected output rt-1.) when time-series frames (Col. 4, lines 7-9, Angelova teaches the input includes a sequence of three video frames ft-1, ft, and ft+1, taken at time t-1, t, and t+1, respectively.) which are included in the training video (Paragraph [0058], Fig. 3, Jenni teaches a sub-sampled digital video 302.) and in which no frame skipping incurs (Col. 4, lines 7-9, Angelova teaches the input includes a sequence of three video frames ft-1, ft, and ft+1, taken at time t-1, t, and t+1, respectively. The Examiner interprets no frame skipping occurs since the frames are each taken 1 second apart.) are input to the plurality (Col. 9, lines 2-6, Col. 5, lines 47-50, Angelova teaches the system 100 may process each frame of the multiple frames using the LSTM layer in the order according to their time of occurrence in the video to generate LSTM output and to update the internal state of the LSTM layer. The LSTM layer 300 includes one or more memory blocks, including an LSTM memory block 320. The LSTM memory block 320 includes an LSTM memory cell 322. The Examiner interprets since there are one or more memory blocks, each of which have a memory cell, there are a “plurality of cells”.) of trained cells (Col. 7, lines 25-29, Angelova teaches the object recognition model may be trained. The Examiner interprets the memory cells are trained since the overall model is trained.). Claims 6, 12 and 18 are rejected under 35 U.S.C. 103(a) as being unpatentable over Angelova et al. (U.S. Patent No. 10,013,640, hereafter referred to as Angelova) in view of Xie et al. (U.S. Patent Pub No. 2021/0051368, hereafter referred to as Xie) in further view of Jenni et al. (U.S. Patent Pub. No. 2023/0276084, hereafter referred to as Jenni) in further view of Rassool (U.S. Patent Pub. No. 2021/0182539, hereafter referred to as Rassool). Regarding Claim 6, Angelova in view of Xie in further view of Jenni discloses the video processing system according to claim 1, wherein the recognition means (Col. 13, lines 20-24, Angelova teaches the processes can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The Examiner is interpreting “recognition means” to be a computer.) inputs the input video (Col. 3, lines 33-35 and lines 43-46, Angelova teaches the consecutive video frames as input.), the first time difference information between the frames of the input video (Paragraphs [0051-53], [0040], Fig. 7, Xie teaches each unit of the prediction network may include an input that receives the dropped frame ratio, also units 702-1 to 702-3 include time inputs 706-1 to 706-3 that receive the time difference associated with each dropped frame ratio. The prediction network is trained.), (Paragraph [0057], Fig. 3, Jenni teaches training a machine-learning model with sample sequences of frame skippings from digital videos as training data.), the second time difference information between the frames of the training video (Paragraphs [0057], [0059], [0064], Fig. 3, Jenni teaches the using the sequence of frame skippings 308 of the digital video as ground truth data for the predicted playback speeds. For example, a sequence of frame skippings of 2, 1, 1, and 2 would translate to ground truth playback speed classifications of 1, 0, 0, and 1 In frame order.), (Abstract, Col. 8, lines 40-48, Angelova teaches processing, using an object recognition model, the multiple frames to generate data that represents a classification of the object to be recognized.). Angelova in view of Xie in further view of Jenni does not explicitly disclose input Rassool is in the same field of art of performing object recognition (such as facial recognition to identify faces) in a video, which may be processed using a recurrent neural network. Further, Rassool teaches input(Paragraphs [0078], [0070], Rassool teaches generating motion information based on differences between the image frames. For example, the system may determine a position of a defined point in the first frame and a position of the corresponding point in the second frame. Motion information may then be calculated as the difference between positions divided by the difference in time between the first and second frames, such as a difference in timestamps associated with respective frames. The motion information may be provided as input to the motion neural network.) (trained using) the motion between the frames of the training video (Paragraph [0063], Fig. 5, Rassool teaches the training data may include motion information of the defined set of points of the face in the video data.). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Angelova in view of Xie in further view of Jenni by inputting motion information such as movement that occurred between frames to the recognition model that is taught by Rassool, to make the invention that indicates how features in the video data move during a sequence of frames; thus, one of ordinary skilled in the art would be motivated to combine the references since motion information extracted from video frames can be used as an additional cue for object recognition and LSTM models are capable of learning motion dependencies in video frames to improve the recognition accuracy (Angelova, Col. 1, lines 14-31). Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. In regards to Claim 12, Angelova in view of Xie in further view of Jenni discloses the video processing apparatus according to claim 7, wherein the recognition means (Col. 13, lines 20-24, Angelova teaches the processes can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The Examiner is interpreting “recognition means” to be a computer.) inputs the input video (Col. 3, lines 33-35 and lines 43-46, Angelova teaches the consecutive video frames as input.), the first time difference information between the frames of the input video (Paragraphs [0051-53], [0040], Fig. 7, Xie teaches each unit of the prediction network may include an input that receives the dropped frame ratio, also units 702-1 to 702-3 include time inputs 706-1 to 706-3 that receive the time difference associated with each dropped frame ratio. The prediction network is trained.), (Paragraph [0057], Fig. 3, Jenni teaches training a machine-learning model with sample sequences of frame skippings from digital videos as training data.), the second time difference information between the frames of the training video (Paragraphs [0057], [0059], [0064], Fig. 3, Jenni teaches the using the sequence of frame skippings 308 of the digital video as ground truth data for the predicted playback speeds. For example, a sequence of frame skippings of 2, 1, 1, and 2 would translate to ground truth playback speed classifications of 1, 0, 0, and 1 In frame order.), (Abstract, Col. 8, lines 40-48, Angelova teaches processing, using an object recognition model, the multiple frames to generate data that represents a classification of the object to be recognized.). Angelova in view of Xie in further view of Jenni does not explicitly disclose input Rassool is in the same field of art of performing object recognition (facial recognition to identify faces) in a video, which may be processed using a recurrent neural network. Further, Rassool teaches inputs(ting) a motion between frames of the input video (to a trained recognition model) (Paragraphs [0078], [0070], Rassool teaches generating motion information based on differences between the image frames. For example, the system may determine a position of a defined point in the first frame and a position of the corresponding point in the second frame. Motion information may then be calculated as the difference between positions divided by the difference in time between the first and second frames, such as a difference in timestamps associated with respective frames. The motion information may be provided as input to the motion neural network.) (trained using) the motion between the frames of the training video (Paragraph [0063], Fig. 5, Rassool teaches the training data may include motion information of the defined set of points of the face in the video data.). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Angelova in view of Xie in further view of Jenni by inputting motion information such as movement that occurred between frames to the recognition model that is taught by Rassool, to make the invention that indicates how features in the video data move during a sequence of frames; thus, one of ordinary skilled in the art would be motivated to combine the references since motion information extracted from video frames can be used as an additional cue for object recognition and LSTM models are capable of learning motion dependencies in video frames to improve the recognition accuracy (Angelova, Col. 1, lines 14-31). Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. In regards to Claim 18, Angelova in view of Xie in further view of Jenni discloses the video processing method according to claim 13, wherein the computer (Col. 2, lines 46-56, Angelova teaches a system of one or more computers can be so configured by virtue of software, firmware, hardware, or a combination of them installed on the system that in operation cause the system to perform the actions) inputs the input video (Col. 3, lines 33-35 and lines 43-46, Angelova teaches the consecutive video frames as input.), the first time difference information between the frames of the input video (Paragraphs [0051-53], [0040], Fig. 7, Xie teaches each unit of the prediction network may include an input that receives the dropped frame ratio, also units 702-1 to 702-3 include time inputs 706-1 to 706-3 that receive the time difference associated with each dropped frame ratio. The prediction network is trained.), (Paragraph [0057], Fig. 3, Jenni teaches training a machine-learning model with sample sequences of frame skippings from digital videos as training data.), the second time difference information between the frames of the training video (Paragraphs [0057], [0059], [0064], Fig. 3, Jenni teaches the using the sequence of frame skippings 308 of the digital video as ground truth data for the predicted playback speeds. For example, a sequence of frame skippings of 2, 1, 1, and 2 would translate to ground truth playback speed classifications of 1, 0, 0, and 1 In frame order.), (Abstract, Col. 8, lines 40-48, Angelova teaches processing, using an object recognition model, the multiple frames to generate data that represents a classification of the object to be recognized.). Angelova in view of Xie in further view of Jenni does not explicitly disclose input Rassool is in the same field of art of performing object recognition (facial recognition to identify faces) in a video, which may be processed using a recurrent neural network. Further, Rassool teaches inputs(ting) a motion between frames of the input video (to a trained recognition model) (Paragraphs [0078], [0070], Rassool teaches generating motion information based on differences between the image frames. For example, the system may determine a position of a defined point in the first frame and a position of the corresponding point in the second frame. Motion information may then be calculated as the difference between positions divided by the difference in time between the first and second frames, such as a difference in timestamps associated with respective frames. The motion information may be provided as input to the motion neural network.) (trained using) the motion between the frames of the training video (Paragraph [0063], Fig. 5, Rassool teaches the training data may include motion information of the defined set of points of the face in the video data.). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Angelova in view of Xie in further view of Jenni by inputting motion information such as movement that occurred between frames to the recognition model that is taught by Rassool, to make the invention that indicates how features in the video data move during a sequence of frames; thus, one of ordinary skilled in the art would be motivated to combine the references since motion information extracted from video frames can be used as an additional cue for object recognition and LSTM models are capable of learning motion dependencies in video frames to improve the recognition accuracy (Angelova, Col. 1, lines 14-31). Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Pertinent Prior Art The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Yoon et al. (U.S. Patent Pub. No. 2024/0331340 A1) teaches an image processing device for estimating object data, which may include estimating object data of a current image frame, based on the past move vector and a weight parameter based on a time difference between the current image frame ad the previous object data detected before the current image frame. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to SYDNEY L BLACKSTEN whose telephone number is (571)272-7120. The examiner can normally be reached 8:30am-4:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Oneal Mistry can be reached at 313-446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SYDNEY L BLACKSTEN/Examiner, Art Unit 2674 /ONEAL R MISTRY/Supervisory Patent Examiner, Art Unit 2674
Read full office action

Prosecution Timeline

Oct 16, 2024
Application Filed
Aug 21, 2026
Non-Final Rejection mailed — §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
100%
Grant Probability
99%
With Interview (+0.0%)
2y 5m (~6m remaining)
Median Time to Grant
Low
PTA Risk
Based on 4 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month