DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of claims: claims 1-20 are pending below.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on June 25th 2025 and June 11th 2026 was filed and considered. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “video prediction model is configured to” in claim 1, “casual temporal attention subblock is configured to” in claim 3, “decouple spatial attention subblock is configured to” in claim 3, “preset multi-layer neural network is configured to” in claim 7.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 8-9 and 16 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by ALVAREZ et al (US 2021/0380143).
Claim 1, similarly claim 8-9 and 16:
ALVAREZ et al (US 2021/0380143) teaches the following subject matter:
A method for training an autonomous driving model, wherein the autonomous driving model comprises
a video prediction model, and the method comprises:
determining, according to at least one of an initial video frame collected by a target vehicle or scenario description metadata of an initial video frame, a scenario context of the initial video frame (paragraph 0037 detail information (also known as key event heuristics) the prioritize key events module 250 may make a decision as to which perspectives/points of view should be included in the augmented image sequence (e.g., the sensor stream), which text/metadata should be displayed, and/or the length of the summarization (e.g., the number of frames). The result is an event summarization that includes a sequence summary that may include a summarized image/video as well as an associated textual description of the key event or events; figure 2 and paragraph 0040-0041);
determining a vehicle movement instruction of the target vehicle according to at least one of the initial video frame or trajectory data of the target vehicle corresponding to the initial video frame (0078 detail event metadata includes key event that includes movement event for trajectory of the vehicle); and
training an initial model using the initial video frame and a control text corresponding to the initial video frame, to obtain the video prediction model, wherein the control text comprises the scenario context and the vehicle movement instruction, and the video prediction model is configured to output a predicted video frame (0068-0071 detail use of learning model/neural network with predictive layer (training of video prediction model) to provide driving specific adjustment for vehicle, where in 0072-0075 further provide prediction for automated driving with image data information of pedestrian or other traffic participants as presented by key events (includes text/metadata, movement/trajectory) stored, prioritized event scenario as taught in 0075).
Regarding claim 8 same to claim 1.
Regarding claim 9, device is taught in 0026, where 0020 further detail processor and memory.
Regarding claim 16, non-transitory computer-readable storage medium is taught in 0021.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 7, 15 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over ALVAREZ et al (US 2021/0380143) in view of Li et al (US 2024/0087222).
Claim 7, similarly claims 15 and 20:
ALVAREZ et al teaches all the subject matter above, but not the following:
The method of claim 1, wherein the video prediction model further comprises a preset multi-layer neural network after a Transformer model, and wherein the preset multi-layer neural network is configured to generate predicted trajectory points of the target vehicle corresponding to the predicted video frame according to a feature map outputted by the Transformer model.
Li et al (US 2024/0087222) teaches the following subject matter:
The method of claim 1, wherein the video prediction model further comprises a preset multi-layer neural network after a Transformer model (0006 teaches neural network then transformer for convert for image feature maps; 0018, 0024, 0054, figure 2 and 0058-0061), and wherein the preset multi-layer neural network is configured to generate predicted trajectory points of the target vehicle corresponding to the predicted video frame according to a feature map outputted by the Transformer model (0251 detail application to controlling and/or interacting components such as vehicle 700 such as ADAS system 738 for autonomous driving information for vehicle maneuvers and trajectories of surrounded environment).
ALVAREZ et al and Li et al are both in the field of image analysis, especially assisting/autonomous drive using predictive frame with use of neural network with metadata and text data (Li et al: paragraph 018) from associated image such that the combine outcome is predictable.
Therefore it would have been obvious to one having ordinary skill before the effective filing date to modify ALVAREZ et al by Li et al with the additional application of transformer-based provide efficient and scalable for sparse representation as disclosed by Li et al paragraph 0040-0041.
Allowable Subject Matter
Claim 2, and its dependent claims 3-6, are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim 10, and its dependent claims 11-14, are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim 17, and its dependent claims 18-19, are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Rust et al (US 2017/0192423) teaches SYSTEM AND METHOD FOR REMOTELY ASSISTING AUTONOMOUS VEHICLE OPERATION - assistance data may include raw sensor data (e.g., live-stream or still images) of a camera of the autonomous vehicle, auditory data from microphones of the autonomous vehicle, a and the like), processed sensor data (e.g., a camera view overlaid with object identification indicators placed by the autonomous vehicle's on-board computer, predicted vehicle trajectory, three-dimensional renderings of the autonomous vehicle's environment), autonomous vehicle analysis (e.g., a text description, generated by the autonomous vehicle (paragraph 0062).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TSUNG-YIN TSAI whose telephone number is (571)270-1671. The examiner can normally be reached 7am-4pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached at (571) 272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/TSUNG YIN TSAI/Primary Examiner, Art Unit 2656