DETAILED ACTION
This is a response to Applicant’s submissions filed on 6/18/2026. Claims 1-20 are pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 6/18/2026 has been entered.
Information Disclosure Statement
The listing of references in paragraph 97 of the specification is not a proper information disclosure statement. 37 CFR 1.98(b) requires a list of all patents, publications, or other information submitted for consideration by the Office, and MPEP § 609.04(a) states, "the list may not be incorporated into the specification but must be submitted in a separate paper." Therefore, unless the references have been cited by the examiner on form PTO-892, they have not been considered.
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
It is noted that Applicant’s amendments to claim 14 have overcome the previous rejection under 35 U.S.C. § 112.
In response to Applicant’s argument that element 622 in Figure 6 is an input rather than output (Applicant’s Remarks; p. 9), it is noted that reference character 622 should have read 606 in the previous office action. Figure 6 appears to include a view of sensor data images of the environment surrounding the vehicle, and a view of a generated map of the environment surrounding the vehicle. Each view should be individually numbered. See objection below.
Drawings
The amended drawings were received on 6/18/2026.
The drawings are objected to because:
Figure 6 appears to include different views of exemplary inputs (620-630) and exemplary outputs (606) of the parallelized model. Each different view must be numbered (see CFR § 1.84(u)(1)), such as by indicating the top collection of inputs 620-630 as figure 6A, and the bottom exemplary outputs as figure 6B.
In the top view of figure 6, it is unclear how reference character 604 illustrates planned outputs generated by the parallelized model as disclosed in paragraph 173, because it appears that reference character 604 merely indicates the centerline markings on a road.
Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Specification
The disclosure is objected to because of the following informality: in paragraph 97, line 2, it is unclear how the water disinfection apparatus disclosed in US application 16/101,432 discloses a real-time ray-tracing hardware accelerator.
Appropriate correction is required.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1-2, 6-12 and 14-20 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Chen et al. (US 2026/0021831), hereinafter Chen.
Regarding claims 1, 11 and 20, Chen discloses a system, comprising: one or more memories storing instructions (Chen; para. 220: memory 1102 stores a program instruction corresponding to … the autonomous driving method); and one or more processors that are coupled to the one or more memories (Chen; fig. 11: processor 1102) and, when executing the instructions, are configured to: receive sensor data (Chen; fig. 2: sensing data) and vehicle information that is distinct from the sensor data (Chen; fig. 2: ego vehicle data/navigation data); and process the sensor data and the vehicle information via a machine learning model (Chen; fig. 2: neural network) in which a plurality of modules execute in parallel (Chen; fig. 2: head networks 1-N) based on one or more cross-attention features (Chen; para. 97: the first cross-attention network is used to perform feature fusion on the first query matrix, the second key matrix, and the second value matrix, and the second cross-attention network is used to perform feature fusion on the second query matrix, the first key matrix, and the first value matrix. The N types of feature data are generated based on an output result of the first cross-attention network and an output result of the second cross-attention network.; Chen; para. 99: the N types of feature data include a prediction feature, a decision feature, a road feature, and an ego vehicle feature. Each type of feature data is subsequently input into a corresponding head network, to perform a corresponding autonomous driving task.) that combine information (Chen; para. 97: the first cross-attention network is used to perform feature fusion on the first query matrix, the second key matrix, and the second value matrix, and the second cross-attention network is used to perform feature fusion on the second query matrix, the first key matrix, and the first value matrix) from at least one of the sensor data or the vehicle information with learned prior information (Chen; para. 96: the first self-attention network is used to perform feature extraction on ego vehicle data and obstacle data in one piece of input data, to obtain a first query (Query) matrix, a first key (Key) matrix, and a first value (Value) matrix (that is, Q, K, and V that are output by the first self-attention network in FIG. 4). The second self-attention network is used to perform feature extraction on road topology data and navigation data in one piece of input data, to obtain a second query (Query) matrix, a second key (Key) matrix, and a second value (Value) matrix (that is, Q, K, and V that are output by the second self-attention network in FIG. 4); para. 80: the input data of the neural network is vectorized data. Specifically, a vectorized form of the ego vehicle data and the obstacle data is represented by locations, orientation, sizes, and the like of objects such as the ego vehicle and an obstacle at different moments in a past time period (for example, elapsed 2 s counted from a collection moment). A vectorized form of the road topology data is represented by attributes such as a center line, a road side line, a traffic light, and a speed limit plate of a road section (for example, 10 m to 20 m near the ego vehicle) at different moments in a past time period. A vectorized form of the navigation data is represented by a driving road, a driving lane, a driving trajectory, and the like of the ego vehicle in a past time period) to generate at least a planned motion for the vehicle (Chen; paras. 85-86: The post-processing converts a predicted trajectory of an obstacle, an ego vehicle decision, a planned trajectory of an ego vehicle, and the like that are output by the head network into final control data of the ego vehicle … The ego vehicle controls traveling of the ego vehicle based on the control data.).
Regarding claims 2 and 12, Chen discloses processing the sensor data and the vehicle information via the machine learning model comprises: generating one or more tokens based on the sensor data (Chen; para. 96: the first self-attention network is used to perform feature extraction on ego vehicle data and obstacle data in one piece of input data, to obtain a first query (Query) matrix, a first key (Key) matrix, and a first value (Value) matrix (that is, Q, K, and V that are output by the first self-attention network in FIG. 4). The second self-attention network is used to perform feature extraction on road topology data and navigation data in one piece of input data, to obtain a second query (Query) matrix, a second key (Key) matrix, and a second value (Value) matrix (that is, Q, K, and V that are output by the second self-attention network in FIG. 4)); processing the one or more tokens and one or more learned tokens associated with the learned prior information via one or more cross-attention layers to generate the one or more cross-attention features (Chen; para. 97: the first cross-attention network is used to perform feature fusion on the first query matrix, the second key matrix, and the second value matrix, and the second cross-attention network is used to perform feature fusion on the second query matrix, the first key matrix, and the first value matrix. The N types of feature data are generated based on an output result of the first cross-attention network and an output result of the second cross-attention network.); and processing each cross-attention feature included in the one or more cross- attention features via one of the modules included in the plurality of modules (Chen; para. 99: the N types of feature data include a prediction feature, a decision feature, a road feature, and an ego vehicle feature. Each type of feature data is subsequently input into a corresponding head network, to perform a corresponding autonomous driving task.).
Regarding claims 6 and 14, Chen discloses the plurality of modules includes at least one of a neural network to generate maps, a neural network to predict motions of objects, a neural network to predict occupancy of the objects within an environment, or a neural network to generate planned motions of the vehicle (Chen; para. 88: the neural network includes … N parallel head networks; para. 108: Each of the N head networks outputs at least one type of output result.; para. 110: an output result corresponding to the prediction task includes a predicted trajectory of an obstacle and predicted trajectory distribution of the obstacle; para. 112: an output result corresponding to the planning task includes a planned road of an ego vehicle, a planned trajectory of the ego vehicle, and planned trajectory distribution of the ego vehicle).
Regarding claim 7, Chen discloses performing one or more operations to train the machine learning model, wherein the one or more operations cause one or more parameters of the plurality of modules and one or more values of one or more learned tokens to be updated (Chen; fig. 5: S550; para. 155: The parameter in the first head network may be updated in either of the following two manners: (1) Only the parameter in the first head network is updated, and a structure of the first head network is not adjusted; and (2) A structure of the first head network is adjusted, and after adjustment, the parameter in the first head network is updated.).
Regarding claim 8, Chen discloses the vehicle information includes at least one of one or more commands used to control the vehicle, controller area network (CAN) bus information, or a history of one or more trajectories of the vehicle (Chen; para. 80: a vectorized form of the ego vehicle data and the obstacle data is represented by locations, orientation, sizes, and the like of objects such as the ego vehicle and an obstacle at different moments in a past time period (for example, elapsed 2 s counted from a collection moment) … A vectorized form of the navigation data is represented by a driving road, a driving lane, a driving trajectory, and the like of the ego vehicle in a past time period.).
Regarding claims 9 and 17, Chen discloses processing the sensor data and the vehicle information via the machine learning model further generates at least one of a map of an environment, a predicted motion of one or more objects, or a predicted occupancy of the one or more objects within the environment (Chen; para. 108: Each of the N head networks outputs at least one type of output result.; para. 110: an output result corresponding to the prediction task includes a predicted trajectory of an obstacle and predicted trajectory distribution of the obstacle; para. 112: an output result corresponding to the planning task includes a planned road of an ego vehicle, a planned trajectory of the ego vehicle, and planned trajectory distribution of the ego vehicle).
Regarding claims 10 and 19, Chen discloses performing one or more operations to control the vehicle based on the planned motion (Chen; paras. 85-86: The post-processing converts a predicted trajectory of an obstacle, an ego vehicle decision, a planned trajectory of an ego vehicle, and the like that are output by the head network into final control data of the ego vehicle … The ego vehicle controls traveling of the ego vehicle based on the control data.).
Regarding claim 15, Chen discloses the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of performing one or more operations to train the machine learning model, wherein the one or more operations cause one or more parameters of the plurality of modules to be updated (Chen; fig. 5: S550) in parallel (Chen; para. 145: a parameter in at least one head network corresponding to the K types of output results is updated, to implement … multi-task joint training) based on a computed loss (Chen; para. 167: the K accumulated loss values are weighted to obtain one loss value corresponding to a single-task training/multi-task joint training process, and back propagation is performed based on the loss value, to update parameters in the entire neural network).
Regarding claim 16, Chen discloses a first cross-attention layer included in the one or more cross-attention layers generates first cross-attention features included in the one or more cross-attention features (Chen; para. 97: the first cross-attention network is used to perform feature fusion on the first query matrix, the second key matrix, and the second value matrix, and the second cross-attention network is used to perform feature fusion on the second query matrix, the first key matrix, and the first value matrix. The N types of feature data are generated based on an output result of the first cross-attention network and an output result of the second cross-attention network.) based on the vehicle information, the one or more tokens (Chen; para. 96: the first self-attention network is used to perform feature extraction on ego vehicle data and obstacle data in one piece of input data, to obtain a first query (Query) matrix, a first key (Key) matrix, and a first value (Value) matrix (that is, Q, K, and V that are output by the first self-attention network in FIG. 4). The second self-attention network is used to perform feature extraction on road topology data and navigation data in one piece of input data, to obtain a second query (Query) matrix, a second key (Key) matrix, and a second value (Value) matrix (that is, Q, K, and V that are output by the second self-attention network in FIG. 4)), and a first learned token included in the one or more learned tokens (Chen; para. 80: a vectorized form of the ego vehicle data and the obstacle data is represented by locations, orientation, sizes, and the like of objects such as the ego vehicle and an obstacle at different moments in a past time period (for example, elapsed 2 s counted from a collection moment) … A vectorized form of the navigation data is represented by a driving road, a driving lane, a driving trajectory, and the like of the ego vehicle in a past time period.).
Regarding claim 18, Chen discloses the sensor data includes at least one of image data, light detection and ranging (LIDAR) data, or radio detection and ranging (RADAR) data (Chen; para. 174: the ego vehicle senses an ambient environment by using a device such as a camera, a sensor, or a lidar).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 3 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Cherian et al. (US 2025/0187198), hereinafter Cherian.
Chen discloses processing camera data (Chen; para. 174: the ego vehicle senses an ambient environment by using a device such as a camera) using a neural network (Chen; para. 176: Perform feature extraction and feature fusion on the input data by using a mixture of experts MOE network in the neural network).
Chen does not explicitly disclose generating the one or more tokens comprises processing the sensor [e.g., camera] data via one of a spatiotemporal transformer model, an autoregressive transformer model, or a QueryTransformer (QFormer) model.
Cherian, in the same field of endeavor (autonomous vehicle controls), discloses generating one or more tokens comprises processing camera data via one of a spatiotemporal transformer model (Cherian; para. 83: The pre-processed sequence of video frames 402a may then be inputted to a spatio-temporal transformer 404. The spatio-temporal transformer 404 transforms each of the pre-processed sequence of video frames 402a frames into a spatio-temporal scene graph 406 (G) of the sequence of video frames 102 to capture spatio-temporal information of the sequence of video frames 102.), an autoregressive transformer model, or a QueryTransformer (QFormer) model.
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, with a reasonable expectation of success, to have modified the neural network that processes the camera data of Chen to process the video frames using a spatiotemporal transformer, as disclosed by Cherian, with the motivation of accurately determining the 3D spatio-semantic relationships of the objects in the scene (Cherian; para. 81) using a computationally efficient and feasible navigation method (Cherian; para. 8).
Claim(s) 4, 5 and 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Zhou et al. (US 2025/0166366), hereinafter Zhou.
Regarding claim 4, Chen discloses the invention substantially as claimed as described above.
Chen does not explicitly disclose each module included in the plurality of modules comprises a decoder model.
Zhou, in the same field of endeavor (autonomous vehicle control), discloses a module included in a plurality of modules (Zhou; fig. 1) comprises a decoder model (Zhou; para. 53: The system uses the trajectory prediction neural network 114 to generate a trajectory prediction output 108 by processing tokens representing both sensor data 102 and perception outputs 106. The trajectory prediction neural network 114 includes a scene encoder 202, a trajectory decoder 204, and a shared multilayer perceptron (MLP) 224.).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, with a reasonable expectation of success, to have modified the head networks of Chen to include an encoder and decoder, as disclosed by Zhou, with the motivation of reducing the number of tokens needed to represent the environment thereby increasing the accuracy of trajectory predictions and increasing training efficiency (Zhou; para. 23).
Regarding claim 5, Chen discloses the invention substantially as claimed as described above.
Chen does not explicitly disclose each module included in the plurality of modules comprises an encoder model and a decoder model.
Zhou discloses a module included in a plurality of modules (Zhou; fig. 1) comprises an encoder model and a decoder model (Zhou; para. 53: The system uses the trajectory prediction neural network 114 to generate a trajectory prediction output 108 by processing tokens representing both sensor data 102 and perception outputs 106. The trajectory prediction neural network 114 includes a scene encoder 202, a trajectory decoder 204, and a shared multilayer perceptron (MLP) 224.).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, with a reasonable expectation of success, to have modified the head networks of Chen to include an encoder and decoder, as disclosed by Zhou, with the motivation of reducing the number of tokens needed to represent the environment thereby increasing the accuracy of trajectory predictions and increasing training efficiency (Zhou; para. 23).
Regarding claim 13, Chen discloses the invention substantially as claimed as described above.
Chen does not explicitly disclose each module included in the plurality of modules comprises a multi-layer perceptron.
Zhou discloses a module included in a plurality of modules (Zhou; fig. 1) comprises a multi-layer perceptron (Zhou; para. 53: The system uses the trajectory prediction neural network 114 to generate a trajectory prediction output 108 by processing tokens representing both sensor data 102 and perception outputs 106. The trajectory prediction neural network 114 includes a scene encoder 202, a trajectory decoder 204, and a shared multilayer perceptron (MLP) 224.).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, with a reasonable expectation of success, to have modified the head networks of Chen to include a multilayer perceptron, as disclosed by Zhou, with the motivation of mapping multiple module outputs to a single prediction for one or more agents in the environment (Zhou; para. 69) thereby reducing the amount of post-processed data.
Supplemental References
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Tang et al. (US 2023/0152791) disclose using a cross-attention based anomaly detection engine to find anomalies from multi-modality data comprising vehicle system data and sensed environmental conditions so a controller can automatically adjust navigation (or other vehicle tasks) to account for a defect detected in the vehicle’s operation.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOSEPH THOMPSON whose telephone number is (571)272-3660. The examiner can normally be reached Mon-Thurs 9:00AM-3:00PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Erin Bishop can be reached at (571)270-3713. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JOSEPH THOMPSON/Examiner, Art Unit 3665
/Erin D Bishop/Supervisory Patent Examiner, Art Unit 3665