DETAILED ACTION
This action is in response to Applicant’s response filed on 06/12/2026. The Amendment filed on 06/12/2026 has been entered. Claims 1, 11, and 20 are amended. Claims 1-20 are now pending in the present application. This Action is made FINAL.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
No information disclosure statement was submitted with the present application filed on 04/12/2024, or with the amendment filed on 06/12/2026.
Priority
The present application claims benefit of provisional application filed on April 19, 2023.
Response to Arguments
Applicant's arguments filed 06/12/2026 have been fully considered but they are not persuasive. In summary, Applicant argues the Abdo (US 20220319054 A1; hereinafter Abdo) and Wu (“Location prediction on trajectory data: A review”; copy previously provided by examiner; hereinafter Wu) fail to teach the amended independent claims 1, 11, and 20; therefore, the 103 rejections should be traversed. Examiner has addressed the Applicant’s specific arguments below.
(1) Applicant argues that Abdo provides a method for generating scene flow labels for point clouds using object labels and relies on point clouds obtained at different time points.
Examiner response: Applicant’s characterization of Abdo as providing a method for generating a scene flow label for point clouds using object labels is acknowledged but does not account for the entirety of the teachings relied upon from Abdo, nor distinguish the claimed subject matter. In addition to generating scene flow labels as ground-truth training data (Abdo [0009]-[0015]; [0037]-[0042]), Abdo further teaches a trained scene flow neural prediction neural network, including a scene encoder and decoder, that generate and process embeddings to generate predicted motion vectors ([0024]-[0029]; [0074]-[0087]; [0089]-[0095]). Accordingly, Applicant’s characterization addresses only the training-label generation aspect of Abdo and does not account for the additional scene-flow prediction teachings relied upon in the rejection
Further, Applicant’s reliance on Abdo’s use of point clouds from different time points does not establish that Abdo lacks position and movement information corresponding to a particular point in time. Abdo expressly teaches that a motion vector represents motion of a point “as of the time that the given point cloud was generated” (Abdo [0010]), and further teaches that its velocity components represent velocity “at the most recent time point” ([0026]) or “at the current time point” ([0060]). Thus, although preceding information is used to determine the movement state ([0061]-[0062]), the resulting movement state corresponds to the current time point. As currently written, the claims do not require that a movement state corresponding to a first time point be derived exclusively from information obtained at that time point.
(2) Applicant argues that Abdo expresses motion vectors in respective directions in the reference frame of the laser sensor at the most recent time point and transforms information between sensor reference frames, whereas the claimed obstacles are transformed based on respective local coordinate systems generated based on the position and movement state of each obstacle at a single time point. Applicant further contends that each local coordinate system is used as the basis for generating the local position and local movement state of the corresponding obstacle.
Respectfully, Applicant’s asserted distinction is not persuasive with respect to the teachings relied upon from Abdo. While Applicant is correct that Abdo expresses the motion-vector components along directions of the current sensor reference frame ([0025-0027]; [0060]), and uses information from different time points, the resulting position and movement information nevertheless corresponds to the current time point, as discussed above in response to Applicant argument (1). Abdo is not limited to merely transforming an entire point cloud between sensor reference frames. Abdo identifies regions containing respective objects ([0050]-[0052]), determines preceding and current object poses represented by transformation matrices having translational and rotational components, determines an object-specific rigid-body transform, and applies the transform to points corresponding to the respective object ([0067]; [0070]-[0073]). Thus, Abdo uses object-specific pose and transformation information as a basis for transforming positional information corresponding to the respective object. Accordingly, Abdo teaches both movement information corresponding to a particular time point and object-specific transformation of spatial information. Applicant’s characterization of Abdo as merely integrating point clouds from different time points into a common sensor coordinate system therefore does not account for the additional teachings of Abdo relied upon in the rejection.
(3) Applicant further argues that Wu does not teach generating a tensor corresponding to obstacle position and movement state at a specific time point and that the combination of Abdo and Wu therefor fails to teach the claimed features found in claims 1, 11, and 20.
Applicant’s argument that Wu does not independently disclose generating an obstacle tensor from the claimed obstacle position and movement-state information is likewise not persuasive as to the rejection previously presented. The prior rejection did not solely rely upon Wu teaching the entirety of the claimed obstacle-tensor generation. Rather, Abdo was relied upon for spatial representations of objects and corresponding motion information, while Wu was relied upon for the use of tensor structures representing location and temporal information for prediction. The rejection proposed applying Wu’s known tensor representation to the position and movement information taught by Abdo. Therefore, obviousness is determined based on the combined teaching of the references rather than whether Wu independently describes Abdo’s scene-flow data. Thus, Applicant’s argument regarding what Wu fails to disclose individually does not address the combined teachings of Abdo and Wu upon which the rejection was based.
Applicant’s arguments have been fully considered. To the extent discussed above, the arguments are not persuasive with respect to the Abdo and Wu references relied upon in the prior rejection. However, Applicant has amended the independent claims to further limit the claimed subject matter. Accordingly, the prior rejection is withdrawn in view of the amended claim language, and the amended claims are rejected on the new grounds set forth below. Specifically, the newly added limitation “wherein a position coordinate of each of the obstacles is an origin of the corresponding local coordinate system” is addressed by the additional teachings of Gao (US 20210150350 A1), as set forth in the new ground of rejection below.
Examiner Note
The previous 103 rejections for claims 11 and 20 found in the prior non-final office action mailed on 03/24/2026 included the following erroneous language from the Examiner, “Abdo further teaches the interpreted structure of "the scene encoding generating apparatus" identified in the U.S.C. 112(f) interpretation, previously discussed in the present office action, for performing the limitations found in the present claim ([0098]; [0101]; [0103]; [106]).” No U.S.C. 112(f) interpretation was or is applied to claims 11 and 20. Accordingly, claims 11 and 20 are both interpreted under their broadest reasonable interpretations.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 11, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Abdo (US 20220319054 A1; hereinafter Abdo), in view of Gao (US 20210150350 A1; hereinafter Gao), and in further view of Wu (“Location prediction on trajectory data: A review”; copy provided by examiner; hereinafter Wu).
Regarding claim 1 (Currently Amended),
Abdo teaches: A scene encoding generating apparatus, comprising:
(Abstract “Methods, systems, and apparatus, including computer programs encoded on computer storage media, for predicting scene flow”; [0024] “The scene flow prediction system 150 processes the point clouds 132 to generate a scene flow output.”
a communication interface; ([0106] “a computing system that includes…a client computer having a graphical user interface…components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network")
and a processor, coupled to the communication interface, and the processor is configured to execute the following operations ([0098] "…'data processing apparatus' refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers."; [0101] "As used in this specification, an 'engine,' or 'software engine,' refers to a software implemented input/output system…Each engine can be implemented on any appropriate type of computing device…that includes one or more processors…"; see also [0106] regarding interconnection of the components of the computing system.):
receiving a position and a movement state in a first time point of each of a plurality of obstacles ([0103] "central processing unit will receive instructions and data"; Abdo teaches receiving current point cloud representing an observed scene at a current time point, wherein each point has a location in a specified coordinate system (Abstract “obtaining a current point cloud representing an observed scene at a current time point; obtaining object label data that identifies a first three-dimensional region in the observed scene”; [0027] “A point cloud generally includes multiple points that represent a sensor measurement of a scene in an environment captured by one or more sensors. Each point has a location in a specified coordinate system, e.g., a three-dimensional coordinate system centered at the sensor, and can optionally be associated with additional features, e.g., intensity, second return, and so on.”; also see [0021] and [0025]), and motion vectors comprising velocity components that characterize motion of corresponding points at the most recent/current time point (see [0010], [0026], and [0060]-[0062] as discussed in greater detail in the above Response to Arguments section). Although Abdo determines such velocity using displacement between current and preceding positions ([0061]-[0062]), the resulting motion vector expressly represents the movement of the corresponding current point at the current time point. Abdo further teaches object-label data identifying respective three-dimensional regions containing different object ([0050]-[0052]), including potential obstacles ([0011]). Thus, Abdo teaches position and movement state information corresponding to respective objects/obstacles at a first time point.);
generating a local coordinate system corresponding to each of the obstacles based on the position and the movement state corresponding to each of the obstacles, (Abdo teaches object-specific spatial processing in which a preceding pose and a current pose are determined for a given object, with each pose represented by transformation matrix having three-dimensional translation and rotational components ([0067]-[0070]). Abdo’s framework determines an object-specific rigid-body transform from the preceding and current poses ([0071]-[0072]) and applies the transform to points corresponding to the respective object ([0073]). Abdo separately determines corresponding movement information from displacement between current and preceding positions ([0061]-[0062]).
Thus, under the broadest reasonable interpretation, Abdo teaches the claimed generation of a local coordinate system corresponding to each obstacle in so far as Abdo establishes, for each respective object, an object-specific spatial transformation defined from the object’s pose information and used to transform spatial coordinates of points corresponding to that object. The object-specific transformation therefore provides the spatial relationship by which coordinates corresponding to the respective object are transformed relative to that object’s pose. However, Abdo does not expressly teach that the position coordinate of the respective obstacle is the origin of the corresponding local coordinate system (i.e. the amended portion of the limitation).)
transforming the position and the movement state corresponding to each of the obstacles into the local coordinate system of the corresponding obstacle to generate a local position and a local movement state of the corresponding obstacle (Abdo teaches applying the above object-specific rigid-body transform to positions of points corresponding to the respective object ([0071]-[0073]) and determining corresponding motion vectors/velocity information ([0058]-[0062]). Accordingly, Abdo teaches transforming spatial information corresponding to each respective object according to the object-specific spatial relationship discussed above and determining corresponding movement information. Abdo does not expressly teach performing such transformation in an obstacle local coordinate system having the obstacle position coordinate as its origin (i.e. according to the amended previous limitation).)
generating a first [point cloud] corresponding to the obstacles based on the local positions and the local movement states corresponding to the obstacles, wherein the first [point cloud] corresponds to the first time point (As previously discussed, Abdo teaches generating point clouds representations at a current point containing spatial information corresponding to objects ([0021]; [0025]; [0038]-[0040]) and corresponding current motion information ([0010]; [0026]; [0058]-[0062]).)
and inputting the first (Abdo teaches processing point clouds using a scene encoder neural network to generate embeddings (see [0076]-[0082] and [0092] “The system processes…point clouds through an encoder neural network to generate respective embeddings for each of the most recent and earlier point clouds…”), wherein an embedding is “an ordered collection of numerical values, e.g., a vector, a matrix, or higher-dimensional feature map” ([0079]). Examiner interprets an encoded feature embedding representative of scene data to be equivalent to a first scene encoding.), and the first scene encoding is configured to be inputted into a decoder to generate a trajectory prediction corresponding to the obstacles (Abdo teaches processing the embeddings through a decoder to generate flow embeddings and predicted motion vectors (see [0083]-[0087]; [0093]- [0094] “The system processes the respective embeddings for each of the first and second point clouds through a decoder neural network to generate a flow embedding feature map (step 506). The flow embedding feature map includes a respective flow embedding for each grid cell of a spatial grid over the most recent point cloud. The system generates a respective predicted motion vector for each point in the most recent point cloud using the flow embedding feature map.”) and further teaches using the scene-flow outputs to estimate trajectories of objects based on the motion vectors ([0011]; [0032]).
Abdo fails to explicitly disclose wherein a position coordinate of each of the obstacles is an origin of the corresponding local coordinate system and an obstacle tensor generated from the resulting local obstacle information.
In a related art, Gao teaches: generating a local coordinate system corresponding to each of the obstacles wherein a position coordinate of each of the obstacles is an origin of the corresponding local coordinate system (Gao teaches generating trajectory predictions for surrounding agents, wherein a surrounding agent may be “a vehicle, bicycle, pedestrian, ship, drone, or any other moving object in an environment” (Gao [0019]-[0020]). Gao further teaches that scene data characterizing an agent’s observed trajectory can specify its location and can additionally include “the heading of the agent, the velocity of the agent” ([0031]). More specifically, Gao teaches an agent-relative coordinate system, stating “the coordinates in the respective vectors that are generated by the system are in a coordinate system that is relative to the position of the single target agent for which the prediction is being generated” and the system can “can normalize the coordinates of all vectors to be centered around the location of the target agent at the most recent item step at which the location of the target agent was observed” ([0076]). Under the broadest reasonable interpretation, a coordinate system that is defined relative to a particular target agent and centered at the location of that target agent, as taught by Gao in paragraph [0076], constitutes a “local coordinate system corresponding to” that agent, because the coordinate system expresses spatial information relative to the respective agent rather than solely with respect to a common global reference frame. Thus, Gao expressly teaches the obstacle/agent-relative nature of the claimed local coordinate system and teaches that the coordinate system is centered at the current position of the respective obstacle/agent, such that the position coordinate of the respective obstacle/agent reasonably corresponds to the origin of the local coordinate system.);
transforming the position (Gao teaches generating vectors representing an observed trajectory, wherein “The respective vector for each of the time intervals generally includes coordinates… of the position of the agent along the trajectory at a beginning of the time interval and coordinates of the position of the agent along the trajectory at an end of the time interval” ([0068]) and further teaches that these vector coordinates are expressed in the coordinate system “relative to the position of the single target agent for which the prediction is being generated” and may be normalized to be “centered around the location of the target agent at the most recent item step” ([0076]).
Accordingly, Gao expressly teaches transforming/normalizing agent positional information into a respective agent-local coordinate system centered at the current position of the agent.).
Gao further teaches the relationship between the resulting representation, encoding, and trajectory prediction ([0061] “The vectors defining the polylines in the vectorized representation 250 can then be processed by an encoder neural network to generate respective polyline features… [and] generate a trajectory prediction for a given one of the agents… using a trajectory decoder neural network.”), and similarly teaches processing the respective agent polylines using an encoder neural network to generate polyline features and then generating a predicted trajectory for the agent from those features ([0078]-[0080]).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Abdo to represent obstacle information using the agent-relative coordinate system taught by Gao, centered around the respective agent’s current position (Gao [0076]), to predictably normalize the obstacle information relative to the respective obstacle and facilitate efficient and accurate predictive processing, consistent with the computational benefits taught by Gao in paragraphs [0012] and [0055].
Abdo, as modified by Gao, teaches the obstacle position and movement information, an obstacle-local coordinate representation, and encoder/trajectory-prediction processing, but fails to explicitly disclose generating a first obstacle tensor based on the resulting local positions and local movement states and using that obstacle tensor as the input to the scene encoder.
In a related art, Wu’s study teaches: generating a first obstacle tensor corresponding to [location and temporal information] (Wu teaches location prediction methods based on trajectory data, wherein “Trajectory data characterizes the locations and times of moving objects” (p. 109, left column, lines 7-8; also see Abstract). Wu further teaches preference-based prediction using tensor structures (sub-section 4.1.4 found on p. 116 right column through p. 118). Specifically, Wu describes a tensor factorization for multidimensional information, wherein “X is a tensor,” with corresponding user, location, activity, and time matrices, and explains that the method “integrates user, activity, location, and temporal information to predict location by tensor factorization” (sub-section 4.1.4 found on p. 116 right column through p. 117 left column end of first paragraph). Thus, Wu teaches using a multidimensional tensor structure to jointly represent different categories of location and temporal/trajectory related information for predictive processing.).
Abdo, Gao, and Wu are analogous and reasonably pertinent to the problem addressed by the claimed invention because each concerns processing position, movement/trajectory, and temporal information associated with moving objects for predictive purposes. Thus, a person of ordinary skill seeking to represent and process object position and movement information for prediction would have had reason to consider the teachings of the three references together.
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to represent the local obstacle position and movement information of Abdo, as modified by Gao, using the multidimensional tensor representation taught by Wu. Wu teaches using a tensor to jointly represent multiple categories of location and temporal information for prediction (Wu sub-section 4.1.4 found on p. 116 right column through p. 117 left column end of first paragraph). Accordingly, applying Wu’s known tensor representation to the local position and movement state information of the modified Abdo system would have predictably organized the different spatial and temporal/movement attributes of the obstacles into a unified multidimensional representation suitable for subsequent scene encoding and prediction.
Accordingly, Abdo teaches the claimed obstacle position and movement-state information and scene encoder/decoder processing; Gao teaches an agent-relative coordinate system centered at the respective agent’s current position, corresponding under the broadest reasonable interpretation to the claimed local coordinate system, and further teaches trajectory prediction; and Wu teaches a multidimensional tensor representation of location and temporal/trajectory related information. Thus, the combined teachings of the references teach or suggest the limitations of claim 1.
Regarding claim 11 (Currently Amended),
Abdo teaches: A scene encoding generating method, being adapted for use in a scene encoding generating apparatus (Abstract “Methods…and apparatus, including computer programs encoded on computer storage media, for predicting scene flow”; [0097])
The remaining limitations found in claim 11 equally mirror limitations found in claim 1 and are rejected based on the same known prior art (Abdo, Gao, and Wu) and motivations to combine as seen above, in 35 U.S.C. 103 claim 1 rejection. Please refer to claim 1 rejection above for basis of rejection.
Regarding claim 20 (Currently Amended),
Abdo teaches: A non-transitory computer readable storage medium, having a computer program stored therein, wherein the computer program comprises a plurality of codes, the computer program executes a scene encoding generating method after being loaded into an electronic apparatus, the scene encoding generating method comprises ([0097] “Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus.”; Abstract and Abdo teachings found with respect to claim 1 discuss the scene encoding aspects).
The remaining limitations found in claim 20 equally mirror limitations found in claim 1 and 11 and are rejected based on the same known prior art (Abdo, Gao, and Wu) and motivations to combine as seen above, in 35 U.S.C. 103 claim 1 rejection. Please refer to claim 1 rejection above for basis of rejection.
Claims 2 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Abdo (US 20220319054 A1), in view of Gao (US 20210150350 A1), in view of Wu (“Location prediction on trajectory data: A review”; copy provided by examiner), and in further view of Vaswani (“Attention Is All You Need”; copy provided by examiner; hereinafter Vaswani).
Regarding claims 2 and 12 (Original Claims),
Abdo, Gao, and Wu teach: the scene encoding generating apparatus of claim 1, including inputting the first obstacle tensor into a scene encoder to generate a first scene encoding, wherein the first scene encoding corresponds to the first time point corresponding to the first obstacle tensor.
Gao further teaches: an encoder comprising a self-attention neural network having one or more self-attention layers that performs attention calculations on input features to generate output features (Gao [0093]-[0097]), and further teaches temporally related agent information, including agent locations corresponding to a current time step and one or more preceding time steps ([0030]-[0031]), and generating respective vectors for time intervals of an observed agent trajectory ([0065]-[0070]).
However, Abdo, Gao, and Wu fail to explicitly disclose wherein the scene encoder comprises a time attention layer, the time attention layer is configured to perform an attention calculation based on a first input tensor corresponding to the first time point and at least one second input tensor corresponding to at least one second time point to generate a first output tensor.
In a related art, Vaswani teaches: an encoder containing self-attention layers (p. 5, subsection 3.2.3, line 4, “The encoder contains self-attention layers”), wherein “Each position in the encoder can attend to all positions in the previous layer of the encoder” (p. 5, subsection 3.2.3, line 6-7). Vaswani further teaches performing an attention calculation on a set of queries simultaneously, with the queries packed into a matrix Q and the keys and values packed into matrices K and V, to generate a matrix of outputs according to Attention(Q,K,V ) = softmax((
Q
K
T
)/ √
d
k
)V (p. 3, subsection 3.2.1, lines 5-7, “…we compute the attention function on a set of queries simultaneously, packed together into a matrix Q. The keys and values are also packed together into matrices K and V . We compute the matrix of outputs as: Attention(Q,K,V ) = softmax((
Q
K
T
)/ √
d
k
)V”). Vaswani further explains that an attention function maps a query and a set of key-value pairs to an output, wherein the query, keys, calues, and output are vectors (p. 2, section “3 Model Architecture”, lines 2-4, “Here, the encoder maps an input sequence of symbol representations (
x
1
,...,
x
n
) to a sequence of continuous representations z = (
z
1
,...,
z
n
). Given z, the decoder then generates an output sequence (
y
1
,...,
y
m
) of symbols one element at a time.”; “An attention function can be described as mapping a query and a set of key-value pairs to an output, where the query, keys, values, and output are all vectors. The output is computed as a weighted sum of the values…” (p.3, subsection 3.2, lines 1-3). Thus, Vaswani teaches performing an attention calculation in which a representation corresponding to one position of an input sequence attends to representations corresponding to other positions of the sequence to generate an output representation.).
Therefore, it would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to apply the attention technique taught by Vaswani to the input tensors taught by Abdo, as modified by Gao and Wu, corresponding to different time points, such that an attention calculation is performed based on a first input tensor corresponding to a first time point and at least one second input tensor corresponding to at least one second time point to generate an output tensor. Gao already teaches self-attention processing and agent information corresponding to current and preceding time points, and applying Vaswani’s attention technique to the temporally corresponding tensors of the modified Abdo system would predictably enable the scene encoder to account for relationships between obstacle information at different time points while providing the efficient parallel processing taught by Vaswani.
Vaswani is reasonably pertinent to the modified Abdo system because Gao already utilizes self-attention in an encoder for processing agent information used in trajectory prediction.
Claims 9-10, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Abdo (US 20220319054 A1), in view of Gao (US 20210150350 A1), in view of Wu (“Location prediction on trajectory data: A review”; copy provided by examiner), and in further view of Mangalam (US 20210295531 A1).
Regarding claims 9 and 19 (Original Claims),
Abdo, Gao, and Wu teach: The scene encoding generating apparatus of claim 1, including a first encoded feature embedding based on a time point as a first scene encoding and generating the trajectory prediction.
Abdo, Gao, and Wu fail to explicitly disclose: concatenating the first scene encoding corresponding to the first time point and at least one second scene encoding corresponding to at least one second time point to generate an output scene encoding, wherein the output scene encoding is configured to be inputted into the decoder to generate the trajectory prediction corresponding to the obstacles.
In a related art, Mangalam teaches: concatenating the first scene encoding corresponding to the first time point and at least one second scene encoding corresponding to at least one second time point to generate an output scene encoding, wherein the output scene encoding is configured to be inputted into the decoder to generate the trajectory prediction corresponding to the obstacles (Mangalam teaches a system for trajectory prediction based on scene data having a plurality of pedestrians, (Abstract) i.e. obstacles. Mangalam further teaches pedestrians have been observed over a period of time, which is indicative of a first and at least a second time, to create motion history that may be a collection of trajectories indicating the position and the trajectory of the pedestrians, and/or over a previous time ([0023]). Next Mangalam teaches encoding the past trajectory of all pedestrians yields motion history of the one or more pedestrians and the motion history (which represents a first and at least a second time point) are concatenated together with a future endpoint to produce an output parameter to be input into a latent decoder to yield “ground truth endpoints Ĝ.sup.k” ([0040]-[0041]). Mangalam further teaches Ĝ.sup.k (ground truth endpoints) are used in the “future trajectory module” to determine future trajectory points for at least one of the plurality of pedestrians ([0043]).)
Thus, under the broadest reasonable interpretation of the claim, Mangalam teaches concatenating the first scene encoding corresponding to the first time point and at least one second scene encoding corresponding to at least one second time point (e.g. encoded motion history of pedestrians at a first time point and at least a second time point with a future endpoint is concatenated) to generate an output scene encoding (e.g. encoded output parameter), wherein the output scene encoding is configured to be inputted into the decoder (e.g. output parameter is input into a latent decoder) to generate the trajectory prediction corresponding to the obstacles (e.g. the decoder yields ““ground truth endpoints Ĝ.sup.k” which is subsequently used to generate the future trajectory points corresponding to at least one of the plurality of pedestrians.
Therefore, it would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the scene encoding generating apparatus taught by Abdo, previously modified by Gao and Wu, including a trajectory prediction framework, to incorporate the trajectory prediction techniques taught by Mangalam in order to provide a more accurate prediction model relating to pedestrian movement, resulting in improvements to downstream components described by Mangalam, including better path planning and decision-making (Mangalam, [0004]). All of these inventions lie in the same field of endeavor of image analysis for making predictions (e.g. trajectory, scene flow, etc.) based on obstacles (e.g. pedestrians, objects, etc.) and their corresponding time points, movement, and location in a scene.
Regarding claim 10 (Original Claim),
Abdo, Gao, Wu, and Mangalam teach: the scene encoding generating apparatus of claim 9.
Abdo, Gao, and Wu previously taught, in claim 1, “inputting the first obstacle tensor into a scene encoder to generate a first scene encoding”, and “the first obstacle tensor corresponds to the first time point” (refer back to claim 1 for further details of Abdo, Gao, and Wu teachings).
Mangalam previously taught, in claim 9, “at least one second scene encoding corresponding to at least one second time point” by modeling the encoder, decoder, and prediction techniques around motion history across a first and at least a second time point (refer back to claim 9 for further details of Mangalam teachings).
A person of ordinary skill in the art could have modified the teachings of a scene encoding generating apparatus by Abdo, previously modified by Gao, Wu, and Mangalam, that inputs at least one first obstacle tensor corresponding to the at least one first second time point into the scene encoder to generate at least one first scene encoding to incorporate the teachings of at least one second scene encoding corresponding to at least one second time point, taught by Mangalam’s framework that uses motion history of pedestrians across multiple periods of time (e.g. at a first time point and at least a second time point). The combinations would have achieved the predictable result of at least one second scene encoding being generated after inputting at least one second obstacle tensor corresponding to the at least one second time point into the scene encoder, thereby increasing the accuracy of the predictions made by the scene encoding generating by basing predictions on data derived from multiple time points.
Therefore, it would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Abdo, previously modified by Gao, Wu, and Mangalam, to incorporate the teachings of at least one second scene encoding corresponding to at least one second time point taught by Mangalam to yield the predictable result of a more robust and accurate model of predicting trajectory by accounting for obstacles at least two time points. The combination of these known techniques would also predictably result in better path planning and decision-making for vehicles, a goal identified by Mangalam (Mangalam, [0004]). All of these inventions lie in the same field of endeavor of image analysis for making predictions (e.g. trajectory, scene flow, etc.) based on obstacles (e.g. pedestrians, objects, etc.) and their corresponding time points, movement, and location in a scene.
Allowable Subject Matter
Claims 3-8, and 13-18 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form, including all of the limitations of the base claim and any intervening claims.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SAMUEL DAVID BAYNES whose telephone number is (571)272-0607. The examiner can normally be reached Monday - Friday 8:00 am - 5:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Stephen R Koziol can be reached at (408)918-7630. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SDB/
Samuel Baynes
Examiner, Art Unit 2665
/Stephen R Koziol/Supervisory Patent Examiner, Art Unit 2665