DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments and amendments in the Amendment filed May 8, 2026 (herein “Amendment”), with respect to the objections to claims 1–20 have been fully considered and are persuasive. The objections to claims 1–20 have been withdrawn.
Applicant’s arguments and amendments in the Amendment with respect to the rejection of claims 4, 6, 9, 13, 15, 18 and claims depending therefrom under 35 U.S.C. 112(b) have been
fully considered and are persuasive. The rejection of claims 4, 6, 9, 13, 15, 18 and claims depending therefrom under 35 U.S.C. 112(b) has been withdrawn.
Applicant’s Terminal Disclaimer filed and approved on May 8, 2026 has obviated the double patenting rejection set forth against co-pending application 18/421,539 and therefore the double patenting rejection is withdrawn.
Applicant’s arguments and amendments in the Amendment with respect to the rejections of claims 1, 10 and 19, on the grounds of nonstatutory double patenting as being unpatentable over claims 1, 7 and 13 of co-pending Application no. 19/314,642 have been fully considered and are persuasive in part. The newly amended limitations in claims 1, 10 and 19 are not recited in 19/314,642. However, in view of Hori et al., “Attention-Based Multimodal Fusion for Video Description,” arXiv:1701.03126v2 [cs.CV] 9 Mar 2017,
https://doi.org/10.48550/arXiv.1701.03126 (herein “Hori”), the new limitations do not patentably distinguish from the claims of 1, 7 and 13 and therefore, the double patenting rejection is re-issued, with an additional reliance upon Hori.
Applicant’s arguments and amendments in the Amendment with respect to the rejections of claims 1, 10 and 19, and various claims depending therefrom under 35 U.S.C. 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, new grounds of rejection are made in view of Hori et al., “Attention-Based Multimodal Fusion for Video Description,” arXiv:1701.03126v2 [cs.CV] 9 Mar 2017, https://doi.org/10.48550/arXiv.1701.03126.
Claim Objections
Claims 9 and 18 are objected to because of the following informalities: the limitation “the one or more trajectories sequences” should instead recite “the one or more trajectory sequences.” Appropriate correction is required.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or
improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg,
140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d
2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van
Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619
(CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely
online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1, 10 and 19 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 7 and 13 of copending Application No. 19/314,642 (herein “‘642 application”) in view of Salzmann et al., “Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data,” arXiv:2001.03093v5 [cs.RO], January 13, 2021, https://doi.org/10.48550/arXiv.2001.03093 (herein “Salzmann”), further in view of Hori et al., “Attention-Based Multimodal Fusion for Video Description,” arXiv:1701.03126v2 [cs.CV] 9 Mar 2017, https://doi.org/10.48550/arXiv.1701.03126 (herein “Hori”).
This is a provisional nonstatutory double patenting rejection because the patentably indistinct claims have not in fact been patented. Although the claims at issue are not identical, they are not patentably distinct from each other because the claims of the ‘642 application recite most of the limitations of the present application with correspondence to the claims is set forth below:
Regarding claims 1, 10 and 19, claims 1, 7 and 13 of the ‘642 application correspond as follows, with deficiencies of claims 1, 7 and 13 of the ‘642 application noted in curly brackets {}:
Claims 1, 10 and 19 of the present
application
Claims 1, 7 and 13 of the ‘642 application
[claim 1 only: A computer implemented
method for tracking one or more individuals
[Claim 1 only: A computer implemented
method for tracking one or more individuals
during a sporting event, the method
comprising:]
during a sporting event, the method
comprising:]
[claim 10 only: A system for tracking one or more individuals during a sporting event, the
system comprising:]
[Claim 7 only: A system for tracking one or more individuals during a sporting event, the
system comprising:]
[Claim 19 only: A non-transitory computer readable medium configured to store processor-readable instructions, wherein when executed by a processor, the
instructions perform operations comprising:]
[Claim 13 only: A non-transitory computer readable medium configured to store processor-readable instructions, wherein when executed by a processor, the
instructions perform operations comprising:]
receiving, as an input, {geospatial} data of a
sporting event;
receiving, as an input, broadcast tracking data
of a sporting event
receiving, as an input, labeled event data
based on sports broadcast footage of the sporting event;
receiving, as an input, … labeled event data of the sporting event;
performing multi-object tracking of one or more agents of the geospatial data to
determine one or more vectors;
performing multi-object tracking of one or more agents of the broadcast tracking data to
determine one or more vectors;
inputting the labeled event data and the one or more vectors into a diffusion model, wherein the diffusion [model includes: an event
encoder; and a tracking decoder];
inputting the labeled event data and one or more vectors into a diffusion model;
and determining, using the diffusion model, one or more trajectory sequences for the one or more agents [by: determining, by the event decoder, event embeddings by embedding the labeled event data; and applying, by the tracking decoder, attention to embed and fuse the one or more vectors with the event
embeddings].
determining, using the diffusion model, one or more trajectory sequences for the one or more agents;
Claims 1, 7 and 13 of the ‘642 application, while reciting “broadcast tracking data of a sporting event” does not recite the data specifically as being “geospatial data” however, Salzmann teaches “geospatial data” on page 7, LIDAR data incorporated into the event/agent representation vectors.
Further, claims 1, 7 and 13 of the ‘642 application do not recite, where Hori teaches wherein the model includes: an event encoder; and a tracking decoder (Hori figs. 3 and 4, pages 3–5, attention based multimodal fusion model including an encoder for feature extraction (bottom two squared sections) and a decoder network comprised of the attention-based multimodal fusion, where page 5, section 5.2 teaches the data being decoded as annotated sports video data of spatiotemporal activity, thus the decoding being a tracking of the sports video) and by: determining, by the event encoder, event embeddings by embedding the labeled event data (Hori fig. 4, pages 3 and 5, input sequence of feature vectors obtained using an encoder which sections 5.1 and 5.2 teaches extracts visual features from a Youtube2Text video corpus comprised of annotated video); and applying, by the tracking decoder, attention to embed and fuse the one or more vectors with the event embeddings (Hori fig. 4, pages 4–5,
decoder performing attention-based multimodal fusion on different modality embedding vectors, including both annotated video and other modality vectors, where page 2 left column teaches the additional modality as motion features derived from the video).
Therefore, taking the teachings of claims 1, 7 and 13 of the ‘642 application and Salzmann together as a whole, it would have been obvious to a “PHOSITA” before the effective filing date of the claimed invention to have modified claims 1, 7 and 13 of the ‘642 application to have the geographic data teachings cited above as disclosed by Salzmann at least because doing so would allow for modeling agents with different perception ranges. See Salzmann, bottom of page 5.
Further, taking the teachings of claims 1, 7 and 13 of the ‘642 application and Hori as a whole, it would have been obvious to a PHOSITA before the effective filing date of the claimed invention to have modified claims 1, 7 and 13 of the ‘642 application to have the encoder-decoder network with attention-based multimodal fusion as disclosed by Hori at least because doing so would provide a robustness strategy for models producing summary output for multimodal input. See Hori page 2, top right column.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C.
102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35
U.S.C. 103 are summarized as follows:
Determining the scope and contents of the prior art.
Ascertaining the differences between the prior art and the claims at issue.
Resolving the level of ordinary skill in the pertinent art.
Considering objective evidence present in the application indicating obviousness or nonobviousness.
3. Claims 1–4, 6, 8, 10–13, 15, 17 and 19–20 are rejected under 35 U.S.C. 103 as being unpatentable over Salzmann et al., “Trajectron++: Dynamically-Feasible Trajectory
Forecasting With Heterogeneous Data,” arXiv:2001.03093v5 [cs.RO], January 13, 2021, https://doi.org/10.48550/arXiv.2001.03093 (herein “Salzmann” – an earlier version of this
reference was cited by Applicant in the IDS filed 9/18/2024), in view of Gu et al., "Stochastic Trajectory Prediction via Motion Indeterminacy Diffusion," 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 2022, pp. 17092-17101, doi: 10.1109/CVPR52688.2022.01660 (herein “Gu”) in view of
Zhu et al., "Event Tactic Analysis Based on Broadcast Sports Video," in IEEE Transactions on Multimedia, vol. 11, no. 1, pp. 49-67, Jan. 2009, doi: 10.1109/TMM.2008.2008918 (herein “Zhu”), further in view of Hori et al., “Attention-
Based Multimodal Fusion for Video Description,” arXiv:1701.03126v2 [cs.CV] 9 Mar 2017,
https://doi.org/10.48550/arXiv.1701.03126 (herein “Hori”).
Regarding claims 1, 10 and 19, with claim 1 as exemplary, substantive differences between the claims noted in curly brackets {}, and deficiencies of Salzmann noted in
square brackets [], Salzmann teaches {A computer implemented method for tracking one or more agents during a [sporting] event, the method comprising: - claim 1 / A system for tracking one or more agents during a [sporting] event, the system comprising: a memory configured to store processor-readable instructions; and a processor operatively connected to the memory, and configured to execute the instructions to perform operations comprising: - claim 10 / A non-transitory computer readable medium configured to store processor-readable instructions, wherein when executed by a processor, the instructions perform operations comprising: - claim 19}(Salzmann abstract, pages 5 and 9, Trajectron++ model for predicting trajectories based on tracking pedestrians in a street scene, is implemented in PyTorch on a computer running Ubuntu (instructions), the computer containing a CPU and GPUs, where CPUs and GPUs are understood to have computer readable memory for executing processes)
receiving, as an input, geospatial data of a [sporting] event (Salzmann page 7, first full paragraph, page 2, additional information of LIDAR data (geospatial data) is included (received input) into the Trajectron++ model framework and encoded as a vector to be added to the backbone of representation vectors ex, the LIDAR data being of the environment in which the dynamics (events) of the agents is taking place);
receiving, as an input, labeled event data based on [sports broadcast footage] of the [sporting] event (Salzmann page 5, section 4, fig. 2, agents are classified (car, bus, pedestrian), thus labeled and are nodes in a created (receiving) spatiotemporal graph, along with edges that represent interactions (events));
performing multi-object tracking of one or more agents of the received geospatial data to determine one or more vectors (Salzmann page 6, Modeling Agent History and Encoding Agent Interactions sections, each agent (multi-object) has its current state and its history encoded (tracking), the agents being “of the received geospatial data” since the agents are tracked in the environment the LIDAR data represents, and encodings are made from the agent interactions to produce a single node representation vector ex (vector));
inputting the labeled event data and the one or more vectors into a [diffusion] model, [wherein the diffusion model includes: an event encoder; and a tracking decoder] (Salzmann page 7, Producing Dynamically-Feasible Trajectories section, the backbone representation ex is fed (inputting) into a decoder (model)); and
determining, using the [diffusion] model, one or more trajectory sequences for the one or more agents (Salzmann page 7, Producing Dynamically-Feasible Trajectories section, the decoder produces trajectories in position space) [by: determining, by the event encoder, event embeddings by embedding the labeled event data; and applying, by the tracking decoder, attention to embed and fuse the one or more vectors with the event embeddings].
While Salzmann teaches receiving data tracking pedestrians in a street scene, nonetheless, Salzmann does not explicitly teach “sporting event.”
Further, Salzmann does not teach “sports broadcast footage” or that its model is a “diffusion model.”
Gu teaches a diffusion model (Gu page 17094, section 3.2, trajectories are determined by way of a diffusion process).
Gu further teaches wherein the diffusion model includes: an event encoder (Gu page 17094, fig. 2, section 3.1, temporal-social encoder); and a tracking decoder (Gu page 17094,
fig. 2, and section 3.2, a decoder takes yk trajectories, and decodes using a Markov chain and Gaussian transitions); and by: determining, by the event encoder, event embeddings by embedding the labeled event data (Gu page 17094, fig. 2, section 3.1, temporal-social encoder maps a history path and social interaction into a state embedding, where different pedestrians are labeled by number as shown).
Zhu teaches “sporting event” and “sports broadcast footage” (Zhu fig. 1, page 52, section III, attack event extraction performed on input broadcast soccer video (sports broadcast footage) of a soccer game (sporting event)).
Hori teaches and applying, by the tracking decoder, attention to embed and fuse the one or more vectors with the event embeddings (Hori fig. 4, pages 4–5, decoder performing attention-based multimodal fusion on different modality embedding vectors, including both annotated video and other modality vectors, where page 2 left column teaches the additional modality as motion features derived from the video).
Therefore, taking the teachings of Salzmann and Gu together as a whole, it would have been obvious to a “PHOSITA” before the effective filing date of the claimed invention to have modified the model of Salzmann to include diffusion and thus be a diffusion model as disclosed by Gu at least because doing so would allow for predicting trajectories with a flexible indeterminacy that is capable of adapting to dynamic environment. See Gu page 17093, left column.
Further, taking the teachings of Salzmann as modified by Gu and Zhu together as a whole, it would have been obvious to a person having ordinary skill in the art (herein “PHOSITA”) before the effective filing date of the claimed invention to have modified the pedestrian event data and LIDAR data of Salzmann to be sporting event data and based on sports
broadcast data of the sporting event as disclosed by Zhu at least because doing so would allow for discovering tactic patterns amongst professional athletes, in order to improve team performance during a game. See Zhu abstract, section I.
Still further, taking the teachings of Salzmann as modified above and Hori as a whole, it would have been obvious to a PHOSITA before the effective filing date of the claimed invention to have modified the tracking processing of Salzmann to have the decoder network with
attention-based multimodal fusion as disclosed by Hori at least because doing so would provide a robustness strategy for models producing summary output for multimodal input. See Hori page 2, top right column.
Regarding claims 2, 11 and 20, with claim 2 as exemplary, Salzmann does not explicitly teach, but Zhu teaches wherein the labeled event data includes a sequential stream of one or more major events throughout a sport event, the major events including at least one of a pass, shot, tackle, foul, turnover, penalty, goal, score, or substitution from the sporting event (Zhu pages 52–53, section III A, fig. 2, web-casting text describing the event that has happened in a game with a timestamp and brief description such as “goal” or “headed in” for a scoring event, also called a “shot”), and wherein the geospatial data includes one or more of the sports broadcast footage, in-venue footage, global positioning system (GPS) data, near field communication (NFC) data, or radio-frequency identification (RFID) data (Zhu page 53, section B, broadcast soccer videos are analyzed and input into the analysis model).
Further, taking the teachings of Salzmann as modified by Gu and Zhu together as a whole, it would have been obvious to a person having ordinary skill in the art (herein “PHOSITA”) before the effective filing date of the claimed invention to have modified the pedestrian event data and LIDAR data of Salzmann to be major events in a sports game like a
goal or score or shot, and broadcast soccer videos as disclosed by Zhu at least because doing so would allow for discovering tactic patterns amongst professional athletes, in order to improve team performance during a game. See Zhu abstract, section I.
Regarding claims 3 and 12, with claim 3 as exemplary, Salzmann teaches wherein the one or more vector includes at least one of an agent two dimensional coordinates on a sporting event's field, an agent position, an agent team, an indicator indicating the agent is a ball, or player visibility information (given that the claim only requires “at least one,” Salzmann teaches on page 5, the spatiotemporal graph which is converted to the representation vector ex includes 2D world positions of agents (an agent position)).
Regarding claims 4 and 13, with claim 4 as exemplary, Salzmann teaches wherein the labeled event data is represented as a two dimensional spatiotemporal grid, the grid representing a stacking of each of the one or more of the labeled events of the agents (Salzmann page 5,
section 4, fig. 2, a scene including agents as nodes and edges as the agents’ interactions (events) is abstracted as a spatiotemporal graph (grid), arranging in a graph pattern (stacking) the agents and their interactions).
Regarding claims 6 and 15, with claim 6 as exemplary, Salzmann does not explicitly teach but Gu teaches wherein the event encoder encodes the labeled event data (Gu page 17094, fig. 2, section 3.1, temporal-social encoder maps a history path and social interaction into a state embedding, where different pedestrians are labeled by number as shown) and the tracking decoder conditionally decodes the one or more trajectory sequences (Gu page 17094, fig. 2, and section 3.2, a decoder takes yk trajectories, and decodes using a Markov chain and Gaussian transitions (conditionally)).
Therefore, taking the teachings of Salzmann and Gu together as a whole, it would have been obvious to a “PHOSITA” before the effective filing date of the claimed invention to have modified the model of Salzmann to include diffusion with an encoder and decoder as disclosed by Gu at least because doing so would allow for predicting trajectories with a flexible indeterminacy that is capable of adapting to dynamic environment. See Gu page 17093, left column.
Regarding claims 8 and 17, with deficiencies of Salzmann noted in square brackets [], Salzmann teaches further including: generating the one or more trajectory sequences (Salzmann page 7, Producing Dynamically-Feasible Trajectories section, the decoder produces trajectories in position space) based on the [fused] one or more vectors (Salzmann page 6, Modeling Agent History and Encoding Agent Interactions sections, each agent (multi-object) has its current state and its history encoded (tracking), the agents being “of the received
geospatial data” since the agents are tracked in the environment the LIDAR data represents, and encodings are made from the agent interactions to produce a single node representation vector ex (vector), where page 7 teaches the vector ex is fed into the decoder to obtain the trajectories) [and event embeddings.]
Salzmann does not explicitly teach where Gu teaches based on event embeddings (Gu page 17094, fig. 2, and section 3.2, decoder on right side of fig. 2 receives as an input the output of the encoder which is interaction embeddings (event embeddings) and the decoder outputs trajectories).
Salzmann does not explicitly teach where Hori teaches based on fused vectors (Hori fig.
4, pages 4–5, decoder performing attention-based multimodal fusion on different modality
embedding vectors, including both annotated video and other modality vectors, where page 2 left column teaches the additional modality as motion features derived from the video).
Therefore, taking the teachings of Salzmann and Gu together as a whole, it would have been obvious to a “PHOSITA” before the effective filing date of the claimed invention to have modified the model of Salzmann to include trajectories based on event embeddings as disclosed by Gu at least because doing so would allow for predicting trajectories with a flexible indeterminacy that is capable of adapting to dynamic environment. See Gu page 17093, left column.
Further, taking the teachings of Salzmann as modified above and Hori as a whole, it would have been obvious to a PHOSITA before the effective filing date of the claimed invention to have modified the tracking processing of Salzmann to have the decoder network with
attention-based multimodal fusion as disclosed by Hori at least because doing so would provide a robustness strategy for models producing summary output for multimodal input. See Hori page 2, top right column.
Claims 5 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Salzmann in view of Gu in view of Zhu in view of Hori, as set forth above regarding claims 1 and 10 from which claims 5 and 14 respectively depend, further in view of Alcorn et al., "baller2vec: A Multi-Entity Transformer For Multi-Agent Spatiotemporal
Modeling," arXiv:2102.03291v3, September 28, 2021, https://doi.org/10.48550/arXiv.2102.03291 (herein “Alcorn”).
Regarding claims 5 and 14, with claim 5 as exemplary, Salzmann does not explicitly teach, but Alcorn teaches wherein the diffusion model applies spatiotemporal axial attention on the labeled event data and the one or more vectors (Alcorn page 2, fig. 1, transformer model
applies self-attention, given the locations of the players (spatio) through time (temporal) using a self-attention mask tensor (multiple dimensional/axial) on multi-entity (one or more vectors) sequential data (event data)), where self-attention is applied across temporal and spatial axis, separately (Alcorn pages 3–4, section 2.3 teaches the self-attention mask as a tensor (multidimensional/separately) applied to an agent value at a time step (temporal and spatial axis)).
Therefore, taking the teachings of Salzmann as modified by Gu and Zhu and Alcorn
together as a whole, it would have been obvious to a “PHOSITA” before the effective filing date of the claimed invention to have modified to have the model of Salzmann apply self-attention across temporal and spatial axis teachings cited above as disclosed by Alcorn at least because doing so would allow for capturing idiosyncratic qualities of players, thus providing a more detailed analysis of game play for a particular sport. See Alcorn, bottom of page 2.
Claims 9 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Salzmann in view of Gu in view of Zhu in view of Hori, as set forth above regarding claims 6 and 15 from which claims 9 and 18 respectively depend, further in view of Lohit et al., "Recovering Trajectories of Unmarked Joints in 3D Human Actions Using Latent Space Optimization," 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 2021, pp. 2341-2350, doi: 10.1109/WACV48630.2021.00239
(herein “Lohit”).
Regarding claims 9 and 18, with claim 9 as exemplary, and with deficiencies of Salzmann noted in square brackets [], Salzmann teaches wherein the [diffusion] model further includes a second tracking decoder (Salzmann fig. 2, page 5, the decoder as shown consisting
of two branches, where each branch is considered its own decoder, therefore including a second decoder, and where each decodes tracking data).
Salzman as modified above does not explicitly teach where Gu teaches a diffusion model (Gu page 17094, section 3.2, trajectories are determined by way of a diffusion process).
Salzmann as modified above does not explicitly teach, but Lohit teaches a transpose temporal convolution, the transpose temporal convolution being configured to expand the one or more trajectory sequences to their initial temporal dimensionality (Lohit page 2344, section 4, temporal convolutional autoencoder with transposed convolutional layers that minimizes the loss between an input sequence (initial temporal dimensionality) and the output of the decoder, also the autoencoder is trained on only a subset of joints observed (subset of dimensionality) such that, as disclosed in section 5, the incomplete action sequence (trajectory sequences) is projected to the range space of the generator which is the complete set of human action sequences, and thus an expansion to an initial dimensionality in time).
Therefore, taking the teachings of Salzmann and Gu together as a whole, it would have been obvious to a “PHOSITA” before the effective filing date of the claimed invention to have modified the model of Salzmann to include diffusion and thus be a diffusion model as disclosed by Gu at least because doing so would allow for predicting trajectories with a flexible indeterminacy that is capable of adapting to dynamic environment. See Gu page 17093, left column.
Further taking the teachings of Salzmann as modified above and Lohit together as a whole, it would have been obvious to a “PHOSITA” before the effective filing date of the claimed invention to have modified the trajectory processing of Salzmann with the transpose temporal convolution operations disclosed in Lohit at least because doing so would provide
robust activity analysis by recovering actions and dynamics of missing joints in human movement. Lohit Abstract.
Allowable Subject Matter
Claim 7 and 16 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims, and that any claim objections and the obviousness non-statutory double patenting issues given above are addressed.
Specifically, claims 7 and 16 recite “tokenizing the labeled event data using a linear projection; applying sinusoidal positional embeddings to specify temporal occurrences of the labeled event data,” which is not taught or suggested in Salzmann, Gu, Zhu, Hori or Alcorn. Further, additionally cited, but not used in any of the above rejections, Strudel et al., US PgPub No. US 2024/0119261, while setting forth using a diffusion model that uses a sequence of discrete tokens, and applies a linear projection to convert tokens, does not teach that the linear projection is used for the tokenizing, much less tokenizing the labeled event data, as claimed. Further, none of Salzmann, Gu, Zhu, Hori, Alcorn, or Strudel, or any of the other cited art of record, whether considered alone or in an obvious combination to a PHOSITA teaches or suggests “applying sinusoidal positional embeddings to specify temporal occurrences of the labeled event data,” and all other limitations from claims 7 and 16, and thus, these claims are allowable over the cited art of record.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHELLE M KOETH whose telephone number is (571)272-5908. The examiner can normally be reached Monday-Thursday, 09:00-17:00, Friday 09:00-13:00, EDT/EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent Rudolph can be reached at 571-272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
MICHELLE M. KOETH
Primary Examiner
Art Unit 2671
/MICHELLE M KOETH/Primary Examiner, Art Unit 2671