Prosecution Insights
Last updated: October 02, 2026
Application No. 18/318,524

TRANSFORMER MODEL FOR JOURNEY SIMULATION AND PREDICTION

Final Rejection §103
Filed
May 16, 2023
Examiner
LAI, DYLAN HONG
Art Unit
2144
Tech Center
2100 — Computer Architecture & Software
Assignee
Adobe Inc.
OA Round
2 (Final)
Grant Probability
Favorable
3-4
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-55.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
13 currently pending
Career history
13
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim(s) 1-4, 6-7, 9-10, and 15-16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Agent models of customer journeys on retail high streets by Paul M. Torrens, hereafter Torrens, in view of US 20240169396 A1 by Zhao et al. hereafter Zhao. Regarding claim 1, Torrens teaches: A computer-implemented method comprising: obtaining, by a model training engine (system that results in a map), training data from a plurality of journeys (movement traces) ((Torrens) pg. 108-109, Section 6.10, Paragraph 1, lines 2-6, “Generally, such systems use closed circuit television (CCTV) cameras to identify and track individual customers as they move around within a store. The end result is either a “heatmap” or a map of movement traces. These traces essentially illustrate how customers traverse the store by steering as a response to path-planning, locomotion, and interactions with staff and products.”.), each journey comprising a sequence (at each timestep) of events (choice-points/interactions with staff and products), wherein the training data from each journey indicates corresponding customer (discrete agent) interaction (discrete state)with each event in the sequence of events of the journey; ((Torrens) pg. 103, Section 6.4, paragraph 2, lines 10-15, “Attention to fine-scale detail would imply that customers could be exposed to a huge number of choice-points while on the customer journey, and a deviation from inertia would suggest that the transition probabilities for agent states relative to those choice-points would need to be assessed at each timestep in the model (from a discrete span of time t → t + 1), for each discrete state, for each discrete agent.”) generating, […] prior to deploying the input journey(Until an actual customer begins to interact), simulated customer interaction (simulation scenarios) with the input journey ((Torrens) pg. 98, Section 6, paragraph 1, lines 9-24, “… a common gap in retailers’ control is the customer journey: a store may realign its layout, but until an actual customer begins to interact with that new design, the retailer cannot reasonably assess what implications result. This is where simulation scenarios based on agent-based models can come into particular usefulness, in allowing the retailer to experiment with synthetic versions of the customers that they seek to influence…” ‘Until an actual customer begins to interact’ indicates that the simulation is before that so is prior to deploying an actual customer journey.) by predicting a probability of customers triggering (transition probabilities) each event of the input sequence of the plurality of events (huge number of touchpoints) of the input journey […] ((Torrens) pg. 103, Section 6.4, paragraph 2, lines 10-15, “Attention to fine-scale detail would imply that customers could be exposed to a huge number of choice-points while on the customer journey, and a deviation from inertia would suggest that the transition probabilities for agent states relative to those choice-points would need to be assessed at each timestep in the model (from a discrete span of time t → t + 1), for each discrete state, for each discrete agent.” A huge number implies there is a plurality.) Torrens does not teach: training, by the model training engine using the training data, a machine learning model to simulate customer interaction with an input journey comprising an input sequence of a plurality of events; predicting … based at least on a positional embedding determined for each event of the input sequence of the plurality of events; and causing, by a user interface component, display of results of the simulated customer interaction with the input journey. Zhao teaches: Training a machine learning model using a model training engine (model training function) to simulate (generate an inference) customer behavior (user behavior information) ((Zhao) Paragraph [0051-0052] “The model training function 209 trains a machine learning model and provides model data to a behavior model registration function 211 … An inference data generation function 214 is based on inputs regarding those features from the target feature engineering function 212, as well as the machine learning model defined by the behavior model registration function 211, and user behavior information from the training data registration function 210, and generates an inference 215); the machine learning model (transformer layers) doing the predicting being based on at least a positional embedding (influenced by a positional embedding) determined for each event of the input sequence of the plurality of events; ((Zhao) (Paragraph [0069] “The data on which the transformer layers 309a-309d operate is influenced by a positional embedding 312.”) and causing, by a user interface component (display), display of results of the simulated customer interaction with the input journey; ((Zhao) Paragraph [0036] “According to embodiments of this disclosure, an electronic device 101 is included in the network configuration 100. The electronic device 101 can include at least one of a bus 110, a processor 120, a memory 130, an input/output (I/O) interface 150, a display 160, a communication interface 170, or a sensor 180.” A display can cause display of results.) Zhao and Torrens are analogous art because they are in the same field of invention: predicting customer behavior. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined training a machine learning model, the machine learning model using positional embedding, and displaying results on a user interface component as taught by Zhao, with the simulation of customer interactions by predicting the probability of customers triggering events as Torrens teaches. The motivation for the combination would have been to take advantage of machine learnings ability to discover patterns that ordinary human analysis might miss, while keeping information about the ordering of events by using positional embedding, and to have a way to view results for review. This application of the technique of using a machine learning model using positional embedding to simulate customer interaction and addition of a display user interface component would yield the predictable result that is the invention specified in claim 1 of the instant application. The Examiner notes that this motivation applies to all dependent and/or otherwise subsequently addressed claims. Regarding claim 2, Torrens, in view of Zhao, teaches the material disclosed in claim 1, and Torrens additionally teaches: wherein the results of the simulated customer interaction with the input journey (simulation scenarios) comprises the probability of the customers triggering (transition probabilities) each event of the input sequence of the plurality of events of the input journey (huge number of choice-points). ((Torrens) pg. 103, Section 6.4, paragraph 2, lines 10-15, “Attention to fine-scale detail would imply that customers could be exposed to a huge number of choice-points while on the customer journey, and a deviation from inertia would suggest that the transition probabilities for agent states relative to those choice-points would need to be assessed at each timestep in the model (from a discrete span of time t → t + 1), for each discrete state, for each discrete agent.”.) Regarding claim 3, Torrens, in view of Zhao, teaches the material disclosed in claim 1, and Torrens additionally teaches: wherein the results of the simulated customer interaction (simulation scenarios) with the input journey comprises a prediction of one or more predicted events (Estimates of the likely flow) based on the input journey (potential paths). ((Torrens) pg. 107, Section 6.7, Paragraph 2, lines 6-10, “Estimates of the likely flow of pedestrians between these locations can be built from input–output models or from spatial interaction models, leaving potential paths that they may have traversed open to estimation by path-planning heuristics”) Regarding claim 4, Torrens, in view of Zhao, teaches the material disclosed in claim 1, and Zhao additionally teaches: wherein the machine learning model (deep learning model) is a transformer-based (transformer structure) machine learning model ((Zhao) Paragraph [0034] “A deep learning model structure may use a transformer structure”) Zhao and Torrens are analogous art because they are in the same field of invention: predicting customer behavior. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined a machine learning model that uses a transformer structure as taught by Zhao, with the simulation of customer interactions as Torrens teaches. The motivation for the combination would have been to take advantage of machine learnings ability to discover patterns that ordinary human analysis might miss, as well as a transformer structure to keep track of ordering during machine learning. This application of the technique of using a machine learning model using a transformer structure to simulate customer interaction would yield the predictable result that is the invention specified in claim 4 of the instant application. Regarding claim 6, Torrens, in view of Zhao, teaches the material disclosed in claim 1, and Zhao additionally teaches: obtaining customer data for each customer of the plurality of journeys of the training data, wherein the customer data comprises demographic data (geographic data); ((Zhao) Paragraph [0071] “In some cases, the feature data used in the present disclosure includes the following information” and diagram below paragraph [0071].) and further training the machine learning model to simulate (generate an inference) the customer interaction (user behavior) with the input journey using the customer data (user behavior information and feature data). ((Zhao) Paragraph [0051-0052]) Zhao and Torrens are analogous art because they are in the same field of invention: predicting customer behavior. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined training a machine learning model that uses geographic data as part of its dataset as taught by Zhao, with the simulation of customer interactions as Torrens teaches. The motivation for the combination would have been to take advantage of machine learnings ability to discover patterns that ordinary human analysis might miss, while geographic data is used as datapoints that are helpful due to being unchanged by typical user behavior. This application of the technique of using a machine learning model trained using a dataset that includes geographic data to simulate customer interaction would yield the predictable result that is the invention specified in claim 6 of the instant application. Regarding claim 7, Torrens, in view of Zhao, teaches the material disclosed in claim 1, and Zhao additionally teaches: wherein the training data further comprises data indicating a time of customer interactions (time of day) with each event in the sequence of events of each journey of the plurality of journeys; ((Zhao) Paragraph [0071] and diagram below paragraph [0071].) and further training the machine learning model to simulate (generate an inference) the customer interaction (user behavior) with the input journey using the time of the customer interactions (time of day). ((Zhao) Paragraph [0051-0052]) Zhao and Torrens are analogous art because they are in the same field of invention: predicting customer behavior. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined training a machine learning model that uses time of interaction data as part of its dataset as taught by Zhao, with the simulation of customer interactions as Torrens teaches. The motivation for the combination would have been to take advantage of machine learnings ability to discover patterns that ordinary human analysis might miss, while time of interaction data is used as datapoints that are helpful due to being relatively unaffected by typical user behavior. This application of the technique of using a machine learning model trained using a dataset that includes time of interaction data to simulate customer interaction would yield the predictable result that is the invention specified in claim 7 of the instant application. Regarding claim 9, Torrens teaches: obtaining a journey (map of movements traces), the journey comprising a sequence of a plurality of events (interactions with staff and products); ((Torrens) pg. 108-109, Section 6.10, Paragraph 1, lines 2-6, “Generally, such systems use closed circuit television (CCTV) cameras to identify and track individual customers as they move around within a store. The end result is either a “heatmap” or a map of movement traces. These traces essentially illustrate how customers traverse the store by steering as a response to path-planning, locomotion, and interactions with staff and products.”.) predicting , […] prior to deploying the journey(Until an actual customer begins to interact), a predicted event (Estimates of the likely flow) ((Torrens) pg. 107, Section 6.7, Paragraph 2, lines 6-10) based on the sequence of the plurality of events of the journey by predicting a probability of customers triggering (transition probabilities) each event of the sequence of the plurality of events (huge number of touchpoints) of the journey […] ((Torrens) pg. 103, Section 6.4, paragraph 2, lines 10-15, “Attention to fine-scale detail would imply that customers could be exposed to a huge number of choice-points while on the customer journey, and a deviation from inertia would suggest that the transition probabilities for agent states relative to those choice-points would need to be assessed at each timestep in the model (from a discrete span of time t → t + 1), for each discrete state, for each discrete agent.” A huge number implies there is a plurality.) Torrens does not teach: A non-transitory computer-readable medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising: …by a machine learning model… …based at least on positional embeddings determined for each event of the sequence of the plurality of events; and causing display of the predicted event. Zhao teaches: A non-transitory computer-readable medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising: ((Zhao) Paragraph [0006], “In a third embodiment, a non-transitory computer readable medium contains instructions that when executed cause…”) the machine learning model (transformer layers) doing the predicting being based on at least a positional embedding (influenced by a positional embedding) determined for each event of the input sequence of the plurality of events; ((Zhao) (Paragraph [0069] “The data on which the transformer layers 309a-309d operate is influenced by a positional embedding 312.”) and causing display of the predicted event; ((Zhao) Paragraph [0036] “According to embodiments of this disclosure, an electronic device 101 is included in the network configuration 100. The electronic device 101 can include at least one of a bus 110, a processor 120, a memory 130, an input/output (I/O) interface 150, a display 160, a communication interface 170, or a sensor 180.” A display can cause display predicted results.) Zhao and Torrens are analogous art because they are in the same field of invention: predicting customer behavior. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined a non-transitory computer readable medium storing instructions, predicting using a machine learning model that uses positional embedding, and displaying results on a user interface component as taught by Zhao, with predicting of a predicted event by predicting the probability of customers triggering events as Torrens teaches. The motivation for the combinations would have been to have a physical medium to run the instruction son, and to take advantage of machine learnings ability to discover patterns that ordinary human analysis might miss, while keeping information about the ordering of events by using positional embedding, and to have a way to view results for review. This application of the technique of using a machine learning model using positional embedding to predict a predicted event, and addition of a non-transitory computer readable medium and display user interface component would yield the predictable result that is the invention specified in claim 9 of the instant application. The Examiner notes that this motivation applies to all dependent and/or otherwise subsequently addressed claims. Regarding claim 10, Torrens, in view of Zhao, teaches the material disclosed in claim 1, and Zhao additionally teaches: wherein the machine learning model (deep learning model) is a transformer-based (transformer structure) machine learning model ((Zhao) Paragraph [0034] “A deep learning model structure may use a transformer structure”) Zhao and Torrens are analogous art because they are in the same field of invention: predicting customer behavior. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined a machine learning model that uses a transformer structure as taught by Zhao, with predicting a predicted event as Torrens teaches. The motivation for the combination would have been to take advantage of machine learnings ability to discover patterns that ordinary human analysis might miss, as well as a transformer structure to keep track of ordering during machine learning. This application of the technique of using a machine learning model using a transformer structure to predict a predicted event would yield the predictable result that is the invention specified in claim 10 of the instant application. Regarding claim 15, Torrens teaches: obtaining… a journey (movement traces), the journey comprising a sequence of a plurality of events (interaction with staff and products); ((Torrens) pg. 108-109, Section 6.10, Paragraph 1, lines 2-6, “Generally, such systems use closed circuit television (CCTV) cameras to identify and track individual customers as they move around within a store. The end result is either a “heatmap” or a map of movement traces. These traces essentially illustrate how customers traverse the store by steering as a response to path-planning, locomotion, and interactions with staff and products.”.) predicting, … prior to deploying the journey(Until an actual customer begins to interact) ((Torrens) pg. 98, Section 6, paragraph 1, lines 9-24, “… a common gap in retailers’ control is the customer journey: a store may realign its layout, but until an actual customer begins to interact with that new design, the retailer cannot reasonably assess what implications result. This is where simulation scenarios based on agent-based models can come into particular usefulness, in allowing the retailer to experiment with synthetic versions of the customers that they seek to influence…” ‘Until an actual customer begins to interact’ indicates that the predicting is before that so is prior to deploying an actual customer journey.), a probability of customers triggering (transition probabilities) each event of the sequence of the plurality of events (huge number of touchpoints) of the journey […] ((Torrens) pg. 103, Section 6.4, paragraph 2, lines 10-15, “Attention to fine-scale detail would imply that customers could be exposed to a huge number of choice-points while on the customer journey, and a deviation from inertia would suggest that the transition probabilities for agent states relative to those choice-points would need to be assessed at each timestep in the model (from a discrete span of time t → t + 1), for each discrete state, for each discrete agent.” A huge number implies there is a plurality.) Torrens does not teach: A computing system comprising: a processor; and a non-transitory computer-readable medium having stored thereon instructions that when executed by the processor, cause the processor to perform operations including: …by a machine learning model… …based at least on a positional embedding determined for each event of the sequence of the plurality of events; and causing, by a user interface component, display of the probability of the customers triggering each event of the sequence of the plurality of events of the journey. Zhao teaches: a processor; and a non-transitory computer-readable medium having stored thereon instructions that when executed by the processor, cause the processor to perform operations ((Zhao) Paragraph [0006], “In a third embodiment, a non-transitory computer readable medium contains instructions that when executed cause at least one processor to …”) the machine learning model (transformer layers) doing the predicting being based on at least a positional embedding (influenced by a positional embedding) determined for each event of the input sequence of the plurality of events; ((Zhao) (Paragraph [0069] “The data on which the transformer layers 309a-309d operate is influenced by a positional embedding 312.”) and causing, by a user interface component (display), display of the probability of the customers triggering each event of the sequence of the plurality of events of the journey. ((Zhao) Paragraph [0036] “According to embodiments of this disclosure, an electronic device 101 is included in the network configuration 100. The electronic device 101 can include at least one of a bus 110, a processor 120, a memory 130, an input/output (I/O) interface 150, a display 160, a communication interface 170, or a sensor 180.” A display can cause display predicted results.) Zhao and Torrens are analogous art because they are in the same field of invention: predicting customer behavior. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined a processor, a non-transitory computer readable medium storing instructions, predicting using a machine learning model that uses positional embedding, and displaying results on a user interface component as taught by Zhao, with predicting the probability of customers triggering events as Torrens teaches. The motivation for the combinations would have been to have a physical medium to run the instruction son, and to take advantage of machine learnings ability to discover patterns that ordinary human analysis might miss, while keeping information about the ordering of events by using positional embedding, and to have a way to view results for review. This application of the technique of using a machine learning model using positional embedding to predict a probability of customers triggering each event of a journey, and addition of a processor, non-transitory computer readable medium, and display user interface component would yield the predictable result that is the invention specified in claim 15 of the instant application. The Examiner notes that this motivation applies to all dependent and/or otherwise subsequently addressed claims. Regarding claim 16, Torrens, in view of Zhao, teaches the material disclosed in claim 15, and Zhao additionally teaches: wherein the machine learning model (deep learning model) is a transformer-based (transformer structure) machine learning model ((Zhao) Paragraph [0034] “A deep learning model structure may use a transformer structure”) Zhao and Torrens are analogous art because they are in the same field of invention: predicting customer behavior. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined a machine learning model that uses a transformer structure as taught by Zhao, with the prediction of probabilities of customers triggering each event of a sequence of events in a journey as Torrens teaches. The motivation for the combination would have been to take advantage of machine learnings ability to discover patterns that ordinary human analysis might miss, as well as a transformer structure to keep track of ordering during machine learning. This application of the technique of using a machine learning model using a transformer structure to predict a probability of customers triggering each event would yield the predictable result that is the invention specified in claim 16 of the instant application. Claim(s) 5, 8, 11-14, and 17-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Torrens , in view of Zhao, in further view of CASPR: Customer Activity Sequence-based Prediction and Representation by Pin-Jung Chen, et al., hereafter CASPR. Regarding claim 5, Torrens, in view of Zhao, teaches the material disclosed in claim 1. Torrens, in view of Zhao, does not teach: training an embedding layer of the machine learning model to generate an embedding for each event in the input sequence of the plurality of events of the input journey; and training a transformer encoder layer of the machine learning model to generate the positional embedding for each embedding of each event in the input sequence of the plurality of events of the input journey. CASPR teaches: training an embedding layer (learn-able categorical embedding layer) of the machine learning model to generate an embedding (vector representation) for each event (each sub-dataset D_E) in the input sequence of the plurality of events (dataset D) of the input journey; ((CASPR) Page 5, Figure 2; Page 4, section 5.1, “We then batch the sub-datasets D_E and pass them through a learn-able categorical embedding layer to convert all categorical activity attributes into a vector representation”) and training a transformer encoder layer (encoding transformer) of the machine learning model to generate the positional embedding (encoded output/encoded embedding representation) for each embedding (input to the transformer is … a dense embedding representation obtained after passing through an embedding layer described in section 5.1) of each event in the input sequence of the plurality of events of the input journey. ((CASPR) page 5, section 5.2; Page 5, Figure 2; Page 2, section 2, “Just as language models can understand the semantics of words in a sentence, the CASPR transformer understands the semantics of events w.r.t. their position in a customer’s journey.” The encoded output understands the semantics of events with respect to positions of a customer journey so is a positional embedding.) CASPR, Torrens, and Zhao are analogous art because they are in the same area of invention: simulating customer journeys. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined training an embedding layer to generate an embedding to input into a transformer encoding layer to generate positional embedding as taught by CASPR with the machine learning model using positional embedding as Torrens, in view of Zhao, teaches. The motivation for the combination would have been to help the model understand the structure and semantic relationships between events is a sequence. This simple substitution of the embedding layer generating an embedding to use as input for a transformer encoding layer to generate a positional embedding, as taught by CASPR, into the general use of positional embedding as Torrens, in view of Zhao, teaches would yield the predictable result that is the invention specified in claim 5 of the instant application. Regarding claim 8, Torrens, in view of Zhao, teaches the material disclosed in claim 1, and additionally Torrens teaches: Generating the simulated customer interaction (simulation scenarios) with the input journey; ((Torrens) pg. 98, Section 6, paragraph 1, lines 9-24) Torrens, in view of Zhao, does not teach: generating, by an embedding layer, an embedding for each event in the input sequence of the plurality of events of the input journey; generating, by a transformer encoder layer, the positional embedding for each embedding of each event in the input sequence of the plurality of events of the input journey; and generating the simulated customer interaction with the input journey based at least on the positional embedding for each embedding of each event in the input sequence of the plurality of events of the input journey. CASPR teaches: generating, by an embedding layer (learn-able categorical embedding layer), an embedding (vector representation) for each event (each sub-dataset D_E) in the input sequence of the plurality of events of the input journey(dataset D); ((CASPR) Page 5, Figure 2;Page 4, section 5.1, “We then batch the sub-datasets D_E and pass them through a learn-able categorical embedding layer to convert all categorical activity attributes into a vector representation”) generating, by a transformer encoder layer (encoding transformer), the positional embedding (encoded output/encoded embedding representation) for each embedding (input to the transformer is … a dense embedding representation obtained after passing through an embedding layer described in section 5.1) of each event in the input sequence of the plurality of events of the input journey(Each sub-dataset D_E in dataset D) ((CASPR) page 5, section 5.2; Page 5, Figure 2; Page 2, section 2, “Just as language models can understand the semantics of words in a sentence, the CASPR transformer understands the semantics of events w.r.t. their position in a customer’s journey.” The encoded output understands the semantics of events with respect to positions of a customer journey so is a positional embedding.); and generating the simulated customer interaction with the input journey (predicting churn/lifetime value/product recommendation) based at least on the positional embedding (encoded output/encoded embedding representation) for each embedding (generated embedding representations) of each event in the input sequence of the plurality of events of the input journey (Each sub-dataset D_E in dataset D) ((CASPR) Page 4, section 5, “The generated embedding representations for each entity can then be used for a range of downstream tasks such as predicting churn, lifetime value or detecting fraudulent accounts.”; Page 6, section 5.3; ) CASPR, Torrens, and Zhao are analogous art because they are in the same area of invention: simulating customer journeys. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined to generate an embedding to input into a transformer encoding layer to generate positional embedding to generate simulated customer interaction as taught by CASPR with the machine learning model using positional embedding to simulate customer interaction as Torrens, in view of Zhao, teaches. The motivation for the combination would have been to help the model understand the structure and semantic relationships between events is a sequence. This simple substitution of the embedding layer generating an embedding to use as input for a transformer encoding layer to generate a positional embedding to generate simulated customer interaction, as taught by CASPR, into the general use of positional embedding to simulate customer interaction as Torrens, in view of Zhao, teaches would yield the predictable result that is the invention specified in claim 8 of the instant application. Regarding claim 11, Torrens, in view of Zhao, teaches the material disclosed in claim 9, and Torrens additionally teaches: obtaining training data from a plurality of journeys (map of movements traces), ((Torrens) pg. 108-109, Section 6.10, Paragraph 1, lines 2-6, “Generally, such systems use closed circuit television (CCTV) cameras to identify and track individual customers as they move around within a store. The end result is either a “heatmap” or a map of movement traces. These traces essentially illustrate how customers traverse the store by steering as a response to path-planning, locomotion, and interactions with staff and products.”.), each journey of the plurality of journeys comprising a corresponding sequence of events(choice-points), wherein the training data indicates customer (discrete agent) interactions (interactions with staff and products) with each event(choice-point/discrete state) in the corresponding sequence of events (choice-points) of each journey of the plurality of journeys; ((Torrens) pg. 103, Section 6.4, paragraph 2, lines 10-15, “Attention to fine-scale detail would imply that customers could be exposed to a huge number of choice-points while on the customer journey, and a deviation from inertia would suggest that the transition probabilities for agent states relative to those choice-points would need to be assessed at each timestep in the model (from a discrete span of time t → t + 1), for each discrete state, for each discrete agent.”) Torrens does not teach: training, using the training data, the machine learning model to predict the predicted event based on the sequence of the plurality of events of the journey, the training comprising: training an embedding layer of the machine learning model to generate an embedding for each event of the sequence of the plurality of events of the journey; and training a transformer encoder layer of the machine learning model to generate the positional embedding for each embedding of each event of the sequence of the plurality of events of the journey. CASPR teaches: training, using the training data, the machine learning model to predict the predicted event (predicting product recommendation) based on the sequence of the plurality of events of the journey (Each sub-dataset D_E in dataset D), ((CASPR) Page 4, section 5, “The generated embedding representations for each entity can then be used for a range of downstream tasks such as predicting churn, lifetime value or detecting fraudulent accounts.”; Page 6, section 5.3; ) the training comprising: training an embedding layer (learn-able categorical embedding layer) of the machine learning model to generate an embedding (vector representation) for each event (each sub-dataset D_E) in the sequence of the plurality of events (dataset D) of the journey; ((CASPR) Page 5, Figure 2; Page 4, section 5.1, “We then batch the sub-datasets D_E and pass them through a learn-able categorical embedding layer to convert all categorical activity attributes into a vector representation”) and training a transformer encoder layer (encoding transformer) of the machine learning model to generate the positional embedding (encoded output/encoded embedding representation) for each embedding (input to the transformer is … a dense embedding representation obtained after passing through an embedding layer described in section 5.1) of each event in the sequence of the plurality of events of the journey. ((CASPR) page 5, section 5.2; Page 5, Figure 2; Page 2, section 2, “Just as language models can understand the semantics of words in a sentence, the CASPR transformer understands the semantics of events w.r.t. their position in a customer’s journey.” The encoded output understands the semantics of events with respect to positions of a customer journey so is a positional embedding.) CASPR, Torrens, and Zhao are analogous art because they are in the same area of invention: simulating customer journeys. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined training an embedding layer to generate an embedding to input into a transformer encoding layer to generate positional embedding to predict a predicted event as taught by CASPR with the machine learning model using positional embedding to predict a predicted event as Torrens, in view of Zhao, teaches. The motivation for the combination would have been to help the model understand the structure and semantic relationships between events is a sequence. This simple substitution of the embedding layer generating an embedding to use as input for a transformer encoding layer to generate a positional embedding to predict a predicted event, as taught by CASPR, into the general use of positional embedding to predict a predicted event as Torrens, in view of Zhao, teaches would yield the predictable result that is the invention specified in claim 11 of the instant application. Regarding claim 12, Torrens, in view of Zhao, teaches the material disclosed in claim 11, and Zhao additionally teaches: obtaining customer data for each customer of the plurality of journeys of the training data, wherein the customer data comprises demographic data (geographic data); ((Zhao) Paragraph [0071] “In some cases, the feature data used in the present disclosure includes the following information” and diagram below paragraph [0071].) and further training the machine learning model to predict (generate an inference) the predicted event (user behavior) using the customer data (user behavior information and feature data). ((Zhao) Paragraph [0051-0052]) Zhao and Torrens are analogous art because they are in the same field of invention: predicting customer behavior. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined training a machine learning model that uses geographic data as part of its dataset as taught by Zhao, with the prediction of a predicted event as Torrens teaches. The motivation for the combination would have been to take advantage of machine learnings ability to discover patterns that ordinary human analysis might miss, while geographic data is used as datapoints that are helpful due to being unchanged by typical user behavior. This application of the technique of using a machine learning model trained using a dataset that includes geographic data to predict a predicted event would yield the predictable result that is the invention specified in claim 12 of the instant application. Regarding claim 13, Torrens, in view of Zhao and CASPR, teaches the material disclosed in claim 11, and Zhao additionally teaches: wherein the training data further comprises data indicating a time of customer interactions (time of day) with each event in the sequence of events of each journey of the plurality of journeys; ((Zhao) Paragraph [0071] and diagram below paragraph [0071].) and further training the machine learning model to predict (generate an inference) the predicted event (user behavior) using the time of the customer interactions (time of day). ((Zhao) Paragraph [0051-0052]) Zhao and Torrens are analogous art because they are in the same field of invention: predicting customer behavior. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined training a machine learning model that uses time of interaction data as part of its dataset as taught by Zhao, with the prediction of a predicted event as Torrens teaches. The motivation for the combination would have been to take advantage of machine learnings ability to discover patterns that ordinary human analysis might miss, while time of interaction data is used as datapoints that are helpful due to being relatively unaffected by typical user behavior. This application of the technique of using a machine learning model trained using a dataset that includes time of interaction data to predict a predicted event would yield the predictable result that is the invention specified in claim 13 of the instant application. Regarding claim 14, Torrens, in view of Zhao, teaches the material disclosed in claim 9, and additionally Torrens teaches: predicting the predicted event (Estimates of the likely flow) … ((Torrens) pg. 107, Section 6.7, Paragraph 2, lines 6-10) Torrens, in view of Zhao, does not teach: generating, by an embedding layer, an embedding for each event in the input sequence of the plurality of events of the input journey; generating, by a transformer encoder layer, the positional embedding for each embedding of each event in the input sequence of the plurality of events of the input journey; and predicting the predicted event based at least on the positional embedding for each embedding of each event in the sequence of the plurality of events of the journey. CASPR teaches: generating, by an embedding layer (learn-able categorical embedding layer), an embedding (vector representation) for each event (each sub-dataset D_E) in the input sequence of the plurality of events of the input journey(dataset D); ((CASPR) Page 5, Figure 2;Page 4, section 5.1, “We then batch the sub-datasets D_E and pass them through a learn-able categorical embedding layer to convert all categorical activity attributes into a vector representation”) generating, by a transformer encoder layer (encoding transformer), the positional embedding (encoded output/encoded embedding representation) for each embedding (input to the transformer is … a dense embedding representation obtained after passing through an embedding layer described in section 5.1) of each event in the input sequence of the plurality of events of the input journey(Each sub-dataset D_E in dataset D) ((CASPR) page 5, section 5.2; Page 5, Figure 2; Page 2, section 2, “Just as language models can understand the semantics of words in a sentence, the CASPR transformer understands the semantics of events w.r.t. their position in a customer’s journey.” The encoded output understands the semantics of events with respect to positions of a customer journey so is a positional embedding.); and predicting the predicted event (predicting product recommendation) based at least on the positional embedding (encoded output/encoded embedding representation) for each embedding (generated embedding representations) of each event in the input sequence of the plurality of events of the input journey (Each sub-dataset D_E in dataset D) ((CASPR) Page 4, section 5, “The generated embedding representations for each entity can then be used for a range of downstream tasks such as predicting churn, lifetime value or detecting fraudulent accounts.”; Page 6, section 5.3; ) CASPR, Torrens, and Zhao are analogous art because they are in the same area of invention: simulating customer journeys. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined to generate an embedding to input into a transformer encoding layer to generate positional embedding to predict a predicted event as taught by CASPR with the machine learning model using positional embedding to predict a predicted event as Torrens, in view of Zhao, teaches. The motivation for the combination would have been to help the model understand the structure and semantic relationships between events is a sequence. This simple substitution of the embedding layer generating an embedding to use as input for a transformer encoding layer to generate a positional embedding to predict a predicted event, as taught by CASPR, into the general use of positional embedding to predict a predicted event as Torrens, in view of Zhao, teaches would yield the predictable result that is the invention specified in claim 14 of the instant application. Regarding claim 17, Torrens, in view of Zhao, teaches the material disclosed in claim 15, and Torrens additionally teaches: obtaining training data from a plurality of journeys (map of movements traces), ((Torrens) pg. 108-109, Section 6.10, Paragraph 1, lines 2-6, “Generally, such systems use closed circuit television (CCTV) cameras to identify and track individual customers as they move around within a store. The end result is either a “heatmap” or a map of movement traces. These traces essentially illustrate how customers traverse the store by steering as a response to path-planning, locomotion, and interactions with staff and products.”.), each journey of the plurality of journeys comprising a corresponding sequence of events(choice-points), wherein the training data indicates customer (discrete agent) interactions (interactions with staff and products) with each event(choice-point/discrete state) in the corresponding sequence of events (choice-points) of each journey of the plurality of journeys; ((Torrens) pg. 103, Section 6.4, paragraph 2, lines 10-15, “Attention to fine-scale detail would imply that customers could be exposed to a huge number of choice-points while on the customer journey, and a deviation from inertia would suggest that the transition probabilities for agent states relative to those choice-points would need to be assessed at each timestep in the model (from a discrete span of time t → t + 1), for each discrete state, for each discrete agent.”) … predicting the probability of the customers triggering each event (transition probability) of the sequence of the plurality of events of the journey…((Torrens) pg. 103, Section 6.4, paragraph 2, lines 10-15) Torrens does not teach: training, by the model training engine using the training data, the machine learning model […], the training comprising: training an embedding layer of the machine learning model to generate an embedding for each event of the sequence of the plurality of events of the journey; and training a transformer encoder layer of the machine learning model to generate the positional embedding for each embedding of each event of the sequence of the plurality of events of the journey. CASPR teaches: training, using the training data, the machine learning model to predict (Each sub-dataset D_E in dataset D), ((CASPR) Page 4, section 5, “The generated embedding representations for each entity can then be used for a range of downstream tasks such as predicting churn, lifetime value or detecting fraudulent accounts.”; Page 6, section 5.3; ) the training comprising: training an embedding layer (learn-able categorical embedding layer) of the machine learning model to generate an embedding (vector representation) for each event (each sub-dataset D_E) in the sequence of the plurality of events (dataset D) of the journey; ((CASPR) Page 5, Figure 2; Page 4, section 5.1, “We then batch the sub-datasets D_E and pass them through a learn-able categorical embedding layer to convert all categorical activity attributes into a vector representation”) and training a transformer encoder layer (encoding transformer) of the machine learning model to generate the positional embedding (encoded output/encoded embedding representation) for each embedding (input to the transformer is … a dense embedding representation obtained after passing through an embedding layer described in section 5.1) of each event in the sequence of the plurality of events of the journey. ((CASPR) page 5, section 5.2; Page 5, Figure 2; Page 2, section 2, “Just as language models can understand the semantics of words in a sentence, the CASPR transformer understands the semantics of events w.r.t. their position in a customer’s journey.” The encoded output understands the semantics of events with respect to positions of a customer journey so is a positional embedding.) CASPR, Torrens, and Zhao are analogous art because they are in the same area of invention: simulating customer journeys. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined training an embedding layer to generate an embedding to input into a transformer encoding layer to generate positional embedding to predict as taught by CASPR with the machine learning model using positional embedding to predict a probability of customers triggering each event as Torrens, in view of Zhao, teaches. The motivation for the combination would have been to help the model understand the structure and semantic relationships between events is a sequence. This simple substitution of the embedding layer generating an embedding to use as input for a transformer encoding layer to generate a positional embedding to predict a probability of customers triggering each event, as taught by CASPR, into the general use of positional embedding to predict a probability of customers triggering each event as Torrens, in view of Zhao, teaches would yield the predictable result that is the invention specified in claim 17 of the instant application. Regarding claim 18, Torrens, in view of Zhao, teaches the material disclosed in claim 17, and Zhao additionally teaches: obtaining customer data for each customer of the plurality of journeys of the training data, wherein the customer data comprises demographic data (geographic data); ((Zhao) Paragraph [0071] “In some cases, the feature data used in the present disclosure includes the following information” and diagram below paragraph [0071].) and further training the machine learning model to predict the probability of customers triggering each event of the sequence of the plurality of events of the journey using the customer data (user behavior information and feature data). ((Zhao) Paragraph [0051-0052]) Zhao and Torrens are analogous art because they are in the same field of invention: predicting customer behavior. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined training a machine learning model that uses geographic data as part of its dataset as taught by Zhao, with predicting a probability of customers triggering each event as Torrens teaches. The motivation for the combination would have been to take advantage of machine learnings ability to discover patterns that ordinary human analysis might miss, while geographic data is used as datapoints that are helpful due to being unchanged by typical user behavior. This application of the technique of using a machine learning model trained using a dataset that includes geographic data to predict a probability of customers triggering each event would yield the predictable result that is the invention specified in claim 18 of the instant application. Regarding claim 19, Torrens, in view of Zhao and CASPR, teaches the material disclosed in claim 17, and Zhao additionally teaches: wherein the training data further comprises data indicating a time of customer interactions (time of day) with each event in the sequence of events of each journey of the plurality of journeys; ((Zhao) Paragraph [0071] and diagram below paragraph [0071].) and further training the machine learning model to predict the probability of customers triggering each event of the sequence of the plurality of events of the journey using the time of the customer interactions (time of day). ((Zhao) Paragraph [0051-0052]) Zhao and Torrens are analogous art because they are in the same field of invention: predicting customer behavior. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined training a machine learning model that uses time of interaction data as part of its dataset as taught by Zhao, with predicting a probability of customers triggering each event as Torrens teaches. The motivation for the combination would have been to take advantage of machine learnings ability to discover patterns that ordinary human analysis might miss, while time of interaction data is used as datapoints that are helpful due to being relatively unaffected by typical user behavior. This application of the technique of using a machine learning model trained using a dataset that includes time of interaction data to predict a probability of customers triggering each event would yield the predictable result that is the invention specified in claim 19 of the instant application. Regarding claim 20, Torrens, in view of Zhao, teaches the material disclosed in claim 15, and additionally Torrens teaches: predicting a probability of customers triggering (transition probabilities) each event of the sequence of the plurality of events (huge number of touchpoints) of the journey […] ((Torrens) pg. 103, Section 6.4, paragraph 2, lines 10-15, A huge number implies there is a plurality.) Torrens, in view of Zhao, does not teach: generating, by an embedding layer, an embedding for each event in the input sequence of the plurality of events of the input journey; generating, by a transformer encoder layer, the positional embedding for each embedding of each event in the input sequence of the plurality of events of the input journey; and predicting … based at least on the positional embedding for each embedding of each event in the sequence of the plurality of events of the journey. CASPR teaches: generating, by an embedding layer (learn-able categorical embedding layer), an embedding (vector representation) for each event (each sub-dataset D_E) in the input sequence of the plurality of events of the input journey(dataset D); ((CASPR) Page 5, Figure 2;Page 4, section 5.1, “We then batch the sub-datasets D_E and pass them through a learn-able categorical embedding layer to convert all categorical activity attributes into a vector representation”) generating, by a transformer encoder layer (encoding transformer), the positional embedding (encoded output/encoded embedding representation) for each embedding (input to the transformer is … a dense embedding representation obtained after passing through an embedding layer described in section 5.1) of each event in the input sequence of the plurality of events of the input journey(Each sub-dataset D_E in dataset D) ((CASPR) page 5, section 5.2; Page 5, Figure 2; Page 2, section 2, “Just as language models can understand the semantics of words in a sentence, the CASPR transformer understands the semantics of events w.r.t. their position in a customer’s journey.” The encoded output understands the semantics of events with respect to positions of a customer journey so is a positional embedding.); and predicting based at least on the positional embedding (encoded output/encoded embedding representation) for each embedding (generated embedding representations) of each event in the input sequence of the plurality of events of the input journey (Each sub-dataset D_E in dataset D) ((CASPR) Page 4, section 5, “The generated embedding representations for each entity can then be used for a range of downstream tasks such as predicting churn, lifetime value or detecting fraudulent accounts.”; Page 6, section 5.3; ) CASPR, Torrens, and Zhao are analogous art because they are in the same area of invention: simulating customer journeys. Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the references in front of them, to have combined to generate an embedding to input into a transformer encoding layer to generate positional embedding to predict as taught by CASPR with the machine learning model using positional embedding to predict a probability of customers triggering each event as Torrens, in view of Zhao, teaches. The motivation for the combination would have been to help the model understand the structure and semantic relationships between events is a sequence. This simple substitution of the embedding layer generating an embedding to use as input for a transformer encoding layer to generate a positional embedding to predict, as taught by CASPR, into the general use of positional embedding to specifically predict a probability of customers triggering each event as Torrens, in view of Zhao, teaches would yield the predictable result that is the invention specified in claim 20 of the instant application. Response to Arguments Applicant’s arguments, see page 13, lines 7-11, filed 06/10/2026, with respect to claims 1-3, 6, 7, 9, and 15 have been fully considered and are persuasive. The independent claims 1, 9, and 15 have been amended to each contain limitations found in claims 4, 10, and 16 respectively. The amended limitations, “generating simulated customer interaction … based at least on a positional embedding determined for each event of the input sequence of the plurality of events”, integrate the abstract ideas into a practical application. The rejections based on 35 U.S.C. 101 of claims 1-3, 6, 7, 9 and 15 has been withdrawn. Applicant’s arguments with respect to claim(s) 1, 9, and 15 and their dependent claims have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Patents and/or related publications are cited in the Notice of References Cited (Form PTO-892) attached to this action to further show the state of the art with respect to predicting and simulating customer journeys, transformer machine learning models, and positional embedding and encoding. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DYLAN H LAI whose telephone number is (571)272-8628. The examiner can normally be reached Monday - Friday 7:30am-5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tamara Kyle can be reached at 5712524241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. D. H. L. Examiner Art Unit 2144 /TAMARA T KYLE/Supervisory Patent Examiner, Art Unit 2144
Read full office action

Prosecution Timeline

May 16, 2023
Application Filed
Mar 12, 2026
Non-Final Rejection mailed — §103
May 21, 2026
Interview Requested
Jun 04, 2026
Examiner Interview Summary
Jun 04, 2026
Applicant Interview (Telephonic)
Jun 10, 2026
Response Filed
Sep 04, 2026
Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
Grant Probability
Moderate
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month