Prosecution Insights
Last updated: October 04, 2026
Application No. 18/683,017

METHOD AND SYSTEM FOR ANALYSING OPERATION OF A ROBOT

Non-Final OA §101§103
Filed
Feb 12, 2024
Priority
Aug 11, 2021 — DE 10 2021 208 769.8 +1 more
Examiner
RAMIREZ BRAVO, BEATRIZ A
Art Unit
Tech Center
Assignee
Kuka Deutschland GmbH
OA Round
1 (Non-Final)
64%
Grant Probability
Moderate
1-2
OA Rounds
1y 11m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 64% of resolved cases
64%
Career Allowance Rate
64 granted / 100 resolved
+4.0% vs TC avg
Strong +29% interview lift
Without
With
+28.9%
Interview Lift
resolved cases with interview
Typical timeline
4y 6m
Avg Prosecution
19 currently pending
Career history
120
Total Applications
across all art units

Statute-Specific Performance

§101
15.1%
-24.9% vs TC avg
§103
59.2%
+19.2% vs TC avg
§102
11.0%
-29.0% vs TC avg
§112
13.3%
-26.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 100 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of Claims Claims 1-10 have been cancelled by Applicant in Preliminary Amendment. New claims 11-24 have been added. Claims 11-24 are currently pending. Information Disclosure Statement The Information Disclosure Statements submitted by Applicant on 2/12/2024, 2/12/2024, 2/29/2024, and 2/29/2024 have been considered. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Such claim limitations are: “means for performing a training phase” in claim 23 “means for performing a monitoring phase” in claim 23 [EXAMINER NOTE: Sufficient structure for performing the limitations above has been identified in pg. 7 of Applicant’s specification stating “A system and/or a means within the meaning of the present invention may be designed in hardware and/or in software, and in particular may comprise at least one data-connected or signal-connected, in particular digital, processing unit, in particular microprocessor unit (CPU), graphic card (CPU) having a memory and/or bus system or the like and/or one or multiple programs or program modules. The processing unit may be designed to process commands that are implemented as a program stored in a memory system, to detect input signals from a data bus and/or to output output signals to a data bus. A storage system may comprise one or a plurality of, in particular different, storage media, in particular optical, magnetic, solid-state, and/or other non-volatile media. The program may be such that it embodies or is capable of executing the methods described herein, so that the processing unit can execute the steps of such methods and thus in particular identify the pattern group(s) and/or analyze, monitor and/or modify operation of the robot. In one embodiment, a computer program product may comprise, in particular be, a storage medium, in particular computer-readable and/or non-volatile, for storing a program or instructions or with a program stored thereon or with instructions stored thereon. In one embodiment, execution of said program or instructions by a system, in particular a computer or an arrangement of a plurality of computers, causes the system, in particular the computer or computers, to carry out a method described herein or one or more steps thereof, or the program or instructions are adapted to do so.] Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 24 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. (Step 1) Claim 24 recites “A computer program or computer program product comprising program code store on a non-transient, computer-readable medium, the program code configured, when executed on a computer, to cause the computer to…”. Examiner notes that “non-transient, computer-readable medium” is not the same as “non-transitory computer-readable medium”. Furthermore, there is no definition for “non-transient, computer-readable medium” in Applicant’s specification. Therefore, claim 24 is rejected under 35 U.S.C. 101 for being directed to non-statutory subject matter as regarded as directed to signal per se. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 11, 12, 16, 17, 18, 19, 20, 21, 22, 23, and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Azzalini et al., “A Minimally Supervised Approach Based on Variational Autoencoders for Anomaly Detection in Autonomous Robots”, (April 2021) in view of Mandlekar et al. (US 20220055689 A1, filed Aug. 20, 2020 and published Feb. 24, 2022) Regarding claim 11, Mandlekar teaches a method for analyzing an operation of a robot, the method comprising: (a) performing a training phase, including: obtaining with a robot controller a first data set having at least one temporal characteristic of at least one state parameter of a first robot (Azzalini, Abstract, teaches detection of anomalies and faults is a crucial ability for fully autonomous robots. This letter proposes a new deep learning-based minimally supervised method for detecting anomalies in autonomous robots. We contribute a new Variational Auto-Encoder architecture able to model very long multivariate sensor logs exploiting a new incremental training method, which induces a progress-based latent space that can be used to detect anomalies both at runtime and offline; Azzalini, pg. 2988, Section III.C teaches represent the temporal dependency of multivariate time series collected from robot sensors), and training an artificial neural network (Azzalini, Section I, col. 2, teaches we also introduce a new incremental method for training VAEs, which induces a progress-based latent space that can be used to detect anomalies both online (at runtime) and offline… Only few(even just one) labeled nominal executions are then required to partition the learned latent space into nominal and anomalous regions. This minimally supervised approach provides a big advantage over semi-supervised approaches in practical settings, where collecting several nominal runs of a robot performing a task could be hard, since a human expert is usually required to supervise the system in order to label runs as nominal.), However, Azzalini does not distinctly disclose: the artificial neural network including: a first autoencoder having an encoder that maps the first data set to temporal characteristic patterns and corresponding activation, and a decoder that reconstructs the first data set using the mapped temporal characteristic patterns, and a second autoencoder having an encoder that maps the temporal characteristic patterns and corresponding activation to pattern groups, and a decoder that reconstructs the temporal characteristic patterns and corresponding activation using the pattern group; (b) performing a monitoring phase, including: obtaining a second data set having at least one temporal characteristic of the at least one state parameter of the first robot or a second robot, and identifying with a computer at least one of the pattern groups of the trained second autoencoder within the second data set Nevertheless, Mandlekar teaches: the artificial neural network including: a first autoencoder having an encoder that maps the first data set to temporal characteristic patterns and corresponding activation, and a decoder that reconstructs the first data set using the mapped temporal characteristic patterns, and a second autoencoder having an encoder that maps the temporal characteristic patterns and corresponding activation to pattern groups, and a decoder that reconstructs the temporal characteristic patterns and corresponding activation using the pattern groups (Mandlekar, [0089] teaches Learning from suboptimal data: In an embodiment, the low-level goal-conditioned controller operates for a small number of timesteps, so the controller has no need to account for suboptimal actions. This is because if the goal is to reach a state s.sub.2 from s.sub.1, and T is sufficiently small, then a policy may only be able to improve by reaching s.sub.2 in less than T steps, which is a negligible improvement for small values of T. By contrast, the value learning component of the goal selection mechanism explicitly accounts for suboptimal solution approaches by evaluating the expected task returns of each goal and selecting the goal with the highest return.; Mandlekar, [0090] further teaches Learning from off-policy datasets: Policy learning from arbitrary off-policy data can be challenging. Some embodiments of IRIS deal with this issue by constraining learning to occur within the distribution of training data. In some examples, the goal-conditioned controller directly imitates sequences from the training data, and the generative goal model is also trained to propose goal observations from the training data. Finally, the value learning component of the goal selection mechanism mitigates extrapolation error by making sure that the Q-network is queried on state-action pairs that lie within the training distribution.; Mandlekar, [0568] 13. Teaches The system of any of clauses 8 to 12, wherein the neural network is a variational autoencoder.; Mandlekar [0569] 14. further teaches The system of any of clauses 8 to 13, wherein the system uses a recurrent neural network to determine a set of actions that, as a result of being performed by the robot, reposition the robot from the current position to an intermediate goal.; Mandlekar, [0570] 15. Teaches A machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least: use a first neural network to determine a set of intermediate goal proposals based on a current position of a robot, the first neural network trained using a set of demonstrations of task performance; select, based at least in part on a value function, an intermediate goal from the set of intermediate goal proposals; use a second neural network to determine a set of actions that, as a result of being performed by the robot, reposition the robot from the current position to the selected intermediate goal; perform the task by at least performing the set of actions.; Mandlekar, [0085] teaches In one example, the cVAE is a conditional generative model that is trained on pairs of current and future observations (s.sub.t, s.sub.t+T) [i.e., “temporal characteristics”] sampled from trajectories in the dataset (lines 5-7 in Algorithm 1). An encoder maps a current and future observation to the parameters of a latent Gaussian distribution μ.sub.b, σ.sub.g=E.sub.Ø(s.sub.t+T, s.sub.t) [i.e., “temporal characteristics”] the decoder is trained to reconstruct the future observation from the current observation and a latent sampled from the encoder distribution ś.sub.t+T=D.sub.Ø(z, s.sub.t), z˜N(μ.sub.g, σ.sub.G). The encoder distribution is regularized with a KL-loss KL(N μ.sub.g, σ.sub.G)∥N(0,1) with weight β.sub.g [15] to encourage the encoder distribution to match a prior latent distribution p(z)=N(0,1) so that at test-time, the decoder can be used as a conditional generative model by sampling latents z˜N (0,1) and passing them through the decoder.; Mandlekar, [0123] further teaches In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows identification of data points in new dataset 1212 that deviate from normal patterns of new dataset 1212.; Mandlekar, [0115] teaches In at least one embodiment, inference and/or training logic 1115 may include, without limitation, one or more arithmetic logic unit(s) (“ALU(s)”) 1110, including integer and/or floating point units, to perform logical and/or mathematical operations based, at least in part on, or indicated by, training and/or inference code (e.g., graph code), a result of which may produce activations (e.g., output values from layers or neurons within a neural network) stored in an activation storage 1120 that are functions of input/output and/or weight parameter data stored in code and/or data storage 1101 and/or code and/or data storage 1105.); and (b) performing a monitoring phase, including: obtaining a second data set having at least one temporal characteristic of the at least one state parameter of the first robot or a second robot (Mandlekar, [0198] teaches In at least one embodiment, vehicle 1400 may include CPU(s) 1418 (e.g., discrete CPU(s), or dCPU(s)), that may be coupled to SoC(s) 1404 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, CPU(s) 1418 may include an X86 processor, for example. CPU(s) 1418 may be used to perform any of a variety of functions, including arbitrating potentially inconsistent results between ADAS sensors and SoC(s) 1404, and/or monitoring status and health of controller(s) 1436 and/or an infotainment system on a chip (“infotainment SoC”) 1430, for example.; Mandlekar [0236], teaches In at least one embodiment, server(s) 1478 may receive data from vehicles and apply data to up-to-date real-time neural networks for real-time intelligent inferencing. [Note: data to up-to-date real-time as second data having at least one “temporal characteristic”.]), and identifying with a computer at least one of the pattern groups of the trained second autoencoder within the second data set (Mandlekar, [0123] teaches in at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows identification of data points in new dataset 1212 that deviate from normal patterns of new dataset 1212.). Before the effective filing date of the claimed invention, it would have been obvious to one or ordinary skill in the art to have modified the method for anomaly detection in robotics, as taught by Azzalini, with the system that learns from task demonstrations of a robot comprising two variational autoencoders, as taught by Mandlekar, in order to allow for selective imitation of local sequences in the dataset. (Mandlekar paragraphs [0082] and [0083]) Regarding claim 12, Azzalini in view of Mandlekar teaches all of the limitations of claim 11, and Mandlekar further teaches wherein the first autoencoder includes at least one variational autoencoder (Mandlekar, [0568] 13. Teaches The system of any of clauses 8 to 12, wherein the neural network is a variational autoencoder.; Mandlekar [0569] 14. further teaches The system of any of clauses 8 to 13, wherein the system uses a recurrent neural network to determine a set of actions that, as a result of being performed by the robot, reposition the robot from the current position to an intermediate goal.; Mandlekar, [0570] 15. teaches A machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least: use a first neural network to determine a set of intermediate goal proposals based on a current position of a robot, the first neural network trained using a set of demonstrations of task performance;). Motivation to combine same as stated in claim 11. Regarding claim 16, Azzalini in view of Mandlekar teaches all of the limitations of claim 11, and Mandlekar further teaches wherein at least one of: the at least one state parameter depends on at least one of" at least one position of a robot-fixed reference, at least one orientation of a robot-fixed reference, or of at least one axial load of the robot; or the at least one state parameter is detected by at least one sensor (Mandlekar, [0177] teaches In at least one embodiment, one or more of SoC(s) 1404 may include a real-time ray-tracing hardware accelerator. In at least one embodiment, real-time ray-tracing hardware accelerator may be used to quickly and efficiently determine positions and extents of objects (e.g., within a world model), to generate real-time visualization simulations, for RADAR signal interpretation, for sound propagation synthesis and/or analysis, for simulation of SONAR systems, for general wave propagation simulation, for comparison to LIDAR data for purposes of localization and/or other functions, and/or for other uses.; Mandlekar, [0181] further teaches in at least one embodiment, neural network may take as its input at least some subset of parameters, such as bounding box dimensions, ground plane estimate obtained (e.g., from another subsystem), output from IMU sensor(s) 1466 that correlates with vehicle 1400 orientation, distance, 3D location estimates of object obtained from neural network and/or other sensors (e.g., LIDAR sensor(s) 1464 or RADAR sensor(s) 1460), among others.). Motivation to combine same as stated in claim 11. Regarding claim 17, Azzalini in view of Mandlekar teaches all of the limitations of claim 16, and the combination further teaches wherein the at least one sensor is a sensor of the robot (Azzalini, pg. 2988, Section III.C teaches represent the temporal dependency of multivariate time series collected from robot sensors). Regarding claim 18, Azzalini in view of Mandlekar teaches all of the limitations of claim 11, and the combination further teaches further comprising marking the identified pattern group in the second data set (Azzalini, Section I, col. 2, teaches An original feature of our approach is that, differently from most approaches for anomaly detection in robotics, it is trained with unlabeled observations, possibly including both nominal and anomalous executions. Only few(even just one) labeled nominal executions are then required to partition the learned latent space into nominal and anomalous regions. This minimally supervised approach provides a big advantage over semi-supervised approaches in practical settings, where collecting several nominal runs of a robot performing a task could be hard, since a human expert is usually required to supervise the system in order to label runs as nominal. Experimental results on datasets collected from real robots show that our method outperforms state-of-the-art methods for anomaly detection in robots both in terms of false positive rate and alert delay.; Azzalini, pg. 2986, col. 2, teaches One significant example of application of AEs to detect anomalies in robots is [14], where the authors propose to convert sensor logs into images and then use a convolutional AE to detect anomalous behaviors resulting from cyber-security attacks.; Azzalini, III.E further teaches online anomaly detection is performed by partitioning the latent space into nominal and anomalous regions according to the provided nominal executions and by testing, at runtime, to which region a new incoming partial run belongs to.). Regarding claim 19, Azzalini in view of Mandlekar teaches all of the limitations of claim 11, and the combination further teaches further comprising at least one of: detecting at least one of a robot anomaly or an event based on the identified pattern group; or classifying the temporal characteristic of the second data set based on the identified pattern group (Azzalini, Section I, col. 2, teaches An original feature of our approach is that, differently from most approaches for anomaly detection in robotics, it is trained with unlabeled observations, possibly including both nominal and anomalous executions. Only few(even just one) labeled nominal executions are then required to partition the learned latent space into nominal and anomalous regions. This minimally supervised approach provides a big advantage over semi-supervised approaches in practical settings, where collecting several nominal runs of a robot performing a task could be hard, since a human expert is usually required to supervise the system in order to label runs as nominal. Experimental results on datasets collected from real robots show that our method outperforms state-of-the-art methods for anomaly detection in robots both in terms of false positive rate and alert delay.; Azzalini, pg. 2986, col. 2, teaches One significant example of application of AEs to detect anomalies in robots is [14], where the authors propose to convert sensor logs into images and then use a convolutional AE to detect anomalous behaviors resulting from cyber-security attacks.; Azzalini, III.E further teaches online anomaly detection is performed by partitioning the latent space into nominal and anomalous regions according to the provided nominal executions and by testing, at runtime, to which region a new incoming partial run belongs to.). Regarding claim 20, Azzalini in view of Mandlekar teaches all of the limitations of claim 11, and the combination further teaches further comprising at least one of: analyzing an operation of the first or second robot (Azzalini, Abstract, teaches detection of anomalies and faults is a crucial ability for fully autonomous robots. This letter proposes a new deep learning-based minimally supervised method for detecting anomalies in autonomous robots. We contribute a new Variational Auto-Encoder architecture able to model very long multivariate sensor logs exploiting a new incremental training method, which induces a progress-based latent space that can be used to detect anomalies both at runtime and offline); monitoring an operation of the first or second robot (Mandlekar, [0059] further teaches in one example, demonstration data is collected for a robot picking up an object, and the system uses the demonstration data to identify a set of intermediate goals that complete the task of picking up the object. The system is particularly applicable to complex tasks that can be divided into a plurality of subtasks.); or modifying an operation of the first or second robot. Motivation to combine same as stated for claim 11. Regarding claim 21, Azzalini in view of Mandlekar teaches all of the limitations of claim 20, and the combination further teaches wherein the at least one of analyzing, monitoring, or modifying is based on at least one of: the identified pattern group; the detected robot anomaly; the detected event; or the classified temporal characteristic of the second data set (Azzalini, Section I, col. 2, teaches An original feature of our approach is that, differently from most approaches for anomaly detection in robotics, it is trained with unlabeled observations, possibly including both nominal and anomalous executions. Only few (even just one) labeled nominal executions are then required to partition the learned latent space into nominal and anomalous regions. This minimally supervised approach provides a big advantage over semi-supervised approaches in practical settings, where collecting several nominal runs of a robot performing a task could be hard, since a human expert is usually required to supervise the system in order to label runs as nominal. Experimental results on datasets collected from real robots show that our method outperforms state-of-the-art methods for anomaly detection in robots both in terms of false positive rate and alert delay.; Azzalini, pg. 2986, col. 2, teaches one significant example of application of AEs to detect anomalies in robots is [14], where the authors propose to convert sensor logs into images and then use a convolutional AE to detect anomalous behaviors resulting from cyber-security attacks.; Azzalini, III.E further teaches online anomaly detection is performed by partitioning the latent space into nominal and anomalous regions according to the provided nominal executions and by testing, at runtime, to which region a new incoming partial run belongs to.; Azzalini, pg. 2989, col. 2 further teaches clusters containing points belonging to nominal executions are considered nominal regions (lines 10-11), while clusters not containing points from nominal executions, outliers, and the rest of the latent space are considered as anomalous regions. Fig. 5 shows how clusters evolve at different slices for the water monitoring robot running example.). Regarding claim 22, Azzalini in view of Mandlekar teaches all of the limitations of claim 21, and the combination further teaches further comprising: marking the identified pattern group in the second data set (Azzalini, pg. 2989, col. 2 further teaches clusters containing points belonging to nominal executions are considered nominal regions (lines 10-11), while clusters not containing points from nominal executions, outliers, and the rest of the latent space are considered as anomalous regions. Fig. 5 shows how clusters evolve at different slices for the water monitoring robot running example.).; wherein the at least one of analyzing, monitoring, or modifying is based on the identified pattern group marked in the second data set (Azzalini, pg. 2989, col. 2, teaches at runtime (Algorithm 3), an incoming incomplete runO that needs to be tested for abnormality is firstly standardized (w.r.t. the mean and standard deviation used for the standardization of the training set) and zero-padded (line 1-2), then it is encoded into its latent representation ˆz (line 5). The cosine similarity between ˆz and the encodings of all the runs in the same slice is computed and if ˆz is within a distance of _ (i.e., DBSCAN’s threshold on the maximum distance between two samples for being considered as neighbors of each other, _ = 0.5 is the default value we use in our experiments) from an encoding belonging to a nominal region, the partial run is considered nominal, while an anomaly is detected otherwise (lines 9-11). [Note: Azzalini, Section I, col. 2, teaches we also introduce a new incremental method for training VAEs, which induces a progress-based latent space that can be used to detect anomalies both online (at runtime) and offline [understood to read on first data set and second dataset].) Regarding claim 23, Azzalini teaches a system for analyzing an operation of a robot, the system comprising: (a) means for performing a training phase, wherein the training phase includes: obtaining a first data set having at least one temporal characteristic of at least one state parameter of a first robot (Azzalini, Abstract, teaches detection of anomalies and faults is a crucial ability for fully autonomous robots. This letter proposes a new deep learning-based minimally supervised method for detecting anomalies in autonomous robots. We contribute a new Variational Auto-Encoder architecture able to model very long multivariate sensor logs exploiting a new incremental training method, which induces a progress-based latent space that can be used to detect anomalies both at runtime and offline; Azzalini, pg. 2988, Section III.C teaches represent the temporal dependency of multivariate time series collected from robot sensors), and training an artificial neural network (Azzalini, Section I, col. 2, teaches we also introduce a new incremental method for training VAEs, which induces a progress-based latent space that can be used to detect anomalies both online (at runtime) and offline… Only few(even just one) labeled nominal executions are then required to partition the learned latent space into nominal and anomalous regions. This minimally supervised approach provides a big advantage over semi-supervised approaches in practical settings, where collecting several nominal runs of a robot performing a task could be hard, since a human expert is usually required to supervise the system in order to label runs as nominal.), However, Azzalini does not distinctly disclose: the artificial neural network including: a first autoencoder having an encoder that maps the first data set to temporal characteristic patterns and corresponding activation, and a decoder that reconstructs the first data set using the mapped temporal characteristic patterns, and a second autoencoder having an encoder that maps the temporal characteristic patterns and corresponding activation to pattern groups, and a decoder that reconstructs the temporal characteristic patterns and corresponding activation using the pattern groups; and (b) means for performing a monitoring phase, wherein the monitoring phase includes: obtaining a second data set having at least one temporal characteristic of the at least one state parameter of the first robot or a second robot, and identifying at least one of the pattern groups of the trained second autoencoder within the second data set. Nevertheless, Mandlekar teaches: the artificial neural network including: a first autoencoder having an encoder that maps the first data set to temporal characteristic patterns and corresponding activation, and a decoder that reconstructs the first data set using the mapped temporal characteristic patterns, and a second autoencoder having an encoder that maps the temporal characteristic patterns and corresponding activation to pattern groups, and a decoder that reconstructs the temporal characteristic patterns and corresponding activation using the pattern groups (Mandlekar, [0089] teaches Learning from suboptimal data: In an embodiment, the low-level goal-conditioned controller operates for a small number of timesteps, so the controller has no need to account for suboptimal actions. This is because if the goal is to reach a state s.sub.2 from s.sub.1, and T is sufficiently small, then a policy may only be able to improve by reaching s.sub.2 in less than T steps, which is a negligible improvement for small values of T. By contrast, the value learning component of the goal selection mechanism explicitly accounts for suboptimal solution approaches by evaluating the expected task returns of each goal and selecting the goal with the highest return.; Mandlekar, [0090] further teaches Learning from off-policy datasets: Policy learning from arbitrary off-policy data can be challenging. Some embodiments of IRIS deal with this issue by constraining learning to occur within the distribution of training data. In some examples, the goal-conditioned controller directly imitates sequences from the training data, and the generative goal model is also trained to propose goal observations from the training data. Finally, the value learning component of the goal selection mechanism mitigates extrapolation error by making sure that the Q-network is queried on state-action pairs that lie within the training distribution.; Mandlekar, [0568] 13. Teaches The system of any of clauses 8 to 12, wherein the neural network is a variational autoencoder.; Mandlekar [0569] 14. further teaches The system of any of clauses 8 to 13, wherein the system uses a recurrent neural network to determine a set of actions that, as a result of being performed by the robot, reposition the robot from the current position to an intermediate goal.; Mandlekar, [0570] 15. Teaches A machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least: use a first neural network to determine a set of intermediate goal proposals based on a current position of a robot, the first neural network trained using a set of demonstrations of task performance; select, based at least in part on a value function, an intermediate goal from the set of intermediate goal proposals; use a second neural network to determine a set of actions that, as a result of being performed by the robot, reposition the robot from the current position to the selected intermediate goal; perform the task by at least performing the set of actions.; Mandlekar, [0085] teaches In one example, the cVAE is a conditional generative model that is trained on pairs of current and future observations (s.sub.t, s.sub.t+T) [i.e., “temporal characteristics”] sampled from trajectories in the dataset (lines 5-7 in Algorithm 1). An encoder maps a current and future observation to the parameters of a latent Gaussian distribution μ.sub.b, σ.sub.g=E.sub.Ø(s.sub.t+T, s.sub.t) [i.e., “temporal characteristics”] the decoder is trained to reconstruct the future observation from the current observation and a latent sampled from the encoder distribution ś.sub.t+T=D.sub.Ø(z, s.sub.t), z˜N(μ.sub.g, σ.sub.G). The encoder distribution is regularized with a KL-loss KL(N μ.sub.g, σ.sub.G)∥N(0,1) with weight β.sub.g [15] to encourage the encoder distribution to match a prior latent distribution p(z)=N(0,1) so that at test-time, the decoder can be used as a conditional generative model by sampling latents z˜N (0,1) and passing them through the decoder.; Mandlekar, [0123] further teaches In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows identification of data points in new dataset 1212 that deviate from normal patterns of new dataset 1212.; Mandlekar, [0115] teaches In at least one embodiment, inference and/or training logic 1115 may include, without limitation, one or more arithmetic logic unit(s) (“ALU(s)”) 1110, including integer and/or floating point units, to perform logical and/or mathematical operations based, at least in part on, or indicated by, training and/or inference code (e.g., graph code), a result of which may produce activations (e.g., output values from layers or neurons within a neural network) stored in an activation storage 1120 that are functions of input/output and/or weight parameter data stored in code and/or data storage 1101 and/or code and/or data storage 1105.); and (b) means for performing a monitoring phase, wherein the monitoring phase includes: obtaining a second data set having at least one temporal characteristic of the at least one state parameter of the first robot or a second robot (Mandlekar, [0198] teaches In at least one embodiment, vehicle 1400 may include CPU(s) 1418 (e.g., discrete CPU(s), or dCPU(s)), that may be coupled to SoC(s) 1404 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, CPU(s) 1418 may include an X86 processor, for example. CPU(s) 1418 may be used to perform any of a variety of functions, including arbitrating potentially inconsistent results between ADAS sensors and SoC(s) 1404, and/or monitoring status and health of controller(s) 1436 and/or an infotainment system on a chip (“infotainment SoC”) 1430, for example.; Mandlekar [0236], teaches In at least one embodiment, server(s) 1478 may receive data from vehicles and apply data to up-to-date real-time neural networks for real-time intelligent inferencing. [Note: data to up-to-date real-time as second data having at least one “temporal characteristic”.]), and identifying at least one of the pattern groups of the trained second autoencoder within the second data set (Mandlekar, [0123] teaches in at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows identification of data points in new dataset 1212 that deviate from normal patterns of new dataset 1212.). Before the effective filing date of the claimed invention, it would have been obvious to one or ordinary skill in the art to have modified the method for anomaly detection in robotics, as taught by Azzalini, with the system that learns from task demonstrations of a robot comprising two variational autoencoders, as taught by Mandlekar, in order to allow for selective imitation of local sequences in the dataset. (Mandlekar paragraphs [0082] and [0083]) Regarding claim 24, Azzalini teaches (a) perform a training phase, including: obtaining a first data set having at least one temporal characteristic of at least one state parameter of a first robot (Azzalini, Abstract, teaches detection of anomalies and faults is a crucial ability for fully autonomous robots. This letter proposes a new deep learning-based minimally supervised method for detecting anomalies in autonomous robots. We contribute a new Variational Auto-Encoder architecture able to model very long multivariate sensor logs exploiting a new incremental training method, which induces a progress-based latent space that can be used to detect anomalies both at runtime and offline; Azzalini, pg. 2988, Section III.C teaches represent the temporal dependency of multivariate time series collected from robot sensors), and training an artificial neural network (Azzalini, Section I, col. 2, teaches we also introduce a new incremental method for training VAEs, which induces a progress-based latent space that can be used to detect anomalies both online (at runtime) and offline… Only few(even just one) labeled nominal executions are then required to partition the learned latent space into nominal and anomalous regions. This minimally supervised approach provides a big advantage over semi-supervised approaches in practical settings, where collecting several nominal runs of a robot performing a task could be hard, since a human expert is usually required to supervise the system in order to label runs as nominal.), However, Azzalini does not distinctly disclose: a computer program or computer program product comprising program code stored on a non-transient, computer-readable medium, … the artificial neural network including: a first autoencoder having an encoder that maps the first data set to temporal characteristic patterns and corresponding activation, and a decoder that reconstructs the first data set using the mapped temporal characteristic patterns, and a second autoencoder having an encoder that maps the temporal characteristic patterns and corresponding activation to pattern groups, and a decoder that reconstructs the temporal characteristic patterns and corresponding activation using the pattern groups; and (b) perform a monitoring phase, including: obtaining a second data set having at least one temporal characteristic of the at least one state parameter of the first robot or a second robot, and identifying at least one of the pattern groups of the trained second autoencoder within the second data set. Nevertheless, Mandlekar teaches: a computer program or computer program product comprising program code stored on a non-transient, computer-readable medium (Mandlekar, [0591] Operations of processes described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. In at least one embodiment, a process such as those processes described herein (or variations and/or combinations thereof) is performed under control of one or more computer systems configured with executable instructions and is implemented as code (e.g., executable instructions, one or more computer programs or one or more applications) executing collectively on one or more processors, by hardware or combinations thereof. In at least one embodiment, code is stored on a computer-readable storage medium, for example, in form of a computer program comprising a plurality of instructions executable by one or more processors. In at least one embodiment, a computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transitory signals (e.g., a propagating transient electric or electromagnetic transmission) but includes non-transitory data storage circuitry (e.g., buffers, cache, and queues) within transceivers of transitory signals. In at least one embodiment, code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media having stored thereon executable instructions (or other memory to store executable instructions) that, when executed (i.e., as a result of being executed) by one or more processors of a computer system, cause computer system to perform operations described herein.), … the artificial neural network including: a first autoencoder having an encoder that maps the first data set to temporal characteristic patterns and corresponding activation, and a decoder that reconstructs the first data set using the mapped temporal characteristic patterns, and a second autoencoder having an encoder that maps the temporal characteristic patterns and corresponding activation to pattern groups, and a decoder that reconstructs the temporal characteristic patterns and corresponding activation using the pattern groups (Mandlekar, [0089] teaches Learning from suboptimal data: In an embodiment, the low-level goal-conditioned controller operates for a small number of timesteps, so the controller has no need to account for suboptimal actions. This is because if the goal is to reach a state s.sub.2 from s.sub.1, and T is sufficiently small, then a policy may only be able to improve by reaching s.sub.2 in less than T steps, which is a negligible improvement for small values of T. By contrast, the value learning component of the goal selection mechanism explicitly accounts for suboptimal solution approaches by evaluating the expected task returns of each goal and selecting the goal with the highest return.; Mandlekar, [0090] further teaches Learning from off-policy datasets: Policy learning from arbitrary off-policy data can be challenging. Some embodiments of IRIS deal with this issue by constraining learning to occur within the distribution of training data. In some examples, the goal-conditioned controller directly imitates sequences from the training data, and the generative goal model is also trained to propose goal observations from the training data. Finally, the value learning component of the goal selection mechanism mitigates extrapolation error by making sure that the Q-network is queried on state-action pairs that lie within the training distribution.; Mandlekar, [0568] 13. Teaches The system of any of clauses 8 to 12, wherein the neural network is a variational autoencoder.; Mandlekar [0569] 14. further teaches The system of any of clauses 8 to 13, wherein the system uses a recurrent neural network to determine a set of actions that, as a result of being performed by the robot, reposition the robot from the current position to an intermediate goal.; Mandlekar, [0570] 15. Teaches A machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least: use a first neural network to determine a set of intermediate goal proposals based on a current position of a robot, the first neural network trained using a set of demonstrations of task performance; select, based at least in part on a value function, an intermediate goal from the set of intermediate goal proposals; use a second neural network to determine a set of actions that, as a result of being performed by the robot, reposition the robot from the current position to the selected intermediate goal; perform the task by at least performing the set of actions.; Mandlekar, [0085] teaches In one example, the cVAE is a conditional generative model that is trained on pairs of current and future observations (s.sub.t, s.sub.t+T) [i.e., “temporal characteristics”] sampled from trajectories in the dataset (lines 5-7 in Algorithm 1). An encoder maps a current and future observation to the parameters of a latent Gaussian distribution μ.sub.b, σ.sub.g=E.sub.Ø(s.sub.t+T, s.sub.t) [i.e., “temporal characteristics”] the decoder is trained to reconstruct the future observation from the current observation and a latent sampled from the encoder distribution ś.sub.t+T=D.sub.Ø(z, s.sub.t), z˜N(μ.sub.g, σ.sub.G). The encoder distribution is regularized with a KL-loss KL(N μ.sub.g, σ.sub.G)∥N(0,1) with weight β.sub.g [15] to encourage the encoder distribution to match a prior latent distribution p(z)=N(0,1) so that at test-time, the decoder can be used as a conditional generative model by sampling latents z˜N (0,1) and passing them through the decoder.; Mandlekar, [0123] further teaches In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows identification of data points in new dataset 1212 that deviate from normal patterns of new dataset 1212.; Mandlekar, [0115] teaches In at least one embodiment, inference and/or training logic 1115 may include, without limitation, one or more arithmetic logic unit(s) (“ALU(s)”) 1110, including integer and/or floating point units, to perform logical and/or mathematical operations based, at least in part on, or indicated by, training and/or inference code (e.g., graph code), a result of which may produce activations (e.g., output values from layers or neurons within a neural network) stored in an activation storage 1120 that are functions of input/output and/or weight parameter data stored in code and/or data storage 1101 and/or code and/or data storage 1105.); and (b) perform a monitoring phase, including: obtaining a second data set having at least one temporal characteristic of the at least one state parameter of the first robot or a second robot (Mandlekar, [0198] teaches In at least one embodiment, vehicle 1400 may include CPU(s) 1418 (e.g., discrete CPU(s), or dCPU(s)), that may be coupled to SoC(s) 1404 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, CPU(s) 1418 may include an X86 processor, for example. CPU(s) 1418 may be used to perform any of a variety of functions, including arbitrating potentially inconsistent results between ADAS sensors and SoC(s) 1404, and/or monitoring status and health of controller(s) 1436 and/or an infotainment system on a chip (“infotainment SoC”) 1430,.; Mandlekar [0236], teaches In at least one embodiment, server(s) 1478 may receive data from vehicles and apply data to up-to-date real-time neural networks for real-time intelligent inferencing. [Note: data to up-to-date real-time as second data having at least one “temporal characteristic”.]), and identifying at least one of the pattern groups of the trained second autoencoder within the second data set (Mandlekar, [0123] teaches in at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows identification of data points in new dataset 1212 that deviate from normal patterns of new dataset 1212.). Claims 13 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Azzalini in view of Mandlekar, as applied to claim 11, and further in view of Li et al., “Spatio-Temporal Graph Dual-Attention Network for Multi-Agent Prediction and Tracking, (July, 2021) Regarding claim 13, Azzalini in view of Mandlekar teaches all of the limitations of claim 11, however, the combination does not distinctly disclose wherein the encoder of the second autoencoder includes at least one attention-based artificial neural network. Nevertheless, Li teaches wherein the encoder of the second autoencoder includes at least one attention-based artificial neural network (Li, Fig. 2 teaches a Graph Dual-Attention Network including an autoencoder with an encoding function and decoding function; Li, pg. 10561, col. 1, further teaches we also employ the multi-head attention mechanism to boost model performance. And, further teaching the multi-head attention mechanism can also be employed by learning different ω [i.e., a weight vector parametrizing the attention function] and fusing the information by averaging or concatenation operations.) Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified the method for anomaly detection in robotics, as taught by Azzalini in view of Mandlekar, to further include the generative neural system for multi-agent trajectory prediction comprising an autoencoder with a graph dual-attention network, as taught by Li, in order to figure out which of the other agents have the most significant influence on a certain agent, as well as the relative importance of different time steps. Furthermore, an effective understanding of the environmental and accurate trajectory prediction of surrounding dynamic obstacles are indispensable for intelligent mobile systems (e.g., autonomous vehicles, social robots). (Li, pg. 10557, col. 2 and Abstract) Regarding claim 14, Azzalini in view of Mandlekar and Li teaches all of the limitations of claim 13, and the combination further teaches wherein the at least one attention-based artificial neural network is at least one multi-head attention block (Li, Fig. 2 teaches a Graph Dual-Attention Network including an autoencoder with an encoding function and decoding function; Li, pg. 10561, col. 1, further teaches we also employ the multi-head attention mechanism to boost model performance. And, further teaching the multi-head attention mechanism can also be employed by learning different ω [i.e., a weight vector parametrizing the attention function] and fusing the information by averaging or concatenation operations.; [Note: Mandlekar teaches first and second autoencoders]). Motivation to combine same as stated in claim 13. Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Azzalini in view of Mandlekar, as applied to claim 11, and further in view of Kosiorek, et al., “Stacked Capsule Autoencoders” (2019) (Applicant Admitted Prior Art in view of page 12 of the Specification) Regarding claim 15, the combination of Mandlekar in view of Li teaches all of the limitations of claim 11, however, Mandlekar does not distinctly disclose wherein the decoder of the second autoencoder has at least one capsule neural network. Nevertheless, Kosiorek teaches wherein the decoder of the second autoencoder has at least one capsule neural network (Kosiorek, pg. 2 teaches Stacked Capsule Autoencoders (Section 2) capture spatial relationships between whole objects and their parts when trained on unlabeled data. The vectors of presence probabilities for the object capsules tend to form tight clusters (cf. Figure 1), and when we assign a class to each cluster we achieve state-of-the-art results for unsupervised classification on SVHN (55%) and MNIST (98.7%). ; Kosiorek, pg. 9, Section 5 teaches the main contribution of our work is a novel method for representation learning, in which highly structured decoder networks are used to train one encoder network that can segment an image into parts and their poses and another encoder network than can compose the parts into coherent wholes. Even though our training objective is not concerned with classification or clustering, SCAE is the only method that achieves competitive results in unsupervised object classification without relying on mutual information. This is significant since, unlike our method, MI-based methods require sophisticated data augmentation.) Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified the method for anomaly detection in robotics, as taught by Azzalini in view of Mandlekar, to further include the unsupervised learning of stacked capsule autoencoders (SCAE), as taught by Kosiorek, as SCAE is the only method that achieves competitive results in unsupervised object classification without relying on mutual information. (Kosiorek, pg. 9, Section 5) Conclusion The following prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Chen et al., “Unsupervised Anomaly Detection of Industrial Robots Using Sliding-Window Convolutional Variational Autoencoder”, (March, 2020) Temlay et al. (US 20220383019 A1) – Fig. 1A, [0100], and [0101] Monti et al., “DAG-Net: Double Attentive Graph Neural Network for Trajectory Forecasting, (Jan. 2021) Rong et al., “Attention-based Sampling Distribution for Motion Planning in Autonomous Driving” (July 2020) Omaima El Alaoui-Elfels et al., “From Autoencoders to Capsule Networks: A Survey” (2020) – teaches stacked capsule auto-encoders as disclosed in Kosiorek et al. Any inquiry concerning this communication or earlier communications from the examiner should be directed to BEATRIZ RAMIREZ BRAVO whose telephone number is 571-272-2156. The examiner can normally be reached Mon. - Fri. 7:30a.m.-5:00p.m.. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, USMAAN SAEED can be reached at 571-272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /B.R.B./ Examiner, Art Unit 2146 /USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146
Read full office action

Prosecution Timeline

Feb 12, 2024
Application Filed
Sep 08, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737587
METAL DETECTION SYSTEM
3y 0m to grant Granted Sep 15, 2026
Patent 12651147
Systems and Methods for Training Conditional Generative Models
2y 10m to grant Granted Jun 09, 2026
Patent 12646619
SYSTEMS AND METHODS FOR GENERATING ALIMENTARY INSTRUCTION SETS BASED ON VIBRANT CONSTITUTIONAL GUIDANCE
6y 4m to grant Granted Jun 02, 2026
Patent 12632791
SYSTEM AND METHOD FOR CONFIGURING AN ARTIFICIAL INTELLIGENCE PIPELINE
2y 5m to grant Granted May 19, 2026
Patent 12632704
NEURAL PROCESSING UNIT INCLUDING POST-PROCESSING UNIT
1y 8m to grant Granted May 19, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
64%
Grant Probability
93%
With Interview (+28.9%)
4y 6m (~1y 11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 100 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month