Prosecution Insights
Last updated: October 01, 2026
Application No. 18/921,174

HUMAN MOTION UNDERSTANDING USING STATE SPACE MODELS

Final Rejection §103
Filed
Oct 21, 2024
Priority
Nov 03, 2023 — provisional 63/547,202 +1 more
Examiner
LIU, GORDON G
Art Unit
2618
Tech Center
2600 — Communications
Assignee
Apple Inc.
OA Round
2 (Final)
83%
Grant Probability
Favorable
3-4
OA Rounds
2m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
581 granted / 701 resolved
+20.9% vs TC avg
Moderate +15% lift
Without
With
+14.8%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 2m
Avg Prosecution
36 currently pending
Career history
720
Total Applications
across all art units

Statute-Specific Performance

§101
7.1%
-32.9% vs TC avg
§103
77.3%
+37.3% vs TC avg
§102
3.4%
-36.6% vs TC avg
§112
2.7%
-37.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 701 resolved cases

Office Action

§103
DETAILED ACTION The Office Action is in response to the Applicants' communication filed on August 4, 2026, which amends the independent claims 1 and 11-12, amends the dependent claims 3 and 14, and presents arguments, is hereby acknowledged. Claims 1-20 are currently pending and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment Applicant’s arguments filed on August 4, 2026, have been fully considered. Applicant argues that by this response, the independent claims 1and 11-12 are hereby amended to add a new limitation “the 2D information based on frames comprising images capturing the continuous time light signal at more than one frame rate” and “, the state space model is a continuous time learnable framework for mapping between continuous time scalar 2D inputs and continuous time scalar 3D outputs” in order to overcome the 35 U.S.C. §103 rejection. Examiner replies that the amended claims with new limitations may overcome the cited portions of the prior arts. However, a newly found art, Delachanal (US 10187607 B1) teaches that the 2D information based on frames comprising images capturing the continuous time light signal at more than one frame rate (See Delachanal: Figs. 1-4, and Col. 3 Lines 66-67~ Col. 4 Lines 1-17, “FIG. 1 illustrates a system 10 that uses a variable capture frame rate for video capture. The system 10 may include one or more of a processor 11, an electronic storage 12, an interface 13 (e.g., bus, wireless interface), an image sensor 14, a motion sensor 15, and/or other components. The image sensor 14 may be configured to generate visual output signals conveying visual information within a field of view of the image sensor 14. The motion sensor 15 may be configured to generate motion output signals conveying motion information of the image sensor 14. First video information defining first video content and second video information defining second video content may be generated based on the visual output signals. The first video content may be captured using a first capture frame rate set to a first value and the second video content may be captured using a second capture frame rate set to a second value. The values of the first capture frame rate and the second capture frame rate may define the numbers of frames captured per a duration of time”; and Col. 8 Lines 43-55, “The combination component 108 may be configured to generate third video information defining third video content based on the first video information, the second video information, and/or other information. The third video information may be generated based on the first value defining a lower number of frames captured per the duration of time than (1) the second value, and (2) the third value. The third video content may include one or more frames of the first video content, one or more frames of the second video content, and/or other frames. The third video content may include a combination of some or all of the first video content, some or all of the second video content, and/or other video content”. Note that the first video is captured at the first frame rate, the second video is captured at the second frame rate, and the third video is generated by combining the first video and the second video). Further, Wang, etc. (US 20120219186 A1) teaches that the state space model is a continuous time learnable framework for mapping between continuous time scalar 2D inputs and continuous time scalar 3D outputs (See Wang: Figs. 1-3, and [0028], “As previously noted, standard SLDS based on first-order Markov assumption does not yield the best performance, particularly when dealing with between action transitions. Presented herein are embodiments of dynamic-system-based framework to model the temporal evolution of a feature for action analysis. In embodiments, the framework utilizes higher-order information to explicitly model the transition within single action primitive or between successive action primitives. In embodiments, the framework, referred to as the Continuous Linear Dynamic System (CLDS) framework, comprises two sets of Linear Dynamic System (LDS) models, one to model the dynamics of individual primitive actions and the other to model the transition between actions. In embodiments, the inference process estimates the best decomposition of a whole sequence by continuously alternating between the two set of models. In embodiments, an approximate Viterbi algorithm may be used in the inference process. Using the CLDS framework, both action type and action boundary may be accurately recognized”; [0094]. “FIG. 5 depicts a method for using a Continuous Linear Dynamic System (CLDS) to detect and label actions in a video according to embodiments of the present invention. As depicted in FIG. 5, the process commences by segmenting (505) input sensor data into time frames. In embodiments, the input sensor data may be video data and the segmented time frames may be image frames, although other sensor data and configurations may be used. Given the segmented input sensor data, a feature for each image frame is generated (510). In embodiments, an image feature for the frame may be an embedded optical flow as previously discussed. These images features are then input into a CLDS model to perform (515) continuous segmentation and recognition. In embodiments, the CLDS model may use one or more of the inference methods discussed above or known to those of ordinary skill in the art’; [0095], “FIG. 6 depicts a block diagram of a Continuous Linear Dynamic System (CLDS) model detector 605 according to embodiments of the present invention. The CLDS model detector 605 receives input sensor data 625 and outputs a sequence of labels 630. As shown in FIG. 6, the CLDS model detector 605 comprises a frame extractor 610, a feature extractor 615, and a CLDS model decoder 620. In embodiments, CLDS model detector performs one or more methods for continuous segmentation and recognition, which include but are not limited to the methods discussed above”. Note that input sensor data 625 is mapped to the continuous 2D scalar input, and output label sequence 630 is mapped to the continuous 3D scalar outputs). The remaining argument of the applicant are mooted in view of the newly found arts. Examiner respectfully further replies that the Applicant's arguments have been fully considered and a new ground of rejections have been made. Accordingly, new grounds of rejection are set forth below. Since the new grounds of rejection are necessitated by Applicant's amendments to the claims, the present action is made final. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Merler, etc. (US 20140010456 A1) in view of Dekel, etc. (US 20220215568 A1), further in view of Cherevatsky, etc. (US 10699421 B1), Hoffberg, etc. (US 6400996 B1), Delachanal (US 10187607 B1) and Wang, etc. (US 20120219186 A1). Regarding claim 1, Merler teaches that a method (See Merler: Figs. 1-2, and [0038], "FIG. 1 and FIG. 2 provide a is a flow diagram that illustrates the method for tracking objects according to an exemplary embodiment of the disclosed subject matter and a computer system having an object tracking system in accordance with an embodiment of the disclosed subject matter, respectively") comprising: at a device having a processor (See Merler: Fig. 2, and [0045], "With reference to FIG. 2, a description of a computer system configured for object tracking according to an embodiment of the presently disclosed subject matter will be described. The system depicted in FIG. 2 can carry out the method depicted in FIG. 1. The system can include a processor 230 in communication with a memory 220"): obtaining two-dimensional (2D) information corresponding to a continuous time light signal providing information about a user in a three-dimensional (3D) environment, the 2D information based on frames comprising images capturing the continuous time light signal at more than one frame rates (See Merler: Fig. 2, and [0046], "The system can include providing video data 211 of the video input 210 containing at least one object 211, the video data corresponding to a frame of video that corresponds to a position t. In some embodiments, the video input 210 can include a plurality of frames. The video data 211 that makes up the frames can be stored in a memory 220, for example in random access memory with the use of the processor 230 and accompanying 1/0 functionality. The video data 211 stored in the memory can comprise all of the data for the video input 210, for example, in a situation in which the video input 210 is pre-recorded video. Alternatively, the video data 211 can comprise one or more frames of the video input 210, for example, in a situation in which video is being captured by a video capture device and streamed in real time to the computer system 200". Note that the video captured device captures images which are streamed in real time to the system is mapped to the 2D information corresponding to a continuous time light signal because the video is a series of frame, and each frame is 2D image captured by the video camera with a specific exposed time period to a continuous light, that is, a frame rate. But Merler does not teach explicitly that the object is in 3D environment, and a second art will be searched and cited below); obtaining discretization information corresponding to the more than one frame rates (See Merler: Fig. 1, and [0043], "Positions 101 through 106 can then be repeated at each position in the interval. That is, the position t can be increased to the next position t+1 (step 107). Steps 101 through 106 can then be applied to the frame in the video input corresponding to position t+1". Note that the time interval t + n is discretization information corresponding to the frame rate, because the frame rate (R) and the time interval Mis mathematically related by the formula R=l/ M, and in digital field, t + n is actually the physical time = t + n M); and determining 3D information about the user by inputting the 2D information and the discretization information into a state space model (See Merler: Figs. 1-2, and [0042], "A final estimate of position and/or velocity of the object at a position can be calculated (step 105) at position t with the use of a Kalman filter 106. A Kalman gain 104 can be used to weight the importance of the noisy measurement component and noisy predictor component of the Kalman filter 106. The Kalman filter can be a general Kalman filter. Alternatively, in some embodiments, the Kalman filter can be a steady-state Kalman filter (i.e, the Kalman gain can be predetermined)"; and [0047], "An object detector 240 can obtain a first location estimate at position t. An object tracker 250 can obtain a second location estimate and a movement estimate at position t. The first location estimate can represent the noisy measurement component 251 of the Kalman filter 106, and the second location estimate and movement estimate can represent the noisy predictor component 252 of the Kalman filter 106. A final estimate 265 of position and/or velocity of the object at a future time t+n, where n>0 can be calculated at each position t with a Kalman filter 260. The final estimate can be calculated with reference to a Kalman gain 267". Note that Kalman filter is a SSM (state space model) that estimates the state of the system from a series of noisy measurement, as shown here that the Kalman filter is used to estimate the position and velocity based on the noisy measurements, that is, Kalman filter is a state space model), the state space model is a continuous time learnable framework for mapping between continuous time scalar 2D inputs and continuous time scalar 3D outputs. However, Merler fails to explicitly disclose that corresponding to a continuous time light signal providing information about a user in a three-dimensional (3D) environment, the 2D information based on frames comprising images capturing the continuous time light signal at more than one frame rate; and the state space model is a continuous time learnable framework for mapping between continuous time scalar 2D inputs and continuous time scalar 3D outputs. However, Dekel teaches that corresponding to a continuous time light signal providing information about a user in a three-dimensional (3D) environment (See Dekel: Fig. 3, and [0055], "FIG. 3 illustrates an example video that may be processed by the systems and operations herein disclosed to determine depth values of features represented within the video. Namely, FIG. 3 illustrates video 310 generated by a monoscopic camera 300. Video 310 includes images 302, 303, 304, 305, 306, 307, and 308-309 (i.e., images 302-309), which may alternatively be referred to as frames or image frames of the video. Images 302-309 may represent static features of an environment (e.g., boxes 314 and 316) and moving features of the environment (e.g., human 312). Static features may include objects and other physical features that are expected to remain stationary for a predetermined period of time, such as building structures, trees, roads, or sidewalks, among other possibilities. Moving features may include objects and other physical features that are expected to move within the predetermined period of time, such as humans, animals, or vehicles, among other possibilities. Thus, some moving features change their positions during video 310, while static features may remain in fixed positions during video 310". Note that the environment with boxes 314 and 316 clearly shows that the user in the 3D environment, and a series of 2D images (continuous time light signals) of the user are captured by the monoscopic camera 300, and this is mapped to the cited limitation of: corresponding to a continuous time light signal providing information about a user in a three-dimensional (3D) environment"). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Merler to have corresponding to a continuous time light signal providing information about a user in a three-dimensional (3D) environment as taught by Dekel in order to allow a machine learning (ML) model to be trained to determine the depth values of the moving features in the target image based on the static depth image and the object mask, so that the ML model is trained to generate the dynamic depth image in an efficient manner (See Dekel: Fig. 5, and [0067], "Thus, optical flow image 504 provides a basis for determining the static depth image using motion parallax between target image 308 and reference image 302. Specifically, if the relative pose of the camera is accounted for, optical flow image 504 may be used to determine the depth of static features. In the case of static features, when the effect of camera pose on the optical flow field is removed, the optical flow field represented in image 504 is attributable to the depths of the static features, thus allowing for determination of these depths"). Merler teaches a method and system that may track one or more objects at each position in an interval in a video input with the use of Kalman filter to estimate the first location, second location and movements of the objects; and final location and movement of the objects using Kalman filter; while Dekel teaches a system and method that may determine the depth information for the object and user in the 3D environment. Therefore, it is obvious to one of ordinary skill in the art to modify Merler by Dekel to have the continuous time light signal (2D images) providing information about a user in the 3D environment. The motivation to modify Merler by Dekel is "Use of known technique to improve similar devices (methods, or products) in the same way". However, Merler, modified by Dekel, fails to explicitly disclose that the 2D information based on frames comprising images capturing the continuous time light signal at more than one frame rate; and the state space model is a continuous time learnable framework for mapping between continuous time scalar 2D inputs and continuous time scalar 3D outputs. However, Cherevatsky teaches that the state space model is a continuous time framework for mapping between continuous time scalar 2D inputs and continuous time scalar 3D outputs (See Cherevatsky: Fig. 3, and Col. 21 Lines 34-58, "At box 325, an initial point cloud is defined from depth image frames captured from the scene using one or more of the depth cameras. For example, where a depth image containing information relating to distances of surfaces of objects within a scene from a perspective of a depth camera is captured, the depth image may be converted into a 3D representation of the physical topography of the scene from that perspective using ranging information for one or more of the pixels provided in the depth image and parameters of the depth camera, e.g., a set of coordinates of the imaging sensor or other components of the depth camera. Two or more depth images captured using RGBD cameras from different perspectives may be further utilized to enhance the quality of the 3D representation of the scene. At box 330, visual cameras having the target object in view within visual image frames captured from the scene are determined. For example, where the 3D bounding region has been defined at box 310, an extent to which a 2D projection of the 3D bounding region appears within the fields of view of each of the imaging devices is determined. At box 332, the extent to which pixels corresponding to the target object are occluded (or not occluded) within the 2D projection of the 3D bounding region is determined, e.g., by comparing depth data for the target cloud points to depth data for other scene points within a frustrum spanned by the 3D bounding region"; and Fig. 7, and Col. 33 Lines 21-24, "At box 752, the probability map for the position of the target object is provided to a Kalman filter or another set of mathematical equations for estimating the position of the target object in a manner that minimizes a mean of the squared errors associated with the position. At box 754, the Kalman filter models motion of the target object based on probability maps determined for all known synchronization points, e.g., synchronization points ranging from 1 to i". Note that Kalman filter (the SSM) is used and 2D images is fused to get the 3D bounding boxes and position of the tracking object, and this is mapped to "mapping between 2D inputs and 3D outputs". However, Cherevatsky does not teach that the SSM is a learnable framework, and another art will be searched below). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Merler to have the state space model is a continuous time framework for mapping between continuous time 2D scalar inputs and continuous time scalar 3D outputs as taught by Cherevatsky in order to improve upon the computer-based tracking of target objects (See Cherevatsky: Fig. 1, and Col. 15 Lines 28-34, "Therefore, by using visual images and depth images to determine positions in 3D space, and training tracking algorithms to recognize objects based on such determined positions, some implementations of the systems and methods of the present disclosure may improve upon the computer-based tracking of target objects, thereby solving a fundamental computer vision problem"). Merler teaches a method and system that may track one or more objects at each position in an interval in a video input with the use of Kalman filter to estimate the first location, second location and movements of the objects; and final location and movement of the objects using Kalman filter; while Cherevatsky teaches a system and method that may determine the 3D outputs (location and bounding box) from 2D continuous input images using the Kalman filters. Therefore, it is obvious to one of ordinary skill in the art to modify Merler by Cherevatsky to mapping the 2D inputs into 3D outputs using the Kalman filters. The motivation to modify Merler by Cherevatsky is "Use of known technique to improve similar devices (methods, or products) in the same way". However, Merler, modified by Dekel and Cherevatsky, fails to explicitly disclose that the 2D information based on frames comprising images capturing the continuous time light signal at more than one frame rate; and the state space model is a continuous time learnable framework for mapping between continuous time scalar 2D inputs and continuous time scalar 3D outputs. However, Hoffberg teaches that the state space model is the state space model is a continuous time learnable framework (See Hoffberg: Fig. 30, and Col. 15 Lines 33-64, "Methods employing other than fractal-based algorithms may also be used. See, e.g., Liu, Y., "Pattern recognition using Hilbert space", Proceedings of the SPIE--The International Society for Optical Engineering, 1825:63-77 (1992), which describes a learning approach, the Hilbert learning. This approach is similar to Fractal learning, but the Fractal part is replaced by Hilbert space. Like the Fractal learning, the first stage is to encode an image to a small vector in the internal space of a learning system. The next stage is to quantize the internal parameter space. The internal space of a Hilbert learning system is defined as follows: a pattern can be interpreted as a representation of a vector in a Hilbert space. Any vectors in a Hilbert space can be expanded. If a vector happens to be in a subspace of a Hilbert space where the dimension L of the subspace is low (order of 10), the vector can be specified by its norm, an L-vector, and the Hermitian operator which spans the Hilbert space, establishing a mapping from an image space to the internal space P. This mapping converts an input image to a 4-tuple: tin P=(Norm, T, N, L-vector), where Tis an operator parameter space, N is a set of integers which specifies the boundary condition. The encoding is implemented by mapping an input pattern into a point in its internal space. The system uses local search algorithm, i.e., the system adjusts its internal data locally. The search is first conducted for an operator in a parameter space of operators, then an error function delta (t) is computed. The algorithm stops at a local minimum of delta (t). Finally, the input training set divides the internal space by a quantization procedure. See also, Liu, Y., "Extensions of fractal theory", Proceedings of the SPIE--The International Society for Optical Engineering, 1966:255-68(1993)"; and Fig. 15, and Col. 50 Lines 59-62, "The interface may therefore provide a model of the user, which is employed in a predictive algorithm. The model parameters may be static (once created) or dynamic, and may be adaptive to the user or alterations in the use pattern". Note that the adaptive pattern recognition algorithm map the 2D inputs into high dimensional output (3D outputs) with parameters adapted to the input data, and this parameter adaptive to the inputs is mapped to the learnable framework). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Merler to have the state space model is a continuous time learnable framework as taught by Hoffberg in order to provide the operator with small number of high probability choices (See Hoff berg: Fig. 15, and Col. 52 Lines 7-14, "A major theme of the present invention is the use of intelligent, adaptive pattern recognition in order to provide the operator with a small number of high probability choices, which may be complex, without the need for explicit definition of each atomic instruction comprising the desired action. The interface system predicts a desired action based on the user input, a past history of use, a context of use, and a set of predetermined or adaptive rules"). Merler teaches a method and system that may track one or more objects at each position in an interval in a video input with the use of Kalman filter to estimate the first location, second location and movements of the objects; and final location and movement of the objects using Kalman filter; while Hoffberg teaches a system and method that may map the 2D input images to high dimensional outputs such as 3D outputs using the adaptive pattern recognition with parameters learned. Therefore, it is obvious to one of ordinary skill in the art to modify Merler by Hoffberg to have Kalman filter parameters learned. The motivation to modify Merler by Hoffberg is "Use of known technique to improve similar devices (methods, or products) in the same way". However, Merler, modified by Dekel, Cherevatsky and Hoffberg, fails to explicitly disclose that the 2D information based on frames comprising images capturing the continuous time light signal at more than one frame rate; and the state space model is a continuous time learnable framework for mapping between continuous time scalar 2D inputs and continuous time scalar 3D outputs. However, Delachanal teaches that the 2D information based on frames comprising images capturing the continuous time light signal at more than one frame rate (See Delachanal: Figs. 1-4, and Col. 3 Lines 66-67~ Col. 4 Lines 1-17, “FIG. 1 illustrates a system 10 that uses a variable capture frame rate for video capture. The system 10 may include one or more of a processor 11, an electronic storage 12, an interface 13 (e.g., bus, wireless interface), an image sensor 14, a motion sensor 15, and/or other components. The image sensor 14 may be configured to generate visual output signals conveying visual information within a field of view of the image sensor 14. The motion sensor 15 may be configured to generate motion output signals conveying motion information of the image sensor 14. First video information defining first video content and second video information defining second video content may be generated based on the visual output signals. The first video content may be captured using a first capture frame rate set to a first value and the second video content may be captured using a second capture frame rate set to a second value. The values of the first capture frame rate and the second capture frame rate may define the numbers of frames captured per a duration of time”; and Col. 8 Lines 43-55, “The combination component 108 may be configured to generate third video information defining third video content based on the first video information, the second video information, and/or other information. The third video information may be generated based on the first value defining a lower number of frames captured per the duration of time than (1) the second value, and (2) the third value. The third video content may include one or more frames of the first video content, one or more frames of the second video content, and/or other frames. The third video content may include a combination of some or all of the first video content, some or all of the second video content, and/or other video content”. Note that the first video is captured at the first frame rate, the second video is captured at the second frame rate, and the third video is generated by combining the first video and the second video). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Merler to have the 2D information based on frames comprising images capturing the continuous time light signal at more than one frame rate as taught by Delachanal in order to enable system to change capture frame rates based on the determined motion (See Delachanal berg: Fig. 1, and Col. 7 Lines 3-13, " The motion component 104 may be configured to determine motion of the image sensor 14 and/or motion of one or more objects within the field of view of the image sensor 14. Determining the motion of the image sensor 14 and/or the object(s) within the field of view of the image sensor 14 may enable system 10 to change one or more capture frame rates based on the determined motion(s). For example, one or more capture frame rates may be changed based on acceleration/deceleration of the image sensor 14 and/or acceleration/deceleration of the object(s) within the field of view of the image sensor 14 "). Merler teaches a method and system that may track one or more objects at each position in an interval in a video input with the use of Kalman filter to estimate the first location, second location and movements of the objects; and final location and movement of the objects using Kalman filter; while Delachanal teaches a system and method that may change the frame rate to capture the image based on the detected motion information. Therefore, it is obvious to one of ordinary skill in the art to modify Merler by Delachanal to have multiple frame rate for capturing the continuous time 2D input images. The motivation to modify Merler by Delachanal is "Use of known technique to improve similar devices (methods, or products) in the same way". However, Merler, modified by Dekel, Cherevatsky, Hoffberg and Delachanal, fails to explicitly disclose that the state space model is a continuous time learnable framework for mapping between continuous time scalar 2D inputs and continuous time scalar 3D outputs. However, Wang teaches that the state space model is a continuous time learnable framework for mapping between continuous time scalar 2D inputs and continuous time scalar 3D outputs (See Wang: Figs. 1-3, and [0028], “As previously noted, standard SLDS based on first-order Markov assumption does not yield the best performance, particularly when dealing with between action transitions. Presented herein are embodiments of dynamic-system-based framework to model the temporal evolution of a feature for action analysis. In embodiments, the framework utilizes higher-order information to explicitly model the transition within single action primitive or between successive action primitives. In embodiments, the framework, referred to as the Continuous Linear Dynamic System (CLDS) framework, comprises two sets of Linear Dynamic System (LDS) models, one to model the dynamics of individual primitive actions and the other to model the transition between actions. In embodiments, the inference process estimates the best decomposition of a whole sequence by continuously alternating between the two set of models. In embodiments, an approximate Viterbi algorithm may be used in the inference process. Using the CLDS framework, both action type and action boundary may be accurately recognized”; [0094]. “FIG. 5 depicts a method for using a Continuous Linear Dynamic System (CLDS) to detect and label actions in a video according to embodiments of the present invention. As depicted in FIG. 5, the process commences by segmenting (505) input sensor data into time frames. In embodiments, the input sensor data may be video data and the segmented time frames may be image frames, although other sensor data and configurations may be used. Given the segmented input sensor data, a feature for each image frame is generated (510). In embodiments, an image feature for the frame may be an embedded optical flow as previously discussed. These images features are then input into a CLDS model to perform (515) continuous segmentation and recognition. In embodiments, the CLDS model may use one or more of the inference methods discussed above or known to those of ordinary skill in the art’; [0095], “FIG. 6 depicts a block diagram of a Continuous Linear Dynamic System (CLDS) model detector 605 according to embodiments of the present invention. The CLDS model detector 605 receives input sensor data 625 and outputs a sequence of labels 630. As shown in FIG. 6, the CLDS model detector 605 comprises a frame extractor 610, a feature extractor 615, and a CLDS model decoder 620. In embodiments, CLDS model detector performs one or more methods for continuous segmentation and recognition, which include but are not limited to the methods discussed above”. Note that input sensor data 625 is mapped to the continuous 2D scalar input, and output label sequence 630 is mapped to the continuous 3D scalar outputs). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Merler to have the state space model is a continuous time learnable framework for mapping between continuous time scalar 2D inputs and continuous time scalar 3D outputs as taught by Wang in order to enable a continuous linear dynamic system (LDS) to gain flexibility during inference process provided with action LDS and transition LDS (See Wang: Fig. 1, and [0061], " Having both the Action LDS and the Transition LDS, CLDS gains more flexibility during the inference process. In embodiment, an observation-dependent transition probability (through the hidden layer X) may be calculated using the corresponding set of LDSs, determined by both the current and the next action labels. In embodiments, the inference process of CLDS may also use approximate Viterbi algorithm, where F is selected by "). Merler teaches a method and system that may track one or more objects at each position in an interval in a video input with the use of Kalman filter to estimate the first location, second location and movements of the objects; and final location and movement of the objects using Kalman filter; while Wang teaches a system and method that may comprise two sets of Linear Dynamic System (LDS) models, one to model the dynamics of individual primitive actions and the other to model the transitions between actions to extract features from the input image and infer the motion label outputs (mapped to scalar 3D outputs). Therefore, it is obvious to one of ordinary skill in the art to modify Merler by Wang to use the continuous time LDS model to map the 2D inputs to 3D outputs. The motivation to modify Merler by Wang is "Use of known technique to improve similar devices (methods, or products) in the same way". Regarding claim 2, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 1 as outlined above. Further, Merler teaches that the method of claim 1, wherein the discretization information comprises delta information corresponding to time periods between the frames (See Merler: Fig. 1, and [0043], "Positions 101 through 106 can then be repeated at each position in the interval. That is, the position t can be increased to the next position t+1 (step 107). Steps 101 through 106 can then be applied to the frame in the video input corresponding to position t+1". Note that the time interval t + n is discretization information corresponding to the frame rate, because the frame rate (R) and the time interval M between two consecutive frames is mathematically related by the formula R=l/ M, and in digital field, t + n is actually the physical time= t + n M, for example, if the first frame is at time t, the next frame is t+1, in physical time is t + M with n = 1). Regarding claim 3, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 1 as outlined above. Further, Merler teaches that the method of claim 1, wherein the discretization information comprises information (See Merler: Fig. 1, and [0047], "An object detector 240 can obtain a first location estimate at position t. An object tracker 250 can obtain a second location estimate and a movement estimate at position t. The first location estimate can represent the noisy measurement component 251 of the Kalman filter 106, and the second location estimate and movement estimate can represent the noisy predictor component 252 of the Kalman filter 106. A final estimate 265 of position and/or velocity of the object at a future time t+n, where n>0 can be calculated at each position t with a Kalman filter 260. The final estimate can be calculated with reference to a Kalman gain 267". Note that the time interval t + n is discretization information associated with the at least one or more frame rates, because the frame rate (R) and the time interval M between two consecutive frames is mathematically related by the formula M = 1/R, and in digital field, t + n is actually the physical time= t + n M = t + n (1/R), for example, if the first frame is at time t, the next frame is t+1, in physical time is t + 1/R with n = 1) associated with the more frame rates (See Delachanal: Figs. 1-4, and Col. 3 Lines 66-67~ Col. 4 Lines 1-17, “FIG. 1 illustrates a system 10 that uses a variable capture frame rate for video capture. The system 10 may include one or more of a processor 11, an electronic storage 12, an interface 13 (e.g., bus, wireless interface), an image sensor 14, a motion sensor 15, and/or other components. The image sensor 14 may be configured to generate visual output signals conveying visual information within a field of view of the image sensor 14. The motion sensor 15 may be configured to generate motion output signals conveying motion information of the image sensor 14. First video information defining first video content and second video information defining second video content may be generated based on the visual output signals. The first video content may be captured using a first capture frame rate set to a first value and the second video content may be captured using a second capture frame rate set to a second value. The values of the first capture frame rate and the second capture frame rate may define the numbers of frames captured per a duration of time”; and Col. 8 Lines 43-55, “The combination component 108 may be configured to generate third video information defining third video content based on the first video information, the second video information, and/or other information. The third video information may be generated based on the first value defining a lower number of frames captured per the duration of time than (1) the second value, and (2) the third value. The third video content may include one or more frames of the first video content, one or more frames of the second video content, and/or other frames. The third video content may include a combination of some or all of the first video content, some or all of the second video content, and/or other video content”. Note that the first video is captured at the first frame rate, the second video is captured at the second frame rate, and the third video is generated by combining the first video and the second video). Regarding claim 4, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 1 as outlined above. Further, Cherevatsky teaches that the method of claim 1, wherein the 2D information comprises information associated with 2D locations of joints of the user (See Cherevatsky: Fig. 7, and Col. 34 Lines 20-35, "At box 770, the tracklet for the target object over the tracking period is defined based on the probability maps and the point clouds defined from the visual image frames and the depth image frames captured at the prior synchronization points i. For example, a voting algorithm may be used to estimate a joint object position probability distribution in 3D space based on representations of the target object in 2D images captured by the plurality of imaging devices, and recognized therein using a tracking algorithm, such as an OpenCV tracker or a KCF tracker. Such representations may be projected onto the point clouds, and a tracklet of the positions of the target object may be determined accordingly, such as by assigning scores to each of the points in 3D space at various times, aggregating scores for such points, and selecting a best candidate based on the aggregated scores". Note that the joint object position is mapped to the information associated with 2D locations of joints of the user). Regarding claim 5, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 1 as outlined above. Further, Merler teaches that the method of claim 1, wherein the 3D information provides a 3D model representing at least a portion of the user (See Merler: Fig. 1, and [0039], "With reference to FIG. 1, the method of some embodiments can include providing a video input that contains at least one object. The video input can comprise video data that can correspond, for example, a plurality of frames. The object can be any object that is desired to be tracked for which an object detector can be provided. For example, the object can be a human face, the full body of a person, a bottle, a ball, a car, an airplane, a bicycle, or any other object in an image or video frame for which an object detector can be created". Note that the user face is mapped to a portion of the user). Regarding claim 6, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 1 as outlined above. Further, Dekel teaches that the method of claim 1, wherein the 3D information provides a 3D representation of at least one joint of the user at a specified location within the 3D environment (See Dekel: Figs. 8A-D, and [0092], "FIG. 8B illustrates a visual representation of an additional object inserted into target image 308 in a depth-aware manner. Specifically, a visual representation of box 804 is inserted into target image 308 at a position in front of human 312 to generate image 802. Since box 804 is inserted in front of human 312, the proper occlusion between portions of human 312 and box 804 may be generated and rendered to accurately represent the position of box 804. Namely, box 804, when placed in front of human 312, may occlude the feet and portions of the legs of actor 312. Accordingly, these portions of human 312 might not be rendered in image 802. Such occlusions may be determined based on the depth values represented by dynamic depth image 412, such that nearby objects occlude, and are thus rendered instead of, corresponding portions of image features that are more distant". Note that the user 312 is at a specific position in the 3D environment and the box is placed in the front of the user, and this is mapped to the limitation of "a 3D representation of at least one joint of the user at a specified location within the 3D environment"). Regarding claim 7, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 1 as outlined above. Further, Cherevatsky teaches that the method of claim 1, wherein the 3D information provides information associated with an action performed by the user (See Cherevatsky: Figs. SA-D, and Col. 28 Lines 57-67 ~ Col. 29 Lines 1-14, "Whether an item is sufficiently represented within imaging data (e.g., visual image frames and/or depth image frames) captured by an imaging device, such as one of the imaging devices 525-1, 525-2 of FIGS. SA and SB, may be determined by calculating a portion or share of a 2D representation of a 3D bounding region having a target object therein that is visible within a field of view of the imaging device, as well as portion or share of the pixels corresponding to the target object within the 2D representation of the 3D bounding region that are occluded from view by one or more other objects. For example, as is shown in FIG. SC, a visual image 530-1 captured at time t.sub.1using the imaging device 525-1, e.g., from a top view of the materials handling facility 520, depicts an operator 580 (e.g., a customer) using a hand 583 to interact with an item 585 (e.g., a medium-sized bottle) on one of the shelves 572-2 in the shelving unit 570. A visual image 530-2 captured at time t.sub.1using the imaging device 525-2, e.g., from a front view of the shelving unit 570, also depicts the operator 580 interacting with the item 585 using the hand 583. A 2D box 535-1 corresponding to a representation of a 3D bounding region in the visual image 530-1 is shown centered on the hand 583, while a 2D box 535-2 corresponding to a representation of the 3D bounding region in the visual image 530-2 is also shown centered on the hand 583". Note that the user interacting with the items in the shelves is mapped to "information associated with an action performed by the user"). Regarding claim 8, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 1 as outlined above. Further, Hoffberg teaches that the method of claim 7, wherein 3D information provides information associated with a number of times the action is performed by the user (See Hoffberg: Figs. 30-31, and Col. 61 Lines 46-64, "In formulating a group preference, individual dislikes may be weighted more heavily than likes, so that the resulting selection is tolerable by all and preferable to most group members. Thus, instead of a best match to a single preference profile for a single user, a group system provides a most acceptable match for the group. It is noted that this method is preferably used in groups of limited size, where individual preference profiles may be obtained, in circumstances where the group will interact with the device a number of times, and where the subject source program material is the subject of preferences. Where large groups are present, demographic profiles may be employed, rather than individual preferences. Where the device is used a small number of times by the group or members thereof, the training time may be very significant and weigh against automation of selection. Where the source material has little variety, or is not the subject of strong preferences, the predictive power of the device as to a desired selection is limited". Note that the number of time of user interaction is mapped to the limitation of "information associated with a number of times the action is performed by the user"). Regarding claim 9, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 1 as outlined above. Further, Cherevatsky teaches that thee method of claim 1, wherein the 3D information comprises information associated with a 3D mesh (See Cherevatsky: Figs. 8A-M, and Col. 1 Lines 22-34, "In dynamic environments such as materials handling facilities, transportation centers, financial institutions or like structures in which diverse collections of people, objects or machines enter and exit from such environments at regular or irregular times or on predictable or unpredictable schedules, it is frequently difficult to detect and track small and/or fast-moving objects using digital cameras. Most systems for detecting and tracking objects in three-dimensional (or "3D") space are limited to the use of a single digital camera and involve both the generation of a 3D mesh (e.g., a polygonal mesh) from depth imaging data captured from such objects and the patching of portions of visual imaging data onto faces of the 3D mesh". Note that the generation of 3D mesh and patching the image data onto the faces of the 3D mesh are mapped to the current cited limitation of "information associated with a 3D mesh"). Regarding claim 10, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 1 as outlined above. Further, Merler teaches that the method of claim 1, wherein the continuous time light signal is captured by an image sensor (See Merler: Fig. 2, and [0004], "The rise of unstructured multimedia content on sites such as Youtube has fostered interest techniques to recognize objects and activities in less structured, unconstrained and more realistic domains. Along the same lines, there is interest in automated video surveillance systems that can detect categories of activity in a field of video of a surveillance camera"; and [0046], "The system can include providing video data 211 of the video input 210 containing at least one object 211, the video data corresponding to a frame of video that corresponds to a position t. In some embodiments, the video input 210 can include a plurality of frames. The video data 211 that makes up the frames can be stored in a memory 220, for example in random access memory with the use of the processor 230 and accompanying 1/0 functionality. The video data 211 stored in the memory can comprise all of the data for the video input 210, for example, in a situation in which the video input 210 is pre-recorded video. Alternatively, the video data 211 can comprise one or more frames of the video input 210, for example, in a situation in which video is being captured by a video capture device and streamed in real time to the computer system 200". Note that the surveillance camera capturing the video and streamed to the system, and surveillance camera is mapped to the camera sensor). Regarding claim 11, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 1 as outlined above. Further, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach that a non-transitory computer-readable medium comprising instructions that when executed by a processor cause the processor to perform operations (See Merler: Figs. 1-2, and [0038], "FIG. 1 and FIG. 2 provide a is a flow diagram that illustrates the method for tracking objects according to an exemplary embodiment of the disclosed subject matter and a computer system having an object tracking system in accordance with an embodiment of the disclosed subject matter, respectively") comprising: obtaining two-dimensional (2D) information corresponding to a continuous time light signal providing information about a user in a three-dimensional (3D) environment (See Dekel: Fig. 3, and [0055], "FIG. 3 illustrates an example video that may be processed by the systems and operations herein disclosed to determine depth values of features represented within the video. Namely, FIG. 3 illustrates video 310 generated by a monoscopic camera 300. Video 310 includes images 302, 303, 304, 305, 306, 307, and 308-309 (i.e., images 302-309), which may alternatively be referred to as frames or image frames of the video. Images 302-309 may represent static features of an environment (e.g., boxes 314 and 316) and moving features of the environment (e.g., human 312). Static features may include objects and other physical features that are expected to remain stationary for a predetermined period of time, such as building structures, trees, roads, or sidewalks, among other possibilities. Moving features may include objects and other physical features that are expected to move within the predetermined period of time, such as humans, animals, or vehicles, among other possibilities. Thus, some moving features change their positions during video 310, while static features may remain in fixed positions during video 310". Note that the environment with boxes 314 and 316 clearly shows that the user in the 3D environment, and a series of 2D images (continuous time light signals) of the user are captured by the monoscopic camera 300, and this is mapped to the cited limitation of: corresponding to a continuous time light signal providing information about a user in a three-dimensional (3D) environment"), the 2D information based on frames comprising images capturing the continuous time light signal at more than one frame rates (See Merler: Fig. 2, and [0046], "The system can include providing video data 211 of the video input 210 containing at least one object 211, the video data corresponding to a frame of video that corresponds to a position t. In some embodiments, the video input 210 can include a plurality of frames. The video data 211 that makes up the frames can be stored in a memory 220, for example in random access memory with the use of the processor 230 and accompanying 1/0 functionality. The video data 211 stored in the memory can comprise all of the data for the video input 210, for example, in a situation in which the video input 210 is pre-recorded video. Alternatively, the video data 211 can comprise one or more frames of the video input 210, for example, in a situation in which video is being captured by a video capture device and streamed in real time to the computer system 200". Note that the video captured device captures images which are streamed in real time to the system is mapped to the 2D information corresponding to a continuous time light signal because the video is a series of frame, and each frame is 2D image captured by the video camera with a specific exposed time period to a continuous light, that is, a frame rate. But Merler does not teach explicitly that the object is in 3D environment, and a second art will be searched and cited below) capturing the continuous time light signal at more than one frame rates (See Delachanal: Figs. 1-4, and Col. 3 Lines 66-67~ Col. 4 Lines 1-17, “FIG. 1 illustrates a system 10 that uses a variable capture frame rate for video capture. The system 10 may include one or more of a processor 11, an electronic storage 12, an interface 13 (e.g., bus, wireless interface), an image sensor 14, a motion sensor 15, and/or other components. The image sensor 14 may be configured to generate visual output signals conveying visual information within a field of view of the image sensor 14. The motion sensor 15 may be configured to generate motion output signals conveying motion information of the image sensor 14. First video information defining first video content and second video information defining second video content may be generated based on the visual output signals. The first video content may be captured using a first capture frame rate set to a first value and the second video content may be captured using a second capture frame rate set to a second value. The values of the first capture frame rate and the second capture frame rate may define the numbers of frames captured per a duration of time”; and Col. 8 Lines 43-55, “The combination component 108 may be configured to generate third video information defining third video content based on the first video information, the second video information, and/or other information. The third video information may be generated based on the first value defining a lower number of frames captured per the duration of time than (1) the second value, and (2) the third value. The third video content may include one or more frames of the first video content, one or more frames of the second video content, and/or other frames. The third video content may include a combination of some or all of the first video content, some or all of the second video content, and/or other video content”. Note that the first video is captured at the first frame rate, the second video is captured at the second frame rate, and the third video is generated by combining the first video and the second video); obtaining discretization information corresponding to the more than one frame rates (See Merler: Fig. 1, and [0043], "Positions 101 through 106 can then be repeated at each position in the interval. That is, the position t can be increased to the next position t+1 (step 107). Steps 101 through 106 can then be applied to the frame in the video input corresponding to position t+1". Note that the time interval t + n is discretization information corresponding to the frame rate, because the frame rate (R) and the time interval Mis mathematically related by the formula R=l/ M, and in digital field, t + n is actually the physical time = t + n M); and determining 3D information about the user by inputting the 2D information and the discretization information into a state space model (See Merler: Figs. 1-2, and [0042], "A final estimate of position and/or velocity of the object at a position can be calculated (step 105) at position t with the use of a Kalman filter 106. A Kalman gain 104 can be used to weight the importance of the noisy measurement component and noisy predictor component of the Kalman filter 106. The Kalman filter can be a general Kalman filter. Alternatively, in some embodiments, the Kalman filter can be a steady-state Kalman filter (i.e, the Kalman gain can be predetermined)"; and [0047], "An object detector 240 can obtain a first location estimate at position t. An object tracker 250 can obtain a second location estimate and a movement estimate at position t. The first location estimate can represent the noisy measurement component 251 of the Kalman filter 106, and the second location estimate and movement estimate can represent the noisy predictor component 252 of the Kalman filter 106. A final estimate 265 of position and/or velocity of the object at a future time t+n, where n>0 can be calculated at each position t with a Kalman filter 260. The final estimate can be calculated with reference to a Kalman gain 267". Note that Kalman filter is a SSM (state space model) that estimates the state of the system from a series of noisy measurement, as shown here that the Kalman filter is used to estimate the position and velocity based on the noisy measurements, that is, Kalman filter is a state space model), the state space model is a continuous time learnable framework (See Hoffberg: Fig. 30, and Col. 15 Lines 33-64, "Methods employing other than fractal-based algorithms may also be used. See, e.g., Liu, Y., "Pattern recognition using Hilbert space", Proceedings of the SPIE--The International Society for Optical Engineering, 1825:63-77 (1992), which describes a learning approach, the Hilbert learning. This approach is similar to Fractal learning, but the Fractal part is replaced by Hilbert space. Like the Fractal learning, the first stage is to encode an image to a small vector in the internal space of a learning system. The next stage is to quantize the internal parameter space. The internal space of a Hilbert learning system is defined as follows: a pattern can be interpreted as a representation of a vector in a Hilbert space. Any vectors in a Hilbert space can be expanded. If a vector happens to be in a subspace of a Hilbert space where the dimension L of the subspace is low (order of 10), the vector can be specified by its norm, an L-vector, and the Hermitian operator which spans the Hilbert space, establishing a mapping from an image space to the internal space P. This mapping converts an input image to a 4-tuple: tin P=(Norm, T, N, L-vector), where Tis an operator parameter space, N is a set of integers which specifies the boundary condition. The encoding is implemented by mapping an input pattern into a point in its internal space. The system uses local search algorithm, i.e., the system adjusts its internal data locally. The search is first conducted for an operator in a parameter space of operators, then an error function delta (t) is computed. The algorithm stops at a local minimum of delta (t). Finally, the input training set divides the internal space by a quantization procedure. See also, Liu, Y., "Extensions of fractal theory", Proceedings of the SPIE--The International Society for Optical Engineering, 1966:255-68(1993)"; and Fig. 15, and Col. 50 Lines 59-62, "The interface may therefore provide a model of the user, which is employed in a predictive algorithm. The model parameters may be static (once created) or dynamic, and may be adaptive to the user or alterations in the use pattern". Note that the adaptive pattern recognition algorithm map the 2D inputs into high dimensional output (3D outputs) with parameters adapted to the input data, and this parameter adaptive to the inputs is mapped to the learnable framework) for (See Cherevatsky: Fig. 3, and Col. 21 Lines 34-58, "At box 325, an initial point cloud is defined from depth image frames captured from the scene using one or more of the depth cameras. For example, where a depth image containing information relating to distances of surfaces of objects within a scene from a perspective of a depth camera is captured, the depth image may be converted into a 3D representation of the physical topography of the scene from that perspective using ranging information for one or more of the pixels provided in the depth image and parameters of the depth camera, e.g., a set of coordinates of the imaging sensor or other components of the depth camera. Two or more depth images captured using RGBD cameras from different perspectives may be further utilized to enhance the quality of the 3D representation of the scene. At box 330, visual cameras having the target object in view within visual image frames captured from the scene are determined. For example, where the 3D bounding region has been defined at box 310, an extent to which a 2D projection of the 3D bounding region appears within the fields of view of each of the imaging devices is determined. At box 332, the extent to which pixels corresponding to the target object are occluded (or not occluded) within the 2D projection of the 3D bounding region is determined, e.g., by comparing depth data for the target cloud points to depth data for other scene points within a frustrum spanned by the 3D bounding region"; and Fig. 7, and Col. 33 Lines 21-24, "At box 752, the probability map for the position of the target object is provided to a Kalman filter or another set of mathematical equations for estimating the position of the target object in a manner that minimizes a mean of the squared errors associated with the position. At box 754, the Kalman filter models motion of the target object based on probability maps determined for all known synchronization points, e.g., synchronization points ranging from 1 to i". Note that Kalman filter (the SSM) is used and 2D images is fused to get the 3D bounding boxes and position of the tracking object, and this is mapped to "mapping between continuous time 2D scalar inputs and continuous time scalar 3D outputs". However, Cherevatsky does not teach that the SSM is a learnable framework, and another art will be searched below) mapping between continuous time scalar 2D inputs and continuous time scalar 3D outputs (See Wang: Figs. 1-3, and [0028], “As previously noted, standard SLDS based on first-order Markov assumption does not yield the best performance, particularly when dealing with between action transitions. Presented herein are embodiments of dynamic-system-based framework to model the temporal evolution of a feature for action analysis. In embodiments, the framework utilizes higher-order information to explicitly model the transition within single action primitive or between successive action primitives. In embodiments, the framework, referred to as the Continuous Linear Dynamic System (CLDS) framework, comprises two sets of Linear Dynamic System (LDS) models, one to model the dynamics of individual primitive actions and the other to model the transition between actions. In embodiments, the inference process estimates the best decomposition of a whole sequence by continuously alternating between the two set of models. In embodiments, an approximate Viterbi algorithm may be used in the inference process. Using the CLDS framework, both action type and action boundary may be accurately recognized”; [0094]. “FIG. 5 depicts a method for using a Continuous Linear Dynamic System (CLDS) to detect and label actions in a video according to embodiments of the present invention. As depicted in FIG. 5, the process commences by segmenting (505) input sensor data into time frames. In embodiments, the input sensor data may be video data and the segmented time frames may be image frames, although other sensor data and configurations may be used. Given the segmented input sensor data, a feature for each image frame is generated (510). In embodiments, an image feature for the frame may be an embedded optical flow as previously discussed. These images features are then input into a CLDS model to perform (515) continuous segmentation and recognition. In embodiments, the CLDS model may use one or more of the inference methods discussed above or known to those of ordinary skill in the art’; [0095], “FIG. 6 depicts a block diagram of a Continuous Linear Dynamic System (CLDS) model detector 605 according to embodiments of the present invention. The CLDS model detector 605 receives input sensor data 625 and outputs a sequence of labels 630. As shown in FIG. 6, the CLDS model detector 605 comprises a frame extractor 610, a feature extractor 615, and a CLDS model decoder 620. In embodiments, CLDS model detector performs one or more methods for continuous segmentation and recognition, which include but are not limited to the methods discussed above”. Note that input sensor data 625 is mapped to the continuous 2D scalar input, and output label sequence 630 is mapped to the continuous 3D scalar outputs). Regarding claim 12, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 1 as outlined above. Further, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach that an electronic device (See Merler: Figs. 1-2, and [0038], "FIG. 1 and FIG. 2 provide a is a flow diagram that illustrates the method for tracking objects according to an exemplary embodiment of the disclosed subject matter and a computer system having an object tracking system in accordance with an embodiment of the disclosed subject matter, respectively") comprising: a non-transitory computer-readable storage medium (See Merler: Figs. 1-2, and [0045], "With reference to FIG. 2, a description of a computer system configured for object tracking according to an embodiment of the presently disclosed subject matter will be described. The system depicted in FIG. 2 can carry out the method depicted in FIG. 1. The system can include a processor 230 in communication with a memory 220"); and one or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium comprises program instructions that, when executed on the one or more processors, cause the electronic device to perform operations (See Merler: Figs. 1-2, and [0015], "In one embodiment, the system can include executable code stored in the memory and configured to instruct the processor to obtain a first location estimate of an object in the video input at each position in the interval. The executable code can be configured to instruct the processor to obtain a second location estimate and a movement estimate of an object in the video input at each position in the interval, and to calculate a final estimate of a position and/or velocity of the object at a future time at each position in the interval with the use of a Kalman filter. The first location estimate can represent the noisy measurement component of the Kalman filter, and the second location estimate and the movement estimate can represent the noisy predictor component of the Kalman filter''; and [0046], "The video data 211 that makes up the frames can be stored in a memory 220, for example in random access memory with the use of the processor 230 and accompanying 1/0 functionality". Note that the processor be configured to execute codes and instructions stored in the memory is mapped to the current cited limitation) comprising: obtaining two-dimensional (2D) information corresponding to a continuous time light signal providing information about a user in a three-dimensional (3D) environment (See Dekel: Fig. 3, and [0055], "FIG. 3 illustrates an example video that may be processed by the systems and operations herein disclosed to determine depth values of features represented within the video. Namely, FIG. 3 illustrates video 310 generated by a monoscopic camera 300. Video 310 includes images 302, 303, 304, 305, 306, 307, and 308-309 (i.e., images 302-309), which may alternatively be referred to as frames or image frames of the video. Images 302-309 may represent static features of an environment (e.g., boxes 314 and 316) and moving features of the environment (e.g., human 312). Static features may include objects and other physical features that are expected to remain stationary for a predetermined period of time, such as building structures, trees, roads, or sidewalks, among other possibilities. Moving features may include objects and other physical features that are expected to move within the predetermined period of time, such as humans, animals, or vehicles, among other possibilities. Thus, some moving features change their positions during video 310, while static features may remain in fixed positions during video 310". Note that the environment with boxes 314 and 316 clearly shows that the user in the 3D environment, and a series of 2D images (continuous time light signals) of the user are captured by the monoscopic camera 300, and this is mapped to the cited limitation of: corresponding to a continuous time light signal providing information about a user in a three-dimensional (3D) environment"), the 2D information based on frames comprising images (See Merler: Fig. 2, and [0046], "The system can include providing video data 211 of the video input 210 containing at least one object 211, the video data corresponding to a frame of video that corresponds to a position t. In some embodiments, the video input 210 can include a plurality of frames. The video data 211 that makes up the frames can be stored in a memory 220, for example in random access memory with the use of the processor 230 and accompanying I/O functionality. The video data 211 stored in the memory can comprise all of the data for the video input 210, for example, in a situation in which the video input 210 is pre-recorded video. Alternatively, the video data 211 can comprise one or more frames of the video input 210, for example, in a situation in which video is being captured by a video capture device and streamed in real time to the computer system 200". Note that the video captured device captures images which are streamed in real time to the system is mapped to the 2D information corresponding to a continuous time light signal because the video is a series of frame, and each frame is 2D image captured by the video camera with a specific exposed time period to a continuous light, that is, a frame rate. But Merler does not teach explicitly that the object is in 3D environment, and a second art will be searched and cited below) capturing the continuous time light signal at more than one frame rates (See Delachanal: Figs. 1-4, and Col. 3 Lines 66-67~ Col. 4 Lines 1-17, “FIG. 1 illustrates a system 10 that uses a variable capture frame rate for video capture. The system 10 may include one or more of a processor 11, an electronic storage 12, an interface 13 (e.g., bus, wireless interface), an image sensor 14, a motion sensor 15, and/or other components. The image sensor 14 may be configured to generate visual output signals conveying visual information within a field of view of the image sensor 14. The motion sensor 15 may be configured to generate motion output signals conveying motion information of the image sensor 14. First video information defining first video content and second video information defining second video content may be generated based on the visual output signals. The first video content may be captured using a first capture frame rate set to a first value and the second video content may be captured using a second capture frame rate set to a second value. The values of the first capture frame rate and the second capture frame rate may define the numbers of frames captured per a duration of time”; and Col. 8 Lines 43-55, “The combination component 108 may be configured to generate third video information defining third video content based on the first video information, the second video information, and/or other information. The third video information may be generated based on the first value defining a lower number of frames captured per the duration of time than (1) the second value, and (2) the third value. The third video content may include one or more frames of the first video content, one or more frames of the second video content, and/or other frames. The third video content may include a combination of some or all of the first video content, some or all of the second video content, and/or other video content”. Note that the first video is captured at the first frame rate, the second video is captured at the second frame rate, and the third video is generated by combining the first video and the second video); obtaining discretization information corresponding to the more than one frame rates (See Merler: Fig. 1, and [0043], "Positions 101 through 106 can then be repeated at each position in the interval. That is, the position t can be increased to the next position t+1 (step 107). Steps 101 through 106 can then be applied to the frame in the video input corresponding to position t+1". Note that the time interval t + n is discretization information corresponding to the frame rate, because the frame rate (R) and the time interval Mis mathematically related by the formula R=l/ M, and in digital field, t + n is actually the physical time = t + n M); and determining 3D information about the user by inputting the 2D information and the discretization information into a state space model (See Merler: Figs. 1-2, and [0042], "A final estimate of position and/or velocity of the object at a position can be calculated (step 105) at position t with the use of a Kalman filter 106. A Kalman gain 104 can be used to weight the importance of the noisy measurement component and noisy predictor component of the Kalman filter 106. The Kalman filter can be a general Kalman filter. Alternatively, in some embodiments, the Kalman filter can be a steady-state Kalman filter (i.e, the Kalman gain can be predetermined)"; and [0047], "An object detector 240 can obtain a first location estimate at position t. An object tracker 250 can obtain a second location estimate and a movement estimate at position t. The first location estimate can represent the noisy measurement component 251 of the Kalman filter 106, and the second location estimate and movement estimate can represent the noisy predictor component 252 of the Kalman filter 106. A final estimate 265 of position and/or velocity of the object at a future time t+n, where n>0 can be calculated at each position t with a Kalman filter 260. The final estimate can be calculated with reference to a Kalman gain 267". Note that Kalman filter is a SSM (state space model) that estimates the state of the system from a series of noisy measurement, as shown here that the Kalman filter is used to estimate the position and velocity based on the noisy measurements, that is, Kalman filter is a state space model), the state space model is a continuous time learnable framework (See Hoffberg: Fig. 30, and Col. 15 Lines 33-64, "Methods employing other than fractal-based algorithms may also be used. See, e.g., Liu, Y., "Pattern recognition using Hilbert space", Proceedings of the SPIE--The International Society for Optical Engineering, 1825:63-77 (1992), which describes a learning approach, the Hilbert learning. This approach is similar to Fractal learning, but the Fractal part is replaced by Hilbert space. Like the Fractal learning, the first stage is to encode an image to a small vector in the internal space of a learning system. The next stage is to quantize the internal parameter space. The internal space of a Hilbert learning system is defined as follows: a pattern can be interpreted as a representation of a vector in a Hilbert space. Any vectors in a Hilbert space can be expanded. If a vector happens to be in a subspace of a Hilbert space where the dimension L of the subspace is low (order of 10), the vector can be specified by its norm, an L-vector, and the Hermitian operator which spans the Hilbert space, establishing a mapping from an image space to the internal space P. This mapping converts an input image to a 4-tuple: tin P=(Norm, T, N, L-vector), where Tis an operator parameter space, N is a set of integers which specifies the boundary condition. Thee needing is implemented by mapping an input pattern into a point in its internal space. The system uses local search algorithm, i.e., the system adjusts its internal data locally. The search is first conducted for an operator in a parameter space of operators, then an error function delta (t) is computed. The algorithm stops at a local minimum of delta (t). Finally, the input training set divides the internal space by a quantization procedure. See also, Liu, Y., "Extensions of fractal theory", Proceedings of the SPIE--The International Society for Optical Engineering, 1966:255-68(1993)"; and Fig. 15, and Col. 50 Lines 59-62, "The interface may therefore provide a model of the user, which is employed in a predictive algorithm. The model parameters may be static (once created) or dynamic, and may be adaptive to the user or alterations in the use pattern". Note that the adaptive pattern recognition algorithm map the 2D inputs into high dimensional output (3D outputs) with parameters adapted to the input data, and this parameter adaptive to the inputs is mapped to the learnable framework) for mapping (See Cherevatsky: Fig. 3, and Col. 21 Lines 34-58, "At box 325, an initial point cloud is defined from depth image frames captured from the scene using one or more of the depth cameras. For example, where a depth image containing information relating to distances of surfaces of objects within a scene from a perspective of a depth camera is captured, the depth image may be converted into a 3D representation of the physical topography of the scene from that perspective using ranging information for one or more of the pixels provided in the depth image and parameters of the depth camera, e.g., a set of coordinates of the imaging sensor or other components of the depth camera. Two or more depth images captured using RGBD cameras from different perspectives may be further utilized to enhance the quality of the 3D representation of the scene. At box 330, visual cameras having the target object in view within visual image frames captured from the scene are determined. For example, where the 3D bounding region has been defined at box 310, an extent to which a 2D projection of the 3D bounding region appears within the fields of view of each of the imaging devices is determined. At box 332, the extent to which pixels corresponding to the target object are occluded (or not occluded) within the 2D projection of the 3D bounding region is determined, e.g., by comparing depth data for the target cloud points to depth data for other scene points within a frustrum spanned by the 3D bounding region"; and Fig. 7, and Col. 33 Lines 21-24, "At box 752, the probability map for the position of the target object is provided to a Kalman filter or another set of mathematical equations for estimating the position of the target object in a manner that minimizes a mean of the squared errors associated with the position. At box 754, the Kalman filter models motion of the target object based on probability maps determined for all known synchronization points, e.g., synchronization points ranging from 1 to i". Note that Kalman filter (the SSM) is used and 2D images is fused to get the 3D bounding boxes and position of the tracking object, and this is mapped to "mapping between continuous time 2D scalar inputs and continuous time scalar 3D outputs". However, Cherevatsky does not teach that the SSM is a learnable framework, and another art will be searched below) between continuous time scalar 2D inputs and continuous time scalar 3D outputs (See Wang: Figs. 1-3, and [0028], “As previously noted, standard SLDS based on first-order Markov assumption does not yield the best performance, particularly when dealing with between action transitions. Presented herein are embodiments of dynamic-system-based framework to model the temporal evolution of a feature for action analysis. In embodiments, the framework utilizes higher-order information to explicitly model the transition within single action primitive or between successive action primitives. In embodiments, the framework, referred to as the Continuous Linear Dynamic System (CLDS) framework, comprises two sets of Linear Dynamic System (LDS) models, one to model the dynamics of individual primitive actions and the other to model the transition between actions. In embodiments, the inference process estimates the best decomposition of a whole sequence by continuously alternating between the two set of models. In embodiments, an approximate Viterbi algorithm may be used in the inference process. Using the CLDS framework, both action type and action boundary may be accurately recognized”; [0094]. “FIG. 5 depicts a method for using a Continuous Linear Dynamic System (CLDS) to detect and label actions in a video according to embodiments of the present invention. As depicted in FIG. 5, the process commences by segmenting (505) input sensor data into time frames. In embodiments, the input sensor data may be video data and the segmented time frames may be image frames, although other sensor data and configurations may be used. Given the segmented input sensor data, a feature for each image frame is generated (510). In embodiments, an image feature for the frame may be an embedded optical flow as previously discussed. These images features are then input into a CLDS model to perform (515) continuous segmentation and recognition. In embodiments, the CLDS model may use one or more of the inference methods discussed above or known to those of ordinary skill in the art’; [0095], “FIG. 6 depicts a block diagram of a Continuous Linear Dynamic System (CLDS) model detector 605 according to embodiments of the present invention. The CLDS model detector 605 receives input sensor data 625 and outputs a sequence of labels 630. As shown in FIG. 6, the CLDS model detector 605 comprises a frame extractor 610, a feature extractor 615, and a CLDS model decoder 620. In embodiments, CLDS model detector performs one or more methods for continuous segmentation and recognition, which include but are not limited to the methods discussed above”. Note that input sensor data 625 is mapped to the continuous 2D scalar input, and output label sequence 630 is mapped to the continuous 3D scalar outputs).. Regarding claim 13, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 12 as outlined above. Further, Merler teaches that the electronic device of claim 12, wherein the discretization information comprises delta information corresponding to time periods between the frames (See Merler: Fig. 1, and [0043], "Positions 101 through 106 can then be repeated at each position in the interval. That is, the position t can be increased to the next position t+1 (step 107). Steps 101 through 106 can then be applied to the frame in the video input corresponding to position t+1". Note that the time interval t + n is discretization information corresponding to the frame rate, because the frame rate (R) and the time interval M between two consecutive frames is mathematically related by the formula R=l/ M, and in digital field, t + n is actually the physical time= t + n M, for example, if the first frame is at time t, the next frame is t+1, in physical time is t + M with n = 1). Regarding claim 14, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 12 as outlined above. Further, Merler teaches that the electronic device of claim 12, wherein the discretization information comprises information associated with the more than one frame rates (See Merler: Fig. 1, and [0047], "An object detector 240 can obtain a first location estimate at position t. An object tracker 250 can obtain a second location estimate and a movement estimate at position t. The first location estimate can represent the noisy measurement component 251 of the Kalman filter 106, and the second location estimate and movement estimate can represent the noisy predictor component 252 of the Kalman filter 106. A final estimate 265 of position and/or velocity of the object at a future time t+n, where n>0 can be calculated at each position t with a Kalman filter 260. The final estimate can be calculated with reference to a Kalman gain 267". Note that the time interval t + n is discretization information associated with the at least one or more frame rates, because the frame rate (R) and the time interval M between two consecutive frames is mathematically related by the formula M = 1/R, and in digital field, t + n is actually the physical time= t + n M = t + n (1/R), for example, if the first frame is at time t, the next frame is t+1, in physical time is t + 1/R with n = 1). Regarding claim 15, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 12 as outlined above. Further, Cherevatsky teaches that the electronic device of claim 12, wherein the 2D information comprises information associated with 2D locations of joints of the user (See Cherevatsky: Fig. 7, and Col. 34 Lines 20-35, "At box 770, the tracklet for the target object over the tracking period is defined based on the probability maps and the point clouds defined from the visual image frames and the depth image frames captured at the prior synchronization points i. For example, a voting algorithm may be used to estimate a joint object position probability distribution in 3D space based on representations of the target object in 2D images captured by the plurality of imaging devices, and recognized therein using a tracking algorithm, such as an OpenCV tracker or a KCF tracker. Such representations may be projected onto the point clouds, and a tracklet of the positions of the target object may be determined accordingly, such as by assigning scores to each of the points in 3D space at various times, aggregating scores for such points, and selecting a best candidate based on the aggregated scores". Note that the joint object position is mapped to the information associated with 2D locations of joints of the user). Regarding claim 16, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 12 as outlined above. Further, Merler teaches that the electronic device of claim 12, wherein the 3D information provides a 3D model representing at least a portion of the user (See Merler: Fig. 1, and [0039], "With reference to FIG. 1, the method of some embodiments can include providing a video input that contains at least one object. The video input can comprise video data that can correspond, for example, a plurality of frames. The object can be any object that is desired to be tracked for which an object detector can be provided. For example, the object can be a human face, the full body of a person, a bottle, a ball, a car, an airplane, a bicycle, or any other object in an image or video frame for which an object detector can be created". Note that the user face is mapped to a portion of the user). Regarding claim 17, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 12 as outlined above. Further, Dekel teaches that the electronic device of claim 12, wherein the 3D information provides a 3D representation of at least one joint of the user at a specified location within the 3D environment (See Dekel: Figs. 8A-D, and [0092], "FIG. 8B illustrates a visual representation of an additional object inserted into target image 308 in a depth-aware manner. Specifically, a visual representation of box 804 is inserted into target image 308 at a position in front of human 312 to generate image 802. Since box 804 is inserted in front of human 312, the proper occlusion between portions of human 312 and box 804 may be generated and rendered to accurately represent the position of box 804. Namely, box 804, when placed in front of human 312, may occlude the feet and portions of the legs of actor 312. Accordingly, these portions of human 312 might not be rendered in image 802. Such occlusions may be determined based on the depth values represented by dynamic depth image 412, such that nearby objects occlude, and are thus rendered instead of, corresponding portions of image features that are more distant". Note that the user 312 is at a specific position in the 3D environment and the box is placed in the front of the user, and this is mapped to the limitation of "a 3D representation of at least one joint of the user at a specified location within the 3D environment"). Regarding claim 18, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 12 as outlined above. Further, Cherevatsky teaches that the electronic device of claim 12, wherein the 3D information provides information associated with an action performed by the user (See Cherevatsky: Figs. SA-D, and Col. 28 Lines 57-67 ~ Col. 29 Lines 1-14, "Whether an item is sufficiently represented within imaging data (e.g., visual image frames and/or depth image frames) captured by an imaging device, such as one of the imaging devices 525-1, 525-2 of FIGS. SA and SB, may be determined by calculating a portion or share of a 2D representation of a 3D bounding region having a target object therein that is visible within a field of view of the imaging device, as well as portion or share of the pixels corresponding to the target object within the 2D representation of the 3D bounding region that are occluded from view by one or more other objects. For example, as is shown in FIG. SC, a visual image 530-1 captured at time t.sub.1 using the imaging device 525-1, e.g., from a top view of the materials handling facility 520, depicts an operator 580 (e.g., a customer) using a hand 583 to interact with an item 585 (e.g., a medium-sized bottle) on one of the shelves 572-2 in the shelving unit 570. A visual image 530-2 captured at time t.sub.1using the imaging device 525-2, e.g., from a front view of the shelving unit 570, also depicts the operator 580 interacting with the item 585 using the hand 583. A 2D box 535-1 corresponding to a representation of a 3D bounding region in the visual image 530-1 is shown centered on the hand 583, while a 2D box 535-2 corresponding to a representation of the 3D bounding region in the visual image 530-2 is also shown centered on the hand 583". Note that the user interacting with the items in the shelves is mapped to "information associated with an action performed by the user"). Regarding claim 19, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 12 as outlined above. Further, Hoffberg teaches that the electronic device of claim 18, wherein 3D information provides information associated with a number of times the action is performed by the user (See Hoffberg: Figs. 30-31, and Col. 61 Lines 46-64, "In formulating a group preference, individual dislikes may be weighted more heavily than likes, so that the resulting selection is tolerable by all and preferable to most group members. Thus, instead of a best match to a single preference profile for a single user, a group system provides a most acceptable match for the group. It is noted that this method is preferably used in groups of limited size, where individual preference profiles may be obtained, in circumstances where the group will interact with the device a number of times, and where the subject source program material is the subject of preferences. Where large groups are present, demographic profiles may be employed, rather than individual preferences. Where the device is used a small number of times by the group or members thereof, the training time may be very significant and weigh against automation of selection. Where the source material has little variety, or is not the subject of strong preferences, the predictive power of the device as to a desired selection is limited". Note that the number of time of user interaction is mapped to the limitation of "information associated with a number of times the action is performed by the user''). Regarding claim 20, Merler, Dekel, Cherevatsky, Hoffberg, Delachanal and Wang teach all the features with respect to claim 12 as outlined above. Further, Cherevatsky teaches that the electronic device of claim 12, wherein the 3D information comprises information associated with a 3D mesh (See Cherevatsky: Figs. 8A-M, and Col. 1 Lines 22-34, "In dynamic environments such as materials handling facilities, transportation centers, financial institutions or like structures in which diverse collections of people, objects or machines enter and exit from such environments at regular or irregular times or on predictable or unpredictable schedules, it is frequently difficult to detect and track small and/or fast-moving objects using digital cameras. Most systems for detecting and tracking objects in three-dimensional (or "3D") space are limited to the use of a single digital camera and involve both the generation of a 3D mesh (e.g., a polygonal mesh) from depth imaging data captured from such objects and the patching of portions of visual imaging data onto faces of the 3D mesh". Note that the generation of 3D mesh and patching the image data onto the faces of the 3D mesh are mapped to the current cited limitation of "information associated with a 3D mesh"). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to GORDON G LIU whose telephone number is (571)270-0382. The examiner can normally be reached Monday - Friday 8:00-5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Devona E Faulk can be reached at 571-272-7515. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /GORDON G LIU/ Primary Examiner, Art Unit 2618
Read full office action

Prosecution Timeline

Oct 21, 2024
Application Filed
Apr 10, 2026
Non-Final Rejection mailed — §103
Jul 12, 2026
Interview Requested
Jul 21, 2026
Examiner Interview Summary
Jul 21, 2026
Applicant Interview (Telephonic)
Aug 04, 2026
Response Filed
Sep 09, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743838
PIXEL GENERATION TECHNIQUE
2y 8m to grant Granted Sep 22, 2026
Patent 12743821
RECORDING MEDIUM AND INFORMATION PROCESSING DEVICE
2y 3m to grant Granted Sep 22, 2026
Patent 12744018
IMAGE OUTPUT CONTROL DEVICE AND METHOD
2y 1m to grant Granted Sep 22, 2026
Patent 12725544
SCREEN DISPLAY DRIVING METHOD AND APPARATUS, SCREEN INFORMATION CONFIGURATION METHOD AND APPARATUS, MEDIUM AND DEVICE
2y 1m to grant Granted Sep 01, 2026
Patent 12725367
PICTURE DISPLAY METHOD, SYSTEM, AND APPARATUS, DEVICE, AND STORAGE MEDIUM
2y 2m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
83%
Grant Probability
98%
With Interview (+14.8%)
2y 2m (~2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 701 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month