Prosecution Insights
Last updated: October 02, 2026
Application No. 18/492,662

SYSTEMS AND ASSOCIATED METHODS FOR REAL-TIME FEATURE DETECTION OF AN ENVIRONMENT

Final Rejection §103
Filed
Oct 23, 2023
Examiner
HAUSMANN, MICHELLE M
Art Unit
2671
Tech Center
2600 — Communications
Assignee
Brightai Corporation
OA Round
2 (Final)
77%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
677 granted / 883 resolved
+14.7% vs TC avg
Strong +21% interview lift
Without
With
+21.1%
Interview Lift
resolved cases with interview
Typical timeline
2y 12m
Avg Prosecution
25 currently pending
Career history
907
Total Applications
across all art units

Statute-Specific Performance

§101
14.0%
-26.0% vs TC avg
§103
67.3%
+27.3% vs TC avg
§102
6.4%
-33.6% vs TC avg
§112
7.3%
-32.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 883 resolved cases

Office Action

§103
DETAILED ACTION Response to Amendment Claims 1-25 are pending. Claims 1-25 are amended directly or by dependency on an amended claim. Response to Arguments Applicant’s arguments with respect to the 35 USC 103 rejections of claim(s) 1-25 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Applicant’s arguments, see page 9, filed 14 July, 2026, with respect to the non-statutory double patenting rejections of claims 1-25 have been fully considered and are persuasive. The non-statutory double patenting rejections of claims 1-25 have been withdrawn. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ebrahimi Afrouzi et al. (US 20220066456 A1) [relies on content published 3 March, 2022 and does not rely on the priority earlier than this date] in view of Lee (US 11550276 B1). Regarding claims 1 and 20, Ebrahimi Afrouzi et al. disclose a method comprising, and system comprising: a robot traversing an environment; a plurality of sensors deployed on the robot (The data collected by the camera may be bundled with data collected by one or more of an OTS, an encoder, an IMU, a gyroscope, etc. The robot may also include a 3D or 2D LIDAR for measuring distances to objects as the robot moves within the environment, [0306]); one or more processors associated with the robot; and one or more storage devices that store instructions, that, when executed by the one or more processors, cause the one or more processors to (processors, storage, instructions, [0269], 1488]): receiving, by one or more processors, a collection of sensor data from a plurality of sensors deployed on a robot traversing an environment, the plurality of sensors associated with a plurality of sensor types (“In embodiments, information is received from sensors and is used in real time by AI algorithms. Decisions actuate the robot without buffer delays based on the real time information. Examples of sensors include, but are not limited to, inertial measurement unit (IMU), gyroscope, optical tracking sensor (OTS), depth camera, obstacle sensor, floor sensor, edge detection sensor, debris sensor, acoustic sensor, speech recognition, camera, image sensor, time of flight (TOF) sensor, TSOP sensor, laser sensor, light sensor, electric current sensor, optical encoder, accelerometer, compass, speedometer, proximity sensor, range finder, LIDAR, LADAR, radar sensor, ultrasonic sensor, piezoresistive strain gauge, capacitive force sensor, electric force sensor, piezoelectric force sensor, optical force sensor, capacitive touch-sensitive surface or other intensity sensors, global positioning system (GPS), etc.”, [0242], The pose estimator may include an Extended Kalman Filter (EKF) that uses odometry, IMU, and LIDAR data, [0244]), tracking by the one or more processors, a position of the robot within the environment, wherein the tracking is performed using data received from a position sensor associated with the robot (The data collected by the camera may be bundled with data collected by one or more of an OTS, an encoder, an IMU, a gyroscope, etc. The robot may also include a 3D or 2D LIDAR for measuring distances to objects as the robot moves within the environment, [0306], In some embodiments, the processor may obtain a first stream of spatial data from a first sensor indicative of the position of the robot within the environment. In some embodiments, the processor may obtain a second stream of spatial data from a second sensor indicative of the position of the robot within the environment, [0423], While an IMU may detect an inertial acceleration after the robot has accelerated a desired cruise speed, the accelerometer may not be helpful in detecting motion with a constant speed. Therefore, in such cases, odometry information from the wheel encoder may be more useful, [0448], The pose estimator may include an Extended Kalman Filter (EKF) that uses odometry, IMU, and LIDAR data. SLAM may build a map based on scan matching. The pose estimator and SLAM may pass information to one another in a feedback loop. The SLAM updated may estimate the pose of the robot, [0244], In case of the LIDAR being covered (i.e., not available), the processor of the robot may use gyroscope data to continue mapping and covering hard surfaces since a gyroscope performs better on hard surfaces. The processor may switch to OTS (optical track sensor) for carpeted areas since OTS performance and accuracy is better in those areas. For example, a mapped area may be generated using LIDAR data, coverage on hard surface by the robot may be executed using only gyroscope sensor, and coverage on carpet by the robot may be executed using an OTS sensor, [0587], the processor may couple LIDAR or camera measurements with IMU, OTS, etc. data, [0589]); recognizing, by the one or more processors, a feature associated with the plurality of sensor types within the environment (In some embodiments, a video that is in red, green, blue (RGB) format may be converted to a video in a different format, such as YCoCg color space format, [0266], In embodiments, a kernel may consist of multiple layers of feature maps, each designed to detect a different feature. All neurons in a single feature map share the same parameters and allow the network to recognize a feature pattern regardless of where the feature pattern is within the input. This is important for object detection. For example, once the network learns that an object positioned in a dwelling is a chair, the network will be able to recognize the chair regardless of where the chair is located in the future, [0274], In some embodiments, a camera of the robot (the camera used for SLAM or another camera) captures images or video while the robot navigates around the environment. Using object recognition, the processor may identify the TV within the images captured and may associate a location within the floor map with the TV, [0400], The robot also includes a camera. The processor of the robot may use data collected by the camera to track a location of features, such as a light fixture, a corner, and an edge. In some embodiments, the camera may be slightly recessed and angled rearward. In some embodiments, the processor uses the location of features to localize the robot, [0589]), using deep learning algorithms operating on the collection of sensor data from the sensor type (In embodiments, there may be a high number of layers in the network (i.e., deep network) or there may be a low number of layers (i.e., shallow network), [0285], “In some embodiments, the AP signal strength data collected by sensors of the robot are fed into the deep neural network model along with accurate LIDAR measurements. In some embodiments, the LIDAR data and AP signal strength data are combined into a data structure then provided to the neural network such that a pattern may be learned and the processor may infer probabilities of a location of the robot based on the AP signal strength data collected”, [0309], reward system of trajectory measurement and observation algorithm are transmitted to the database for input into the Deep Q-Network for reinforcement learning, [0343], Some embodiments provide an image sensor and image processor coupled to the robot and use deep learning to analyze images captured by the image sensor and identify objects in the images, either locally or via the cloud, [1193]); mapping, by the one or more processors, the feature to the position of the robot (SLAM updated may estimate the pose of the robot, [0244], during relocalization a camera of the robot may capture local images and the processor may attempt to locate the robot within the state-space by searching the known map to find a pattern similar to its current observation, [0260], “In some embodiments, the processor may not know the correspondence between data points a priori when merging images and may start by matching nearby points. The processor may then update the most likely correspondence and iterate on. In some embodiments, the processor of the robot may localize the robot against the environment based on feature detection and matching. This may be synonymous to pose estimation or determining the position of cameras and other sensors of the robot relative to a known three dimensional object in the scene, [0339], a camera of the robot may capture an image comprising a television that the processor may use in identifying the room the robot is within, [0588], the processor uses the location of features to localize the robot, [0589]); aggregating, by the one or more processors, the feature using a weighting algorithm to generate an aggregated feature (Since the robot is moving, the most recent measurements captured by the robot may be given more weight as they are more relevant. For instance, data at a current timestamp t is given more weight than older measurements captured at t−1, t−2, t−i. In some embodiments, the position of the robot may be a multidimensional array or tensor and the kernel may be a set of parameters organized in a multidimensional array. The two multidimensional arrays may be convolved to produce a feature map, [0324], processor adjusts the weight given to classification based on the collection of past experiences of robots and classification based on the experiences of the respective robot itself, [0385], “weighted sums computed by hidden layers of the network are propagated to the output layer which may present probabilities to describe a classification, an object detection (to be tracked), a feature detection (to be tracked), etc.”, [0524], “In some embodiments, the weight assigned to readings may be proportional to the size of the overlap area identified. For example, data points corresponding to a moving object captured in one or two frames overlapping with several other frames captured without the moving object may be assigned a low weight as they likely do not fall within the adjustment range and are not consistent with data points collected in other overlapping frames and would likely be rejected for having low assigned weight”, [0926]); and creating, by the one or more processors and based on the aggregated feature, an output predicting the feature present in the environment (In embodiments, x is a first function and is the input to the network, w is a second function called a kernel, and the output of the network is a feature map, [0297], “In some embodiments, low level features are processed in real time. In some embodiments, different outputs may each require a different speed of response from the robot. For instance, an output indicating probabilities of a distance of the robot from an object. This requires fast response from the robot to avoid a collision”, [0318], the processor stitches images and creates a spatial representation of the scene after correcting images with preprocessing, [0339], In some embodiments, computer vision may be used to help with the labeling. For instance, the processor of the robot may recognize cabinetry, an oven, and a dishwasher in a same room and may therefore assume and label the room as the kitchen. Bedrooms, bathrooms, etc. may similarly be identified and labelled. In some embodiments, the processor may use history cubes to determine elements with direction. For example, directions that doors open may be determined using images of a same door at various time stamps. In some embodiments, an architectural plan may be generated by combination of a SLAM generated map and computer vision. In embodiments, additional data may be added to the map by a user or the processor, including labels for each room, specific measurement, notes, etc., [0361], In some embodiments, the processor may use object recognition to identify different objects in the stream of images and may label objects and associate locations in the map with the labelled objects. In some embodiments, the processor may label dynamic obstacles, such as humans and pets, in the map. In some embodiments, the dynamic obstacles have a half life that is determine based on a probability of their presence, [0381], “In some embodiments, the processor classifies the type, size, texture, and nature of objects. In some embodiments, such object classifications are provided as input to the Q-SLAM navigational stack, which then returns as output a decision on how to handle the object with the particular classifications. For example, a decision of the Q-SLAM navigational stack of an autonomous car may be very conservative when an object has even the slightest chance of being a living being, and may therefore decide to avoid the object. In the context of a robotic vacuum cleaner, the Q-SLAM navigational stack may be extra conservative in its decision of handling an object when the object has the slightest chance of being pet bodily waste.” [0514], “The output may be in the form of probabilities of possible outcomes, the outcomes being high-level features such as object type, scene, distance measurement, or displacement of a camera”, [0527]). Ebrahimi Afrouzi et al. has multiple embodiments described. It would have been obvious at the time of filing to one of ordinary skill in the art to combine the embodiments above as the combination would have predictable results, and Ebrahimi Afrouzi et al. indicate “In embodiments, the processor executes deep learning to improve perception, improve trajectory such that it follows the planned path, improve coverage, improve obstacle detection and prevention, make decisions that are more human-like, and to improve operation of the robot in situations where data becomes unavailable (e.g., due to a malfunctioning sensor)” ([0271]), “In embodiments, DNN and CNN are advantageous as there are several different tools that may be used to a necessary degree. For example, proper weight initialization may break symmetries or advantageously choosing ELU or ReLu where negative values or those close to a value of zero are important or using leaky ReLu to advantageously increase performance for a more real-time experience or use of sparsification technique by selecting FTRL over Adam optimization” ([0279]) “In some embodiments, the processor may use the SLAM data to add accurate measurement to the generated architectural plan” ([0361]) “In embodiments, the SLAM algorithm is superior to SLAM methods described in prior art as it is less likely to lose localization of the robot. For example, using traditional SLAM methods, localization of the robot may be lost if the robot is randomly picked up and moved to a different room during a work session. However, using the SLAM algorithm described herein, localization is not lost” ([0485]) providing several computational performance benefits and accuracy improvements when embodiments are combined. Ebrahimi Afrouzi et al. does not disclose embedding, by the one or more processors and using deep learning algorithms, the collection of sensor data from the plurality of sensor types into a single vector space to produce a feature vector. Lee teaches embedding, by the one or more processors and using deep learning algorithms (deep learning-based multi-data fusion include extracting feature vectors from each sensory input using CNN, concatenating each feature vector into a single feature vector, col. 1, lines 45-55, machine learning model may be a model which accepts sensor data collected by cameras and/or other sensors as inputs, deep learning, col. 6, lines 1-20), the collection of sensor data from the plurality of sensor types (For example, if a home has two sensors, a camera and a microphone, these sensors can have different hardware specifications or different ranges and capabilities. Sensors can include audio sensors, visual sensors, motion sensors, heat sensors, etc. The sensors can be used to calculate the location, direction, type, etc. of an activity performed within a particular area of interest, col. 2, lines 10-20, Each sensor can have different specifications, including what properties the sensor can detect, the distance at which the sensor can detect its specific properties, the mechanism by which the sensor detects properties, etc., col. 2, lines 38-41, machine learning model may be a model which accepts sensor data collected by cameras and/or other sensors as inputs, col. 6, lines 1-20, one or more cameras, one or more proximity sensors, one or more gyroscopes, one or more accelerometers, one or more magnetometers, a global positioning system (GPS) unit, an altimeter, one or more sonar or laser sensors, col. 13, lines 55-65) into a single vector space to produce a feature vector (Location information for each sensor is represented in a 2D map-based feature embedding space, col. 1, lines 40-50, “In the technique described in this application, the extracted feature vectors from the multiple sensors are represented in the form of a 2D map that represents a target location. For example, the map can represent the floor plan of a home, a factory, an office building, etc. The extracted feature vectors are then integrated into a single feature vector and used as input to a classifier. The installation location of each sensor is generally known, and each sensor has its own properties and constraints; and thus proper weights can be assigned to each sensor according to where the activity happens and the properties of each sensor to obtain better activity classification recognition”, col. 2, lines 1-15, create a map-based embedding space for data fusion of the sensor data from each of the multiple sensors, col. 4, lines 45-55, Once the raw feature vectors are extracted from the data received from each of multiple different sensors, they are integrated together to form one integrated feature vector, col. 6, lines 60-65) recognizing, by the one or more processors and using the feature vector, a feature associated with the plurality of sensor types within the environment (The extracted feature vectors are then integrated into a single feature vector and used as input to a classifier. The installation location of each sensor is generally known, and each sensor has its own properties and constraints; and thus proper weights can be assigned to each sensor according to where the activity happens and the properties of each sensor to obtain better activity classification recognition, col. 2, lines 5-15, The map can be represented as a feature vector. For example, the map can be represented as a 2D feature vector. A feature vector is a vector of numerical features that represent a particular item, such as the map, an activity, a location, etc., col. 6, lines 45-55, In some implementations, steps (312) and (314) can be performed concurrently or as a single, continuous step such that the system 200 detects a particular activity based on the integrated feature vector. Detecting a particular activity can include classifying, by the activity classifier 120, the received sensor data, using the activity classification model and based on the integrated feature vector. For example, feature vector integrator 220 can provide the labelled, integrated feature vector to the activity classifier 120 and receive activity classification data that indicates that a user has entered the kitchen from an adjoining room, col. 9, lines 35-45); mapping, by the one or more processors, the feature to the position of the robot (In the technique described in this application, the extracted feature vectors from the multiple sensors are represented in the form of a 2D map that represents a target location. For example, the map can represent the floor plan of a home, a factory, an office building, etc., col. 2, lines 1-5, For example, the first floor of the home can be presented in a 4×4 two-dimensional grid. The microphone can have a set range that is represented in the map based on the area that the microphone can cover, and the camera can have a set range that is represented in the map based on the area the camera can cover, col. 2, lines 55-60, System 200 implements the location-based multi-sensor fusion technique described in this application to perform human activity recognition. As described above, the technique utilizes a map of the target location (i.e., the first floor of the residential home) and is provided with the installed locations of each sensor of the multiple sensors. While the following description is drafted in the context of a home, it is understood that the disclosure can be directed to various types of property, such as office buildings, public buildings, etc. The map can be represented as a feature vector. For example, the map can be represented as a 2D feature vector. A feature vector is a vector of numerical features that represent a particular item, such as the map, an activity, a location, etc., col. 6, lines 38-51, robotic devices 490, and are configured to communicate sensor and image data to the one or more user devices, col. 19, lines 55-56); aggregating, by the one or more processors, the feature using a weighting algorithm to generate an aggregated feature (The extracted feature vectors are then integrated into a single feature vector and used as input to a classifier. The installation location of each sensor is generally known, and each sensor has its own properties and constraints; and thus proper weights can be assigned to each sensor according to where the activity happens and the properties of each sensor to obtain better activity classification recognition, col. 2, lines 5-15, Training module 110 allows the activity classifier to learn by changing the weights applied to different variables to emphasize or deemphasize the importance of the variable within the model. By changing the weights applied to variables within the model, training module 110 allows the model to learn which types of information (e.g., which sensor inputs, what locations, etc.) should be more heavily weighted to produce a more accurate activity classifier, col. 5, lines 44-52); and creating, by the one or more processors and based on the aggregated feature, an output predicting the feature present in the environment (generating, using the extracted feature vectors, an integrated feature vector, detecting a particular activity based on the integrated feature vector, and in response to detecting the particular activity, performing a monitoring action, abstract, In the technique described in this application, the extracted feature vectors from the multiple sensors are represented in the form of a 2D map that represents a target location. For example, the map can represent the floor plan of a home, a factory, an office building, etc. The extracted feature vectors are then integrated into a single feature vector and used as input to a classifier, col. 2, lines 1-10, train the classifier to create a map-based embedding space for data fusion of the sensor data from each of the multiple sensors according to the sensor information, col. 4, lines 45-55, Activity classifier 120 receives a labelled, integrated feature vector as input, and outputs activity classifications and locations based on the labelled, integrated feature vector, col. 5, lines 60-65, Feature vector integrator 220 integrates the extracted feature vectors from the feature extractors 212 into a single, integrated feature vector, col. 7, lines 25-27, In some implementations, detecting a particular activity based on the integrated feature vector includes providing, as input to a machine learning model using convolutional neural networks, the received sensor data. For example, the system 200 can provide the sensor data from sensors 202 to activity classifier 120 to output an activity classification, col. 8, lines 15-20). Ebrahimi Afrouzi et al. and Lee are in the same art of autonomous devices (Ebrahimi Afrouzi et al., [0003]; Lee, col. 1, lines 35-36, col. 13, lines 30-35). The combination of Lee with Ebrahimi Afrouzi et al. will enable embedding, by the one or more processors and using deep learning algorithms, the collection of sensor data from the plurality of sensor types into a single vector space to produce a feature vector. It would have been obvious at the time of filing to one of ordinary skill in the art combine the embedding of Lee with the invention of Ebrahimi Afrouzi et al. as this was known at the time of invention, the combination would have predictable results, and as Lee states “Each sensor can have different specifications, including what properties the sensor can detect, the distance at which the sensor can detect its specific properties, the mechanism by which the sensor detects properties, etc. The distance at which the sensor can detect a property, or the range of the sensor, can be specific to the type of sensor or the particular sensor”, col. 2, lines 35-50 and “In some implementations, the monitoring action includes taking a photo or video. For example, the system 200 can take a photo of a user turning off a stove and the user can later check to make sure that they have turned off the stove for safety reasons. In some implementations, the monitoring action includes activating a building automation system to perform an action. For example, the system 200 can activate a building automation system to automatically turn off a stove when it is detected that a user has left the house and the stove is still on in the kitchen” (col. 9, line 65 – col. 10, line 10) suggesting a safety benefit to combining inventions. Claim(s) 2-5 and 21-24 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ebrahimi Afrouzi et al. (US 20220066456 A1) and Lee (US 11550276 B1) as applied to claims 1 and 20 above, further in view of Philbin et al. (US 20210101624 A1). Regarding claims 2 and 21, Ebrahimi Afrouzi et al. and Lee disclose the method and system of claims 1 and 20. Ebrahimi Afrouzi et al. further indicate sensor data comprises two-dimensional RGB camera feed data and wherein recognizing the feature associated with the plurality of sensor types comprises feeding the two-dimensional camera feed data into an RGB recognizer algorithm (In some embodiments, a video that is in red, green, blue (RGB) format may be converted to a video in a different format, such as YCoCg color space format, [0266], In embodiments, a kernel may consist of multiple layers of feature maps, each designed to detect a different feature. All neurons in a single feature map share the same parameters and allow the network to recognize a feature pattern regardless of where the feature pattern is within the input. This is important for object detection. For example, once the network learns that an object positioned in a dwelling is a chair, the network will be able to recognize the chair regardless of where the chair is located in the future, [0274], In some embodiments, a camera of the robot (the camera used for SLAM or another camera) captures images or video while the robot navigates around the environment. Using object recognition, the processor may identify the TV within the images captured and may associate a location within the floor map with the TV, [0400], RGB, SLAM, [0464], The robot also includes a camera. The processor of the robot may use data collected by the camera to track a location of features, such as a light fixture, a corner, and an edge. In some embodiments, the camera may be slightly recessed and angled rearward. In some embodiments, the processor uses the location of features to localize the robot, [0589]). Ebrahimi Afrouzi et al. and Lee do not disclose processing RGB output into an ensemble predictor to predict the feature. Philbin et al. teach sensor data comprises two-dimensional RGB camera feed data (image sensors (e.g., red-green-blue (RGB), [0036]) and wherein the recognizing step comprises feeding the two-dimensional camera feed data into an RGB recognizer algorithm and processing RGB output into an ensemble predictor to predict the feature (In some examples, operation 408 may comprise an ensemble voting technique such as, for example, majority voting, plurality voting, weighted voting (e.g., where certain pipelines are attributed more votes, functionally) and/or an averaging technique such as simple averaging, weighted averaging, and/or the like. In other words, a first occupancy map may indicate that there is a 0.9 likelihood that a portion of the environment is occupied and a second occupancy map may indicate that there is a 0.8 likelihood that the portion is occupied. The techniques may comprise using the likelihoods in a voting technique to determine whether to indicate that the portion is occupied or unoccupied and/or averaging the likelihoods to associate an averaged likelihood therewith, [0076]). Ebrahimi Afrouzi et al. and Philbin et al. are in the same art of collision avoidance (Ebrahimi Afrouzi et al., [0307]; Philbin et al., abstract). The combination of Philbin et al. with Ebrahimi Afrouzi et al. and Lee will enable processing RGB output into an ensemble predictor to predict the feature. It would have been obvious at the time of filing to one of ordinary skill in the art combine the ensemble predictor of Philbin et al. with the invention of Ebrahimi Afrouzi et al. and Lee as this was known at the time of invention, the combination would have predictable results, and as Philbin et al. state “To safely operate, an autonomous vehicle may include multiple sensors and various systems for detecting and tracking events surrounding the autonomous vehicle and may take these events into account when controlling the autonomous vehicle. For example, the autonomous vehicle may detect and track every object within a 360-degree view of a set of cameras, LIDAR sensors, radar, and/or the like to control the autonomous vehicle safely” ([0002]) and “The techniques discussed herein may improve the safety of a vehicle by preventing invalid or risky trajectories from being implemented by the vehicle. In at least some examples, such techniques may further prevent collisions due to providing redundancy in such a way as to mitigate errors in any system or subsystem associated with the trajectory generation components (perception, prediction, planning, etc.). Moreover, the techniques may reduce the amount of computational bandwidth, memory, and/or power consumed for collision avoidance in comparison to former techniques. The accuracy of the collision avoidance system may also be higher than an accuracy of the primary perception system, thereby reducing an overall error rate of trajectories implemented by the autonomous vehicle by filtering out invalid trajectories” ([0022]) suggesting a safety benefit that would result from combing inventions. Regarding claims 3 and 22, Ebrahimi Afrouzi et al. and Lee and Philbin et al. disclose the method and system of claims 2 and 21. Ebrahimi Afrouzi et al. further indicate recognizing the feature associated with the plurality of sensor types comprises image segmentation based on color or contrast (processor of the robot may perform segmentation wherein an object captured in an image is separated from other objects and the background of the image, the processor identifies the object based on the characteristics and features of the object. Characteristics of the object, for example, may include shape, color, size, presence of a leaf, and positioning of the leaf., [0382], In some embodiments, classification of an area may be based on commonalities and differences. Commonalities may include, for example, objects, floor types, patterns on walls, corners, ceiling, painting on the walls, windows, doors, power outlets, light fixtures, furniture, appliances, brightness, curtains, and other commonalities and how each of these commonalities relate to one another. Examples of different commonalities observed for an area include a bed, the color of the walls and the tile flooring. Based on these observed commonalities, the processor may classify the area, [0430], For instance, a first image may be segmented using fixed segmentation, whereas other images may be segmented based on entropy and contrast, [0550]). Regarding claims 4 and 23, Ebrahimi Afrouzi et al. and Lee and Philbin et al. disclose the method of claims 3 and 22. Ebrahimi Afrouzi et al. and Lee and Philbin et al. further indicate the deep learning algorithm is a convolutional neural network (CNN) (Ebrahimi Afrouzi et al., [0273]-[0279], [0292], [0313], [0321], [0393]; Lee, col. 1, lines 50-51; Philbin et al., [0051]). Regarding claims 5 and 24, Ebrahimi Afrouzi et al. and Lee and Philbin et al. disclose the method and system of claims 4 and 23. Philbin et al. further indicate executing the CNN to extract features using an encoding algorithm and decoding the features corresponding to the two-dimensional RGB camera feed data (camera feed can be RGB, see [0036], ML can be CNN, see [0051], In some examples, the ML model(s) 302(1)-(n) may comprise an encoder-decoder network, although other architectures are contemplated. In an example that uses an encoder-decoder network with convolutional layers, the encoder layer(s) may use average pooling with a pooling size of (2,2) and the decoder may comprise bilinear up-sampling. Following the decoder, the architecture may comprise a single linear convolution layer that generates logits and a final layer may apply a softmax to produce the final output probabilities associated with the different object classifications (e.g., pedestrian, cyclist, motorcyclist, vehicle, other may be labeled ground), [0089], For example, FIG. 6A depicts an ML model 600 that is trained to determine occupancy maps based at least in part on lidar data and/or radar data may include an encoder comprising a set of five blocks consisting of a pair of convolutional layers with batch normalization followed by an average pooling layer. In some examples, the convolutional layers may comprise ReLU activations, although other activations are contemplated (e.g., sigmoid, hyperbolic tangent, leaky ReLU, parameteric ReLU, softmax, Swish) The decoder may include five blocks consisting of three convolutional layers with batch normalization. The network may additionally or alternatively comprise a skip connection 602 from the fourth block of the encoder to the second block of the decoder, [0090], Continuing an additional or alternate example of variations in the architectures of the ML model(s) 302(1)-(n), FIG. 6B depicts an ML model 604 trained to determine occupancy maps based at least in part on image data may comprise an encoder-decoder network built on top of a ResNet (or other vision) backbone. For example, the ResNet block may comprise three layers, although an additional or alternate ResNet or other vision backbone component may be used. In some examples, the encoder and decoder may four blocks on top of ResNet blocks and, in at least some examples, the architecture 604 may comprise an orthographic feature transform layer between the encoder and decoder. The images are in perspective view even though the output is in a top-down view. The orthographic feature transform layer may convert from pixel space to top-down space. In some examples, the orthographic feature transform layer may comprise a series of unbiased fully connected layers with ReLU activations, although other activations are contemplated (e.g., sigmoid, hyperbolic tangent, leaky ReLU, parameteric ReLU, softmax, Swish). In an example where the image data 304 comprises images from different cameras, the architecture 604 may be configured to receive images from each camera view through a shared encoder and the architecture may be trained to learn a separate orthographic transformation for each view, add together the projected features, and pass the result through a single decoder, [0091]). Claim(s) 6 and 25 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ebrahimi Afrouzi et al. (US 20220066456 A1) and Lee (US 11550276 B1) and Philbin et al. (US 20210101624 A1) as applied to claim 5 and 24 above, further in view of Guo et al. (US 20230177637 A1). Regarding claims 6 and 25, Ebrahimi Afrouzi et al. and Lee and Philbin et al. disclose the method and system of claims 5 and 24. Ebrahimi Afrouzi et al. and Lee and Philbin et al. do not disclose the CNN comprises convolution layers to produce feature vectors. Guo et al. teach a CNN comprising convolution layers to produce feature vectors (The teacher network 530 comprises an encoder 532 having a plurality of convolution layers and pooling layers that reduce the dimensionality of the input images 520 to provide encoded intermediate perception outputs (e.g., a feature vector) at an encoded layer 534 (i.e., a bottleneck layer). In at least one embodiment, the encoder 532 includes densely connected convolutional layers (e.g., DenseNet169). The teacher network 530 further comprises a decoder 536 having a plurality of convolution or deconvolution layers and unpooling or upsampling layers that increase the dimensionality of the encoded intermediate perception outputs from the encoded layer 534 to generate an output logit (e.g., having dimensions h1×w1×N, where h1 and w1 are the height and width, respectively, of the perspective projection images 520). The output logit is normalized by the output layer 540 (e.g., softmax) to provide the final perception outputs of the teacher network 530 (e.g., having dimensions h1×w1), [0044] The student network 550 comprises an encoder 552 having a plurality of convolution layers and pooling layers that reduce the dimensionality of the input image 510 to provide encoded intermediate perception outputs (e.g., a feature vector) at an encoded layer 554 (i.e., a bottleneck layer). In at least one embodiment, the encoder 552 include densely connected convolutional layers (e.g., DenseNet121). The student network 550 further comprises a decoder 556 having a plurality of convolution or deconvolution layers and unpooling or upsampling layers that increase the dimensionality of the encoded intermediate perception outputs from the encoded layer 554 to generate an output logit (e.g., having dimensions h2×w2×N, where h2 and w2 are the height and width, respectively, of the omnidirectional image 510). The output logit is normalized by the output layer 560 (e.g., softmax) to provide the final perception outputs of the student network 550 (e.g., having dimensions h2×w2), [0049]). Ebrahimi Afrouzi et al. and Guo et al. are in the same art of autonomous devices (Ebrahimi Afrouzi et al., [0003]; Guo et al., [0003], [0056]). The combination of Guo et al. with Ebrahimi Afrouzi et al. and Lee and Philbin et al. will enable using a CNN comprising convolution layers to produce feature vectors. It would have been obvious at the time of filing to one of ordinary skill in the art combine the convolution layers of Guo et al. with the invention of Ebrahimi Afrouzi et al. and Lee and Philbin et al. as this was known at the time of invention, the combination would have predictable results, and as Guo et al. state “By way of this training, the student model learns to perform the same machine perception task, except in the omnidirectional image domain, using limited or no suitably labeled training data in the omnidirectional image domain” (abstract) demonstrating an improvement to training efficiency and decreasing a need for human man hours by requiring less labeled data. Claim(s) 7-14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ebrahimi Afrouzi et al. (US 20220066456 A1) and Lee (US 11550276 B1) and Philbin et al. (US 20210101624 A1) as applied to claim 2 above, further in view of Papi et al. (US 12416730 B1). Regarding claim 7, Ebrahimi Afrouzi et al. and Lee and Philbin et al. disclose the method of claim 2. Ebrahimi Afrouzi et al. and Philbin et al. do not disclose the sensor data further comprises two-dimensional infrared sensor data and wherein the recognizing step further comprises feeding the two-dimensional infrared sensor data an infrared recognizer algorithm; and combining the RGB output and infrared output into an ensemble predictor to predict the feature. Papi et al. teach sensor data further comprises two-dimensional infrared sensor data and wherein the recognizing step further comprises feeding the two-dimensional infrared sensor data an infrared recognizer algorithm; and combining the RGB output and infrared output into an ensemble predictor to predict the feature (Object detection and tracking systems may use machine-learned transformer models with self-attention for detecting, classifying, and/or tracking objects in an environment. Techniques described herein may include receiving sensor data generated by different sensor modalities of a vehicle, determining different bounding shapes based on the different sensor modalities, and using a machine-learned transformer model to determine associated and/or combined bounding shapes, abstract, Various examples herein relate to receiving sensor data (and/or object detections or bounding shapes determined based on the sensor data) from with different sensor modalities. As used herein, a sensor modality may refer to a type of sensor data and/or to a type of sensor configured to capture or process sensor data. Examples of sensor modalities may include, but are not limited to, lidar, radar, vision (e.g., image and/or video), sonar, depth, time-of-flight, audio, cameras (e.g., RGB, IR, intensity, depth, etc.), and the like, col. 6, lines 40-50, Each combined object detection may include an updated/refined set of attributes (e.g., location, size dimensions, yaw, classification, intent, etc.) based on the attributes of the associated object detections from the different sensor modalities. The ML transformer model 106 may be trained to determine an optimal set of attributes for each combined object detection, so that the combined object detection represents the corresponding object in the environment 114 more accurately than any of the individual object detections from the different sensor modalities, col. 9, lines 35-60, Although discussed in the context of neural networks, any type of machine-learning can be used consistent with this disclosure. For example, machine-learning algorithms can include ensemble, col. 25, lines 5-45). Ebrahimi Afrouzi et al. and Papi et al. are in the same art of autonomous devices (Ebrahimi Afrouzi et al., [0003]; Papi et al., col. 5, lines 50-67). The combination of Papi et al. with Ebrahimi Afrouzi et al. and Lee and Philbin et al. will enable combining the RGB output and infrared output into an ensemble predictor to predict the feature. It would have been obvious at the time of filing to one of ordinary skill in the art combine the data combination of Papi et al. with the invention of Ebrahimi Afrouzi et al. and Lee and Philbin et al. as this was known at the time of invention, the combination would have predictable results, and as Papi et al. state, “For at least these reasons, the techniques described herein also can improve the safe operation of autonomous vehicles. For instance, the disclosed techniques, among other things, improve an autonomous vehicle's ability to detect, classify, and track certain objects in an environment. Being able to detect, classify, and track objects may be critical for the overall safety and quality of autonomous driving. The technologies disclosed herein can classify objects based on a combination of sensor modalities, such as vision (e.g., images), lidar, and/or radar data. For instance, using the technologies described herein, an object can be tracked with high certainty as to the object's location, size, velocity, yaw, and classification, etc. This is due to the ability to process and analyze object detections and/or bounding shapes from different sensor modalities in an ML transformer model, and determine associated (e.g., combined) object detections with improved accuracy over the object detections generated by the individual sensor modalities” (Col. 5, lines 50-67) thereby providing an accuracy benefit and therefore likely safety improvement when the inventions are combined. Regarding claim 8, Ebrahimi Afrouzi et al. and Lee and Philbin et al. and Papi et al. disclose the method of claim 7. Ebrahimi Afrouzi et al. further indicate recognizing the feature associated with the plurality of sensor types comprises image segmentation based on color or contrast and further based on heat gradients (processor of the robot may perform segmentation wherein an object captured in an image is separated from other objects and the background of the image, the processor identifies the object based on the characteristics and features of the object. Characteristics of the object, for example, may include shape, color, size, presence of a leaf, and positioning of the leaf, [0382], In some embodiments, classification of an area may be based on commonalities and differences. Commonalities may include, for example, objects, floor types, patterns on walls, corners, ceiling, painting on the walls, windows, doors, power outlets, light fixtures, furniture, appliances, brightness, curtains, and other commonalities and how each of these commonalities relate to one another. Examples of different commonalities observed for an area include a bed, the color of the walls and the tile flooring. Based on these observed commonalities, the processor may classify the area, [0430], For instance, a first image may be segmented using fixed segmentation, whereas other images may be segmented based on entropy and contrast, [0550], “In some embodiments, the user interface may display information about a current state of the robot or previous states of the robot or its environment. Examples may include a heat map of dirt or debris sensed over an area, visual indications of classifications of floor surfaces in different areas of the map, visual indications of a path that the robot has taken during a current cleaning session or other type of work session, visual indications of a path that the robot is currently following and has computed to plan further movement in the future, and visual indications of a path that the robot has taken between two points in the environment, like between a point A and a point B on different sides of a room or a house in a point-to-point traversal mode”, [1417] ) [heat map interpreted as heat gradients]. Regarding claim 9, Ebrahimi Afrouzi et al. and Lee and Philbin et al. and Papi et al. disclose the method of claim 8. Ebrahimi Afrouzi et al. and Philbin et al. and Papi et al. further indicate the deep learning algorithm is a convolutional neural network (CNN) (Ebrahimi Afrouzi et al., [0273]-[0279], [0292], [0313], [0321], [0393]; Philbin et al., [0051]; Papi et al., col. 3, line 34 - col. 4, line 3, col. 25, lines 5-45). Regarding claim 10, Ebrahimi Afrouzi et al. and Lee and Philbin et al. and Papi et al. disclose the method of claim 9. Philbin et al. and Papi et al. further indicate executing the CNN to extract features using an encoding algorithm and decoding the features corresponding to the two-dimensional RGB camera feed data and the two-dimensional infrared sensor data (Philbin et al., camera feed can be RGB, infrared, see [0036], ML can be CNN, see [0051], In some examples, the ML model(s) 302(1)-(n) may comprise an encoder-decoder network, although other architectures are contemplated. In an example that uses an encoder-decoder network with convolutional layers, the encoder layer(s) may use average pooling with a pooling size of (2,2) and the decoder may comprise bilinear up-sampling. Following the decoder, the architecture may comprise a single linear convolution layer that generates logits and a final layer may apply a softmax to produce the final output probabilities associated with the different object classifications (e.g., pedestrian, cyclist, motorcyclist, vehicle, other may be labeled ground), [0089], For example, FIG. 6A depicts an ML model 600 that is trained to determine occupancy maps based at least in part on lidar data and/or radar data may include an encoder comprising a set of five blocks consisting of a pair of convolutional layers with batch normalization followed by an average pooling layer. In some examples, the convolutional layers may comprise ReLU activations, although other activations are contemplated (e.g., sigmoid, hyperbolic tangent, leaky ReLU, parameteric ReLU, softmax, Swish) The decoder may include five blocks consisting of three convolutional layers with batch normalization. The network may additionally or alternatively comprise a skip connection 602 from the fourth block of the encoder to the second block of the decoder, [0090], Continuing an additional or alternate example of variations in the architectures of the ML model(s) 302(1)-(n), FIG. 6B depicts an ML model 604 trained to determine occupancy maps based at least in part on image data may comprise an encoder-decoder network built on top of a ResNet (or other vision) backbone. For example, the ResNet block may comprise three layers, although an additional or alternate ResNet or other vision backbone component may be used. In some examples, the encoder and decoder may four blocks on top of ResNet blocks and, in at least some examples, the architecture 604 may comprise an orthographic feature transform layer between the encoder and decoder. The images are in perspective view even though the output is in a top-down view. The orthographic feature transform layer may convert from pixel space to top-down space. In some examples, the orthographic feature transform layer may comprise a series of unbiased fully connected layers with ReLU activations, although other activations are contemplated (e.g., sigmoid, hyperbolic tangent, leaky ReLU, parameteric ReLU, softmax, Swish). In an example where the image data 304 comprises images from different cameras, the architecture 604 may be configured to receive images from each camera view through a shared encoder and the architecture may be trained to learn a separate orthographic transformation for each view, add together the projected features, and pass the result through a single decoder, [0091]; Papi et al., Object detection and tracking systems may use machine-learned transformer models with self-attention for detecting, classifying, and/or tracking objects in an environment. Techniques described herein may include receiving sensor data generated by different sensor modalities of a vehicle, determining different bounding shapes based on the different sensor modalities, and using a machine-learned transformer model to determine associated and/or combined bounding shapes, abstract, Various examples herein relate to receiving sensor data (and/or object detections or bounding shapes determined based on the sensor data) from with different sensor modalities. As used herein, a sensor modality may refer to a type of sensor data and/or to a type of sensor configured to capture or process sensor data. Examples of sensor modalities may include, but are not limited to, lidar, radar, vision (e.g., image and/or video), sonar, depth, time-of-flight, audio, cameras (e.g., RGB, IR, intensity, depth, etc.), and the like, col. 6, lines 40-50, Each combined object detection may include an updated/refined set of attributes (e.g., location, size dimensions, yaw, classification, intent, etc.) based on the attributes of the associated object detections from the different sensor modalities. The ML transformer model 106 may be trained to determine an optimal set of attributes for each combined object detection, so that the combined object detection represents the corresponding object in the environment 114 more accurately than any of the individual object detections from the different sensor modalities, col. 9, lines 35-60) [Philbin et al. teaches concepts of CNN encoding, decoding, RGB data, infrared data, Papi et al. teach concept of combining RGB and infrared data]. Regarding claim 11, Ebrahimi Afrouzi et al. and Lee and Philbin et al. and Papi et al. disclose the method of claim 7. Papi et al. further indicate the sensor data further comprises LIDAR three-dimensional point cloud data fed into a point cloud recognizer, and wherein recognizing the feature associated with the plurality of sensor types comprises combining an output of the point cloud recognizer with the RGB output and the infrared output in the ensemble predictor to predict the feature (Object detection and tracking systems may use machine-learned transformer models with self-attention for detecting, classifying, and/or tracking objects in an environment. Techniques described herein may include receiving sensor data generated by different sensor modalities of a vehicle, determining different bounding shapes based on the different sensor modalities, and using a machine-learned transformer model to determine associated and/or combined bounding shapes, abstract, Various examples herein relate to receiving sensor data (and/or object detections or bounding shapes determined based on the sensor data) from with different sensor modalities. As used herein, a sensor modality may refer to a type of sensor data and/or to a type of sensor configured to capture or process sensor data. Examples of sensor modalities may include, but are not limited to, lidar, radar, vision (e.g., image and/or video), sonar, depth, time-of-flight, audio, cameras (e.g., RGB, IR, intensity, depth, etc.), and the like, col. 6, lines 40-50, Each combined object detection may include an updated/refined set of attributes (e.g., location, size dimensions, yaw, classification, intent, etc.) based on the attributes of the associated object detections from the different sensor modalities. The ML transformer model 106 may be trained to determine an optimal set of attributes for each combined object detection, so that the combined object detection represents the corresponding object in the environment 114 more accurately than any of the individual object detections from the different sensor modalities, col. 9, lines 35-60, Although discussed in the context of neural networks, any type of machine-learning can be used consistent with this disclosure. For example, machine-learning algorithms can include ensemble, col. 25, lines 5-45) [LIDAR, RGB listed as optional modalities] Regarding claim 12, Ebrahimi Afrouzi et al. and Lee and Philbin et al. and Papi et al. disclose the method of claim 11. Papi et al. further indicate the ensemble predictor comprises a linear rule-based model which combines the output of the point cloud recognizer, the RGB output and the projected infrared output (perception component may use a machine-learned transformer model with self-attention to determine associated and/or combined object detections (e.g., bounding shapes) representing the objects in the environment, col. 2, lines 18 - 56, transformer model to receive input bounding shapes from the various sensor modalities, and to output a corresponding set (or stream) of object detections with a one-to-one constraint to solve over segmentation, col. 4, lines 4 - 25, Each of the ML lidar pipeline 122, the ML image pipeline 124, and/or the ML radar pipeline 126 may include one or more machine-learned model(s) incorporating any combination of machine-learning components (e.g., multilayer perceptrons, feedforward neural networks, attention components, etc.), col. 8, line 42 - col. 9, line 3, include linear regression, logistic regression, Linear Discriminant Analysis (LDA), association rule learning algorithms e.g., perceptron, col. 25, lines 5-45). Regarding claim 13, Ebrahimi Afrouzi et al. and Lee and Philbin et al. and Papi et al. disclose the method of claim 12. Philbin et al. and Papi et al. further indicate the ensemble predictor weights the output of the point cloud recognizer, the RGB output and the infrared output (Philbin et al., In some examples, operation 408 may comprise an ensemble voting technique such as, for example, majority voting, plurality voting, weighted voting (e.g., where certain pipelines are attributed more votes, functionally) and/or an averaging technique such as simple averaging, weighted averaging, and/or the like; Papi et al., Various examples herein relate to receiving sensor data (and/or object detections or bounding shapes determined based on the sensor data) from with different sensor modalities. As used herein, a sensor modality may refer to a type of sensor data and/or to a type of sensor configured to capture or process sensor data. Examples of sensor modalities may include, but are not limited to, lidar, radar, vision (e.g., image and/or video), sonar, depth, time-of-flight, audio, cameras (e.g., RGB, IR, intensity, depth, etc.), and the like, col. 6, lines 40-50, Each combined object detection may include an updated/refined set of attributes (e.g., location, size dimensions, yaw, classification, intent, etc.) based on the attributes of the associated object detections from the different sensor modalities. The ML transformer model 106 may be trained to determine an optimal set of attributes for each combined object detection, so that the combined object detection represents the corresponding object in the environment 114 more accurately than any of the individual object detections from the different sensor modalities, col. 9, lines 35-60, Although discussed in the context of neural networks, any type of machine-learning can be used consistent with this disclosure. For example, machine-learning algorithms can include ensemble, col. 25, lines 5-45). Regarding claim 14, Ebrahimi Afrouzi et al. and Lee and Philbin et al. and Papi et al. disclose the method of claim 13. Ebrahimi Afrouzi et al. and Lee and Philbin et al. and Papi et al. further indicate an inertial measurement unit (IMU) associated with the robot, wherein the IMU is configured to determine a position of the robot in the environment, and wherein tracking the position of the robot is performed using the IMU (Ebrahimi Afrouzi et al., [0238], [0244], [0245], [0420]; Lee, For instance, the robotic devices 490 may navigate within the home using one or more cameras, one or more proximity sensors, one or more gyroscopes, one or more accelerometers, one or more magnetometers, a global positioning system (GPS) unit, an altimeter, one or more sonar or laser sensors, and/or any other types of sensors that aid in navigation about a space; Philbin et al., [0045], [vehicle has robotic control, [0023]; Papi et al., col. 7, lines 18-60, col. 22, lines 40-65 [robotic col. 6, lines 1-15]). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHELLE ENTEZARI whose telephone number is (571)270-5084. The examiner can normally be reached 10-7 M-F. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent M Rudolph can be reached at (571) 272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MICHELLE M ENTEZARI/Primary Examiner, Art Unit 2671
Read full office action

Prosecution Timeline

Show 3 earlier events
Jun 15, 2026
Interview Requested
Jun 30, 2026
Applicant Interview (Telephonic)
Jun 30, 2026
Examiner Interview Summary
Jul 14, 2026
Response Filed
Aug 12, 2026
Final Rejection mailed — §103
Sep 10, 2026
Interview Requested
Sep 24, 2026
Applicant Interview (Telephonic)
Sep 24, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725301
SEMANTIC VISUAL FEATURE SHARING
2y 11m to grant Granted Sep 01, 2026
Patent 12718585
ACCURACY FOR OBJECT DETECTION
2y 11m to grant Granted Aug 25, 2026
Patent 12711781
IDENTIFYING BIDIRECTIONAL CHANNELIZATION ZONES AND LANE DIRECTIONALITY
3y 5m to grant Granted Aug 18, 2026
Patent 12700257
CASCADED DETECTION OF FACIAL ATTRIBUTES
3y 0m to grant Granted Aug 04, 2026
Patent 12688722
SIMULATION OF LABEL DATA TO OPTIMIZE THE VISUAL DOCUMENT UNDERSTANDING BY USING PDFS ANNOTATION AWARE METHODOLOGY
2y 10m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
77%
Grant Probability
98%
With Interview (+21.1%)
2y 12m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 883 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month