Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed 7/06/2026 have been fully considered but they are not persuasive.
Regarding 35 U.S.C 101 rejection, applicant argues in page 10-11 “On page 4 of the Office Action, claims 1-21 were rejected under 35 U.S.C. § 101 because the claimed invention is allegedly directed to abstract idea without significantly more. Applicant respectfully traverses these rejections.
Applicant submits that the claims are not directed to an abstract idea. Even assuming arguendo that the claims recite an abstract idea, the claims integrate any such idea into a practical application by reciting specific graphical user interface (GUI) mechanisms that improve computer functionality.
The Office Action characterizes the claims as reciting a mental process of determining values associated with metrics. However, the claims do not merely recite determining values. Rather, they recite specific GUI functionality that visually presents model-estimated location information and ground truth locations, and dynamically adjusts visual presentations based on user feedback. These are not steps that can practically be performed in the human mind.” The applicant claims the office characterizes the claims as reciting a mental process. Based on the argument the applicant is referring to the previous claim limitations of claim 1 “determine values associated with metrics based on the obtained out”. The claim recited an abstract idea as the limitation does not claim any GUI or other process that cannot be performed by the human mind as explained in the previous office action (See MPEP 2106.04(a)(2)III). Further applicant argues page 11-15 “2. The Claims Are Directed to a Specific Improvement in GUI Technology Even assuming arguendo that the claims recite an abstract idea under Step 2A, Prong 1 (which Applicant does not concede), Applicant respectfully submits that the claims integrate any such idea into a practical application under Step 2A, Prong 2. Specifically, the claims recite additional elements that provide a specific improvement to the functioning of computers, namely, a specific graphical user interface that improves how computers display and present ML model evaluation information.
Accordingly, like the claims in Trading Technologies and Core Wireless, the present claims are directed to a specific improvement in GUI technology that provides concrete benefits (e.g., in speed, accuracy, and usability). The claimed functionality is not an abstract idea, but rather a specific technical improvement in how computers display and present ML model evaluation information. Applicant respectfully submits that claim 1 is eligible under 35 U.S.C. § 101 as the claim integrates any alleged abstract idea into a practical application. Applicant respectfully requests reconsideration and withdrawal of the rejection. Independent claims 9 and 17 include similar claim language and are patentable for at least the same reasons. “ Applicant argues how the specification shows that the claims improve upon the technology. However when determining any improvement the specification is consulted to determine if the claims purport the proposed improvement (See MPEP 2106.05(a)). Further argument is unclear as to whether the applicant is arguing the amended claim limitations. Based on the current limitations the amendments continue fail to provide significantly more than the abstract idea as the GUI elements only display information regarding the model. Independent claims do not provide what type of feedback is given to the model nor how the model is updated based on the feedback. The claim limitation “dynamically adjust, based on user feedback, an individual visual presentation of the model-estimated location information or the ground truth location based on the location-error value for the object, thereby providing feedback for validating ML model performance” only limits the claim to adjusting the location of the estimates based on the user changes. Further the amended claim limitations have not been examined and thus the argument is moot and not convincing. See updated 101 rejection.
Regarding 35 U.S.C 103 rejection, applicant argues in page 15-18 “Applicant respectfully disagrees with the assertion that the claims, as previously presented, are obvious over the references of record (alone or in combination) suggested by the Office Action. However, to further prosecution and without admission, claim 1 has been amended to recite, in relevant part:
(Emphasis added). Independent claims 9 and 17 have been similarly amended. The references of record, including Zaremba, Arroyo, and Xie, alone or in any combination, fail to teach or suggest such features.” The applicant claims that it was stated that Arroyo does not teach the cited limitations. However the office action only depended on Arroyo for the specific areas that it was quoted. An examiner can use the prior art as a puzzle piece and combine art based on the elements that are needed (See MPEP 2141). Further the updated 103 rejection does not depend on Xie. The applicant further argues amended claim limitations and the amended limitations have not been examined, thus the argument is moot and not convincing.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-21 rejected under 35 U.S.C. 101 because the claimed invention is directed to abstract idea without significantly more. The subject matter eligibility test for products and process is describe below for claim 1 in view of dependent claims.
Regarding claim 1:
Step 1: Is the claim to a process machine manufacture or composition of matter?
Yes – Claim 1 recites a method and that falls under the statutory categories.
Step 2A Prong 1: Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes – The claim recites the following:
“determining, for one or more of the objects, a location-error value by comparing the model- estimated location information indicated by the output for the object to a ground truth location of the object;” - The limitations recites a mental process of determine location error value by comparing model estimated location to ground truth location of the object (see MPEP 2106.04(a)(2)III).
Step 2 Prong 2: Does the claim recite additional elements that integrate the judicial exception into a particular application? No –
The claim includes the additional element(s):
“A method implemented by a system of one or more processors, the method comprising: obtaining information associated with a machine learning (ML) model, wherein the ML model is associated with autonomous or semi-autonomous operation of a vehicle;”
The additional elements fall under Insignificant Extra-Solution Activity as mere data gathering by obtaining information associated with a ML model. See MPEP 2106.5(g).
“obtaining validation data, wherein the validation data includes one or more video sequences obtained from image sensors of an end-user vehicle;”
The additional elements fall under Insignificant Extra-Solution Activity as mere data gathering by obtaining validation data. See MPEP 2106.5(g).
“obtaining output via computing a forward pass-through the ML model using the validation data, wherein the output indicates, at least, model-estimated location information associated with objects detected via the ML model in the validation data;”
The additional elements fall under “apply it” as using a generic computer to obtain output by passing by processing the validation data by the ML model. See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)).
“generating user interface information based on one or more of the determined location- error value[[s]] or obtained output, wherein the user interface information causes a user interface to:”
The additional elements fall under “apply it” as using a generic computer to generate user interface information. See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)).
“visually present the model-estimated location information and the ground truth location of the object proximate to the end-user vehicle, enabling a user to assess ML model accuracy;”
The additional elements fall under “apply it” as using a generic computer to generate user interface information based on model estimations and ground truth location so that user can access ML model accuracy. See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)).
“dynamically adjust, based on user feedback, an individual visual presentation of the model-estimated location information or the ground truth location based on the location-error value for the object, thereby providing feedback for validating ML model performance.”
The additional elements fall under “apply it” as using a generic computer to update the visual presentation based on the model-estimated location or the ground truth location to provide feedback for validation. See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)).
Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception?
No - The claim does not include additional elements that are sufficient to amount to a significantly more than the judicial exemption. As an order whole, the claim is directed to obtaining results by using machine learning model and displaying results. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements of obtaining and generating fall under using generic computer to apply an exemption and mere data gathering. The method does not improve on the function of a computer, transforms an article into another article, nor is it applied by a particular machine, making the claim not patent eligible.
Regarding claim 2:
Step 2A Prong 2, Step 2B: The additional element(s):
“The method of claim 1, wherein the validation data further includes one or more of velocities of the end-user vehicle.”
The additional elements fall under Insignificant Extra-Solution Activity. See MPEP 2106.5(g).
Regarding claim 3:
Step 2A Prong 2, Step 2B: The additional element(s):
“The method of claim 1, wherein the user interface information includes a graphical representation of objects detected via the ML model.”
The additional elements fall under Insignificant Extra-Solution Activity. See MPEP 2106.5(g). The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application
Regarding claim 4:
Step 2A Prong 2, Step 2B: The additional element(s):
“The method of claim 3, wherein the user interface information further includes error information associated with the objects.”
The additional elements fall under Insignificant Extra-Solution Activity. See MPEP 2106.5(g). The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application
Regarding claim 5:
Step 2A Prong 2, Step 2B: The additional element(s):
“The method of claim 1, wherein the user interface: presents a graphical depiction of the end-user vehicle; presents ground truth locations of the objects which are proximate to the end-user vehicle
The additional elements fall under “apply it” as using a generic computer to present a graphical depiction of the end-user vehicle and preset ground truth locations . See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)). The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application
Regarding claim 6:
Step 2A Prong 2, Step 2B: The additional element(s):
“The method of claim [[5]] 1, wherein each adjusted visual presentation includes a color whose radius is selected based on the location-error value for the object.”
The additional elements fall under “apply it” as using a generic computer to include color in the presentation. See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)).
The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application.
Regarding claim 7:
Step 2A Prong 2, Step 2B: The additional element(s):
“The method of claim [[5]] 1, wherein the user interface visually presents a particular video sequence and wherein the ground truth locations of the objects are dynamically updated based on the video sequence.”
The additional elements fall under Insignificant Extra-Solution Activity. See MPEP 2106.5(g).
The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application.
Regarding claim 8:
Step 2A Prong 2, Step 2B: The additional element(s):
“The method of claim 7, wherein the individual visual presentations of the ground truth locations are dynamically updated based on the video sequence.”
The additional elements fall under Insignificant Extra-Solution Activity. See MPEP 2106.5(g).
The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application.
Claims 9-16 recites a system and are analogous to the method of claims 1-8. Therefore, the rejections of claim 1-8 above applies to claims 9-16.
Claims 17-21 recite a computer readable medium product and are analogous to the method of claims 1 and 5-8. Therefore, the rejections of claim 1 and 5-8 above applies to claims 17-21.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 3, 9, 11, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Jarquin Arroyo et al. (US20220111864A1) (“Arroyo”) in view of Crego et al. (US11810365B1) (“Crego”) and further in view of Gallo et al. (US20210150088A1) (“Gallo”).
Regarding claim 1 and analogous claims 9 and 17, Arroyo teaches A method implemented by a system of one or more processors, the method comprising: obtaining information associated with a machine learning (ML) model, wherein the ML model is associated with autonomous or semi-autonomous operation of a vehicle (Arroyo Para. 0022 line 1-12, As a general matter, it is desirable that a given sensing-model instance (and, more broadly, a given sensing system) of a given autonomous vehicle provide the control system of the vehicle with an accurate and robust representation of the current environment in which the autonomous vehicle is operating. Better data makes for better decisions. In many current implementations, the development process for a given sensing system of an ADS of an autonomous vehicle makes use of datasets captured from real-world environments with multiple sensors. These datasets are often annotated with what are referred to as "ground-truth labels" to aid in evaluating the accuracy of the sensing tasks.
para. 0033 line 1-10, FIG. 1 depicts an example geographic snapshot 100, in accordance with at least one embodiment. In particular, the geographic snapshot 100 depicts a moment in time in a given geographic area, and further depicts a number of entities ( a person, a bicycle, a number of cars, and so on) that have been detected and had ground-truth labels associated with them by a trained instance of a machine learning model. Identifying objects using, e.g., machine learning and applying labels to the identified objects is known to those of skill in the art [wherein the ML model is associated with autonomous or semi-autonomous operation of a vehicle;].
Para 0036, FIG. 2 depicts an example model-instance-and dataset-evaluation system 200, in accordance with at least one embodiment. Depicted on the left side of FIG. 2 is a set of automated-driving datasets 202 including, as examples a dataset 204, a dataset 206, and a dataset 208. Any number of automated driving datasets could be utilized in connection with a given embodiment. Moreover, in at least one embodiment, the depicted model-instance-and-dataset-evaluation system 200 processes one automated-driving dataset 202 at a time, and a plurality of automated-driving datasets 202 are depicted in FIG. 2 as an illustration of the cross-domain capabilities of the model-instance-and-dataset-evaluation system 200 [obtaining information associated with a machine learning (ML) model]);
obtaining validation data, wherein the validation data includes one or more video sequences obtained from image sensors of an end-user vehicle (Arroyo Para 0036, FIG. 2 depicts an example model-instance-and dataset-evaluation system 200, in accordance with at least one embodiment. Depicted on the left side of FIG. 2 is a set of automated-driving datasets 202 including, as examples a dataset 204, a dataset 206, and a dataset 208. Any number of automated driving datasets could be utilized in connection with a given embodiment. Moreover, in at least one embodiment, the depicted model-instance-and-dataset-evaluation system 200 processes one automated-driving dataset 202 at a time, and a plurality of automated-driving datasets 202 are depicted in FIG. 2 as an illustration of the cross-domain capabilities of the model-instance-and-dataset-evaluation system 200 [obtaining validation data, wherein].
Para 0068 line 1-11, As can be seen in FIG. 3, the depicted example dataframe transform tree 300 includes, at its center, a global node 302. In the context of the dataset being an automated driving dataset, the global node 302 may represent a center of mass of the corresponding vehicle. The dataframe transform tree 300 further includes a camera 308 associated with an image 328, a camera 310 associated with an image 330, a camera 312 associated with an image 332, a camera 314 associated with an image 334, a camera 316 associated with an image 336, a camera 318 associated with an image 338, and a camera 320 associated with an image 340.
Para. 0084, As a general matter, datasets usually contain ground-truth annotations on a per-frame basis to support the supervised training of sensing-model instances, the validation of the trained model instances on the particular subset of data, and the like. As described elsewhere in the present disclosure, an annotation is often represented by (i) a bounding box positioned and oriented in a specific frame, (ii) a category label, (iii) an instance id, and (if needed) (iv) additional metadata. Embodiments of the present disclosure provide access to not only individual annotations within the dataset content but also to their properties and transformations [wherein the validation data includes one or more video sequences obtained from image sensors of an end-user vehicle;]);
obtaining output via computing a forward pass-through the ML model using the validation data, wherein the output indicates, at least, model-estimated location information associated with objects detected via the ML model in the validation data (Arroyo Para. 0035 line 1-16, Moreover, also in the geographic snapshot 100, a street 104 includes a vehicle 108 (car) inside a bounding box 110 and accompanied by a label 112 that reads "car 00 (driving)." Similarly, a vehicle 114 is shown in a bounding box 116 with an accompanying label 118 that reads "car 01 (driving)." The geographic snapshot 100 also includes a parking area 106 that includes four depicted parking spots. A vehicle 132 (car) is depicted in a bounding box 134 and having a label 136 that reads "car 04 (parked)." Furthermore, a vehicle 138 (car) is depicted in a bounding box 140 and having a label 142 that reads "car 05 (parked)." Finally, a vehicle 144 (car) is depicted in a bounding box 146 and having a label 148 that reads "car 06 (parked)." The arrangement shown in FIG. 1 is provided by way of example to illustrate the type of output that a machine-learning-model instance may provide during training and operation.
Para. 0036, FIG. 2 depicts an example model-instance-and dataset- evaluation system 200, in accordance with at least one embodiment. Depicted on the left side of FIG. 2 is a set of automated-driving datasets 202 including, as examples a dataset 204, a dataset 206, and a dataset 208. Any number of automated driving datasets could be utilized in connection with a given embodiment. Moreover, in at least one embodiment, the depicted model-instance-and-dataset-evaluation system 200 processes one automated-driving dataset 202 at a time, and a plurality of automated-driving datasets 202 are depicted in FIG. 2 as an illustration of the cross-domain capabilities of the model-instance-and-dataset-evaluation system 200 [obtaining output via computing a forward pass-through the ML model using the validation data, wherein the output indicates, at least,].
Para. 0063, As can be seen in FIG. 4, depth data (e.g., a lidar "point cloud") and visible-light data (e.g., a camera "image") are provided to various feature-extraction processes under the heading of "independent feature extraction," as is known in the art. The output of that portion of the processing is fed into a module to conduct sensor fusion in order to maximize fault tolerance. The result of the fusion process is a detection network, which can then be used to produce outputs similar to those shown in FIG. 1-i.e., class labels and bounding boxes. Also depicted in the model instance training process 400 as being provided as output is an aleatoric uncertainty estimator, as is known in the art. In at least one embodiment, it is those types of outputs that may be synthesized and summarized and displayed by the model instance- and-dataset-evaluation system 200 as the test results 260 of FIG. 2 [at least, model-estimated location information associated with objects detected via the ML model in the validation data].);
and generating user interface information based on one or more of the determined location- error value[[s]] or obtained output, wherein the user interface information causes a user interface to: visually present the model-estimated location information and the ground truth location of the object proximate to the end-user vehicle, enabling a user to assess ML model accuracy (Arroyo
PNG
media_image1.png
530
981
media_image1.png
Greyscale
Para 0039 line1-7, In various embodiments, a user 210 may interact with the system via any suitable user interface (e.g., graphical, text-based, command-line, and/or the like). In the embodiments that are primarily described in the present disclosure, the user 210 interacts with the model-instance and-dataset-evaluation system 200 using command-line instructions, as described more fully below [, wherein the user interface information causes a user interface to:].
Para 0078, In various examples, and using Python as an example programming environment, systems (such as the model-instance-and-dataset-evaluation system 200) in accordance with embodiments of the present disclosure can be integrated with other Python modules such as 'matplotlib' and 'pytorch visualization' to further explore the content of a given automated-driving dataset. In many instances, this type of dataset exploration is highly beneficial to sensing model developers and sensing-model evaluators. This sort of exploration can often help such developers and evaluators determine data-manipulation needs, such as corrections in data labels, among many other examples that could be listed here. The following input and output illustrates an example for retrieving camera information with annotations and plotting it with matplotlib. An example output is displayed in FIG. 7 [enabling a user to assess ML model accuracy;].
para 0085, The ensuing portion of the present disclosure provides an example of using a system such as the model instance-and-dataset-evaluation system 200 to selectively work with annotations, including retrieving the content of the annotations, and also including obtaining the transform for one of the ground-truth labels from a sensor point of view using what is referred to herein as a "transform( ) function" (not to be confused with the above-introduced "get_transform( )" function). An example of this is illustrated in the following sequence of two example-input-and example output pairs:
Para 0090, The above input-and-output pair demonstrate an example of retrieving the content of a given example annotation. Below is an example of using transform() to obtaining the transform for one of the ground-truth labels from a sensor point of view. The example sensor in this case is LIDAR sensor 304.
Para 0104, FIG. 9 depicts an example visualization output 900, in accordance with at least one embodiment. In particular, the visualization output 900 includes a left-hand frame from the perspective of a "top camera" oriented vertically above the relevant area, and also includes a right-hand frame from the perspective of a different available camera, showing a different perspective of the same moment in time. It can be seen that a given annotation can visually appear on a user interface in different coordinates in a given 2D plane (i.e., the depicted images) while still corresponding to a common 3D, real-world location [and generating user interface information based on one or more of the][ or obtained output].
Para 0110, Additionally, another set of functions that are provided by systems in accordance with some embodiments of the present disclosure are referred to here as dataset-label exploration functions and dataset-label-synchronization functions. Thus, in addition to customization of labels as described above, embodiments provide users with capabilities of exploration and statistical analysis of existing labels in a given dataset. In some embodiments, one or more mapping functions are also provided to enable users to integrate label differences across datasets. This can be quite helpful in order to harmonize labels across datasets for purposes such as cross-domain training of a given machine learning- model instance. Following are several examples of dataset-label-exploration-and-synchronization functions, one or more of which may be provided in various embodiments [visually present the model-estimated location information and the ground truth location of the object proximate to the end-user vehicle,]);
Arroyo does not explicitly teach visually present the model-estimated location information and the ground truth location of the object proximate to the end-user vehicle, enabling a user to assess ML model accuracy;
and dynamically adjust, based on user feedback, an individual visual presentation of the model-estimated location information or the ground truth location based on the location-error value for the object, thereby providing feedback for validating ML model performance.
However Crego teaches determining, for one or more of the objects, a location-error value by comparing the model- estimated location information indicated by the output for the object to a ground truth location of the object values associated with metrics based on the obtained output (Crego Col 5 line 43-52, In some examples, a model training system may train the model in operation 102 using ground truth data to provide precise measurements of erroneous or inaccurate data output by the perception system. For instance, the model training system may provide library of labeled images to the perception system for analysis, and may compare the outputs of the perception system (e.g., the attributes of the objects data to the known object attributes from the labeled images, to 50 determine the errors within the outputs of the perception system.
Col 6 line 8-11, Thus, in various implementations the model training system may use ground truth data (e.g., labeled images), log data, or a combination of ground truth and log data, to train the model with perception error data.
Col 8 line 23-35, At operation 108, an autonomous vehicle (or simulation system) may use the perception error probability distribution to determine an object presence boundary for the object detected in operation 104. The object presence boundary ( or contour line) may be represented as an imaginary line generated by the autonomous vehicle that surrounds and/or encompasses an object detected in the environment. The precise location and shape of a contour line may correspond to the probability that the object inhabits the physical space at the boundary. For instance, an example image 130, shown in association with operation 108, depicts a top-down representation of a vehicle 122 and two separate contour lines132-134 surrounding the vehicle 122.
Col 11 line 40-59
In some implementations, the model training system 202 may train machine-learned models to output perception error probability distributions based on ground truth data, which may include a repository of labeled images 214. In such examples, the labeled images 214 may include scenes of environments having various objects that may be detected and identified by a perception system 210. Each of the labeled images 214 may include metadata associated with the labeled image identifying each object and various object attributes ( e.g., classification, size, position, distance away from the vehicle, etc.). The model training system 202 may input the labeled images 214 into the perception system 210, and may calculate the differences between the object attributes output by the perception system 210 and the corresponding object attributes in the ground truth metadata associated with the labeled images 214, to determine the perception error associated with the attribute of each object within each labeled image. The individual perception errors output by the perception system 210 may be analyzed to determine a perception error distribution [determining, for one or more of the objects, a location-error value by comparing the model- estimated location information indicated by the output for the object for the object to a ground truth location of the object].
Col 12 line 3-15, In some instances, the model training system 202 may use trend analyses, outlier detection, and the like, to determine which of the object attributes in the log data 216 are likely to be accurate, and which are likely to be inaccurate/erroneous. Based on these analyses, the model training system 202 may identify a number of individual perception errors within the log data 216, and may use the individual perception errors to determine a perception error distribution. In various implementations, to train machine-learned models to output perception error probability distributions, the model training system 202 may use ground truth data (e.g., labeled images 214), log data 216, and/or a combination of ground truth and log data [values associated with metrics based on the obtained output]);
Arroyo and Crego are considered to be analogous to the claim invention because they are in the same field of machine learning for autonomous vehicles. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to have modified Arroyo in view of Crego to incorporate determining perception error of objects. Doing so to improve estimates of confidence associated with underlying perception of object and allowing planning risk based decision and route selection. (Crego col 5 line 1-7, Additionally, mixture models may be trained based on combinations of perception error and prediction error, which provide the planning system improved estimates of the confidence associated with the underlying perceptions of objects and agent predictions, thereby allowing the planning system to make improved risk-based decisions and route selections.).
Gallo teaches and dynamically adjust, based on user feedback, an individual visual presentation of the model-estimated location information or the ground truth location based on the location-error value for the object, thereby providing feedback for validating ML model performance (Gallo para 0091, Once the model has been generated/trained, step 822 may be performed to actually use the model to determine the orientation of the symbols in a drawing. In step 834, since the orientation of the symbols are usually aligned with wall directions, there are limited directions for the symbols—e.g., four (4) directions (left, right, up, down) or more detailed 360-degree directions. Thus, it is enough to use a classification model (at step 834) as well as the object orientation model (at step 836) for symbol orientation prediction. For example, the nearest wall of the detected symbols can also be queried in the floor plan drawing and the direction of the wall can be used to further validate the predicted orientation. Accordingly, at step 822, the orientation of the object symbol instances is determined based on the ML model trained in step 830.
Gallo para 0093, At step 838, BIM symbol elements may be filtered. Such a filtering may be based on domain knowledge and/or errors/low confidence in detection/orientation output. For example, domain knowledge may be used to determine that if a bathtub is located in a kitchen, or a duplex/light switch is not attached to a wall, it does not make sense. Accordingly, at step 838, such errors or low confidence predictions may be filtered out. In this regard, a confidence threshold level may be defined and used to determine the level of confidence/accuracy that is tolerable (e.g., a user may adjust the threshold level or it may be predefined within the system). Of note, is that filtering step 838 may be performed at multiple locations (e.g., at the current location in the flow of FIG. 8 or after further steps are performed [based on the location-error value for the object, thereby]).
para 0094, User Interaction to Provide Feedback and Refine the Object Detection Model
Para 0095, After the orientation is determined (i.e., at step 822) and symbol element filtering is performed at step 838, user interaction may be provided to obtained feedback and refine the object detection model at step 840. In this regard, the floor plan drawing with the placed BIM elements may be presented to the user. For example, the filtered and detected symbols may be presented to the user with different colors for different levels of confidence. User feedback is then received. In addition, users may adjust the confidence threshold to filter out some detection. Based on the user feedback at step 840, labels, orientation or other information may be corrected as necessary at step 842. In this regard, users may also fine tune the bounding boxes of the detected symbols and correct the wrong labels of some symbols. User feedback may also update the symbol orientation. The user's feedback (i.e., the updated symbol labels 808 and updated symbol orientations 848 can be treated as a ground-truth and can be used to retrain object detection model (i.e., at step 804) and the symbol orientation classification model (i.e., at step 830) [and dynamically adjust, based on user feedback, an individual visual presentation of the model-estimated location information] [, thereby providing feedback for validating ML model performance].).
Arroyo and Gallo are considered to be analogous to the claim invention because they are in the same field of machine learning with user feedback. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to have modified Arroyo in view of Gallo to incorporate determining perception error of objects. Doing so to improve the object detection model by giving feedback to refine the presentation (Gallo para. 0095, After the orientation is determined (i.e., at step 822) and symbol element filtering is performed at step 838, user interaction may be provided to obtained feedback and refine the object detection model at step 840. In this regard, the floor plan drawing with the placed BIM elements may be presented to the user. For example, the filtered and detected symbols may be presented to the user with different colors for different levels of confidence. User feedback is then received. In addition, users may adjust the confidence threshold to filter out some detection. Based on the user feedback at step 840, labels, orientation or other information may be corrected as necessary at step 842. In this regard, users may also fine tune the bounding boxes of the detected symbols and correct the wrong labels of some symbols. User feedback may also update the symbol orientation. The user's feedback (i.e., the updated symbol labels 808 and updated symbol orientations 848 can be treated as a ground-truth and can be used to retrain object detection model (i.e., at step 804) and the symbol orientation classification model (i.e., at step 830)).
Regarding claim 3 and analogous 11, Arroyo in view of Crego and Gallo teach the method according to claim 1.
Arroyo, Crego and Gallo are combined with the same rationale used in claim 1 and analogous claims 9 and 17.
Arroyo teaches wherein the user interface information includes a graphical representation of objects detected via the ML model (Arroyo Para 0095, Moreover, in embodiments of the present disclosure, the model-instance-and-dataset-evaluation system 200 enables users to provide training labels in the frame of reference of the sensing algorithms in either in 2D or 3D. The following portion of this disclosure illustrates an example of extracting a 3D object (box) (identified by, as an example, the camera 312) and a 2D object (Rect) targeted to the local coordinate frame of reference of the image 328 associated with the camera 308.
Example Input
Para 0100, >>rect=geo.Rect.from_annotation(annotation, data_frame.transform_tree, 'IMAGE_328')
para 0101, >>print(rect)
Example Output
Para 0102, Rect (frame=IMAGE_328, label=None, instance_ id=l 79, xc=591.46, yc=-68.85, w=12.83, h=7.66, orientation=0.00)
Para 0103, A visualization of the resulting 2D/3D labels is displayed in the example visualization output 800 of FIG. 8. In that visualization are 2D and 3D label annotations rendered after applied to sensor input from the perspective of image 328, which is the projection onto the 2D image plane of the 3D location of the camera 308. Certainly many other examples could be provided as well [includes a graphical representation of objects detected via the ML model].).
Claim(s) 2, and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Arroyo in view of Crego and further in view of Gallo and Zaremba et al. (US11919545B2) (“Zaremba”).
Regarding claim 2 and analogous 10, Arroyo in view of Crego and Gallo teach the method according to claim 1.
Arroyo, Crego and Gallo are combined with the same rationale used in claim 1 and analogous claims 9 and 17.
Zaremba further teaches wherein the validation data further includes one or more of velocities of the end-user vehicle (Zaremba Col 7 line 63-67 and col 8 line 1-11, The traffic scenario system 116 extracts videos/video frames representing different traffic scenarios from videos for training and validating ML models. The system determines different types of scenarios that are extracted for training/validation of models so as to provide a comprehensive coverage of various types of traffic scenarios. The system receives a filter based on various attributes including (1) vehicle attributes such as speed, turn direction, and so on and (2) traffic attributes describing behavior of traffic entities, for example, whether a pedestrian has intent to cross the street, and (3) road attributes, for example, whether there is an intersection or a cross walk coming up. The system applies the filter to video frames to identify sets of video frames representing different scenarios. The video frames classified according to various traffic scenarios are used for ML model validation or training [wherein the validation data further includes one or more of velocities of the end-user vehicle].).
Arroyo and Zaremba are considered to be analogous to the claim invention because they are in the same field of machine learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to have modified Arroyo in view of Zaremba to incorporate speed and speed and direction of end-user vehicle. Doing so to determine different scenarios and fill out them out for validation or training (Zaremba Col 7 line 63-67 and col 8 line 1-11, The traffic scenario system 116 extracts videos/video frames representing different traffic scenarios from videos for training and validating ML models. The system determines different types of scenarios that are extracted for training/validation of models so as to provide a comprehensive coverage of various types of traffic scenarios. The system receives a filter based on various attributes including (1) vehicle attributes such as speed, turn direction, and so on and (2) traffic attributes describing behavior of traffic entities, for example, whether a pedestrian has intent to cross the street, and (3) road attributes, for example, whether there is an intersection or a cross walk coming up. The system applies the filter to video frames to identify sets of video frames representing different scenarios. The video frames classified according to various traffic scenarios are used for ML model validation or training).
Claim(s) 4 analogous 12 are rejected under 35 U.S.C. 103 as being unpatentable over Arroyo in view of Crego and further in view of Gallo and Wang et al. (US 20170185868 A1) (“Wang”).
Regarding claim 4 and analogous 12, Arroyo in view of Crego and Gallo teach the method according to claim 3.
Arroyo, Crego and Gallo are combined with the same rationale used in claim 1 and analogous claims 9 and 17.
Arroyo does not explicitly teach wherein the user interface information further includes error information associated with the objects.
However Wang teaches wherein the user interface information further includes error information associated with the objects (Wang para 0054 line 7-13, The screenshot 710 displays an LRP image 714 in which there is a lack of illumination. The user may view this image 714 to verify the classification of the fault. In addition, the erroneous registration identifier 716 and the suggested corrected registration identifier 718 may be displayed to allow the user to verify or choose the correct number [associated with the objects].
Para 0057, FIG. 9 shows an exemplary user interface screenshot 902. The user interface includes an upper panel 904, a left panel 906, a central panel 908 and a right panel 910. Upper panel 904 enables the user to pick a particular day ( or date) to inspect. Left panel 906 displays a list of sensors ordered by the error rate calculated by the present framework. The user may select one of the displayed sensors for further inspection. Central panel 908 presents an array of erroneous LPR records and suggested corrections determined by the present framework of the selected sensor. The user may select one of the displayed records for inspection and verification. Right panel 910 displays lane distributions 912a-c. Lane distribution 912a represents proportions of un-recognized, falsely-recognized and other types of errors; lane distribution 912b represents proportions of unrecognized errors; and lane distribution 912c represents proportions of falsely-recognized errors [user interface information further includes error information].).
Arroyo and Wang are considered to be analogous to the claim invention because they are in the same field of machine learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to have modified Arroyo in view of Wang to include generating error information. Doing so to allow the user to determine the correct identification (Wang para 0054 line 7-13, The screenshot 710 displays an LRP image 714 in which there is a lack of illumination. The user may view this image 714 to verify the classification of the fault. In addition, the erroneous registration identifier 716 and the suggested corrected registration identifier 718 may be displayed to allow the user to verify or choose the correct number.).
Claim(s) 5, 7, 8, 13, 15, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Arroyo in view of Crego and further in view of Gallo and Yang (US20220092349A1) (“Yang”).
Regarding claim 5 and analogous claims 13 and 18, Arroyo in view of Crego and Gallo teach the method according to claim 1.
Arroyo, Crego and Gallo are combined with the same rationale used in claim 1 and analogous claims 9 and 17.
Arroyo does not explicitly teach wherein the user interface: presents a graphical depiction of the end-user vehicle (Yang
PNG
media_image2.png
716
1164
media_image2.png
Greyscale
[presents a graphical depiction of the end-user vehicle]
Para 0080, Referring now to FIG . 3 , FIG . 3 illustrates an example of a display 300 of data indicating results of evaluating differences in performance between testing data sets , in accordance with some embodiments of the present disclosure . The display 300 may correspond to a user interface that presents a view of other representation of one or more samples from the training data set 122 , the testing data set 124A , and / or the testing data set 124B , along other corresponding information which may include a presentation 302 of metadata , at least some of which may correspond to the evaluation data 126. For example , the presentation 302 includes a presentation 304 indicating a reliance score corresponding to a depiction 306 of one or more correlated samples . In at least one embodiment , the presentation 302 may be updated to correspond to the depiction 306 in a display region of the interface as the depiction 306 is changed to correspond to one or more other correlated samples ( e.g. , using video playback and / or frame or time based user selection ) . In other examples , the presentation 304 may be without a corresponding depiction and / or may be displayed using graphs, charts, and / or other forms of presentation.).
Arroyo and Yang are considered to be analogous to the claim invention because they are in the same field of machine learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to have modified Arroyo in view of Yang to include user interface to display end-user vehicle. Doing so to allow the user display the data and evaluate difference in performance between testing dataset and include corresponding database on user selection. (Yang para 0010, FIG . 3 illustrates an example of a display of data indicating results of evaluating differences in performance between testing data sets , in accordance with some embodiments of the present disclosure.
Para 0080, Referring now to FIG . 3 , FIG . 3 illustrates an example of a display 300 of data indicating results of evaluating differences in performance between testing data sets , in accordance with some embodiments of the present disclosure . The display 300 may correspond to a user interface that presents a view of other representation of one or more samples from the training data set 122 , the testing data set 124A , and / or the testing data set 124B , along other corresponding information which may include a presentation 302 of metadata , at least some of which may correspond to the evaluation data 126. For example , the presentation corresponding to a depiction 306 of one or more correlated samples . In at least one embodiment , the presentation 302 display region of the interface as the depiction 306 is changed to correspond to one or more other correlated samples ( e.g. , using video playback and / or frame or time based user selection ) . In other examples , the presentation 304 may be without a corresponding depiction and / or may be displayed using graphs , charts , and / or other forms of presentation .).
Regarding claim 7 and analogous claims 15 and 20, Arroyo in view of Crego and Gallo teach the method according to claim 1.
Arroyo, Crego and Gallo are combined with the same rationale used in claim 1 and analogous claims 9 and 17.
Arroyo and Yang are combined with the same rationale used in claim 5 and analogous claims 13 and 18.
Yang teaches wherein the user interface visually presents a particular video sequence and wherein the ground truth locations of the objects are dynamically updated based on the video sequence (Yang Para 0080, Referring now to FIG . 3 , FIG . 3 illustrates an example of a display 300 of data indicating results of evaluating differences in performance between testing data sets , in accordance with some embodiments of the present disclosure . The display 300 may correspond to a user interface that presents a view of other representation of one or more samples from the training data set 122 , the testing data set 124A , and / or the testing data set 124B , along other corresponding information which may include a presentation 302 of metadata , at least some of which may correspond to the evaluation data 126. For example , the presentation 302 includes a presentation 304 indicating a reliance score corresponding to a depiction 306 of one or more correlated samples . In at least one embodiment , the presentation 302 may be updated to correspond to the depiction 306 in a display region of the interface as the depiction 306 is changed to correspond to one or more other correlated samples ( e.g. , using video playback and / or frame or time based user selection ) . In other examples , the presentation 304 may be without a corresponding depiction and / or may be displayed using graphs, charts, and / or other forms of presentation [wherein the user interface visually presents a particular video sequence].
para 0081, By way of example , the depiction 306 includes an indicator 310 of a human trajectory , an indicator 312 of ground truth lane center , and an indicator 314 of a predicted lane center or predicted trajectory made using the MLM 118 ( e.g. , overlaid on an image representing one or more corresponding samples ) . The depiction 306 also includes indicators 320 of vehicle tire locations and indicators 322 of lane boundaries . A depiction 340 provides another example of how one or more of such indicators may be presented using a top down view corresponding to the one or more correlated samples [and wherein the ground truth locations of the objects are dynamically updated based on the video]).
Regarding claim 8 and analogous claims 16 and 21, Arroyo in view of Crego and Gallo teach the method according to claim 1.
Arroyo, Crego and Gallo are combined with the same rationale used in claim 1 and analogous claims 9 and 17.
Arroyo and Yang are combined with the same rationale used in claim 5 and analogous claims 13 and 18.
Yang teaches wherein the individual visual presentations of the ground truth locations are dynamically updated based on the video sequence (Yang Para 0080, Referring now to FIG . 3 , FIG . 3 illustrates an example of a display 300 of data indicating results of evaluating differences in performance between testing data sets , in accordance with some embodiments of the present disclosure . The display 300 may correspond to a user interface that presents a view of other representation of one or more samples from the training data set 122 , the testing data set 124A , and / or the testing data set 124B , along other corresponding information which may include a presentation 302 of metadata , at least some of which may correspond to the evaluation data 126. For example , the presentation 302 includes a presentation 304 indicating a reliance score corresponding to a depiction 306 of one or more correlated samples . In at least one embodiment , the presentation 302 may be updated to correspond to the depiction 306 in a display region of the interface as the depiction 306 is changed to correspond to one or more other correlated samples ( e.g. , using video playback and / or frame or time based user selection ) . In other examples , the presentation 304 may be without a corresponding depiction and / or may be displayed using graphs, charts, and / or other forms of presentation.
para 0081, By way of example , the depiction 306 includes an indicator 310 of a human trajectory , an indicator 312 of ground truth lane center , and an indicator 314 of a predicted lane center or predicted trajectory made using the MLM 118 ( e.g. , overlaid on an image representing one or more corre sponding samples ) . The depiction 306 also includes indicators 320 of vehicle tire locations and indicators 322 of lane boundaries . A depiction 340 provides another example of how one or more of such indicators may be presented using a top down view corresponding to the one or more correlated samples [wherein the individual visual presentations of the ground truth locations are dynamically updated based on the video sequence]).
Claim(s) 6, 14, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Arroyo in view of Crego and further in view of Gallo and Z. Huang, W. Li, X. -G. Xia and R. Tao, "A General Gaussian Heatmap Label Assignment for Arbitrary-Oriented Object Detection," in IEEE Transactions on Image Processing, vol. 31, pp. 1895-1910, 2022, doi: 10.1109/TIP.2022.3148874 (“Huang”).
Regarding claim 6 and analogous claims 14 and 19, Arroyo in view of Crego and Gallo teach the method according to claim 1.
Arroyo, Crego and Gallo are combined with the same rationale used in claim 1 and analogous claims 9 and 17.
Arroyo does not teach explicitly wherein each adjusted visual presentation includes a color whose radius is selected based on the location-error value for the object.
Huang wherein each adjusted visual presentation includes a color whose radius is selected based on the location-error value for the object (Huang page 1899 Fig. 5,
PNG
media_image3.png
298
732
media_image3.png
Greyscale
Page 1899 A. OLA Strategy para 7, Third, the Spatial and Scale Extents of the Candidate Regions Using the Above Strategy Need to be Carefully Studied: First, a bounding box centered at the Gaussian peak location (called C-BBox) is computed based on the assigned labels. Then, it is assumed that many bounding boxes of different sizes centered at the other Gaussian candidate locations are generated. At a location, if there exists a bounding box whose Intersection over Union (IoU) with the C-BBox is greater than the threshold TIoU, this location is selected as a positive location. As shown in Fig. 5, these positive locations form a subset of the original Gaussian candidate locations (appearing as a smaller ellipse that is co-centered with the original Gaussian ellipse),
Page 1902, Therefore, the classification sub-task is affected by the OBB regression error. In the training process, in order to obtain a higher classification accuracy, the model parameters will be jointly adjusted to approach the optimal results of not only the classification sub-task but also the OBB regression task [whose radius is selected based on the error].
Page 1905,
PNG
media_image4.png
350
728
media_image4.png
Greyscale
Page 1905, Third, based on the baseline, i.e., Vanilla-AF, the proposed OLA and ORC are used to make the positive candidate region conform to the shape and direction characteristics of the objects. This improvement makes the mAP increase by 2.56. The object candidates of ORC are further improved by OWAM, i.e., ORC-OWAM, and mAP is further improved by 0.54. For the non-Gaussian center prior objects analyzed previously, like the harbor (HA), the performance improves more. The visualized feature maps of the CNN output layer in Fig. 9 verify this claim. Further using the proposed JOL, the mAP increases by 1.52 [wherein each adjusted visual presentation includes a color]).
Arroyo and Huang are considered to be analogous to the claim invention because they are in the same field of machine learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to have modified Arroyo in view of Huang to properly represent an object shape and direction on a heatmap. Doing so to fit the Gaussian center prior to fit the characteristics of different object thought neural network learning (Huang Abstract, Recently, many arbitrary-oriented object detection (AOOD) methods have been proposed and attracted widespread attention in many fields. However, most of them are based on anchor-boxes or standard Gaussian heatmaps. Such label assignment strategy may not only fail to reflect the shape and direction characteristics of arbitrary-oriented objects, but also have high parameter-tuning efforts. In this paper, a novel AOOD method called General Gaussian Heatmap Label Assignment (GGHL) is proposed. Specifically, an anchor-free object-adaptation label assignment (OLA) strategy is presented to define the positive candidates based on two-dimensional (2D) oriented Gaussian heatmaps, which reflect the shape and direction features of arbitrary-oriented objects. Based on OLA, an oriented-boundingbox (OBB) representation component (ORC) is developed to indicate OBBs and adjust the Gaussian center prior weights to fit the characteristics of different objects adaptively through neural network learning. Moreover, a joint-optimization loss (JOL) with area normalization and dynamic confidence weighting is designed to refine the misalign optimal results of different subtasks. Extensive experiments on public datasets demonstrate that the proposed GGHL improves the AOOD performance with low parameter-tuning and time costs. Furthermore, it is generally applicable to most AOOD methods to improve their performance including lightweight models on embedded platforms.).
Pertinent Prior Art
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Nix et al. (US20180365913A1) teaches a user interface that allows the user to provide feedback to the system of autonomous vehicle.
Sambo et al. (US11403851B2) teaches generating birds eye view to include vehicles and displaying the results.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALFREDO CAMPOS whose telephone number is (571)272-4504. The examiner can normally be reached 7:00 - 4:00 pm M - F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J. Huntley can be reached at (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALFREDO CAMPOS/Examiner, Art Unit 2129
/IMAD KASSIM/Primary Examiner, Art Unit 2129