Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Drawings
The drawings are objected to under 37 CFR 1.83(a) because the flowchart boxes in Fig. 4 lack descriptive text and therefore do not illustrate the steps recited in the claims. The drawings must show every feature of the invention specified in the claims. Therefore, the descriptive text must be shown in each of the steps of the flowchart boxes. No new matter should be entered.
Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
In addition, new corrected drawings in compliance with 37 CFR 1.121(d) are required in this application because the lines and/or characters and/or objects in the drawings (see Figs 2-3, for example) are too small, unclear, or faint and do not meet the requirements of 37 CFR 1.84. Applicant is advised to employ the services of a competent patent draftsperson outside the Office, as the U.S. Patent and Trademark Office no longer prepares new drawings. The corrected drawings are required in reply to the Office action to avoid abandonment of the application. The requirement for corrected drawings will not be held in abeyance.
Allowable Subject Matter
Claim 3 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-2 and 4-9 are rejected under 35 U.S.C. 103 as being unpatentable over Yang et al. (US 2021/0272304) in view of Garimella et al. (US 2021/0004611).
Regarding claim 1, xxx teaches a method for predicting a state of an environment of a vehicle, comprising: determining an occupancy grid of the environment of the vehicle for a current state of the environment of the vehicle; (see Yang at the Abstract, for example, which discloses that in various examples, a deep neural network (DNN) is trained to accurately predict, in deployment, distances to objects and obstacles using image data alone; see Yang at [0007] which further discloses systems and methods are disclosed that accurately and robustly predict distances to objects or obstacles in an environment using a deep neural network (DNN) trained with sensor data; see Yang at [0037] which discloses that although the present disclosure may be described with respect to an example autonomous vehicle 1400 (alternatively referred to herein as "vehicle 1400", "ego-vehicle 1400", or "autonomous vehicle 1400," an example of which is described with respect to FIGS. 14A-14D, this is not intended to be limiting; see Yang at [0135] which discloses that: One or more of the controller(s) 1436 may receive inputs (e.g., represented by input data) from an instrument cluster 1432 of the vehicle 1400 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 1434, an audible annunciator, a loudspeaker, and/or via other components of the vehicle 1400. The outputs may include information such as vehicle velocity, speed, time, map data (e.g., the HD map 1422 of FIG. 14C), location data (e.g., the vehicle’s 1400 location, such as on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and status of objects as perceived by the controller(s) 1436, etc.)
determining a digital map for the current state of the environment of the vehicle; (see Yang at [0135], for example, which discloses that: One or more of the controller(s) 1436 may receive inputs (e.g., represented by input data) from an instrument cluster 1432 of the vehicle 1400 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 1434, an audible annunciator, a loudspeaker, and/or via other components of the vehicle 1400. The outputs may include information such as vehicle velocity, speed, time, map data (e.g., the HD map 1422 of FIG. 14C), location data (e.g., the vehicle’s 1400 location, such as on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and status of objects as perceived by the controller(s) 1436, etc.)
determining a list of objects present in the environment of the vehicle in the current state of the environment of the vehicle; (see Yang at [0135], for example, which discloses that: One or more of the controller(s) 1436 may receive inputs (e.g., represented by input data) from an instrument cluster 1432 of the vehicle 1400 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 1434, an audible annunciator, a loudspeaker, and/or via other components of the vehicle 1400. The outputs may include information such as vehicle velocity, speed, time, map data (e.g., the HD map 1422 of FIG. 14C), location data (e.g., the vehicle’s 1400 location, such as on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and status of objects as perceived by the controller(s) 1436, etc. Also, see Yang at [0205] which discloses that the LIDAR sensor(s) 1464 may be capable of providing a list of objects and their distances for a 360-degree field of view.)
encoding the occupancy grid to a first occupancy grid representation in a latent space for the occupancy grid; (see Yang at the Abstract, for example, which discloses that the DNN (deep neural network) may be trained with ground truth data that is generated and encoded using sensor data from any number of depth predicting sensors, such as, without limitation, RADAR sensors, LIDAR sensors, and/or SONAR sensors. Also, see Yang at [0009] which discloses that the ground truth data encoding pipeline may use sensor data from depth sensor(s) to—automatically, without manual annotation, in embodiments—encode ground truth data corresponding to training image data in order to train the DNN to make accurate predictions from image data alone. Absent a specific definition of the term “latent space” in the specification, the Examiner has interpreted this term based on a broadest reasonable interpretation. Examiner notes that “latent space” corresponds to a mathematical vector space used in the representation of data, as used in machine learning, such as in a neural network.)
encoding the digital map to a first map representation in a latent space for the digital map; (see Yang at [0058] which discloses that once a final distance value(s) has been selected for an object 306, one or more pixels of the image 302 may be encoded with the final depth value(s) to generate the ground truth depth map 222 and that the ground truth depth map 222 may represent the ground truth distance(s) encoded onto an image—e.g., a depth map image.)
encoding the list of objects to a first object list representation in a latent space for the list of objects; and (see Yang at [0043] which discloses that: As a non-limiting embodiment, to generate the ground truth data for training the machine learning model(s) 104, ground truth encoding 110 may be performed according to the process for ground truth encoding 110 of FIG. 2. For example, the sensor data 102—such as image data representative of one or more images—may be used by an object detector 214 to detect objects and/or obstacles represented by the image data. Examiner notes that using ground truth encoding by an object detector to detect objects and/or obstacles corresponds to encoding the list of objects to a first object list representation in a latent space for the list of objects. Examiner shows a teaching based on a broadest reasonable interpretation of the claimed language.)
predicting, for each of one or more points in time of future states of the environment of the vehicle, (i) a respective further occupancy grid representation in the latent space for the occupancy grid, (see Yang at [0072] which discloses that the machine learning model(s) 104 may further be trained to intrinsically predict the locations of bounding shapes as the object detection(s) 116; see Yang at [0105] which discloses object detections and depth predictions based on outputs of machine learning model(s); see Yang at [0144] which discloses that cameras with a field of view that include portions of the environment to the side of the vehicle 1400 (e.g., side-view cameras) may be used for surround view, providing information used to create and update the occupancy grid, as well as to generate side impact collision warnings. Examiner notes that predicting the locations of bounding shapes during object detection and updating the occupancy grid corresponds to predicting, for each of one or more points in time of future states of the environment of the vehicle, (i) a respective further occupancy grid representation in the latent space for the occupancy grid.)
(ii) a respective further map representation in the latent space for the digital map, (see Yang at [0093] which discloses that the method 600, at block B608, includes generating fifth data representative of ground truth information, the fifth data generated based at least in part on converting the depth information to a depth map. For example, the depth information corresponding to the bounding shapes may be used to generate the ground truth depth map 222 (FIG. 2); see Yang at [0094] which discloses that the method 600, at block B610, includes training a neural network to compute the predicted depth map using the fifth data. Examiner notes that generating fifth data representative of ground truth information and converting the depth information into a depth map by way of training using a neural network corresponds to predicting (ii) a respective further map representation in the latent space for the digital map.)
and (iii) a respective further object list representation in the latent space for the list of objects; (see Yang at [0059] which discloses updating and optimizing the machine learning model(s) 104 for predicting locations of objects and/or obstacles; see Yang at [0204] which discloses that in some examples, the LIDAR sensor(s) 1464 may be capable of providing a list of objects and their distances for a 360-degree field of view. Also, see Yang at [0205] which discloses that the LIDAR sensor(s) 1464 may be capable of providing a list of objects and their distances for a 360-degree field of view. Also, see Yang at [0233] which discloses that the deep-learning infrastructure may receive periodic updates from the vehicle 1400, such as a sequence of images and/or objects that the vehicle 1400 has located in that sequence of images (e.g., via computer vision and/or other machine learning object classification techniques.).
Yang does not expressly disclose wherein, for each of the one or more points in time, the respective further occupancy grid representation is predicted prior to predicting the respective further map representation and predicting the respective further object list representation and the respective further occupancy grid representation is used for predicting the respective further map representation and predicting the respective further object list representation which in a related art Garimella discloses (see Garimella, at the Abstract, which discloses that techniques for determining predictions on a top-down representation of an environment based on vehicle action(s) are discussed herein and that a multi-channel image representing a top-down view of the object(s) and the environment can be generated based on the sensor data, map data, and/or action data. Garimella, at the Abstract, further discloses that multiple images can be generated representing the environment over time and input into a prediction system configured to output prediction probabilities associated with possible locations of the object(s) in the future, which may be based on the actions of the autonomous vehicle. Also, see Garimella at [0016] which discloses that in some examples, an intensity of the heat map may represent a probability that a cell or pixel will be occupied by any object at the specified instance in time (e.g., an occupancy grid). Examiner notes that the multi-channel image representing a top-down view of the object(s) and the environment generated based on sensor data, map data, and/or action data corresponds to occupancy grid representation, map representation, and object list representation of an environment of the vehicle. Further, see Garimella at [0025] which discloses that the machine learning model can include a convolutional neural network (CNN), which may include one or more recurrent neural network (RNN) layers, such as, but not limited to, long short-term memory (LSTM) layers. See Garimella at [0105-0107] which discloses occupancy and unoccupancy map of objects in the future. Examiner notes that the occupancy and occupancy map of objects in the future correspond to the prediction of further occupancy grid representation for the occupancy grid and map representation for the digital map, as well as an object list representation for the list of objects. See Grimaldi at [0149] which discloses that the order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and/or in parallel to implement the processes. Examiner notes that being able to implement any of the prediction operations in any order related to the locations of objects and determining the occupancy and unoccupancy map of objects in future corresponds to the recited features of the foregoing clause. Examiner further notes that the use of one or more recurrent neural network (RNN) layers, such as, but not limited to, long short-term memory (LSTM) layers corresponds to using a predicted further occupancy grid representation for predicting the respective further map representation and predicting the respective further object list representation.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Yang to wherein, for each of the one or more points in time, the respective further occupancy grid representation is predicted prior to predicting the respective further map representation and predicting the respective further object list representation and the respective further occupancy grid representation is used for predicting the respective further map representation and predicting the respective further object list representation, as taught by Garimella.
One would have been motivated to make such a modification to determine how a particular entity is likely to behave in the future, as suggested by Garimella at [0001].
Regarding claim 2, the modified Yang teaches the method according to claim 1, wherein, for each of the one or more points in time, the prediction of the respective further map representation occurs prior to the prediction of the respective further object list representation and is used to predict the respective further object list representation (see Garimella at [0025] which discloses that the machine learning model can include a convolutional neural network (CNN), which may include one or more recurrent neural network (RNN) layers, such as, but not limited to, long short-term memory (LSTM) layers. Also, see Grimaldi at [0149] which discloses that the order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and/or in parallel to implement the processes. Examiner notes that the use of one or more recurrent neural network (RNN) layers, such as, but not limited to, long short-term memory (LSTM) layers corresponds to using a predicted respective further map representation for predicting the respective further object list representation.)
Regarding claim 4, the modified Yang teaches the method according to claim 1, further comprising: planning a behavior of the vehicle for each of the one or more points in time by determining a respective behavior representation in a latent space for the behavior using the respective further occupancy grid representation, the respective further map representation, and the respective further object list representation predicted for the point in time (see Garimella at [0014] which discloses that the prediction probabilities can be generated or determined based on particular candidate actions, and the prediction probabilities can be evaluated to select or determine a candidate action to control the autonomous vehicle; see Garimella at [0018] which discloses that candidate actions can be evaluated to determine a risk, cost, and/or reward associated with the candidate action, and a candidate action can be selected or determined based at least in part on evaluating the candidate actions and that the autonomous vehicle can be controlled based at least in part on a selected or determined candidate action. Examiner maps controlling the autonomous vehicle based at least in part on a selected or determined candidate action to the planning a behavior of the vehicle. Examiner maps candidate action to respective behavior.)
Regarding claim 5, the modified Yang teaches the method according to claim 1, further comprising:
predicting the respective further occupancy grid representation using a neural occupancy grid predictive network; predicting the respective further map representation using a neural map predictive network; predicting the respective further object list representation using a neural object list predictive network; (see Yang at [0135] which discloses that the outputs may include information such as vehicle velocity, speed, time, map data (e.g., the HD map 1422 of FIG. 14C), location data (e.g., the vehicle's 1400 location, such as on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and status of objects as perceived by the controller(s) 1436, etc.; see Yang at [0094] in conjunction with Fig. 6 which discloses that the method 600, at block B610, includes training a neural network to compute the predicted depth map using the fifth data and that for example, the machine learning model(s) 104 may be trained, using the ground truth depth map 222 as ground truth, to generate the distance(s) 106 corresponding to objects and/or obstacles depicted in images. Examiner notes that a predicted depth map may correspond to predicting the respective further map representation using a neural map predictive network. Further, see Yang at [0138] which discloses that prediction probabilities 802 refer to a series of eight frames (labeled t1 -t8) illustrating an output of the prediction component 228, whereby the prediction probabilities 802 are not based in part on action data, and that in a first frame of the prediction probabilities 802 (illustrated as frame ti), the scenario represents a vehicle 806 and objects 808, 810, and 812. Examiner notes that the depicted objects 808, 810, and 812 may correspond to the predicted respective further object list representation.)
and training the neural occupancy grid predictive network, the neural map predictive network, and the neural object list predictive network by: determining occupancy grid costs by decoding the respective further occupancy grid representation to a corresponding respective further occupancy grid and comparing it to ground truth information for the occupancy grid for the respective point in time, (see Garimella at [0076] which discloses that the evaluation component 238 can determine whether a trajectory for the vehicle 202 traverses through regions associated with prediction probabilities (which may include dilated prediction probabilities) associated with the vehicle 202, that the evaluation component 238 can determine costs, risks, and/or rewards at individual time steps in the future and/or cumulatively for some or all time steps associated with a candidate action, and that accordingly, the evaluation component 238 can compare costs, risks, and/or rewards for different candidate actions and can select an action for controlling the vehicle; see Garimella at [0105] which discloses that the output logits from the machine learned component 232 can be compared against training data 258 (e.g., ground truth representing an occupancy map) using a sigmoid cross entropy loss.)
and/or by encoding the ground truth information for the occupancy grid for the respective point in time to an occupancy grid ground truth and comparing it with the respective further occupancy grid representation; determining map costs by decoding the respective further map representation to a corresponding respective further digital map and comparing it to a ground truth information for the digital map for the respective point in time, and/or by encoding the ground truth information for the digital map for the respective point in time to a map ground truth and comparing it to the further map representation; (Examiner notes that Applicant has used the phrase “and/or” in the instant claim. The Patent Trial and Appeal Board (PTAB) has held that use of the phrase “and/or” within a claim is not indefinite. According to the PTAB, “and/or” is not wrong, but it’s not preferred verbiage (see Ex Parte Gross, Appeal No. 2011-004811). Nevertheless, during patent examination, the pending claims must be given their broadest reasonable interpretation (BRI) consistent with the specification (see MPEP § 2111; Phillips v. AWH Corp., 415 F.3d 1303, 1316, 75 USPQ2d 1321, 1329 (Fed. Cir. 2005)). Based upon this guidance from the MPEP and the Federal Circuit Court of Appeals, the Examiner interprets the phrase “and/or” under its broadest reasonable interpretation of “or” for purposes of examination of the instant Application.)
and/or determining object list costs by decoding the respective further object list representation to a corresponding respective further list of objects for the respective point in time and comparing it with a ground truth information for the list of objects for the respective point in time and/or by encoding the ground truth information for the list of objects for the respective point in time to an object list ground truth and comparing it with the further object list representation (Examiner notes that Applicant has used the phrase “and/or” in the instant claim. The Patent Trial and Appeal Board (PTAB) has held that use of the phrase “and/or” within a claim is not indefinite. According to the PTAB, “and/or” is not wrong, but it’s not preferred verbiage (see Ex Parte Gross, Appeal No. 2011-004811). Nevertheless, during patent examination, the pending claims must be given their broadest reasonable interpretation (BRI) consistent with the specification (see MPEP § 2111; Phillips v. AWH Corp., 415 F.3d 1303, 1316, 75 USPQ2d 1321, 1329 (Fed. Cir. 2005)). Based upon this guidance from the MPEP and the Federal Circuit Court of Appeals, the Examiner interprets the phrase “and/or” under its broadest reasonable interpretation of “or” for purposes of examination of the instant Application.)
Regarding claim 6, the modified Yang teaches the method according to claim 1, wherein a computer program includes instructions that, when executed by a processor, cause the processor to carry out the method (see Yang at [0089] which discloses that each block of method 600, described herein, comprises a computing process that may be performed using any combination of hardware, firmware, and/or software, that for instance, various functions may be carried out by a processor executing instructions stored in memory, that the method 600 may also be embodied as computer-usable instructions stored on computer storage media, and that the method 600 may be provided by a standalone application, a service or hosted service (standalone or in combination with another hosted service), or a plug-in to another product, to name a few. Examiner notes that computer-usable instructions maps to computer program.)
Regarding claim 7, the modified Yang teaches a method for controlling a vehicle, comprising: predicting a state of an environment of the vehicle according to the method of claim 1; and controlling the vehicle based on the predicted state of the environment (see Garimella at [0001] which discloses that prediction techniques can be used to determine future states of entities in an environment and that prediction techniques can be used to determine how a particular entity is likely to behave in the future; see Yang at [0094], for example, which discloses training a neural network to compute the predicted depth map; see Yang at [0134], for example, which discloses that the controller(s) 1436 may provide the signals for controlling one or more components and/or systems of the vehicle 1400 in response to sensor data received from one or more sensors (e.g., sensor inputs); see Garimella at [0013] which discloses controlling a vehicle based on the prediction probabilities, in accordance with examples of the disclosure; see Garimella at [0054] which discloses that the process 100 can include evaluating the candidate action and/or controlling the autonomous vehicle 106 based at least in part on the candidate actions.)
Regarding claim 8, the modified Yang teaches a vehicle control device configured to perform the method according to claim 1 (see Yang at [0089] for example which discloses that each block of method 600, described herein, comprises a computing process that may be performed using any combination of hardware, firmware, and/or software. Examiner maps hardware to vehicle control device.)
Regarding claim 9, the modified Yang teaches a non-transitory computer-readable medium that stores instructions that, when executed by a processor, cause the processor to carry out the method according to claim 1 (see Garimella at [0097] which discloses that the processor(s) 216 of the vehicle 202 and the processor(s) 244 of the computing device(s) 242 can be any suitable processor capable of executing instructions to process data and perform operations as described herein; see Garimella at [0098] which discloses that memory 218 and 246 are examples of non-transitory computer-readable media and that the memory 218 and 246 can store an operating system and one or more software applications, instructions, programs, and/or data to implement the methods described herein and the functions attributed to the various systems.)
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ROY RHEE whose telephone number is 313-446-6593. The examiner can normally be reached M-F 8:30 am to 5:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, Applicant may contact the Examiner via telephone or use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kito Robinson, can be reached on 571-270-3921. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, one may visit: https://patentcenter.uspto.gov. In addition, more information about Patent Center may be found at https://www.uspto.gov/patents/apply/patent-center. Should you have questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ROY RHEE/Primary Examiner, Art Unit 3664