DETAILED ACTION
In view of the Appeal Brief filed on 28 May 2026, PROSECUTION IS HEREBY REOPENED. New grounds of rejection are set forth below.
To avoid abandonment of the application, appellant must exercise one of the following two options:
(1) file a reply under 37 CFR 1.111 (if this Office action is non-final) or a reply under 37 CFR 1.113 (if this Office action is final); or,
(2) initiate a new appeal by filing a notice of appeal under 37 CFR 41.31 followed by an appeal brief under 37 CFR 41.37. The previously paid notice of appeal fee and appeal brief fee can be applied to the new appeal. If, however, the appeal fees set forth in 37 CFR 41.20 have been increased since they were previously paid, then appellant must pay the difference between the increased fees and the amount previously paid.
A Supervisory Patent Examiner (SPE) has approved of reopening prosecution by signing below:
{ 4 }
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
Claims 1-20 are pending in this application.
Claims 1-20 are presented for examination.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 10-14, 16, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Austin et al. (US Publication 2021/0394793 A1) in view of Tusch et al. (US Publication 2021/0279475 A1).
Regarding claim 1, Austin teaches a method performed by one or more computers, the method comprising: obtaining sensor data (i) that is captured by multiple sensors of an autonomous vehicle (Austin: Para. 38, 53, 63; the autonomous vehicle continuously takes at least one of camera, LiDAR and radar images of the surrounding environment to monitor for road users; images may be timestamped by an image processor and analyzed for changes in motion, head pose and body posture in order to determine a gaze direction of the road user; computer vision techniques may be applied to the image data to identify road users, such as pedestrians, bicyclists and non-autonomous vehicles) and (ii) that characterizes an agent that is in a vicinity of the autonomous vehicle in an environment at a current time point, wherein the sensor data comprises at least an image patch from a camera image captured by a camera sensor and at least a portion of a point cloud captured by a laser sensor (Austin: Para. 38, 53, 63; the autonomous vehicle continuously takes at least one of camera, LiDAR and radar images of the surrounding environment to monitor for road users; images may be timestamped by an image processor and analyzed for changes in motion, head pose and body posture in order to determine a gaze direction of the road user; computer vision techniques may be applied to the image data to identify road users, such as pedestrians, bicyclists and non-autonomous vehicles); and processing the sensor data comprising at least the image patch from the camera image captured by the camera and at least the portion of the point cloud captured by the laser sensor using a gaze prediction neural network to generate a gaze prediction that predicts a gaze of the agent at the current time point (Austin: Para. 42, 66; determine the position, body posture and head pose of the road user in order to determine the gaze direction of the road user; determine the gaze direction, the computer system may use the cameras images, LiDAR data; gaze recognition examples through machine learning based techniques for example, convolutional neural networks), wherein the gaze prediction neural network comprises (Austin: Para. 66; gaze recognition examples through machine learning based techniques for example, convolutional neural networks).
Austin doesn’t explicitly teach an embedding subnetwork that is configured to process at least the image patch from the camera image captured by the camera sensor and at least the portion of the point cloud captured by the laser sensor to generate an embedding characterizing the agent; and a gaze subnetwork that is configured to process the embedding generated by processing at least the image patch from the camera image captured by the camera sensor and at least the portion of the point cloud captured by the laser sensor to generate the gaze prediction.
However Austin in view of Tusch teaches an embedding subnetwork that is configured to process at least the image patch from the camera image captured by the camera sensor and at least the portion of the point cloud captured by the laser sensor to generate an embedding characterizing the agent; and a gaze subnetwork that is configured to process the embedding generated by processing at least the image patch from the camera image captured by the camera sensor and at least the portion of the point cloud captured by the laser sensor to generate the gaze prediction.
Austin teaches computer vision techniques applied to image data identify pedestrians (Austin: Para. 63). Austin teaches feature and gaze recognition through machine learning based techniques such as convolution neural networks (Austin: Para. 66). Neural network computer vision often relies on embedding to identify a person in an image. Tusch teaches a computer vision engine using data from multiple sensors to track an object moving through the environment. The digital representation includes feature vectors that define the appearance of a generalized person (Tusch: Para. 161, 164). This feature vector is used to analyze the person’s trajectory, pose, or predicting the person’s intent (Tusch: Para. 165). Embedding is used in tools like computer vision to recognize features in sensor data. Tusch teaches computer vision of multiple sensors where digital representation includes feature vectors. Austin teaches computer vision applied to image data to identify features. It would be obvious to one of ordinary skill in the art to know convolutional neural network machine learning computer vision techniques would routinely use embedding where the digital representation includes feature vectors.
Austin teaches an autonomous vehicle identifies a road user by an image recorded by vehicle camera (Austin: Para. 38). The vehicle takes the camera image and uses computer vision and object recognition to identify road users, such as a pedestrian (Austin: Para. 63). The image patch is just the relevant part of the camera image. It is well known in the art of computer vision to identify a pedestrian from a camera image taken by an autonomous vehicle when the pedestrian is only in a portion of the image. Tusch teaches a computer vision engine using data from multiple sensors to track an object moving through the environment. The digital representation includes feature vectors that define the appearance of a generalized person (Tusch: Para. 161, 164). This feature vector is used to analyze the person’s trajectory, pose, or predicting the person’s intent (Tusch: Para. 165). Embedding is used in tools like computer vision to recognize features in sensor data. Tusch teaches computer vision of multiple sensors where digital representation includes feature vectors. Austin teaches a LIDAR system, that sends thousands of laser pulses every second, to create a 3D point cloud that detect pedestrian candidate regions via a bounded point cloud that are used to determine the position, body posture, and head pose of the road user in order to determine the gaze direction of the road user (Austin: Para. 41-42). Austin teaches a computer system, using machine learning based technique of convolutional neural networks, performing gaze detection by matching gaze patterns to the detected facial area, locating the person’s eyes, and center of pupil, to determine the gaze direction (Austin: Para. 66). Therefore Austin in view of Tusch teaches uses a embedding subnetwork with camera and LIDAR data in order to determine the gaze direction of the road user.
It would have been obvious to one having ordinary skill in the art to modify the road user’s gaze direction determination through camera image analysis and LIDAR point cloud analysis (Austin: Para. 43) with the embedding vector in a computer vision system (Tusch: Para. 25) with a reasonable expectation of success because a computer vision systems creating vectors of define key points of image data provides real time analytics on detected people (Tusch: Para. 1, 25).
Regarding claim 10, Austin doesn’t explicitly teach wherein the embedding subnetwork is configured to: process at least the image patch from the camera image captured by the camera sensor to generate a respective initial camera embedding; process at least the portion of the point cloud captured by the laser sensor to generate a respective initial point cloud embedding; and combine the respective initial embeddings to generate the embedding characterizing the agent, wherein the combining the respective initial embeddings comprises summing, averaging, or concatenating the respective initial embeddings.
However Austin in view of Tusch teaches wherein the embedding subnetwork is configured to: process at least the image patch from the camera image captured by the camera sensor to generate a respective initial camera embedding; process at least the portion of the point cloud captured by the laser sensor to generate a respective initial point cloud embedding; and combine the respective initial embeddings to generate the embedding characterizing the agent, wherein the combining the respective initial embeddings comprises summing, averaging, or concatenating the respective initial embeddings.
Tusch teaches a computer vision engine using data from multiple sensors to track an object moving through the environment. The digital representation includes feature vectors that define the appearance of a generalized person (Tusch: Para. 161, 164). This feature vector is used to analyze the person’s trajectory, pose, or predicting the person’s intent (Tusch: Para. 165). Embedding is used in tools like computer vision to recognize features in sensor data. Austin teaches computer vision techniques applied to image data identify pedestrians (Austin: Para. 63). Austin teaches a correlation module is configured to concatenate the views from each autonomous vehicle, correlating these views with the trajectories of each autonomous vehicle and each road user (Austin: Para. 97). It would be obvious to one of ordinary skill in the art to know convolutional neural network machine learning computer vision techniques would routinely use embedding where the digital representation includes feature vectors.
It would have been obvious to one having ordinary skill in the art to modify the road user’s gaze direction determination through camera image analysis and LIDAR point cloud analysis (Austin: Para. 43) with the embedding vector in a computer vision system (Tusch: Para. 25) with a reasonable expectation of success because a computer vision systems creating vectors of define key points of image data provides real time analytics on detected people (Tusch: Para. 1, 25).
Regarding claim 11, Austin teaches the method of claim 10, wherein the sensor data comprises an image patch depicting the agent generated from an image of the environment captured by the camera sensor and a portion of a point cloud generated by the laser sensor (Austin: Para. 43; computer of the vehicle is configured to use data gathered by camera image analysis, LiDAR 3D point cloud analysis and radar and/or millimeter wave radar images to determine the gaze direction of the road user; Both LiDAR and camera recognition processes can be performed based on trained and/or predefined libraries of data).
Regarding claim 12, Austin teaches the method of claim 10, wherein the gaze prediction neural network has been trained on one or more auxiliary tasks (Austin: Para. 66; gaze recognition examples through machine learning based techniques, for example, convolutional neural networks).
Austin doesn’t explicitly teach wherein the one or more auxiliary tasks include one or more auxiliary tasks that measure respective initial gaze predictions made directly from each of the initial embeddings.
However Tusch, in the same field of endeavor, teaches wherein the one or more auxiliary tasks include one or more auxiliary tasks that measure respective initial gaze predictions made directly from each of the initial embeddings (Tusch: Para. 122-127, 131; edge layer includes a computer-vision system or engine that (a) generates from a pixel stream a digital representation of a person or other object; edge layer can infer or describe a person's behaviour or intent by analysing one or more of the trajectory, pose, gesture, identity of that person).
It would have been obvious to one having ordinary skill in the art to modify the road user’s gaze direction determination through camera image analysis and LIDAR point cloud analysis (Austin: Para. 43) with the embedding vector in a computer vision system (Tusch: Para. 25) with a reasonable expectation of success because a computer vision systems creating vectors of define key points of image data provides real time analytics on detected people (Tusch: Para. 1, 25).
Regarding claim 13, Austin teaches the method of claim 1, wherein the gaze prediction neural network has been trained on one or more auxiliary tasks (Austin: Para. 66; gaze recognition examples through machine learning based techniques, for example, convolutional neural networks; once a person's eyes are located, a sufficiently powerful camera may track the center of the pupil to detect gaze direction).
Regarding claim 14, Austin teaches the method of claim 13, wherein the one or more auxiliary tasks include a heading prediction task (Austin: Para. 82; an environment mapping module incorporates a global model and then focuses on determining the gaze direction, wherein the trajectory module predicts the future path of the road user).
Regarding claim 16, Austin teaches a system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising: obtaining sensor data (i) that is captured by multiple sensors of an autonomous vehicle (Austin: Para. 43; computer of the vehicle is configured to use data gathered by camera image analysis, LiDAR 3D point cloud analysis) and (ii) that characterizes an agent that is in a vicinity of the autonomous vehicle in an environment at a current time point, wherein the sensor data comprises at least an image patch from a camera image captured by a camera and a laser sensor (Austin: Para. 38, 53, 63; the autonomous vehicle continuously takes at least one of camera, LiDAR and radar images of the surrounding environment to monitor for road users; images may be timestamped by an image processor and analyzed for changes in motion, head pose and body posture in order to determine a gaze direction of the road user; computer vision techniques may be applied to the image data to identify road users, such as pedestrians, bicyclists and non-autonomous vehicles); and processing the sensor data comprising at least the image patch from the camera image captured by the camera sensor and at least the portion of the point cloud captured by the laser sensor using a gaze prediction neural network to generate a gaze prediction that predicts a gaze of the agent at the current time point (Austin: Para. 42, 66; determine the position, body posture and head pose of the road user in order to determine the gaze direction of the road user; determine the gaze direction, the computer system may use the cameras images, LiDAR data; gaze recognition examples through machine learning based techniques for example, convolutional neural networks), wherein the gaze prediction neural network comprises (Austin: Para. 66; gaze recognition examples through machine learning based techniques for example, convolutional neural networks).
Austin doesn’t explicitly teach an embedding subnetwork that is configured to process at least the image patch from the camera image captured by the camera sensor and at least the portion of the point cloud captured by the laser sensor to generate an embedding characterizing the agent; and a gaze subnetwork that is configured to process the embedding generated by processing at least the image patch from the camera image captured by the camera sensor and at least the portion of the point cloud and captured by the laser sensor to generate the gaze prediction.
However Austin in view of Tusch teaches an embedding subnetwork that is configured to process at least the image patch from the camera image captured by the camera sensor and at least the portion of the point cloud captured by the laser sensor to generate an embedding characterizing the agent; and a gaze subnetwork that is configured to process the embedding generated by processing at least the image patch from the camera image captured by the camera sensor and at least the portion of the point cloud and captured by the laser sensor to generate the gaze prediction.
Austin teaches computer vision techniques applied to image data identify pedestrians (Austin: Para. 63). Austin teaches feature and gaze recognition through machine learning based techniques such as convolution neural networks (Austin: Para. 66). Neural network computer vision often relies on embedding to identify a person in an image. Tusch teaches a computer vision engine using data from multiple sensors to track an object moving through the environment. The digital representation includes feature vectors that define the appearance of a generalized person (Tusch: Para. 161, 164). This feature vector is used to analyze the person’s trajectory, pose, or predicting the person’s intent (Tusch: Para. 165). Embedding is used in tools like computer vision to recognize features in sensor data. Tusch teaches computer vision of multiple sensors where digital representation includes feature vectors. Austin teaches computer vision applied to image data to identify features. It would be obvious to one of ordinary skill in the art to know convolutional neural network machine learning computer vision techniques would routinely use embedding where the digital representation includes feature vectors.
Austin teaches an autonomous vehicle identifies a road user by an image recorded by vehicle camera (Austin: Para. 38). The vehicle takes the camera image and uses computer vision and object recognition to identify road users, such as a pedestrian (Austin: Para. 63). The image patch is just the relevant part of the camera image. It is well known in the art of computer vision to identify a pedestrian from a camera image taken by an autonomous vehicle when the pedestrian is only in a portion of the image. Tusch teaches a computer vision engine using data from multiple sensors to track an object moving through the environment. The digital representation includes feature vectors that define the appearance of a generalized person (Tusch: Para. 161, 164). This feature vector is used to analyze the person’s trajectory, pose, or predicting the person’s intent (Tusch: Para. 165). Embedding is used in tools like computer vision to recognize features in sensor data. Tusch teaches computer vision of multiple sensors where digital representation includes feature vectors. Austin teaches a LIDAR system, that sends thousands of laser pulses every second, to create a 3D point cloud that detect pedestrian candidate regions via a bounded point cloud that are used to determine the position, body posture, and head pose of the road user in order to determine the gaze direction of the road user (Austin: Para. 41-42). Austin teaches a computer system, using machine learning based technique of convolutional neural networks, performing gaze detection by matching gaze patterns to the detected facial area, locating the person’s eyes, and center of pupil, to determine the gaze direction (Austin: Para. 66). Therefore Austin in view of Tusch teaches uses a embedding subnetwork with camera and LIDAR data in order to determine the gaze direction of the road user.
It would have been obvious to one having ordinary skill in the art to modify the road user’s gaze direction determination through camera image analysis and LIDAR point cloud analysis (Austin: Para. 43) with the embedding vector in a computer vision system (Tusch: Para. 25) with a reasonable expectation of success because a computer vision systems creating vectors of define key points of image data provides real time analytics on detected people (Tusch: Para. 1, 25).
Regarding claim 20, Austin teaches one or more non-transitory computer storage media encoded with computer program instructions that when executed by a plurality of computers cause the plurality of computers to perform operations comprising: obtaining sensor data (i) that is captured by multiple sensors of an autonomous vehicle (Austin: Para. 43; computer of the vehicle is configured to use data gathered by camera image analysis, LiDAR 3D point cloud analysis) and (ii) that characterizes an agent that is in a vicinity of the autonomous vehicle in an environment at a current time point, wherein the sensor data comprises at least an image patch from a camera image captured by a camera sensor and at least a portion of a point cloud captured by a laser sensor (Austin: Para. 38, 53, 63; the autonomous vehicle continuously takes at least one of camera, LiDAR and radar images of the surrounding environment to monitor for road users; images may be timestamped by an image processor and analyzed for changes in motion, head pose and body posture in order to determine a gaze direction of the road user; computer vision techniques may be applied to the image data to identify road users, such as pedestrians, bicyclists and non-autonomous vehicles); and processing the sensor data comprising at least the image patch from the camera image captured by the camera sensor and at least the portion of the point cloud captured by the laser sensor using a gaze prediction neural network to generate a gaze prediction that predicts a gaze of the agent at the current time point (Austin: Para. 42, 66; determine the position, body posture and head pose of the road user in order to determine the gaze direction of the road user; determine the gaze direction, the computer system may use the cameras images, LiDAR data; gaze recognition examples through machine learning based techniques for example, convolutional neural networks), wherein the gaze prediction neural network comprises (Austin: Para. 66; gaze recognition examples through machine learning based techniques for example, convolutional neural networks).
Austin doesn’t explicitly teach an embedding subnetwork that is configured to process at least the image patch from the camera image captured by the camera sensor and at least the portion of the point cloud captured by the laser sensor to generate an embedding characterizing the agent; and a gaze subnetwork that is configured to process the embedding generated by processing at least the image patch from the camera image captured by the camera sensor and at least the portion of the point cloud captured by the laser sensor to generate the gaze prediction.
However Austin in view of Tusch teaches an embedding subnetwork that is configured to process at least the image patch from the camera image captured by the camera sensor and at least the portion of the point cloud captured by the laser sensor to generate an embedding characterizing the agent; and a gaze subnetwork that is configured to process the embedding generated by processing at least the image patch from the camera image captured by the camera sensor and at least the portion of the point cloud captured by the laser sensor to generate the gaze prediction.
Austin teaches computer vision techniques applied to image data identify pedestrians (Austin: Para. 63). Austin teaches feature and gaze recognition through machine learning based techniques such as convolution neural networks (Austin: Para. 66). Neural network computer vision often relies on embedding to identify a person in an image. Tusch teaches a computer vision engine using data from multiple sensors to track an object moving through the environment. The digital representation includes feature vectors that define the appearance of a generalized person (Tusch: Para. 161, 164). This feature vector is used to analyze the person’s trajectory, pose, or predicting the person’s intent (Tusch: Para. 165). Embedding is used in tools like computer vision to recognize features in sensor data. Tusch teaches computer vision of multiple sensors where digital representation includes feature vectors. Austin teaches computer vision applied to image data to identify features. It would be obvious to one of ordinary skill in the art to know convolutional neural network machine learning computer vision techniques would routinely use embedding where the digital representation includes feature vectors.
Austin teaches an autonomous vehicle identifies a road user by an image recorded by vehicle camera (Austin: Para. 38). The vehicle takes the camera image and uses computer vision and object recognition to identify road users, such as a pedestrian (Austin: Para. 63). The image patch is just the relevant part of the camera image. It is well known in the art of computer vision to identify a pedestrian from a camera image taken by an autonomous vehicle when the pedestrian is only in a portion of the image. Tusch teaches a computer vision engine using data from multiple sensors to track an object moving through the environment. The digital representation includes feature vectors that define the appearance of a generalized person (Tusch: Para. 161, 164). This feature vector is used to analyze the person’s trajectory, pose, or predicting the person’s intent (Tusch: Para. 165). Embedding is used in tools like computer vision to recognize features in sensor data. Tusch teaches computer vision of multiple sensors where digital representation includes feature vectors. Austin teaches a LIDAR system, that sends thousands of laser pulses every second, to create a 3D point cloud that detect pedestrian candidate regions via a bounded point cloud that are used to determine the position, body posture, and head pose of the road user in order to determine the gaze direction of the road user (Austin: Para. 41-42). Austin teaches a computer system, using machine learning based technique of convolutional neural networks, performing gaze detection by matching gaze patterns to the detected facial area, locating the person’s eyes, and center of pupil, to determine the gaze direction (Austin: Para. 66). Therefore Austin in view of Tusch teaches uses a embedding subnetwork with camera and LIDAR data in order to determine the gaze direction of the road user.
It would have been obvious to one having ordinary skill in the art to modify the road user’s gaze direction determination through camera image analysis and LIDAR point cloud analysis (Austin: Para. 43) with the embedding vector in a computer vision system (Tusch: Para. 25) with a reasonable expectation of success because a computer vision systems creating vectors of define key points of image data provides real time analytics on detected people (Tusch: Para. 1, 25).
Claims 2-9, 15 and 17-19 are rejected under 35 U.S.C. 103 as being unpatentable over Austin et al. (US Publication 2021/0394793 A1) in view of Tusch et al. (US Publication 2021/0279475 A1) and in further view of Benou et al. (US Publication 2022/0332349 A1).
Regarding claim 2, Austin teaches the method of claim 1, further comprising: determining, from the gaze prediction, an awareness signal that indicates whether the agent is aware of a presence of one or more entities in the environment (Austin: Para. 48; a controller of a vehicle computing system in order to provide a road user the eHMI notification in the field of view indicated by his/her gaze direction).
Austin and Tusch don’t explicitly teach using the awareness signal to determine a future trajectory of the autonomous vehicle after the current time point.
However Benou, in the same field of endeavor, teaches using the awareness signal to determine a future trajectory of the autonomous vehicle after the current time point (Benou: Para. 434; cause the host vehicle to perform at least one of: maintaining a current speed of the host vehicle or maintaining a current heading direction of the host vehicle).
It would have been obvious to one having ordinary skill in the art to modify the road user’s gaze direction determination through camera image analysis and LIDAR point cloud analysis (Austin: Para. 43) by adding the future autonomous vehicle action determination based on the direction of the pedestrian’s gaze (Benou: Para. 434) with a reasonable expectation of success because when the pedestrian is looking away from the host vehicle, the host vehicle may navigate by using a greater margin of safety than it otherwise might use if the pedestrian was looking at or in the direction of the host vehicle as taught by Benou (Benou: Para. 397).
It would have been obvious to one having ordinary skill in the art to modify the road user’s gaze direction determination through camera image analysis and LIDAR point cloud analysis (Austin: Para. 43) with the embedding vector in a computer vision system (Tusch: Para. 25) by adding the future autonomous vehicle action determination based on the direction of the pedestrian’s gaze (Benou: Para. 434) with a reasonable expectation of success because when the pedestrian is looking away from the host vehicle, the host vehicle may navigate by using a greater margin of safety than it otherwise might use if the pedestrian was looking at or in the direction of the host vehicle as taught by Benou (Benou: Para. 397).
Regarding claim 3, Austin teaches the method of claim 2, wherein the awareness signal indicates whether the agent is aware of a presence of the autonomous vehicle (Austin: Para. 48; a controller of a vehicle computing system in order to provide a road user the eHMI notification in the field of view indicated by his/her gaze direction).
Regarding claim 4, Austin teaches the method of claim 2, wherein the awareness signal indicates whether the agent is aware of a presence of one or more other agents in the environment (Austin: Para. 48; a controller of a vehicle computing system in order to provide a road user the eHMI notification in the field of view indicated by his/her gaze direction).
Regarding claim 5, Austin and Tusch don’t explicitly teach providing an input comprising the awareness signal to a machine learning model that is used by a planning system of the autonomous vehicle to plan the future trajectory of the autonomous vehicle.
However Benou, in the same field of endeavor, teaches providing an input comprising the awareness signal to a machine learning model that is used by a planning system of the autonomous vehicle to plan the future trajectory of the autonomous vehicle (Benou: Para. 434; the recognized gesture may be indicative of the pedestrian turning to look toward the host vehicle; cause the host vehicle to perform at least one of: maintaining a current speed of the host vehicle or maintaining a current heading direction of the host vehicle).
It would have been obvious to one having ordinary skill in the art to modify the road user’s gaze direction determination through camera image analysis and LIDAR point cloud analysis (Austin: Para. 43) with the embedding vector in a computer vision system (Tusch: Para. 25) by adding the future autonomous vehicle action determination based on the direction of the pedestrian’s gaze (Benou: Para. 434) with a reasonable expectation of success because when the pedestrian is looking away from the host vehicle, the host vehicle may navigate by using a greater margin of safety than it otherwise might use if the pedestrian was looking at or in the direction of the host vehicle as taught by Benou (Benou: Para. 397).
Regarding claim 6, Austin and Tusch don’t explicitly teach wherein the gaze prediction comprises a predicted gaze direction in a horizontal plane and a predicted gaze direction in a vertical axis.
However Benou, in the same field of endeavor, teaches wherein the gaze prediction comprises a predicted gaze direction in a horizontal plane and a predicted gaze direction in a vertical axis (Benou: Para. 397; looking direction of the pedestrian may be estimated based on the rotational angle and pitch angle; head pose may be represented by a rotational angle (also called yaw angle) in a horizontal plane; a pitch angle in a vertical plane).
It would have been obvious to one having ordinary skill in the art to modify the road user’s gaze direction determination through camera image analysis and LIDAR point cloud analysis (Austin: Para. 43) with the embedding vector in a computer vision system (Tusch: Para. 25) by adding the future autonomous vehicle action determination based on the direction of the pedestrian’s gaze (Benou: Para. 434) with a reasonable expectation of success because when the pedestrian is looking away from the host vehicle, the host vehicle may navigate by using a greater margin of safety than it otherwise might use if the pedestrian was looking at or in the direction of the host vehicle as taught by Benou (Benou: Para. 397).
Regarding claim 7, Austin and Tusch don’t explicitly teach wherein determining, from the gaze prediction, the awareness signal of the presence of the one or more entities in the environment comprises: determining that the predicted gaze direction in the vertical axis is horizontal; determining that the one or more entities is within a predetermined range centered at the predicted gaze direction in the horizontal plane.
However Benou, in the same field of endeavor, teaches wherein determining, from the gaze prediction, the awareness signal of the presence of the one or more entities in the environment comprises: determining that the predicted gaze direction in the vertical axis is horizontal (Benou: Para. 397; head pose may be represented by a rotational angle (also called yaw angle) in a horizontal plane parallel to the ground surface); determining that the predicted gaze direction in the vertical axis is horizontal; determining that the one or more entities is within a predetermined range centered at the predicted gaze direction in the horizontal plane (Benou: Para. 397; a pitch angle in a vertical plane that extends from the pedestrian's nose and the back of the head).
It would have been obvious to one having ordinary skill in the art to modify the road user’s gaze direction determination through camera image analysis and LIDAR point cloud analysis (Austin: Para. 43) with the embedding vector in a computer vision system (Tusch: Para. 25) by adding the future autonomous vehicle action determination based on the direction of the pedestrian’s gaze (Benou: Para. 434) with a reasonable expectation of success because when the pedestrian is looking away from the host vehicle, the host vehicle may navigate by using a greater margin of safety than it otherwise might use if the pedestrian was looking at or in the direction of the host vehicle as taught by Benou (Benou: Para. 397).
In the following limitation, Austin teaches in response, determining that the agent is aware of the presence of the one or more entities in the environment (Austin: Para. 48; a controller of a vehicle computing system in order to provide a road user the eHMI notification in the field of view indicated by his/her gaze direction).
Regarding claim 8, Austin teaches the method of claim 2, wherein the awareness signal comprises one or more of an active awareness signal and a historical awareness signal (Austin: Para. 73; environment mapping module may access a database of stored sets of images associated with poses, body posture, walking speeds, and the like, and may match each stitched image to a stored image to determine the gaze direction; predict the trajectory of the road user from the gaze direction), wherein the active awareness signal indicates whether the agent is aware of the presence of the one or more entities in the environment at the current time point (Austin: Para. 48; a controller of a vehicle computing system in order to provide a road user the eHMI notification in the field of view indicated by his/her gaze direction) wherein the historical awareness signal (i) is determined from one or more gaze predictions at one or more previous time points in a previous time window that precedes the current time point (Austin: Para. 73; environment mapping module may access a database of stored sets of images associated with poses, body posture, walking speeds, and the like, and may match each stitched image to a stored image to determine the gaze direction; predict the trajectory of the road user from the gaze direction) and (ii) indicates whether the agent is aware of the presence of the one or more entities in the environment during the previous time window (Austin: Para. 73; environment mapping module may access a database of stored sets of images associated with poses, body posture, walking speeds, and the like, and may match each stitched image to a stored image to determine the gaze direction; predict the trajectory of the road user from the gaze direction).
Regarding claim 9, Austin and Tusch don’t explicitly teach using both the gaze prediction and the awareness signal to determine a future trajectory of the autonomous vehicle after the current time point.
However Benou, in the same field of endeavor, teaches using both the gaze prediction and the awareness signal to determine a future trajectory of the autonomous vehicle after the current time point (Benou: Para. 434; the recognized gesture may be indicative of the pedestrian turning to look toward the host vehicle; cause the host vehicle to perform at least one of: maintaining a current speed of the host vehicle or maintaining a current heading direction of the host vehicle).
It would have been obvious to one having ordinary skill in the art to modify the road user’s gaze direction determination through camera image analysis and LIDAR point cloud analysis (Austin: Para. 43) with the embedding vector in a computer vision system (Tusch: Para. 25) by adding the future autonomous vehicle action determination based on the direction of the pedestrian’s gaze (Benou: Para. 434) with a reasonable expectation of success because when the pedestrian is looking away from the host vehicle, the host vehicle may navigate by using a greater margin of safety than it otherwise might use if the pedestrian was looking at or in the direction of the host vehicle as taught by Benou (Benou: Para. 397).
Regarding claim 15, Austin teaches the method of claim 1, wherein the gaze prediction neural network comprises a regression output layer and a classification output layer (Austin: Para. 43, 66; gaze recognition examples through machine learning based techniques, for example, convolutional neural networks; LiDAR and camera recognition processes can be performed based on trained and/or predefined libraries of data, with known and recognizable shapes and edges of obstacles (e.g. vehicles, cyclists, etc.)).
Austin and Tusch don’t explicitly teach wherein the regression output layer is configured to generate a predicted gaze direction in a horizontal plane and the classification output layer is configured to generate a predicted gaze direction in a vertical axis.
However Benou, in the same field of endeavor, teaches wherein the regression output layer is configured to generate a predicted gaze direction in a horizontal plane (Benou: Para. 397; head pose may be represented by a rotational angle (also called yaw angle) in a horizontal plane parallel to the ground surface) and the classification output layer is configured to generate a predicted gaze direction in a vertical axis (Benou: Para. 397; a pitch angle in a vertical plane that extends from the pedestrian's nose and the back of the head).
It would have been obvious to one having ordinary skill in the art to modify the road user’s gaze direction determination through camera image analysis and LIDAR point cloud analysis (Austin: Para. 43) with the embedding vector in a computer vision system (Tusch: Para. 25) by adding the future autonomous vehicle action determination based on the direction of the pedestrian’s gaze (Benou: Para. 434) with a reasonable expectation of success because when the pedestrian is looking away from the host vehicle, the host vehicle may navigate by using a greater margin of safety than it otherwise might use if the pedestrian was looking at or in the direction of the host vehicle as taught by Benou (Benou: Para. 397).
Regarding claim 17, Austin teaches the system of claim 16, the operations further comprise: determining, from the gaze prediction, an awareness signal that indicates whether the agent is aware of a presence of one or more entities in the environment (Austin: Para. 48; a controller of a vehicle computing system in order to provide a road user the eHMI notification in the field of view indicated by his/her gaze direction).
Austin and Tusch don’t explicitly teach using the awareness signal to determine a future trajectory of the autonomous vehicle after the current time point.
However Benou, in the same field of endeavor, teaches using the awareness signal to determine a future trajectory of the autonomous vehicle after the current time point (Benou: Para. 434; cause the host vehicle to perform at least one of: maintaining a current speed of the host vehicle or maintaining a current heading direction of the host vehicle).
It would have been obvious to one having ordinary skill in the art to modify the road user’s gaze direction determination through camera image analysis and LIDAR point cloud analysis (Austin: Para. 43) with the embedding vector in a computer vision system (Tusch: Para. 25) by adding the future autonomous vehicle action determination based on the direction of the pedestrian’s gaze (Benou: Para. 434) with a reasonable expectation of success because when the pedestrian is looking away from the host vehicle, the host vehicle may navigate by using a greater margin of safety than it otherwise might use if the pedestrian was looking at or in the direction of the host vehicle as taught by Benou (Benou: Para. 397).
Regarding claim 18, Austin teaches the system of claim 17, wherein the awareness signal indicates whether the agent is aware of a presence of the autonomous vehicle (Austin: Para. 48; a controller of a vehicle computing system in order to provide a road user the eHMI notification in the field of view indicated by his/her gaze direction).
Regarding claim 19, Austin teaches the system of claim 17, wherein the awareness signal indicates whether the agent is aware of a presence of one or more other agents in the environment (Austin: Para. 48; a controller of a vehicle computing system in order to provide a road user the eHMI notification in the field of view indicated by his/her gaze direction).
Response to Arguments
Applicant’s arguments with respect to claims 1-20 have been considered but are moot because the arguments do not apply to the references being used in the current rejection.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LAURA E LINHARDT whose telephone number is (571) 272-8325. The examiner can normally be reached on M-TR, M-F: 8am-4pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Angela Ortiz can be reached on (571) 272-1206. The fax phone number for the organization where this application or proceeding is assigned is (571) 273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at (866) 217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call (800) 786-9199 (IN USA OR CANADA) or (571) 272-1000.
/L.E.L./Examiner, Art Unit 3663
/ANGELA Y ORTIZ/Supervisory Patent Examiner, Art Unit 3663