DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
This Final Office Action is in response to the applicant’s amendment/response of 20 April 2026.
Claims 79-80, 82, 88, 90-91, 95-96, 99, 111-112, 114, 116, 119, 121, 123, 135, 137, and 139-
148 are currently pending and addressed below.
Response to Arguments
Applicant's arguments/amendments with respect to the rejection of claims under 35 U.S.C.
101 have been fully considered but they are not persuasive.
Specifically, applicant argued:
First, Applicant respectfully submits that the Office has improperly oversimplified the above-identified subject matter by ignoring the related wherein clauses. For example, amended claim 79 does not merely “detect a first semantic feature represented in the first image” but instead further defines the first semantic feature in reciting “wherein the first semantic feature is associated with a predetermined object type classification, wherein the first semantic features corresponds to a front side of an object along the road segment.” . . . Accordingly, Applicant respectfully submits that the human mind is incapable of either with or without a physical aid, detecting a semantic feature in an image, identifying a position descriptor associated with the semantic feature, or determining a position of one semantic feature relative to another semantic feature at least because semantic features, position descriptors, and position determined based on semantic features and position descriptors are specific data elements the human mind is ill-equipped to understand and process. . . That is the first semantic feature and the second semantic feature are both associated a common object along a road segment, each of which detected based on respective images captured by a same host vehicle at different points relative to the object. Applicant respectfully submits that the human mind is ill-equipped to perform these separate detections on a common object . . . Accordingly, amended independent claim 111 does not recite a mental process . . .
Firstly, Applicant respectfully submits that the independent claims implement the alleged judicial exception in conjunction with a particular machine that is integral to the claim . . . In the present application, as discussed above, Applicant’s specification explains how utilizing both front and rear side analyses of an object relative to a host vehicle traveling along a same trajectory improves object detection and feature association . . . Accordingly, the amended independent claim 79 recites significantly more than any alleged abstract idea under Step 2B . . .
The Examiner’s response:
Applicant asserts “First, Applicant respectfully submits that the Office has improperly oversimplified the above-identified subject matter by ignoring the related wherein clauses. For example, amended claim 79 does not merely “detect a first semantic feature represented in the first image” but instead further defines the first semantic feature in reciting “wherein the first semantic feature is associated with a predetermined object type classification, wherein the first semantic features corresponds to a front side of an object along the road segment.” . . . Accordingly, Applicant respectfully submits that the human mind is incapable of either with or without a physical aid, detecting a semantic feature in an image, identifying a position descriptor associated with the semantic feature, or determining a position of one semantic feature relative to another semantic feature at least because semantic features, position descriptors, and position determined based on semantic features and position descriptors are specific data elements the human mind is ill-equipped to understand and process. . . That is the first semantic feature and the second semantic feature are both associated a common object along a road segment, each of which detected based on respective images captured by a same host vehicle at different points relative to the object. Applicant respectfully submits that the human mind is ill-equipped to perform these separate detections on a common object . . . Accordingly, amended independent claim 111 does not recite a mental process . . .”. However, the Examiner respectfully disagrees. The Examiner submits the limitations “detect a first semantic feature . . .wherein the first semantic feature corresponds to a front side of an object along the road segment”, “identify, . . . , at least one position descriptor associated with the first semantic feature represented in the first image captured by the forward-facing camera”, “detect a second semantic feature represented in the second image, . . . , wherein the second semantic feature corresponds to a rear side of the object along the road segment . . . “, “identify, . . . , at least one position descriptor associated with the second semantic feature represented in the second image captured by the rearward-facing camera”, and “determine, . . . , at least one position descriptor associated with the second semantic feature, . . . “, under its broadest reasonable interpretation, can reasonably be performed by a human mentally or with aid of pen and paper. For example, a human mind could reasonably detect/label a feature from an image(s) with semantic description. Further, applicant asserts “ . . . the human mind is ill-equipped to perform these separate detections on a common object and that the Office has not demonstrated otherwise”. The Examiner respectfully disagrees because a human mind could reasonably determine/detect a common object/feature from different images (e.g. images captured by different cameras).
Further, applicant asserts “Firstly, Applicant respectfully submits that the independent claims implement the alleged judicial exception in conjunction with a particular machine that is integral to the claim . . . In the present application, as discussed above, Applicant’s specification explains how utilizing both front and rear side analyses of an object relative to a host vehicle traveling along a same trajectory improves object detection and feature association . . . Accordingly, the amended independent claim 79 recites significantly more than any alleged abstract idea under Step 2B . . .” however, the Examiner respectfully disagrees. The improvement in “object detection and feature association” is an improved abstract idea, and cannot constitute the additional element in the claim that might integrate the abstract idea (e.g. mental process) to a practical application. Accordingly, even in combination, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea .
Furthermore, applicant asserts “the claims do not merely specify the transmission and reception of data over a network but instead recite a specific coordination of drive information between a host vehicle, a first plurality of other vehicles, and a second plurality of other vehicles by an entity remotely-located relative to the host vehicle . . .” however, the Examiner respectfully disagrees because applicant arguments are not commensurate with the scope of the claim language that does not require any “coordination of drive information . . .”. Additionally, the limitation “align the drive information . . . “ is recited at a high-level of generality and/or directed to insignificant extra-solution activity, therefore, the claim is not indicative of an inventive concept and does not amount to significantly more than the abstract idea.
The Examiner notes that the rejection has been modified reflecting the amendments most recently submitted by the applicant.
Applicant’s arguments/amendments with respect to the rejection of claims under 35 U.S.C. 103 have been considered but are moot because the new ground of rejection does not rely on the combination of references applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 79-80, 82, 88, 90-91, 95-96, 99, 111-112, 114, 116, 119, 121, 123, 135, 137, and 139-148 are rejected under 35 U.S.C. 101
Regarding claim 79:
Step 1: Statutory Category - Yes
The claim is directed toward a system which falls within one of the four statutory categories. MPEP 2106.3.
Step 2A Prong 1: Judicial Exception – Yes
Independent claim 79 includes limitations that recites an abstract idea. The claim recites “detect a first semantic feature represented in the first image…wherein the first semantic feature corresponds to a front side of an object along the road segment”, “identify, …, at least one position descriptor associated with the first semantic feature represented in the first image captured by the forward-facing camera”, “detect a second semantic feature represented in the second image, wherein the second semantic feature is associated with a predetermined object type classification,... ”, “identify, …, at least one position descriptor associated with the second semantic feature represented in the second image captured by the rearward-facing camera”, and “determine, based on the at least one position descriptor associated with the first semantic feature, the at least one position descriptor associated with the second semantic feature, and the position information, a position the first semantic feature relative the second semantic feature”, which given their broadest reasonable interpretation, the claim covers performance of the limitations in the human mind. For example, “detect a first semantic feature…” and “detect a second semantic feature…” steps in the context of this claim encompasses a human analyzing one or more image(s) and interprets elements such as objects, actions/events, and the image’s context to understand what the image represents.
Step 2A Prong 2: Practical Application – No
Claim 79 is evaluated whether as a whole it integrates the recited judicial exception into a practical application. As noted in the 2019 PEG, it must be determined whether any additional elements in the claim beyond the abstract idea integrate the exception into a practical application in a manner that imposes a meaningful limit on the judicial exception. The courts have indicated that additional elements merely using a computer to implement an abstract idea, adding insignificant extra solution activity, or generally linking use of a judicial except ion to a particular technological environment or field of use do not integrate a judicial exception into a “practical application”.
The claim does not include additional elements that are sufficient enough to amount to integrating the judicial exception into a practical application, for example, the claimed elements “receive a first image captured by a forward-facing camera onboard the host vehicle…”, “receive a second image captured by a rearward-facing camera onboard the host vehicle…”, “receive position information indicative of a position of the forward-facing camera when the first image was captured and indicative of a position of the rearward-facing camera when the second image was captured;”, “cause transmission of drive information for the road segment to an entity remotely-located relative to the host vehicle…”, “receive, in addition to the drive information transmitted by the host vehicle, drive information from a first plurality of other vehicles…”, and “align the drive information received from the first plurality of other vehicles with the drive information received from the second plurality of other vehicles…” are recited at a high-level of generality and amounts to mere pre- or post- solution actions, which is a form of insignificant extra-solution activity. Claim 79 recites the additional elements of “forward-facing camera”, and “rearward-facing camera” are recited at a high-level of generality and merely links the abstract idea to a particular technological environment. Additionally, the “at least one processor”, “the entity remotely-located” and “trained machine learning model” are recited at a high-level of generality and amount to no more than mere instructions apply the exception using a generic computer. The claim is recited at a high-level of generality and merely automates the aforementioned step(s). The claim is directed to the abstract idea.
Accordingly, even in combination, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea .
Step 2B:
Claim 79 is evaluated as to whether the claim as a whole amount to significantly more
than the recited exception, i.e., whether any additional element, or combination of additional elements, adds an inventive concept to the claim.
The claim does not include additional elements that are sufficient enough to provide an
inventive concept in Step 2B, for example, the claimed elements “receive a first image captured by a forward-facing camera onboard the host vehicle…”, “receive a second image captured by a rearward-facing camera onboard the host vehicle…”, “receive position information indicative of a position of the forward-facing camera when the first image was captured and indicative of a position of the rearward-facing camera when the second image was captured;”, “cause transmission of drive information for the road segment to an entity remotely-located relative to the host vehicle…”, “receive, in addition to the drive information transmitted by the host vehicle, drive information from a first plurality of other vehicles…”, and “align the drive information received from the first plurality of other vehicles with the drive information received from the second plurality of other vehicles…” are well-understood, routine and conventional activity in the art. See MPEP 2106.05(d), II, “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information);”.
As discussed with respect to step 2A Prong 2, the additional elements of “forward-facing camera”, and “rearward-facing camera” are recited at a high-level of generality and merely links the abstract idea to a particular technological environment. Additionally, the “at least one processor”, “the entity remotely-located”, and “trained machine learning model” are recited at a high-level of generality and amount to no more than mere instructions apply the exception using a generic computer. The claim is recited at a high-level of generality and merely automates the aforementioned step(s). Accordingly, the claim is not patent eligible.
Regarding claims 99 and 135 , claim 99 recites a method and claim 135 recites a non-transitory computer-readable medium which fall within at least one of the four statutory categories. Claims 99 and 135 recite similar limitations as indicated above with respect to claim 79. Hence, the claim is not eligible for the same reasons as discussed above with respect to claim 79. All other limitations not discussed are the same as those discussed above with to claim 79. Discussion is omitted for brevity.
Claims 80, 82, 88, 90-91, 95-96, 139-140, 143-144, and 147-148 are also rejected under 35 U.S.C. 101 by virtue of their dependency to the independent claims.
Claims 80, 82, 88, 90-91, 95-96, 139-140, 143-144, and 147-148 do not recite additional elements that integrate the judicial exception into a practical application, because the additional elements are directed toward additional aspects of judicial exception and/or well-understood, routine and conventional additional elements that do not integrate the judicial exception into a practical application. For example, claim 82 recites “where in the X-Y-Z position is determined based on an one ego motion of the host vehicle and based on at least one of:…” further the abstract idea.
The dependent claims are rejected under 35 U.S.C. 101 under similar rationale as their independent claims.
Regarding claim 111:
Step 1: Statutory Category - Yes
The claim is directed toward a system which falls within one of the four statutory categories. MPEP 2106.3.
Step 2A Prong 1: Judicial Exception – Yes
Independent claim 111 includes limitations that recites an abstract idea. The claim recites “detect at least one object represented in the first image”, “identify, …, at least one front side two-dimensional feature point, the at least one front side two-dimensional feature point being associated with the at least one object represent in the first image”, “detect a representation of the at least one object in the second image, wherein the representation of the at least one object in the second image is detected after the host vehicle passes the at least one object along the road segment”, “identify, …, at least one rear side two-dimensional feature point, the at least one rear side two-dimensional feature point being associated with the at least one object represented in the second image”, and “determine, based on the at least one front side two-dimensional feature point, that at least one rear side two-dimensional feature point, and the position information, that the at least one front side two-dimensional feature point and the at least one rear side two-dimensional feature point are associated with a common object” which given their broadest reasonable interpretation, the claim covers performance of the limitations in the human mind. For example, “detect a first semantic feature…” and “detect a second semantic feature…” steps in the context of this claim encompasses a human analyzing one or more image(s) and interprets elements such as objects, actions/events, and the image’s context to understand what the image represents.
Step 2A Prong 2: Practical Application – No
Claim 111 is evaluated whether as a whole it integrates the recited judicial exception into a practical application. As noted in the 2019 PEG, it must be determined whether any additional elements in the claim beyond the abstract idea integrate the exception into a practical application in a manner that imposes a meaningful limit on the judicial exception. The courts have indicated that additional elements merely using a computer to implement an abstract idea, adding insignificant extra solution activity, or generally linking use of a judicial except ion to a particular technological environment or field of use do not integrate a judicial exception into a “practical application”.
The claim does not include additional elements that are sufficient enough to amount to integrating the judicial exception into a practical application, for example, the claimed elements “receive a first image captured by a forward-facing camera onboard the host vehicle…”, “receive a second image captured by a rearward-facing camera onboard the host vehicle…”, “receive position information indicative of a position of the forward-facing camera when the first image was captured and indicative of a position of the rearward-facing camera when the second image was captured;”, “cause transmission of drive information for the road segment to an entity remotely-located relative to the host vehicle…”,“receive, in addition to the drive information transmitted by the host vehicle, drive information from a first plurality of other vehicles that travel along the road segment in the first direction along with drive information from a second plurality of other vehicles…”, and “align the drive information received from the first plurality of other vehicles with the drive information received from the second plurality of other vehicles…” are recited at a high-level of generality and amounts to mere pre- or post- solution actions, which is a form of insignificant extra-solution activity. Claim 111 recites the additional elements of “forward-facing camera”, and “rearward-facing camera” are recited at a high-level of generality and merely links the abstract idea to a particular technological environment. Additionally, the “at least one processor”, “the entity remotely-located”, and “trained machine learning model” are recited at a high-level of generality and amount to no more than mere instructions apply the exception using a generic computer. The claim is recited at a high-level of generality and merely automates the aforementioned step(s).
Accordingly, even in combination, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea .
Step 2B:
Claim 111 is evaluated as to whether the claim as a whole amount to significantly more
than the recited exception, i.e., whether any additional element, or combination of additional elements, adds an inventive concept to the claim.
The claim does not include additional elements that are sufficient enough to provide an
inventive concept in Step 2B, for example, the claimed elements “receive a first image captured by a forward-facing camera onboard the host vehicle…”, “receive a second image captured by a rearward-facing camera onboard the host vehicle…”, “receive position information indicative of a position of the forward-facing camera when the first image was captured and indicative of a position of the rearward-facing camera when the second image was captured;”, “cause transmission of drive information for the road segment to an entity remotely-located relative to the host vehicle…”,“receive, in addition to the drive information transmitted by the host vehicle, drive information from a first plurality of other vehicles that travel along the road segment in the first direction along with drive information from a second plurality of other vehicles…”, and “align the drive information received from the first plurality of other vehicles with the drive information received from the second plurality of other vehicles…” are well-understood, routine and conventional activity in the art. See MPEP 2106.05(d), II, “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information);”.
As discussed with respect to step 2A Prong 2, additional elements of “forward-facing camera”, and “rearward-facing camera” are recited at a high-level of generality and merely links the abstract idea to a particular technological environment. Additionally, the “at least one processor”, “the entity remotely-located”, and “trained machine learning model” are recited at a high-level of generality and amount to no more than mere instructions apply the exception using a generic computer. The claim is recited at a high-level of generality and merely automates the aforementioned step(s). Accordingly, the claim is not patent eligible.
Regarding claims 123 and 137 , the claim 123 recites a method and claim 137 recites a non-transitory computer-readable medium which fall within at least one of the four statutory categories. Claims 123 and 137 recite similar limitations as indicated above with respect to claim 111. Hence, the claim is not eligible for the same reasons as discussed above with respect to claim 111. All other limitations not discussed are the same as those discussed above with to claim 111. Discussion is omitted for brevity.
Claims 112, 114, 116, 119, 121, 141-142, and 145-146 are also rejected under 35 U.S.C. 101 by virtue of their dependency to the independent claims.
Claims 112, 114, 116, 119, 121, 141-142, and 145-146 do not recite additional elements that integrate the judicial exception into a practical application, because the additional elements are directed toward additional aspects of judicial exception and/or well-understood, routine and conventional additional elements that do not integrate the judicial exception into a practical application. For example, claim 114 recites “wherein the correlation is further based on at least one of the position information…” further the abstract idea.
The dependent claims are rejected under 35 U.S.C. 101 under similar rationale as their independent claims.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 79-80, 82, 88, 90 – 91, 95-96, 99, 111 – 112, 114, 116, 119, 121, 123, 135, 137 and 139 – 148 are rejected under 35 U.S.C. 103 as being unpatentable over Ma et al. “Exploiting Sparse Semantic HD Maps for Self-Driving Vehicle Localization”, 2019 (Cited in IDS filed on 05/25/2022) in view Zou et al. (US 20170300763 A1), in view of Singh et al. (US 20220182498 A1), in view of Fridman (US 20180025235 A1, cited previously in the Office Action of 11/18/2024), and further in view of Lindemann (US 20190147257 A1).
a. Regarding claim 79, and similarly with respect to claims 99 and 135, Ma et al. discloses A host vehicle-based sparse map feature harvester system, comprising: (Title, “Exploiting Sparse Semantic HD Maps for Self-Driving Vehicle Localization”) at least one processor comprising circuitry and a memory, wherein the memory includes instructions that when executed by the circuitry cause the at least one processor to: (Section IV(B)(b) “FFT-conv is used to accelerate the computation speed by a factor of 20 over the state-of-the-art GEMM-based spatial GPU correlation implementations”)
receive a first image captured by a forward-facing camera onboard the host vehicle, as the host vehicle travels along a road segment in a first direction, (Section IV(A) “Our localization system exploits a wide variety of sensors: GPS, IMU, wheel encoders, LiDAR, and cameras. These sensors are available in most self-driving vehicles. The GPS provides a coarse location with several meters accuracy; an IMU captures vehicle dynamic measurements; the wheel encoders measure the total travel distance; the LiDAR accurately perceives the geometry of the surrounding area through a sparse point cloud; images capture dense and rich appearance information.”) wherein the first image is representative of an environment in a forward direction relative to the first direction; (Fig. 3 “Our system detects signs in the camera images”)
detect a first semantic feature represented in the first image, wherein the first semantic feature is associated with a predetermined object type classification, (Fig. 3 “Our system detects signs in the camera images”, and section IV(A)(b) “we run an image-based semantic segmentation algorithm that performs dense semantic labeling of traffic signs.”)
identify, using at least one trained machine learning model, at least one position descriptor associated with the first semantic feature represented in the first image captured by the forward-facing camera; (Fig. 2 “We first detect signs in 2D using semantic segmentation in the camera frame, and then use the LiDAR points to localize the signs in 3D.” and Fig. 3 “an illustration of the neural network’s input and output”)
Ma et al. fails to explicitly disclose receive a second image captured by a rearward-facing camera onboard the host vehicle, as the host vehicle travels along the road segment in the first direction, wherein the second image is representative of an environment in a backward direction relative to the first direction; detect a second semantic feature represented in the second image, wherein the second semantic feature is associated with a predetermined object type classification, . . .; identify, using the at least one trained machine learning model, at least one position descriptor associated with the second semantic feature represented in the second image captured by the rearward-facing camera; cause transmission of drive information for the road segment to an entity remotely-located relative to the host vehicle, wherein the drive information includes the at least one position descriptor associated with the first semantic feature, the at least one position descriptor associated with the second semantic feature,
Zou et al. teaches receive a second image captured by a rearward-facing camera onboard the host vehicle, (Fig. 1, 130d) as the host vehicle travels along the road segment in the first direction, wherein the second image is representative of an environment in a backward direction relative to the first direction; ( [0047] “ In particular, the method 600 provides for coordinating multi-camera fusion of images from a plurality of cameras (e.g., the cameras 130a, 130b, 130c, 130d).”, [0048] “the processing system 110 receives an image from each of the cameras 130. At block 604, for each of the cameras 130, following occurs: the top view generation engine 212 generates a top view of the road based on the image; the lane boundaries detection engine 214 detects lane boundaries of a lane of the road based on the top view; and the road feature detection engine 216 detects a road feature within the lane boundaries of the lane of the road using machine learning. The road feature detection engine can detect multiple road features.”)
detect a second semantic feature represented in the second image, wherein the second semantic feature is associated with a predetermined object type classification, . . . ([0040] “The feature extraction 402 uses a neural network, as described herein, to extract road features, for example, using feature maps. The feature extraction 402 outputs the road features to the classification 404 to classify the road features, such as based on road features stored in the road feature database 218. The classification 404 outputs the road features 406a, 406b, 406c, 406d, etc., which can be a speed limit indicator, a bicycle lane indicator, a railroad indicator, a school zone indicator, a direction indicator, or other road feature.” and [0041] “It should be appreciated that the road feature detection engine 216, using the feature detection 402 and classification 404, can detect multiple road features (e.g., road features 406a, 406b, 406e, 406d, etc.) in parallel as one step and in real-time.”)
identify, using the at least one trained machine learning model, at least one position descriptor associated with the second semantic feature represented in the second image captured by the rearward-facing camera; ([0034] “The road feature detection engine 216 searches within the top view, as defined by the lane boundaries, to detect road features. The road feature detection engine 216 can determine a type of road feature (e.g., a straight arrow, a left-turn arrow, etc.) as well as a location of the road feature (e.g., arrow ahead, bicycle lane to the left, etc.).” and [0041] “It should be appreciated that the road feature detection engine 216, using the feature detection 402 and classification 404, can detect multiple road features (e.g., road features 406a, 406b, 406e, 406d, etc.) in parallel as one step and in real-time.”)
cause transmission of drive information for the road segment to an entity remotely-located relative to the host vehicle, (Fig. 2, [0034] “the road feature detection engine 216 uses the lane boundaries to detect road features within the lane boundaries of the lane of the road using machine learning and/or computer vision techniques. The road feature detection engine 216 searches within the top view, as defined by the lane boundaries, to detect road features. The road feature detection engine 216 can determine a type of road feature (e.g., a straight arrow, a left-turn arrow, etc.) as well as a location of the road feature (e.g., arrow ahead, bicycle lane to the left, etc.).” and [0035] “The road feature database 218 can be updated when road features are detected, and the road feature database 218 can be accessible by other vehicles, such as from a cloud computing environment over a network or from the vehicle 100 directly (e.g., using direct short-range communications (DSCR)). This enables crowd-sourcing of road features.”) wherein the drive information includes the at least one position descriptor associated with the first semantic feature, the at least one position descriptor associated with the second semantic feature, ([0034] “The road feature detection engine 216 can determine a type of road feature (e.g., a straight arrow, a left-turn arrow, etc.) as well as a location of the road feature (e.g., arrow ahead, bicycle lane to the left, etc.).”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the vehicle detection system of Ma et al. to incorporate multiple cameras including a rearward camera and a cloud/crowdsource system as taught by Zou et al. for the purpose of allowing the vehicle to detect features of the environment on all sides as well as inform other vehicles, increasing coverage of features.
Ma et al. in combination with Zou et al. fails to explicitly disclose receive position information indicative of a position of the forward- facing camera when the first image was captured and indicative of a position of the rearward-facing camera when the second image was captured; and wherein the drive information includes … the position information.
Singh et al. teaches receive position information indicative of a position of the forward- facing camera when the first image was captured and indicative of a position of the rearward-facing camera when the second image was captured; and wherein the drive information includes … the position information. ([0075] “The sensors may also include an image sensor (imaging means) that captures an image of surroundings of the vehicle 200. The example of FIG. 2 includes the first front camera 203, the side camera 204, the rear camera 208, and the second front camera 205, as image sensors.” and [0078] “the system 700 may receive data from each of the sensors described above. These data may be configured to allow determination or estimation of, for example, a position of each of objects around the vehicle 200 with respect to the vehicle 200 (or with respect to the corresponding one of the sensors), a distance from the vehicle 200 (or from each sensor) to the corresponding one of the objects, a type of each of the objects, and a behavior (e.g., a movement direction and speed of an object) of each of the objects.)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the feature location information/detection of Ma et al. in combination with Zou et al. to use position information of the respective sensor/camera when detecting features as taught by Singh et al. for the purpose of providing accurate and precise location information of the feature.
Ma et al. in combination with Zou et al. and Singh et al. fails to explicitly disclose the entity remotely-located relative to the host vehicle being configured to: determine, based on the at least one position descriptor associated with the first semantic feature, the at least one position descriptor associated with the second semantic feature, and the position information, a position of the first semantic feature relative to the second semantic feature; receive, in addition to the drive information transmitted by the host vehicle, drive information from a first plurality of other vehicles that travel along the road segment and drive information from a second plurality of other vehicles that travel along the road segment, the drive information from the first plurality of other vehicles being representative of the environment in the backward direction relative to the first direction; and align the drive information received from the first plurality of other vehicles with the drive information received from the second plurality of other vehicles based, at least in part, on the relationship between the first semantic feature and the second semantic feature.
Fridman teaches the entity remotely-located relative to the host vehicle being configured to: (Figure 12, 1230) determine, based on the at least one position descriptor associated with the first semantic feature, the at least one position descriptor associated with the second semantic feature, and the position information, a position of the first semantic feature relative to the second semantic feature; ([0017] “determining a line representation of a road surface feature extending along a road segment, where the line representation of the road surface feature is configured for use in autonomous vehicle navigation, may comprise receiving, by a server, a first set of drive data including position information associated with the road surface feature, and receiving, by a server, a second set of drive data including position information associated with the road surface feature. The position information may be determined based on analysis of images of the road segment. The method may further comprise segmenting the first set of drive data into first drive patches and segmenting the second set of drive data into second drive patches; longitudinally aligning the first set of drive data with the second set of drive data within corresponding patches; and determining the line representation of the road surface feature based on the longitudinally aligned first and second drive data in the first and second draft patches.”)
receive, in addition to the drive information transmitted by the host vehicle, drive information from a first plurality of other vehicles that travel along the road segment in the first direction and drive information from a second plurality of other vehicles that travel along the road segment in a second direction opposite to the first direction (Figure 12, [0053] “a system that uses crowd sourcing data received from a plurality of vehicles for autonomous vehicle navigation”, [0278] “Each vehicle may be similar to vehicles disclosed in other embodiments (e.g., vehicle 200), and may include components or devices included in or associated with vehicles disclosed in other embodiments. Each vehicle may be equipped with an image capture device or camera (e.g., image capture device 122 or camera 122). Each vehicle may communicate with a remote server 1230 via one or more networks (e.g., over a cellular network and/or the Internet, etc.) through wireless communication paths 1235, as indicated by the dashed lines. Each vehicle may transmit data to server 1230 and receive data from server 1230. For example, server 1230 may collect data from multiple vehicles travelling on the road segment 1200 at different times, and may process the collected data to generate an autonomous vehicle road navigation model, or an update to the model. Server 1230 may transmit the autonomous vehicle road navigation model or the update to the model to the vehicles that transmitted data to server 1230.”, and see at least [0277] and [0279])
align the drive information received from the first plurality of other vehicles with the drive information received from the second plurality of other vehicles based, at least in part, on the relationship between the first semantic feature and the second semantic feature. (Figure 15, [0292] “the data (e.g. ego-motion data, road markings data, and the like) may be shown as a function of position S (or S.sub.1 or S.sub.2) along the drive. Server 1230 may identify landmarks for the sparse map by identifying unique matches between landmarks 1501, 1503, and 1505 of drive 1510 and landmarks 1507 and 1509 of drive 1520. Such a matching algorithm may result in identification of landmarks 1511, 1513, and 1515. One skilled in the art would recognize, however, that other matching algorithms may be used. For example, probability optimization may be used in lieu of or in combination with unique matching. As described in further detail below with respect to FIG. 29, server 1230 may longitudinally align the drives to align the matched landmarks. For example, server 1230 may select one drive (e.g., drive 1520) as a reference drive and then shift and/or elastically stretch the other drive(s) (e.g., drive 1510) for alignment.”, and [0293] “aligned landmark data for use in a sparse map. In the example of FIG. 16, landmark 1610 comprises a road sign. The example of FIG. 16 further depicts data from a plurality of drives 1601, 1603, 1605, 1607, 1609, 1611, and 1613. In the example of FIG. 16, the data from drive 1613 consists of a “ghost” landmark, and the server 1230 may identify it as such because none of drives 1601, 1603, 1605, 1607, 1609, and 1611 include an identification of a landmark in the vicinity of the identified landmark in drive 1613.”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the feature location information/detection of Ma et al. in combination with Zou et al. and Singh to incorporate crowdsourcing data and aligning data received from other vehicles as taught by Fridman for the purpose of allowing the vehicle to navigate one or more roads.
However, Ma et al. in combination with Zou et al., Singh, and Fridman may be alleged to no explicitly disclose wherein the first semantic feature corresponds to a front side of an object along the road segment; and wherein the second semantic feature corresponds to a rear side of the object along the road segment, the first semantic feature and the second semantic feature being associated with a common object, the second semantic feature being detected after the host vehicle passes the object;
Lindemann teaches wherein the first semantic feature corresponds to a front side of an object along the road segment; (Figure 1, and Figures 5a – 5d) wherein the second semantic feature corresponds to a rear side of the object along the road segment, (Figures 5a – 5d) the first semantic feature and the second semantic feature being associated with a common object, (101, Figures 1-2) the second semantic feature being detected after the host vehicle passes the object; ([0030] “FIG. 2 shows an exemplary image 200, taken by a camera pointing rearward with respect to the vehicle, with the traffic sign 101 shown in FIG. 1 after it has been passed. The vehicle is moving in the lane of the road 102, situated on the left in the image, toward the viewer, as indicated by the arrow 103. The content of the traffic sign 101 is now clearly recognizable—it shows a speed limit of 80 km/h. A process, applied to the image, for traffic sign recognition can thus identify the traffic sign that applies to the opposite direction and link it to a placement site that was captured as the vehicle drove past. This data set can then be stored in a database and processed further.”, [0035] “a vehicle 104 is driving in the right-hand lane of a road 102. The direction of travel is indicated by the arrow 103, the index v at the arrow 103 representing a driving speed of the vehicle 104. Gray fields 106 and 108 indicate the capturing regions of a respective camera of the vehicle (not shown in the figure) that is oriented to the front or the rear in the driving direction. The capturing regions 106 and 108 extend beyond the gray region in the continuation of the lateral boundary lines. In the capturing region 106 of the camera oriented to the front, a traffic sign 101 that applies to the opposite direction is placed next to the lane, without the content of said traffic sign being able to be uniquely determined from the rear. According to an aspect of the method, the traffic sign is marked as a candidate and tracked in subsequent images, i.e. the respective position thereof in a respectively latest image is determined,”, and [0037] “the vehicle 104 has passed the traffic sign 101, and the traffic sign 101 has entered the capturing region of the camera oriented to the rear. The content of the traffic sign is now recognizable on the images thereof and can be determined using an apparatus for traffic sign recognition (not shown in the figure). The placement site of the traffic sign 101 can be determined either from the known optical properties of the camera oriented to the front or the camera oriented to the rear and the known geographic position of the vehicle, e.g. for the candidate when leaving the capturing region of the camera oriented to the front, or for the candidate when entering the capturing region of the camera oriented to the rear. As soon as a traffic sign recognition has been successfully performed, the position of the now confirmed candidate can be assigned to the traffic sign.”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the feature location information/detection of Ma et al. in combination with Zou et al., Singh, and Fridman to incorporate detecting a common feature from different cameras as taught by Lindemann for the purpose of accurately recognizing the road feature (e.g. traffic sign), providing accurate information to the host vehicle for navigation purposes.
b. Regarding claim 80, and similarly with respect to claim 143, Ma et al. in view of Zou et al., Singh et al., Fridman, and Lindemann discloses The system of claim 79,
Ma et al. discloses the at least one position descriptor associated with the first semantic feature includes an X-Y image position relative to the first image and the at least one position descriptor associated with the second semantic feature includes an X-Y image position relative to the (Fig. 2 “We first detect signs in 2D using semantic segmentation in the camera frame”). Examiner Notes: The plurality of images captures the X-Y position of the traffic signs.
Zou et al. teaches the at least one position descriptor associated with … the at least one position descriptor associated with the second semantic feature includes ([0034] “The road feature detection engine 216 can determine a type of road feature (e.g., a straight arrow, a left-turn arrow, etc.) as well as a location of the road feature (e.g., arrow ahead, bicycle lane to the left, etc.).”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the vehicle detection system of Ma et al. in combination with Zou et al., Singh et al., Fridman, and Lindemann to incorporate multiple cameras including a rearward camera as taught by Zou et al. for the purpose of allowing the vehicle to detect features of the environment on all sides, increasing coverage of features. Further, It would have been obvious to one of ordinary skill in the art, when in the combination, to perform detecting features in 2D using semantic segmentation as taught by Ma et al. with the second image captured by the rearward facing camera of Zou et al.
c. Regarding claim 82, and similarly with respect to claim 144, Ma et al. in view of Zou et al., Singh et al., Fridman, and Lindemann discloses The system of claim 80
Ma et al. discloses wherein the X-Y-Z position is determined based on an one ego motion of the host vehicle and based on at least one of: (Fig.2, “then use the LiDAR points to localize the signs in 3D. Mapping can aggregate information from multiple passes through
the same area using the ground truth pose information”) tracking a change in image position of the first semantic feature between the first image and at least one additional image or tracking a change in image position of the second semantic feature between the second image and the at least one additional image, the ego motion being determined based on at least one of a plurality of images or an output of at least one eqo motion sensor, the at least one eqo motion sensor including at least one of a speedometer, an accelerometer, or a GPS receiver. (Section IV “We formulate the localization problem as a histogram filter taking as input the structured outputs of our sign and lane detection neural networks, as well as GPS, IMU, and wheel odometry information, and outputting a probability histogram over the vehicle’s pose,
expressed in world coordinates.” and Section IV(A) “Let Gt be the GPS readings at time t and let L and T represent the lane graph and traffic sign maps respectively. We compute an estimate of the vehicle dynamics Xt from both IMU and the wheel encoders smoothed through an extended Kalman filter, which is updated at 100Hz. The localization task is formulated as a histogram filter aiming to maximize the agreement between the observed and mapped lane graphs and traffic signs while respecting vehicle dynamics”)
d. Regarding claim 88, Ma et al. in view of Zou et al., Singh et al., Fridman, and Lindemann discloses The system of claim 79,
Zou et al. teaches wherein at least one of the first semantic feature or the second semantic feature includes one of: a front side of a speed limit sign, a yield sign, a pole, a painted directional arrow, a traffic light, a billboard, or a building. ([0034] “The road feature detection engine 216 can determine a type of road feature (e.g., a straight arrow, a left-turn arrow, etc.) as well as a location of the road feature (e.g., arrow ahead, bicycle lane to the left, etc.).”) and [0035] “The road features can be predefined in a database of road features (e.g., road feature database 218). Examples of road features include a speed limit indicator, a bicycle lane indicator, a railroad indicator, a school zone indicator, and a direction indicator (e.g., left-turn arrow, straight arrow, right-turn arrow, straight and left-turn arrow, straight and right-turn arrow, etc.), and the like. “)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the semantic labeling system of features of Ma et al. in combination with Zou et al., Singh et al., Fridman, and Lindemann to incorporate specific labels of features including speed limit signs etc. as taught by Zou et al. for the purpose of precisely labeling feature types for accurate feature detection.
e. Regarding claim 90, and similarly with respect to claim 139, Ma et al. in view of Zou et al., Singh et al., Fridman, and Lindemann discloses The system of claim 79,
Zou et al. teaches wherein the drive information further includes one or more descriptors associated with each of the first and second semantic features. ([0034] “The road feature detection engine 216 can determine a type of road feature (e.g., a straight arrow, a left-turn arrow, etc.) as well as a location of the road feature (e.g., arrow ahead, bicycle lane to the left, etc.).”), and [0035] “The road features can be predefined in a database of road features (e.g., road feature database 218). Examples of road features include a speed limit indicator, a bicycle lane indicator, a railroad indicator, a school zone indicator, and a direction indicator (e.g., left-turn arrow, straight arrow, right-turn arrow, straight and left-turn arrow, straight and right-turn arrow, etc.), and the like.”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the semantic labeling system of features of Ma et al. in combination with Zou et al., Singh et al., Fridman, and Lindemann to incorporate specific descriptors of features including speed limit signs etc. as taught by Zou et al. for the purpose of precisely labeling feature types for accurate feature detection.
f. Regarding claim 91, and similarly with respect to claim 140, Ma et al. in view of Zou et al., Singh et al., Fridman, and Lindemann discloses The system of claim 90,
Zou et a. teaches wherein the one or more descriptors include at least one of a height or a width, a bounding box, or a type classification. ([0045] “The road feature detection engine 216 can perform feature extraction from the top view (within the lane boundaries) using a neural network and then perform a classification of the road features (e.g., determine what kind of road feature).”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the semantic labeling system of features of Ma et al. in combination with Zou et al., Singh et al., Fridman, and Lindemann to incorporate specific classification of features (e.g. classification of road features) as taught by Zou et al. for the purpose of precisely labeling feature types for accurate feature detection.
g. Regarding claim 95, Ma et al. in view of Zou et al., Singh et al., Fridman, and Lindemann discloses The system of claim 94,
Fridman teaches wherein the entity remotely-located relative to the host vehicle is further configured to align the drive information received from the first plurality of other vehicles with the drive information received from the second plurality of other vehicles based on the position information. (Figure 15, [0292] “the data (e.g. ego-motion data, road markings data, and the like) may be shown as a function of position S (or S.sub.1 or S.sub.2) along the drive. Server 1230 may identify landmarks for the sparse map by identifying unique matches between landmarks 1501, 1503, and 1505 of drive 1510 and landmarks 1507 and 1509 of drive 1520. Such a matching algorithm may result in identification of landmarks 1511, 1513, and 1515. One skilled in the art would recognize, however, that other matching algorithms may be used. For example, probability optimization may be used in lieu of or in combination with unique matching. As described in further detail below with respect to FIG. 29, server 1230 may longitudinally align the drives to align the matched landmarks. For example, server 1230 may select one drive (e.g., drive 1520) as a reference drive and then shift and/or elastically stretch the other drive(s) (e.g., drive 1510) for alignment.”, and [0293] “aligned landmark data for use in a sparse map. In the example of FIG. 16, landmark 1610 comprises a road sign. The example of FIG. 16 further depicts data from a plurality of drives 1601, 1603, 1605, 1607, 1609, 1611, and 1613. In the example of FIG. 16, the data from drive 1613 consists of a “ghost” landmark, and the server 1230 may identify it as such because none of drives 1601, 1603, 1605, 1607, 1609, and 1611 include an identification of a landmark in the vicinity of the identified landmark in drive 1613.”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the feature location information/detection of Ma et al. in combination with Zou et al., Singh, Fridman, and Lindemann to incorporate crowdsourcing data and aligning data received from other vehicles as taught by Fridman for the purpose of allowing the vehicle to navigate one or more roads.
h. Regarding claim 96, and similarly with respect to claims 116 and 145, Ma et al. in view of Zou et al., Singh et al., Fridman, and Lindemann discloses The system of claim 79,
Singh et al. teaches wherein the position information includes at least one indicator of position determined based on an output of a GPS sensor associated with the host vehicle, at least one indicator of position determined based on an ego motion of the host vehicle, or a combination of the output of the GPS sensor and ego motion of the host vehicle. ([0051] “The vehicle may receive a global position and its state (global speed, direction, acceleration, etc.). The vehicle may also identify itself or determine its position on the map based on its condition and position.”, [0076] “The sensors may also include a position sensor that acquires position information on the vehicle. The example of FIG. 2 includes the GPS and the INS 207 as position sensors.” and [0078] “the system 700 may receive data from each of the sensors described above. These data may be configured to allow determination or estimation of, for example, a position of each of objects around the vehicle 200 with respect to the vehicle 200 (or with respect to the corresponding one of the sensors), a distance from the vehicle 200 (or from each sensor) to the corresponding one of the objects, a type of each of the objects, and a behavior (e.g., a movement direction and speed of an object) of each of the objects.)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the feature location information/detection of Ma et al. in combination with Zou et al., Singh et al., Fridman, and Lindemann to use position and motion information of the respective sensor/camera/vehicle when detecting features as taught by Singh et al. for the purpose of providing accurate and precise location information of the feature.
i. Regarding claim 111, and similarly with respect to claims 123 and 137, Ma et al. discloses A host vehicle-based sparse map feature harvester system, (Title, “Exploiting Sparse Semantic HD Maps for Self-Driving Vehicle Localization”) comprising: at least one processor comprising circuitry and a memory, wherein the memory includes instructions that when executed by the circuitry cause the at least one processor to: (Section IV(B)(b) “FFT-conv is used to accelerate the computation speed by a factor of 20 over the state-of-the-art GEMM-based spatial GPU correlation implementations”)
receive a first image captured by a forward-facing camera onboard the host vehicle, as the host vehicle travels along a road segment in a first direction, (Section IV(A) “Our localization system exploits a wide variety of sensors: GPS, IMU, wheel encoders, LiDAR, and cameras. These sensors are available in most self-driving vehicles. The GPS provides a coarse location with several meters accuracy; an IMU captures vehicle dynamic measurements; the wheel encoders measure the total travel distance; the LiDAR accurately perceives the geometry of the surrounding area through a sparse point cloud; images capture dense and rich appearance information.”) wherein the first image is representative of an environment forward of the host vehicle; (Fig. 3 “Our system detects signs in the camera images”)
detect at least one object represented in the first image; (Fig. 3 “Our system detects signs in the camera images”, and section IV(A)(b) “we run an image-based semantic segmentation algorithm that performs dense semantic labeling of traffic signs.”)
identify, using at least one trained machine learning model, at least one front side two-dimensional feature point, the at least one front side two-dimensional feature point being associated with the at least one object represented in the first image; (Fig. 2 “We first detect signs in 2D using semantic segmentation in the camera frame, and then use the LiDAR points to localize the signs in 3D.”, and Figure 3)
wherein the drive information includes the at least one front side two-dimensional feature point, (Fig. 2 “We first detect signs in 2D using semantic segmentation in the camera frame, and then use the LiDAR points to localize the signs in 3D.”)
Ma et al. fails to explicitly disclose receive a second image captured by a rearward-facing camera onboard the host vehicle, as the host vehicle travels along the road segment in the first direction, wherein the second image is representative of an environment behind the host vehicle; detect a representation of the at least one object in the second image, . . . ; identify at least one rear side
Zou et al. teaches receive a second image captured by a rearward-facing camera onboard the host vehicle, (Fig. 1, 130d) as the host vehicle travels along the road segment in the first direction, wherein the second image is representative of an environment behind the host vehicle; ([0047] “ In particular, the method 600 provides for coordinating multi-camera fusion of images from a plurality of cameras (e.g., the cameras 130a, 130b, 130c, 130d).”, [0048] “the processing system 110 receives an image from each of the cameras 130. At block 604, for each of the cameras 130, following occurs: the top view generation engine 212 generates a top view of the road based on the image; the lane boundaries detection engine 214 detects lane boundaries of a lane of the road based on the top view; and the road feature detection engine 216 detects a road feature within the lane boundaries of the lane of the road using machine learning. The road feature detection engine can detect multiple road features.”)
detect a representation of the at least one object in the second image, . . .; ([0040] “The feature extraction 402 uses a neural network, as described herein, to extract road features, for example, using feature maps. The feature extraction 402 outputs the road features to the classification 404 to classify the road features, such as based on road features stored in the road feature database 218. The classification 404 outputs the road features 406a, 406b, 406c, 406d, etc., which can be a speed limit indicator, a bicycle lane indicator, a railroad indicator, a school zone indicator, a direction indicator, or other road feature.”, [0041] “It should be appreciated that the road feature detection engine 216, using the feature detection 402 and classification 404, can detect multiple road features (e.g., road features 406a, 406b, 406e, 406d, etc.) in parallel as one step and in real-time.” and [0049] “At block 606, a fusion engine (not shown) first fuses the lane boundary information from each of the cameras, then based on consolidated lane information, fuses road features from each of the cameras. Fusing the road features from each of the cameras provides for uninterrupted and accurate road feature detecting by predicting road feature locations when a road feature passes from the FOV of one camera to the FOV of another camera, when a road feature is partially obstructed in the FOV of one or more cameras, and the like.”)
identify, using at least one trained machine learning model, at least one rear side two-dimensional feature point, the at least one rear side two-dimensional feature point being associated with the at least one object represented in the second image; ([0034] “The road feature detection engine 216 searches within the top view, as defined by the lane boundaries, to detect road features. The road feature detection engine 216 can determine a type of road feature (e.g., a straight arrow, a left-turn arrow, etc.) as well as a location of the road feature (e.g., arrow ahead, bicycle lane to the left, etc.).” and [0041] “It should be appreciated that the road feature detection engine 216, using the feature detection 402 and classification 404, can detect multiple road features (e.g., road features 406a, 406b, 406e, 406d, etc.) in parallel as one step and in real-time.” and [0049] “At block 606, a fusion engine (not shown) first fuses the lane boundary information from each of the cameras, then based on consolidated lane information, fuses road features from each of the cameras. Fusing the road features from each of the cameras provides for uninterrupted and accurate road feature detecting by predicting road feature locations when a road feature passes from the FOV of one camera to the FOV of another camera, when a road feature is partially obstructed in the FOV of one or more cameras, and the like.”)
cause transmission of drive information for the road segment to an entity remotely-located relative to the host vehicle, (Fig. 2, [0034] “the road feature detection engine 216 uses the lane boundaries to detect road features within the lane boundaries of the lane of the road using machine learning and/or computer vision techniques. The road feature detection engine 216 searches within the top view, as defined by the lane boundaries, to detect road features. The road feature detection engine 216 can determine a type of road feature (e.g., a straight arrow, a left-turn arrow, etc.) as well as a location of the road feature (e.g., arrow ahead, bicycle lane to the left, etc.).” and [0035] “. The road feature database 218 can be updated when road features are detected, and the road feature database 218 can be accessible by other vehicles, such as from a cloud computing environment over a network or from the vehicle 100 directly (e.g., using direct short-range communications (DSCR)). This enables crowd-sourcing of road features.”) wherein the drive information includes the at least one front side two- dimensional feature point, the at least one rear side two-dimensional feature point, ([0024] “The cameras 130 capture images external to the vehicle 100. Each of the cameras 130 has a field-of-view (FOV) 131a, 131b, 131c, 131d (collectively referred to herein as “FOV 131”). The FOV is the area observable by a camera. For example, the camera 130a has an FOV 131a, the camera 131b has an FOV 131b, the camera 130c has an FOV 131c, and the camera 131d has an FOV 131d. The captured images can be the entire FOV for the camera or can be a portion of the FOV of the camera.”, [0033] “Detecting lane boundaries enables the detection of road features within the lane boundaries. In particular, the final image (e.g., curve consolidation 312, also referred to as “boundary image 312”) is used to detect road features, although the road features can be detected using other images” and [0034] “The road feature detection engine 216 can determine a type of road feature (e.g., a straight arrow, a left-turn arrow, etc.) as well as a location of the road feature (e.g., arrow ahead, bicycle lane to the left, etc.).”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the vehicle detection system of Ma et al. to incorporate multiple cameras including a rearward camera and a cloud/crowdsource system as taught by Zou et al. for the purpose of allowing the vehicle to detect features of the environment on all sides as well as inform other vehicles, increasing coverage of features. Further, It would have been obvious to one of ordinary skill in the art, when in the combination, to perform detecting features in 2D using semantic segmentation as taught by Ma et al. with the second image captured by the rearward facing camera of Zou et al.
Ma et al. in combination with Zou et al. fails to explicitly disclose receive position information indicative of a position of the forward- facing camera when the first image was captured and indicative of a position of the rearward-facing camera when the second image was captured; and wherein the drive information includes …the position information.
Singh et al. teaches receive position information indicative of a position of the forward-facing camera when the first image was captured and indicative of a position of the rearward-facing camera when the second image was captured; and wherein the drive information includes … the position information. ([0075] “The sensors may also include an image sensor (imaging means) that captures an image of surroundings of the vehicle 200. The example of FIG. 2 includes the first front camera 203, the side camera 204, the rear camera 208, and the second front camera 205, as image sensors.” and [0078] “the system 700 may receive data from each of the sensors described above. These data may be configured to allow determination or estimation of, for example, a position of each of objects around the vehicle 200 with respect to the vehicle 200 (or with respect to the corresponding one of the sensors), a distance from the vehicle 200 (or from each sensor) to the corresponding one of the objects, a type of each of the objects, and a behavior (e.g., a movement direction and speed of an object) of each of the objects.)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the feature location information/detection of Ma et al. in combination with Zou et al. to use position information of the respective sensor/camera when detecting features as taught by Singh et al. for the purpose of providing accurate and precise location information of the feature.
Ma et al. in combination with Zou et al., and Singh et al. fails to explicitly disclose the entity remotely-located relative to the host vehicle being configured to: determine, based on the at least one front side two-dimensional feature point, at least one rear side two-dimensional feature point, and the position information that the at least one rear side two-dimensional feature point are associated with a common object; receive, in addition to the drive information transmitted by the host vehicle, drive information from a first plurality of other vehicles that travel along the road segment and drive information from a second plurality of other vehicles that travel along the road segment, the drive information from the first plurality of other vehicles being representative of the environment in the backward direction relative to the first direction; and align the drive information received from the first plurality of other vehicles with the drive information received from the second plurality of other vehicles based, at least in part, on the determination that at least one front side two-dimensional feature point and the at least one rear side two-dimensional feature point are associated with a common object.
Fridman teaches the entity remotely-located relative to the host vehicle being configured to: (Figure 12, 1230) the entity remotely-located relative to the host vehicle being configured to: determine, based on the at least one front side two-dimensional feature point, at least one rear side two-dimensional feature point, and the position information that the at least one rear side two-dimensional feature point are associated with a common object; ([0017] “determining a line representation of a road surface feature extending along a road segment, where the line representation of the road surface feature is configured for use in autonomous vehicle navigation, may comprise receiving, by a server, a first set of drive data including position information associated with the road surface feature, and receiving, by a server, a second set of drive data including position information associated with the road surface feature. The position information may be determined based on analysis of images of the road segment. The method may further comprise segmenting the first set of drive data into first drive patches and segmenting the second set of drive data into second drive patches; longitudinally aligning the first set of drive data with the second set of drive data within corresponding patches; and determining the line representation of the road surface feature based on the longitudinally aligned first and second drive data in the first and second draft patches.”)
receive, in addition to the drive information transmitted by the host vehicle, drive information from a first plurality of other vehicles that travel along the road segment and (Figure 12, [0053] “a system that uses crowd sourcing data received from a plurality of vehicles for autonomous vehicle navigation”, and see at least paragraph [0277]) drive information from a second plurality of other vehicles that travel along the road segment, the drive information from the first plurality of other vehicles being representative of the environment in the backward direction relative to the first direction; and (Figure 12, [0053] “a system that uses crowd sourcing data received from a plurality of vehicles for autonomous vehicle navigation”, and see at least paragraph [0277]) align the drive information received from the first plurality of other vehicles with the drive information received from the second plurality of other vehicles based, at least in part, on the determination that at least one front side two-dimensional feature point and the at least one rear side two-dimensional feature point are associated with a common object.(Figure 15, [0292] “the data (e.g. ego-motion data, road markings data, and the like) may be shown as a function of position S (or S.sub.1 or S.sub.2) along the drive. Server 1230 may identify landmarks for the sparse map by identifying unique matches between landmarks 1501, 1503, and 1505 of drive 1510 and landmarks 1507 and 1509 of drive 1520. Such a matching algorithm may result in identification of landmarks 1511, 1513, and 1515. One skilled in the art would recognize, however, that other matching algorithms may be used. For example, probability optimization may be used in lieu of or in combination with unique matching. As described in further detail below with respect to FIG. 29, server 1230 may longitudinally align the drives to align the matched landmarks. For example, server 1230 may select one drive (e.g., drive 1520) as a reference drive and then shift and/or elastically stretch the other drive(s) (e.g., drive 1510) for alignment.”, and [0293] “aligned landmark data for use in a sparse map. In the example of FIG. 16, landmark 1610 comprises a road sign. The example of FIG. 16 further depicts data from a plurality of drives 1601, 1603, 1605, 1607, 1609, 1611, and 1613. In the example of FIG. 16, the data from drive 1613 consists of a “ghost” landmark, and the server 1230 may identify it as such because none of drives 1601, 1603, 1605, 1607, 1609, and 1611 include an identification of a landmark in the vicinity of the identified landmark in drive 1613.”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the feature location information/detection of Ma et al. in combination with Zou et al. and Singh to incorporate crowdsourcing data and aligning data received from other vehicles as taught by Fridman for the purpose of allowing the vehicle to navigate one or more roads.
However, Ma et al. in combination with Zou et al., Singh, and Fridman may be alleged to not explicitly disclose wherein the representation of the at least one object in the second image is detected after the host vehicle passes the at least one object along the road segment;
Lindemann teaches wherein the representation of the at least one object in the second image is detected after the host vehicle passes the at least one object along the road segment; (Figure 5d, [0030] “FIG. 2 shows an exemplary image 200, taken by a camera pointing rearward with respect to the vehicle, with the traffic sign 101 shown in FIG. 1 after it has been passed. The vehicle is moving in the lane of the road 102, situated on the left in the image, toward the viewer, as indicated by the arrow 103. The content of the traffic sign 101 is now clearly recognizable—it shows a speed limit of 80 km/h. A process, applied to the image, for traffic sign recognition can thus identify the traffic sign that applies to the opposite direction and link it to a placement site that was captured as the vehicle drove past. This data set can then be stored in a database and processed further.”, [0035] “a vehicle 104 is driving in the right-hand lane of a road 102. The direction of travel is indicated by the arrow 103, the index v at the arrow 103 representing a driving speed of the vehicle 104. Gray fields 106 and 108 indicate the capturing regions of a respective camera of the vehicle (not shown in the figure) that is oriented to the front or the rear in the driving direction. The capturing regions 106 and 108 extend beyond the gray region in the continuation of the lateral boundary lines. In the capturing region 106 of the camera oriented to the front, a traffic sign 101 that applies to the opposite direction is placed next to the lane, without the content of said traffic sign being able to be uniquely determined from the rear. According to an aspect of the method, the traffic sign is marked as a candidate and tracked in subsequent images, i.e. the respective position thereof in a respectively latest image is determined,”, and [0037] “the vehicle 104 has passed the traffic sign 101, and the traffic sign 101 has entered the capturing region of the camera oriented to the rear. The content of the traffic sign is now recognizable on the images thereof and can be determined using an apparatus for traffic sign recognition (not shown in the figure). The placement site of the traffic sign 101 can be determined either from the known optical properties of the camera oriented to the front or the camera oriented to the rear and the known geographic position of the vehicle, e.g. for the candidate when leaving the capturing region of the camera oriented to the front, or for the candidate when entering the capturing region of the camera oriented to the rear. As soon as a traffic sign recognition has been successfully performed, the position of the now confirmed candidate can be assigned to the traffic sign.”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the feature location information/detection of Ma et al. in combination with Zou et al., Singh, and Fridman to incorporate detecting a common feature from different cameras as taught by Lindemann for the purpose of accurately recognizing the road feature (e.g. traffic sign), providing accurate information to the host vehicle for navigation purposes.
j. Regarding claim 112, and similarly with respect to claim 141, Ma et al. in view of Zou et al., Singh et al., Fridman, and Lindemann discloses The system of claim 111,
Ma et al. discloses two-dimensional feature point … associated with the (Fig. 2 “We first detect signs in 2D using semantic segmentation in the camera frame, and then use the LiDAR points to localize the signs in 3D.”)
Zou et al. teaches wherein the drive information further includes an indicator that the at least one front side ([0049] “At block 606, a fusion engine (not shown) first fuses the lane boundary information from each of the cameras, then based on consolidated lane information, fuses road features from each of the cameras. Fusing the road features from each of the cameras provides for uninterrupted and accurate road feature detecting by predicting road feature locations when a road feature passes from the FOV of one camera to the FOV of another camera, when a road feature is partially obstructed in the FOV of one or more cameras, and the like.”). Examiner Notes: See “an indicator” that the features are associated with the common object as the object appearing in the predicted location during fusion.
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify 2D semantic feature identification and positioning system of Ma et al. in combination with Zou et a., Singh et al., Fridman, and Lindemann to incorporate fusing the front-view and rear-view camera information as taught by Zou et al. for the purpose of providing a consistent feature identification/positioning to share within the cloud server.
k. Regarding claim 114, Ma et al. in view of Zou et al., Singh et al., Fridman, and Lindemann discloses The system of claim 111,
Fridman teaches wherein the entity remotely-located relative to the host vehicle is further configured to align the drive information received from the first plurality of other vehicles with the drive information received from the second plurality of other vehicles. (Figure 15, [0292] “the data (e.g. ego-motion data, road markings data, and the like) may be shown as a function of position S (or S.sub.1 or S.sub.2) along the drive. Server 1230 may identify landmarks for the sparse map by identifying unique matches between landmarks 1501, 1503, and 1505 of drive 1510 and landmarks 1507 and 1509 of drive 1520. Such a matching algorithm may result in identification of landmarks 1511, 1513, and 1515. One skilled in the art would recognize, however, that other matching algorithms may be used. For example, probability optimization may be used in lieu of or in combination with unique matching. As described in further detail below with respect to FIG. 29, server 1230 may longitudinally align the drives to align the matched landmarks. For example, server 1230 may select one drive (e.g., drive 1520) as a reference drive and then shift and/or elastically stretch the other drive(s) (e.g., drive 1510) for alignment.”, and [0293] “aligned landmark data for use in a sparse map. In the example of FIG. 16, landmark 1610 comprises a road sign. The example of FIG. 16 further depicts data from a plurality of drives 1601, 1603, 1605, 1607, 1609, 1611, and 1613. In the example of FIG. 16, the data from drive 1613 consists of a “ghost” landmark, and the server 1230 may identify it as such because none of drives 1601, 1603, 1605, 1607, 1609, and 1611 include an identification of a landmark in the vicinity of the identified landmark in drive 1613.”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the feature map of Ma et al. in combination with Zou et al., Singh et al., Fridman, and Lindemann to incorporate the teachings of Fridman for the same reasons stated in the motivation of claim 111.
l. Regarding claim 119, and similarly with respect to claim 146, Ma et al. in view of Zou et al., Singh et al., Fridman, and Lindemann discloses The system of claim 111,
Zou et al. teaches wherein the memory further includes instructions that when executed by the circuitry cause the at least one processor to detect the representation of the at least one object in the second image based on a predetermined relationship between the forward-facing camera and the rearward-facing camera and based on an ego motion of the host vehicle. ([0040] “The feature extraction 402 uses a neural network, as described herein, to extract road features, for example, using feature maps. The feature extraction 402 outputs the road features to the classification 404 to classify the road features, such as based on road features stored in the road feature database 218. The classification 404 outputs the road features 406a, 406b, 406c, 406d, etc., which can be a speed limit indicator, a bicycle lane indicator, a railroad indicator, a school zone indicator, a direction indicator, or other road feature.”, [0041] “It should be appreciated that the road feature detection engine 216, using the feature detection 402 and classification 404, can detect multiple road features (e.g., road features 406a, 406b, 406e, 406d, etc.) in parallel as one step and in real-time.” and [0049] “At block 606, a fusion engine (not shown) first fuses the lane boundary information from each of the cameras, then based on consolidated lane information, fuses road features from each of the cameras. Fusing the road features from each of the cameras provides for uninterrupted and accurate road feature detecting by predicting road feature locations when a road feature passes from the FOV of one camera to the FOV of another camera, when a road feature is partially obstructed in the FOV of one or more cameras, and the like.” and [0049] “At block 606, a fusion engine (not shown) first fuses the lane boundary information from each of the cameras, then based on consolidated lane information, fuses road features from each of the cameras. Fusing the road features from each of the cameras provides for uninterrupted and accurate road feature detecting by predicting road feature locations when a road feature passes from the FOV of one camera to the FOV of another camera, when a road feature is partially obstructed in the FOV of one or more cameras, and the like.”).
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the camera detection system of Ma et al. in combination with Zou et a., Singh et al., Fridman, and Lindemann to incorporate detection of a feature using a plurality of cameras including a rearward-facing camera as taught by Zou et al. for the purpose of allowing the vehicle to detect features of the environment on all sides, increasing coverage of features.
m. Regarding claim 121, and similarly with respect to claim 142, Ma et al. in view of Zou et al., Singh et al., Fridman, and Lindemann discloses The system of claim 111,
Ma et al. discloses wherein the at least one front side two-dimensional feature point identified by the at least one processor includes an X-Y position relative to the first image and
the at least one (Fig. 2 “We first detect signs in 2D using semantic segmentation in the camera frame”). Examiner Notes: The plurality of images captures the X-Y position of the traffic signs.
Zou et al. teaches the at least one rear side ([0024] “The cameras 130 capture images external to the vehicle 100. Each of the cameras 130 has a field-of-view (FOV) 131a, 131b, 131c, 131d (collectively referred to herein as “FOV 131”). The FOV is the area observable by a camera. For example, the camera 130a has an FOV 131a, the camera 131b has an FOV 131b, the camera 130c has an FOV 131c, and the camera 131d has an FOV 131d. The captured images can be the entire FOV for the camera or can be a portion of the FOV of the camera.”, and [0034] “The road feature detection engine 216 can determine a type of road feature (e.g., a straight arrow, a left-turn arrow, etc.) as well as a location of the road feature (e.g., arrow ahead, bicycle lane to the left, etc.).”)
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the vehicle detection system of Ma et al. in combination with Zou et al., Singh et al., Fridman, and Lindemann to incorporate multiple cameras including a rearward camera as taught by Zou et al. for the purpose of allowing the vehicle to detect features of the environment on all sides, increasing coverage of features. Further, It would have been obvious to one of ordinary skill in the art, when in the combination, to perform detecting features in 2D using semantic segmentation as taught by Ma et al. with the second image captured by the rearward facing camera of Zou et al.
n. Regarding claim 147, Ma et al. in combination with Zou et al., Singh et al., Fridman, and Lindemann discloses The system of claim 79,
Zou et al. teaches wherein the first plurality of other vehicles and the second plurality of other vehicles share at least one common vehicle. ([0019] “To detect road features, the present techniques implement a deep learning network to enable multiple road feature detection and classification in parallel as one step and in real-time. In some examples, the road features can be fused with other in-vehicle sensors/data (e.g., long range sensors, other cameras, LIDAR sensors, maps, etc.) to improve detection and classification accuracy and robustness. In additional examples, the road features can be used for self-mapping and crowdsourcing to generate and/or update a road feature database.”, [0023] “a vehicle 100 including a processing system 110 for road feature detection, according to aspects of the present disclosure. In addition to the processing system 110, the vehicle 100 includes a display 120, a sensor suite 122, and cameras 130a, 130b, 130c, 130d (collectively referred to herein as “cameras 130”). The vehicle 100 can be a car, truck, van, bus, motorcycle, boat, plane, or another suitable vehicle 100.”, [0024] “The cameras 130 capture images external to the vehicle 100. Each of the cameras 130 has a field-of-view (FOV) 131a, 131b, 131c, 131d (collectively referred to herein as “FOV 131”). The FOV is the area observable by a camera. For example, the camera 130a has an FOV 131a, the camera 131b has an FOV 131b, the camera 130c has an FOV 131c, and the camera 131d has an FOV 131d. The captured images can be the entire FOV for the camera or can be a portion of the FOV of the camera.”).
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the vehicle detection system of Ma et al. in combination with Zou et al., Singh et al., Fridman, and Lindemann to incorporate multiple cameras including a rearward camera and a cloud/crowdsource system as taught by Zou et al. for the purpose of allowing vehicles to detect features of the environment on all sides as well as inform other vehicles, increasing coverage of features
o. Regarding claim 148, Ma et al. in combination with Zou et al., Singh et al., Fridman, and Lindemann discloses The system of claim 147,
Zou et al. teaches wherein the at least one common vehicle includes at least one forward facing sensor and at least one backward facing sensor. ([0019] “To detect road features, the present techniques implement a deep learning network to enable multiple road feature detection and classification in parallel as one step and in real-time. In some examples, the road features can be fused with other in-vehicle sensors/data (e.g., long range sensors, other cameras, LIDAR sensors, maps, etc.) to improve detection and classification accuracy and robustness. In additional examples, the road features can be used for self-mapping and crowdsourcing to generate and/or update a road feature database.”, [0023] “a vehicle 100 including a processing system 110 for road feature detection, according to aspects of the present disclosure. In addition to the processing system 110, the vehicle 100 includes a display 120, a sensor suite 122, and cameras 130a, 130b, 130c, 130d (collectively referred to herein as “cameras 130”). The vehicle 100 can be a car, truck, van, bus, motorcycle, boat, plane, or another suitable vehicle 100.”, [0024] “The cameras 130 capture images external to the vehicle 100. Each of the cameras 130 has a field-of-view (FOV) 131a, 131b, 131c, 131d (collectively referred to herein as “FOV 131”). The FOV is the area observable by a camera. For example, the camera 130a has an FOV 131a, the camera 131b has an FOV 131b, the camera 130c has an FOV 131c, and the camera 131d has an FOV 131d. The captured images can be the entire FOV for the camera or can be a portion of the FOV of the camera.”).
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention with reasonable expectations of success to modify the vehicle detection system of Ma et al. in combination with Zou et al., Singh et al., Fridman, and Lindemann to incorporate multiple cameras including a rearward camera and a cloud/crowdsource system as taught by Zou et al. for the purpose of allowing vehicles to detect features of the environment on all sides as well as inform other vehicles, increasing coverage of features
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MISA HUYNH NGUYEN whose telephone number is (571)270-5604. The examiner can normally be reached Monday-Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Anne Antonucci can be reached at (313) 446-6519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MISA H NGUYEN/Examiner, Art Unit 3666
/ANNE MARIE ANTONUCCI/Supervisory Patent Examiner, Art Unit 3666