DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1 thru 10 have been examined.
Information Disclosure Statement
The information disclosure statement filed 6/6/2025 regarding the List of Copending Applications fails to comply with the provisions of 37 CFR 1.97, 1.98 and MPEP § 609 because the application number entered under “U.S. Serial No.” is not a US application, but instead it is the application number of the German patent application (the priority document). It has been placed in the application file, the other documents of the IDS have been considered. Applicant is advised that the date of any re-submission of any item of information contained in this information disclosure statement or the submission of any missing element(s) will be the date of submission for purposes of determining compliance with the requirements based on the time of filing the statement, including all certification requirements for statements under 37 CFR 1.97(e). See MPEP § 609.05(a).
Specification
Applicant is reminded of the proper language and format for an abstract of the disclosure.
The abstract should be in narrative form and generally limited to a single paragraph on a separate sheet within the range of 50 to 150 words in length. The abstract should describe the disclosure sufficiently to assist readers in deciding whether there is a need for consulting the full patent text for details.
The language should be clear and concise and should not repeat information given in the title. It should avoid using phrases which can be implied, such as, “The disclosure concerns,” “The disclosure defined by this invention,” “The disclosure describes,” etc. In addition, the form and legal phraseology often used in patent claims, such as “means” and “said,” should be avoided.
The use of the term BLUETOOTH P[0042], which is a trade name or a mark used in commerce, has been noted in this application. The term should be accompanied by the generic terminology; furthermore the term should be capitalized wherever it appears or, where appropriate, include a proper symbol indicating use in commerce such as ™, SM , or ® following the term.
Although the use of trade names and marks used in commerce (i.e., trademarks, service marks, certification marks, and collective marks) are permissible in patent applications, the proprietary nature of the marks should be respected and every effort made to prevent their use in any manner which might adversely affect their validity as commercial marks.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: executing by a processing device in claim 1; capturing image data by a capture device, and generating a text description message by a visual question answering system in claims 1 and 9.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. The processing device is interpreted as a processor unit (processor circuit), and can comprise a microprocessor and/or microcontroller and/or a FPGA (Field Programmable Gate Array) and/or a DSP (Digital Signal Processor), a CPU (Central Processing Unit), a GPU (Graphical Processing Unit), or an NPU (Neural Processing Unit) that can each be used as a microprocessor P[0052]. The capture device is interpreted as a camera P[0009]. The visual question answering system is interpreted as a Neural Network and can be “a VQA system from the prior art” P[0014].
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 3 and 4 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 3 recites “a distance” in line 4, while claim 1 also recites “a distance” in line 17. It is unclear if this is a new distance or the same distance. Based on the context of claims 1 and 3, the examiner assumes it is “a particular distance” (or similar) for continued examination.
Claim 3 recites “a trajectory” in lines 4 and 5, while claim 1 also recites “a trajectory” in line 18. It is unclear if this is a new trajectory or the same trajectory. The examiner assumes it is the same trajectory for continued examination.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1 thru 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Park Patent Application Publication Number 2022/0292971 A1 in view of Wang et al German patent document DE 11 2018 000 899 T5 (translation cited).
Regarding claim 1 Park teaches the claimed method of actuation at least one system component of a system according to surroundings information generated by the system, an operation flowchart of an autonomous driving vehicle system (Figure 35), comprising:
the claimed executing a process by a processing device, “In operation S3501, when a user/driver requests to start autonomous driving, the autonomous driving vehicle system operates the vehicle in an autonomous driving mode in operation S3503.” (P[0513] and Figure 35), “the autonomous driving system 3300 according to the embodiment includes a sensor unit 3302, a vehicle operation information acquisition unit 3304, a vehicle control command generation unit 3306, a vehicle operation control unit 3308, a communication unit 3310, an AI accelerator 3312, a memory 3314, and a processor 3316” (P[0499] and Figure 33), and “The processor 3316 controls each block to provide an autonomous driving system according to the embodiment and when an autonomous driving disengagement event occurs, generates time synchronization to acquire data from each sensors of the sensor unit 3302. Further, when the autonomous driving disengagement event occurs, the processor 3316 controls to determine labels for sensor data acquired by the sensor unit 3302 and transmit the labeled data sets to the server through the communication unit 3310.” (P[0506] and Figure 33), including:
the claimed capturing image data by a capture device of the system and feeding the image data into the capture device, “The sensor unit 3302 may include a vision sensor such as a camera that uses a charge coupled device (CCD) and a complementary metal oxide semiconductor (CMOS) sensor and a non-vision sensor such as an electromagnetic sensor, an acoustic sensor, a vibration sensor, a radiation sensor, a radio wave sensor, or a thermal sensor.” (P[0500] and Figure 33), and “When occurrence of the autonomous driving disengagement event is sensed in operation S3505 (S3505—Yes), the autonomous driving vehicle system stores sensor data and vehicle operation data acquired at the time when the autonomous driving disengagement event occurs in operation S3507. At this time, in operation S3507, the autonomous driving vehicle system may store sensor data and vehicle operation data acquired for a predetermined time interval (for example, one minute or 30 seconds) including the time when the autonomous driving disengagement event occurs in response to identification of the occurrence of the autonomous driving disengagement event.” (P[0514] and Figure 35),
the claimed upon detection of a trigger, annotating position data of the system on the image data, “In operation S3509, the autonomous driving vehicle system determines whether the autonomous driving event occurred in operation S3505 corresponds to a predetermined condition, when the occurred autonomous driving event corresponds to the predetermined condition (S3509—Yes), in operation S3511, labels the stored sensor data and vehicle operation data and in operation S3513, transmits the labeled data (sensor data and vehicle operation data) to the server.” (P[0516] and Figure 35), and
the claimed generating a text description message for the image data so that the surroundings information based on surroundings of the system is generated, “in operation S3511, labels the stored sensor data and vehicle operation data and in operation S3513, transmits the labeled data (sensor data and vehicle operation data) to the server” (P[0516] and Figure 35), “When in operation S3515, it is necessary to update the autonomous driving software (S3515—Yes), the autonomous driving vehicle system downloads the updated autonomous driving software from the server in operation S3517 and updates the deep learning model with an inference model of the updated autonomous driving software in operation S3519, and then performs the autonomous driving using the updated deep learning model in operation S3521.” (P[0519] and Figure 35), and “The conditions of S3509 are conditions of the event to be labeled among the events and are necessary to update the inference model for the autonomous driving software.” P[0521],
the claimed actuating the at least one system component according to the surroundings information to generate a navigation instruction, “When in operation S3515, it is necessary to update the autonomous driving software (S3515—Yes), the autonomous driving vehicle system downloads the updated autonomous driving software from the server in operation S3517 and updates the deep learning model with an inference model of the updated autonomous driving software in operation S3519, and then performs the autonomous driving using the updated deep learning model in operation S3521.” (P[0519] and Figure 35), the performing of autonomous driving equates to the claimed generating a navigation instruction.
Park does not explicitly teach the claimed generating a text description is performed by a visual question answering system, but according to the disclosure of the applicant, “The visual question answering (VQA) system can be a VQA system from the prior art.” (applicant’s specification P[0014], PGPub P[0017]). Wang et al is one of the prior art documents provided by the applicant that uses a visual question answering (VQA) system.
Wang et al teach, a vehicle with a camera to acquire images and identify objects from the images my multi-modal fusion (translation paragraph bridging pages 2 and 3), “The fast RCNN network can take a whole picture 206 as input and create 2D feature maps using a VGG16 network. In addition, 2D suggestion boxes 232 from the 3D suggestion boxes 226 obtained by projection. Then you can get a 2D ROI pooling layer 234 fixed 2D feature vectors for each 2D suggestion box 232 extract. Next, a fully connected layer 236 the 2D feature vectors as in the 3D branch 144 smooth.” (translation page 7 paragraph 6), and “To the advantages of both inputs (point cloud 202 and picture 206 ) can use a multimodal compact bilinear (MCB) pooling layer 148 be used to combine multimodal features efficiently and expressively. The original bilinear pooling model calculated the outer product between two vectors, allowing a multiplicative interaction between all elements of both vectors. Subsequently, the original bilinear pooling model used a count-sketch projection function to reduce dimensionality and improve the efficiency of bilinear pooling. Examples of original bilinear pooling models are described in A.Fukui, DH Park, D. Yang, A. Rohrbach, T. Darrell and M. Rohrbach, "Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding" in arXiv, 2016, and Y. Gao, O. Beijbom, N. Zhang and T. Darrell, "Compact Bilinear Pooling" in CVPR, 2016, both of which are incorporated herein by reference. The original compact bilinear pooling layer was applied to a visual question and answer task by combining multimodal features from visual and textual representations. The success of multimodal compact bilinear pooling has shown its potential to handle the fusion of features from two very different domains.” (translation paragraph bridging pages 7 and 8). Additionally, the NPL of "Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding" has been provided as a reference for the “Visual Question Answering” of Wang et al.
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the operation of an autonomous driving vehicle system of Park with the visual question answering system of Wang et al in order to, with a reasonable expectation of success, more accurately capture objects of interest and estimate their orientation (Wang et al translation page 2 Background Section).
Regarding claim 2 Park teaches the claimed system is a vehicle, “the autonomous driving vehicle system operates the vehicle in an autonomous driving mode” P[0513].
Regarding claim 3 Park teaches the claimed trigger comprises an obstacle located at a distance to the system, “The event and the condition in step S3509 of FIG. 35 may be set in advance in a design/manufacturing step of the autonomous driving vehicle or at the time when the autonomous driving software is developed. For example, the event may be triggered by a situation when an unexpected obstacle appears during the autonomous driving of the vehicle” P[0520].
Regarding claim 4, the claim limitations are directed to the “time-controlled query” of claim 3. Claim 3 recites that the time-controlled query is one of five different triggers that are required to meet the claim limitations. According to the above rejection of claim 3, Park teaches the claimed trigger is an obstacle P[0520]. Because only one of the options are required in claim 3 to meet the claim limitations and the trigger being an obstacle meets this limitation, the claim 4 limitation of a time-controlled query (not selected trigger of claim 3) is not required. Therefore, Park meets the claim 4 limitation because there are no requirements for the trigger being an obstacle recited in claim 4.
Regarding claim 5 Park does not teach the claimed text description message is generated by the visual question answering system with at least one question answered that is linked to similar features from the image data, and the claimed visual question answering system is trained using a data set which contains image data or features from image data, associated questions and corresponding answers. Wang et al teach,
the claimed text description message is generated by the visual question answering system with at least one question answered that is linked to similar features from the image data, the "Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding" (hereafter Fukui et al) is incorporated by reference (translation paragraph bridging pages 7 and 8) , “The Visual Question Answering (VQA) real-image dataset (Antol et al., 2015) consists of approximately 200,000 MSCOCO images (Lin et al., 2014), with 3 questions per image and 10 answers per question. There are 3 data splits: train (80K images), validation (40K images), and test (80K images).” (Fukui et al page 461 right column first paragraph), because Fukui et al is incorporated by reference into Wang et al, it is included in the citation for Wang et al; and
the claimed visual question answering system is trained using a data set which contains image data or features from image data, associated questions and corresponding answers, “Table 4 compares our approach with the state-of-the-art on VQA test set. Our best single model uses MCB pooling with two attention maps. Additionally, we augment our training data with images and QA pairs from the Visual Genome dataset. We also concatenate the learned word embedding with pretrained GloVe vectors (Pennington et al., 2014). Each model in our ensemble of 7 models uses MCB with attention. Some of the models were trained with data from Visual Genome, and some were trained with two attention maps.” (Fukui et al page 463 right column).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the operation of an autonomous driving vehicle system of Park with the visual question answering system of Wang et al (the VQA imaging system for multiple questions related to the images of Fukui et al) in order to, with a reasonable expectation of success, more accurately capture objects of interest and estimate their orientation (Wang et al translation page 2 Background Section), and improve phrase localization accuracy, indicate better interaction between query phrase representations and visual representations of proposal bounding boxes (Fukui et al page 465 Conclusion).
Regarding claim 6 Park teaches the claimed surroundings information or position data additionally comprise one information to include at least alignment angle, speed, or planned trajectory of the system, “The vehicle operation information acquisition unit 3304 acquires operating information required to drive the vehicle such as a speed, braking, a driving direction, a turn signal, a headlight, or a steering, from an odometer or ECU of the vehicle.” (P[0501] and Figure 33), and
the claimed one information is provided automatically or manually to a user of a further system, “when the autonomous driving disengagement event occurs, the processor 3316 controls to determine labels for sensor data acquired by the sensor unit 3302 and transmit the labeled data sets to the server through the communication unit 3310” P[0506], “The communication unit 3402 communicates with a communication unit of a vehicle or infrastructures installed around the road to transmit and receive data and a training data set generation unit 3404 generates a training data set using labeled data sets acquired from the vehicle.” P[0509], and “In operation S3811, the server generates an inference model through the deep neural network model and in operation S3813, transmits the generated inference model to vehicles connected to the network.” (P[0544] and Figure 38), the vehicle connected to the network equate to the claimed further system.
Regarding claim 7 Park teaches the claimed making the text description message accessible via an edge cloud or cloud infrastructure for a further system or by direct transmission, “when the occurred autonomous driving event corresponds to the predetermined condition (S3509—Yes), in operation S3511, labels the stored sensor data and vehicle operation data and in operation S3513, transmits the labeled data (sensor data and vehicle operation data) to the server” (P[0516] and Figure 35).
Regarding claim 8 Park teaches the claimed functionality of the system is expandable by at least versioning in a backend of the system or receiving update data wirelessly, “When in operation S3515, it is necessary to update the autonomous driving software (S3515—Yes), the autonomous driving vehicle system downloads the updated autonomous driving software from the server in operation S3517 and updates the deep learning model with an inference model of the updated autonomous driving software in operation S3519, and then performs the autonomous driving using the updated deep learning model in operation S3521.” (P[0519] and Figure 35), and “The autonomous driving software updating unit 3408 releases a new version of autonomous driving software to which a newly generated inference model is reflected to store the new version of autonomous driving software in the memory 3410 and the processor 3412 transmits the new version of autonomous driving software stored in the memory 3410 to a vehicle which is connected to a network through the communication unit 3402.” P[0511].
Regarding claim 9 Park teaches the claimed system, an autonomous driving system 3300 (Figure 33), comprising:
the claimed control device comprising a processor to execute program instructions to cause the processor to execute a method, “the autonomous driving system 3300 according to the embodiment includes a sensor unit 3302, a vehicle operation information acquisition unit 3304, a vehicle control command generation unit 3306, a vehicle operation control unit 3308, a communication unit 3310, an AI accelerator 3312, a memory 3314, and a processor 3316” (P[0499] and Figure 33), “The processor 3316 controls each block to provide an autonomous driving system according to the embodiment and when an autonomous driving disengagement event occurs, generates time synchronization to acquire data from each sensors of the sensor unit 3302. Further, when the autonomous driving disengagement event occurs, the processor 3316 controls to determine labels for sensor data acquired by the sensor unit 3302 and transmit the labeled data sets to the server through the communication unit 3310.” (P[0506] and Figure 33), an operation flowchart of an autonomous driving vehicle system (Figure 35), and “In operation S3501, when a user/driver requests to start autonomous driving, the autonomous driving vehicle system operates the vehicle in an autonomous driving mode in operation S3503.” (P[0513] and Figure 35), the method including:
the claimed capturing image data by a capture device of the system and feeding the image data into the capture device, “The sensor unit 3302 may include a vision sensor such as a camera that uses a charge coupled device (CCD) and a complementary metal oxide semiconductor (CMOS) sensor and a non-vision sensor such as an electromagnetic sensor, an acoustic sensor, a vibration sensor, a radiation sensor, a radio wave sensor, or a thermal sensor.” (P[0500] and Figure 33), and “When occurrence of the autonomous driving disengagement event is sensed in operation S3505 (S3505—Yes), the autonomous driving vehicle system stores sensor data and vehicle operation data acquired at the time when the autonomous driving disengagement event occurs in operation S3507. At this time, in operation S3507, the autonomous driving vehicle system may store sensor data and vehicle operation data acquired for a predetermined time interval (for example, one minute or 30 seconds) including the time when the autonomous driving disengagement event occurs in response to identification of the occurrence of the autonomous driving disengagement event.” (P[0514] and Figure 35),
the claimed upon detection of a trigger, annotating position data of the system on the image data, “In operation S3509, the autonomous driving vehicle system determines whether the autonomous driving event occurred in operation S3505 corresponds to a predetermined condition, when the occurred autonomous driving event corresponds to the predetermined condition (S3509—Yes), in operation S3511, labels the stored sensor data and vehicle operation data and in operation S3513, transmits the labeled data (sensor data and vehicle operation data) to the server.” (P[0516] and Figure 35), and
the claimed generating a text description message for the image data so that the surroundings information based on surroundings of the system is generated, “in operation S3511, labels the stored sensor data and vehicle operation data and in operation S3513, transmits the labeled data (sensor data and vehicle operation data) to the server” (P[0516] and Figure 35), “When in operation S3515, it is necessary to update the autonomous driving software (S3515—Yes), the autonomous driving vehicle system downloads the updated autonomous driving software from the server in operation S3517 and updates the deep learning model with an inference model of the updated autonomous driving software in operation S3519, and then performs the autonomous driving using the updated deep learning model in operation S3521.” (P[0519] and Figure 35), and “The conditions of S3509 are conditions of the event to be labeled among the events and are necessary to update the inference model for the autonomous driving software.” P[0521],
the claimed actuating the at least one system component according to the surroundings information to generate a navigation instruction, “When in operation S3515, it is necessary to update the autonomous driving software (S3515—Yes), the autonomous driving vehicle system downloads the updated autonomous driving software from the server in operation S3517 and updates the deep learning model with an inference model of the updated autonomous driving software in operation S3519, and then performs the autonomous driving using the updated deep learning model in operation S3521.” (P[0519] and Figure 35), the performing of autonomous driving equates to the claimed generating a navigation instruction.
Park does not explicitly teach the claimed generating a text description is performed by a visual question answering system, but according to the disclosure of the applicant, “The visual question answering (VQA) system can be a VQA system from the prior art.” (applicant’s specification P[0014], PGPub P[0017]). Wang et al is one of the prior art documents provided by the applicant that uses a visual question answering (VQA) system.
Wang et al teach, a vehicle with a camera to acquire images and identify objects from the images my multi-modal fusion (translation paragraph bridging pages 2 and 3), “The fast RCNN network can take a whole picture 206 as input and create 2D feature maps using a VGG16 network. In addition, 2D suggestion boxes 232 from the 3D suggestion boxes 226 obtained by projection. Then you can get a 2D ROI pooling layer 234 fixed 2D feature vectors for each 2D suggestion box 232 extract. Next, a fully connected layer 236 the 2D feature vectors as in the 3D branch 144 smooth.” (translation page 7 paragraph 6), and “To the advantages of both inputs (point cloud 202 and picture 206 ) can use a multimodal compact bilinear (MCB) pooling layer 148 be used to combine multimodal features efficiently and expressively. The original bilinear pooling model calculated the outer product between two vectors, allowing a multiplicative interaction between all elements of both vectors. Subsequently, the original bilinear pooling model used a count-sketch projection function to reduce dimensionality and improve the efficiency of bilinear pooling. Examples of original bilinear pooling models are described in A.Fukui, DH Park, D. Yang, A. Rohrbach, T. Darrell and M. Rohrbach, "Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding" in arXiv, 2016, and Y. Gao, O. Beijbom, N. Zhang and T. Darrell, "Compact Bilinear Pooling" in CVPR, 2016, both of which are incorporated herein by reference. The original compact bilinear pooling layer was applied to a visual question and answer task by combining multimodal features from visual and textual representations. The success of multimodal compact bilinear pooling has shown its potential to handle the fusion of features from two very different domains.” (translation paragraph bridging pages 7 and 8). Additionally, the NPL of "Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding" has been provided as a reference for the “Visual Question Answering” of Wang et al.
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the operation of an autonomous driving vehicle system of Park with the visual question answering system of Wang et al in order to, with a reasonable expectation of success, more accurately capture objects of interest and estimate their orientation (Wang et al translation page 2 Background Section).
Regarding claim 10 Park teaches the claimed motor vehicle having the system, “the autonomous driving vehicle system operates the vehicle in an autonomous driving mode” P[0513].
Related Art
The examiner points to Zhu et al Patent Number 9,221,396 B1 as related art, but not relied upon for any rejection. Zhu et al is directed to: “FIG. 12 illustrates an example of logic flow 1202 for cross-validating a sensor of the autonomous vehicle 102. Initially, one or more sensors 146 of the autonomous vehicle may capture one or more images of the driving environment (Block 1204). As discussed with reference to FIG. 7 and FIG. 8, these images may include camera images and/or laser point cloud images.”, “The autonomous driving computer system 144 may then select which sensor to use a reference sensor (e.g., one or more of the cameras 314/316) and which sensor to cross-validate (e.g., the laser 304). The autonomous driving computer system 144 may then determine whether, and which, objects appearing in one or more images captured by the reference sensor are in one or more images captured by the sensor to be cross-validated (Block 1206). The autonomous driving computer system 144 may then apply object labels to the objects appearing in both sets of images (Block 1208). Alternatively, the object labels may be applied to objects in the first and second set of images, and then the autonomous driving computer system 144 may determine which objects the first and second set of images have in common.”, “The autonomous driving computer system 144 may then determine state information from the one or more common objects appearing in the one or more images captured by the reference sensor and determine state information for the same objects appearing in the one or more images captured by the sensor being cross-validated (Block 1210). As previously discussed, the state information may include position (relative to the autonomous vehicle 102 and/or driving environment), changes in position, speed (also relative to the autonomous vehicle 102 and/or driving environment), changes in speed, heading, and other such state information.”, “The autonomous driving computer system 144 may then compare the state information for the common objects to determine one or more deviation values (Block 1212). As previously discussed, the deviation values may include deviation values specific to individual state characteristics (e.g., a positional deviation value, a speed deviation value, etc.) but may also include deviation values determined in the aggregate (e.g., a maximum deviation value, an average deviation value, and an average signed deviation value).”, “Alternatively, or in addition, the autonomous driving computer system 144 may compare the object labels applied to the common objects (Block 1214). Comparing the object labels may include comparing the object labels from images captured at approximately the same time and/or comparing the object labels from one or more images captured during a predetermined time interval. As with the comparison of the state information, the autonomous driving computer system 144 may determine various deviation values based on the comparison of the object labels.”, and “The autonomous driving computer system 144 may then determine whether the sensor being cross-validated is having a problem (Block 1216).” (the method of Figure 12, emphasis added).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DALE W HILGENDORF whose telephone number is (571)272-9635. The examiner can normally be reached Monday - Friday 9-5:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jelani Smith can be reached at 571-270-3969. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DALE W HILGENDORF/Primary Examiner, Art Unit 3662