Prosecution Insights
Last updated: August 17, 2026
Application No. 18/791,977

ENVIRONMENTAL TEXT PERCEPTION AND PARKING EVALUATION USING VISION LANGUAGE MODELS

Final Rejection §101§103
Filed
Aug 01, 2024
Priority
Mar 18, 2024 — provisional 63/566,731 +1 more
Examiner
MCCLEARY, CAITLIN RENEE
Art Unit
3669
Tech Center
3600 — Transportation & Electronic Commerce
Assignee
NVIDIA Corporation
OA Round
2 (Final)
59%
Grant Probability
Moderate
3-4
OA Rounds
10m
Est. Remaining
84%
With Interview

Examiner Intelligence

Grants 59% of resolved cases
59%
Career Allowance Rate
73 granted / 123 resolved
+7.3% vs TC avg
Strong +25% interview lift
Without
With
+25.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
40 currently pending
Career history
165
Total Applications
across all art units

Statute-Specific Performance

§101
12.6%
-27.4% vs TC avg
§103
44.0%
+4.0% vs TC avg
§102
13.6%
-26.4% vs TC avg
§112
28.4%
-11.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 123 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are currently pending. Claims 1, 8, 12-13, and 18-20 have been amended. No claims have been cancelled or newly added. Accordingly, claims 1-20 remain pending and have been examined in this application. Examiner's Note Examiner has cited particular paragraphs/columns and line numbers or figures in the references as applied to the claims below for the convenience of the applicant. Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested from the applicant, in preparing the responses, to fully consider the references in their entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the examiner. Applicant is reminded that the Examiner is entitled to give the broadest reasonable interpretation to the language of the claims. Furthermore, the Examiner is not limited to Applicant's definition which is not specifically set forth in the disclosure. Claim Interpretation Use of the word "means" ( or "step for") in a claim with functional language creates a rebuttable presumption that the claim element is to be treated in accordance with 35 U.S.C. 112(-f) (pre-AIA 35 U.S.C. 112, sixth paragraph). The presumption that 35 U.S.C. 112(-f) (pre- AIA 35 U.S.C. 112, sixth paragraph) is invoked is rebutted when the function is recited with sufficient structure, material, or acts within the claim itself to entirely perform the recited function. Absence of the word "means" ( or "step for") in a claim creates a rebuttable presumption that the claim element is not to be treated in accordance with 35 U.S.C. 112(-f) (pre-AIA 35 U.S.C. 112, sixth paragraph). The presumption that 35 U.S.C. 112(-f) (pre-AIA 35 U.S.C. 112, sixth paragraph) is not invoked is rebutted when the claim element recites function but fails to recite sufficiently definite structure, material or acts to perform that function. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre- AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “a control system” in claim 20, “a perception system” in claim 20, each instance of “a system for” in claim 20, each instance of “a system implementing” in claim 20, and each instance of a system incorporating” in claim 20. Claim 20 recites “wherein the method is performed by at least one of” and therefore links the functions performed in claim 19 to the options listed in claim 20. The only limitations in claim 20 that don’t meet the 3-prong requirement of 35 U.S.C. 112(f) are “a system implemented using an edge device”, “a system implemented using a robot”, “a system implemented at least partially in a data center”, and “a system implemented at least partially using cloud computing resources.” Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The above-referenced claim limitations has/have been interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because: “a control system” in claim 20, “a perception system” in claim 20, each instance of “a system for” in claim 20, each instance of “a system implementing” all use a generic placeholder “system” coupled with functional language without reciting sufficient structure to achieve the function. Furthermore, the generic placeholder is not preceded by a structural modifier. Since the claim limitation(s) invokes 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, the claims have been interpreted to cover the corresponding structure described in the specification that achieves the claimed function, and equivalents thereof. A review of the specification shows that the following appears to be the corresponding structure described in the specification for the 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph limitation: Control system: [0048, 0099] Perception system: [0049] Systems: [0048, 0099] For all the units corresponding to a computer (hardware) the software (steps in an algorithm/flowchart) should be included to indicate proper support. If applicant wishes to provide further explanation or dispute the examiner's interpretation of the corresponding structure, applicant must identify the corresponding structure with reference to the specification by page and line number, and to the drawing, if any, by reference characters in response to this Office action. If applicant does not intend to have the claim limitation(s) treated under 35 U.S.C. l 12(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may amend the claim(s) so that it/they will clearly not invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, or present a sufficient showing that the claim recites/recite sufficient structure, material, or acts for performing the claimed function to preclude application of 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. For more information, see MPEP § 2173 et seq. and Supplementary Examination Guidelines for Determining Compliance With 35 U.S. C. 112 and for Treatment of Related Issues in Patent Applications, 76 FR 7162, 7167 (Feb. 9, 2011). Claim Rejections - 35 USC § 101 Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims are either directed to one or more processors, a system, or a method, which are each one of the statutory categories of invention. (Step 1: YES) The examiner has identified claim 1 as the claim that represents the claimed invention for analysis. Claim 1 recites the limitations of: “One or more processors comprising processing circuitry to: identify image data generated using one or more cameras of an ego-machine and representing one or more parking signs; apply, to a vision-language model (VLM), a multimodal prompt comprising the image data representing the one or more parking signs and a text prompt instructing the VLM to generate one or more responses determining whether parking is permitted in one or more candidate parking spaces based at least on the image data; and control, using an Advanced Driver Assistance System (ADAS) of the ego-machine, one or more parking operations of the ego-machine with respect to at least one candidate parking space of the one or more candidate parking spaces based at least on the one or more responses.” Independent claim 1 includes limitations that recite an abstract idea. The examiner submits that the foregoing limitation(s) constitute a “mental process” because under its broadest reasonable interpretation, the claim covers performance of the limitation in the human mind. For example, “identify image data generated using one or more cameras of an ego-machine and representing one or more parking signs” and “applying a multimodal prompt comprising the image data representing the one or more parking signs and a text prompt instructing to generate one or more responses determining whether parking is permitted in one or more candidate parking spaces based at least on the image data” in the context of this claim encompasses a person looking at images of parking signs which have been captured by a camera, receiving a written or verbal question (“Can the vehicle park here?”), determining if parking in the space is permitted or not based on the parking sign in the image data, and making a decision about whether to park in the space or not. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “mental processes” grouping of abstract ideas. (Step2A-Prong 1: YES. The claims are abstract) This judicial exception is not integrated into a practical application. Limitations that are not indicative of integration into a practical application include: (1) Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05.f), (2) Adding insignificant extra-solution activity to the judicial exception (MPEP 2106.05.g), (3) Generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05.h). In particular, the claims recite additional elements of using one or more processors comprising processing circuitry to perform the recited steps. The one or more processors are recited at a high-level of generality (i.e., as generic processors performing generic computer functions) such that it amounts to no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims recite additional elements of using a vision language model (VLM). The VLM is used to generally apply the abstract idea without placing any limits on how the VLM functions. The recitation of a vision language model (VLM) merely uses the VLM as a tool to perform the abstract idea, see MPEP 2106.05(f). The VLM limitation merely confines the use of the abstract idea to a particular technological environment (language models) and thus fails to add an inventive concept to the claims. See MPEP 2106.05(h). The limitation of controlling, using an ADAS, one or more parking operations based on the one or more responses is also considered an additional element. The ADAS is recited at a high-level of generality, and thus is considered insignificant extra-solution activity. The step of controlling one or more parking operations, under broadest reasonable interpretation (see claim 9 for example) is considered a visual or audible output (i.e., a generic output in response to the generating step) and is therefore considered insignificant post-solution activity. Thus, taken alone, the additional elements do not integrate the abstract idea into a practical application. Further, looking at the additional limitation(s) as an ordered combination or as a whole, the limitation(s) add nothing that is not already present when looking at the elements taken individually. For instance, there is no indication that the additional elements, when considered as a whole, reflect an improvement in the functioning of a computer or an improvement to another technology or technical field, apply or use the above-noted judicial exception to effect a particular treatment or prophylaxis for a disease or medical condition, implement/use the above-noted judicial exception with a particular machine or manufacture that idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. (Step 2A-Prong 2: NO. The additional claimed elements are not integrated into a practical application) The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered separately and as an ordered combination, they do not add significantly more (also known as an "inventive concept") to the exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using at least one processor amounts to nothing more than applying the exception using a generic computer component. The additional element of using a VLM also amounts to nothing more than applying the exception using a generic VLM. Generally applying an exception using a generic computer component or generic VLM cannot provide an inventive concept. And as discussed above, the examiner submits that the additional limitations are insignificant extra-solution activities. Further, a conclusion that an additional element is insignificant extra-solution activity in Step 2A should be re-evaluated in Step 2B to determine if they are more than what is well-understood, routine, conventional activity in the field. Hence, the claim is not patent eligible. Thus claim 1 (and similarly claims 13 and 19) is not patent eligible. (Step 2B: NO. The claims do not provide significantly more) Claims 2-12, 14-18, and 20 further define the abstract idea that is present in their respective independent claims and hence are abstract for at least the reasons presented above. The dependent claims do not include any additional elements that integrate the abstract idea into a practical application or are sufficient to amount to significantly more than the judicial exception when considered both individually and as an ordered combination. Therefore, the dependent claims are directed to an abstract idea. Thus, the aforementioned claims are not patent-eligible. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-4, 12-16, and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Yalla (US 2018/0321685 A1) in view of Wen (On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving). Regarding claim 1, Yalla discloses one or more processors comprising processing circuitry (see at least [0019] – controller 150 may be implemented as one or more processors) to: identify image data generated using one or more cameras of an ego-machine and representing one or more parking signs (see at least [0014, 0027-0029] – sensors 212 may use digital photographic imaging, computer vision, and the like to determine the position and/or identities of objects relative to autonomous vehicle 204... detect sign 208); apply a prompt comprising the image data to generate one or more responses determining whether parking is permitted in one or more candidate parking spaces based at least on the image data (see at least [0027-0029, 0031] - sensors 212 may use digital photographic imaging, computer vision, and the like to determine the position and/or identities of objects relative to autonomous vehicle 204... detect sign 208… determine whether or not the autonomous vehicle 204 may be permissioned to park in parking place 206); and control, using an Advanced Driver Assistance System (ADAS) of the ego-machine, one or more parking operations of the ego-machine with respect to at least one candidate parking space of the one or more candidate parking spaces based at least on the one or more responses (see at least [0029, 0031] – when the autonomous vehicle 204 may be permissioned to park in parking place 206 the vehicle control unit 220 initiates a command to position autonomous vehicle 204 in parking place 206… when autonomous vehicle 204 is not permissioned to park in parking place 206, vehicle-parking controller 156 can then initiate a command to eliminate parking place 206 as a parking option and to continue to search for a permissioned parking place). Yalla does not appear to explicitly disclose apply, to a vision-language model (VLM), a multimodal prompt comprising the image data representing the one or more parking signs and a text prompt instructing the VLM to generate one or more responses. Wen, in the same field of endeavor, teaches the following limitations: apply, to a vision-language model (VLM), a multimodal prompt comprising the image data representing the one or more signs and a text prompt instructing the VLM to generate one or more responses (see at least page 7 section 2.1, page 10 section 2.1, page 36 section 4.5 – GPT-4V possesses a commendable ability to recognize traffic signs… see the prompt in Figure 6 including the front-camera view and the response from GPT-4V… Through the aforementioned five tests, it is observed that GPT-4V has initially acquired decision making abilities similar to human drivers. It can combine the states of various traffic elements (e.g., pedestrians, vehicles, traffic lights, road signs, lanes) to provide the final driving strategy. Besides, GPT-4V can make reasonable decisions in diverse driving scenarios such as parking lots, intersections, highways, and ramps. Overall, GPT-4V demonstrates strong adherence to rules and safety awareness with relatively conservative driving strategies.). It would have been obvious to one of ordinary skill in the art before the effective filing date to have incorporated the teachings of Wen into the invention of Yalla with a reasonable expectation of success. The motivation of doing so is that GPT-4V demonstrates superior performance in scene understanding and casual reasoning compared to existing autonomous systems, with the potential to handle out-of-distribution scenarios, recognize intentions, and make informed decisions in real driving contexts (Wen - abstract). In particular, GPT-4V possess a commendable ability to recognize traffic signs, can make reasonable decisions in diverse driving scenarios such as parking lots, intersections, highways, and ramps, and demonstrates strong adherence to rules and safety awareness with relatively conservative driving strategies (Wen – page 7 section 2.1, page 36 section 4.5). Regarding claim 2, Yalla discloses wherein the processing circuitry is further to initiate monitoring for the one or more parking signs based at least on the ego-machine entering a detected parking mode (see at least [0020-0025] - Autonomous vehicle 204 may schedule routing to arrive at City Hall at 1:30 PM to initiate a search for parking within a range of distance from which the passenger may walk (e.g., at a predicted rate of walking speeds) from a parking space to the destination. Upon arrival at City Hall, autonomous vehicle 204 can initiate a search for available parking beginning at the entrance of City Hall and continuing along the perimeter of the nearest city block until an available parking place is determined.). Regarding claim 3, Yalla discloses wherein the processing circuitry is further to detect a parking domain of the ego-machine by performing at least one of (BRI requires only one of the following): using a mapping application (see at least [0020-0025, 0033] - Autonomous vehicle 204 may schedule routing to arrive at City Hall at 1:30 PM to initiate a search for parking within a range of distance from which the passenger may walk (e.g., at a predicted rate of walking speeds) from a parking space to the destination. Upon arrival at City Hall, autonomous vehicle 204 can initiate a search for available parking beginning at the entrance of City Hall and continuing along the perimeter of the nearest city block until an available parking place is determined… Utilizing computer vision, digital photographic imaging, text recognition, onboard database(s) and systems, external database(s), and/or GPS location, autonomous vehicle 204 can identify one or more sign details 210 related to multiple signs to determine permissioned parking relative to multiple classes of restricted parking. Thus, based on sign details 210, Saturday at 11:00 AM would be a permissioned time to park in parking place 206. However, if a temporary additional sign (not depicted), such as a construction sign (e.g., “NO PARKING on Tuesday”), is present adjacent to parking place 206, one or more sensors 212 can detect sign details 210 to provide permission and/or restriction data related to parking place 206. Based on sign details 210 of at least one sign 208, vehicle control unit 220 will initiate a command to park autonomous vehicle 204 in parking place 206 or to initiate a command to continue to search for permissioned parking.), or prompting the VLM to detect the parking domain based at least on one or more frames comprising at least some of the image data. Regarding claim 4, Yalla discloses wherein the processing circuitry is further to detect the one or more parking signs based at least on detecting one or more classes of parking signs associated with a detected parking domain of the ego-machine (see at least [0033] - Utilizing computer vision, digital photographic imaging, text recognition, onboard database(s) and systems, external database(s), and/or GPS location, autonomous vehicle 204 can identify one or more sign details 210 related to multiple signs to determine permissioned parking relative to multiple classes of restricted parking. Thus, based on sign details 210, Saturday at 11:00 AM would be a permissioned time to park in parking place 206. However, if a temporary additional sign (not depicted), such as a construction sign (e.g., “NO PARKING on Tuesday”), is present adjacent to parking place 206, one or more sensors 212 can detect sign details 210 to provide permission and/or restriction data related to parking place 206. Based on sign details 210 of at least one sign 208, vehicle control unit 220 will initiate a command to park autonomous vehicle 204 in parking place 206 or to initiate a command to continue to search for permissioned parking.). Regarding claim 12, Yalla discloses wherein the one or more processors are comprised in at least one of (BRI requires only one of the following): a control system for an autonomous or semi-autonomous machine (see at least [0014, 0019] - autonomy controller 150 including a sensor fusion module 154, an ancillary sensor manager 110, a vehicle-parking controller 156, and a vehicle control unit 113); a perception system for an autonomous or semi-autonomous machine (see at least [0014, 0019] - autonomy controller 150 including a sensor fusion module 154, an ancillary sensor manager 110, a vehicle-parking controller 156, and a vehicle control unit 113); a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational Al operations; a system for generating synthetic data; a system for generating synthetic data using Al; a system for performing one or more generative Al operations; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources. Wen, in the same field of endeavor, also teaches the following limitations: a system implementing one or more vision language models (VLMs) (see at least abstract – VLM, GPT-4V). The motivation to combine Yalla and Wen are the same as in the rejection of claim 1 above. Regarding claims 13 and 19, all the limitations have been analyzed in view of claim 1, and it has been determined that claims 13 and 19 do not teach or define any new limitations beyond those previously recited in claim 1; therefore, claims 13 and 19 are also rejected over the same rationale as claim 1. Regarding claim 14, all the limitations have been analyzed in view of claim 2, and it has been determined that claim 14 does not teach or define any new limitations beyond those previously recited in claim 2; therefore, claim 14 is also rejected over the same rationale as claim 2. Regarding claim 15, all the limitations have been analyzed in view of claim 3, and it has been determined that claim 15 does not teach or define any new limitations beyond those previously recited in claim 3; therefore, claim 15 is also rejected over the same rationale as claim 3. Regarding claim 16, all the limitations have been analyzed in view of claim 4, and it has been determined that claim 16 does not teach or define any new limitations beyond those previously recited in claim 4; therefore, claim 16 is also rejected over the same rationale as claim 4. Regarding claims 18 and 20, all the limitations have been analyzed in view of claim 12, and it has been determined that claims 18 and 20 do not teach or define any new limitations beyond those previously recited in claim 12; therefore, claims 18 and 20 are also rejected over the same rationale as claim 12. Claims 5-7 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Yalla in view of Wen and Chen (CN 116052182 A, a machine translation is attached and being relied upon). Regarding claim 5, Yalla does not appear to explicitly disclose wherein the processing circuitry is further to verify legibility of the one or more parking signs based at least on one or more detected regions of interest representing the one or more parking signs. However, Yalla does disclose wherein the processing circuitry is further to identify the one or more parking signs based at least on one or more detected regions of interest representing the one or more parking signs (see at least [0027-0029, 0031] - sensors 212 may use digital photographic imaging, computer vision, and the like to determine the position and/or identities of objects relative to autonomous vehicle 204... detect sign 208… determine whether or not the autonomous vehicle 204 may be permissioned to park in parking place 206). Chen, in the same field of endeavor, teaches the following limitations: wherein the processing circuitry is further to verify legibility of the one or more detected regions of interest (see at least [0048-0062] – text detection performed on a target object in the current video frame to obtain text detection information… based on the text detection confidence score of the text region to be identified, a quality score for the text region to be identified is determined). It would have been obvious to one of ordinary skill in the art before the effective filing date to have incorporated the teachings of Chen into the invention of Yalla with a reasonable expectation of success for the purpose of reducing the amount of time required to recognize text in video frames in the field of transportation (Chen – [0002-0004, 0034]). Regarding claim 6, Yalla does not appear to explicitly disclose wherein the processing circuitry is further to prompt the VLM to verify legibility of the one or more parking signs. However, Yalla does disclose wherein the processing circuitry is further to prompt to identify the one or more parking signs (see at least [0027-0029, 0031] - sensors 212 may use digital photographic imaging, computer vision, and the like to determine the position and/or identities of objects relative to autonomous vehicle 204... detect sign 208… determine whether or not the autonomous vehicle 204 may be permissioned to park in parking place 206). Wen, in the same field of endeavor, also teaches the following limitations: wherein the processing circuitry is further to prompt the VLM to identify the one or more signs (see at least page 7 section 2.1 – GPT-4V possesses a commendable ability to recognize traffic signs). The motivation to combine Yalla and Wen are the same as in the rejection of claim 1 above. Chen, in the same field of endeavor, teaches the following limitations: wherein the processing circuitry is further to prompt to verify legibility of one or more detected regions of interest (see at least [0048-0062] – text detection performed on a target object in the current video frame to obtain text detection information… based on the text detection confidence score of the text region to be identified, a quality score for the text region to be identified is determined). The motivation to combine Yalla and Chen are the same as in the rejection of claim 5 above. Regarding claim 7, Yalla does not appear to explicitly disclose wherein the processing circuitry is further to: cache the image data representing the one or more parking signs based at least on verifying legibility of the one or more parking signs, and prompt the VLM to evaluate the cached image data in response to detecting the one or more candidate parking spaces. However, Yalla does disclose wherein the processing circuitry is further to: prompt to evaluate the image data in response to detecting the one or more candidate parking spaces (see at least [0025-0034] - In an embodiment, autonomous vehicle 204 utilizes one or more sensors 212 to determine a permissioned parking place 206 relative to multiple classes of restricted parking. Examples of computations or logic that can be used to determine permissions for a parking place include cross-referencing databases of information related to types of parking at various locations. For instance, a central database of disabled person parking locations and their respective GPS positions may be analyzed with data received from one or more sensors 212 to confirm that a parking place is reserved for disabled persons. Examples of restricted parking may be: disabled person parking, commercial parking only, time-restricted parking, loading zone only, employee only, customer parking only, residents only, and the like. Examples of permissioned parking can be: a parking place not subject to a parking restriction, paid parking, parking lots, parking garages, reserved parking for a passenger who possesses a permission to park such as an identifying placard, sign, badge, identification card, sticker, etc. Likewise, various embodiments can employ similar technology for use on city streets, in parking garages, in parking lots, etc.). Wen, in the same field of endeavor, teaches the following limitations: wherein the processing circuitry is further to prompt the VLM to evaluate the image data (see at least page 7 section 2.1, page 10 section 2.1 – GPT-4V possesses a commendable ability to recognize traffic signs… see the prompt in Figure 6 including the front-camera view and the response from GPT-4V). The motivation to combine Yalla and Wen are the same as in the rejection of claim 1 above. Chen, in the same field of endeavor, teaches the following limitations: wherein the processing circuitry is further to: cache the image data representing one or more detected regions of interest based at least on verifying legibility of the one or more detected regions of interest (see at least [0063-0068] – selecting the text region to be identified corresponding to the highest quality score in the current video frame…text region to be recognized corresponding to the target object in the current video frame is stored into the cached image sequence to obtain the updated cached image sequence), and prompt to evaluate the cached image data (see at least [0069-0073] – video frame corresponding to the highest quality score in the updated cached image sequence is selected and input into the text recognition model for text recognition). The motivation to combine Yalla and Chen are the same as in the rejection of claim 5 above. Regarding claim 17, all the limitations have been analyzed in view of claim 5, and it has been determined that claim 17 does not teach or define any new limitations beyond those previously recited in claim 5; therefore, claim 17 is also rejected over the same rationale as claim 5. Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Yalla in view of Wen and Borras (US 2021/0183169 A1). Regarding claim 8, Yalla does not appear to explicitly disclose wherein the processing circuitry is further to include, in the multi-modal prompt to the VLM a representation of one or more characteristics of one or more geo-tagged parking permits. Wen, in the same field of endeavor, teaches the following limitations: wherein the processing circuitry is further to include, in the multi-modal prompt to the VLM a representation of one or more characteristics (see at least page 7 section 2.1, page 10 section 2.1 – GPT-4V possesses a commendable ability to recognize traffic signs… see the prompt in Figure 6 including the front-camera view and the response from GPT-4V). The motivation to combine Yalla and Wen are the same as in the rejection of claim 1 above. Borras, in the same field of endeavor, teaches the following limitations: evaluate whether parking is permitted in the one or more candidate parking spaces based at least on one or more geo-tagged parking permits (see at least Figs. 10, 14, [0047, 0078-0079, 0082-0083] – determine whether the vehicle is parked in parking space based on geo-location coordinates… user has an active parking account… an indicator can be activated in the vehicle or the cellular phone device to indicate that the vehicle is properly parked and payment will be accepted to avoid a parking violation). It would have been obvious to one of ordinary skill in the art before the effective filing date to have incorporated the teachings of Borras into the invention of Yalla with a reasonable expectation of success. The motivation of doing so is to derive high geo-location accuracy determination for dynamically defined tolling lanes and parking spaces for mobile payments (Borras – [0002]). This method would save vehicle occupants time in evaluating whether or not they can park in a space. In autonomous vehicle applications the vehicle needs to be able to navigate complex situations while making dynamic decisions, including dynamic parking decisions. Integrating the determination of whether parking is permitted using the geo-tagged parking permits as taught by Borras into the VLM prompt as taught by Wen is considered obvious because it is applying a known technique (Wen’s GPT-4V) to a particular application of determining whether a vehicle can park based on geo-tagged parking permits (as taught by Borras). This would allow autonomous vehicles that are integrating new technology such as GPT-4V to utilize geo-tagged parking permits. Claims 9-11 are rejected under 35 U.S.C. 103 as being unpatentable over Yalla in view of Wen and Malczyk (US 2020/0278218 A1). Regarding claim 9, Yalla does not appear to explicitly disclose wherein the one or more parking operations of the ego-machine comprise outputting at least one of a visual or an audible representation of whether parking is permitted in the one or more candidate parking spaces. Malczyk, in the same field of endeavor, teaches the following limitations: wherein the one or more parking operations of the ego-machine comprise outputting at least one of a visual or an audible representation of whether parking is permitted in the one or more candidate parking spaces (see at least Fig. 3B, [0041-0042] – display an icon for the parking bay in which parking is or is not permitted). It would have been obvious to one of ordinary skill in the art before the effective filing date to have incorporated the teachings of Malczyk into the invention of Yalla with a reasonable expectation of success for the purpose of enabling the user to easily obtain parking information, such as whether parking is permitted in certain bays and also the price of parking in a bay (Malczyk – [0017, 0041-0042]). Regarding claim 10, Yalla does not appear to explicitly disclose wherein the processing circuitry is further to prompt the VLM to determine a cost to park in the one or more candidate parking spaces for a designated duration of time. Wen, in the same field of endeavor, teaches the following limitations: wherein the processing circuitry is further to prompt the VLM (see at least page 7 section 2.1, page 10 section 2.1 – GPT-4V possesses a commendable ability to recognize traffic signs… see the prompt in Figure 6 including the front-camera view and the response from GPT-4V). The motivation to combine Yalla and Wen are the same as in the rejection of claim 1 above. Malczyk, in the same field of endeavor, teaches the following limitations: wherein the processing circuitry is further to determine a cost to park in the one or more candidate parking spaces for a designated duration of time (see at least Fig. 3B, [0042] – display the price of parking in the bay). The motivation to combine Yalla and Malczyk is the same as in the rejection of claim 9 above. Regarding claim 11, Yalla does not appear to explicitly disclose wherein the processing circuitry is further to output at least one of a visual or an audible representation of a cost to park in the one or more candidate parking spaces determined using the VLM. Wen, in the same field of endeavor, teaches the following limitations: the VLM (see at least abstract – VLM, GPT-4V). The motivation to combine Yalla and Wen are the same as in the rejection of claim 1 above. Malczyk, in the same field of endeavor, teaches the following limitations: wherein the processing circuitry is further to output at least one of a visual or an audible representation of a cost to park in the one or more candidate parking spaces determined (see at least Fig. 3B, [0042] – display the price of parking in the bay). The motivation to combine Yalla and Malczyk is the same as in the rejection of claim 9 above. Response to Arguments As tentatively agreed upon during the interview on June 1, 2026, the amendments appeared to overcome the previous prior art rejections and they have been withdrawn. As noted during the interview, further search and consideration was necessary to determine allowability. The previous prior art rejections have been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Wen. Applicant’s arguments, see pages 19-21 filed 6/11/2026, with respect to 35 U.S.C. 101 rejections have been fully considered but they are not persuasive. Applicant argues that the Office Action interprets that a human can make the determination, but that is unreasonable as the claim expressly requires the VLM to make the determination. The recitation of a vision language model (VLM) merely uses the VLM as a tool to perform the abstract idea, see MPEP 2106.05(f). The VLM limitation merely confines the use of the abstract idea to a particular technological environment (language models) and thus fails to add an inventive concept to the claims. See MPEP 2106.05(h). Applicant argues the controlling step cannot be made in the human mind. This step was not interpreted as part of the abstract idea. Applicant argues that the claimed technique is an improvement in the field of autonomous and semi-autonomous machines and therefore integrate any abstract idea into a practical application. The examiner respectfully disagrees. The claims are implementing one or more processors and a VLM as a tool into a particular application. Furthermore, the claims are not directed to an improvement in the way computers operate. While the claimed limitations certainly purport to present effective means for determining whether parking is permitted based on parking signs, the efficiency comes from the use of the one or more processors and the VLM. The fact that the determination could be performed more efficiently via the one or more processors and the VLM does not materially alter the patent eligibility of the claimed subject matter. The focus of the claims is not on an improvement as computers or VLMs as tools, but on certain independently abstract ideas that use computers and VLMs as tools. Applicant’s arguments, see pages 22-23 filed 6/11/2026, with respect to 35 U.S.C. 112 rejections have been fully considered and are persuasive. The 35 U.S.C. 112 rejections have been withdrawn. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to CAITLIN MCCLEARY whose telephone number is (703)756-1674. The examiner can normally be reached Monday - Friday 10:00 am - 7:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Navid Z Mehdizadeh can be reached at (571) 272-7691. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CAITLIN R MCCLEARY/Examiner, Art Unit 3669 /Erin M Piateski/Supervisory Patent Examiner, Art Unit 3669
Read full office action

Prosecution Timeline

Aug 01, 2024
Application Filed
Dec 18, 2025
Non-Final Rejection mailed — §101, §103
May 15, 2026
Interview Requested
Jun 01, 2026
Applicant Interview (Telephonic)
Jun 01, 2026
Examiner Interview Summary
Jun 11, 2026
Response Filed
Aug 03, 2026
Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12703370
REST CONDUCIVE AUTOMATED VEHICLE OPERATION
2y 3m to grant Granted Aug 11, 2026
Patent 12694727
VEHICLE DATA COLLECTION SYSTEM AND METHODS
4y 4m to grant Granted Jul 28, 2026
Patent 12682761
TECHNOLOGIES FOR OPTIMAL VEHICLE PLATOONING CONTROL OVER STEEP TERRAIN
3y 2m to grant Granted Jul 14, 2026
Patent 12679416
METHOD FOR DRIVING CONTROL BASED ON BOARDING CONGESTION AND A VEHICLE USING THE SAME
2y 7m to grant Granted Jul 14, 2026
Patent 12681486
INFORMATION PROCESSING DEVICE AND MOVING OBJECT
2y 1m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
59%
Grant Probability
84%
With Interview (+25.0%)
2y 10m (~10m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 123 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month