Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
This communication is a Final Office Action in response to Applicant’s amendment for application number 18/921,846 received on 04/07/2026.
In accordance with Applicant’s amendment, claims 1-3, 5-13, 16-19, and 21-24 are amended, currently pending, and have been examined. Claims 4, 14-15, and 20 have been cancelled. Claims 21-24 have been added.
Response to Amendment
The amendment filed on 04/07/2026 has been entered.
Applicant’s amendment necessitated the new ground(s) of rejection set forth in this Office Action.
Upon review of amendments, the 112 claim interpretations previously applied are withdrawn.
Response to Arguments
Response to §101 arguments – Applicant’s arguments with respect to the §101 rejections previously applied to the claims have been considered and are unpersuasive.
Applicant argues (Remarks at pg. 12): “Initially, the Patent Office's analysis appears to improperly abstract isolated verbs such as "identify" or "generate" while ignoring the technological context in which those actions occur. When properly considered, the claims recite a machine-implemented vehicle system performing computationally intensive operations and producing machine control outputs, not a mental process within the meaning of MPEP §2106.04(a).”. In response, Examiner respectfully disagrees and notes that in performing the analysis of the claims to determine if the claims recite abstract ideas, Examiner correctly identified abstract steps. For example, as documented in at least page 6 of the office action mailed 01/07/2026, Examiner notes, regarding the limitation “a control module in communication with the conversion module, the control module configured to receive the plurality of embeddings and an input query from a user specific to a vehicle, identify vehicle information for the vehicle based on the input query, and generate a response specific to the vehicle based on the vehicle information and the plurality of embeddings in response to the input query;”, but for the additional elements recited in the claim limitations, which are analyzed under Steps 2A, Prong 2, and 2B, the steps to “receive the plurality of embeddings and an input query from a user”, “identify vehicle information”, and “generate a response” could be accomplished mentally, such as by human observation, evaluation, judgement, opinion, or with the help of pen and paper. That is, one of ordinary skill in the art would be able to receive the plurality of embeddings and an input query from a user with the help of pen and paper, and identify vehicle information via observation (for example, by looking at the VIN number), and generate a response via judgement or opinion. Therefore, the claim recites abstract steps that fall under the “Mental Processes” abstract idea grouping by setting forth activities that could be performed mentally by a human (including an observation, evaluation, judgment, opinion) (see MPEP § 2106.04(a)(2), subsection III). Therefore, when considered as a whole, but for the recitation of additional elements, the claims are directed to an abstract idea,
Applicant argues (Remarks at pg. 13): “That said, even if the claims are directed to a mental process as alleged by the Patent Office, the pending claims recite several features that integrate the cited abstract idea into a practical application and therefore patent-eligible under Step 2A, prong two. For example, the application identifies a concrete technical problem in vehicle systems. Specifically, conventional vehicle manuals and assistance tools are generic and do not reflect a vehicle's exact configuration, leading to inaccurate guidance, delayed responses, and potentially unsafe operation. See paragraph 34 of the subject application. The amended claims address this problem through a technical solution that includes generating vehicle-specific embeddings, identifying a precise vehicle configuration using a VIN, generating responses constrained by those embeddings (e.g., using a large language model), and using the resulting responses to automatically control vehicle operations. This is not a mere presentation or organization of information, but a system that improves how vehicle systems respond to queries and how vehicle operations are controlled based on vehicle-specific data.”. In response, Examiner respectfully disagrees and notes that the additional elements and the combination of additional elements recited in the claim limitations fail to integrate the abstract idea into a practical application for the reasons set forth in the updated 35 USC 101 rejections (necessitated by amendments) of the instant office action, including using generic computing components and/or instructions as a tool, reciting additional elements at a high level of generality, and performing insignificant extra-solution activity. See 101 rejections below for details.
Applicant argues (Remarks at pg. 14): “Further, even assuming, arguendo, that the pending claims are directed to an abstract idea, they surely recite something significantly more than the abstract idea itself under Step 2B.”. In response, Examiner respectfully disagrees and notes that the additional elements and the combination of additional elements recited in the claim limitations fail to add significantly more to the claims for the reasons set forth in the updated 35 USC 101 rejections (necessitated by amendments) of the instant office action, including using generic computing components and/or instructions as a tool, reciting additional elements at a high level of generality, and performing insignificant extra-solution activity. See 101 rejections below for details.
Applicant argues (Remarks at pg. 15): “Moreover, the Patent Office appears to ignore vehicle control features of the claims because they amount "to insignificant extra-solution activity." Pages 7-8 of the Office Acton. However, vehicle control cannot be dismissed as "insignificant extra-solution activity." It is the technical output of the claimed system and the mechanism by which the claims produce a real-world effect. Without vehicle control, the claimed system would not achieve its stated technical purpose. As such, the claimed vehicle control features cannot be dismissed as post solution activity.”. In response, Examiner respectfully disagrees and reminds the Applicant that, as stated in MPEP 2106.05(g), the term "extra-solution activity" can be understood as activities incidental to the primary process or product that are merely a nominal or tangential addition to the claim. Extra-solution activity includes both pre-solution and post-solution activity. Furthermore, MPEP 2106.05(g) also states “An example of post-solution activity is an element that is not integrated into the claim as a whole, e.g., a printer that is used to output a report of fraudulent transactions, which is recited in a claim to a computer programmed to analyze and manipulate information about credit card transactions in order to detect whether the transactions were fraudulent.”. Examiner’s determination of the steps to/for “control”/”controlling” the vehicle as extra-solution activity are supported by Applicant’s own claims and specification. For example, the limitation of claim 1 to “generate and transmit an electronic notification including the response for the user to a vehicle controller in the vehicle for automatically controlling, without user interaction, at least one vehicle operation of the vehicle based on the response specific to the vehicle.” only requires a notification to be generated and transmitted. Furthermore, Applicant’s claims are directed to vehicle issue resolution services, as supported by Applicant’s own specification, where in par [0002], Applicant discloses “The present disclosure relates to online vehicle issue resolution services, and more particularly to vehicle systems and control methods for vehicle query resolution services.”. Because Applicant’s own specification defines Applicant’s invention as relating to “online vehicle issue resolution services”, any claim element that is not integrated into the claim as a whole, is insignificant extra-solution activity, including “controlling at least one vehicle operation of the vehicle based on the control signal” from claim 13, and “control at least one vehicle operation of the vehicle based on the control signal” from claim 19.
Response to §103 arguments – But for the following, Applicant’s arguments with respect to the §103 rejections previously applied to the claims are moot based on the new grounds of rejections, as necessitated by the amendment.
Applicant argues (Remarks at pg. 18): “In rejecting the claims, the Patent Office appears to suggest that controlling an operation of a vehicle is met by Wolverton's disclosure of presenting directions to the nearest gas station (at paragraph 106) and retrieving information in a vehicle-specific user's guide knowledge base 140, and providing that information to a driver (at paragraph 138). Pages 21 and 26 of the Office Action. However, presenting directions or retrieving and providing information to a driver are merely actions for outputting (voice or graphics) information to the driver. Such actions are unrelated to controlling actual vehicle operations. As such, whether considered alone or in combination, Wolverton and Rempe fail to disclose or suggest (a) automatically controlling, without user interaction, at least one vehicle operation of the vehicle based on the response specific to the vehicle, as recited by amended claim 1, (b) automatically generating a control signal based on the at least one recommendation without user interaction and controlling at least one vehicle operation of the vehicle based on the based on the control signal, as recited by amended claim 13, and (c) generating a control signal based on the at least one recommendation specific to the vehicle, and controlling at least one vehicle operation of the vehicle based on the control signal, as recited by amended claim 19. For this reason, amended claims 1, 13 and 19 are patentable over Wolverton and Rempe.”. In response, Examiner respectfully disagrees and notes that, under BRI, any action performed by the system autonomously can be reasonably interpreted as “controlling at least one vehicle operation”. For Example, as noted in page 21 of the office action mailed 01/07/2026, Wolverton teaches: [0106] In response to the conditional instruction, the method 400 may track the inputs 110 relating to the vehicle 104's fuel level and notify the user of the nearest gas station when the fuel reaches a certain level. To do this, the method 400 may interpret the phrase "running out of gas" as meaning "less than 1/8 of a tank" (based, e.g., on data or rules obtained from the vehicle-specific conversation model 132, the vehicle-specific user's guide knowledge base 140, or the vehicle context model 116). The method 400 may then monitor the vehicle sensors 106 that are associated with the vehicle 104's fuel level, and monitor the vehicle 104's navigation system (or a communicatively coupled Global Positioning System or GPS device, or a smart phone map or navigation application, for example). The method 400 may then present (e.g., by voice or graphic display) directions to the nearest gas station when the fuel level reaches one-eighth of a tank.). One of ordinary skill in the art would reasonably consider the actions to monitor the vehicle sensors 106 that are associated with the vehicle 104's fuel level, and monitor the vehicle 104's navigation system as equivalent to “controlling at least one vehicle operation”.
Claim Objections
Claim 13 objected to because of the following informalities: typographical error. Claim 13 includes the limitation “and controlling at least one vehicle operation of the vehicle based on the based on the control signal.”, where the phrase “based on the” is repeated twice. Appropriate correction is required.
Dependent claims 16-18 inherit the deficiency from independent claim 13 and are therefore objected to.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 21-24 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding claims 21-24: The claims recite a “display”, which is not supported by the specification. When looking to the specification, in at least pars. [0036-0037, 0039, 0044-0045, 0047-0048, 0050, 0059, 0061, 0079], the specification recites a “display module 116”. Furthermore, the specification recites in at least par. [0079] a “user device 118”, capable of displaying information. However, it’s not clear if the “display” from the claims is the same as the “display module 116” or the “user device 118” from the specification. For the purpose of compact prosecution, Examiner is interpreting “display” as any electronic device (including hardware and software) used for displaying information and/or interacting with a user via a graphical user interface (e.g., generic computing component).
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-3, 5-13, 16-19, and 21-24 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-patentable subject matter. The claims are directed to an abstract idea without significantly more. The judicial exception is not integrated into a practical application. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception as further set forth in MPEP 2106.
Step 1: The claimed invention is analyzed to determine if it falls outside one of the four statutory categories of invention. See MPEP 2106.03
Claim(s) 1-3, 5-12, 19, and 21-24 is/are directed to a system (i.e., Machine), and claim(s) 13, and 16-18 is/are directed to a method (i.e., Process). Therefore, the claims are directed to patent eligible categories of invention. Accordingly, the claims satisfy Step 1 of the eligibility inquiry.
As drafted, the limitations recited by claims 1-3, 5-13, 16-19, and 21-24 fall under the “Mental Processes” abstract idea group by setting forth activities that could be performed mentally by a human (including an observation, evaluation, judgment, opinion) (see MPEP § 2106.04(a)(2), subsection III).
Independent claim 1 recites a vehicle system for vehicle query resolution services with the following abstract limitations: “generate a plurality of embeddings for the vehicle specific information and user manuals, wherein the plurality of embeddings are generated based on defined parameters and wherein the defined parameters include at least one of a chunk size parameter or an overlap parameter; identify vehicle information for the vehicle based on the VIN and the query, and generate a response specific to the vehicle based on the vehicle information and the plurality of embeddings in response to the query; to generate an electronic notification including the response to a vehicle controller in the vehicle for automatically controlling, without user interaction, at least one vehicle operation of the vehicle based on the response specific to the vehicle.”. As currently recited, the limitations of claim 1 (but for the recitation of additional elements), could be accomplished mentally, such as by human observation, evaluation, judgement, or with the help of pen and paper. Therefore, claim 1 recites abstract ideas that fall under the “Mental Processes” abstract idea grouping by setting forth activities that could be performed mentally by a human (including an observation, evaluation, judgment, opinion) (see MPEP § 2106.04(a)(2), subsection III).
Independent claim 13 recites a method with the following abstract limitations: “generating a plurality of embeddings based on vehicle specific information and user manuals; receiving an input from a user, the input including a query specific to a vehicle and a vehicle identification number (VIN) unique to the vehicle; identifying vehicle information for the vehicle based on the VIN and the query; generating a response with at least one recommendation specific to the vehicle based on the vehicle information and the plurality of embeddings in response to the query; and in response to the at least one recommendation specific to the vehicle, automatically generating a control signal based on the at least one recommendation without user interaction”. As currently recited, the limitations of claim 13 (but for the recitation of additional elements), could be accomplished mentally, such as by human observation, evaluation, judgement, or with the help of pen and paper. Therefore, claim 13 recites abstract ideas that fall under the “Mental Processes” abstract idea grouping by setting forth activities that could be performed mentally by a human (including an observation, evaluation, judgment, opinion) (see MPEP § 2106.04(a)(2), subsection III).
Independent claim 19 recites a system with the following abstract limitations: “generate a plurality of embeddings for the vehicle specific information and user manuals; the input including a query specific to a vehicle and a vehicle identification number (VIN) unique to the vehicle; identify vehicle information for the vehicle based on the Vin and the query; and generate a response with at least one recommendation specific to the vehicle based on the vehicle information and the plurality of embeddings in response to the query; generate an output with the at least one recommendation specific to the vehicle, in response to a user input indicating approval of the at least one recommendation, generate a control signal based on the at least one recommendation specific to the vehicle”. As currently recited, the limitations of claim 19 (but for the recitation of additional elements), could be accomplished mentally, such as by human observation, evaluation, judgement, or with the help of pen and paper. Therefore, claim 19 recites abstract ideas that fall under the “Mental Processes” abstract idea grouping by setting forth activities that could be performed mentally by a human (including an observation, evaluation, judgment, opinion) (see MPEP § 2106.04(a)(2), subsection III).
Dependent claims 7, 11, and 21 further narrow the abstract idea and introduce further additional elements for consideration.
Dependent claims 2, 3, 5-6, 8-10, 12, 16-18, and 22-24 further narrow the abstract idea and do not introduce further additional elements for consideration.
Accordingly, because the Step 2A Prong One and Prong Two analysis resulted in the conclusion that the claims are directed to an abstract idea, additional analysis under Step 2B of the eligibility inquiry must be conducted in order to determine whether any claim element or combination of elements amount to significantly more than the judicial exception.
Step 2A, Prong 2: An evaluation is made whether a claim recites any additional element, or combination of additional elements, that integrate the judicial exception into a practical application of the exception. See MPEP 2106.04(d).
Regarding the computing additional elements, namely conversion controller, vehicle system controller, alert controller, and vehicle controller from the independent claims, these additional elements have been evaluated but fail to integrate the abstract idea into a practical application because they amount to using generic computing elements or instructions (software) to perform the abstract idea, similar to adding the words “apply it” (or equivalent), which merely serves to link the use of the judicial exception to a particular technological environment (generic computing environment). See MPEP 2106.05(f) and 2106.05(h).
The limitations “large language model configured to receive the plurality of embeddings and an input from a user” from claim 19 provides nothing more than mere instructions to implement an abstract idea on a generic computer, which does not integrate the abstract idea into a practical application. See MPEP 2106.05(f). MPEP 2106.05(f) provides the following considerations for determining whether a claim simply recites a judicial exception with the words “apply it” (or an equivalent), such as mere instructions to implement an abstract idea on a computer: (1) whether the claim recites only the idea of a solution or outcome i.e., the claim fails to recite details of how a solution to a problem is accomplished; (2) whether the claim invokes computers or other machinery merely as a tool to perform an existing process; and (3) the particularity or generality of the application of the judicial exception.
With respect to the limitations for “and control/controlling at least one vehicle operation of the vehicle based on the control signal.” from claims 13 and 19, these activities at most amount to insignificant extra-solution activity (e.g., insignificant application), which does not integrate the abstract idea into a practical application, as noted in MPEP 2106.05(g).
With respect to the limitations for receive vehicle specific information and user manuals, and receive the plurality of embeddings and an input from a user, the input including a query specific to a vehicle and a vehicle identification number (VIN) unique to the vehicle from claims 1 and 19, and transmit an electronic notification including the response to a vehicle controller in the vehicle for automatically controlling, without user interaction, at least one vehicle operation of the vehicle based on the response specific to the vehicle from claim 1, these activities at most amount to insignificant extra-solution activity (e.g., mere data gathering), which does not integrate the abstract idea into a practical application, as noted in MPEP 2106.05(g).
Regarding the computing additional elements, namely crowdsourcing controller from claim 7, display from claim 21 these additional elements have been evaluated but fail to integrate the abstract idea into a practical application because they amount to using generic computing elements or instructions (software) to perform the abstract idea, similar to adding the words “apply it” (or equivalent), which merely serves to link the use of the judicial exception to a particular technological environment (generic computing environment). See MPEP 2106.05(f) and 2106.05(h).
With respect to the vehicle sensor from claim 11, the vehicle sensor has been considered under Step 2A Prong Two, however the vehicle sensor is recited at a high level of generality and fails to provide a technical improvement or otherwise integrate the abstract idea into a practical application.
Accordingly, because the Step 2A Prong One and Prong Two analysis resulted in the conclusion that the claims are directed to an abstract idea, additional analysis under Step 2B of the eligibility inquiry must be conducted in order to determine whether any claim element or combination of elements amount to significantly more than the judicial exception.
Step 2B: The claims are analyzed to determine whether any additional element, or combination of additional elements, is/are sufficient to ensure that the claims amount to significantly more than the judicial exception. This analysis is also termed a search for "inventive concept." See MPEP 2106.05.
Regarding the computing additional elements, namely conversion controller, vehicle system controller, alert controller, and vehicle controller from the independent claims, these additional elements have been evaluated but fail to add significantly more because they amount to using generic computing elements or instructions (software) to perform the abstract idea, similar to adding the words “apply it” (or equivalent), which merely serves to link the use of the judicial exception to a particular technological environment (generic computing environment). See MPEP 2106.05(f) and 2106.05(h). See, e.g., Alice Corp., 134 S. Ct. 2347, 110 USPQ2d 1976; Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015).
The limitations “large language model configured to receive the plurality of embeddings and an input from a user” from claim 19 provides nothing more than mere instructions to implement an abstract idea on a generic computer, which does not add significantly more. See MPEP 2106.05(f). MPEP 2106.05(f) provides the following considerations for determining whether a claim simply recites a judicial exception with the words “apply it” (or an equivalent), such as mere instructions to implement an abstract idea on a computer: (1) whether the claim recites only the idea of a solution or outcome i.e., the claim fails to recite details of how a solution to a problem is accomplished; (2) whether the claim invokes computers or other machinery merely as a tool to perform an existing process; and (3) the particularity or generality of the application of the judicial exception. See, e.g., Alice Corp., 134 S. Ct. 2347, 110 USPQ2d 1976; Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015).
With respect to the limitations for “and control/controlling at least one vehicle operation of the vehicle based on the control signal.” from claims 13 and 19, these activities at most amount to insignificant extra-solution activity (e.g., insignificant application), which does not add significantly more to the abstract idea, as noted in MPEP 2106.05(g). See printing or downloading generated menus, Ameranth 842 F.3d at 1241-42, 120 USPQ2d at 1854-55.
With respect to the limitations for receive vehicle specific information and user manuals, and receive the plurality of embeddings and an input from a user, the input including a query specific to a vehicle and a vehicle identification number (VIN) unique to the vehicle from claims 1 and 19, and transmit an electronic notification including the response to a vehicle controller in the vehicle for automatically controlling, without user interaction, at least one vehicle operation of the vehicle based on the response specific to the vehicle from claim 1, these activities at most amount to insignificant extra-solution activity (e.g., mere data gathering), which does not add significantly more to the abstract idea, as noted in MPEP 2106.05(g). Additionally, the mere data gathering extra-solution activity has been recognized as well-understood, routine, and conventional, and thus insufficient to add significantly more to the abstract idea. See MPEP 2106.05(d) - Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network).
Regarding the computing additional elements, namely crowdsourcing controller from claim 7, display from claim 21 these additional elements have been evaluated but fail to add significantly more because they amount to using generic computing elements or instructions (software) to perform the abstract idea, similar to adding the words “apply it” (or equivalent), which merely serves to link the use of the judicial exception to a particular technological environment (generic computing environment). See MPEP 2106.05(f) and 2106.05(h). See, e.g., Alice Corp., 134 S. Ct. 2347, 110 USPQ2d 1976; Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015).
With respect to the vehicle sensor from claim 11, the vehicle sensor has been considered under Step 2A Prong Two, however the vehicle sensor is recited at a high level of generality and fails to provide a technical improvement or otherwise add significantly more.
In addition, when taken as an ordered combination, the ordered combination adds nothing that is not already present as when the elements are taken individually. Their collective functions merely provide generic computer implementation. Therefore, when viewed as a whole, these additional claim elements do not provide meaningful limitations to amount to significantly more than the abstract idea itself. The ordered combination of elements in the claims (including the limitations inherited from the parent claim(s)) add nothing that is not already present as when the elements are taken individually. There is no indication that the combination of elements improves the functioning of a computer or improves any other technology. Their collective functions merely provide generic computer implementation. Accordingly, the subject matter encompassed by the claims fails to amount to significantly more than the abstract idea itself.
Claim Rejections - 35 USC § 103
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 5-6, 13, 19, and 21-24 are rejected under 35 U.S.C. 103 as being unpatentable over Wolverton et al. (US 20140136187 A1, hereinafter “Wolverton”), in view of Shetty et al. (US 20260100042 A1, hereinafter “Shetty”).
Regarding claim 1: Wolverton teaches a vehicle system for vehicle query resolution services with the following limitations:
a conversion controller ([0125] The illustrative computing system 100 includes at least one processor 712 (e.g. a microprocessor, microcontroller, digital signal processor, etc.), memory 714, and an input/output (I/O) subsystem 716.; [0126] Although not specifically shown, it should be understood that the I/O subsystem 716 typically includes, among other things, an I/O controller, a memory controller, and one or more I/O ports. The processor 712 and the I/O subsystem 716 are communicatively coupled to the memory 714. The memory 714 may be embodied as any type of suitable computer memory device (e.g., volatile memory such as various forms of random access memory);
configured to receive vehicle specific information and user manuals ([0145] Embodiments may also be implemented as instructions stored using one or more machine-readable media, which may be read and executed by one or more processors.; [0062] The knowledge base 140 may include all of the content (and perhaps more) typically found in a vehicle owner's manual, including text, graphics, video, as well as conversational spoken natural-language representations of the text, graphics, and/or video found in the vehicle owner's manual.);
a vehicle system controller in communication with the conversion controller, ([0125] The illustrative computing system 100 includes at least one processor 712 (e.g. a microprocessor, microcontroller, digital signal processor, etc.), memory 714, and an input/output (I/O) subsystem 716.; [0126] Although not specifically shown, it should be understood that the I/O subsystem 716 typically includes, among other things, an I/O controller, a memory controller, and one or more I/O ports. The processor 712 and the I/O subsystem 716 are communicatively coupled to the memory 714. The memory 714 may be embodied as any type of suitable computer memory device (e.g., volatile memory such as various forms of random access memory; [0105] Referring now to FIG. 4, the method 400, which is executable as computerized programs, routines, logic and/or instructions by one or more of the various modules of the computing system 100 to proactively initiate a dialog with the user in response to task and/or condition-based user input 102 and real-time vehicle-related inputs 110, is performed. A human-generated input 102 may be classified as a conditional instruction at block 206 of FIG. 2, if the input 102 contains a task or condition and an action to be performed if the task or condition is satisfied.);
the vehicle system controller configured to receive the plurality of embeddings and an input from a user, (([0142] the vehicle personal assistant 112 may be configured to monitor the current context for and respond to user-specified commands or conditions. For instance, a user may issue a conditional instruction to the vehicle personal assistant 112 to "look for cheap gas when the tank gets below halfway." When the condition is met (i.e., gas tank half full), the vehicle personal assistant 112 initiates a search for nearby gas stations using a relevant website or software application);
the input including a query specific to a vehicle and a vehicle identification number (VIN) unique to the vehicle, identify vehicle information for the vehicle based on the VIN and the query([0055] For example, certain inputs 102 may have a different meaning depending on whether the vehicle 104 is turned on or off, whether the user is situated in the driver's seat or a passenger seat, whether the user is inside the vehicle, standing outside the vehicle, or simply accessing the vehicle personal assistant 112 from a computer located inside a home or office, or even whether the user is driving the vehicle 104 at relatively high speed on a freeway as opposed to being stuck in traffic or on a country road. For instance, if the vehicle personal assistant 112 detects that it is connected to the vehicle 104 and the vehicle 104 is powered on, the reasoner 136 can obtain specific information about the vehicle 104 (e.g., the particular make, model, and options that are associated with the vehicle's VIN or Vehicle Identification Number). Such information can be obtained from the vehicle manufacturer and stored in a vehicle-specific configuration 122 portion of the vehicle context model 116. The reasoner 136 can supplement the inputs 102 with such information, so that any responses provided by the vehicle personal assistant 112 are tailored to the configuration of the vehicle 104 and omit inapplicable or irrelevant information (rather than simply reciting generic statements from the printed version of the owner's manual).;
and generate a response specific to the vehicle based on the vehicle information and the plurality of embeddings in response to the query; ([0095] At block 216, the method 200 retrieves information related to the search request. In doing so, at blocks 218 and 220, the method 200 may search the vehicle-specific user's guide knowledge base 140 and/or one or more third party sources within the vehicle-related search realm 142 (e.g., the approved third party sources 140 and/or the other third party sources 142). The method 200 may determine which source(s) to utilize based on the context of the query. For example, if the user asks a question that may be answered by the vehicle owner's manual (e.g., "How often should I change my windshield wipers?"), the method 200 may primarily or exclusively consult or search the vehicle-specific user's guide knowledge base 140 for the answer. However, if the user asks a question that is beyond the scope of the owner's manual (e.g., "Should I use synthetic oil?"), the method 200 may search the third party sources 142 for the answer.; [0142] the vehicle personal assistant 112 may be configured to monitor the current context for and respond to user-specified commands or conditions. For instance, a user may issue a conditional instruction to the vehicle personal assistant 112 to "look for cheap gas when the tank gets below halfway." When the condition is met (i.e., gas tank half full), the vehicle personal assistant 112 initiates a search for nearby gas stations using a relevant website or software application);
and an alert controller in communication with the vehicle system controller, ([0124] The illustrative computing system 100 is in communication with the vehicle network 108 and, via one or more other networks 744 (e.g., a "cloud"), other computing systems or devices 746.; [0145] Embodiments may also be implemented as instructions stored using one or more machine-readable media, which may be read and executed by one or more processors.; [0125] The illustrative computing system 100 includes at least one processor 712 (e.g. a microprocessor, microcontroller, digital signal processor, etc.), memory 714, and an input/output (I/O) subsystem 716.; [0126] Although not specifically shown, it should be understood that the I/O subsystem 716 typically includes, among other things, an I/O controller, a memory controller, and one or more I/O ports. The processor 712 and the I/O subsystem 716 are communicatively coupled to the memory 714. The memory 714 may be embodied as any type of suitable computer memory device (e.g., volatile memory such as various forms of random access memory);
the alert controller configured to generate and transmit an electronic notification including the response to a vehicle controller in the vehicle ([0142] In yet another example, the vehicle personal assistant 112 may be configured to monitor the current context for and respond to user-specified commands or conditions. For instance, a user may issue a conditional instruction to the vehicle personal assistant 112 to "look for cheap gas when the tank gets below halfway." When the condition is met (i.e., gas tank half full), the vehicle personal assistant 112 initiates a search for nearby gas stations using a relevant website or software application, such as GASBUDDY.COM, and displays a list of the lowest-priced gas stations in the area. As another example, a busy working mom may request the vehicle personal assistant 112 to "look for pizza places with deal coupons within a few miles of my daughter's high school after her game this Tuesday" or "be on the lookout for good deals on tennis balls before next weekend." In each of these scenarios, the vehicle personal assistant 112 may classify the input 102 as a conditional instruction and process it as discussed above. For example, in identifying a good deal on tennis balls, the vehicle personal assistant 112 may determine an average cost for tennis balls in the immediate geographical region, periodically search for tennis ball prices as the vehicle 104 comes within geographic range of sports stores, and present a notification when the condition is met.);
for automatically controlling, without user interaction, at least one vehicle operation of the vehicle based on the response specific to the vehicle. ([0105] Referring now to FIG. 4, the method 400, which is executable as computerized programs, routines, logic and/or instructions by one or more of the various modules of the computing system 100 to proactively initiate a dialog with the user in response to task and/or condition-based user input 102 and real-time vehicle-related inputs 110, is performed. A human-generated input 102 may be classified as a conditional instruction at block 206 of FIG. 2, if the input 102 contains a task or condition and an action to be performed if the task or condition is satisfied. For example, the user may request the vehicle personal assistant 112 to "Tell me where the closest gas station is if I am running out of gas." In some embodiments, the vehicle personal assistant 112 may classify this spoken language input as a conditional instruction because there is both a condition (i.e., running out of gas) and an action associated with the condition (i.e., identify the nearest gas station).; [0106] In response to the conditional instruction, the method 400 may track the inputs 110 relating to the vehicle 104's fuel level and notify the user of the nearest gas station when the fuel reaches a certain level. To do this, the method 400 may interpret the phrase "running out of gas" as meaning "less than 1/8 of a tank" (based, e.g., on data or rules obtained from the vehicle-specific conversation model 132, the vehicle-specific user's guide knowledge base 140, or the vehicle context model 116). The method 400 may then monitor the vehicle sensors 106 that are associated with the vehicle 104's fuel level, and monitor the vehicle 104's navigation system (or a communicatively coupled Global Positioning System or GPS device, or a smart phone map or navigation application, for example). The method 400 may then present (e.g., by voice or graphic display) directions to the nearest gas station when the fuel level reaches one-eighth of a tank.).
Wolverton doesn’t teach:
and generate a plurality of embeddings for the vehicle specific information and user manuals,
wherein the plurality of embeddings are generated based on defined parameters and wherein the defined parameters include at least one of a chunk size parameter or an overlap parameter;
the vehicle system controller configured to receive the plurality of embeddings
and generate a response specific to the vehicle based on the plurality of embeddings in response to the query;
Shetty teaches:
and generate a plurality of embeddings for the vehicle specific information and user manuals, ([0212] where the input indicates that the user is interested in a desired tire pressure for a particular make and model of vehicle, the RAG component 992 may retrieve—using a RAG model performing a vector search in an embedding space, for example—the tire pressure information or the text corresponding thereto from a digital (embedded) version of the user manual for that particular vehicle make and model.);
wherein the plurality of embeddings are generated based on defined parameters and wherein the defined parameters include at least one of a chunk size parameter or an overlap parameter; ([0223] FIG. 9B is a block diagram of an example implementation in which the generative LM 930 includes a transformer encoder-decoder. For example, assume input text such as “Who discovered gravity” is tokenized (e.g., by the tokenizer 910 of FIG. 9A) into tokens such as words, and each token is encoded (e.g., by the embedding component 920 of FIG. 99A) into a corresponding embedding (e.g., of size 512).);
the vehicle system controller configured to receive the plurality of embeddings ([0210] Additionally or alternatively, the input 901 may include numerical sequences, precomputed embeddings (e.g., word or sentence embeddings), and/or structured data (e.g., in tabular formats, JSON, or XML).);
and generate a response specific to the vehicle based on the plurality of embeddings in response to the query; ([0071] the inference server 470 may provide the tokenized multi-modal prompt to the LLM 480 and return the LLM's response back to the detection task manager 420.; [0053] the text portion of the multi-modal prompt may depend on the applicable task and/or the implementation. Continuing with the example in which the stream selector 125a instructs, primes, or otherwise triggers the multi-modal language model(s) 140 to determine whether one of the sensors, sensor feeds, and/or frames of sensor data represents a scene or scene content that is relevant to a designated task, the stream selector 125a may use any known prompt engineering technique (e.g., using system prompts, role-based prompts, instruction prompts, question prompts, multi-turn prompts, etc.) to provide (e.g., via the inference coordinator 135a and/or the inference server 170) designated questions and/or instructions to (e.g., the LLM(s) 180 of) the multi-modal language model(s) 140 corresponding to the applicable detection task. By way of nonlimiting example, a text prompt may comprise a question such as “Which [camera/sensor] can see [the condition being monitoring for]?” or “Can one of the [cameras/sensors] see [the condition being monitoring for]?” Taking environmental text comprehension (e.g., understanding road signs) as an example, a possible text prompt may include a question such as “Which [camera/sensor] can see a road sign?” Taking child presence detection as an example, a possible text prompt may include a question such as “Which [camera/sensor] can see a child?” Taking license plate detection as an example, a possible text prompt may include a question such as “Can any of the [cameras/sensors] see a license plate?” Taking obstacle detection and avoidance as an example, a possible text prompt may include a question such as “Can any of the [cameras/sensors] see a car, road debris, or some other obstacle?” Those of ordinary skill in the art will understand how to adapt a text prompt to the designated detection task, examples of which are provided in more detail below. Whether the choice is expressed as a selection of a particular sensor, a sensor feed generated using the sensor, sensor data generated using the sensor, or otherwise, the response may effectively indicate or otherwise be representative of a selection of both the sensor and its corresponding sensor data (or sensor feed). Generally, a designated output format for sensor selection (e.g., a tag or other identifier representing the selected sensor) may be enforced in any suitable manner (e.g., using prompt engineering, post-processing, training, multiple inferences, etc.).).
It would have been obvious to one of ordinary skill in the art, at the time of applicant’s invention, to combine Wolverton with Shetty’s feature(s) listed above. One would’ve been motivated to do so in order to encode the sequential relationships and context of the tokens in the input sequence (Shetty; [0223]). By incorporating the teachings of Shetty, one would’ve been able to generate embeddings of a specific size to represent the input data.
Regarding claim 2: Wolverton teaches:
wherein the response specific to the vehicle includes at least one recommendation specific to the vehicle. ([0063] vehicle users may be likely to ask questions about the vehicle's tires, and so the question and answer pairs may include a number of answering sentences associated with the words "tire" and "tires." Such answering sentences may include "the recommended tire pressure for the front tires is 33 psi" and "the tires should be rotated every 5,000 miles.").
Regarding claim 3: Wolverton teaches:
further comprising the vehicle controller positioned in the vehicle and in communication with the alert controller, ([0125] The illustrative computing system 100 includes at least one processor 712 (e.g. a microprocessor, microcontroller, digital signal processor, etc.), memory 714, and an input/output (I/O) subsystem 716.; [0126] Although not specifically shown, it should be understood that the I/O subsystem 716 typically includes, among other things, an I/O controller, a memory controller, and one or more I/O ports. The processor 712 and the I/O subsystem 716 are communicatively coupled to the memory 714. The memory 714 may be embodied as any type of suitable computer memory device (e.g., volatile memory such as various forms of random access memory; [0105] Referring now to FIG. 4, the method 400, which is executable as computerized programs, routines, logic and/or instructions by one or more of the various modules of the computing system 100 to proactively initiate a dialog with the user in response to task and/or condition-based user input 102 and real-time vehicle-related inputs 110, is performed. A human-generated input 102 may be classified as a conditional instruction at block 206 of FIG. 2, if the input 102 contains a task or condition and an action to be performed if the task or condition is satisfied.);
the vehicle controller configured to automatically control the at least one vehicle operation of the vehicle based on the at least one recommendation specific to the vehicle. ([0117] At block 504, the method 500 analyzes the input 102 to determine the user's intended meaning thereof; e.g., the goal or objective in the user's mind at the time that the user generated the input 102, or the user's reason for providing the input 102. For example, the method 500 may determine: does the user want concise factual information (e.g., recommended motor oil), a more detailed explanation (e.g., a graphic showing where the oil dipstick is located), or a step-by-step tutorial (e.g., how to change the oil filter)? To do this, the method 500 may use standard (now existing or later developed) intent recognition and/or intent merging techniques, any of the techniques discussed above in connection with the input recognizer/interpreter 130, and/or other suitable methods. For example, the method 500 may extract words from the input 102 that appear to indicate a particular type of dialog: "what is" may be associated with a direct question and answer-type dialog; "how" or "why" may be associated with a how-to or trouble shooting-type dialog while "this" or "that" may be associated with a situation aware-type dialog. Particularly in situation-aware dialogs, adjectives (e.g., "wavy lines" or "orange") may be extracted and assigned a higher importance.; [0124] The illustrative computing system 100 is in communication with the vehicle network 108 and, via one or more other networks 744 (e.g., a "cloud"), other computing systems or devices 746.; [0145] Embodiments may also be implemented as instructions stored using one or more machine-readable media, which may be read and executed by one or more processors.; [0044] The human-generated input monitor 128 detects and receives human-generated inputs 102 from time to time during the operation of the vehicle personal assistant 112. The human-generated input monitor 128 may run continuously (e.g., as a background process), or may be invoked or terminated by a user on demand (e.g., by pressing a button control or saying a specific keyword). In other words, the vehicle personal assistant 112, and thus the human-generated input monitor 128, can be configured to monitor the inputs 102 irrespective of whether the vehicle personal assistant 112 is installed in the vehicle 104 and whether or not the vehicle 104 is in operation. Additionally, the user may choose whether to allow the vehicle personal assistant 112 to make use of certain of the human-generated inputs 102 but not others. For example, the user may allow the vehicle personal assistant 112 to monitor voice inputs 152 but not facial features or expressions 160.).
Regarding claim 5: Wolverton teaches:
wherein the response includes at least one of: a recommendation to obtain a vehicle application or service; ([0103] At block 316, the method 300 determines whether the user has actively or passively responded to the output presented at block 226 (e.g., whether any additional inputs 102 have been detected). If so, the method 300 may continue the dialog by performing a question-and-answer session with the user at block 318. That is, upon receiving a suggestion from the vehicle personal assistant 112, the user may respond with a follow-up question or statement. For example, the vehicle personal assistant 112 may suggest, "You should change your transmission fluid," in which the user may respond "Why?" In such a circumstance, the vehicle personal assistant 112 may respond explaining, for example, "The vehicle owner's manual indicates that your transmission fluid should be changed every 30,000 miles," or "The fluid sensor indicates that your transmission fluid is too viscous or adulterated." In the fog light example above, the user may respond to the suggestion to turn on the fog lights with a question, "How do I turn on the fog lights?" to which the vehicle personal assistant 112 may respond, "Flip the switch to the left of the steering wheel.");
and an answer is currently unavailable. ([Fig. 3] 312 “Provide Suggestion?”, No; [0101] At block 312, the method 300 determines whether suggestion or informational content has been identified at block 306 that appears to pertain to the current context. In doing so, the method 300 may consult a probabilistic or statistical model as described above, to determine the likely relevance of the identified content to the current context. If not, the method 300 returns to block 302 and continues monitoring the vehicle context while the vehicle is powered on (and as long as the vehicle personal assistant 112 is in use).).
Regarding claim 6: Wolverton doesn’t teach:
wherein the vehicle system controller includes a large language model configured to generate the response specific to the vehicle based on the vehicle information and the plurality of embeddings.
Shetty teaches:
wherein the vehicle system controller includes a large language model configured to generate the response specific to the vehicle based on the vehicle information and the plurality of embeddings. ([0054] In some embodiments, the text portion of the multi-modal prompt corresponds to a user command or query, and may be generated using an LLM (e.g., of the LLM(s) 180 of FIG. 1). Taking an example in-cabin embodiment, an occupant may ask a digital assistant a question such as “Is my [wallet or laptop] in the car?” via an input (e.g., audio, touch) interface of the vehicle, and the digital assistant may convert the request to a textual representation and provide the request (or a portion thereof) to the detection task manager 120a to trigger the detection task manager 120a to look for the requested object (or check for some other requested condition). In some embodiments, at block B206, the stream selector 125a of the detection task manager 120a may parse the textual representation of the request to identify the requested condition (e.g., the presence of a specified object) and insert the requested condition into one or more fields of a template text prompt that instructs, primes, or otherwise triggers the multi-modal language model(s) 140 to check for a relevant sensor, sensor feed, and/or frame of sensor data (e.g., “Can one of the [cameras/sensors] see [the specified/parsed/extracted condition]?” In some embodiments, at block B206, the stream selector 125a may prompt the LLM(s) 180 to convert the user command or query into an appropriate text prompt for sensor selection (e.g., using a text prompt such as “Convert the following user input into a text prompt that instructs a VLM to check whether one of the images provided to the VLM shows what the user is asking for: [user input]”). As such, the stream selector 125a may parse the response, extract the generated text prompt, and use it as part of a multi-modal prompt for sensor selection at block B206. These are meant simply as examples, and variations may be implemented within the scope of the present disclosure.; [0055] Continuing with the method 200, at block B208, the stream selector 125a may receive one or more responses from (e.g., the LLM(s) 180 of) the multi-modal language model(s) 140 and interpret the one or more responses.).
It would have been obvious to one of ordinary skill in the art, at the time of applicant’s invention, to combine Wolverton with Shetty’s feature(s) listed above. One would’ve been motivated to do so in order to provide a response identifying a selected sensor that is relevant to the designated task (Shetty; [0055]). By incorporating the teachings of Shetty, one would’ve been able to include a large language model configured to generate responses.
Regarding claim 13: Wolverton teaches:
receiving an input from a user, the input including a query specific to a vehicle and a vehicle identification number (VIN) unique to the vehicle; ([0095] At block 216, the method 200 retrieves information related to the search request. In doing so, at blocks 218 and 220, the method 200 may search the vehicle-specific user's guide knowledge base 140 and/or one or more third party sources within the vehicle-related search realm 142 (e.g., the approved third party sources 140 and/or the other third party sources 142). The method 200 may determine which source(s) to utilize based on the context of the query. For example, if the user asks a question that may be answered by the vehicle owner's manual (e.g., "How often should I change my windshield wipers?"), the method 200 may primarily or exclusively consult or search the vehicle-specific user's guide knowledge base 140 for the answer. However, if the user asks a question that is beyond the scope of the owner's manual (e.g., "Should I use synthetic oil?"), the method 200 may search the third party sources 142 for the answer.; [0055] the reasoner 136 can obtain specific information about the vehicle 104 (e.g., the particular make, model, and options that are associated with the vehicle's VIN or Vehicle Identification Number).);
identifying vehicle information for the vehicle based on the VIN and the query; ([0055] For example, certain inputs 102 may have a different meaning depending on whether the vehicle 104 is turned on or off, whether the user is situated in the driver's seat or a passenger seat, whether the user is inside the vehicle, standing outside the vehicle, or simply accessing the vehicle personal assistant 112 from a computer located inside a home or office, or even whether the user is driving the vehicle 104 at relatively high speed on a freeway as opposed to being stuck in traffic or on a country road. For instance, if the vehicle personal assistant 112 detects that it is connected to the vehicle 104 and the vehicle 104 is powered on, the reasoner 136 can obtain specific information about the vehicle 104 (e.g., the particular make, model, and options that are associated with the vehicle's VIN or Vehicle Identification Number). Such information can be obtained from the vehicle manufacturer and stored in a vehicle-specific configuration 122 portion of the vehicle context model 116.);
generating a response with at least one recommendation specific to the vehicle based on the vehicle information and the plurality of embeddings in response to the query; ([0142] the vehicle personal assistant 112 may be configured to monitor the current context for and respond to user-specified commands or conditions. For instance, a user may issue a conditional instruction to the vehicle personal assistant 112 to "look for cheap gas when the tank gets below halfway." When the condition is met (i.e., gas tank half full), the vehicle personal assistant 112 initiates a search for nearby gas stations using a relevant website or software application);
and in response to the at least one recommendation specific to the vehicle, automatically generating a control signal based on the at least one recommendation without user interaction ([0142] In yet another example, the vehicle personal assistant 112 may be configured to monitor the current context for and respond to user-specified commands or conditions. For instance, a user may issue a conditional instruction to the vehicle personal assistant 112 to "look for cheap gas when the tank gets below halfway." When the condition is met (i.e., gas tank half full), the vehicle personal assistant 112 initiates a search for nearby gas stations using a relevant website or software application, such as GASBUDDY.COM, and displays a list of the lowest-priced gas stations in the area. As another example, a busy working mom may request the vehicle personal assistant 112 to "look for pizza places with deal coupons within a few miles of my daughter's high school after her game this Tuesday" or "be on the lookout for good deals on tennis balls before next weekend." In each of these scenarios, the vehicle personal assistant 112 may classify the input 102 as a conditional instruction and process it as discussed above. For example, in identifying a good deal on tennis balls, the vehicle personal assistant 112 may determine an average cost for tennis balls in the immediate geographical region, periodically search for tennis ball prices as the vehicle 104 comes within geographic range of sports stores, and present a notification when the condition is met.);
and controlling at least one vehicle operation of the vehicle based on the based on the control signal. ([0105] Referring now to FIG. 4, the method 400, which is executable as computerized programs, routines, logic and/or instructions by one or more of the various modules of the computing system 100 to proactively initiate a dialog with the user in response to task and/or condition-based user input 102 and real-time vehicle-related inputs 110, is performed. A human-generated input 102 may be classified as a conditional instruction at block 206 of FIG. 2, if the input 102 contains a task or condition and an action to be performed if the task or condition is satisfied. For example, the user may request the vehicle personal assistant 112 to "Tell me where the closest gas station is if I am running out of gas." In some embodiments, the vehicle personal assistant 112 may classify this spoken language input as a conditional instruction because there is both a condition (i.e., running out of gas) and an action associated with the condition (i.e., identify the nearest gas station).; [0106] In response to the conditional instruction, the method 400 may track the inputs 110 relating to the vehicle 104's fuel level and notify the user of the nearest gas station when the fuel reaches a certain level. To do this, the method 400 may interpret the phrase "running out of gas" as meaning "less than 1/8 of a tank" (based, e.g., on data or rules obtained from the vehicle-specific conversation model 132, the vehicle-specific user's guide knowledge base 140, or the vehicle context model 116). The method 400 may then monitor the vehicle sensors 106 that are associated with the vehicle 104's fuel level, and monitor the vehicle 104's navigation system (or a communicatively coupled Global Positioning System or GPS device, or a smart phone map or navigation application, for example). The method 400 may then present (e.g., by voice or graphic display) directions to the nearest gas station when the fuel level reaches one-eighth of a tank.).
Wolverton doesn’t teach:
generating a plurality of embeddings based on vehicle specific information and user manuals;
generating a response with at least one recommendation specific to the vehicle based on the plurality of embeddings in response to the query;
Shetty teaches:
generating a plurality of embeddings based on vehicle specific information and user manuals; ([0212] where the input indicates that the user is interested in a desired tire pressure for a particular make and model of vehicle, the RAG component 992 may retrieve—using a RAG model performing a vector search in an embedding space, for example—the tire pressure information or the text corresponding thereto from a digital (embedded) version of the user manual for that particular vehicle make and model.);
generating a response with at least one recommendation specific to the vehicle based on the plurality of embeddings in response to the query; ([0071] the inference server 470 may provide the tokenized multi-modal prompt to the LLM 480 and return the LLM's response back to the detection task manager 420.; [0053] the text portion of the multi-modal prompt may depend on the applicable task and/or the implementation. Continuing with the example in which the stream selector 125a instructs, primes, or otherwise triggers the multi-modal language model(s) 140 to determine whether one of the sensors, sensor feeds, and/or frames of sensor data represents a scene or scene content that is relevant to a designated task, the stream selector 125a may use any known prompt engineering technique (e.g., using system prompts, role-based prompts, instruction prompts, question prompts, multi-turn prompts, etc.) to provide (e.g., via the inference coordinator 135a and/or the inference server 170) designated questions and/or instructions to (e.g., the LLM(s) 180 of) the multi-modal language model(s) 140 corresponding to the applicable detection task. By way of nonlimiting example, a text prompt may comprise a question such as “Which [camera/sensor] can see [the condition being monitoring for]?” or “Can one of the [cameras/sensors] see [the condition being monitoring for]?” Taking environmental text comprehension (e.g., understanding road signs) as an example, a possible text prompt may include a question such as “Which [camera/sensor] can see a road sign?” Taking child presence detection as an example, a possible text prompt may include a question such as “Which [camera/sensor] can see a child?” Taking license plate detection as an example, a possible text prompt may include a question such as “Can any of the [cameras/sensors] see a license plate?” Taking obstacle detection and avoidance as an example, a possible text prompt may include a question such as “Can any of the [cameras/sensors] see a car, road debris, or some other obstacle?” Those of ordinary skill in the art will understand how to adapt a text prompt to the designated detection task, examples of which are provided in more detail below. Whether the choice is expressed as a selection of a particular sensor, a sensor feed generated using the sensor, sensor data generated using the sensor, or otherwise, the response may effectively indicate or otherwise be representative of a selection of both the sensor and its corresponding sensor data (or sensor feed). Generally, a designated output format for sensor selection (e.g., a tag or other identifier representing the selected sensor) may be enforced in any suitable manner (e.g., using prompt engineering, post-processing, training, multiple inferences, etc.).).
It would have been obvious to one of ordinary skill in the art, at the time of applicant’s invention, to combine Wolverton with Shetty’s feature(s) listed above. One would’ve been motivated to do so in order to encode the sequential relationships and context of the tokens in the input sequence (Shetty; [0223]). By incorporating the teachings of Shetty, one would’ve been able to generate embeddings of a specific size to represent the input data.
Regarding claim 19: Wolverton teaches:
a conversion controller ([0125] The illustrative computing system 100 includes at least one processor 712 (e.g. a microprocessor, microcontroller, digital signal processor, etc.), memory 714, and an input/output (I/O) subsystem 716.; [0126] Although not specifically shown, it should be understood that the I/O subsystem 716 typically includes, among other things, an I/O controller, a memory controller, and one or more I/O ports. The processor 712 and the I/O subsystem 716 are communicatively coupled to the memory 714. The memory 714 may be embodied as any type of suitable computer memory device (e.g., volatile memory such as various forms of random access memory);
configured to receive vehicle specific information and user manuals ([0145] Embodiments may also be implemented as instructions stored using one or more machine-readable media, which may be read and executed by one or more processors.; [0062] The knowledge base 140 may include all of the content (and perhaps more) typically found in a vehicle owner's manual, including text, graphics, video, as well as conversational spoken natural-language representations of the text, graphics, and/or video found in the vehicle owner's manual.);
a vehicle system controller in communication with the conversion controller, ([0125] The illustrative computing system 100 includes at least one processor 712 (e.g. a microprocessor, microcontroller, digital signal processor, etc.), memory 714, and an input/output (I/O) subsystem 716.; [0126] Although not specifically shown, it should be understood that the I/O subsystem 716 typically includes, among other things, an I/O controller, a memory controller, and one or more I/O ports. The processor 712 and the I/O subsystem 716 are communicatively coupled to the memory 714. The memory 714 may be embodied as any type of suitable computer memory device (e.g., volatile memory such as various forms of random access memory; [0105] Referring now to FIG. 4, the method 400, which is executable as computerized programs, routines, logic and/or instructions by one or more of the various modules of the computing system 100 to proactively initiate a dialog with the user in response to task and/or condition-based user input 102 and real-time vehicle-related inputs 110, is performed. A human-generated input 102 may be classified as a conditional instruction at block 206 of FIG. 2, if the input 102 contains a task or condition and an action to be performed if the task or condition is satisfied.);
the input including a query specific to a vehicle and a vehicle identification number (VIN) unique to the vehicle, ([0095] At block 216, the method 200 retrieves information related to the search request. In doing so, at blocks 218 and 220, the method 200 may search the vehicle-specific user's guide knowledge base 140 and/or one or more third party sources within the vehicle-related search realm 142 (e.g., the approved third party sources 140 and/or the other third party sources 142). The method 200 may determine which source(s) to utilize based on the context of the query. For example, if the user asks a question that may be answered by the vehicle owner's manual (e.g., "How often should I change my windshield wipers?"), the method 200 may primarily or exclusively consult or search the vehicle-specific user's guide knowledge base 140 for the answer. However, if the user asks a question that is beyond the scope of the owner's manual (e.g., "Should I use synthetic oil?"), the method 200 may search the third party sources 142 for the answer.; [0055] the reasoner 136 can obtain specific information about the vehicle 104 (e.g., the particular make, model, and options that are associated with the vehicle's VIN or Vehicle Identification Number).);
identify vehicle information for the vehicle based on the Vin and the query, ([0055] For example, certain inputs 102 may have a different meaning depending on whether the vehicle 104 is turned on or off, whether the user is situated in the driver's seat or a passenger seat, whether the user is inside the vehicle, standing outside the vehicle, or simply accessing the vehicle personal assistant 112 from a computer located inside a home or office, or even whether the user is driving the vehicle 104 at relatively high speed on a freeway as opposed to being stuck in traffic or on a country road. For instance, if the vehicle personal assistant 112 detects that it is connected to the vehicle 104 and the vehicle 104 is powered on, the reasoner 136 can obtain specific information about the vehicle 104 (e.g., the particular make, model, and options that are associated with the vehicle's VIN or Vehicle Identification Number). Such information can be obtained from the vehicle manufacturer and stored in a vehicle-specific configuration 122 portion of the vehicle context model 116.);
and generate a response with at least one recommendation specific to the vehicle based on the vehicle information and the plurality of embeddings in response to the query; ([0142] the vehicle personal assistant 112 may be configured to monitor the current context for and respond to user-specified commands or conditions. For instance, a user may issue a conditional instruction to the vehicle personal assistant 112 to "look for cheap gas when the tank gets below halfway." When the condition is met (i.e., gas tank half full), the vehicle personal assistant 112 initiates a search for nearby gas stations using a relevant website or software application);
and a vehicle controller positioned in the vehicle, ([0125] The illustrative computing system 100 includes at least one processor 712 (e.g. a microprocessor, microcontroller, digital signal processor, etc.), memory 714, and an input/output (I/O) subsystem 716.; [0126] Although not specifically shown, it should be understood that the I/O subsystem 716 typically includes, among other things, an I/O controller, a memory controller, and one or more I/O ports.);
the vehicle controller configured to generate an output with the at least one recommendation specific to the vehicle, ([0089] the context analyzer 124 compares the inputs 110 and/or other data stored in the portions 118, 120, 122 of the vehicle context model 116 to the rules and/or templates 124 to determine the current context of the vehicle and identify suggestions or actions that may be performed by the vehicle personal assistant 112 in response thereto.);
in response to a user input indicating approval of the at least one recommendation, generate a control signal based on the at least one recommendation specific to the vehicle, and control at least one vehicle operation of the vehicle based on the control signal. ([0065] The trouble-shooting portion 182 may also include information that can be used to respond to "how-to" questions. For example, the user may start a dialog with a question, "how often should my tires be rotated?" to which the vehicle personal assistant 112 responds with an answering sentence "every 5,000 miles is recommended." Realizing that the user needs to rotate his or her vehicle's tires, the user may then ask, "how do I rotate the tires?" to which the vehicle personal assistant 112 responds with a more detailed explanation or interactive tutorial gleaned from the trouble-shooting portion 182, e.g., "move the left front wheel and tire to the left rear, the right front to the right rear, the left rear to the right front and the right rear to the left front." In this case, the dialog manager described above may insert pauses between each of the steps of the tutorial, to wait for affirmation from the user that the step has been completed or to determine whether the user needs additional information to complete the task.).
Wolverton doesn’t teach:
and generate a plurality of embeddings for the vehicle specific information and user manuals;
the vehicle system controller including a large language model configured to receive the plurality of embeddings and an input from a user,
and generate a response with at least one recommendation specific to the vehicle based on the vehicle information and the plurality of embeddings in response to the query;
Shetty teaches:
and generate a plurality of embeddings for the vehicle specific information and user manuals; ([0212] where the input indicates that the user is interested in a desired tire pressure for a particular make and model of vehicle, the RAG component 992 may retrieve—using a RAG model performing a vector search in an embedding space, for example—the tire pressure information or the text corresponding thereto from a digital (embedded) version of the user manual for that particular vehicle make and model.);
the vehicle system controller including a large language model configured to receive the plurality of embeddings and an input from a user, ([Figure 1] LLM(s) 180; [0054] In some embodiments, the text portion of the multi-modal prompt corresponds to a user command or query, and may be generated using an LLM (e.g., of the LLM(s) 180 of FIG. 1). Taking an example in-cabin embodiment, an occupant may ask a digital assistant a question such as “Is my [wallet or laptop] in the car?” via an input (e.g., audio, touch) interface of the vehicle, and the digital assistant may convert the request to a textual representation and provide the request (or a portion thereof) to the detection task manager 120a to trigger the detection task manager 120a to look for the requested object (or check for some other requested condition).;
and generate a response with at least one recommendation specific to the vehicle based on the vehicle information and the plurality of embeddings in response to the query; ([Figure 4]; [0070] FIG. 4 is a block diagram illustrating an example multi-modal language model 430 (e.g., which may correspond to the multi-modal language model(s) 140 of FIG. 1). In this example implementation, the multi-modal language model 430 includes an encoder 440 (which may correspond to the encoder(s) 150 of FIG. 1) and a projector 450 (which may correspond to the projector(s) 160 of FIG. 1) hosted and executing on first hardware 410 (e.g., which may correspond to the SoC 110 of FIG. 1) and an LLM 480 (e.g., which may correspond to the LLM(s) 180 of FIG. 1) hosted and executing on second hardware 165 (e.g., which may correspond to the external hardware 165 of FIG. 1). As such, this is an example implementation in which a multi-modal language model (e.g., a VLM) may be split up and hosted by multiple devices.; [0075] If the multi-modal language model(s) 140 detect distress, the control component(s) 190 may trigger some emergency response (e.g., contacting emergency services, displaying or announcing emergency instructions).
It would have been obvious to one of ordinary skill in the art, at the time of applicant’s invention, to combine Wolverton with Shetty’s feature(s) listed above. One would’ve been motivated to do so in order to return the LLM's response back to the detection task manager 420 (Shetty; [0071]). By incorporating the teachings of Shetty, one would’ve been able to use a LLM to analyze the input and generate a response.
Regarding claim 21: Wolverton teaches:
further comprising a display positioned in the vehicle, ([Fig. 8] Display Screen 824);
the display configured to display the recommendation specific to the vehicle and a selectable input to approve the at least one recommendation. ([Fig. 8] Display Screen 824, Display Portion 822: Options to “Ignore”, “Listen”, “View”, and “Save”.).
Regarding claim 22: Wolverton teaches:
the display is configured to receive the query specific to the vehicle from the user; and the vehicle controller is configured to, ([0039] A number of different human-generated inputs 102 may be involved in a dialog between a person and the vehicle personal assistant 112. Generally, the vehicle personal assistant 112 processes and responds to conversational spoken natural language (voice) inputs 152, but it may consider other input forms that may be generated by the user, alternatively or in addition to the voice inputs 152. Such other forms of human-generated inputs 102 may include gestures (e.g., body movements) 154, gaze (e.g., location and/or duration of eye focus) 156, touch (e.g., pressing a button or turning a dial) 158, facial features or expressions (e.g., whether the user appears alert, sleepy, or agitated) 160, media (e.g., text, photographs, video, or recorded sounds supplied by the user) 162, and/or others. Generally speaking, the inputs 102 may include any variety of human-initiated inputs. For example, the inputs 102 may include deliberate or "active" inputs that are intended to cause an event to occur at the computing system 100 (e.g., the act of touching a button control or saying a specific voice command). The inputs 102 may also include involuntary or "passive" inputs that the user may not expressly intend to result in a system event (such as a non-specific vocal or facial expression, tone of voice or loudness).; [0127] In the illustrative computing environment 700, the I/O subsystem 716 is communicatively coupled to a number of hardware components including: at least one touch-sensitive display 718 (e.g., a touchscreen, virtual keypad).);
in response to receiving the query, automatically attach the VIN to the query and transmit the input including the query and the VIN to the vehicle system controller. ([0055] the reasoner 136 can obtain specific information about the vehicle 104 (e.g., the particular make, model, and options that are associated with the vehicle's VIN or Vehicle Identification Number). Such information can be obtained from the vehicle manufacturer and stored in a vehicle-specific configuration 122 portion of the vehicle context model 116. The reasoner 136 can supplement the inputs 102 with such information, so that any responses provided by the vehicle personal assistant 112 are tailored to the configuration of the vehicle 104 and omit inapplicable or irrelevant information (rather than simply reciting generic statements from the printed version of the owner's manual).).
Regarding claim 23: Wolverton teaches:
further comprising a display positioned in the vehicle, ([Fig. 8] Display Screen 824);
the display configured to display the recommendation specific to the vehicle and a selectable input to approve the at least one recommendation. ([Fig. 8] Display Screen 824, Display Portion 822: Options to “Ignore”, “Listen”, “View”, and “Save”.).
Regarding claim 24: Wolverton teaches:
the display is configured to receive the query specific to the vehicle from the user; ([0039] A number of different human-generated inputs 102 may be involved in a dialog between a person and the vehicle personal assistant 112. Generally, the vehicle personal assistant 112 processes and responds to conversational spoken natural language (voice) inputs 152, but it may consider other input forms that may be generated by the user, alternatively or in addition to the voice inputs 152. Such other forms of human-generated inputs 102 may include gestures (e.g., body movements) 154, gaze (e.g., location and/or duration of eye focus) 156, touch (e.g., pressing a button or turning a dial) 158, facial features or expressions (e.g., whether the user appears alert, sleepy, or agitated) 160, media (e.g., text, photographs, video, or recorded sounds supplied by the user) 162, and/or others. Generally speaking, the inputs 102 may include any variety of human-initiated inputs. For example, the inputs 102 may include deliberate or "active" inputs that are intended to cause an event to occur at the computing system 100 (e.g., the act of touching a button control or saying a specific voice command). The inputs 102 may also include involuntary or "passive" inputs that the user may not expressly intend to result in a system event (such as a non-specific vocal or facial expression, tone of voice or loudness).; [0127] In the illustrative computing environment 700, the I/O subsystem 716 is communicatively coupled to a number of hardware components including: at least one touch-sensitive display 718 (e.g., a touchscreen, virtual keypad).);
and the vehicle controller is configured to, in response to receiving the query, automatically attach the VIN to the query and transmit the input including the query and the VIN to the vehicle system controller. ([0055] the reasoner 136 can obtain specific information about the vehicle 104 (e.g., the particular make, model, and options that are associated with the vehicle's VIN or Vehicle Identification Number). Such information can be obtained from the vehicle manufacturer and stored in a vehicle-specific configuration 122 portion of the vehicle context model 116. The reasoner 136 can supplement the inputs 102 with such information, so that any responses provided by the vehicle personal assistant 112 are tailored to the configuration of the vehicle 104 and omit inapplicable or irrelevant information (rather than simply reciting generic statements from the printed version of the owner's manual).).
Claims 7, 8, 10, 11, 16, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Wolverton et al. (US 20140136187 A1, hereinafter “Wolverton”), in view of Shetty et al. (US 20260100042 A1, hereinafter “Shetty”) as applied to claim 1, and 13 above, in further view of Xu et al. (US 20240095460 A1, hereinafter “Xu”).
Regarding claim 7: Wolverton doesn’t teach:
the plurality of embeddings for the vehicle specific user manuals are a plurality of first embeddings;
and the vehicle system further includes a crowdsourcing module in communication with the vehicle system controller, the crowdsourcing controller configured to: receive a plurality of previous queries and a corresponding plurality of responses to the previous queries;
determine subsets of frequently asked queries from the queries, the subsets of frequently asked queries specific to different vehicle models;
and generate a plurality of second embeddings for the subsets of frequently asked queries specific to the different vehicle models and their corresponding responses.
Xu teaches:
the plurality of embeddings for the vehicle specific user manuals are a plurality of first embeddings; ([0003] For a second example, the systems and methods may use the retrieval system(s) to retrieve contextual information related to the speech, such as contextual information from a (fixed or live) text-based knowledge base—such as a manual, a vehicle manual, a machine manual, a document, etc.—that is stored in the database(s).);
and the vehicle system further includes a crowdsourcing controller in communication with the vehicle system controller, the crowdsourcing controller configured to: receive a plurality of previous queries and a corresponding plurality of responses to the previous queries; ([0032] various functions may be carried out by a processor executing instructions stored in memory.; [0099] The vehicle 900 may include a system(s) on a chip (SoC) 904. The SoC 904 may include CPU(s) 906, GPU(s) 908, processor(s) 910, cache(s) 912, accelerator(s) 914, data store(s) 916, and/or other components and features not illustrated. The SoC(s) 904 may be used to control the vehicle 900 in a variety of platforms and systems. For example, the SoC(s) 904 may be combined in a system (e.g., the system of the vehicle 900) with an HD map 922 which may obtain map refreshes and/or updates via a network interface 924 from one or more servers (e.g., server(s) 978 of FIG. 9D).; [0062] As further shown, the language model(s) 124 may further output contextual data 120, which, in some examples, may include at least a portion of the output data 126. As discussed herein, the prompt component 118 may further use at least a portion of the contextual data 120 to generate the prompt data 122. For example, if the user continues to ask questions associated with the vehicle, the prompt component 118 may use the contextual data 120 to continue generating the prompt data 122 for the questions, where the contextual data 120 represents a context associated with outputs to previous questions.; [0070] the vehicle may input additional data into the language model(s) 124, such as contextual data 116 generated by the retrieval component 114 and/or contextual data 120 previously output by the language model(s) 124.);
determine subsets of frequently asked queries from the queries, the subsets of frequently asked queries specific to different vehicle models; ([0001] For a conversational assistant to operate within a vehicle, the conversational assistant may be preloaded with a set of answers to a set of questions that are commonly asked by passengers.; [0036] Referring back to FIG. 1, the process 100 may include a retrieval component 108 generating question/answer data 110 associated with the text data 106. For instance, an information database(s) 112 may store a number of question/answer pairs. As described herein, the number of question/answer pairs may include, but is not limited to, one question/answer pair, one hundred question/answer pairs, five hundred question/answer pairs, one thousand question/answer pairs, and/or any other number of question/answer pairs. In some examples, the question/answer pairs may be associated with a specific type of vehicle, such as a vehicle manufacturer, a vehicle model, and/or a vehicle year. For instance, the question/answer pairs may be generated using a knowledge base—such as an OEM manual—associated with the vehicle. In some examples, the question/answer pairs may be associated with more than one type of vehicle. For instance, the question/answer pairs may include general questions and answers associated with different vehicle manufacturers, different vehicle models, and/or different vehicle years. Still, in some examples, the question/answer pairs may be associated with topics other than vehicles.; [0037] FIG. 3 illustrates an example of question/answer pairs for vehicles);
and generate a plurality of second embeddings for the subsets of frequently asked queries specific to the different vehicle models and their corresponding responses. ([0004 the current systems, in some embodiments, use a language model(s) (e.g., a large language model(s)) to generate outputs that are more natural, conversational, robust, scalable, and accurate.; [0023] The system(s) may then use a retrieval system(s) to retrieve, from the database(s) (or data stores, or other storage or memory types), one or more question/answer pairs that are related to the text data. In some examples, to retrieve the question/answer pair(s), the question/answer pairs stored within the database(s) may be associated with embeddings. For instance, a first question/answer pair may be associated with a first embedding, a second question/answer pair may be associated with a second embedding, a third question/answer pair may be associated with a third embedding, and/or so forth.; [0025] In some embodiments, in addition to or alternatively from storing and/or retrieving question/answer pairs, the system may store intents, sub-intents, tokens, or other classification types corresponding to different answer types. In such examples, embeddings may be matched to a closest intent (e.g., a “tire pressure intent,” a “open gas compartment intent,” etc.), and this information may be used to determine an answer.; [0033] The process 100 may include one or more speech-processing components 102 processing audio data 104. For instance, the vehicle may generate the audio data 104 using one or more microphone(s), where the audio data 104 represents speech (e.g., an utterance) from a user of the vehicle. In some examples, the speech may represent a task being requested by the user, such as a question about the vehicle. As described herein, when the task includes a question, the question may be associated with a feature (e.g., a radio, a display, etc.) of the vehicle, a component (e.g., a window, a door, a tire, an engine, etc.) of the vehicle, a maintenance schedule associated with the vehicle, and/or any other aspect of the vehicle. The vehicle may then process the audio data 104 using the speech-processing component(s) 102. As described herein, the speech-processing component(s) 102 may include, but is not limited to, one or more ASR models, one or more STT models, one or more NLP models, and/or any other type of speech model.
It would have been obvious to one of ordinary skill in the art, at the time of applicant’s invention, to combine modified Wolverton with Xu’s feature(s) listed above. One would’ve been motivated to do so in order to generate an embedding for the transcript (Xu; [0023]). By incorporating the teachings of Xu, one would’ve been able to generate embeddings for the questions and responses.
Regarding claim 8: Wolverton doesn’t teach:
wherein the control module is configured to generate the response specific to the vehicle based on the vehicle information, the plurality of first embeddings, and the plurality of second embeddings in response to the input query.
Xu teaches:
wherein the vehicle system controller is configured to generate the response specific to the vehicle based on the vehicle information, the plurality of first embeddings, and the plurality of second embeddings in response to the query. ([0071] The method 700, at block B708, may include determining, using the language model, an output associated with the question. For instance, the language model(s) 124 may process the text data 106 and the question/answer data 110 (e.g., the prompt data 122) and, based on the processing, output data 126 associated with the question. As described herein, the output data 126 may represent information associated with the question. The vehicle may then provide the information to the passenger. In some examples, the vehicle provides the information by outputting sound associated with the output data 126, where the sound includes one or more words representing the information. In some examples, the vehicle provides the information by displaying content associated with the output data 126, where the content includes one or more words representing the information.; [0057] While the example of FIG. 5 illustrates one example technique that the retrieval component 114 may use to retrieve the contextual information, in other examples, the retrieval component 114 may use additional and/or alternative techniques. For a first example, and as described herein, the information may be separated into different categories. For instance, and using the example of FIG. 5 where the information 504 is from the OEM manual 502, the information 504 may be separated into component categories, such as tires, motor, doors, windows, and/or the like. In such examples, the retrieval component 114 may use the categories to retrieve contextual information that is in a similar category as the transcript 206 represented by the text data 204. For a second example, the retrieval component 114 may match one or more words represented by the text data 204 to one or more words represented by the information 504. The retrieval component 114 may then retrieve contextual information that includes at least a threshold number (e.g., one, two, three, five, ten, etc.) of matching words.).
It would have been obvious to one of ordinary skill in the art, at the time of applicant’s invention, to combine modified Wolverton with Xu’s feature(s) listed above. One would’ve been motivated to do so in order to generate question/answer data 110 representing the retrieved question/answer pair(s) (Xu; [0069]). By incorporating the teachings of Xu, one would’ve been able to generate a response.
Regarding claim 10/17: Wolverton teaches:
wherein the vehicle system controller is configured to: receive at least one condition associated with the vehicle; ([0049] The illustrative vehicle-specific conversation model 132 also includes an acoustic model that is appropriate for the in-vehicle environment as well as other possible environments in which the vehicle personal assistant 112 may be used (such as a garage, a car wash, a line at a drive-thru restaurant, etc.). As such, the acoustic model is configured to account for noise and channel conditions that are typical of these environments, including road noise as well as background noise (e.g., radio, video, or the voices of other vehicle occupants or other persons outside the vehicle).).
Wolverton doesn’t teach:
and generate the response specific to the vehicle based on the vehicle information, the plurality of embeddings, and the at least one condition associated with the vehicle in response to the query.
Xu teaches:
and generate the response specific to the vehicle based on the vehicle information, the plurality of embeddings, and the at least one condition associated with the vehicle in response to the query. ([0071] The method 700, at block B708, may include determining, using the language model, an output associated with the question. For instance, the language model(s) 124 may process the text data 106 and the question/answer data 110 (e.g., the prompt data 122) and, based on the processing, output data 126 associated with the question. As described herein, the output data 126 may represent information associated with the question. The vehicle may then provide the information to the passenger. In some examples, the vehicle provides the information by outputting sound associated with the output data 126, where the sound includes one or more words representing the information. In some examples, the vehicle provides the information by displaying content associated with the output data 126, where the content includes one or more words representing the information.; [0057] While the example of FIG. 5 illustrates one example technique that the retrieval component 114 may use to retrieve the contextual information, in other examples, the retrieval component 114 may use additional and/or alternative techniques. For a first example, and as described herein, the information may be separated into different categories. For instance, and using the example of FIG. 5 where the information 504 is from the OEM manual 502, the information 504 may be separated into component categories, such as tires, motor, doors, windows, and/or the like. In such examples, the retrieval component 114 may use the categories to retrieve contextual information that is in a similar category as the transcript 206 represented by the text data 204. For a second example, the retrieval component 114 may match one or more words represented by the text data 204 to one or more words represented by the information 504. The retrieval component 114 may then retrieve contextual information that includes at least a threshold number (e.g., one, two, three, five, ten, etc.) of matching words.; [0082] Controller(s) 936, which may include one or more system on chips (SoCs) 904 (FIG. 9C) and/or GPU(s), may provide signals (e.g., representative of commands) to one or more components and/or systems of the vehicle 900. For example, the controller(s) may send signals to operate the vehicle brakes via one or more brake actuators 948, to operate the steering system 954 via one or more steering actuators 956, to operate the propulsion system 950 via one or more throttle/accelerators 952. The controller(s) 936 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals, and output operation commands (e.g., signals representing commands) to enable autonomous driving and/or to assist a human driver in driving the vehicle 900. The controller(s) 936 may include a first controller 936 for autonomous driving functions, a second controller 936 for functional safety functions, a third controller 936 for artificial intelligence functionality (e.g., computer vision), a fourth controller 936 for infotainment functionality, a fifth controller 936 for redundancy in emergency conditions, and/or other controllers. In some examples, a single controller 936 may handle two or more of the above functionalities, two or more controllers 936 may handle a single functionality, and/or any combination thereof.).
It would have been obvious to one of ordinary skill in the art, at the time of applicant’s invention, to combine modified Wolverton with Xu’s feature(s) listed above. One would’ve been motivated to do so in order to control one or more components and/or systems of the vehicle 900 (Xu; [0083]). By incorporating the teachings of Xu, one would’ve been able to consider a condition of the vehicle.
Regarding claim 11: Wolverton teaches:
wherein the at least one condition includes at least one of a vehicle condition detected by a vehicle sensor, a weather condition received from an external source, or a road condition received from an external source. ([0023] The vehicle personal assistant may determine a current vehicle-related context of the further human-generated input based on the real-time sensor input, and may interpret the further human-generated input based on the current vehicle-related context of the further human-generated input. The real-time sensor input may indicate a current status of a feature of the vehicle. The real-time sensor input may provide information about a current driving situation of the vehicle, including vehicle location, vehicle speed, vehicle acceleration, fuel status, and/or weather information.).
Regarding claim 16: Wolverton doesn’t teach:
Xu teaches:
the plurality of embeddings for the vehicle specific user manuals are a plurality of first embeddings; ([0003] For a second example, the systems and methods may use the retrieval system(s) to retrieve contextual information related to the speech, such as contextual information from a (fixed or live) text-based knowledge base—such as a manual, a vehicle manual, a machine manual, a document, etc.—that is stored in the database(s).);
the vehicle control method further includes receiving a plurality of previous input queries and a corresponding plurality of responses to the previous input queries, ([0032] various functions may be carried out by a processor executing instructions stored in memory.; [0099] The vehicle 900 may include a system(s) on a chip (SoC) 904. The SoC 904 may include CPU(s) 906, GPU(s) 908, processor(s) 910, cache(s) 912, accelerator(s) 914, data store(s) 916, and/or other components and features not illustrated. The SoC(s) 904 may be used to control the vehicle 900 in a variety of platforms and systems. For example, the SoC(s) 904 may be combined in a system (e.g., the system of the vehicle 900) with an HD map 922 which may obtain map refreshes and/or updates via a network interface 924 from one or more servers (e.g., server(s) 978 of FIG. 9D).; [0062] As further shown, the language model(s) 124 may further output contextual data 120, which, in some examples, may include at least a portion of the output data 126. As discussed herein, the prompt component 118 may further use at least a portion of the contextual data 120 to generate the prompt data 122. For example, if the user continues to ask questions associated with the vehicle, the prompt component 118 may use the contextual data 120 to continue generating the prompt data 122 for the questions, where the contextual data 120 represents a context associated with outputs to previous questions.; [0070] the vehicle may input additional data into the language model(s) 124, such as contextual data 116 generated by the retrieval component 114 and/or contextual data 120 previously output by the language model(s) 124.);
determining subsets of frequently asked queries from the input queries, the subsets of frequently asked queries specific to different vehicle models, ([0001] For a conversational assistant to operate within a vehicle, the conversational assistant may be preloaded with a set of answers to a set of questions that are commonly asked by passengers.; [0036] Referring back to FIG. 1, the process 100 may include a retrieval component 108 generating question/answer data 110 associated with the text data 106. For instance, an information database(s) 112 may store a number of question/answer pairs. As described herein, the number of question/answer pairs may include, but is not limited to, one question/answer pair, one hundred question/answer pairs, five hundred question/answer pairs, one thousand question/answer pairs, and/or any other number of question/answer pairs. In some examples, the question/answer pairs may be associated with a specific type of vehicle, such as a vehicle manufacturer, a vehicle model, and/or a vehicle year. For instance, the question/answer pairs may be generated using a knowledge base—such as an OEM manual—associated with the vehicle. In some examples, the question/answer pairs may be associated with more than one type of vehicle. For instance, the question/answer pairs may include general questions and answers associated with different vehicle manufacturers, different vehicle models, and/or different vehicle years. Still, in some examples, the question/answer pairs may be associated with topics other than vehicles.; [0037] FIG. 3 illustrates an example of question/answer pairs for vehicles);
and generating a plurality of second embeddings for the subsets of frequently asked queries specific to the different vehicle models and their corresponding responses; ([0004 the current systems, in some embodiments, use a language model(s) (e.g., a large language model(s)) to generate outputs that are more natural, conversational, robust, scalable, and accurate.; [0023] The system(s) may then use a retrieval system(s) to retrieve, from the database(s) (or data stores, or other storage or memory types), one or more question/answer pairs that are related to the text data. In some examples, to retrieve the question/answer pair(s), the question/answer pairs stored within the database(s) may be associated with embeddings. For instance, a first question/answer pair may be associated with a first embedding, a second question/answer pair may be associated with a second embedding, a third question/answer pair may be associated with a third embedding, and/or so forth.; [0025] In some embodiments, in addition to or alternatively from storing and/or retrieving question/answer pairs, the system may store intents, sub-intents, tokens, or other classification types corresponding to different answer types. In such examples, embeddings may be matched to a closest intent (e.g., a “tire pressure intent,” a “open gas compartment intent,” etc.), and this information may be used to determine an answer.; [0033] The process 100 may include one or more speech-processing components 102 processing audio data 104. For instance, the vehicle may generate the audio data 104 using one or more microphone(s), where the audio data 104 represents speech (e.g., an utterance) from a user of the vehicle. In some examples, the speech may represent a task being requested by the user, such as a question about the vehicle. As described herein, when the task includes a question, the question may be associated with a feature (e.g., a radio, a display, etc.) of the vehicle, a component (e.g., a window, a door, a tire, an engine, etc.) of the vehicle, a maintenance schedule associated with the vehicle, and/or any other aspect of the vehicle. The vehicle may then process the audio data 104 using the speech-processing component(s) 102. As described herein, the speech-processing component(s) 102 may include, but is not limited to, one or more ASR models, one or more STT models, one or more NLP models, and/or any other type of speech model.
and generating the response specific to the vehicle includes generating the response based on the vehicle information, the plurality of first embeddings, and the plurality of second embeddings in response to the input query. ([0071] The method 700, at block B708, may include determining, using the language model, an output associated with the question. For instance, the language model(s) 124 may process the text data 106 and the question/answer data 110 (e.g., the prompt data 122) and, based on the processing, output data 126 associated with the question. As described herein, the output data 126 may represent information associated with the question. The vehicle may then provide the information to the passenger. In some examples, the vehicle provides the information by outputting sound associated with the output data 126, where the sound includes one or more words representing the information. In some examples, the vehicle provides the information by displaying content associated with the output data 126, where the content includes one or more words representing the information.; [0057] While the example of FIG. 5 illustrates one example technique that the retrieval component 114 may use to retrieve the contextual information, in other examples, the retrieval component 114 may use additional and/or alternative techniques. For a first example, and as described herein, the information may be separated into different categories. For instance, and using the example of FIG. 5 where the information 504 is from the OEM manual 502, the information 504 may be separated into component categories, such as tires, motor, doors, windows, and/or the like. In such examples, the retrieval component 114 may use the categories to retrieve contextual information that is in a similar category as the transcript 206 represented by the text data 204. For a second example, the retrieval component 114 may match one or more words represented by the text data 204 to one or more words represented by the information 504. The retrieval component 114 may then retrieve contextual information that includes at least a threshold number (e.g., one, two, three, five, ten, etc.) of matching words.).
It would have been obvious to one of ordinary skill in the art, at the time of applicant’s invention, to combine modified Wolverton with Xu’s feature(s) listed above. One would’ve been motivated to do so in order to generate an embedding for the transcript (Xu; [0023]), and to generate question/answer data 110 representing the retrieved question/answer pair(s) (Xu; [0069]). By incorporating the teachings of Xu, one would’ve been able to generate embeddings for the questions and responses, and generate a response for the user’s query.
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Wolverton et al. (US 20140136187 A1, hereinafter “Wolverton”), in view of Shetty et al. (US 20260100042 A1, hereinafter “Shetty”), in further view of Xu et al. (US 20240095460 A1, hereinafter “Xu”) as applied to claim 7 above, in further view of Megerian et al. (US 20200272439 A1, hereinafter “Megerian”).
Regarding claim 9: Wolverton doesn’t teach:
wherein the plurality of responses to the previous queries include at least one response annotated by a technician.
Megerian teaches:
wherein the plurality of responses to the previous queries include at least one response annotated by a technician. ([0009] Various Graphical User Interfaces (GUI) may be presented in various formats to present ranked options that users may select from. As discussed herein, a paradigm (also referred to as a diagnosis paradigm) is a logical structure that is used to identify a condition (or several conditions) that may be addressed by one or more action plans. The paradigms provide rankings for the various action plans and identify which action plans are to be displayed in a GUI and what order those action plans are to be presented in. For example, a maintenance technician may be presented with a GUI that shows several action plans for troubleshooting a nonconformance in a device or structure (e.g., an air conditioner, a building, a computer, a vehicle) and several paradigms (e.g., a manufacturer's recommendation, a company policy, previous technician's notes) that recommend one action plan over another action plan.).
It would have been obvious to one of ordinary skill in the art, at the time of applicant’s invention, to combine modified Wolverton with Megerian’s feature(s) listed above. One would’ve been motivated to do so in order to display the highest ranked action plans and/or the highest ranked paradigms more prominently than lower ranked action plans and/or paradigms (i.e., highlighting a preferred option), but users may select lower-ranked action plans or may select action plans based on lower ranked paradigms (Megerian; [0010]). By incorporating the teachings of Megerian, one would’ve been able to use a response annotated by a technician.
Claims 12 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Wolverton et al. (US 20140136187 A1, hereinafter “Wolverton”), in view of Shetty et al. (US 20260100042 A1, hereinafter “Shetty”) as applied to claims 1, and 13 above, in further view of Ricci (US 20190279447 A1, hereinafter “Ricci”).
Regarding claim 12/18: Wolverton teaches:
wherein the vehicle system controller is configured to: receive a plurality of queries; ([0002] According to at least one aspect of this disclosure, a conversational spoken natural language-enabled interactive vehicle user's guide embodied in one or more machine readable storage media is executable by a computing system to receive human-generated spoken natural language input relating to a component of a vehicle.);
identify a set of related queries from the plurality of queries for which no response is available; ([0012] The spoken natural language-enabled interactive vehicle user's guide may determine whether further human-generated input is needed to respond to the human-generated conversational spoken natural language input and solicit further human-generated input in response to determining that further human-generated input is needed. Examiner notes that one of ordinary skill in the art would’ve reasonably considered the determination that more information is needed in order to provide a response as equivalent to having no response for the query.).
Wolverton doesn’t teach:
and in response to the set of related input queries being greater than a threshold, generate a notification for a manufacturer indicating a desired vehicle feature based on the related input queries.
Ricci teaches:
and in response to the set of related queries being greater than a threshold, generate a notification for a manufacturer indicating a desired vehicle feature based on the related queries. ([0350] One or more warnings may be stored in portion 1286. The warnings data 1286 may include warning generated by the vehicle 104, systems of the vehicle 104, manufacturer of the vehicle, federal agency, third party, and/or a user associated with the vehicle. For example, several components of the vehicle may provide health status information (e.g., stored in portion 1278) that, when considered together, may suggest that the vehicle 104 has suffered some type of damage and/or failure. Recognition of this damage and/or failure may be stored in the warnings data portion 1286. The data in portion 1286 may be communicated to one or more parties (e.g., a manufacturer, maintenance facility, user, etc.).; [Fig. 30] method 3000).
It would have been obvious to one of ordinary skill in the art, at the time of applicant’s invention, to combine modified Wolverton with Ricci’s feature(s) listed above. One would’ve been motivated to do so in order to obtain or track health status data of the systems and/or components in portion 1278 (Ricci; [0349]). By incorporating the teachings of Ricci, one would’ve been able to notify the manufacturer when the input queries exceed a threshold.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GABRIEL J TORRES CHANZA whose telephone number is (571)272-3701. The examiner can normally be reached Monday thru Friday 8am - 5pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Brian Epstein can be reached on (571)270-5389. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/G.J.T./Examiner, Art Unit 3625
/BRIAN M EPSTEIN/Supervisory Patent Examiner, Art Unit 3625