DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Applicant claims the benefit of US Provisional Application No. 63/587,726, filed 10/04/2023. Claims 1-20 have been afforded the benefit of this filing date.
Information Disclosure Statement
The IDS dated 02/05/2026 has been considered and placed in the application file.
Claim Objections
Claims 1-20 are objected to because of the following informalities:
Claim 1 and Claim 11 are objected to because "A vehicle-mounted apparatus, installed on a vehicle a user ridden in " in Line 1 is unclear. It is unclear whether a user in currently riding in the vehicle or has previously ridden in the vehicle. The examiner suggests that the claim be changed to read " A vehicle-mounted apparatus, installed on a vehicle in which a user is currently riding in " Claims 2-10 and 12-20 depend on Claims 1 and 11 therefore they are also objected to. Appropriate correction is required.
Claim Interpretation
The claims in this application are given their broadest reasonable interpretation using the
plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification.
Under MPEP 2111.04, " Claim scope is not limited by claim language that suggests or makes optional but does not require steps to be performed, or by claim language that does not limit a claim to a particular structure. However, examples of claim language, although not exhaustive, that may raise a question as to the limiting effect of the language in a claim are:[AltContent: rect]
(A) "adapted to" or "adapted for" clauses;
(B) "wherein" clauses; and
(C) "whereby" clauses.
The determination of whether each of these clauses is a limitation in a claim depends on the specific facts of the case. See, e.g., Griffin v. Bertina, 285 F.3d 1029, 1034, 62 USPQ2d 1431 (Fed. Cir. 2002) (finding that a "wherein" clause limited a process claim where the clause gave "meaning and purpose to the manipulative steps"). In In re Giannelli, 739 F.3d 1375, 1378, 109 USPQ2d 1333, 1336 (Fed. Cir. 2014), the court found that an "adapted to" clause limited a machine claim where "the written description makes clear that 'adapted to,' as used in the [patent] application, has a narrower meaning, viz., that the claimed machine is designed or constructed to be used as a rowing machine whereby a pulling force is exerted on the handles." In Hoffer v. Microsoft Corp., 405 F.3d 1326, 1329, 74 USPQ2d 1481, 1483 (Fed. Cir. 2005), the court held that when a "‘whereby’ clause states a condition that is material to patentability, it cannot be ignored in order to change the substance of the invention." Id. However, the court noted that a "‘whereby clause in a method claim is not given weight when it simply expresses the intended result of a process step positively recited.’" Id. (quoting Minton v. Nat’l Ass’n of Securities Dealers, Inc., 336 F.3d 1373, 1381, 67 USPQ2d 1614, 1620 (Fed. Cir. 2003))”.
Claim 11 recite " A control method being adapted for use in a vehicle-mounted apparatus”. Since “adapted for use” raises a question as to the limiting effect of the language in a claim, the claim scope is not limited by claim language that suggests or makes optional but does not require steps to be performed, or by claim language that does not limit a claim to a particular structure. The broadest reasonable interpretation of a system (or apparatus or product) claim having structure that performs a function, which only needs to occur if a condition precedent is met, requires structure for performing the function should the condition occur. The system claim interpretation differs from a method claim interpretation because the claimed structure must be present in the system regardless of whether the condition is met and the function is actually performed is being adopted for the purposes of this Office Action. Applicant’s comments and/or amendments relating to this issue are invited to clarify the claim language and the prosecution history.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or no obviousness.
Claims 1-20 are rejected under 35 U.S.C. 103 as unpatentable over Wu et al (US Patent Publication US 2025/0200283 A1, hereafter referred to as Wu) in view of Bucker et al (Bucker, Arthur, et al. "Latte: Language trajectory transformer." arXiv preprint arXiv:2208.02918 (2022)).
Regarding Claim 1, Wu teaches a vehicle-mounted apparatus, installed on a vehicle a user ridden in (Wu ¶0259, ¶0260, and Fig 16A disclose a vehicle mounted system installed in a car which a user would be in), comprising:
a camera (Wu Fig 16A and 16B discloses a surround camera attached to the vehicle), configured to capture an environment image around the vehicle (Wu Fig 16B and ¶0034-¶0035 discloses the cameras capturing the surrounding areas of the vehicle including the observations of the environment around the vehicle);
an output interface (Wu ¶0138 and Fig 6 discloses a graphical user interface that processing information for presentation to the user); and
a processor, communicatively connected to the camera and the output interface relatively (Wu Fig 6 and Fig 10 and ¶0174-¶0177 disclose the processor connected to the camera and display), and configured to execute (Wu ¶0174-¶0177 and ¶0182 discloses a processor executing instructions) the following operations:
inputting the environment image (Wu ¶0038 discloses capturing the environment through sensor data including image data, ¶0072, ¶0077 discloses extracting the relevant information from the image) into a composite model (Wu Fig 3 and Fig 4A discloses both perception data and aligned map points being input into a model, therefore the examiner is interpreting that the model can be considered a composite model as it has two different types of data input) to generate an environment text (Wu Fig 4A and ¶0104 discloses the environment observation being input into a model to receive a tokenized description of the region) wherein the environment text is configured to describe the environment image (Wu Fig 5G, 584, ¶0035, discloses the tokenized description being a text string describing the environment);
inputting the environment text (Wu Fig 4A, 412 discloses providing the aligned map data and the perception data as input to a language model)
to generate a first response text corresponding to the vehicle (Wu ¶0064, ¶0207 discloses the environment generation process can then generate environments automatically, in response to human prompts, and modifications to the environment can be made relatively quickly and without significant processing through updating of the tokenized text string); and
generating a control signal (Wu Fig 3, 316 and ¶0080 discloses the control signal based on the tokenized description of the environment) corresponding to the first response text (Wu ¶0064,¶0207, discloses the environment generation process can then generate environments automatically, in response to human prompts, and modifications to the environment can be made relatively quickly and without significant processing through updating of the tokenized text string) to control the output interface (Wu ¶0138 and Fig 6 discloses a graphical user interface that processing information for presentation to the user) to execute an interactive operation corresponding to the user (Wu ¶0056 ¶0233, ¶0241 discloses generating a virtual interactive display or environment interaction by users of a system and interacting with respect to objects in the environment).
Wu does not explicitly disclose into a first language model.
Bucker is in the same field of using language based frameworks to interpret user input. Further, Bucker teaches into a first language model (Bucker Fig 2 and Abstract discloses BERT a first language model).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Wu by incorporating the implementation of the network architecture including two language models that include encoder and decoder capabilities connected to a transformer model as taught by Bucker; to make an invention that can use the network architecture to increase the efficiency of the response to the user; thus one of ordinary skilled in the art would be motivated to combine the references since there is a need for a robot or autonomous entity to have the ability to recognize and understand natural language commands in a given context and map them to the task domain space – where tasks and constraints are largely influenced by context, intent and affordances with objects (Bucker, Abstract and Introduction).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
Regarding Claim 2, Wu in view of Bucker teaches the vehicle-mounted apparatus of claim 1, wherein the composite model (Wu Fig 3 and Fig 4A discloses both perception data and aligned map points being input into a model, therefore the examiner is interpreting that the model can be considered a composite model as it has two different types of data input) is configured to:
input the environment image (Wu ¶0038 discloses capturing the environment through sensor data including image data, ¶0072, ¶0077 discloses extracting the relevant information from the image) into an image encoder to generate a plurality of image features (Wu Fig 2, 210, ¶0053, ¶0058 discloses an feature extraction module can include an encoder that can extract features from the various instances of sensor data and encode those features as embeddings or points in a latent space);
input the image features (Wu Fig 5H discloses inputting features and generating a set of feature vectors corresponding to objects) and a query (Wu ¶0069, ¶0111, ¶0131 discloses a query used to direct the output of the LLM) into a transformer model (Bucker Fig 2 and Abstract and Section II disclose the use of a transformer encoder and decoder pair) to generate a plurality of extracted features corresponding to a feature vector (Wu Fig 5H discloses inputting features and generating a set of feature vectors corresponding to objects), wherein the query comprises the feature vector (Wu ¶0069, ¶0111, ¶0131 and ¶0133 discloses a query used to direct the output of the LLM where the query relates to the feature vectors derived from the image); and
input the extracted features (Wu Fig 5H discloses inputting features and generating a set of feature vectors corresponding to objects) into a second language model (Bucker Fig 2 and Abstract discloses CLIP a second language model) to generate the environment text (Wu Fig 5H discloses output of the trained language model being a tokenized text string, corresponding to the objects, and token descriptors including semantic information about the objects). See Claim 1 for rationale, its parent claim.
Regarding Claim 3, Wu in view of Bucker teaches the vehicle-mounted apparatus of claim 2, wherein the composite model (Wu Fig 3 and Fig 4A discloses both perception data and aligned map points being input into a model, therefore the examiner is interpreting that the model can be considered a composite model as it has two different types of data input) is further configured to:
in response to receiving a real-time coordinate corresponding to the vehicle (Wu Fig 4A 406 and 408, ¶0042, ¶0070, ¶0108 disclose identifying the geolocation of the vehicle including coordinates) capturing the environment image (Wu Fig 16B and ¶0034-¶0035 discloses the cameras capturing the surrounding areas of the vehicle including the observations of the environment around the vehicle), input the real-time coordinate (Wu Fig 4A 406 and 408, ¶0042, ¶0070, ¶0108 disclose identifying the geolocation of the vehicle including coordinates), the image features (Wu Fig 5H discloses inputting features and generating a set of feature vectors corresponding to objects), and the query (Wu ¶0069, ¶0111, ¶0131 and ¶0133 discloses a query used to direct the output of the LLM where the query relates to the feature vectors derived from the image) into the transformer model (Bucker Fig 2 and Abstract and Section II disclose the use of a transformer encoder and decoder pair) to generate the extracted features (Wu ¶0059, ¶0118 discloses generating extracted features). See Claim 1 for rationale, its parent claim.
Regarding Claim 4, Wu in view of Bucker teaches the vehicle-mounted apparatus of claim 2, further comprising:
an input interface (Wu Fig 9, 925, ¶0170 discloses user input interfaces), configured to generate an input data corresponding to the user (Wu Fig 2A 218 discloses a client device that provides additional input and ¶0064 discloses the examples of input given by the user);
wherein the composite model (Wu Fig 3 and Fig 4A discloses both perception data and aligned map points being input into a model, therefore the examiner is interpreting that the model can be considered a composite model as it has two different types of data input) is further configured to:
in response to receiving the input data (Wu Fig 2A 218 discloses a client device that provides additional input and ¶0064 discloses the examples of input given by the user), generate the query based on the input data and the feature vector (Wu ¶0069 discloses basing the search on the query provided by the user and the feature vectors); and
input the image features (Wu Fig 5H discloses inputting features and generating a set of feature vectors corresponding to objects) and the query (Wu ¶0069, ¶0111, ¶0131 discloses a query used to direct the output of the LLM) into the transformer model (Bucker Fig 2 and Abstract and Section II disclose the use of a transformer encoder and decoder pair) to generate the extracted features (Wu ¶0059, ¶0118 discloses generating extracted features). See Claim 1 for rationale, its parent claim.
Regarding Claim 5, Wu in view of Bucker teaches the vehicle-mounted apparatus of claim 2, wherein the second language model (Bucker Fig 2 and Abstract discloses CLIP a second language model) is a decoder corresponding to the image encoder (Bucker Fig 2 and Abstract and Section II disclose the use of a transformer encoder and decoder pair) and the transformer model (Bucker Fig 2 and Abstract and Section II disclose the use of a transformer encoder and decoder pair). See Claim 1 for rationale, its parent claim.
Regarding Claim 6, Wu in view of Bucker teaches the vehicle-mounted apparatus of claim 1, wherein the processor is further configured to execute (Wu ¶0174-¶0177 and ¶0182 discloses a processor executing instructions) the following operations:
transforming a real-time coordinate corresponding to the vehicle (Wu Fig 4A 406 and 408, ¶0042, ¶0070, ¶0108 disclose identifying the geolocation of the vehicle including coordinates) capturing the environment image (Wu Fig 16B and ¶0034-¶0035 discloses the cameras capturing the surrounding areas of the vehicle including the observations of the environment around the vehicle) into a location text (Wu ¶0038, ¶0047, ¶0055, discloses the input to the model being location and the map reflecting the location) ;and
inputting the location text (Wu ¶0038, ¶0047, ¶0055, discloses the input to the model being location and the map reflecting the location) and the environment text (Wu Fig 5G, 584, ¶0035, discloses the tokenized description being a text string describing the environment) into the first language model (Bucker Fig 2 and Abstract discloses BERT a first language model) (Wu Fig 4A, 412 discloses providing the aligned map data and the perception data as input to a language model) to generate the first response text (Wu Fig 4B discloses the output being a tokenized description of at least a region of the physical environment). See Claim 1 for rationale, its parent claim.
Regarding Claim 7, Wu in view of Bucker teaches the vehicle-mounted apparatus of claim 1, further comprising:
an input interface (Wu Fig 9, 925, ¶0170 discloses user input interfaces), configured to generate an input text corresponding to the user (Wu Fig 2A 218 discloses a client device that provides additional input and ¶0064 discloses the examples of input given by the user which can include text);
wherein the first language model (Bucker Fig 2 and Abstract discloses BERT a first language model) is further configured to:
in response to receiving the input text (Wu Fig 2A 218 discloses a client device that provides additional input and ¶0064 discloses the examples of input given by the user which can include text), generate the first response text (Wu ¶0064, ¶0207 discloses the environment generation process can then generate environments automatically, in response to human prompts, and modifications to the environment can be made relatively quickly and without significant processing through updating of the tokenized text string) based on the input text (Wu Fig 2A 218 discloses a client device that provides additional input and ¶0064 discloses the examples of input given by the user which can include text), wherein the first response text responds to the input text (Wu ¶0064, ¶0207 discloses the environment generation process can then generate environments automatically, in response to human prompts, and modifications to the environment can be made relatively quickly and without significant processing through updating of the tokenized text string). See Claim 1 for rationale, its parent claim.
Regarding Claim 8, Wu in view of Bucker teaches the vehicle-mounted apparatus of claim 1, wherein the first language model (Bucker Fig 2 and Abstract discloses BERT a first language model) is further configured to:
generate a driving suggestion for operating the vehicle (Wu ¶0081, ¶0134 discloses a semi-autonomous vehicle operating based on the perception data) based on the environment text (Wu Fig 4A and ¶0104 discloses the environment observation being input into a model to receive a tokenized description of the region); and
take the driving suggestion (Wu ¶0081, ¶0134, discloses a semi-autonomous vehicle operating based on the perception data) as the first response text (Wu ¶0064, ¶0207 discloses the environment generation process can then generate environments automatically, in response to human prompts, and modifications to the environment can be made relatively quickly and without significant processing through updating of the tokenized text string). See Claim 1 for rationale, its parent claim.
Regarding Claim 9, Wu in view of Bucker teaches the vehicle-mounted apparatus of claim 1, wherein the first language model (Bucker Fig 2 and Abstract discloses BERT a first language model) is further configured to:
generate an interactive data corresponding to the vehicle (Wu ¶0073 discloses identifying objects for interaction based on their location in the environment and ¶0232 discloses a virtual interactive display with users) based on the environment text (Wu Fig 4A and ¶0104 discloses the environment observation being input into a model to receive a tokenized description of the region); and
take the interactive data (Wu ¶0073 discloses identifying objects for interaction based on their location in the environment and ¶0232 discloses a virtual interactive display with users) as the first response text (Wu ¶0064, ¶0207 discloses the environment generation process can then generate environments automatically, in response to human prompts, and modifications to the environment can be made relatively quickly and without significant processing through updating of the tokenized text string). See Claim 1 for rationale, its parent claim.
Regarding Claim 10, Wu in view of Bucker teaches the vehicle-mounted apparatus of claim 9, further comprising:
an input interface (Wu Fig 9, 925, ¶0170 discloses user input interfaces), configured to generate an input text corresponding to the user (Wu Fig 2A 218 discloses a client device that provides additional input and ¶0064 discloses the examples of input given by the user);
wherein the first language model (Bucker Fig 2 and Abstract discloses BERT a first language model) is further configured to:
after generating the interactive data (Wu ¶0073 discloses identifying objects for interaction based on their location in the environment and ¶0232 discloses a virtual interactive display with users), in response to receiving the input text (Wu Fig 2A 218 discloses a client device that provides additional input and ¶0064 discloses the examples of input given by the user), generate a second response text based on the input text and the interactive data (Wu ¶0064 discloses generate environments automatically, in response to human prompts, or through a combination of both. Modifications to the environment can be made relatively quickly and without significant processing through updating of the tokenized text string), wherein the second response text responds to the input text (Wu ¶0064 Modifications to the environment can be made relatively quickly and without significant processing through updating of the tokenized text string). See Claim 1 for rationale, its parent claim.
Regarding Claim 11, Wu teaches a control method (Wu Fig 4A, 416, ¶0032, ¶0039 discloses a control system and methods for determining operations for the vehicle) being adapted for use in a vehicle-mounted apparatus, wherein the vehicle-mounted apparatus is installed on a vehicle a user ridden in (Wu ¶0259, ¶0260, and Fig 16A disclose a vehicle mounted system installed in a car which a user would be in), and the control method (Wu Fig 4A, 416, ¶0032, ¶0039 discloses a control system and methods for determining operations for the vehicle) comprises the following steps:
capturing an environment image around the vehicle (Wu Fig 16B and ¶0034-¶0035 discloses the cameras capturing the surrounding areas of the vehicle including the observations of the environment around the vehicle);
inputting the environment image (Wu ¶0038 discloses capturing the environment through sensor data including image data, ¶0072, ¶0077 discloses extracting the relevant information from the image) into a composite model (Wu Fig 3 and Fig 4A discloses both perception data and aligned map points being input into a model, therefore the examiner is interpreting that the model can be considered a composite model as it has two different types of data input) to generate an environment text (Wu Fig 4A and ¶0104 discloses the environment observation being input into a model to receive a tokenized description of the region) wherein the environment text is configured to describe the environment image (Wu Fig 5G, 584, ¶0035, discloses the tokenized description being a text string describing the environment);
inputting the environment text (Wu Fig 4A, 412 discloses providing the aligned map data and the perception data as input to a language model)
to generate a first response text corresponding to the vehicle (Wu ¶0064, ¶0207 discloses the environment generation process can then generate environments automatically, in response to human prompts, and modifications to the environment can be made relatively quickly and without significant processing through updating of the tokenized text string); and
executing an interactive operation corresponding to the user (Wu ¶0056 ¶0233, ¶0241 discloses generating a virtual interactive display or environment interaction by users of a system and interacting with respect to objects in the environment) based on the first response text (Wu ¶0064,¶0207, discloses the environment generation process can then generate environments automatically, in response to human prompts, and modifications to the environment can be made relatively quickly and without significant processing through updating of the tokenized text string).
Wu does not explicitly disclose into a first language model.
Bucker is in the same field of using language based frameworks to interpret user input. Further, Bucker teaches into a first language model (Bucker Fig 2 and Abstract discloses BERT a first language model).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Wu by incorporating the implementation of the network architecture including two language models that include encoder and decoder capabilities connected to a transformer model as taught by Bucker; to make an invention that can use the network architecture to increase the efficiency of the response to the user; thus one of ordinary skilled in the art would be motivated to combine the references since there is a need for a robot or autonomous entity to have the ability to recognize and understand natural language commands in a given context and map them to the task domain space – where tasks and constraints are largely influenced by context, intent and affordances with objects (Bucker, Abstract and Introduction).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
Regarding Claim 12, Wu in view of Bucker teaches the control method of claim 11, wherein the composite model (Wu Fig 3 and Fig 4A discloses both perception data and aligned map points being input into a model, therefore the examiner is interpreting that the model can be considered a composite model as it has two different types of data input) is configured to:
input the environment image (Wu ¶0038 discloses capturing the environment through sensor data including image data, ¶0072, ¶0077 discloses extracting the relevant information from the image) into an image encoder to generate a plurality of image features (Wu Fig 2, 210, ¶0053, ¶0058 discloses an feature extraction module can include an encoder that can extract features from the various instances of sensor data and encode those features as embeddings or points in a latent space);
input the image features (Wu Fig 5H discloses inputting features and generating a set of feature vectors corresponding to objects) and a query (Wu ¶0069, ¶0111, ¶0131 discloses a query used to direct the output of the LLM) into a transformer model (Bucker Fig 2 and Abstract and Section II disclose the use of a transformer encoder and decoder pair) to generate a plurality of extracted features corresponding to a feature vector (Wu Fig 5H discloses inputting features and generating a set of feature vectors corresponding to objects), wherein the query comprises the feature vector (Wu ¶0069, ¶0111, ¶0131 and ¶0133 discloses a query used to direct the output of the LLM where the query relates to the feature vectors derived from the image); and
input the extracted features (Wu Fig 5H discloses inputting features and generating a set of feature vectors corresponding to objects) into a second language model (Bucker Fig 2 and Abstract discloses CLIP a second language model) to generate the environment text (Wu Fig 5H discloses output of the trained language model being a tokenized text string, corresponding to the objects, and token descriptors including semantic information about the objects). See Claim 11 for rationale, its parent claim.
Regarding Claim 13, Wu in view of Bucker teaches the control method of claim 12, wherein the composite model (Wu Fig 3 and Fig 4A discloses both perception data and aligned map points being input into a model, therefore the examiner is interpreting that the model can be considered a composite model as it has two different types of data input) is further configured to:
in response to receiving a real-time coordinate corresponding to the vehicle (Wu Fig 4A 406 and 408, ¶0042, ¶0070, ¶0108 disclose identifying the geolocation of the vehicle including coordinates) capturing the environment image (Wu Fig 16B and ¶0034-¶0035 discloses the cameras capturing the surrounding areas of the vehicle including the observations of the environment around the vehicle), input the real-time coordinate (Wu Fig 4A 406 and 408, ¶0042, ¶0070, ¶0108 disclose identifying the geolocation of the vehicle including coordinates), the image features (Wu Fig 5H discloses inputting features and generating a set of feature vectors corresponding to objects), and the query (Wu ¶0069, ¶0111, ¶0131 and ¶0133 discloses a query used to direct the output of the LLM where the query relates to the feature vectors derived from the image) into the transformer model (Bucker Fig 2 and Abstract and Section II disclose the use of a transformer encoder and decoder pair) to generate the extracted features (Wu ¶0059, ¶0118 discloses generating extracted features). See Claim 11 for rationale, its parent claim.
Regarding Claim 14, Wu in view of Bucker teaches the control method of claim 12, wherein the composite model (Wu Fig 3 and Fig 4A discloses both perception data and aligned map points being input into a model, therefore the examiner is interpreting that the model can be considered a composite model as it has two different types of data input) is further configured to:
in response to receiving an input data (Wu Fig 2A 218 discloses a client device that provides additional input and ¶0064 discloses the examples of input given by the user), generate the query based on the input data and the feature vector (Wu ¶0069 discloses basing the search on the query provided by the user and the feature vectors); and
input the image features (Wu Fig 5H discloses inputting features and generating a set of feature vectors corresponding to objects) and the query (Wu ¶0069, ¶0111, ¶0131 discloses a query used to direct the output of the LLM) into the transformer model (Bucker Fig 2 and Abstract and Section II disclose the use of a transformer encoder and decoder pair) to generate the extracted features (Wu ¶0059, ¶0118 discloses generating extracted features). See Claim 11 for rationale, its parent claim.
Regarding Claim 15, Wu in view of Bucker teaches the control method of claim 12, wherein the second language model (Bucker Fig 2 and Abstract discloses CLIP a second language model) is a decoder corresponding to the image encoder (Bucker Fig 2 and Abstract and Section II disclose the use of a transformer encoder and decoder pair) and the transformer model (Bucker Fig 2 and Abstract and Section II disclose the use of a transformer encoder and decoder pair). See Claim 11 for rationale, its parent claim.
Regarding Claim 16, Wu in view of Bucker teaches the control method of claim 11, further comprising:
transforming a real-time coordinate corresponding to the vehicle (Wu Fig 4A 406 and 408, ¶0042, ¶0070, ¶0108 disclose identifying the geolocation of the vehicle including coordinates) capturing the environment image (Wu Fig 16B and ¶0034-¶0035 discloses the cameras capturing the surrounding areas of the vehicle including the observations of the environment around the vehicle) into a location text (Wu ¶0038, ¶0047, ¶0055, discloses the input to the model being location and the map reflecting the location); and
inputting the location text (Wu ¶0038, ¶0047, ¶0055, discloses the input to the model being location and the map reflecting the location) and the environment text (Wu Fig 5G, 584, ¶0035, discloses the tokenized description being a text string describing the environment) into the first language model (Bucker Fig 2 and Abstract discloses BERT a first language model) (Wu Fig 4A, 412 discloses providing the aligned map data and the perception data as input to a language model) to generate the first response text (Wu Fig 4B discloses the output being a tokenized description of at least a region of the physical environment). See Claim 11 for rationale, its parent claim.
Regarding Claim 17, Wu in view of Bucker teaches the control method of claim 11, wherein the first language model (Bucker Fig 2 and Abstract discloses BERT a first language model) is further configured to:
in response to receiving an input text (Wu Fig 2A 218 discloses a client device that provides additional input and ¶0064 discloses the examples of input given by the user which can include text) corresponding to the user (Wu Fig 2A 218 discloses a client device that provides additional input and ¶0064 discloses the examples of input given by the user which can include text), generate the first response text (Wu ¶0064, ¶0207 discloses the environment generation process can then generate environments automatically, in response to human prompts, and modifications to the environment can be made relatively quickly and without significant processing through updating of the tokenized text string) based on the input text (Wu Fig 2A 218 discloses a client device that provides additional input and ¶0064 discloses the examples of input given by the user which can include text), wherein the first response text responds to the input text (Wu ¶0064, ¶0207 discloses the environment generation process can then generate environments automatically, in response to human prompts, and modifications to the environment can be made relatively quickly and without significant processing through updating of the tokenized text string). See Claim 11 for rationale, its parent claim.
Regarding Claim 18, Wu in view of Bucker teaches the control method of claim 11, wherein the first language model (Bucker Fig 2 and Abstract discloses BERT a first language model) is further configured to:
generate a driving suggestion for operating the vehicle (Wu ¶0081, ¶0134, discloses a semi-autonomous vehicle operating based on the perception data) based on the environment text (Wu Fig 4A and ¶0104 discloses the environment observation being input into a model to receive a tokenized description of the region); and
take the driving suggestion (Wu ¶0081, ¶0134, discloses a semi-autonomous vehicle operating based on the perception data) as the first response text (Wu ¶0064, ¶0207 discloses the environment generation process can then generate environments automatically, in response to human prompts, and modifications to the environment can be made relatively quickly and without significant processing through updating of the tokenized text string). See Claim 11 for rationale, its parent claim.
Regarding Claim 19, Wu in view of Bucker teaches the control method of claim 11, wherein the first language model (Bucker Fig 2 and Abstract discloses BERT a first language model) is further configured to:
generate an interactive data corresponding to the vehicle (Wu ¶0073 discloses identifying objects for interaction based on their location in the environment and ¶0232 discloses a virtual interactive display with users) based on the environment text (Wu Fig 4A and ¶0104 discloses the environment observation being input into a model to receive a tokenized description of the region); and
take the interactive data (Wu ¶0073 discloses identifying objects for interaction based on their location in the environment and ¶0232 discloses a virtual interactive display with users) as the first response text (Wu ¶0064, ¶0207 discloses the environment generation process can then generate environments automatically, in response to human prompts, and modifications to the environment can be made relatively quickly and without significant processing through updating of the tokenized text string). See Claim 11 for rationale, its parent claim.
Regarding Claim 20, Wu in view of Bucker teaches the control method of claim 19, wherein the first language model (Bucker Fig 2 and Abstract discloses BERT a first language model) is further configured to:
after generating the interactive data (Wu ¶0073 discloses identifying objects for interaction based on their location in the environment and ¶0232 discloses a virtual interactive display with users) , in response to receiving the input text (Wu Fig 2A 218 discloses a client device that provides additional input and ¶0064 discloses the examples of input given by the user) corresponding to the user (Wu Fig 2A 218 discloses a client device that provides additional input and ¶0064 discloses the examples of input given by the user), generate a second response text based on the input text and the interactive data (Wu ¶0064 discloses generate environments automatically, in response to human prompts, or through a combination of both. Modifications to the environment can be made relatively quickly and without significant processing through updating of the tokenized text string), wherein the second response text responds to the input text (Wu ¶0064 Modifications to the environment can be made relatively quickly and without significant processing through updating of the tokenized text string). See Claim 11 for rationale, its parent claim.
Reference Cited
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure.
US-20250292687-A1 to Pathak et al. discloses using a vision language model to evaluate signs and the environment around vehicles.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RACHEL ROBERTS whose telephone number is (571)272-6413. The examiner can normally be reached Monday- Friday 7:30am- 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Oneal Mistry can be reached on (313) 446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RACHEL L ROBERTS/Examiner, Art Unit 2674
/ONEAL R MISTRY/Supervisory Patent Examiner, Art Unit 2674