DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 05/13/2026 has been entered.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 05/13/2026 was filed in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Response to Arguments
1. Regarding the claim objection, Applicant has amended claim 19 to address the minor informality. Accordingly, the objection is withdrawn.
2. Regarding the rejection under 35 U.S.C. § 112(a), Applicant has amended independent claims 1, 10, and 16 to remove limitations which were identified as new matter in the prior Office Action. Accordingly, the rejection is withdrawn.
3. Regarding the rejection under 35 U.S.C. § 101, Applicant's arguments filed 05/13/2026 have been fully considered but they are not persuasive.
Applicant argues on pgs. 9-18 of the Remarks that the claims are subject matter eligible under 35 U.S.C. 101.
Step 2A Prong 1:
It is argued on pgs. 10-11 that the claims have been improperly characterized as being directed to an abstract idea without significantly more, and that the amended claims are subject matter eligible.
Applicant first argues in section IV.B.a of the Remarks on pgs. 12-13 that amended independent claim 1 does not recite abstract ideas under Step 2A Prong 1 analysis. Specifically, it is argued that the features recited in claim 1 directed to “generating machine-parsable tokenized map representations expressed in a domain specific language, parsing token sequences encoding map topology and object relationships, identifying inconsistencies or omissions within encoded map relationships, and generating textual recommendations based on language model processing for use in autonomous or semi-autonomous machine operations” are not able to be performed practically in the human mind. It is further argued that amended claim 1 is analogous to claims determined as not reciting mental process under the USPTO October 2019 Guidance, and thus also does not recite mental processes. The Examiner respectfully disagrees with these arguments. Amended claim 1 still recites limitations at a broad level which does not preclude them from being performable as mental processes with the aid of pen and paper. Generating a tokenized description of a map in a domain specific language is performable by a person with pen and paper (i.e. given a syntax for a DSL, a person can write down information about objects and relationships in the map into a DSL). Furthermore, a person can identify issues in the map by parsing the tokenized sequence and determining inconsistencies (i.e. a person can read the DSL they wrote down, and can determine that according to the rules of the language, an omission has occurred, such as a missing road, traffic sign, etc.). A person can then write down one or more textual representations/recommendations regarding the issues (i.e. a person writes down “Add a stop sign at the intersection”). Therefore, the claim recites abstract ideas in the form of mental processes under Step 2A Prong 1.
Step 2A Prong 2:
Applicant further argues on pgs. 13-17 that amended claim 1 integrate the judicial exception into a practical application via a technical improvement. In section IV.B.b it is argued that amended claim 1 integrates the judicial exception into a practical application by providing an improvement to machine-based map processing and validation workflows used by autonomous systems. Section IV.B.c provides further arguments that amended claim 1 recites a technical improvement directed towards a technical problem in the art of conventional approaches being unable to effectively reason about complex spatial relationships and topology structures within large-scale map environments, and that the amended claim overcomes this problem via implementation of a computational pipeline involving the tokenized map representations and language-model based parsing and analysis. Section IV.b.d provides further arguments that amended claim 1 provides a technical solution to a problem in the art. The Examiner respectfully disagrees with the arguments presented. Under Step 2A Prong 2, additional elements are analyzed and viewed with the claim as a whole to determine if the judicial exception is integrated into a practical application. The additional elements present in amended claim 1 are the “autonomous or semi-autonomous machine” which uses the map, “a first language model”, and “a second language model”. Each of these elements is recited at a high level of generality (i.e. no specific ML architecture/sub-components/technical processing not performable via a human is recited). Each of these elements merely amounts to applying the identified mental processes using generic computer components. None of the identified additional elements integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract idea. A person can generate domain-specific language code for a map, can identify issues in the map by reading the code, and can write textual recommendations based on the issues. Therefore, amended claim 1 does not reflect a technical improvement as argued by Applicant.
Regarding dependent claims and other independent claims:
Applicant further argues on pgs. 17-18 that independent claims 10 and 16, and dependent claims 2-9, 11-15, and 17-20 are subject matter eligible. The Examiner disagrees with these arguments for analogous reasons as presented above.
Hence, Applicant’s arguments regarding the rejection under 35 U.S.C. 101 are not persuasive.
4. Regarding the rejections under 35 U.S.C. § 103, Applicant's arguments filed 05/13/3036 have been fully considered but they are not persuasive.
Regarding Independent claim 1:
Applicant first argues that the cited references do not teach or suggest all of the limitations of amended claim 1. Specifically, Applicant argues that the following limitations are not taught individually or in combination by the cited references:
“performing one or more planning, navigation, or control operations associated with an autonomous or semi-autonomous machine using a map, wherein at least a portion of the map was generated…”
“generating, based at least on a first language model processing data associated with at least a section of a preliminary map, a tokenized description of the section of the preliminary map represented using a sequence of tokens, expressed in a domain specific language, that encode information about objects or features and relationships therebetween within the section of the preliminary map” and
“identifying, based in part on the tokenized description, one or more potential issues with respect to the section of the preliminary map by parsing the sequence of tokens and determining at least one of an inconsistency in the relationships or an omission of an expected object, feature, or relationship”.
The Examiner respectfully disagrees with Applicant’s arguments.
Regarding the first limitation identified, Applicant argues that Xie does not disclose the planning, navigation, or control operations using a generated map, and further argues that Levinson does not remedy these deficiencies. The Examiner respectfully disagrees. The combined teachings of Xie and Levinson read on this feature. This feature under the BRI requires that the cited prior art teaches some generation of an updated map (i.e. image) for use in autonomous or semi-autonomous machine systems. Levinson teaches this feature. Specifically, Levinson teaches identifying a system which identifies data changes corresponding to current map data, such as a change in external objects with respect to the map data, and in response generates an updated map (para. 0157 “Data change detector 3653 is configured to detect changes in data sets 3655a and 3655b, which are examples of any number of data sets of 3-D map data. Data change detector 3653 also is configured to generate data identifying a portion of map data that has changed, as well as optionally identifying or classifying an object associated with the changed portion of map data. …At time, T2, however, data change detector 3653 may detect that another number of data sets, including data set 3655b, includes data representing the presence of external objects in portions of map data 3665 of 3-D model data 3661, whereby portions of map data 3665 coincide with portions of map data 3664 at different times. Therefore, data change detector 3653 may detect changes in map data, and may further adaptively modify map data to include the changed map data (e.g., as updated map data).”). Levinson further teaches storing this updated map to a repository for use in autonomous vehicle controlling (para. 0160 “A tile generator 3656 may be configured to generate two-dimensional or three-dimensional map tiles based on map data from data sets 3655a and 3655b. The map tiles may be transmitted for storage in map repository 3605a. Tile generator 3656 may generate map tiles that include indicator data for indicating a portion of the map is an updated portion of map data.”; Levinson teaches using the map for autonomous operations: para. 0160 “Tile generator 3656 may generate map tiles that include indicator data for indicating a portion of the map is an updated portion of map data. Further, an updated map portion may be incorporated into a reference data repository 3605 in an autonomous vehicle…”; see para. 0071, map repository data used for localizing autonomous vehicle relative to reference data, see para. 0056, stating localizer is used for causing autonomous vehicles to drive autonomously). Regarding the further claimed steps of generating tokenized descriptions, identifying potential issues, and generating textual recommendations/representations, while Levinson does not teach these limitations, the combination of Levinson with Xie and with newly applied reference Rahman (see below claim mapping under 103 rejection) does teach these steps. Therefore, the first limitation is taught by the cited art.
Regarding the second limitation, Applicant argues that Xie does not teach or suggest generating a tokenized description represented using a sequence of tokens expressed in a domain specific language that encodes relationships between map objects or features within a preliminary map. Specifically, Applicant argues this feature is not disclosed due to the visual cues generated by Xie being human-readable. The Examiner respectfully disagrees. Under the broadest reasonable interpretation, the term “domain specific language” can be any form of tokens which encode relationships and objects/features for a particular domain (i.e. purpose), as claimed in amended claim 1. Thus, human-readable visual cues text for the purpose of image captioning as disclosed by Xie reads on this feature (Fig. 2, visual cues 130 represents sequence of tokens (words) representing objects and relationships between the objects (see “Objects in this image” and “Attributes”)). If Applicant wishes to distinguish the claimed tokenized description from the cited art, further details about the actual domain specific language used would be needed. Therefore, the second limitation is taught by the cited art.
Regarding the third limitation, Applicant’s arguments have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. This limitation is rejected using a newly applied reference Rahman et al. (see below 103 rejection).
Hence, Applicant arguments with respect to independent claim 1 are not found persuasive.
Regarding independent claims 10 and 16, and dependent claims:
For analogous reasons as presented above, arguments regarding independent claims 10 and 16 and the dependent claims, are not found persuasive.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
5. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1, “A method” is recited, which is directed to one of the four statutory categories of invention (process) (Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite mental processes which fall into the category of abstract idea (Step 2A Prong 1: YES).
The following limitations, under their broadest reasonable interpretation, recite mental processes:
performing one or more planning, navigation, or control operations…using a map, wherein at least a portion of the map was generated, at least by…: a person can use a map to perform operations (e.g. can determine a route from point A to point B)
generating…based at least one…processing data associated with at least a section of a preliminary map, a tokenized description of the section of the preliminary map represented using a sequence of tokens, expressed in a domain specific language, that encode information about objects or features and relationships therebetween within the section of the preliminary map: a person looks at a map, and writes down a tokenized description (e.g. words) describing the map describing objects and relationships between the objects (e.g. which streets intersect, which intersections have lights or stop signs, etc.)
identifying, based in part on the tokenized description, one or more potential issues with respect to the section of the preliminary map by parsing the sequence of tokens and determining at least one of an inconsistency in the relationship or an omission of an expected object, feature, or relationship: a person uses the description to determine issues in the map, such as a missing feature (e.g. missing stop sign, missing street, etc.)
generating, based at least on…processing data corresponding to the one or more potential issues, one or more textual representations or textual recommendations regarding the one or more potential issues: a person analyzes the issues they identified in order to write down textual representations or recommendations regarding the issues
Claim 1 does not contain any additional elements which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). The only additional limitations are “based at least on a first language model” and “based at least on the first language model or a second language model”, and “…associated with an autonomous or semi-autonomous machine…”. These limitations are recited broadly and amounts to mere instructions to implement the judicial exception using a generic computer, which does not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract idea. Accordingly, claim 1 is directed to an abstract idea.
Claim 1 does not contain any additional elements which amount to significantly more than the judicial exception (Step 2B: NO). As discussed above, the additional limitations amount to mere instructions to implement the judicial exception using a generic computer, which do not amount to significantly more than the judicial exception as they do not provide an inventive concept. Therefore, claim 1 is not patent eligible.
Regarding dependent claims 2-9, “The method” is recited, which is directed to one of the four statutory categories of invention (process) (Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite mental processes which fall into the category of abstract idea (Step 2A Prong 1: YES).
The following limitations, under their broadest reasonable interpretation, recite mental processes:
Claim 2:
Claim 2 contains the additional limitation: “wherein the first language model and the second language model are portions of a single language model”, which amounts to mere instructions to implement the judicial exception using a generic computer.
Claim 3:
Claim 3 contains the additional limitation “wherein the identifying of the one or more potential issues is performed using a third language model, wherein the third language model is one of a standalone language model, part of the first language model, part of the second language model, or part of the single language model”, which amounts to mere instructions to implement the judicial exception using a generic computer.
Claim 4:
wherein the one or more potential issues relate to at least one of a correction of an identified error, an addition of information determined to be absent from the preliminary map, or an enhancement to the preliminary map: a person identifies issues such as correcting an error, adding missing information in the map, or enhancing the map.
Claim 5:
generating, based at least on…processing data associated with a second section of the preliminary map, a tokenized description of the second section of the preliminary map: a person writes down a tokenized description of a viewed map
determining, based in part on the tokenized description of the second section of the preliminary map, that there are no potential modifications to be made to the second section of the preliminary map: a person analyzes the description they wrote and determines that no modifications are to be made
providing indication of validation of the second section of the preliminary map: a person writes down the result (e.g. message saying that the map is validated).
Claim 5 contains the additional limitations “based at least on the first language model”, which amounts to mere instructions to implement the judicial exception using a generic computer.
Claim 6:
presenting the one or more textual representations or textual recommendations regarding the one of more potential issues; and based at least on the presenting, receiving data corresponding to one or more data inputs indicating whether to implement one or more modifications to the map data of the preliminary map: a person presents the results to a user, and obtains data indicating whether or not to implement the modifications
Claim 7:
wherein the tokenized description is a tokenized text string representative of the section of a preliminary map, the tokenized text string including a sequence of tokens associated with objects in the section of the preliminary map: a person writes down a description, including tokens describing objects they see in the map
Claim 8:
wherein the tokenized text string is written in a road topology language (RTL) or a domain specific language (DSL): a person writes down the text string in a domain specific language (writes down the description in a specific format)
Claim 9:
wherein the one or more textual representations or textual recommendations are tokenized text strings: a person writes down the representations or recommendations as a series of text tokens on pen and paper.
Claims 2-9 do not contain any additional elements which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). As discussed above, the additional limitations amount to mere instructions to implement the judicial exception using a generic computer, which do not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract idea. Accordingly, claims 2-9 are directed to an abstract idea.
Claims 2-9 do not contain any additional elements which amount to significantly more than the judicial exception (Step 2B: NO). As discussed above, the only additional limitations are mere instructions to implement the judicial exception using a generic computer, which do not amount to significantly more than the judicial exception as they do not provide an inventive concept. Therefore, claims 2-9 are not patent eligible.
Regarding claim 10, “A processor” is recited, which is directed to one of the four statutory categories of invention (machine) (Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite mental processes which fall into the category of abstract idea (Step 2A Prong 1: YES).
The following limitations, under their broadest reasonable interpretation, recite mental processes:
generate… a tokenized description of at least a section of preliminary map data represented using a sequence of tokens expressed in a domain specific language that encode information about objects or features and relationships therebetween within the section of the preliminary map: a person looks at a map, and writes down a tokenized description (e.g. words) describing the map
…determine probability values for individual tokens of the sequence of tokens: a person can write down an accompanying probability of each token (e.g. probability of the token being correct/no modifications needed).
identify, based at least on the probability values, one or more potential modifications with respect to the section of the preliminary map by parsing the sequence of tokens and determining at least one of an inconsistency in the relationship or an omission of an expected object, feature, or relationship: a person uses the values they wrote down to determine whether to identify modifications to be made, such as a missing features (e.g. missing stop sign, missing street, etc.)
update, using the identified one or more potential modifications, the preliminary map data to generate updated map data: a person uses the representations or recommendations to update a map (e.g. adding a drawing to a map to fix a particular issue)
perform one or more planning, navigation, or control operations…using the updated map data: a person can use a map to perform operations (e.g. can determine a route from point A to point B)
Claim 10 does not contain any additional elements which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). The only additional limitations are “A processor, comprising: one or more circuits to” “using a first language model” and “use a second language model”, and “…associated with an autonomous or semi-autonomous machine”. These limitations are recited broadly and amount to mere instructions to implement the judicial exception using a generic computer, which do not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract idea. Accordingly, claim 10 is directed to an abstract idea.
Claim 10 does not contain any additional elements which amount to significantly more than the judicial exception (Step 2B: NO). As discussed above, the additional limitations are mere instructions to implement the judicial exception using a generic computer, which do not amount to significantly more than the judicial exception as they do not provide an inventive concept. Therefore, claim 10 is not patent eligible.
Regarding dependent claims 11-15, “The processor” is recited, which is directed to one of the four statutory categories of invention (machine) (Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite mental processes which fall into the category of abstract idea (Step 2A Prong 1: YES).
The following limitations, under their broadest reasonable interpretation, recite mental processes:
Claim 11:
generate, based at least on processing the one or more potential modifications, one or more textual representations or textual recommendations regarding the one or more potential modifications: a person analyzes the modifications they identified, and then write down representations or recommendations regarding the issues.
Claim 11 contains the additional limitation “use at least one language model to”, which amounts to mere instructions to implement the judicial exception using a generic computer.
Claim 12:
wherein the one or more potential modifications relate to at least one of a correction of an identified error, an addition of information determined to be absent from the preliminary map data, or an enhancement to the preliminary map data: a person identifies issues such as correcting an error, adding missing information in the map, or enhancing the map.
Claim 13:
wherein the tokenized description is a tokenized text string representative of the section of a preliminary map data, the tokenized text string including a sequence of tokens associated with objects in the section of the preliminary map data: a person writes down a description, including tokens describing objects they see in the map
Claim 14:
Claim 14 contains the additional limitation: “wherein the first language model and the second language model are portions of a single language model”, which amounts to mere instructions to implement the judicial exception using a generic computer.
Claim 15:
wherein the processor is comprised in at least one of:a system for performing simulation operations;a system for performing simulation operations to test or validate autonomous machine applications;a system for performing digital twin operations;a system for performing light transport simulation;a system for rendering graphical output;a system for performing deep learning operations;a system for performing generative Al operations using a large language model (LLM);a system implemented using an edge device;a system for generating or presenting virtual reality (VR) content;a system for generating or presenting augmented reality (AR) content;a system for generating or presenting mixed reality (MR) content;a system incorporating one or more Virtual Machines (VMs);a system implemented at least partially in a data center;a system for performing hardware testing using simulation;a system for performing generative operations using a language model (LM);a system for synthetic data generation;a collaborative content creation platform for 3D assets; or a system implemented at least partially using cloud computing resources: this limitation is recited at a high level of generality, and amounts to mere instructions to implement the judicial exception using a generic computer.
Claims 11-15 do not contain any additional elements which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). As discussed above, the additional limitations amount to mere instructions to implement the judicial exception using a generic computer, which do not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract idea. Accordingly, claims 11-15 are directed to an abstract idea.
Claims 11-15 do not contain any additional elements which amount to significantly more than the judicial exception (Step 2B: NO). As discussed above, the only additional limitations are mere instructions to implement the judicial exception using a generic computer, which do not amount to significantly more than the judicial exception as they do not provide an inventive concept. Therefore, claims 11-15 are not patent eligible.
Regarding claim 16, “A system” is recited, which is directed to one of the four statutory categories of invention (machine) (Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite mental processes which fall into the category of abstract idea (Step 2A Prong 1: YES).
The following limitations, under their broadest reasonable interpretation, recite mental processes:
identify one or more modifications to be performed with respect to at least a section of a map, represented using a sequence of tokens, by parsing the sequence of tokens and determining at least one of an inconsistency in the relatinoships or an omission of an expected object, feature, or relationship, the sequence of tokens being expressed in a domain specific language that encode information about objects or features and relationships therebetween within the section of the map: a person looks at a map, and writes down a tokenized description (e.g. words) describing the map: a person looks at a map, identifies modifications to be made to the map based on writing down a tokenized description of the map
update, using the identified one or more modifications, map data of the map to generate updated map data: a person uses the representations or recommendations to update a map (e.g. adding a drawing to a map to fix a particular issue)
perform one or more planning, navigation, or control operations…using the updated map data: a person can use a map to perform operations (e.g. can determine a route from point A to point B)
Claim 16 does not contain any additional elements which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). The only additional limitations are “A system comprising: one or more processors to” and “use a language model”, and “…associated with an autonomous or semi-autonomous machine”. These limitations are recited broadly and amounts to mere instructions to implement the judicial exception using a generic computer, which does not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract idea. Accordingly, claim 16 is directed to an abstract idea.
Claim 16 does not contain any additional elements which amount to significantly more than the judicial exception (Step 2B: NO). As discussed above, the additional limitations amount to mere instructions to implement the judicial exception using a generic computer, which do not amount to significantly more than the judicial exception as they do not provide an inventive concept. Therefore, claim 16 is not patent eligible.
Regarding dependent claims 17-20, “The system” is recited, which is directed to one of the four statutory categories of invention (machine) (Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite mental processes which fall into the category of abstract idea (Step 2A Prong 1: YES).
The following limitations, under their broadest reasonable interpretation, recite mental processes:
Claim 17:
generate the tokenized description of the section of the map: a person writes down a tokenized description describing the map
Claim 17 contains the additional limitation “use a second language model to”, which amounts to mere instructions to implement the judicial exception using a generic computer.
Claim 18:
generate, based at least on processing the one or more modifications, one or more textual representations or textual recommendations regarding the one or more modifications: a person writes down representations or recommendations regarding the modifications.
Claim 18 contains the additional limitation “use a third language model to”, which amounts to mere instructions to implement the judicial exception using a generic computer.
Claim 19:
Claim 19 contains the additional limitation “wherein the at least two of the first, second, and third language models, the second language model, and the third language model are portions of a single language model”, which amounts to mere instructions to implement the judicial exception using a generic computer.
Claim 20:
wherein the simulation system comprises at least one of:a system for performing simulation operations;a system for performing simulation operations to test or validate autonomous machine applications;a system for performing digital twin operations;a system for performing light transport simulation;a system for rendering graphical output;a system for performing deep learning operations;a system for performing generative AI operations using a large language model (LLM);a system implemented using an edge device;a system for generating or presenting virtual reality (VR) content;a system for generating or presenting augmented reality (AR) content;a system for generating or presenting mixed reality (MR) content;a system incorporating one or more Virtual Machines (VMs);a system implemented at least partially in a data center;a system for performing hardware testing using simulation; a system for performing generative operations using a language model (LM);a system for synthetic data generation;a collaborative content creation platform for 3D assets; or a system implemented at least partially using cloud computing resources: this limitation is recited at a high level of generality, and amounts to mere instructions to implement the judicial exception using a generic computer.
Claims 17-20 do not contain any additional elements which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). As discussed above, the additional limitations amount to mere instructions to implement the judicial exception using a generic computer, which do not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract idea. Accordingly, claims 17-20 are directed to an abstract idea.
Claims 17-20 do not contain any additional elements which amount to significantly more than the judicial exception. As discussed above, the only additional limitations are mere instructions to implement the judicial exception using a generic computer, which do not amount to significantly more than the judicial exception as they do not provide an inventive concept. Therefore, claims 17-20 are not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
6. Claims 1-5 and 7-20 are rejected under 35 U.S.C. 103 as being unpatentable over Xie et al. (US 2023/0394855 A1, hereinafter Xie) in view of Levinson & Sibley (US 2017/0248963 A1, hereinafter Levinson) and further in view of Rahman et al. (NPL “A Bidirectional LSTM Language Model for Code Evaluation and Repair”, hereinafter Rahman).
Regarding claim 1, Xie discloses generating, based at least on a first language model processing data associated with at least a section of a preliminary map (para. 0019 “Architecture 100 intakes an image 102 and has a vision language model 112, a captioner 114, and an object detector 116. Object detector 116 detects a plurality of objects 118 in image 102. Vision language model 112 and captioner 114 produce visual information 120 from image 102 and plurality of objects 118.”), a tokenized description of the section of the preliminary map represented using a sequence of tokens, expressed in a domain specific language, that encode information about objects or features and relationships therebetween within the section of the preliminary map (para. 0020 “Visual information 120 comprises text that describes what is contained within image 102, such as image tags 122, an initial image caption 124, and object information 126.”; para. 0028 “A baseline graph 304 is generated for an image 302, using either tags for objects automatically detected in image 302”; Fig. 2, visual cues 130 represents objects and relationships between the objects (see “Objects in this image” and “Attributes”)); identifying, based in part on the tokenized description, one or more potential issues with respect to the section of the preliminary map by parsing the sequence of tokens… (para. 0026 “Visual clues 130 are provided to generative language model 140, which in some examples, comprises an autoregressive language model that uses deep learning to produce human-like text. Generative language model 140 produces crisp language descriptions that are informative to a user, without being cluttered with irrelevant information from visual clues 130.”); and generating, based at least on a the first language model or a second language model processing data corresponding to the one or more potential issues (based on a second language model: para. 0027 “In some examples, vision language model 150 directly selects selected story caption 154 from among plurality of image story caption candidates 144, whereas in other examples, vision language model 150 scores plurality of image story caption candidates 144 and a down selection component 152 selects selected story caption 154 based on at least the scores from vision language model 150.”), one or more textual representations or textual recommendations regarding the one or more issues (para. 0025 “A vision language model 150 selects a selected story caption 154”).
Xie does not specifically disclose performing one or more planning, navigation, or control operations associated with an autonomous or semi-autonomous machine using a map, wherein at least a portion of the map was generated at least by…
Levinson teaches performing one or more planning, navigation, or control operations associated with an autonomous or semi-autonomous machine using a map, wherein at least a portion of the map was generated at least by… (Levinson teaches generating updated maps via detecting necessary corrections to reflect current map: para. 0066 “For example, perception engine 366 may be able to detect and classify external objects as pedestrians, bicyclists, dogs, other vehicles, etc. (e.g., perception engine 366 is configured to classify the objects in accordance with a type of classification, which may be associated with semantic information, including a label) …; para. 0157 “Data change detector 3653 is configured to detect changes in data sets 3655a and 3655b, which are examples of any number of data sets of 3-D map data. Data change detector 3653 also is configured to generate data identifying a portion of map data that has changed, as well as optionally identifying or classifying an object associated with the changed portion of map data. …At time, T2, however, data change detector 3653 may detect that another number of data sets, including data set 3655b, includes data representing the presence of external objects in portions of map data 3665 of 3-D model data 3661, whereby portions of map data 3665 coincide with portions of map data 3664 at different times. Therefore, data change detector 3653 may detect changes in map data, and may further adaptively modify map data to include the changed map data (e.g., as updated map data).”; Levinson teaches updating the map: para. 0160 “A tile generator 3656 may be configured to generate two-dimensional or three-dimensional map tiles based on map data from data sets 3655a and 3655b. The map tiles may be transmitted for storage in map repository 3605a. Tile generator 3656 may generate map tiles that include indicator data for indicating a portion of the map is an updated portion of map data.”; Levinson teaches using the map for autonomous operations: para. 0160 “Tile generator 3656 may generate map tiles that include indicator data for indicating a portion of the map is an updated portion of map data. Further, an updated map portion may be incorporated into a reference data repository 3605 in an autonomous vehicle…”; see para. 0071, map repository data used for localizing autonomous vehicle relative to reference data, see para. 0056, stating localizer is used for causing autonomous vehicles to drive autonomously).
Xie and Levinson are considered to be analogous to the claimed invention as
they both are in the same field of image processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Xie to incorporate the teachings of Levinson in order to specifically perform one or more planning, navigation, or control operations associated with an autonomous or semi-autonomous machine using a map, wherein at least a portion of the map was generated. Doing so would be beneficial, as this would enhance the accuracy of localization functions for autonomous vehicles (Levinson, para. 0146).
Xie in view of Levinson discloses identifying potential issues with respect to the section of the preliminary map by parsing the sequence of tokens (see above claim mapping), but does not specifically disclose […and ] determining at least one of an inconsistency in the relationship or an omission of an expected object, feature, or relationship.
Rahman teaches the use of a language model for determining at least one of an inconsistency in the relationship or an omission of an expected object, feature, or relationship (Fig. 2, Bi-LSTM language model used; pg. 7, section 4, 2nd para. “Our proposed model is a sequence-to-sequence language model that predicts next words in incorrect codes on the basis of probability. The Softmax activation layer (as defined in Equation (11)) is used to transform the output vector S(z)…for probabilities. The Softmax layer generates the probability for each word (token or ID) if the probability is too low (less than 0.1), which is considered an error candidate and immediately mark the entire line. At the same time, the model generates a possible correct word instead of the error. To predict the correct word (token or ID), the model (BiLSTM) calculates the code sequences (both forward and backward) to find the best possible word on the basis of the highest probability…”; see Tables 3-4, logical and syntactical inconsistencies are detected for input code).
Xie, Levinson, and Rahman are considered to be analogous to the claimed invention as Xie and Rahman are in the same field of language processing and Levinson is in the same field of image processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Xie in view of Levinson to incorporate the teachings of Rahman in order to specifically determine at least one of an inconsistency in the relationship or an omission of an expected object, feature, or relationship. Doing so would be beneficial, as this would allow for logical errors/inconsistencies to be detected and corrected automatically via machine learning without manual search, which can be hard to identify manually (Rahman, Abstract).
Regarding claim 2, Xie in view of Levinson and Rahman discloses wherein the first language model and the second language model are portions of a single language model (Xie, para. 0025 “In some examples, vision language model 112 and vision language model 150 both comprise a common vision language model.”).
Regarding claim 3, Xie in view of Levinson and Rahman discloses wherein identifying of the one or more potential issues is performed using a third language model, wherein the third language model is one of a standalone model, part of the first language model, part of the second language model, or part of the single language model (Xie, standalone model: para. 0026 “Visual clues 130 are provided to generative language model 140, which in some examples, comprises an autoregressive language model that uses deep learning to produce human-like text.”).
Regarding claim 4, Xie in view of Levinson and Rahman discloses wherein the one or more potential issues relate to at least one of a correction of an identified error, an addition of information determined to be absent from the preliminary map, or an enhancement to the preliminary map (Xie, para. 0018 “A rich semantic representation of an input image, such as image tags, object attributes and locations, and captions, is constructed as a structured textual prompt, termed “visual clues”, using a vision foundation model.”; para. 0021 “A generative language model 140 generates a plurality of image story caption candidates 144, which includes story captions 141, 142, and 143, from visual information 120, which includes, from visual clues 130.”).
Regarding claim 5, Xie in view of Levinson and Rahman discloses generating, based at least on the first language model processing data associated with a second section of the preliminary map, a tokenized description of the second section of the preliminary map (Xie, para. 0019 “Architecture 100 intakes an image 102 and has a vision language model 112, a captioner 114, and an object detector 116. Object detector 116 detects a plurality of objects 118 in image 102. Vision language model 112 and captioner 114 produce visual information 120 from image 102 and plurality of objects 118.”; para. 0020 “Visual information 120 comprises text that describes what is contained within image 102, such as image tags 122, an initial image caption 124, and object information 126.”; para. 0028 “A baseline graph 304 is generated for an image 302, using either tags for objects automatically detected in image 302”); determining, based in part on the tokenized description of the second section of the preliminary map, that there are no potential modifications to be made to the second section of the preliminary map (Rahman, uses probability values to determine whether or not sequence of tokens is correct: pg. 7, section 4, 2nd para. “Our proposed model is a sequence-to-sequence language model that predicts next words in incorrect codes on the basis of probability. The Softmax activation layer (as defined in Equation (11)) is used to transform the output vector S(z)…for probabilities. The Softmax layer generates the probability for each word (token or ID) if the probability is too low (less than 0.1), which is considered an error candidate and immediately mark the entire line. At the same time, the model generates a possible correct word instead of the error. To predict the correct word (token or ID), the model (BiLSTM) calculates the code sequences (both forward and backward) to find the best possible word on the basis of the highest probability…”; see Tables 3-4, logical and syntactical inconsistencies are detected for input code; probability values which are not less than 0.1 are not considered error candidates and thus do not require potential modifications; Xie teaches sequence of tokens for section of preliminary map, see claim mapping for claim 1); and providing indication of validation of the second section of the preliminary map (pg. 8, evaluation metrics generates for sequence of tokens, reflecting how many suggestions were generated; pg. 7, section 4.1, “…To balance the evaluation of the experimental results, we selected an equal number of correct (50%) and incorrect (50%) source codes from each type of problem (GCD and IS)…”; correct source codes would yield no suggestions; Xie teaches sequence of tokens for section of preliminary map).
Xie, Levinson, and Rahman are considered to be analogous to the claimed invention as Xie and Rahman are in the same field of language processing and Levinson is in the same field of image processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Xie in view of Levinson to incorporate the teachings of Rahman in order to specifically determine based in part on the tokenized description that there are no potential modifications to be made and to provide indication of validation. Doing so would be beneficial, as this would allow for logical errors/inconsistencies to be detected and corrected automatically via machine learning without manual search, which can be hard to identify manually (Rahman, Abstract).
Regarding claim 7, Xie in view of Levinson and Rahman discloses wherein the tokenized description is a tokenized text string representative of the section of a preliminary map (Xie, para. 0020 “Visual information 120 comprises text that describes what is contained within image 102, such as image tags 122, an initial image caption 124, and object information 126.”), the tokenized text string including a sequence of tokens associated with objects in the section of the preliminary map (Xie, para. 0020 “In some examples, image tags 122 includes one or more tags identifying objects within image 102, and object information 126 includes additional tags, captions, attributes, and locations for objects within image 102. Visual information 120 also includes visual clues 130 that is based on at least image tags 122, initial image caption 124, and object information 126. Visual clues 130 is a semantic representation of image 102 and comprises semantic components from object and attribute tags to localized detection regions and region captions.”).
Regarding claim 8, Xie in view of Levinson and Rahman discloses wherein the tokenized text string is written in a road topology language (RTL) or a domain specific language (Xie, embedded visual cues reads on a domain specific language: para. 0020 “Visual clues 130 is a semantic representation of image 102 and comprises semantic components from object and attribute tags to localized detection regions and region captions.”).
Regarding claim 9, Xie in view of Levinson and Rahman discloses wherein the one or more textual representations or textual recommendations are tokenized text strings (Xie, para. 0021 “A vision language model 150 selects a selected story caption 154, which was previously story caption 141”; Fig. 2, selected caption 154).
Regarding claim 10, Xie discloses A processor (Fig. 6, 614), comprising: one or more circuits to (para. 0084 “Processor(s) 614 may include any quantity of processing units that read data from various entities, such as memory 612 or I/O components 620. Specifically, processor(s) 614 are programmed to execute computer-executable instructions for implementing aspects of the disclosure. … Moreover, in some examples, the processor(s) 614 represent an implementation of analog techniques to perform the operations described herein. … One skilled in the art will understand and appreciate that computer data may be presented in a number of ways, such as visually in a graphical user interface (GUI), audibly through speakers, wirelessly between computing devices 600, across a wired connection,”): generate, using a first language model, a tokenized description of at least a section of preliminary map data represented using a sequence of tokens expressed in a domain specific language that encode information about objects or features and relationships therebetween within the section of the preliminary map (para. 0019 “Architecture 100 intakes an image 102 and has a vision language model 112, a captioner 114, and an object detector 116. Object detector 116 detects a plurality of objects 118 in image 102. Vision language model 112 and captioner 114 produce visual information 120 from image 102 and plurality of objects 118.”; para. 0020 “Visual information 120 comprises text that describes what is contained within image 102, such as image tags 122, an initial image caption 124, and object information 126.”; para. 0028 “A baseline graph 304 is generated for an image 302, using either tags for objects automatically detected in image 302”; Fig. 2, visual cues 130 represents objects and relationships between the objects (see “Objects in this image” and “Attributes”)); use a second language model (para. 0027 “In some examples, vision language model 150 directly selects selected story caption 154 from among plurality of image story caption candidates 144, whereas in other examples, vision language model 150 scores plurality of image story caption candidates 144 and a down selection component 152 selects selected story caption 154 based on at least the scores from vision language model 150.”)…
Xie does not specifically disclose update, using the identified one or more potential modifications, the preliminary map data to generate updated map data; and perform one or more planning, navigation, or control operations associated with an autonomous or semi-autonomous machine using the updated map data.
Levinson teaches update, using the identified one or more potential modifications (para. 0066 “For example, perception engine 366 may be able to detect and classify external objects as pedestrians, bicyclists, dogs, other vehicles, etc. (e.g., perception engine 366 is configured to classify the objects in accordance with a type of classification, which may be associated with semantic information, including a label). Based on the classification of these external objects, the external objects may be labeled as dynamic objects or static objects. For example, an external object classified as a tree may be labeled as a static object, while an external object classified as a pedestrian may be labeled as a static object. External objects labeled as static may or may not be described in map data. Examples of external objects likely to be labeled as static include traffic cones, cement barriers arranged across a roadway, lane closure signs, newly-placed mailboxes or trash cans adjacent a roadway, etc. Examples of external objects likely to be labeled as dynamic include bicyclists, pedestrians, animals, other vehicles, etc. If the external object is labeled as dynamic, and further data about the external object may indicate a typical level of activity and velocity, as well as behavior patterns associated with the classification type.”; para. 0109 “Classifier 2360 is configured to identify an object and to classify that object by classification type (e.g., as a pedestrian, cyclist, etc.) and by energy/activity (e.g. whether the object is dynamic or static), whereby data representing classification is described by a semantic label.”), the preliminary map data to generate updated map data (para. 0157 “Data change detector 3653 is configured to detect changes in data sets 3655a and 3655b, which are examples of any number of data sets of 3-D map data. Data change detector 3653 also is configured to generate data identifying a portion of map data that has changed, as well as optionally identifying or classifying an object associated with the changed portion of map data. …At time, T2, however, data change detector 3653 may detect that another number of data sets, including data set 3655b, includes data representing the presence of external objects in portions of map data 3665 of 3-D model data 3661, whereby portions of map data 3665 coincide with portions of map data 3664 at different times. Therefore, data change detector 3653 may detect changes in map data, and may further adaptively modify map data to include the changed map data (e.g., as updated map data).”; para. 0159 “As shown, map data 3692 stored map repository 3605a is associated with, or linked to, indication data (“delta data”) 3694 that indicated that an associated portion of map data has changed. Further to the example shown, indication data 3694 may identify a set of traffic cones, as changed portions of map data 3665, disposed in a physical environment associated with 3-D model 3661 through which an autonomous vehicle travels.”; para. 0160 “A tile generator 3656 may be configured to generate two-dimensional or three-dimensional map tiles based on map data from data sets 3655a and 3655b. The map tiles may be transmitted for storage in map repository 3605a. Tile generator 3656 may generate map tiles that include indicator data for indicating a portion of the map is an updated portion of map data.”); and perform one or more planning, navigation, or control operations associated with an autonomous or semi-autonomous machine using the updated map data (para. 0160 “Tile generator 3656 may generate map tiles that include indicator data for indicating a portion of the map is an updated portion of map data. Further, an updated map portion may be incorporated into a reference data repository 3605 in an autonomous vehicle…”; see para. 0071, map repository data used for localizing autonomous vehicle relative to reference data, see para. 0056, stating localizer is used for causing autonomous vehicles to drive autonomously).
Xie and Levinson are considered to be analogous to the claimed invention as
they both are in the same field of image processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Xie to incorporate the teachings of Levinson in order to specifically update using the one or more potential modifications, preliminary map data to generate updated map data and to perform one or more planning, navigation, or control operations associated with an autonomous or semi-autonomous machine using the updated map data. Doing so would be beneficial, as this would enhance the accuracy of localization functions for autonomous vehicles (Levinson, para. 0146).
Xie in view of Levinson does not specifically disclose [use a second language model] to determine probability values for individual tokens of the sequence of tokens; identify, based at least one the probability values, one or more potential modifications with respect to the section of the preliminary map by parsing the sequence of tokens and determining at least one of an inconsistency in the relationships or an omission of an expected object, feature, or relationship to be made with respect to the section of the preliminary map data.
Rahman teaches to determine probability values for individual tokens of the sequence of tokens (pg. 7, section 4, 2nd para. “Our proposed model is a sequence-to-sequence language model that predicts next words in incorrect codes on the basis of probability. The Softmax activation layer (as defined in Equation (11)) is used to transform the output vector S(z)…for probabilities.); identify, based at least one the probability values, one or more potential modifications with respect to the section of the preliminary map by parsing the sequence of tokens and determining at least one of an inconsistency in the relationships or an omission of an expected object, feature, or relationship to be made with respect to the section of the preliminary map data (pg. 7, section 4, 2nd para. “Our proposed model is a sequence-to-sequence language model that predicts next words in incorrect codes on the basis of probability. The Softmax activation layer (as defined in Equation (11)) is used to transform the output vector S(z)…for probabilities. The Softmax layer generates the probability for each word (token or ID) if the probability is too low (less than 0.1), which is considered an error candidate and immediately mark the entire line. At the same time, the model generates a possible correct word instead of the error. To predict the correct word (token or ID), the model (BiLSTM) calculates the code sequences (both forward and backward) to find the best possible word on the basis of the highest probability…”; see Tables 3-4, logical and syntactical inconsistencies are detected for input code based on calculated probability values of tokens; Xie teaches sequence of tokens with respect to the section of preliminary map, see above claim mapping).
Xie, Levinson, and Rahman are considered to be analogous to the claimed invention as Xie and Rahman are in the same field of language processing and Levinson is in the same field of image processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Xie in view of Levinson to incorporate the teachings of Rahman in order to specifically use a second language model to determine probability values for individual tokens of the sequence of tokens, and to identify based at least on the probability values, one or more potential modifications with respect to the section of the preliminary map by parsing the sequence of tokens and determining at least one of an inconsistency in the relationships or an omission of an expected object, feature, or relationship. Doing so would be beneficial, as this would allow for logical errors/inconsistencies to be detected and corrected automatically via machine learning without manual search, which can be hard to identify manually (Rahman, Abstract).
Regarding claim 11, Xie in view of Levinson and Rahman discloses wherein the one or more circuits are further to use at least one language model to generate, based at least on processing the one or more potential modifications, one or more textual representations or textual recommendations regarding the one or more modifications (Rahman, language model used, see Fig. 2 BiLSTM, to generate textual recommendations regarding the modifications based on probability values: pg. 7, section 4, 2nd para. “Our proposed model is a sequence-to-sequence language model that predicts next words in incorrect codes on the basis of probability. The Softmax activation layer (as defined in Equation (11)) is used to transform the output vector S(z)…for probabilities. The Softmax layer generates the probability for each word (token or ID) if the probability is too low (less than 0.1), which is considered an error candidate and immediately mark the entire line. At the same time, the model generates a possible correct word instead of the error. To predict the correct word (token or ID), the model (BiLSTM) calculates the code sequences (both forward and backward) to find the best possible word on the basis of the highest probability…”; see Tables 3-4, logical and syntactical inconsistencies are detected for input code based on calculated probability values of tokens).
Xie, Levinson, and Rahman are considered to be analogous to the claimed invention as Xie and Rahman are in the same field of language processing and Levinson is in the same field of image processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Xie in view of Levinson to incorporate the teachings of Rahman in order to specifically use a second language model to determine use at least one language model to generate, based at least on processing the one or more potential modifications, one or more textual representations or textual recommendations regarding the one or more potential modifications. Doing so would be beneficial, given the same rationale as claim 10.
Regarding claim 12 Xie in view of Levinson and Rahman discloses wherein the one or more potential modifications relate to at least one of a correction of an identified error, an addition of information determined to be absent from the preliminary map data, or an enhancement to the preliminary map data (Xie teaches token sequence with respect to a section of preliminary map data, see above claim mapping for claim 10; Rahman teaches modifications including a correction of an identified error in a token sequence: pg. 7, section 4, 2nd para. “Our proposed model is a sequence-to-sequence language model that predicts next words in incorrect codes on the basis of probability. The Softmax activation layer (as defined in Equation (11)) is used to transform the output vector S(z)…for probabilities. The Softmax layer generates the probability for each word (token or ID) if the probability is too low (less than 0.1), which is considered an error candidate and immediately mark the entire line. At the same time, the model generates a possible correct word instead of the error. To predict the correct word (token or ID), the model (BiLSTM) calculates the code sequences (both forward and backward) to find the best possible word on the basis of the highest probability…”; see Tables 3-4, logical and syntactical inconsistencies are detected for input code based on calculated probability values of tokens).
Xie, Levinson, and Rahman are considered to be analogous to the claimed invention as Xie and Rahman are in the same field of language processing and Levinson is in the same field of image processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Xie in view of Levinson to incorporate the teachings of Rahman in order to specifically have the one or more potential modifications relate to at least one of a correction of an identified error, an addition of information determined to be absent from the preliminary map data, or an enhancement to the preliminary map data. Doing so would be beneficial, given the same rationale as claim 10.
Regarding claim 13, Xie in view of Levinson and Rahman discloses wherein the tokenized description is a tokenized text string representative of the section of the preliminary map data (Xie, para. 0020 “Visual information 120 comprises text that describes what is contained within image 102, such as image tags 122, an initial image caption 124, and object information 126.”), the tokenized text string including a sequence of tokens associated with objects in the section of the preliminary map data (Xie, para. 0020 “In some examples, image tags 122 includes one or more tags identifying objects within image 102, and object information 126 includes additional tags, captions, attributes, and locations for objects within image 102. Visual information 120 also includes visual clues 130 that is based on at least image tags 122, initial image caption 124, and object information 126. Visual clues 130 is a semantic representation of image 102 and comprises semantic components from object and attribute tags to localized detection regions and region captions.”).
Regarding claim 14, Xie in view of Levinson and Rahman discloses wherein the first language model and the second language model are portions of a single language model (Xie, para. 0025 “In some examples, vision language model 112 and vision language model 150 both comprise a common vision language model.”).
Regarding claim 15, Xie in view of Levinson and Rahman discloses wherein the processor is comprised in at least one of: a system for performing simulation operations; a system for performing simulation operations to test or validate autonomous machine applications; a system for performing digital twin operations; a system for performing light transport simulation; a system for rendering graphical output; a system for performing deep learning operations; a system for performing generative Al operations using a large language model (LLM);a system implemented using an edge device; a system for generating or presenting virtual reality (VR) content; a system for generating or presenting augmented reality (AR) content; a system for generating or presenting mixed reality (MR) content; a system incorporating one or more Virtual Machines (VMs);a system implemented at least partially in a data center; a system for performing hardware testing using simulation; a system for performing generative operations using a language model (LM);a system for synthetic data generation; a collaborative content creation platform for 3D assets; or a system implemented at least partially using cloud computing resources (Xie, a system for performing generative operations using a language model: Fig. 6 discloses system, Fig. 2, “Generative Language Model 140”).
Regarding claim 16, Xie discloses A system (Fig. 6) comprising: one or more processors (Fig. 6, 614) to use a language model (para. 0027 “In some examples, vision language model 150 directly selects selected story caption 154 from among plurality of image story caption candidates 144, whereas in other examples, vision language model 150 scores plurality of image story caption candidates 144 and a down selection component 152 selects selected story caption 154 based on at least the scores from vision language model 150.”) … with respect to at least a section of a map, represented using a sequence of tokens, by parsing the sequence of tokens… the sequence of tokens being expressed in a domain specific language that encode information about objects or features and relationships therebetween within the section of the map (para. 0019 “Architecture 100 intakes an image 102 and has a vision language model 112, a captioner 114, and an object detector 116. Object detector 116 detects a plurality of objects 118 in image 102. Vision language model 112 and captioner 114 produce visual information 120 from image 102 and plurality of objects 118.”; para. 0020 “Visual information 120 comprises text that describes what is contained within image 102, such as image tags 122, an initial image caption 124, and object information 126.”; para. 0028 “A baseline graph 304 is generated for an image 302, using either tags for objects automatically detected in image 302”; Fig. 2, visual cues 130 represents objects and relationships between the objects (see “Objects in this image” and “Attributes”)).
Xie does not specifically disclose update, using the identified one or more modifications, map data of the map to generate updated map data; and perform one or more planning, navigation, or control operations associated with an autonomous or semi-autonomous machine using the updated map data.
Levinson teaches update, using the identified one or more potential modifications (para. 0066 “For example, perception engine 366 may be able to detect and classify external objects as pedestrians, bicyclists, dogs, other vehicles, etc. (e.g., perception engine 366 is configured to classify the objects in accordance with a type of classification, which may be associated with semantic information, including a label). Based on the classification of these external objects, the external objects may be labeled as dynamic objects or static objects. For example, an external object classified as a tree may be labeled as a static object, while an external object classified as a pedestrian may be labeled as a static object. External objects labeled as static may or may not be described in map data. Examples of external objects likely to be labeled as static include traffic cones, cement barriers arranged across a roadway, lane closure signs, newly-placed mailboxes or trash cans adjacent a roadway, etc. Examples of external objects likely to be labeled as dynamic include bicyclists, pedestrians, animals, other vehicles, etc. If the external object is labeled as dynamic, and further data about the external object may indicate a typical level of activity and velocity, as well as behavior patterns associated with the classification type.”; para. 0109 “Classifier 2360 is configured to identify an object and to classify that object by classification type (e.g., as a pedestrian, cyclist, etc.) and by energy/activity (e.g. whether the object is dynamic or static), whereby data representing classification is described by a semantic label.”), the preliminary map data to generate updated map data (para. 0157 “Data change detector 3653 is configured to detect changes in data sets 3655a and 3655b, which are examples of any number of data sets of 3-D map data. Data change detector 3653 also is configured to generate data identifying a portion of map data that has changed, as well as optionally identifying or classifying an object associated with the changed portion of map data. …At time, T2, however, data change detector 3653 may detect that another number of data sets, including data set 3655b, includes data representing the presence of external objects in portions of map data 3665 of 3-D model data 3661, whereby portions of map data 3665 coincide with portions of map data 3664 at different times. Therefore, data change detector 3653 may detect changes in map data, and may further adaptively modify map data to include the changed map data (e.g., as updated map data).”; para. 0159 “As shown, map data 3692 stored map repository 3605a is associated with, or linked to, indication data (“delta data”) 3694 that indicated that an associated portion of map data has changed. Further to the example shown, indication data 3694 may identify a set of traffic cones, as changed portions of map data 3665, disposed in a physical environment associated with 3-D model 3661 through which an autonomous vehicle travels.”; para. 0160 “A tile generator 3656 may be configured to generate two-dimensional or three-dimensional map tiles based on map data from data sets 3655a and 3655b. The map tiles may be transmitted for storage in map repository 3605a. Tile generator 3656 may generate map tiles that include indicator data for indicating a portion of the map is an updated portion of map data.”); and perform one or more planning, navigation, or control operations associated with an autonomous or semi-autonomous machine using the updated map data (para. 0160 “Tile generator 3656 may generate map tiles that include indicator data for indicating a portion of the map is an updated portion of map data. Further, an updated map portion may be incorporated into a reference data repository 3605 in an autonomous vehicle…”; see para. 0071, map repository data used for localizing autonomous vehicle relative to reference data, see para. 0056, stating localizer is used for causing autonomous vehicles to drive autonomously).
Xie and Levinson are considered to be analogous to the claimed invention as
they both are in the same field of image processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Xie to incorporate the teachings of Levinson in order to specifically update using the one or more potential modifications, preliminary map data to generate updated map data and to perform one or more planning, navigation, or control operations associated with an autonomous or semi-autonomous machine using the updated map data. Doing so would be beneficial, as this would enhance the accuracy of localization functions for autonomous vehicles (Levinson, para. 0146).
Xie in view of Levinson does not specifically disclose to use a language model to identify one or more modifications to be performed with respect to at least a section of a map, by…determining at least one of an inconsistency in the relationships or an omission of an expected object, feature, or relationship….
Rahman teaches to identify one or more modifications to be performed with respect to at least a section of a map, by…determining at least one of an inconsistency in the relationships or an omission of an expected object, feature, or relationship…. (Xie teaches token sequence with respect to a section of preliminary map data, see above claim mapping for claim 10; Rahman teaches modifications including a correction of an identified error in a token sequence: pg. 7, section 4, 2nd para. “Our proposed model is a sequence-to-sequence language model that predicts next words in incorrect codes on the basis of probability. The Softmax activation layer (as defined in Equation (11)) is used to transform the output vector S(z)…for probabilities. The Softmax layer generates the probability for each word (token or ID) if the probability is too low (less than 0.1), which is considered an error candidate and immediately mark the entire line. At the same time, the model generates a possible correct word instead of the error. To predict the correct word (token or ID), the model (BiLSTM) calculates the code sequences (both forward and backward) to find the best possible word on the basis of the highest probability…”; see Tables 3-4, logical and syntactical inconsistencies are detected for input code based on calculated probability values of tokens).
Xie, Levinson, and Rahman are considered to be analogous to the claimed invention as Xie and Rahman are in the same field of language processing and Levinson is in the same field of image processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Xie in view of Levinson to incorporate the teachings of Rahman in order to specifically use a second language model to determine probability values for individual tokens of the sequence of tokens, and to identify based at least on the probability values, one or more potential modifications with respect to the section of the preliminary map by parsing the sequence of tokens and determining at least one of an inconsistency in the relationships or an omission of an expected object, feature, or relationship. Doing so would be beneficial, as this would allow for logical errors/inconsistencies to be detected and corrected automatically via machine learning without manual search, which can be hard to identify manually (Rahman, Abstract).
Regarding claim 17, Xie in view of Levinson and Rahman discloses wherein the one or more processors are further to use a second language model to generate the tokenized description of the section of the map (Xie, para. 0019 “Architecture 100 intakes an image 102 and has a vision language model 112, a captioner 114, and an object detector 116. Object detector 116 detects a plurality of objects 118 in image 102. Vision language model 112 and captioner 114 produce visual information 120 from image 102 and plurality of objects 118.”; para. 0020 “Visual information 120 comprises text that describes what is contained within image 102, such as image tags 122, an initial image caption 124, and object information 126.”; para. 0028 “A baseline graph 304 is generated for an image 302, using either tags for objects automatically detected in image 302”).
Regarding claim 18, Xie in view of Levinson and Rahman discloses wherein the one or more processors are further to use a third language model to generate, based at least on processing the one or more modifications, one or more textual representations or textual recommendations regarding the one or more modifications (Rahman, language model used, see Fig. 2 BiLSTM, to generate textual recommendations regarding the modifications based on probability values: pg. 7, section 4, 2nd para. “Our proposed model is a sequence-to-sequence language model that predicts next words in incorrect codes on the basis of probability. The Softmax activation layer (as defined in Equation (11)) is used to transform the output vector S(z)…for probabilities. The Softmax layer generates the probability for each word (token or ID) if the probability is too low (less than 0.1), which is considered an error candidate and immediately mark the entire line. At the same time, the model generates a possible correct word instead of the error. To predict the correct word (token or ID), the model (BiLSTM) calculates the code sequences (both forward and backward) to find the best possible word on the basis of the highest probability…”; see Tables 3-4, logical and syntactical inconsistencies are detected for input code based on calculated probability values of tokens).
Xie, Levinson, and Rahman are considered to be analogous to the claimed invention as Xie and Rahman are in the same field of language processing and Levinson is in the same field of image processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Xie in view of Levinson to incorporate the teachings of Rahman in order to specifically use a third language model to generate, based at least on processing the one or more modifications, one or more textual representations or textual recommendations regarding the one or more modifications. Doing so would be beneficial, given the same rationale as claim 16.
Regarding claim 19, Xie in view of Levinson and Rahman discloses wherein at least two of: i) the first, second, and third language models, ii) the second language model, and iii) the third language model are portions of a single language model (Xie, para. 0025 “In some examples, vision language model 112 and vision language model 150 both comprise a common vision language model.”).
Regarding claim 20, Xie in view of Levinson and Rahman discloses wherein the system is comprised in least one of: a system for performing simulation operations; a system for performing simulation operations to test or validate autonomous machine applications; a system for performing digital twin operations; a system for performing light transport simulation; a system for rendering graphical output; a system for performing deep learning operations; a system for performing generative AI operations using a large language model (LLM);a system implemented using an edge device; a system for generating or presenting virtual reality (VR) content; a system for generating or presenting augmented reality (AR) content; a system for generating or presenting mixed reality (MR) content; a system incorporating one or more Virtual Machines (VMs);a system implemented at least partially in a data center ;a system for performing hardware testing using simulation; a system for performing generative operations using a language model (LM);a system for synthetic data generation; a collaborative content creation platform for 3D assets; or a system implemented at least partially using cloud computing resources (Xie, a system for performing generative operations using a language model: Fig. 6 discloses system, Fig. 2, “Generative Language Model 140).
7. Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Xie in view Levinson and Rahman, and further in view of Barut et al. (US 12,045,288, hereinafter Barut).
Regarding claim 6, Xie in view of Levinson and Rahnam discloses presenting the one or more textual representations or textual recommendations regarding the one or more potential issues (Xie, para. 0022 “Example applications for leveraging the capability of architecture 100, which may be instructed using caption focus 132, include visual storytelling, automatic advertisement generation, social media posting, background explanation, accessibility, and machine learning (ML) training data annotation.”; para. 0023 “For automatic advertisement generation, a seller uploads image 104 and caption focus 132 may be “Write a product description to sell in an online marketplace.” In some examples, caption focus 132 may indicate a number of different objects within image 104 for which to generate an advertisement. For a social media posting, the actual posting may be performed by a bot, and caption focus 132 may be “Social media post.” In some applications, the user may wish to edit selected story caption 154.”).
Xie in view of Levinson and Rahman does not specifically disclose based at least on the presenting, receiving data corresponding to one or more inputs indicating whether to implement one or more modifications to the map data of the preliminary map.
Barut teaches based at least on the presenting, receiving data corresponding to one or more inputs indicating whether to implement one or more modifications to the map data of the preliminary map (Col. 16 Lines 8-29 “In some examples, the ranking component 320 may determine that two different skills are equally applicable for processing the input data. In such examples, the decider engine 332 may determine that disambiguation should occur. Accordingly, the routing plan 334 may include sending the input data to a dialog skill 352 that may output (via TTS) one or more questions (e.g., a disambiguation request) used to prompt the user to disambiguate between the two equally likely (or approximately equally likely) interpretations of the input data. For example, it may be unclear, based on a user's request, whether the user intended to invoke a movie playback skill or a music playback skill, as a movie and a soundtrack for the movie may be identified using the same name. …the dialog skill 352 may inquire whether the user intended to play the movie or the soundtrack.”).
Xie, Levinson, Rahman, and Barut are considered to be analogous to the claimed invention as they are in the same field of natural language processing or image processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Xie in view of Levinson and Rahman to incorporate the teachings of Barut in order to receive data corresponding to one or more inputs indicating whether to implement one or more modifications to the preliminary map data based at least on the presenting. Doing so would be beneficial, as this would allow for disambiguation to occur when it is unclear which modification should be performed (Col. 16 Lines 8-29), ensuring that the desired action is performed which improves user experience.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Arcadinho et al. (US 2023/0325154 A1): generating tokenized outputs using language models which conform to a computer language grammar (Fig. 2)
Bose & Ochoa (US 2024/0338248 A1): generating a set of instructions corresponding to an image (Fig. 2, para. 0045)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CODY DOUGLAS HUTCHESON whose telephone number is (703)756-1601. The examiner can normally be reached M-F 8:00AM-5:00PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir can be reached at (571)-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CODY DOUGLAS HUTCHESON/Examiner, Art Unit 2659
/PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659