Prosecution Insights
Last updated: October 02, 2026
Application No. 18/513,849

SYSTEM AND METHOD FOR TRAINING A MACHINE LEARNING MODEL

Non-Final OA §101§103§112
Filed
Nov 20, 2023
Priority
Dec 02, 2022 — GB 2218156.4
Examiner
COHEN, ZARED ORION
Art Unit
Tech Center
Assignee
Sony Group Corporation
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
8 currently pending
Career history
2
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “a data obtaining unit configured to obtain training data” in claim 1 “an event identifying unit configured to identify” in claim 1 “a list generating unit configured to generate a list of identified events in the training data” in claim 1 “a dataset generating unit configured to generate a dataset” in claim 1 “and a training unit configured to train a machine learning model using the generated dataset” in claim 1 Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 3 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 3 recites the limitation "the gameplay" in line 2. There is insufficient antecedent basis for this limitation in the claim. It is unclear what “the gameplay” is referring to. For the purposes of examination, the Examiner has interpreted these instances and all subsequent instances as the gameplay of a game from Claim 2. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim(s) 1- 13 are rejected under 35 U.S.C. 101 because the claims are directed towards an abstract idea without significantly more. Regarding Claim 1: Subject Matter Eligibility Analysis Step 1: Claim 1 recites a system and is thus a machine, one of the four statutory categories of patentable subject matter. Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 1 recites: A system for generating a training dataset for a machine learning process (This limitation is a mental process as it encompasses a human mentally generating a dataset and is thus an evaluation.) an event identifying unit configured to identify, based upon one or more corresponding indicators, the occurrence of an event of interest in the training data (This limitation is a mental process as it encompasses a human mentally identifying events and is thus an evaluation.) a list generating unit configured to generate a list of identified events in the training data, wherein identified events are added to the list with a probability that is inversely proportional to the frequency of the occurrence of that event within the training data (This limitation is a mental process as it encompasses a human mentally generating a list of events and is thus an evaluation.) a dataset generating unit configured to generate a dataset comprising information about the events contained in the generated list (This limitation is a mental process as it encompasses a human mentally generating a dataset and is thus an evaluation.) Therefore, claim 1 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 1 further recites additional elements of: and training a machine learning model (This element does not integrate the abstract idea into a practical application because it recites insignificant extra-solution activity of training a model (see MPEP 2106.05(g)).) the system comprising: a data obtaining unit configured to obtain training data comprising a plurality of events of interest and the behaviour of an agent corresponding to those events (This element does not integrate the abstract idea into a practical application because it recites the insignificant extra-solution activity of receiving data (see MPEP 2106.05(g)).) wherein identified events are added to the list with a probability that is inversely proportional to the frequency of the occurrence of that event within the training data (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).) and a training unit configured to train a machine learning model using the generated dataset, wherein the machine learning model is trained to generate behaviour for an agent corresponding to events within the generated dataset (This element does not integrate the abstract idea into a practical application because it recites insignificant extra-solution activity of training a model (see MPEP 2106.05(g)).) Therefore, claim 1 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 1 do not provide significantly more than the abstract idea itself, taken alone and in combination because training a machine learning model is the well understood, routine, and conventional activity of iteratively training a model (US 2021/0125108 A1, Metzler et al. page 11, paragraph 0052, “a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position”). the system comprising: a data obtaining unit configured to obtain training data comprising a plurality of events of interest and the behaviour of an agent corresponding to those events is the well understood, routine, and conventional activity of "transmitting or receiving data over a network" (see MPEP 2106.05(d)(II); OIP Techs., Inc., V. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network)). wherein identified events are added to the list with a probability that is inversely proportional to the frequency of the occurrence of that event within the training data uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)). and a training unit configured to train a machine learning model using the generated dataset, wherein the machine learning model is trained to generate behaviour for an agent corresponding to events within the generated dataset is the well understood, routine, and conventional activity of iteratively training a model (US 2021/0125108 A1, Metzler et al. page 11, paragraph 0052, “a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position”). Therefore, claim 1 is subject-matter ineligible. Regarding Claim 2: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 2 recites the same abstract idea as claim 1. Therefore, claim 2 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 2 further recites additional elements of: wherein the training data comprises videos of gameplay of a game, logs of inputs provided by users, screenshots of gameplay of a game, and/or a log of events within gameplay of a game (This element does not integrate the abstract idea into a practical application because it further modifies the insignificant extra-solution activity of receiving training data from Claim 1 (see MPEP 2106.05(d)).) Therefore, claim 2 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 2 do not provide significantly more than the abstract idea itself, taken alone and in combination because wherein the training data comprises videos of gameplay of a game, logs of inputs provided by users, screenshots of gameplay of a game, and/or a log of events within gameplay of a game is the well understood, routine, and conventional activity of "transmitting or receiving data over a network" (see MPEP 2106.05(d)(II); OIP Techs., Inc., V. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network)). Therefore, claim 2 is subject-matter ineligible. Regarding Claim 3: Regarding Claim 3: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 3 recites the same abstract idea as claim 1. Therefore, claim 3 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 3 further recites additional elements of: wherein the indicators comprise one or more of game parameters, user inputs, image features of the gameplay, audio features of the gameplay, and/or entries in an event log (This element does not integrate the abstract idea into a practical application because it recites a technological environment in which to apply a judicial exception (see MPEP 2106.05(h)).) Therefore, claim 3 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 3 do not provide significantly more than the abstract idea itself, taken alone and in combination because wherein the indicators comprise one or more of game parameters, user inputs, image features of the gameplay, audio features of the gameplay, and/or entries in an event log specifies a technological environment to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(h)). Therefore, claim 3 is subject-matter ineligible. Regarding Claim 4: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 4 recites the same abstract idea as claim 1. Therefore, claim 4 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 4 further recites additional elements of: The system of claim 1, wherein the probability of adding an event to the list is additionally proportional to a defined significance of the corresponding indicator and/or event within the training data (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).) Therefore, claim 4 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 4 do not provide significantly more than the abstract idea itself, taken alone and in combination because The system of claim 1, wherein the probability of adding an event to the list is additionally proportional to a defined significance of the corresponding indicator and/or event within the training data uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)). Therefore, claim 4 is subject-matter ineligible. Regarding Claim 5: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 5 recites: wherein the dataset generating unit is configured to generate the dataset by additionally sampling the training data obtained by the data obtaining unit. (This limitation is a mental process as it encompasses a human mentally generating a dataset and is thus an evaluation.) Therefore, claim 5 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 5 does not further recite any additional elements. Therefore, claim 5 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 5 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 5 is subject-matter ineligible. Regarding Claim 6: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 6 recites: wherein: the list generating unit is configured to generate a plurality of lists of identified events in the training data, each list comprising a different set of identified events (This limitation is a mental process as it encompasses a human mentally generating a list of events and is thus an evaluation.) the dataset generating unit is configured to generate a plurality of datasets each corresponding to a respective one of the plurality of lists (This limitation is a mental process as it encompasses a human mentally generating a plurality of datasets and is thus an evaluation.) Therefore, claim 6 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 6 further recites additional elements of: and the training unit is configured to use each of these datasets for training the machine learning model (This element does not integrate the abstract idea into a practical application because it recites insignificant extra-solution activity of training a model (see MPEP 2106.05(g)).) Therefore, claim 6 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 6 do not provide significantly more than the abstract idea itself, taken alone and in combination because and the training unit is configured to use each of these datasets for training the machine learning model is the well understood, routine, and conventional activity of iteratively training a model (US 2021/0125108 A1, Metzler et al. page 11, paragraph 0052, “a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position”). Therefore, claim 6 is subject-matter ineligible. Regarding Claim 7: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 7 recites the same abstract idea as claim 6. Therefore, claim 7 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 7 further recites additional elements of: wherein the training unit is configured to use a respective subset of the plurality of datasets for training respective ones of two or more machine learning models (This element does not integrate the abstract idea into a practical application because it recites the insignificant extra-solution activity of training a model (see MPEP 2106.05(g)).) Therefore, claim 7 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 7 do not provide significantly more than the abstract idea itself, taken alone and in combination because wherein the training unit is configured to use a respective subset of the plurality of datasets for training respective ones of two or more machine learning models is the well understood, routine, and conventional activity of iteratively training a model (US 2021/0125108 A1, Metzler et al. page 11, paragraph 0052, “a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position”). Therefore, claim 7 is subject-matter ineligible. Regarding Claim 8: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 8 recites the same abstract idea as claim 1. Therefore, claim 8 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 8 further recites additional elements of: wherein the indicators are predefined for the training data (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).) Therefore, claim 8 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 8 do not provide significantly more than the abstract idea itself, taken alone and in combination because wherein the indicators are predefined for the training data uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)). Therefore, claim 8 is subject-matter ineligible. Regarding Claim 9: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 9 recites: wherein the event identifying unit is configured to identify indicators for events of interest based upon one or more labelled examples in the training data. (This limitation is a mental process as it encompasses a human mentally identifying indicators and is thus an evaluation.) Therefore, claim 9 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 9 does not further recite any additional elements. Therefore, claim 9 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 9 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 9 is subject-matter ineligible. Regarding Claim 10: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 10 recites the same abstract idea as claim 1. Therefore, claim 10 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 10 further recites additional elements of: wherein the probabilities for respective events are updated in response to the addition of identified events to the list (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).) Therefore, claim 10 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 10 do not provide significantly more than the abstract idea itself, taken alone and in combination because wherein the probabilities for respective events are updated in response to the addition of identified events to the list uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)). Therefore, claim 10 is subject-matter ineligible. Regarding Claim 11: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 11 recites the same abstract idea as claim 1. Therefore, claim 11 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 11 further recites additional elements of: wherein the training unit is configured to train the machine learning model using an imitation learning method (This element does not integrate the abstract idea into a practical application because it recites a technological environment in which to apply a judicial exception (see MPEP 2106.05(h)).) Therefore, claim 11 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 11 do not provide significantly more than the abstract idea itself, taken alone and in combination because wherein the training unit is configured to train the machine learning model using an imitation learning method specifies a technological environment to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(h)). Therefore, claim 11 is subject-matter ineligible. Regarding Claim 12: Subject Matter Eligibility Analysis Step 1: Claim 12 recites a method and is thus a process, one of the four statutory categories of patentable subject matter. Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 12 recites: A method for generating a training dataset for a machine learning process (This limitation is a mental process as it encompasses a human mentally generating a dataset and is thus an evaluation.) identifying, based upon one or more corresponding indicators, the occurrence of an event of interest in the training data (This limitation is a mental process as it encompasses a human mentally identifying events and is thus an evaluation.) generating a list of identified events in the training data (This limitation is a mental process as it encompasses a human mentally generating a list of events and is thus an evaluation.) generating a dataset comprising information about the events contained in the generated list (This limitation is a mental process as it encompasses a human mentally generating a dataset and is thus an evaluation.) Therefore, claim 12 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 12 further recites additional elements of: and training a machine learning model (This element does not integrate the abstract idea into a practical application because it recites insignificant extra-solution activity of training a model (see MPEP 2106.05(g)).) the method comprising: obtaining training data comprising a plurality of events of interest and the behaviour of an agent corresponding to those events (This element does not integrate the abstract idea into a practical application because it recites the insignificant extra-solution activity of receiving training data (see MPEP 2106.05(d)).) wherein identified events are added to the list with a probability that is inversely proportional to the frequency of the occurrence of that event within the training data (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).) and training a machine learning model using the generated dataset, wherein the machine learning model is trained to generate behaviour for an agent corresponding to events within the generated dataset (This element does not integrate the abstract idea into a practical application because it recites insignificant extra-solution activity of training a model (see MPEP 2106.05(g)).) Therefore, claim 12 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 12 do not provide significantly more than the abstract idea itself, taken alone and in combination because and training a machine learning model is the well understood, routine, and conventional activity of iteratively training a model (US 2021/0125108 A1, Metzler et al. page 11, paragraph 0052, “a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position”). the method comprising: obtaining training data comprising a plurality of events of interest and the behaviour of an agent corresponding to those events is the well understood, routine, and conventional activity of "transmitting or receiving data over a network" (see MPEP 2106.05(d)(II); OIP Techs., Inc., V. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network)). wherein identified events are added to the list with a probability that is inversely proportional to the frequency of the occurrence of that event within the training data uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)). and training a machine learning model using the generated dataset, wherein the machine learning model is trained to generate behaviour for an agent corresponding to events within the generated dataset is the well understood, routine, and conventional activity of iteratively training a model (US 2021/0125108 A1, Metzler et al. page 11, paragraph 0052, “a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position”). Therefore, claim 12 is subject-matter ineligible. Regarding Claim 13: Subject Matter Eligibility Analysis Step 1: Claim 13 recites a non-transitory machine-readable storage medium and is thus a machine, one of the four statutory categories of patentable subject matter. Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 13 recites: generating a training dataset for a machine learning process (This limitation is a mental process as it encompasses a human mentally generating a dataset and is thus an evaluation.) identifying, based upon one or more corresponding indicators, the occurrence of an event of interest in the training data (This limitation is a mental process as it encompasses a human mentally identifying events and is thus an evaluation.) generating a list of identified events in the training data (This limitation is a mental process as it encompasses a human mentally generating a list of events and is thus an evaluation.) generating a dataset comprising information about the events contained in the generated list (This limitation is a mental process as it encompasses a human mentally generating a dataset and is thus an evaluation.) Therefore, claim 13 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 13 further recites additional elements of: A non-transitory machine-readable storage medium which stores computer software which, when executed by a computer, causes the computer to perform a method for (This element does not integrate the abstract idea into a practical application because it recites generic computing components on which to perform the abstract idea (see MPEP 2106.05(f)).) and training a machine learning model (This element does not integrate the abstract idea into a practical application because it recites insignificant extra-solution activity of training a model (see MPEP 2106.05(g)).) the method comprising: obtaining training data comprising a plurality of events of interest and the behaviour of an agent corresponding to those events (This element does not integrate the abstract idea into a practical application because it recites the insignificant extra-solution activity of receiving data (see MPEP 2106.05(d)).) wherein identified events are added to the list with a probability that is inversely proportional to the frequency of the occurrence of that event within the training data (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).) and training a machine learning model using the generated dataset, wherein the machine learning model is trained to generate behaviour for an agent corresponding to events within the generated dataset (This element does not integrate the abstract idea into a practical application because it recites insignificant extra-solution activity of training a model (see MPEP 2106.05(g)).) Therefore, claim 13 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 13 do not provide significantly more than the abstract idea itself, taken alone and in combination because A non-transitory machine-readable storage medium which stores computer software which, when executed by a computer, causes the computer to perform a method for uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)).) and training a machine learning model is the well understood, routine, and conventional activity of iteratively training a model (US 2021/0125108 A1, Metzler et al. page 11, paragraph 0052, “a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position”). the method comprising: obtaining training data comprising a plurality of events of interest and the behaviour of an agent corresponding to those events is the well understood, routine, and conventional activity of "transmitting or receiving data over a network" (see MPEP 2106.05(d)(II); OIP Techs., Inc., V. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network)). wherein identified events are added to the list with a probability that is inversely proportional to the frequency of the occurrence of that event within the training data uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)). and training a machine learning model using the generated dataset, wherein the machine learning model is trained to generate behaviour for an agent corresponding to events within the generated dataset is the well understood, routine, and conventional activity of iteratively training a model (US 2021/0125108 A1, Metzler et al. page 11, paragraph 0052, “a classifier is trained using a conventional iterative machine learning training process that determines weights for each result list position”). Therefore, claim 13 is subject-matter ineligible. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim(s) 1-5, 8, 11-13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Angus et al. (“Unlimited Road-scene Synthetic Annotation (URSA) Dataset”) in view of Cao et al. (US 2021/0398014 A1) in further view of Naeem et al. (“Classification of movie reviews using term frequency-inverse document frequency and optimized machine learning algorithms”). Regarding Claim 1, Angus teaches A system for generating a training dataset for a machine learning process, and training a machine learning model, the system comprising: a data obtaining unit configured to obtain training data comprising a plurality of events of interest (Angus, page 987, “Our proposed data generation scheme can be separated into three steps. First, we uniquely identify super-pixels in an image by correlating File path of a drawable, Model name, Shader index, and Sampler (FMSS) data that we parse using the open-source Codewalker tool [16].” Examiner notes the events of interest are the identified super-pixels (FMSS) in the image.) an event identifying unit configured to identify, based upon one or more corresponding indicators, the occurrence of an event of interest in the training data (Angus, page 987, “The FMSS sectioning results in 1,178,355 total in-game sections (hereafter refered to as FMSS), of which 56,540 are relevant for generating our road dataset. FMSS that occur indoors, for example, are not considered. Second, we create a user interface to collect annotations for every relevant FMSS via AMT.” Examiner notes that the indicators are the in-game sections that are relevant for generating the dataset and the relevant FMSS data are the events of interest.) a list generating unit configured to generate a list of identified events in the training data (Angus, page 987, “we devise a view selection mechanism that allows us to generate a near optimal number of frames required to label the 56,540 FMSS. We exploit the paths used by the in-game AI drivers that form a graph G = (V,E). Each vertex, along with position, contains information about the road type, which is used to choose only major roads for view selection (i.e. no dirt roads, alleyways). The edges define the structure of the road network. The goal is to find a minimal set (M,L) ⊆ V ×E such that all FMSS used in outdoor scenes are existent in at least one scene. Here M is the set of vertices to view from and L is the edge to look down.” Examiner notes that the list of identified events is the minimal set.) a dataset generating unit configured to generate a dataset comprising information about the events contained in the generated list (Angus, page 988, “The 3,388 viewpoints are used to generate an equal number of dashcam-style scene snapshots, with each snapshot sectioned into FMSS as shown in Fig 1. We design a web-based annotation GUI that allows annotators to classify each FMSS into one of the 28 road scene classes (Fig. 2).” Angus, page 987, “Similar to Richter et al. [15], we ease the pixel labelling burden by identifying super-pixels that uniquely correlate to a portion of an in-game object. As mentioned previously this unique identifier is based on FMSS.” Examiner notes that the scene snapshots comprise the dataset and the information about the events is the unique identifier based on FMSS.) and a training unit configured to train a machine learning model using the generated dataset (Angus, page 989, “To validate the quality of our generated data in comparison to other synthetic datasets, we devise several experiments involving training two deep neural network models.” Angus, page 989, “Closeness To Real Data: Here we train the neural networks for 110K iterations on each synthetic dataset with VGG-16 ImageNet weight initialization.”) Angus does not teach “the behaviour of an agent corresponding to those events” or “a machine learning model is trained to generate behaviour for an agent corresponding to events within the generated dataset”. However, Cao teaches the behaviour of an agent corresponding to those events (Cao, paragraph 0046, “The low-level imitation learning model 302 may receive a sequence of frames 306. Each frame 306 may be an RGB image. During training, imitation learning observes expert demonstrations, which consist of the RGB images (e.g., frames 306) and the expert actions.” Examiner notes that the events are the RGB images and the behaviour of the agents are the expert actions.) wherein the machine learning model is trained to generate behaviour for an agent corresponding to events within the generated dataset (Cao, paragraph 0046, “During training, imitation learning observes expert demonstrations, which consist of the RGB images (e.g., frames 306) and the expert actions. During execution, the agent observes the RGB images (e.g., frames 306) and produces the actions itself.”) Cao and Angus are analogous to the claimed invention because they involve creating specialized training datasets for machine learning models from unlabeled data. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Angus to use the training datasets to generate behavior from an agent corresponding to the data like Cao. Doing so is advantageous because “Autonomous agents (e.g., vehicles, robots, etc.) employ machine learning models to navigate through an environment” (Cao, paragraph 0003) and “It is desirable to improve a machine learning model's ability to model phase transitions. A phase transition may occur when a change in an agent's state causes a change in an optimal action, such as an optimal trajectory for arriving at a destination” (Cao, paragraph 0003).” Cao and Angus do not, but Naeem teaches wherein identified events are added to the list with a probability that is inversely proportional to the frequency of the occurrence of that event within the training data (Naeem, page 8, “In this approach, weights are assigned to every term in a document based on term frequency and inverse document frequency (Neethu & Rajasree, 2013; Biau & Scornet, 2016). Terms having higher weights are supposed to be more important than terms having lower weights. The weight for each term is based on the Eq. (1). PNG media_image1.png 59 144 media_image1.png Greyscale where TFt,d is the number of occurrences of term t in document d, Df,t is the number of documents having the term t and N is the total number of documents in the dataset. TF-IDF is a kind of scoring measurement approach which is widely used in summarization and information retrieval. TF calculates the frequency of a token and gives higher importance to more common tokens in a given document (Vishwanathan & Murty, 2002). On the other hand, IDF calculates the tokens which are rare in a corpus. In this way, if uncommon words appear in more than one document, they are considered meaningful and important. In a set of documents D, IDF weighs a token x using the Eq. (2). PNG media_image2.png 33 136 media_image2.png Greyscale where n(x) denotes frequency of x in D and N/n(x) denotes the inverse frequency. TF-IDF is calculated using TF and IDF as shown in Eq. (3). PNG media_image3.png 26 149 media_image3.png Greyscale TF-IDF is applied to calculate the weights of important terms and the final output of TF-IDF is in the form of a weight matrix. Values gradually increase to the count in TF-IDF but are balanced with the frequency of the word in dataset (Zhang et al., 2008).”) Cao, Angus, and Naeem are analogous to the claimed invention because they involve creating specialized training datasets for machine learning models. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Angus and Cao to add events to the list with a value inversely proportional to the event frequency in the data like Naeem. Doing so would be advantageous because “Experimental results indicate that the SVM obtains the highest accuracy when used with TF-IDF features” (Naeem, page 1), wherein TF-IDF is term frequency-inverse document frequency. Regarding Claim 2, Angus, Cao, and Naeem teach The system of claim 1, wherein the training data comprises videos of gameplay of a game, logs of inputs provided by users, screenshots of gameplay of a game, and/or a log of events within gameplay of a game (Angus, page 985, “Utilizing open-source tools and resources found in single-player modding communities, we provide a method for persistent, ground truth, asset annotation of a game world. By collecting a synthetic dataset containing upwards of 1,000,000 images, we demonstrate real time, on-demand, ground truth data annotation capability of our method.”) Regarding Claim 3, Angus, Cao, and Naeem teach The system of claim 1, wherein the indicators comprise one or more of game parameters, user inputs, image features of the gameplay, audio features of the gameplay, and/or entries in an event log (Angus, page 987, “Our proposed data generation scheme can be separated into three steps. First, we uniquely identify super-pixels in an image by correlating File path of a drawable, Model name, Shader index, and Sampler (FMSS) data that we parse using the open-source Codewalker tool [16].” Examiner notes the FMSS indicators are image features of gameplay.) Regarding Claim 4, Angus, Cao, and Naeem teach The system of claim 1, wherein the probability of adding an event to the list is additionally proportional to a defined significance of the corresponding indicator and/or event within the training data (Angus, page 988, “However, each super-pixel in the scene may be influenced by multiple FMSS (e.g. partially transparent textures). To remedy this problem, we identify super-pixels in each scene by weighing the influence of different FMSS at a pixel-level. Pixel influence is calculated by magnitude of contribution an FMSS to a given pixel. Each pixel is then assigned to the FMSS tuple with the highest influence.”) Regarding Claim 5, Angus, Cao, and Naeem teach The system of claim 1, wherein the dataset generating unit is configured to generate the dataset by additionally sampling the training data obtained by the data obtaining unit. (Angus, page 988, “The 3,388 viewpoints are used to generate an equal number of dashcam-style scene snapshots, with each snapshot sectioned into FMSS as shown in Fig 1. We design a web-based annotation GUI that allows annotators to classify each FMSS into one of the 28 road scene classes (Fig. 2).” Angus, page 987, “The FMSS sectioning results in 1,178,355 total in-game sections (hereafter refered to as FMSS), of which 56,540 are relevant for generating our road dataset.”) Regarding Claim 8, Angus, Cao, and Naeem teach The system of claim 1, wherein the indicators are predefined for the training data (Angus, page 987, “First, we uniquely identify super-pixels in an image by correlating File path of a drawable, Model name, Shader index, and Sampler (FMSS) data that we parse using the open-source Codewalker tool [16].” Examiner notes that the indicators are the FMSS data which include a File path, Model name, Shader index, and Sampler data. Examiner further notes that each of these elements are predefined values as the game uses these elements to load data or define known values.) Regarding Claim 11, Angus, Cao, and Naeem teach The system of claim 1. Angus nor Naeem teach “wherein the training unit is configured to train the machine learning model using an imitation learning method” However, Cao teaches wherein the training unit is configured to train the machine learning model using an imitation learning method (Cao, paragraph 0045, “FIG. 3A is a block diagram illustrating an example of a hierarchical reinforcement and imitation learning (H-REIL) framework 300, in accordance with aspects of the present disclosure. As shown in FIG. 3A, the H-REIL framework 300 includes a low-level imitation learning model 302 and a high-level reinforcement learning model 304.”) Cao Angus, and Naeem are analogous to the claimed invention because they involve creating specialized training datasets for machine learning models from unlabeled data. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Angus and Naeem to use an imitation learning method to train the machine learning model like Cao. Doing so is advantageous because “Autonomous agents (e.g., vehicles, robots, etc.) employ machine learning models to navigate through an environment” (Cao, paragraph 0003) and “It is desirable to improve a machine learning model's ability to model phase transitions. A phase transition may occur when a change in an agent's state causes a change in an optimal action, such as an optimal trajectory for arriving at a destination” (Cao, paragraph 0003).” Regarding Claim 12, Angus teaches A method for generating a training dataset for a machine learning process, and training a machine learning model, the method comprising: obtaining training data comprising a plurality of events of interest (Angus, page 987, “Our proposed data generation scheme can be separated into three steps. First, we uniquely identify super-pixels in an image by correlating File path of a drawable, Model name, Shader index, and Sampler (FMSS) data that we parse using the open-source Codewalker tool [16].” Examiner notes the events of interest are the identified super-pixels (FMSS) in the image.) identifying, based upon one or more corresponding indicators, the occurrence of an event of interest in the training data (Angus, page 987, “The FMSS sectioning results in 1,178,355 total in-game sections (hereafter refered to as FMSS), of which 56,540 are relevant for generating our road dataset. FMSS that occur indoors, for example, are not considered. Second, we create a user interface to collect annotations for every relevant FMSS via AMT.” Examiner notes that the indicators are the in-game sections that are relevant for generating the dataset and the relevant FMSS data are the events of interest.) generating a list of identified events in the training data (Angus, page 987, “we devise a view selection mechanism that allows us to generate a near optimal number of frames required to label the 56,540 FMSS. We exploit the paths used by the in-game AI drivers that form a graph G = (V,E). Each vertex, along with position, contains information about the road type, which is used to choose only major roads for view selection (i.e. no dirt roads, alleyways). The edges define the structure of the road network. The goal is to find a minimal set (M,L) ⊆ V ×E such that all FMSS used in outdoor scenes are existent in at least one scene. Here M is the set of vertices to view from and L is the edge to look down.” Examiner notes that the list of identified events is the minimal set.) generating a dataset comprising information about the events contained in the generated list (Angus, page 988, “The 3,388 viewpoints are used to generate an equal number of dashcam-style scene snapshots, with each snapshot sectioned into FMSS as shown in Fig 1. We design a web-based annotation GUI that allows annotators to classify each FMSS into one of the 28 road scene classes (Fig. 2).” Angus, page 987, “Similar to Richter et al. [15], we ease the pixel labelling burden by identifying super-pixels that uniquely correlate to a portion of an in-game object. As mentioned previously this unique identifier is based on FMSS.” Examiner notes that the scene snapshots comprise the dataset and the information about the events is the unique identifier based on FMSS.) and training a machine learning model using the generated dataset (Angus, page 989, “To validate the quality of our generated data in comparison to other synthetic datasets, we devise several experiments involving training two deep neural network models.” Angus, page 989, “Closeness To Real Data: Here we train the neural networks for 110K iterations on each synthetic dataset with VGG-16 ImageNet weight initialization.”) Angus does not teach “the behaviour of an agent corresponding to those events” or “the machine learning model is trained to generate behaviour for an agent corresponding to events within the generated dataset”. However, Cao teaches the behaviour of an agent corresponding to those events (Cao, paragraph 0046, “The low-level imitation learning model 302 may receive a sequence of frames 306. Each frame 306 may be an RGB image. During training, imitation learning observes expert demonstrations, which consist of the RGB images (e.g., frames 306) and the expert actions.” Examiner notes that the events are the RGB images and the behaviour of the agents are the expert actions.) wherein the machine learning model is trained to generate behaviour for an agent corresponding to events within the generated dataset (Cao, paragraph 0046, “During training, imitation learning observes expert demonstrations, which consist of the RGB images (e.g., frames 306) and the expert actions. During execution, the agent observes the RGB images (e.g., frames 306) and produces the actions itself.” Cao and Angus are analogous to the claimed invention because they involve creating specialized training datasets for machine learning models from unlabeled data. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Angus to use the training datasets to generate behavior from an agent corresponding to the data like Cao. Doing so is advantageous because “Autonomous agents (e.g., vehicles, robots, etc.) employ machine learning models to navigate through an environment” (Cao, paragraph 0003) and “It is desirable to improve a machine learning model's ability to model phase transitions. A phase transition may occur when a change in an agent's state causes a change in an optimal action, such as an optimal trajectory for arriving at a destination” (Cao, paragraph 0003).” Cao and Angus do not, but Naeem teaches wherein identified events are added to the list with a probability that is inversely proportional to the frequency of the occurrence of that event within the training data (Naeem, page 8, “In this approach, weights are assigned to every term in a document based on term frequency and inverse document frequency (Neethu & Rajasree, 2013; Biau & Scornet, 2016). Terms having higher weights are supposed to be more important than terms having lower weights. The weight for each term is based on the Eq. (1). PNG media_image1.png 59 144 media_image1.png Greyscale where TFt,d is the number of occurrences of term t in document d, Df,t is the number of documents having the term t and N is the total number of documents in the dataset. TF-IDF is a kind of scoring measurement approach which is widely used in summarization and information retrieval. TF calculates the frequency of a token and gives higher importance to more common tokens in a given document (Vishwanathan & Murty, 2002). On the other hand, IDF calculates the tokens which are rare in a corpus. In this way, if uncommon words appear in more than one document, they are considered meaningful and important. In a set of documents D, IDF weighs a token x using the Eq. (2). PNG media_image2.png 33 136 media_image2.png Greyscale where n(x) denotes frequency of x in D and N/n(x) denotes the inverse frequency. TF-IDF is calculated using TF and IDF as shown in Eq. (3). PNG media_image3.png 26 149 media_image3.png Greyscale TF-IDF is applied to calculate the weights of important terms and the final output of TF-IDF is in the form of a weight matrix. Values gradually increase to the count in TF-IDF but are balanced with the frequency of the word in dataset (Zhang et al., 2008).”) Cao, Angus, and Naeem are analogous to the claimed invention because they involve creating specialized training datasets for machine learning models. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Angus and Cao to add events to the list with a value inversely proportional to the event frequency in the data like Naeem. Doing so would be advantageous because “Experimental results indicate that the SVM obtains the highest accuracy when used with TF-IDF features” (Naeem, page 1), wherein TF-IDF is term frequency-inverse document frequency. Regarding Claim 13, Angus teaches a method for generating a training dataset for a machine learning process, and training a machine learning model, the method comprising: obtaining training data comprising a plurality of events of interest (Angus, page 987, “Our proposed data generation scheme can be separated into three steps. First, we uniquely identify super-pixels in an image by correlating File path of a drawable, Model name, Shader index, and Sampler (FMSS) data that we parse using the open-source Codewalker tool [16].” Examiner notes the events of interest are the identified super-pixels (FMSS) in the image.) identifying, based upon one or more corresponding indicators, the occurrence of an event of interest in the training data (Angus, page 987, “The FMSS sectioning results in 1,178,355 total in-game sections (hereafter refered to as FMSS), of which 56,540 are relevant for generating our road dataset. FMSS that occur indoors, for example, are not considered. Second, we create a user interface to collect annotations for every relevant FMSS via AMT.” Examiner notes that the indicators are the in-game sections that are relevant for generating the dataset and the relevant FMSS data are the events of interest.) generating a list of identified events in the training data (Angus, page 987, “we devise a view selection mechanism that allows us to generate a near optimal number of frames required to label the 56,540 FMSS. We exploit the paths used by the in-game AI drivers that form a graph G = (V,E). Each vertex, along with position, contains information about the road type, which is used to choose only major roads for view selection (i.e. no dirt roads, alleyways). The edges define the structure of the road network. The goal is to find a minimal set (M,L) ⊆ V ×E such that all FMSS used in outdoor scenes are existent in at least one scene. Here M is the set of vertices to view from and L is the edge to look down.” Examiner notes that the list of identified events is the minimal set.) generating a dataset comprising information about the events contained in the generated list (Angus, page 988, “The 3,388 viewpoints are used to generate an equal number of dashcam-style scene snapshots, with each snapshot sectioned into FMSS as shown in Fig 1. We design a web-based annotation GUI that allows annotators to classify each FMSS into one of the 28 road scene classes (Fig. 2).” Angus, page 987, “Similar to Richter et al. [15], we ease the pixel labelling burden by identifying super-pixels that uniquely correlate to a portion of an in-game object. As mentioned previously this unique identifier is based on FMSS.” Examiner notes that the scene snapshots comprise the dataset and the information about the events is the unique identifier based on FMSS.) and training a machine learning model using the generated dataset (Angus, page 989, “To validate the quality of our generated data in comparison to other synthetic datasets, we devise several experiments involving training two deep neural network models.” Angus, page 989, “Closeness To Real Data: Here we train the neural networks for 110K iterations on each synthetic dataset with VGG-16 ImageNet weight initialization.”) Angus does not, but Cao teaches A non-transitory machine-readable storage medium which stores computer software which, when executed by a computer, causes the computer to perform (Cao, paragraph 0089, “The steps of a method or algorithm described in connection with the present disclosure may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module may reside in storage or machine readable medium, including random access memory (RAM), read only memory (ROM), flash memory, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, a hard disk, a removable disk, a CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. A software module may comprise a single instruction, or many instructions, and may be distributed over several different code segments, among different programs, and across multiple storage media. A storage medium may be coupled to a processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor.”) the behaviour of an agent corresponding to those events (Cao, paragraph 0046, “The low-level imitation learning model 302 may receive a sequence of frames 306. Each frame 306 may be an RGB image. During training, imitation learning observes expert demonstrations, which consist of the RGB images (e.g., frames 306) and the expert actions.” Examiner notes that the events are the RGB images and the behaviour of the agents are the expert actions.) wherein the machine learning model is trained to generate behaviour for an agent corresponding to events within the generated dataset (Cao, paragraph 0046, “During training, imitation learning observes expert demonstrations, which consist of the RGB images (e.g., frames 306) and the expert actions. During execution, the agent observes the RGB images (e.g., frames 306) and produces the actions itself.” Angus teaches a method for generating a training dataset for a machine learning model using images taken from gameplay of a game. Cao teaches a computer medium comprising instructions to run a method for generating a training dataset for an imitation machine learning model using data from an agent. Angus and Cao both teach generating training datasets for machine learning models from unlabeled data and thus are considered analogous to the claimed invention. It would have been obvious to one with ordinary skill in the art to modify Angus to implement the aforementioned method on a computer medium as Cao did as this would be applying a known technique to a known device to yield the predictable result of running software instructions on a computer medium successfully (see MPEP 2143(d)). It would have also been obvious to one of ordinary skill in the art before the effective filing date to modify Angus to use the training datasets to generate behavior from an agent corresponding to the data like Cao. Doing so is advantageous because “Autonomous agents (e.g., vehicles, robots, etc.) employ machine learning models to navigate through an environment” (Cao, paragraph 0003) and “It is desirable to improve a machine learning model's ability to model phase transitions. A phase transition may occur when a change in an agent's state causes a change in an optimal action, such as an optimal trajectory for arriving at a destination” (Cao, paragraph 0003).” Cao and Angus do not, but Naeem teaches wherein identified events are added to the list with a probability that is inversely proportional to the frequency of the occurrence of that event within the training data (Naeem, page 8, “In this approach, weights are assigned to every term in a document based on term frequency and inverse document frequency (Neethu & Rajasree, 2013; Biau & Scornet, 2016). Terms having higher weights are supposed to be more important than terms having lower weights. The weight for each term is based on the Eq. (1). PNG media_image1.png 59 144 media_image1.png Greyscale where TFt,d is the number of occurrences of term t in document d, Df,t is the number of documents having the term t and N is the total number of documents in the dataset. TF-IDF is a kind of scoring measurement approach which is widely used in summarization and information retrieval. TF calculates the frequency of a token and gives higher importance to more common tokens in a given document (Vishwanathan & Murty, 2002). On the other hand, IDF calculates the tokens which are rare in a corpus. In this way, if uncommon words appear in more than one document, they are considered meaningful and important. In a set of documents D, IDF weighs a token x using the Eq. (2). PNG media_image2.png 33 136 media_image2.png Greyscale where n(x) denotes frequency of x in D and N/n(x) denotes the inverse frequency. TF-IDF is calculated using TF and IDF as shown in Eq. (3). PNG media_image3.png 26 149 media_image3.png Greyscale TF-IDF is applied to calculate the weights of important terms and the final output of TF-IDF is in the form of a weight matrix. Values gradually increase to the count in TF-IDF but are balanced with the frequency of the word in dataset (Zhang et al., 2008).”) Cao, Angus, and Naeem are analogous to the claimed invention because they involve creating specialized training datasets for machine learning models. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Angus and Cao to add events to the list with a value inversely proportional to the event frequency in the data like Naeem. Doing so would be advantageous because “Experimental results indicate that the SVM obtains the highest accuracy when used with TF-IDF features” (Naeem, page 1), wherein TF-IDF is term frequency-inverse document frequency. Claim(s) 6-7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Angus in view of Cao in further view of Naeem in further view of Atsmon et al. (US 20230202511 A1). Regarding Claim 6, Angus, Cao, and Naeem teach The system of claim 1, wherein: the list generating unit is configured to generate a plurality of lists of identified events in the training data, each list comprising a different set of identified events (“Angus, 987, “The goal is to find a minimal set (M,L) ⊆ V ×E such that all FMSS used in outdoor scenes are existent in at least one scene. Here M is the set of vertices to view from and L is the edge to look down. We partition V into three sets: PNG media_image4.png 66 175 media_image4.png Greyscale Vsimp consists of simple road sections, where finding the direction of travel is trivial. Vcomp consists of more complex interchanges such as merge lanes or intersections, whereas Vdead consists of dead ends to be ignored.”) Angus, Cao, and Naeem do not teach “the dataset generating unit is configured to generate a plurality of datasets each corresponding to a respective one of the plurality of lists” or “the training unit is configured to use each of these datasets for training the machine learning model.” However, Atsmon teaches the dataset generating unit is configured to generate a plurality of datasets each corresponding to a respective one of the plurality of lists (Atsmon, paragraph 0070, “Optionally, other machine learning model 320 is trained using a plurality of recorded data sets, each recorded while a vehicle traverses a physical scene.” Atsmon, paragraph 0070, “Optionally, each of the plurality recorded data sets comprises a recorded driving scenario 331 and a plurality of recorded driving commands 332. Optionally, other machine learning model 320 computes classification 321B, indicative of a likelihood that recorded driving scenario 331 includes an interesting driving scenario.”) and the training unit is configured to use each of these datasets for training the machine learning model (Atsmon, paragraph 0070, “Optionally, other machine learning model 320 is trained using a plurality of recorded data sets, each recorded while a vehicle traverses a physical scene.”) Cao, Angus, Naeem, and Atsmon are analogous to the claimed invention because they involve creating specialized training datasets for machine learning models from unlabeled data. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Angus, Cao, and Naeem to create multiple datasets corresponding to the different lists to train a machine learning model like Atsmon. Doing so is advantageous because “increasing an amount of datasets used to train a machine learning model typically increases accuracy of an output of the machine learning model” (Atsmon, paragraph 0040). Regarding Claim 7, Angus, Cao, and Naeem teach the system of claim 6. Angus, Cao, and Naeem do not teach “wherein the training unit is configured to use a respective subset of the plurality of datasets for training respective ones of two or more machine learning models.” However, Atsmon teaches wherein the training unit is configured to use a respective subset of the plurality of datasets for training respective ones of two or more machine learning models (Atsmon, paragraph 0091, “In 210, processing unit 101 optionally provides at least some of plurality of simulated driving scenarios 311 to one or more autonomous driving model.” Atsmon, paragraph 0091, “Optionally, processing unit 101 provides the at least some of plurality of simulated driving scenarios 311 for one or more purposes selected from a group of purposes comprising: training the one or more autonomous driving model.” Examiner notes that each dataset corresponds to one of the multiple driving scenarios, which are used to train one or more driving models.) Cao, Angus, Naeem, and Atsmon are analogous to the claimed invention because they involve creating specialized training datasets for machine learning models from unlabeled data. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Angus, Cao, and Naeem to train multiple models using the datasets like Atsmon because “In the field of autonomous driving, it is common practice for a system, for example an Autonomous Driving System (ADS) or an Advance Driver-Assistance System (ADAS), to include one or more machine learning models. Such machine learning models may serve as a system's corner-stone for learning how to function well on the road” (Atsmon, paragraph 0004). Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Angus in view of Cao in further view of Naeem in further view of Dareddy et al. (GB2579603A). Regarding Claim 9, Angus, Cao, and Naeem teach the system of claim 1. Angus, Cao, and Naeem do not teach “wherein the event identifying unit is configured to identify indicators for events of interest based upon one or more labelled examples in the training data.” However, Dareddy teaches wherein the event identifying unit is configured to identify indicators for events of interest based upon one or more labelled examples in the training data. (Dareddy, Figure 6. Dareddy, page 21, lines 5-8, “the training method further comprises a step S608 of manually labelling each cluster (for a given signal) with a label indicating an event associated with that cluster. Here, manually labelling means that e.g. a data scientist or developer must provide an input to indicate a label that is to be given each cluster.” Dareddy, page 24, lines 15-18, “The different clusters may then be labelled (e.g. 'combat', 'ambling', etc.) and used to train a machine learning model. The machine learning model may be trained to identify, for a given frame of button presses, a corresponding label for that frame.” Examiner notes the event identifying unit is the machine learning model trained to identify labels corresponding to events.) Angus, Cao, Naeem, and Dareddy are analogous to the claimed invention because they involve creating specialized training datasets for training machine learning models from unlabeled data. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Angus, Cao, and Naeem to implement labeled examples in the training data to configure the event identifying unit to identify indicators for events of interest. Doing so would be advantageous because “Training the models in this way can be advantageous in that the machine learning models can be trained in a more bespoke manner. For example, generating feature representations using a pre-trained model such as e.g. DenseNet may be inefficient because the pre-trained model will likely have been trained using thousands of images that have no relevance to a particular video game. As a result, the use of such a pre-trained model may be excessive in terms of the memory required to store it and the time taken to execute it (requiring the use of a GPU, for example)” (Dareddy, page 23, lines 26-32). Claim(s) 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Angus in view of Cao in further view of Naeem in further view of Xie et al. (“Predicting Player Disengagement and First Purchase with Event-Frequency Based Data Representation”). Regarding Claim 10, Angus, Cao, and Naeem teach the system of claim 1. Angus, Cao, and Naeem do not teach “wherein the probabilities for respective events are updated in response to the addition of identified events to the list.” However, Xie teaches wherein the probabilities for respective events are updated in response to the addition of identified events to the list. (Xie, page 231, Table III. Xie, page 233, “For each player, their total activities (the sum of all event frequency features) in both month 1 and month 2 will be calculated separately and sorted.” Examiner notes that the probabilities are the total frequencies of each event, which is automatically calculated and updated with the addition of new events.) Angus, Cao, Naeem, and Xie are analogous to the claimed invention because they involve creating specialized training datasets from unlabeled data. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Angus, Cao, and Naeem to implement frequency-based analysis for the events in the dataset. This approach is advantageous because “this data representation is able to provide generality because it solely relies on the frequency of game events rather than any prior knowledge of them” (Xie, page 230) and “the results indicated that the event frequency based data representation can be successfully applied on both games and achieves a competitive performance” (Xie, page 230). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Zared O. Cohen whose telephone number is (571)270-0531. The examiner can normally be reached M-Th, 8am to 5pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle Bechtold can be reached at (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Z.O.C./Examiner, Art Unit 2148 /MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148
Read full office action

Prosecution Timeline

Nov 20, 2023
Application Filed
Sep 01, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month