Prosecution Insights
Last updated: August 17, 2026
Application No. 18/961,655

CONTENT MODERATION FOR ARTIFICIAL INTELLIGENCE (AI) SYSTEMS

Non-Final OA §103
Filed
Nov 27, 2024
Examiner
LEE, EUNICE SOMIN
Art Unit
2656
Tech Center
2600 — Communications
Assignee
Amazon Technologies Inc.
OA Round
1 (Non-Final)
88%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 88% — above average
88%
Career Allowance Rate
38 granted / 43 resolved
+26.4% vs TC avg
Strong +26% interview lift
Without
With
+26.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
9 currently pending
Career history
55
Total Applications
across all art units

Statute-Specific Performance

§101
24.6%
-15.4% vs TC avg
§103
63.1%
+23.1% vs TC avg
§102
8.2%
-31.8% vs TC avg
§112
2.5%
-37.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 43 resolved cases

Office Action

§103
DETAILED ACTION This communication is in response to the Application filed on November 27, 2024. Claims 1 - 20 are pending and have been examined. Claims 1, 5 and 13 are independent. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on February 4, 2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Drawings The drawings filed on November 27, 2024 have been accepted and considered by the Examiner. Double Patenting Note The Examiner notes that previously issued patents U.S. 10,962,939, 11,423,265, 12,586,578 and 12,647,492 and previously published patent applications U.S. Patent Application 2024/0428783 were analyzed for Double Patenting. However, based on the current claim scope no Double patenting was found. Claim Rejections - 35 USC § 103 The following is a quotation of pre-AIA 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action: (a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made. Claims 1 - 3, 5 - 20 are rejected under 35 U.S.C. 103(a) as being unpatentable over Mathur et al., (U.S. Patent Application Publication 2026/0057218), hereinafter referred to as Mathur, in view of Amatriain-Rubio et al., (U.S. Patent Application Publication 2025/0005288), hereinafter referred to as Rubio. Regarding Claim 1, Mathur teaches: 1. A computer-implemented method comprising: receiving text data representing a natural language user input; [Mathur, “While not meant to be particularly limited, the LLM 200 and/or encoder 114 can include a neural network machine learning architecture that is capable of processing large amounts of text data and generating high-quality natural language responses.” Par. 0022; “The input 202 denotes an input provided by a user (i.e., the claimed “user input”) (or upstream system) and can be represented as a sequence of tokens, individual words or sub-words (i.e., the claimed “text data”), from which input embeddings 204 can be generated.” Par. 0026; “Different approaches to chunking an input text (i.e., the claimed “receiving text data representing a natural language user input”) and/or multimodal content are possible, such as fixed size chunking, heuristic chunking, and semantic chunking, and all such configurations and combinations thereof are within the contemplated scope of this disclosure.” Par. 0048] determining a first prompt including the text data and a first request to generate a response to the natural language user input; [Mathur, “While not meant to be particularly limited, the LLM 200 and/or encoder 114 can include a neural network machine learning architecture that is capable of processing large amounts of text data and generating high-quality natural language responses.” Par. 0022; “The input 202 denotes an input provided by a user (i.e., the claimed “user input”) (or upstream system) and can be represented as a sequence of tokens, individual words or sub-words (i.e., the claimed “text data”), from which input embeddings 204 can be generated.” Par. 0026; “Different approaches to chunking an input text (i.e., the claimed “text data representing a natural language user input”) and/or multimodal content are possible, such as fixed size chunking, heuristic chunking, and semantic chunking, and all such configurations and combinations thereof are within the contemplated scope of this disclosure.” Par. 0048; “Thus, the dynamic multimodal prompt generation (i.e., the claimed “determining a first prompt”) system described herein streamlines the content moderation process by ensuring that content moderation decisions are more responsive to policy changes and emerging content trends, and improves policy design itself by identifying policy gaps and ambiguities in those policies.” Par. 0017; “Returning to the multimodal prompt generation system 102, in some embodiments, the multimodal prompt generation system 102 dynamically builds a multimodal prompt 120 (i.e., the claimed “determining a first prompt”),” Par. 0036] generating, by a first language model and based on the first prompt, a first number of tokens corresponding to a first portion of the response; [Mathur, “In some embodiments, the prompt 120 (i.e., the claimed “first prompt”) is passed to large language model (i.e., the claimed “first language model”) 106 (i.e., the claimed “generating, by a first language model and based on the first prompt”).” Par. 0039; “In other words, the prompt 120 passed to the large language model 106 (i.e., the claimed “generating, by a first language model and based on the first prompt”) will dynamically change in response to changes in policy,” Par. 0039; “At its core, a large language model consists of an encoder and a decoder. The encoder takes in a sequence of input tokens, such as words or characters, and produces a sequence of hidden representations for each token that capture the contextual information of the input sequence. The decoder then uses these hidden representations, along with a sequence of target tokens, to generate a sequence of output tokens (i.e., the claimed “first number of tokens corresponding to a first portion of the response”).” Par. 0023] processing, using a second language model, the first number of tokens to determine that the first portion of the response corresponds to a non-moderated content category, [Mathur, Mathur, “Notably, in some embodiments, large language model 106 is pre-trained to classify content 110 and/or to review and apply one or more polices to content 110 (e.g., to determine whether some content violates a policy).” Par. 0039; “The decoder then uses these hidden representations, along with a sequence of target tokens, to generate a sequence of output tokens (i.e., the claimed “first number of tokens”).” Par. 0023; “For example, consider a scenario where the large language model 106 (i.e., the claimed “second language model”) is trained to provide a predictive score, where a score of 0 indicates complete confidence in no policy violation (i.e., the claimed “processing, using a second language model, the first number of tokens to determine that the first portion of the response corresponds to a non-moderated content category”) and a score of 1 indicates complete confidence in a policy violation for a particular content 110—retrieved policy chunk 112 pair.” Par. 0041; “For example, prompt template 122 might include instructions to the large language model 106 (i.e., the claimed “second language model”) to consider content 110 in the context of the retrieved policy chunks 112 and to determine whether the content 110 violates any portion(s) of the retrieved policy chunks 112 (i.e., the claimed “processing, using a second language model, the first number of tokens to determine that the first portion of the response corresponds to a non-moderated content category”).” Par. 0037; “In some embodiments, the EBR module 116 can be configured for rule-based exclusions, thereby allowing the EBR module 116 to handle unique and/or specific scenarios such as geographical carve outs. For example, if the content 110 is posted from country Z, Policy T does not apply (i.e., the claimed “non-moderated content category”), and a rule-based exclusion can be enforced such that the EBR module 116 will not include policy chunks from Policy T in the retrieved policy chunks 112 (e.g., if country Z, then ignore policy T) (i.e., the claimed “non-moderated content category”).” Par. 0031] wherein the second language model is configured to determine whether inputted tokens correspond to one or more of a set of moderated content categories; [Mathur, Mathur, “Notably, in some embodiments, large language model 106 is pre-trained to classify content 110 and/or to review and apply one or more polices to content 110 (e.g., to determine whether some content violates a policy).” Par. 0039; “For example, consider a scenario where the large language model 106 (i.e., the claimed “second language model”) is trained to provide a predictive score, where a score of 0 indicates complete confidence in no policy violation and a score of 1 indicates complete confidence in a policy violation (i.e., the claimed “second language model is configured to determine whether inputted tokens correspond to one or more of a set of moderated content categories”) for a particular content 110—retrieved policy chunk 112 pair.” Par. 0041; “In some embodiments, prompt template 122 contains content moderation instructions for the large language model 106, such as instructions on moderating content based on any attached or referenced policy documents (i.e., the claimed “one or more of a set of moderated content categories”).” Par. 0037; “At block 502, the method includes receiving, by a multimodal prompt generation system, a request for a decision (e.g., a content moderation decision) for content.” Par. 0068; “Content moderation refers to the practice of monitoring and applying a set of rules, guidelines, and policies to user-generated content submissions. This process determines whether a particular piece of content should be published, flagged, restricted, modified, and/or removed from a platform. Content moderation typically covers a wide range of issues, including but not limited to the identification of illegal content, hate speech and discrimination, violence and graphic content, harassment and bullying, misinformation and disinformation, copyright infringement, spam, and scams (collectively, “flagged” content).” Par. 0011; Mathur teaches categorical embedding space: “At block 506, the method includes retrieving, by an embedding based retrieval (EBR) module of the prompt generation system, K retrieved chunks (e.g., policy chunks) from a database (e.g., a policy database) having a plurality of source documents (e.g., policies), the K retrieved chunks having a Kth closest distance to the embedding in an embedding space (i.e., categorial embedding space). In some embodiments, the retrieved chunks are associated with multiple source documents (e.g., multiple separate policies) of the plurality of source documents in the database.” Par. 0070; “This disclosure introduces a dynamic multimodal prompt generation system and framework for efficient content moderation. In some embodiments, a collection of policy documents is segmented into a plurality of so-called policy chunks, which are stored in a policy database. Later, a selection of those policy chunks is combined with content to dynamically generate a content moderation prompt that is unique to the respective policy chunks and content.” Par. 0014] in response to the first portion of the response corresponding to the non-moderated content category, causing presentation of the first number of tokens; [Mathur, Mathur, “Notably, in some embodiments, large language model 106 is pre-trained to classify content 110 and/or to review and apply one or more polices to content 110 (e.g., to determine whether some content violates a policy).” Par. 0039; “In some embodiments, the content moderation service 100 is integrated with and/or otherwise incorporated within or alongside the external system(s). For example, the external system(s) can include publishing services and/or social media platforms for posting content 110, such as a connections network. In this manner, the content moderation service 100 can serve as an external and/or internal check (i.e., the claimed “in response to the first portion of the response corresponding to the non-moderated content category”) for provisional content prior to live publication (i.e., the claimed “causing presentation of the first number of tokens”) in the service or platform.” Par. 0046; “The encoder takes in a sequence of input tokens, such as words or characters, and produces a sequence of hidden representations for each token that capture the contextual information of the input sequence. The decoder then uses these hidden representations, along with a sequence of target tokens, to generate a sequence of output tokens (i.e., the claimed “first number of tokens”).” Par. 0023; “For example, consider a scenario where the large language model 106 (i.e., the claimed “second language model”) is trained to provide a predictive score, where a score of 0 indicates complete confidence in no policy violation (i.e., the claimed “processing, using a second language model, the first number of tokens to determine that the first portion of the response corresponds to a non-moderated content category”) and a score of 1 indicates complete confidence in a policy violation for a particular content 110—retrieved policy chunk 112 pair.” Par. 0041; “For example, prompt template 122 might include instructions to the large language model 106 (i.e., the claimed “second language model”) to consider content 110 in the context of the retrieved policy chunks 112 and to determine whether the content 110 violates any portion(s) of the retrieved policy chunks 112 (i.e., the claimed “processing, using a second language model, the first number of tokens to determine that the first portion of the response corresponds to a non-moderated content category”).” Par. 0037] wherein generating, by the first language model and based on the first prompt, a second number of tokens corresponding to a second portion of the response, wherein the second number is larger than the first number; [Mathur, “Different approaches to chunking an input text and/or multimodal content are possible, such as fixed size chunking, heuristic chunking, and semantic chunking, and all such configurations and combinations thereof are within the contemplated scope of this disclosure.” Par. 0048; “In some embodiments, finding optimal values for parameters such as the number of tokens in a chunk (i.e., the claimed “first number of tokens corresponding to the first portion of a response”, “second number tokens corresponding to the second portion of a response”, etc.), overlap values, and/or a maximum token limit, etc., are found iteratively by initializing those values as desired and then adjusting those values in view of empirical performance data.” Par. 0049; “If a paragraph is very large (e.g., exceeds a predetermined threshold length), a chunk can be split (i.e., splitting into the claimed “first number of tokens corresponding to the first portion of a response” and “second number tokens corresponding to the second portion of a response”),” Par. 0050; “In any case, context preservation can be estimated using heuristics and/or semantic analysis, as explained below (e.g., extend a chunk by a few tokens (i.e., extension enables the claimed “second number is larger than the first number”) to reach end of sentence, extend a chunk by a few tokens (i.e., extension enables the claimed “second number is larger than the first number”) to avoid cutting out material of high semantic similarity, etc.).” Par. 0049; Mathur, “In some embodiments, the prompt 120 (i.e., the claimed “first prompt”) is passed to large language model (i.e., the claimed “first language model”) 106 (i.e., the claimed “generating, by a first language model and based on the first prompt”).” Par. 0039; “In other words, the prompt 120 passed to the large language model 106 (i.e., the claimed “generating, by a first language model and based on the first prompt”) will dynamically change in response to changes in policy,” Par. 0039; “At its core, a large language model consists of an encoder and a decoder. The encoder takes in a sequence of input tokens, such as words or characters, and produces a sequence of hidden representations for each token that capture the contextual information of the input sequence. The decoder then uses these hidden representations, along with a sequence of target tokens, to generate a sequence of output tokens (i.e., the claimed “second number of tokens corresponding to a second portion of the response”).” Par. 0023 processing, using the second language model, the second number of tokens to determine that the second portion of the response corresponds to the non-moderated content category; and Mathur, Mathur, “Notably, in some embodiments, large language model 106 is pre-trained to classify content 110 and/or to review and apply one or more polices to content 110 (e.g., to determine whether some content violates a policy).” Par. 0039; “For example, consider a scenario where the large language model 106 (i.e., the claimed “second language model”) is trained to provide a predictive score, where a score of 0 indicates complete confidence in no policy violation (i.e., the claimed “processing, using a second language model, the second number of tokens to determine that the second portion of the response corresponds to a non-moderated content category”) and a score of 1 indicates complete confidence in a policy violation for a particular content 110—retrieved policy chunk 112 pair.” Par. 0041; “For example, prompt template 122 might include instructions to the large language model 106 (i.e., the claimed “second language model”) to consider content 110 in the context of the retrieved policy chunks 112 and to determine whether the content 110 violates any portion(s) of the retrieved policy chunks 112 (i.e., the claimed “processing, using a second language model, the second number of tokens to determine that the second portion of the response corresponds to a non-moderated content category”).” Par. 0037; “In some embodiments, the EBR module 116 can be configured for rule-based exclusions, thereby allowing the EBR module 116 to handle unique and/or specific scenarios such as geographical carve outs. For example, if the content 110 is posted from country Z, Policy T does not apply (i.e., the claimed “non-moderated content category”), and a rule-based exclusion can be enforced such that the EBR module 116 will not include policy chunks from Policy T in the retrieved policy chunks 112 (e.g., if country Z, then ignore policy T) (i.e., the claimed “non-moderated content category”).” Par. 0031] in response to the second portion of the response corresponding to the non-moderated content category, causing presentation of the second number of tokens. [Mathur, “Notably, in some embodiments, large language model 106 is pre-trained to classify content 110 and/or to review and apply one or more polices to content 110 (e.g., to determine whether some content violates a policy).” Par. 0039;“In some embodiments, the content moderation service 100 is integrated with and/or otherwise incorporated within or alongside the external system(s). For example, the external system(s) can include publishing services and/or social media platforms for posting content 110, such as a connections network. In this manner, the content moderation service 100 can serve as an external and/or internal check (i.e., the claimed “in response to the second portion of the response corresponding to the non-moderated content category”) for provisional content prior to live publication (i.e., the claimed “causing presentation of the second number of tokens”) in the service or platform.” Par. 0046; “The encoder takes in a sequence of input tokens, such as words or characters, and produces a sequence of hidden representations for each token that capture the contextual information of the input sequence. The decoder then uses these hidden representations, along with a sequence of target tokens, to generate a sequence of output tokens (i.e., the claimed “second number of tokens”).” Par. 0023; “For example, consider a scenario where the large language model 106 (i.e., the claimed “second language model”) is trained to provide a predictive score, where a score of 0 indicates complete confidence in no policy violation (i.e., the claimed “processing, using a second language model, the second number of tokens to determine that the second portion of the response corresponds to a non-moderated content category”) and a score of 1 indicates complete confidence in a policy violation for a particular content 110—retrieved policy chunk 112 pair.” Par. 0041; “For example, prompt template 122 might include instructions to the large language model 106 (i.e., the claimed “second language model”) to consider content 110 in the context of the retrieved policy chunks 112 and to determine whether the content 110 violates any portion(s) of the retrieved policy chunks 112 (i.e., the claimed “processing, using a second language model, the second number of tokens to determine that the second portion of the response corresponds to a non-moderated content category”).” Par. 0037] Mathur fails to explicitly teach category. However, Rubio teaches: Regarding Claim 1, Rubio teaches: 1. A computer-implemented method comprising: receiving text data representing a natural language user input; [Rubio, “The disclosed technologies are generative in that one or more generative models (e.g., LLMs) are used to machine-generate and output responses to user requests in a conversational natural language manner (i.e., the claimed “natural language user input”).” Par. 0028; “Any dialog, thread, or thread portion can include one or more different types of digital content, including natural language text (i.e., the claimed “text data”), audio, video, digital imagery, hyperlinks, and/or multimodal content such as web pages.” Par. 0038] determining a first prompt including the text data and a first request to generate a response to the natural language user input; [Rubio, “A generative language model is a particular type of generative model that generates new text in response to model input (i.e., the claimed “natural language user input”). The model input includes a task description, also referred to as a prompt (i.e., the claimed “first prompt”). The task description can include instructions and/or examples of digital content. A task description can be in the form of natural language text (i.e., the claimed “text data”), such as a question or a statement (i.e., the claimed “natural language user input”), and can include non-text forms of content, such as digital imagery and/or digital audio.” Par. 0022; “As such, thread classification prompt generator 104 and plan execution prompt generator 112 are each specially configured to cause one or more large language models to generate and output thread portions that are responsive to user-generated thread portions in accordance with specific parameters, instructions, and constraints that are applicable to a specific task to be performed by the one or more large language models, such as thread classification or plan execution.” Par. 0045] generating, by a first language model and based on the first prompt, a first number of tokens corresponding to a first portion of the response; [Rubio, “For example, a thread can include a first thread portion (i.e., the claimed “first portion”), such as a question received from a user of a computing device, and a second thread portion, such as natural language text, audio, video, and/or imagery machine-generated by the user assistance system in response to the user's question.” Par, 0037; “A generative language model (i.e., the claimed “first language model”) is a particular type of generative model that generates new text in response (i.e., the claimed “first portion of the response”) to model input (i.e., the claimed “natural language user input”). The model input includes a task description, also referred to as a prompt (i.e., the claimed “first prompt”). The task description can include instructions and/or examples of digital content. A task description can be in the form of natural language text (i.e., the claimed “text data”), such as a question or a statement (i.e., the claimed “natural language user input”), and can include non-text forms of content, such as digital imagery and/or digital audio.” Par. 0022] processing, using a second language model, the first number of tokens to determine that the first portion of the response corresponds to a non-moderated content category, [Rubio, “For example, a thread can include a first thread portion (i.e., the claimed “first portion”), such as a question received from a user of a computing device, and a second thread portion, such as natural language text, audio, video, and/or imagery machine-generated by the user assistance system in response to the user's question.” Par, 0037; “The example thread classification prompt also constrains the large language model (i.e., the claimed “second language model”) to a set number of possible categories into which the user input is to be classified (e.g., the large language model is required to pick only one category (i.e., the claimed “non-moderated content category”)), specifies the applicable thread context data (e.g., user profile, dialog history, previous user input, categories, and job recommendations), and specifies the output format for the thread (i.e., the claimed “portion”) classification to be produced by the large language model (e.g., natural language text) (i.e., the claimed “first portion response corresponds to a non-moderated content category”).” Par. 0099; “As such, thread classification prompt generator 104 and plan execution prompt generator 112 are each specially configured to cause one or more large language models to generate and output thread portions that are responsive to user-generated thread portions in accordance with specific parameters, instructions, and constraints that are applicable to a specific task to be performed by the one or more large language models, such as thread classification or plan execution.” Par. 0045] wherein the second language model is configured to determine whether inputted tokens correspond to one or more of a set of moderated content categories; [Rubio, “For example, a thread can include a first thread portion (i.e., the claimed “first portion”), such as a question received from a user of a computing device, and a second thread portion, such as natural language text, audio, video, and/or imagery machine-generated by the user assistance system in response to the user's question.” Par, 0037; “The example thread classification prompt also constrains the large language model (i.e., the claimed “second language model”) to a set number of possible categories (i.e., the claimed “one or more of a set of moderated content categories”) into which the user input is to be classified (e.g., the large language model is required to pick only one category), specifies the applicable thread context data (e.g., user profile, dialog history, previous user input, categories, and job recommendations), and specifies the output format for the thread (i.e., the claimed “portion”) classification to be produced by the large language model (e.g., natural language text).” Par. 0099] in response to the first portion of the response corresponding to the non-moderated content category, causing presentation of the first number of tokens; [Rubio, “The disclosed technologies are generative in that one or more generative models (e.g., LLMs) are used to machine-generate and output responses to user requests (i.e., the claimed “causing presentation of the first number of tokens”) in a conversational natural language manner.” Par. 0028; “For example, a classification prompt as used herein can include instructions to cause a large language model to output a classification (e.g., the large language model operates in a discriminative manner), while a plan execution prompt as used here can include instructions to cause a large language model to execute a plan (e.g., a multi-step prompt) to machine-generate and output (i.e., the claimed “causing presentation of the first number of tokens”) one or more thread portions (i.e., the claimed “first portion of the response”) (e.g., the large language model operates in a generative manner).” Par. 0044; “As such, thread classification prompt generator 104 and plan execution prompt generator 112 are each specially configured to cause one or more large language models to generate and output thread portions that are responsive to user-generated thread portions in accordance with specific parameters, instructions, and constraints that are applicable to a specific task to be performed by the one or more large language models, such as thread classification or plan execution.” Par. 0045] generating, by the first language model and based on the first prompt, a second number of tokens corresponding to a second portion of the response, wherein the second number is larger than the first number; [Rubio, “As such, thread classification prompt generator 104 and plan execution prompt generator 112 are each specially configured to cause one or more large language models (i.e., the claimed “first language model”) to generate and output thread portions that are responsive to user-generated thread portions in accordance with specific parameters, instructions, and constraints that are applicable to a specific task to be performed by the one or more large language models, such as thread classification or plan execution.” Par. 0045] processing, using the second language model, the second number of tokens to determine that the second portion of the response corresponds to the non-moderated content category; and [Rubio, “As such, thread classification prompt generator 104 and plan execution prompt generator 112 are each specially configured to cause one or more large language models to generate and output thread portions that are responsive to user-generated thread portions in accordance with specific parameters, instructions, and constraints that are applicable to a specific task to be performed by the one or more large language models (i.e., the claimed “second language model”), such as thread classification (i.e., the claimed “non-moderated content category”) or plan execution.” Par. 0045] in response to the second portion of the response corresponding to the non-moderated content category, causing presentation of the second number of tokens. [Rubio, “As such, thread classification prompt generator 104 and plan execution prompt generator 112 are each specially configured to cause one or more large language models to generate and output thread portions (i.e., the claimed “first portion of the response”, “second portion of the response”, etc.) that are responsive to user-generated thread portions (i.e., the claimed “first portion”, “second portion”, etc.) in accordance with specific parameters, instructions, and constraints that are applicable to a specific task to be performed by the one or more large language models, such as thread classification or plan execution.” Par. 0045; For example, a classification prompt as used herein can include instructions to cause a large language model to output a classification (e.g., the large language model operates in a discriminative manner), while a plan execution prompt as used here can include instructions to cause a large language model to execute a plan (e.g., a multi-step prompt) to machine-generate and output (i.e., the claimed “causing presentation of the second number of tokens”) one or more thread portions (i.e., the claimed “second portion of the response”) (e.g., the large language model operates in a generative manner).” Par. 0044] Mathur and Rubio pertain technologies that are “portion based”/ “thread based”/ “chunk based” for efficiency and are analogous to the instant application. Accordingly, it would have been obvious to one of ordinary skill in technologies that are “portion based”/ “thread based”/ “chunk based” for efficiency art to modify Mathur’s teachings of “machine learning, online platforms, and content moderation, and specifically to dynamic multimodal prompt generation for efficient content moderation” (Mathur, Par. 0001) with the explicit teachings of “categories” (Rubio, Par. 0099) taught by Rubio in order prevent “generating content that is nonsensical or unfaithful to the provided source content” (Rubio, Par. 0025). Regarding Claim 2, Mathur in view of Rubio has been discussed above. The combination further teaches: determining a confidence score associated with the second language model processing of the second number of tokens; [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1] determining that the confidence score satisfies a condition; [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Mathur, “For example, consider a scenario where the large language model 106 is trained to provide a predictive score, where a score of 0 indicates complete confidence in no policy violation and a score of 1 indicates complete confidence in a policy violation for a particular content 110—retrieved policy chunk 112 pair. In some embodiments, scores (i.e., the claimed “confidence score”) above a predetermined threshold (e.g., above 0.51, above 0.6, above 0.9, etc.) (i.e., the claimed “confidence score satisfies a condition”) indicate partial confidence of various degrees in a policy violation and scores below a predetermined threshold (e.g., below 0.49, below 0.4, below 0.25, etc.) indicate partial confidence of various degrees in no policy violation being present.” Par. 0041] based on the confidence score satisfying the condition, determining a third number of tokens to be processed by the second language model, [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Mathur, “For example, a retrieved policy chunk 112 which is frequently associated with (again, against any predetermined threshold (i.e., the claimed “satisfies a condition”) ambiguous predictive scores can be re-chunked (i.e., re-chunking enables the claimed “third number of tokens to be processed”) to increase or decrease the scope (e.g., amount of text, number of video/audio tokens, etc.) of the respective retrieved policy chunk 112.” Par. 0042] wherein the third number is larger than the second number; [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Claim is directed to repeating the subject matter for third number. However, third number/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] generating, based on the first language model processing the first prompt, the third number of tokens corresponding to a third portion of the response; [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Claim is directed to repeating the subject matter for third number. However, third number/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] processing, using the second language model, the third number of tokens to determine that the third portion of the response corresponds to the non-moderated content category; and [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Claim is directed to repeating the subject matter for third number. However, third number/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] in response to the third portion of the response corresponding to the non-moderated content category, causing presentation of the third number of tokens. [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Claim is directed to repeating the subject matter for third number. However, third number/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] Regarding Claim 3, Mathur in view of Rubio has been discussed above. The combination further teaches: determining a second prompt including the first number of tokens and a second request to determine whether the first portion of the response corresponds to one of a set of moderated content categories, [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Claim is directed to repeating the subject matter for second prompt and second request. However, second prompt/second request/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] wherein processing, using the second language model, the first number of tokens comprises processing the second prompt using the second language model; [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Claim is directed to repeating the subject matter for second request. However, second request/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] determining, based on the second language model processing the second prompt, embedding data corresponding to the second prompt; and [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Mathur, “In some embodiments, the multimodal prompt generation system 102 includes an encoder 114 (e.g., a large language model (i.e., the claimed “second language model”) encoder) that generates embeddings (i.e., the claimed “embedding data”) for the content 110,” Par. 0020; Claim is directed to repeating the subject matter for second prompt. However, second prompt/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] determining a third prompt including the first number of tokens, the second number of tokens and a third request to determine whether the first portion of the response and the second portion of the response correspond to one of the set of moderated content categories, [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Claim is directed to repeating the subject matter for third request. However, third request/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] wherein processing, using the second language model, the second number of tokens comprises processing, using the second language model, the embedding data and a third portion of the third prompt. [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Mathur, “In some embodiments, the multimodal prompt generation system 102 includes an encoder 114 (e.g., a large language model (i.e., the claimed “second language model”) encoder) that generates embeddings (i.e., the claimed “embedding data”) for the content 110,” Par. 0020; Claim is directed to repeating the subject matter for third portion and third prompt. However, third prompt/third portion/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] Regarding Claims 5 and 13, Mathur in view of Rubio has been discussed above. The combination further teaches: at least one processor; and [Mathur, “The computer system 400 includes at least one processing device 402, which generally includes one or more processors or processing units for performing a variety of functions, such as, for example, completing any portion of the content moderation service 100 described previously. Components of the computer system 400 also include a system memory 404, and a bus 406 that couples various system components including the system memory 404 to the processing device 402. The system memory 404 may include a variety of computer system readable media.” Par. 0063] at least one memory including instructions that, when executed by the at least one processor, cause the system to: [Mathur, “The computer system 400 includes at least one processing device 402, which generally includes one or more processors or processing units for performing a variety of functions, such as, for example, completing any portion of the content moderation service 100 described previously. Components of the computer system 400 also include a system memory 404, and a bus 406 that couples various system components including the system memory 404 to the processing device 402. The system memory 404 may include a variety of computer system readable media.” Par. 0063; “The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.” Par. 0094 receiving user input data; [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1] processing, using a first generative model, the user input data to generate first tokens corresponding to a first portion of a response to the user input data; [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Rubio, “A generative language model is a particular type of generative model (i.e., the claimed “first generative model”, “second generative model”, etc.) that generates new text in response to model input. The model input includes a task description, also referred to as a prompt.” Par. 0022; Mathur, “According to the techniques described herein, personal data provided by users may be processed to determine prompts when using a generative AI (i.e., generative AI includes the claimed “first generative model, second generative model, etc.) feature at the request of the user,” Par. 0084] determining that the first tokens correspond to a first content category; [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Rubio, “The example thread classification prompt also constrains the large language model (i.e., the claimed “second language model”) to a set number of possible categories into which the user input is to be classified (e.g., the large language model is required to pick only one category (i.e., the claimed “first content category”)), specifies the applicable thread context data (e.g., user profile, dialog history, previous user input, categories, and job recommendations), and specifies the output format for the thread (i.e., the claimed “portion”) classification to be produced by the large language model (e.g., natural language text).” Par. 0099] in response to the first tokens corresponding to the first content category, [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Rubio, “The example thread classification prompt also constrains the large language model (i.e., the claimed “second language model”) to a set number of possible categories into which the user input is to be classified (e.g., the large language model is required to pick only one category (i.e., the claimed “first content category”)), specifies the applicable thread context data (e.g., user profile, dialog history, previous user input, categories, and job recommendations), and specifies the output format for the thread (i.e., the claimed “portion”) classification to be produced by the large language model (e.g., natural language text).” Par. 0099] sending the first tokens to a system component for further processing; [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Mathur, “This disclosure introduces a dynamic multimodal prompt generation system and framework for efficient content moderation.” Par. 0014; “In some embodiments, a positional encoding 208 can be generated to encode the position of each token (i.e., the claimed “first tokens”) in input 202 as a set of numbers. These numbers can be fed (i.e., the claimed “sending”) into the encoder 206 with the input embeddings 204,” Par. 0026] while sending the first tokens to the system component, [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Mathur, “This disclosure introduces a dynamic multimodal prompt generation system and framework for efficient content moderation.” Par. 0014; “In some embodiments, a positional encoding 208 can be generated to encode the position of each token (i.e., the claimed “first tokens”) in input 202 as a set of numbers. These numbers can be fed (i.e., the claimed “sending”) into the encoder 206 with the input embeddings 204,” Par. 0026; “In some embodiments, finding optimal values for parameters such as the number of tokens in a chunk (i.e., the claimed “first number of tokens corresponding to the first portion of a response”, “second number tokens corresponding to the second portion of a response”, etc.), overlap values, and/or a maximum token limit, etc., are found iteratively by initializing those values as desired and then adjusting those values in view of empirical performance data.” Par. 0049; “If a paragraph is very large (e.g., exceeds a predetermined threshold length), a chunk can be split (i.e., splitting into the claimed “first number of tokens corresponding to the first portion of a response” and “second number tokens corresponding to the second portion of a response”),” Par. 0050] processing, using the first generative model, to generate second tokens corresponding to a second portion of the response to the user input data; [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; “In some embodiments, a positional encoding 208 can be generated to encode the position of each token (i.e., the claimed “first tokens”) in input 202 as a set of numbers. These numbers can be fed (i.e., the claimed “sending”) into the encoder 206 with the input embeddings 204,” Par. 0026; “In some embodiments, finding optimal values for parameters such as the number of tokens in a chunk (i.e., the claimed “first number of tokens corresponding to the first portion of a response”, “second number tokens corresponding to the second portion of a response”, etc.), overlap values, and/or a maximum token limit, etc., are found iteratively by initializing those values as desired and then adjusting those values in view of empirical performance data.” Par. 0049; “If a paragraph is very large (e.g., exceeds a predetermined threshold length), a chunk can be split (i.e., splitting into the claimed “first number of tokens corresponding to the first portion of a response” and “second number tokens corresponding to the second portion of a response”),” Par. 0050] determining the second tokens correspond to the first content category; and [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Mathur, “This disclosure introduces a dynamic multimodal prompt generation system and framework for efficient content moderation.” Par. 0014; “In some embodiments, a positional encoding 208 can be generated to encode the position of each token (i.e., the claimed “first tokens”) in input 202 as a set of numbers. These numbers can be fed (i.e., the claimed “sending”) into the encoder 206 with the input embeddings 204,” Par. 0026; “In some embodiments, finding optimal values for parameters such as the number of tokens in a chunk (i.e., the claimed “first number of tokens corresponding to the first portion of a response”, “second number tokens corresponding to the second portion of a response”, etc.), overlap values, and/or a maximum token limit, etc., are found iteratively by initializing those values as desired and then adjusting those values in view of empirical performance data.” Par. 0049; “If a paragraph is very large (e.g., exceeds a predetermined threshold length), a chunk can be split (i.e., splitting into the claimed “first number of tokens corresponding to the first portion of a response” and “second number tokens corresponding to the second portion of a response”),” Par. 0050; “The example thread classification prompt also constrains the large language model (i.e., the claimed “second language model”) to a set number of possible categories into which the user input is to be classified (e.g., the large language model is required to pick only one category (i.e., the claimed “first content category”)), specifies the applicable thread context data (e.g., user profile, dialog history, previous user input, categories, and job recommendations), and specifies the output format for the thread (i.e., the claimed “portion”) classification to be produced by the large language model (e.g., natural language text).” Par. 0099] in response to the second tokens corresponding to the first content category, [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Mathur, “This disclosure introduces a dynamic multimodal prompt generation system and framework for efficient content moderation.” Par. 0014; “In some embodiments, a positional encoding 208 can be generated to encode the position of each token (i.e., the claimed “first tokens”) in input 202 as a set of numbers. These numbers can be fed (i.e., the claimed “sending”) into the encoder 206 with the input embeddings 204,” Par. 0026; “In some embodiments, finding optimal values for parameters such as the number of tokens in a chunk (i.e., the claimed “first number of tokens corresponding to the first portion of a response”, “second number tokens corresponding to the second portion of a response”, etc.), overlap values, and/or a maximum token limit, etc., are found iteratively by initializing those values as desired and then adjusting those values in view of empirical performance data.” Par. 0049; “If a paragraph is very large (e.g., exceeds a predetermined threshold length), a chunk can be split (i.e., splitting into the claimed “first number of tokens corresponding to the first portion of a response” and “second number tokens corresponding to the second portion of a response”),” Par. 0050; “The example thread classification prompt also constrains the large language model (i.e., the claimed “second language model”) to a set number of possible categories into which the user input is to be classified (e.g., the large language model is required to pick only one category (i.e., the claimed “first content category”)), specifies the applicable thread context data (e.g., user profile, dialog history, previous user input, categories, and job recommendations), and specifies the output format for the thread (i.e., the claimed “portion”) classification to be produced by the large language model (e.g., natural language text).” Par. 0099] sending the second tokens to the system component for further processing. [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Mathur, “This disclosure introduces a dynamic multimodal prompt generation system and framework for efficient content moderation.” Par. 0014; “In some embodiments, a positional encoding 208 can be generated to encode the position of each token (i.e., the claimed “first tokens”) in input 202 as a set of numbers. These numbers can be fed (i.e., the claimed “sending”) into the encoder 206 with the input embeddings 204,” Par. 0026; “In some embodiments, finding optimal values for parameters such as the number of tokens in a chunk (i.e., the claimed “first number of tokens corresponding to the first portion of a response”, “second number tokens corresponding to the second portion of a response”, etc.), overlap values, and/or a maximum token limit, etc., are found iteratively by initializing those values as desired and then adjusting those values in view of empirical performance data.” Par. 0049; “If a paragraph is very large (e.g., exceeds a predetermined threshold length), a chunk can be split (i.e., splitting into the claimed “first number of tokens corresponding to the first portion of a response” and “second number tokens corresponding to the second portion of a response”),” Par. 0050] Regarding Claims 6 and 14, Mathur in view of Rubio has been discussed above. The combination further teaches: wherein the at least one memory includes further instructions that, when executed by the at least one processor, further cause the system to: [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5] based on determining the first tokens correspond to the first content category, [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5] determining a number of the second tokens to be processed, [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5] wherein the second tokens include a larger number of tokens than the first tokens. [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5] Regarding Claims 7 and 15, Mathur in view of Rubio has been discussed above. The combination further teaches: wherein the at least one memory includes further instructions that, when executed by the at least one processor, further cause the system to: [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5] determining, using a trained model, that the second tokens correspond to the first content category, [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5; Mathur, “Notably, in some embodiments, large language model 106 is pre-trained to classify content 110 and/or to review and apply one or more polices to content 110 (e.g., to determine whether some content violates a policy).” Par. 0039; “For example, consider a scenario where the large language model 106 is trained (i.e., the claimed “trained model”) to provide a predictive score, where a score of 0 indicates complete confidence in no policy violation and a score of 1 indicates complete confidence in a policy violation for a particular content 110—retrieved policy chunk 112 pair. In some embodiments, scores above a predetermined threshold (e.g., above 0.51, above 0.6, above 0.9, etc.) indicate partial confidence of various degrees in a policy violation and scores below a predetermined threshold (e.g., below 0.49, below 0.4, below 0.25, etc.) indicate partial confidence of various degrees in no policy violation being present (i.e., the claimed “categories”).” Par. 0041; Rubio, “Examples of recommendation systems 182 include trained machine learning-based scoring models, ranking models, and/or classification models, such as job recommender models and connection recommender models.” Par. 0105] wherein the trained model is configured to determine whether inputted tokens correspond to one or more of a set of moderated content categories; and [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5; Mathur, “Notably, in some embodiments, large language model 106 is pre-trained to classify content 110 and/or to review and apply one or more polices to content 110 (e.g., to determine whether some content violates a policy).” Par. 0039; “For example, consider a scenario where the large language model 106 is trained (i.e., the claimed “trained model”) to provide a predictive score, where a score of 0 indicates complete confidence in no policy violation and a score of 1 indicates complete confidence in a policy violation for a particular content 110—retrieved policy chunk 112 pair. In some embodiments, scores above a predetermined threshold (e.g., above 0.51, above 0.6, above 0.9, etc.) indicate partial confidence of various degrees in a policy violation and scores below a predetermined threshold (e.g., below 0.49, below 0.4, below 0.25, etc.) indicate partial confidence of various degrees in no policy violation being present (i.e., the claimed “categories”).” Par. 0041; Rubio, “Examples of recommendation systems 182 include trained machine learning-based scoring models, ranking models, and/or classification models, such as job recommender models and connection recommender models.” Par. 0105] based on the trained model processing of the second tokens, determining a number of tokens to be subsequently processed by the trained model. [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5; Mathur, “Notably, in some embodiments, large language model 106 is pre-trained to classify content 110 and/or to review and apply one or more polices to content 110 (e.g., to determine whether some content violates a policy).” Par. 0039; “For example, consider a scenario where the large language model 106 is trained (i.e., the claimed “trained model”) to provide a predictive score, where a score of 0 indicates complete confidence in no policy violation and a score of 1 indicates complete confidence in a policy violation for a particular content 110—retrieved policy chunk 112 pair. In some embodiments, scores above a predetermined threshold (e.g., above 0.51, above 0.6, above 0.9, etc.) indicate partial confidence of various degrees in a policy violation and scores below a predetermined threshold (e.g., below 0.49, below 0.4, below 0.25, etc.) indicate partial confidence of various degrees in no policy violation being present (i.e., the claimed “categories”).” Par. 0041; Rubio, “Examples of recommendation systems 182 include trained machine learning-based scoring models, ranking models, and/or classification models, such as job recommender models and connection recommender models.” Par. 0105] Regarding Claims 8 and 16, Mathur in view of Rubio has been discussed above. The combination further teaches: wherein the at least one memory includes further instructions that, when executed by the at least one processor, further cause the system to: [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5] processing, using the first generative model for a first generation step to generate the first tokens; and [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5; Rubio, “A generative language model is a particular type of generative model (i.e., the claimed “first generative model”) that generates new text in response to model input.” Par. 0022; “A large language model (LLM) is a type of generative language model that is trained in an unsupervised way on massive amounts of unlabeled data, such as publicly available texts extracted from the Internet, using deep learning techniques. A large language model can be configured to perform one or more natural language processing (NLP) tasks, such as generating text (i.e., the claimed “generation steps”), classifying text, answering questions in a conversational manner, and translating text from one language to another.” Par. 0024; Mathur, “According to the techniques described herein, personal data provided by users may be processed to determine prompts when using a generative AI (i.e., generative AI includes the claimed “first generative model, second generative model, etc.) feature at the request of the user,” Par. 0084] based on determining that the first tokens correspond to a non-moderated content category, [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5] processing, using the first generative model for a plurality of generation steps to generate the second tokens. [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5; Rubio, “A generative language model is a particular type of generative model (i.e., the claimed “first generative model”) that generates new text in response to model input.” Par. 0022; “A large language model (LLM) is a type of generative language model that is trained in an unsupervised way on massive amounts of unlabeled data, such as publicly available texts extracted from the Internet, using deep learning techniques. A large language model can be configured to perform one or more natural language processing (NLP) tasks, such as generating text (i.e., the claimed “plurality of generation steps”), classifying text, answering questions in a conversational manner, and translating text from one language to another.” Par. 0024; Mathur, “According to the techniques described herein, personal data provided by users may be processed to determine prompts when using a generative AI (i.e., generative AI includes the claimed “first generative model, second generative model, etc.) feature at the request of the user,” Par. 0084] Regarding Claims 9 and 17, Mathur in view of Rubio has been discussed above. The combination further teaches: wherein the at least one memory includes further instructions that, when executed by the at least one processor, further cause the system to: [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5] determining a first prompt including a first request to determine whether the first tokens correspond to one or more of a set of moderated content categories, [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5] determining, using a second generative model and the first prompt, that the first tokens correspond to the first content category; and [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5; “In some embodiments, one or more models of directive generative thread-based user assistance system includes multiple generative models (i.e., the claimed “first generative model”, “second generative model”, etc.),” Par. 0073] storing embedding data corresponding to the second generative model processing of the first prompt. [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5; “In some implementations, the neural network-based architecture includes one or more input layers that receive model inputs, generate one or more embeddings (i.e., the claimed “embedding data”) based on the model inputs, and pass the one or more embeddings to one or more other layers of the neural network. In other implementations, the one or more embeddings are generated based on the model input by a pre-processor, the embeddings are input to the neural network model, and the neural network model generates output based on the embeddings.” Par. 0069; “In some embodiments, one or more models of directive generative thread-based user assistance system includes multiple generative models (i.e., the claimed “first generative model”, “second generative model”, etc.),” Par. 0073] Regarding Claims 10 and 18, Mathur in view of Rubio has been discussed above. The combination further teaches: wherein the at least one memory includes further instructions that, when executed by the at least one processor, further cause the system to: [Mathur, see mapping applied to claims 1,5,9; Rubio, see mapping applied to claims 1,5,9] determining a second prompt including a second request to determine whether the second tokens correspond to one of the set of moderated content categories; and [Mathur, see mapping applied to claims 1,5,9; Rubio, see mapping applied to claims 1,5,9; Claim is directed to repeating the subject matter for second prompt and second request. However, second prompt/second request/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] processing, using the second generative model, the embedding data and a portion of the second prompt to determine that the second tokens correspond to the first content category. [Mathur, see mapping applied to claims 1,5,9; Rubio, see mapping applied to claims 1,5,9; “In some embodiments, one or more models of directive generative thread-based user assistance system includes multiple generative models (i.e., the claimed “first generative model”, “second generative model”, etc.),” Par. 0073; Claim is directed to repeating the subject matter for second prompt and second request. However, second prompt/second request/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] Regarding Claims 11 and 19, Mathur in view of Rubio has been discussed above. The combination further teaches: wherein the at least one memory includes further instructions that, when executed by the at least one processor, further cause the system to: [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5] processing, using first generative model, to generate third tokens corresponding to a first response including the first tokens and second tokens; [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5; Claim is directed to repeating the subject matter for third tokens. However, third tokens/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] determining that the third tokens correspond to a second content category; and [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5; Claim is directed to repeating the subject matter for third tokens and another category. However, third tokens/another category/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] in response to determining that the third tokens correspond to the second content category, [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5; Claim is directed to repeating the subject matter for third tokens and another category. However, third tokens/another category/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] processing the user input data using the first generative model to generate fourth tokens corresponding to a second response. [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5; Claim is directed to repeating the subject matter for third tokens and another response. However, fourth tokens/another response/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] Regarding Claims 12 and 20, Mathur in view of Rubio has been discussed above. The combination further teaches: wherein the at least one memory includes further instructions that, when executed by the at least one processor, further cause the system to: [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5] processing, using first generative model, to generate third tokens corresponding to a first response including the first tokens and second tokens; [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5; Claim is directed to repeating the subject matter for third tokens. However, third tokens/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] determining that the third tokens correspond to a second content category; [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5; Claim is directed to repeating the subject matter for third tokens and another category. However, third tokens/another category/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] determining instructions to generate an output corresponding to a third content category instead of the second content category; and [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5; Claim is directed to repeating the subject matter for another category. However, another category/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] processing, using the first generative model, the user input data and the instructions to generate a second response to the user input data. [Mathur, see mapping applied to claims 1,5; Rubio, see mapping applied to claims 1,5; Claim is directed to repeating the subject matter for another response. However, another response/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] Claims 4 are rejected under 35 U.S.C. 103(a) as being unpatentable over Mathur in view of Rubio as applied in claim 1 above, and in further view of Torok et al., (U.S. Patent Application Publication 2026/0004786), hereinafter referred to as Torok. Regarding Claim 4, Mathur in view of Rubio has been discussed above. The combination further teaches: generating, based on the first language model processing the first prompt, a third number of tokens corresponding to a third portion of the response; [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Claim is directed to repeating the subject matter for third number of tokens and third portion. However, third number of tokens/third portion/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] processing, using the second language model, the third number of tokens to determine that the third portion of the response corresponds to a first moderated content category; [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Claim is directed to repeating the subject matter for third number of tokens and third portion. However, third number of tokens/third portion/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] determining first data corresponding to the first moderated content category, the first data including instructions to generate an output corresponding to the non-moderated content category instead of the first moderated content category; [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; “For example, if the content 110 is posted from country Z, Policy T does not apply (i.e., the claimed “non-moderated content category”), and a rule-based exclusion can be enforced such that the EBR module 116 will not include policy chunks from Policy T (i.e., the claimed “first moderated content category”) in the retrieved policy chunks 112 (e.g., if country Z, then ignore policy T) (i.e., the claimed “non-moderated content category”).” Par. 0031; Claim is directed to repeating the subject matter. However, repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] in response to the third portion of the response corresponding to the first moderated content category, [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Claim is directed to repeating the subject matter. However, repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] determining a second prompt including the text data, the first data and a second request to generate a response to the natural language user input based on the first data; [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Claim is directed to repeating the subject matter for second prompt and second request. However, second prompt/second request/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] processing, using the first language model, the second prompt to determine a second response to the natural language user input; and [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Claim is directed to repeating the subject matter for second prompt and second response. However, second prompt/second response/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] causing presentation of the second response. [Mathur, see mapping applied to claim 1; Rubio, see mapping applied to claim 1; Claim is directed to repeating the subject matter for second response. However, second response/repeating steps known from prior art is straightforward, amounts to the normal use of the teachings of Mathur in view of Rubio and are rejected under similar rationale.] The combination fails to teach ceasing generation of further tokens by the first language model. However, Torok teaches: ceasing generation of further tokens by the first language model; [Torok, “The prompt, however, causes the LM (i.e., the claimed “first language model”) to stop (i.e., the claimed “ceasing”) generating tokens,” Par. 0044] Mathur, Rubio and Torok pertain efficiency of language models and are analogous to the instant application. Accordingly, it would have been obvious to one of ordinary skill in efficiency of language models art to modify Mathur’s teachings of “machine learning, online platforms, and content moderation, and specifically to dynamic multimodal prompt generation for efficient content moderation” (Mathur, Par. 0001) with the explicit teachings of “categories” (Rubio, Par. 0099) taught by Rubio and “LM (i.e., the claimed “first language model”) to stop (i.e., the claimed “ceasing”) generating tokens” (Torok, Par. 0044) taught by Torok in order prevent “generating content that is nonsensical or unfaithful to the provided source content” (Rubio, Par. 0025) and “efficient and effective adaptation to specific tasks” (Torok, Par. 0170). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Austraat et al., (U.S. Patent 12,657,538) teaches content moderation. Bazari et al., (U.S. Patent Application Publication 2026/0161825) teaches content moderation. Bonacci et al., (U.S. Patent 12,657,536) teaches content moderation. Luo et al., (U.S. Patent 12,627,715) teaches content moderation. Any inquiry concerning this communication or earlier communications from the examiner should be directed to EUNICE LEE whose telephone number is 571-272-1886. The examiner can normally be reached M-F 8:00 AM - 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached on 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /EUNICE LEE/Examiner, Art Unit 2656 /BHAVESH M MEHTA/ Supervisory Patent Examiner, Art Unit 2656
Read full office action

Prosecution Timeline

Nov 27, 2024
Application Filed
Jun 26, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12706111
RECEIVE-SIDE AUDIO PROCESSING FOR CALLS IN A WEB CONFERENCING CLIENT
2y 6m to grant Granted Aug 11, 2026
Patent 12694867
TECHNIQUES FOR IMPROVED AUDIO PROCESSING USING COMBINATIONS OF CLIPPING ENGINES AND ACOUSTIC MODELS
3y 8m to grant Granted Jul 28, 2026
Patent 12694230
MULTI-STAGE MULTI-HOP NATURAL LANGUAGE AND MODEL EXECUTION PLAN GENERATION
2y 10m to grant Granted Jul 28, 2026
Patent 12676160
SOUND SIGNAL PROCESSING METHOD AND ELECTRONIC DEVICE
2y 10m to grant Granted Jul 07, 2026
Patent 12657404
EFFICIENT AND EFFECTIVE SYSTEM AND METHOD TO BUILD MULTI-LINGUAL LARGE LANGUAGE MODELS
2y 5m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
88%
Grant Probability
99%
With Interview (+26.3%)
2y 7m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 43 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month