Prosecution Insights
Last updated: October 02, 2026
Application No. 18/876,837

CONTENT RECOMMENDATION BASED ON EMBEDDING SUMMARIZATION

Non-Final OA §101§103
Filed
Dec 19, 2024
Priority
Aug 02, 2022 — CN 202210921856.7 +1 more
Examiner
HOANG, SON T
Art Unit
2169
Tech Center
2100 — Computer Architecture & Software
Assignee
Microsoft Technology Licensing, LLC
OA Round
1 (Non-Final)
84%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
775 granted / 926 resolved
+28.7% vs TC avg
Strong +35% interview lift
Without
With
+34.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
14 currently pending
Career history
939
Total Applications
across all art units

Statute-Specific Performance

§101
16.0%
-24.0% vs TC avg
§103
55.0%
+15.0% vs TC avg
§102
12.0%
-28.0% vs TC avg
§112
6.1%
-33.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 926 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status This instant application No. 18/876,837 has claims 1-15 pending. Priority / Filing Date Applicant’s claim for priority of PCT/US2023/027718 (filed on July 14, 2023) and CHINA 202210921856.7 (filed on August 2, 2022) is acknowledged. The effective filing date of the instant application is August 2, 2022. Abstract The abstract of the disclosure is objected due to the use of implied language. Note that in the abstract, the language should be clear and concise and should not repeat information given in the title. It should avoid using phrases which can be implied, such as, “The disclosure concerns,” “The disclosure defined by this invention,” “The disclosure describes,” etc… See MPEP § 608.01(b). Note that in the abstract, Applicant recites “The present disclosure proposes content recommendation based on embedding summarization…” on lines 1-2. This citation clearly provokes the use of implied language and repeats the title. Revision and/or correction are required (e.g., removal of the entire first sentence of the abstract). Drawings The drawings filed on December 19, 2024 are acceptable for examination purposes. Information Disclosure Statement As required by M.P.E.P. 609(C), the Applicant’s submissions of the Information Disclosure Statements filed on 10 April 2025 and 17 March 2026 are acknowledged by the Examiner and the cited references have been considered in the examination of the claims now pending. As required by M.P.E.P. 609 C(2), a copy of the PTOL-1449 initialed and dated by the Examiner is attached to the instant Office action. Claim Objections Claim 2 recites “the embedding corresponding to the basic input” and it is not clear whether the term refers to “a pooling embedding” previously recited in claim 1. Claim 12 recites multiple instances of “a transformer layer” and it is not clear whether the citation refers to “a transformer layer” previously recited in claim 11. Claim 13 recites “the embedding sequence corresponding to the context input embedding sequence” and this creates an indefinite nested phrase since it is not clear whether these embedding sequences are different or the same. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 14-15 are rejected under 35 U.S.C. 101 because the claimed invention is directed to nonstatutory subject matters. a. Regarding claim 14, an apparatus comprising a processor and a memory are being recited in the claim. However, each of the claimed components can be interpreted by a person of ordinary skills in the art as software modules that carry out the claimed functions (e.g., virtual processor, virtual memory). Furthermore, in accordance with Applicant’s specification ([Page 26, Lines 14-17] of Specification), Applicant states the each component of the apparatus can be implemented by software consisting of data structures and computer programs, which impart functionality when employed as a computer component. As such, the claim is not limited to statutory subject matter and is therefore non-statutory. Applicant is suggested to include at least one hardware component (e.g. hardware processor and/or hardware memory) within the claimed apparatus to overcome the issue raised. One example is as follows: “An apparatus for content recommendation…comprising: a hardware processor; and a hardware memory storing…” b. Claim 15 recites a “computer program product…comprising a computer program…” is being recited. However, data structures not claimed as embodied in non-transitory computer-readable media are descriptive material per se and are not statutory because they are not capable of causing functional change in the computer. See, e.g., Warmerdam, 33 F.3d at 1361, 31 USPQ2d at 1760 (claim to a data structure per se held nonstatutory). Such claimed data structures do not define any structural and functional interrelationships between the data structure and other claimed aspects of the invention which permit the data structure's functionality to be realized. Applicant is noted that a claimed non-transitory computer-readable medium encoded with a data structure defines structural and functional interrelationships between the data structure and the computer software and hardware components which permit the data structure's functionality to be realized, and is thus statutory. One example is as follows: “A computer program product…, comprising a non-transitory computer-readable medium having stored thereon a computer program that is executed by a processor for:..” The claimed invention in claims 1-13 are directed to a judicial exception (i.e., an abstract idea) without significantly more. Applicant is noted that even if claims 14-15 were amended to direct the claims to a hardware apparatus or a non-transitory computer-readable medium, claims 14-15 would still be rejected for the similar reasons presented in claim 1. As such, rejections for claims 14-15 are also presented hereon for consistency purposes. a. Claims 1, 14, and 15 recite in each claim elements that are directed to an abstract idea (“Courts have examined claims that required the use of a computer and still found that the underlying, patent-ineligible invention could be performed via pen and paper or in a person’s mind.” Versata Dev. Group v. SAP Am., Inc., 793 F.3d 1306, 1335, 115 USPQ2d 1681, 1702 (Fed. Cir. 2015)). Each of the claims recites steps/instructions for: Obtaining text inputs (basic and context inputs); Generating embedding sequences; Performing mathematical operations (pooling and summarization); Generating a combined representation; and Predicting a click probability for content recommendation. These limitations fall into judicial exceptions under the following categories: Mathematical Concepts: Generating numerical embeddings, pooling vectors (e.g., mean/max pooling), and summarizing vectors represent mathematical calculations and statistical algorithms. Certain Methods of Organizing Human Activity: Predicting click probability for content recommendation is a fundamental commercial practice and user behavior modeling (targeted content/advertising). Mental Processes / Generic Data Manipulation: Collecting inputs, mathematically aggregating data, and calculating an outcome can be performed theoretically or mentally using pen and paper. Under step 2A – prong 2 of the abstract idea analysis, the claims each is not integrated into a practical application because the claims do not recite an improvement to the functioning of a computer or other technology. Each claim also does not impose meaningful technical limitations on the abstract concepts: Each claim recites the steps/instructions using broad, high-level functional language (e.g., generating, performing a pooling operation, obtaining, predicting) without detailing specific structural implementations, hardware configurations, or algorithmic constraints that overcome a specific technological bottleneck. The end result of predicting a click probability of the candidate content item being clicked merely represents an abstract commercial outcome or data calculation, which amounts to no more than the generic application of the mathematical concept. The addition of generic processor and memory or computer program simply sets the environment to a generic computer and does not integrate the abstract data manipulation into a practical technological improvement. Under step 2B of the abstract idea analysis, each claim fails to recite significantly more than the abstract idea itself. The generic computer components recites perform routine, well-understood, and convention (WURC) computer operations of data retrieval, storage, and execution of mathematical functions. The sequence of steps such as receiving input, converting input to vector/embeddings, pooling vectors, filtering/summarizing vectors, and generating an output score represents WURC data processing steps in natural language processing and information retrieval. Each claim does not recite specific implementations and does not alter the physical or technical operation of the processor or data store. For these reasons, there is no inventive concept in each claims, thus, the claims are ineligible. b. Claims 2, 3, and 7 recite standard transformer embeddings and self-attention. These claims merely recite WURC deep learning and NLP operations, such as standard token/segment/position embeddings (claims 2, 3), generic self-attention mechanism (claims 3, 7) and standard classification predefined encoding/tokens (claim 7). These additions merely describe generic mathematical modeling techniques and do not provide a technological improvement or transform the abstract idea into eligible subject matter. c. Claims 4-6 recite context summarization and/or similarity filtering. These claims specify filtering context embeddings based on similarity ranking against the pooled basic input. While these steps describe mathematical data reduction, they are recited purely at a high level of abstraction (i.e., calculating similarity, ranking, selecting top-ranked embeddings) and amount to generic mathematical optimization and data sorting. In the absence of specific hardware/software integration showing a concrete technical solution to computational constraints in the claim body, these limitations remain part of the abstract mathematical idea and routing data manipulation. d. Claims 8-9 recite query and search history inputs. These claims merely limit the fields of data to a query and historical search results for the query. Gathering or filtering specific types of data constitutes standard extra-solution data gathering and further underscores the abstract method of organizing human activity (e.g., search and commercial content retrieval). e. Claim 10 recites model optimization/training. Reciting that the model is gradually optimized during training describes the WURC mathematical nature of ML algorithms (e.g., gradient descent and/or backpropagation) and does not provide specific technical mechanisms toward eligibility. f. Claims 11-13 recite multi-stage training and transformer layer configurations. Reciting different numbers of transformers across a first-stage and second-stage training pipeline (fewer transformers in the second stage) merely claims higher-level mathematical modeling configurations and algorithmic tuning. The claims lack technical details regarding how the computing system’s physical architecture is specifically enhanced. These steps remain within the realm of abstract algorithm design and conventional training techniques. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-10, and 14-15 are rejected under AIA 35 U.S.C. 103 as being unpatentable over Chen et al. (“End-to-End User Behavior Retrieval in Click-Through Rate Prediction Model”, published on August 10, 2021; hereinafter Chen) in view of Devlin et al. (“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, published on May 24, 2019; hereinafter Devlin). Regarding claims 1, 14, and 15, Chen clearly shows and discloses a method for content recommendation based on embedding summarization (Abstract); an apparatus for content recommendation based on embedding summarization, comprising: a processor; and a memory storing computer-executable instructions that, when executed, cause the processor to implement the method; and a computer program product for content recommendation based on embedding summarization, comprising a computer program that is executed by a processor for implementing the method (Figure 2) comprising: obtaining an input, the input including a basic input and a context input corresponding to the basic input, the basic input including at least a candidate content item (As shown in Figure 2, our model takes user/item-side features as input and outputs the click probability of a certain user-item pair. H𝑙𝑢, H𝑠𝑢, 𝑥𝑢, 𝑥𝑡 and 𝑥𝑐 are raw input features, [Section 4.1]); generating an embedding sequence corresponding to the basic input and an embedding sequence corresponding to the context input (We use 𝒆𝑖 ∈ R𝑑×1 to represent the embedding vectors of item 𝑖. All the embedding vectors of user behavior items are then packed together into a matrix 𝑬𝑠 ∈ R𝐿×𝑑, [Section 4.2]. Figure 2 shows ETA (End-to-end Target Attention) model. 𝒆𝑘+1 ∈ R𝑑×1 and 𝒆𝑡 ∈ R𝑑×1 represent the embedding vectors of behavior item 𝑘+1 and target item 𝑡); generating a pooling embedding corresponding to the basic input through performing a pooling operation on the embedding sequence corresponding to the basic input (Avg-Pooling DNN: The simplest way to utilize user behavior sequence is average pooling which encodes the various length of user sequences into fixed-size hidden vector. This baseline can be regarded as variant of DIN by replacing the target attention with average pooling, which is similar to YouTube [6]. This baseline is mainly used to show the necessity of target attention when compared with DIN. • DIN [36]: DIN is proposed to model personalized user interests with different target items by an attention mechanism, which is called as target attention. However, DIN only utilizes the short-term user behavior sequence, [Section 5.2]. It is noted that sum/mean pooling are conventional in sequence modeling, [Section 4.2]); obtaining a representative embedding sequence corresponding to the context input through performing a summary operation on the embedding sequence corresponding to the context input with the pooling embedding (At the first stage, an auxiliary task is designed to retrieve the top-𝑘 similar items from long-term user behavior sequence, [Abstract]. Our ETA uses SimHash to convert the inner product of two vectors into hamming distance calculation, which is shown in Figure 2. 𝑬′𝑠 ∈ R𝑘×𝑑 consists of top 𝑘 rows selected from 𝑬𝑠1 which have the largest hamming distances with target item 𝑬𝑡 ∈ R1×𝑑, [Section 4.4]); generating an input representation of the text input based at least on the embedding sequence corresponding to the basic input and the representative embedding sequence corresponding to the context input (Then the hidden vectors are concatenated together and are fed into the MLP (Multi-layer Perception) part, [Section 4.1]. After the top-k closest keys to query are selected, the normal attention and back propagation are conducted on the original embedding vectors of these top-k items, [Section 4.5.1]); and predicting a click probability of the candidate content item being clicked based on the input representation (At the last layer of MLP, sigmoid function is used to map the hidden vector into a scalar 𝑝(𝑦𝑗 |H𝑙𝑢,H𝑠𝑢, 𝑥𝑢, 𝑥𝑡 , 𝑥𝑐 ; 𝜃) which represents the click probability of a certain user-item pair. This probability can be used as the ranking score for the downstream tasks, [Section 4.1]). Devlin then additionally or alternatively discloses: the input is text input (To make BERT handle a variety of down-stream tasks, our input representation is able to unambiguously represent both a single sentence and a pair of sentences (e.g., h Question, Answer i) in one token sequence. Throughout this work, a “sentence” can be an arbitrary span of contiguous text, rather than an actual linguistic sentence. A “sequence” refers to the input token sequence to BERT, which may be a single sentence or two sentences packed together, [Section 3]); generating an embedding sequence corresponding to the basic input and an embedding sequence corresponding to the context input (For a given token, its input representation is constructed by summing the corresponding token, segment, and position embeddings. A visualization of this construction can be seen in Figure 2, [Section 3]); generating a pooling embedding corresponding to the basic input through performing a pooling operation on the embedding sequence corresponding to the basic input (The first token of every sequence is always a special classification token ([CLS]). The final hidden state corresponding to this token is used as the aggregate sequence representation for classification tasks, [Section 3]); generating a text input representation of the text input based at least on the embedding sequence and the representative embedding sequence (BERT instead uses the self-attention mechanism to unify these two stages, as encoding a concatenated text pair with self-attention effectively includes bidirectional cross attention between two sentences, [Section 3.2]). It would have been obvious to an ordinary person skilled in the art at the time of the invention was effectively filed to incorporate the teachings of Devlin with the teachings of Chen for the purpose of improving semantic representation and search relevance to evaluate text-rich candidate items semantically while preserving low-latency execution benefits. Regarding claim 2, Devlin further discloses the generating an embedding sequence corresponding to the basic input and an embedding sequence corresponding to the context input comprises: obtaining a basic input token sequence corresponding to the basic input and a context input token sequence corresponding to the context input (To make BERT handle a variety of down-stream tasks, our input representation is able to unambiguously represent both a single sentence and a pair of sentences (e.g., h Question, Answer i) in one token sequence. Throughout this work, a “sentence” can be an arbitrary span of contiguous text, rather than an actual linguistic sentence. A “sequence” refers to the input token sequence to BERT, which may be a single sentence or two sentences packed together, [Section 3]); and generating the embedding corresponding to the basic input and the embedding sequence corresponding to the context input based on at least one of a token embedding sequence, a segment embedding sequence, and a position embedding sequence corresponding to the basic input token sequence and the context input token sequence (For a given token, its input representation is constructed by summing the corresponding token, segment, and position embeddings. A visualization of this construction can be seen in Figure 2, [Section 3]). Regarding claim 3, Devlin further discloses the generating an embedding sequence corresponding to the basic input and an embedding sequence corresponding to the context input comprises: obtaining a basic input token sequence corresponding to the basic input and a context input token sequence corresponding to the context input (To make BERT handle a variety of down-stream tasks, our input representation is able to unambiguously represent both a single sentence and a pair of sentences (e.g., h Question, Answer i) in one token sequence. Throughout this work, a “sentence” can be an arbitrary span of contiguous text, rather than an actual linguistic sentence. A “sequence” refers to the input token sequence to BERT, which may be a single sentence or two sentences packed together, [Section 3]); generating an initial basic input embedding sequence and an initial context input embedding sequence based on at least one of a token embedding sequence, a segment embedding sequence, and a position embedding sequence corresponding to the basic input token sequence and the context input token sequence (For a given token, its input representation is constructed by summing the corresponding token, segment, and position embeddings. A visualization of this construction can be seen in Figure 2, [Section 3]); and generating the embedding sequence corresponding to the basic input and the embedding sequence corresponding to the context input based on the initial basic input embedding sequence and the initial context input embedding sequence through a self-attention mechanism (BERT instead uses the self-attention mechanism to unify these two stages, as encoding a concatenated text pair with self-attention effectively includes bidirectional cross attention between two sentences, [Section 3.2]). Regarding claim 4, Chen further discloses the number of embeddings in the representative embedding sequence corresponding to the context input is less than the number of embeddings in the embedding sequence corresponding to the context input (Thus we can first retrieval top-𝑘 items from the behavior sequence and conduct the multi-head target attention on these 𝑘 behaviors. 𝑬′𝑠 ∈ R𝑘×𝑑 consists of top 𝑘 rows selected from 𝑬𝑠1 which have the largest hamming distances with target item 𝑬𝑡 ∈ R1×𝑑, [Section 4.4]. It is clear that the long sequence length L is reduced down to a smaller subset k satisfying k < L). Regarding claim 5, Chen further discloses the representative embedding sequence corresponding to the context input includes embeddings relevant to the basic input that are selected from the embedding sequence corresponding to the context input (From the perspective of the CTR model, the retrieval part is transparent but can ensure the model use the most closet items to conduct multi-head attention, [Section 4.5.1]. It is clear that the selected items are relevant or closes to the target candidate item). Regarding claim 6, Chen further discloses the embedding sequence corresponding to the context input comprises a plurality of context input embeddings, and the performing a summary operation comprises: calculating similarity between the pooling embedding and each context input embedding in the plurality of context input embeddings, to obtain a plurality of similarities (As shown in Figure 2, we use SimHash function and hamming distance to calculate the similarity of two embedding vectors instead of inner product, [Section 4.4]); ranking the plurality of context input embeddings based on the plurality of similarities, to obtain a plurality of ranked context input embeddings (Thus, the similarity between the embedding vectors can be replaced by the similarity between the hashed fingerprints. A 𝑑-dimensional embedding vector can be encoded into 𝑚-bit number. Then the similarity between two fingerprints can be measured by hamming distance, [Section 4.4]); selecting a top-ranked predetermined number of context input embeddings from the plurality of ranked context input embeddings (The top-𝑘 retrieval layer can find the top-k similar user behavior items with target item more efficiently by hamming distance compared with the inner-product based models. The hamming distance for two integers is defined as the number of different bits positions at which the corresponding bits are different, [Section 4.4]); and combining the selected context input embeddings into the representative embedding sequence corresponding to the context input (In the first stage, an auxiliary task is designed to retrieve the top-𝑘 similar items from long-term user behavior sequence, such that the top-𝑘 similar items are prepared in advance. In the second stage, the target attention mechanism is conducted between target item and 𝑘 items selected in the first stage, [Section 1]). Regarding claim 7, Devlin further discloses the generating a text input representation comprises: generating an embedding corresponding to a classification predefined encoding (The first token of every sequence is always a special classification token ([CLS]). The final hidden state corresponding to this token is used as the aggregate sequence representation for classification tasks, [Section 3]); generating a classification embedding, a basic input embedding sequence, and a representative context input embedding sequence based on the embedding corresponding to the classification predefined encoding, the embedding sequence corresponding to the basic input, and the representative embedding sequence corresponding to the context input through a self-attention mechanism (Sentence pairs are packed together into a single sequence. We differentiate the sentences in two ways. First, we separate them with a special token ([SEP]). Second, we add a learned embedding to every token indicating whether it belongs to sentence A or sentence B, [Section 3]. Fine-tuning is straightforward since the self-attention mechanism in the Transformer allows BERT to model many downstream tasks—whether they involve single text or text pairs—by swapping out the appropriate inputs and outputs, [Section 3.2]); and taking the classification embedding as the text input representation (To fine-tune on GLUE, we represent the input sequence (for single sentence or sentence pairs) as described in Section 3, and use the final hidden vector corresponding to the first input token ([CLS]) as the aggregate representation, [Section 4.1]). Regarding claim 8, Chen further discloses the basic input further includes a query (UBR4CTR is also a two-stage method which utilizes the long-term user behavior sequence in CTR prediction task. In UBR4CTR, a query is prepared by a feature selection model to retrieve the most similar behavior items, [Section 5.1]). Regarding claim 9, Chen then discloses the context input includes historical search results for the query (Industrial dataset: This dataset is collected from our own online RS, which is one of the top-tier mobile Apps in our country. There are three advantages for our industrial dataset. (i) Our dataset contains impression interaction, which indicates an item is displayed to an user but not clicked by the user. Impression interaction is naturally the negative sample of CTR model, [Section 5.1]). Regarding claim 10, Chen further discloses the click probability is predicted through a click probability predicting model, and the representative embedding sequence corresponding to the context input is gradually optimized during training of the click probability predicting model (As long as the input embedding vector 𝒆𝑘 ∈ R1×𝑑 is updated, the signature of SimHash is updated correspondingly. The Locality-sensitive properties ensure that the top-k nearest keys to query in each iteration are selected using the latest embedding of CTR model seamlessly. Thus the gap of goal between retrieval and CTR model is much smaller than those other retrieval methods, e.g., offline inverted index based method shown in Table 2, [Section 4.5.1]). Claims 11-13 are rejected under AIA 35 U.S.C. 103 as being unpatentable over Chen in view of Devlin and further in view of Wang et al. (Pub. No. US 2022/0198276, filed on December 20, 2021; hereinafter Wang). Regarding claim 11, Wang then discloses the click probability is predicted through a click probability predicting model, training of the click probability predicting model includes a first-stage training and a second-stage training (the whole process of a method for a pre-trained language model automatic compression based on multilevel knowledge distillation of the present invention is divided into three stages, the first stage is to construct multilevel knowledge distillation, and to distill a knowledge structure of a large model on three different levels: a self-attention unit, a hidden layer state and an embedded layer; and the second stage is to train a knowledge distillation network of meta-learning to generate a general compression architecture of a plurality of pre-trained language models, [0047]), and the numbers of transformers included in a transformer layer in the click probability predicting model are different in the first-stage training and in the second-stage training (Transformer layer distillation includes knowledge distillation based on self-attention and knowledge distillation based on hidden layer state, as shown in FIG. 2. The distillation based on self-attention can focus on rich language knowledge, [0049]. The specific task fine tuning module constructs a downstream task network on the pre-trained model distillation network generated by the automatic compression component, performs fine tuning on a downstream task scenario by using a feature layer and an output layer of the distillation network, and outputs a final fine-tuned student model, that is, a compressed model of the pre-trained language model, which includes downstream tasks and is required by the logged-in user, [0092]. It is clear that the large transformer model is sampled via a layer sampling vector to create a student model containing different number of transformer modules, which is subsequently fine-tuned on downstream tasks). It would have been obvious to an ordinary person skilled in the art at the time of the invention was effectively filed to incorporate the teachings of Wang with the teachings of Chen, as modified by Devlin, for the purpose of constructing multilevel knowledge distillation and distilling a knowledge structure of a large model at multiple different levels to enhance fine tuning on a downstream task. Regarding claim 12, Wang further discloses a transformer layer in the click probability predicting model in the first-stage training includes a first number of transformers, a transformer layer in the click probability predicting model in the second-stage training includes a second number of transformers, and the second number is less than the first number (It is worth noting that, BERTbase has a total of 12 Transformer modules. In order to prevent the number (referring to the number of elements being 1 in the layer sampling vector) of transferred Transformer modules during layer sampling from being too small, it is proposed to increase layer sampling constraint conditions, that is, every time a distillation network structure is generated, in the layer sampling stage of all the Transformer layers of the BERT, the constraint conditions are constructed so that the number of elements being 1 in the vector obtained from the final layer sampling is not less than 6, [0065], [0100]-[0101]). Regarding claim 13, Wang further discloses in the first-stage training, the text input representation is generated based at least on the embedding sequence corresponding to the basic input and the embedding sequence corresponding to the context input embedding sequence (self-attention distribution knowledge, hidden state knowledge and embedded layer knowledge are encoded as a distillation network, as shown in FIG. 2. A large model is compressed into a small model by using knowledge distillation, and large-scale self-attention knowledge of the large model is transferred to the small model to the greatest extent, [0048]-[0049]. It is clear that the full, uncompressed teacher model evaluates all input tokens across the full self-attention matrix prior to layer distillation. This indicates representation generation from full input sequences in the initial stage). Relevant Prior Art The following references are deemed relevant to the claims: Shou et al. (Pub. No. US 2025/0165544) teaches a historical content item sequence of a user may be obtained. A topic and a text of each historical content item in the historical content item sequence may be identified, to obtain a topic sequence and a text sequence corresponding to the historical content item sequence. A comprehensive topic representation may be generated based on the topic sequence. A comprehensive text representation may be generated based on the text sequence. A user interest representation of the user may be generated based on the comprehensive topic representation and the comprehensive text representation. Stoffel et al. (Pub. No. US 2021/0350247) teaches generation and execution of hybrid decision trees for machine learning systems and processing environments. The hybrid decision tree including a plurality of nodes, each node comprising at least one directional associative link to or from another node, each node associated with a function performed on an input variable, wherein each directional associative link is agnostic to a hierarchical layer of the associated nodes in the hybrid decision tree; and means for executing functions associated with each node of the hybrid decision tree according to the corresponding directional associative links. Contact Information Any inquiry concerning this communication or earlier communications from the Examiner should be directed to Son Hoang whose telephone number is (571) 270-1752. The Examiner can normally be reached on Monday – Friday (7:00 AM – 4:00 PM). If attempts to reach the Examiner by telephone are unsuccessful, the Examiner’s supervisor, Sherief Badawi can be reached on (571) 272-9782. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SON T HOANG/Primary Examiner, Art Unit 2169 August 22, 2026
Read full office action

Prosecution Timeline

Dec 19, 2024
Application Filed
Aug 26, 2026
Non-Final Rejection mailed — §101, §103
Sep 08, 2026
Applicant Interview (Telephonic)
Sep 08, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737422
INTELLIGENT CLUSTERING SYSTEMS AND METHODS USEFUL FOR DOMAIN PROTECTION
2y 3m to grant Granted Sep 15, 2026
Patent 12730815
SYSTEMS AND METHODS FOR ANALYZING DISTRIBUTED SYSTEM DATA STREAMS USING DECLARATIVE SPECIFICATION, DETECTION, AND EVALUATION OF HAPPENED-BEFORE RELATIONSHIPS
1y 6m to grant Granted Sep 08, 2026
Patent 12717824
ARTIFICIAL INTELLIGENCE APPARATUS AND CHEMICAL MATERIAL SEARCH METHOD THEREOF
1y 7m to grant Granted Aug 25, 2026
Patent 12688256
CENTRALIZED REPOSITORY AND DATA SHARING HUB FOR ESTABLISHING MODEL SUFFICIENCY
4y 2m to grant Granted Jul 21, 2026
Patent 12688229
MEDIA FILE RECOMMENDATIONS FOR A SEARCH ENGINE
1y 6m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
84%
Grant Probability
99%
With Interview (+34.6%)
2y 11m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 926 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month