Prosecution Insights
Last updated: August 08, 2026
Application No. 18/309,088

SYSTEMS AND METHODS FOR TRAINING AND LEVERAGING A MULTI-HEADED MACHINE LEARNING MODEL FOR PREDICTIVE ACTIONS IN A COMPLEX PREDICTION DOMAIN

Final Rejection §103
Filed
Apr 28, 2023
Priority
Feb 01, 2023 — provisional 63/482,612
Examiner
WENG, PEI YONG
Art Unit
2141
Tech Center
2100 — Computer Architecture & Software
Assignee
UnitedHealth Group Incorporated
OA Round
2 (Final)
80%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
513 granted / 645 resolved
+24.5% vs TC avg
Strong +23% interview lift
Without
With
+23.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
26 currently pending
Career history
664
Total Applications
across all art units

Statute-Specific Performance

§101
13.2%
-26.8% vs TC avg
§103
54.6%
+14.6% vs TC avg
§102
21.2%
-18.8% vs TC avg
§112
7.3%
-32.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 645 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION This action is responsive to the following communication: Amendment filed May 21, 2026. This Action is made Final. Claims 1-20 are pending in the case. Claims 1, 12 and 18 are independent claims. Response to Arguments Applicant’s arguments have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 4-8, 12 and 15-18 are rejected under 35 U.S.C. 103 as being unpatentable over Wang (“Cross-Media Keyphrase Prediction: A Unified Framework with Multi-Modality Multi-Head Attention and Image Wordings”, 2020) in view of Jin et al. (hereinafter Jin) U.S. Patent Pub. 2020/0293716 and in further view of Lee et al. (hereinafter Lee) U.S. Patent Pub. 2020/0364407. With respect to independent claim 1, Wang teaches computer-implemented method, the computer-implemented method comprising: receiving, by one or more processors, a text input and contextual data indicative of a predictive category for the text input (see e.g., Section 3 – “we use an open-source toolkit (Smith, 2007) to extract OCR texts in form of a word sequence. It is then appended into the post text with a delimited token hsepi to notify the change of text genres, which is shown to be a simple yet effective design to combine OCR features.”); generating, by the one or more processors and using a multi-headed composite model, an output embedding for the text input based on the predictive category (see e.g., Section 3.2 – “Our design of multi-head attention is inspired by its prototype in Transformer (Vaswani et al., 2017). We extend it to capture multiple forms of crossmodality interactions for a multimedia post, which is therefore named as M3H-Att, short for MultiModality Multi-Head Attention. Compared to its original use as a self-attention over texts only, we instead operate on three modalities (text, attribute, and vision) in a pairwise co-attention manner.”), wherein: the multi-headed composite model comprises a model body, a plurality of model heads, and a gate function, the text input is processed with at least one model head of the plurality of model heads (see e.g., Fig. 3 and Section 3.1-3.3 The examiner notes that a “gate function” is a data flow filter function.), and the gate function is configured to select the at least one model head based on the predictive category for the text input (see e.g., Fig. 2 and Section 3 – “our proposed crossmedia keyphrase prediction model in Figure 2. We first encode a text-image tweet into three modalities: text, attribute, and vision (§3.1), and propose a Multi-Modality Multi-Head Attention (M3H-Att) to capture their intricate interactions (§3.2).”); and providing, by the one or more processors, a predictive label for the text input based on the output embedding (see e.g., Section 3.1 “It will be fed into a keyphrase classifier and generator for the unified prediction. Notably, this indicates that our M3HAtt’s great potential to serve as a generic module for benefiting other cross-media applications.”). Wang does not expressly show a plurality of predictive categories of a predictive domain and a first model head of the plurality of model heads corresponds to the first predictive category, wherein the first predictive category is different from a second predictive category of the plurality of predictive categories that corresponds to a second model head of the plurality of model heads and select category corresponding to the first predictive category of the first model head. However, Wang teaches the plurality of model heads (see e.g., Fig. 2 and Section 3 – “our proposed crossmedia keyphrase prediction model in Figure 2. We first encode a text-image tweet into three modalities: text, attribute, and vision (§3.1), and propose a Multi-Modality Multi-Head Attention (M3H-Att) to capture their intricate interactions (§3.2).”). Furthermore, Jin teaches a plurality of text report categories with category specific evaluation output (see e.g. Abstract Para [4]-[6]-“sorting text report categories, including: obtaining a behavior history and personal information of a user after a text report request of the user is received; inputting the behavior history and the personal information of the user to a personalized model, to obtain a personalized evaluation result of each text report category, where the personalized model is a classification model trained using supervised learning techniques on a plurality of user behavior history samples and personal information samples with each associated with a respective label indicating a category of a text report; and determining, based on the personalized evaluation result, an order of different categories of text reports that can be submitted by the user”) Wang as well as Jin are directed towards machine learning. Therefore, Wang as well as Jin are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Wang with the teachings of Jin by using the model in Wang as the model architecture for the multiple models trained based on different categories in Jin. Jin provides as additional motivation for combination (see e.g., Para [4]-[6] – to provide personalized model based on a specific category). This motivation for combination also applies to the remaining claims which depend on this combination. Wang-Jin does not expressly show the text input is processed with the gate function. However, Lee teaches similar feature (Para [40]-[50] – “a gate mechanism is used so that the small-scale classification task expressed from relatively sufficient data is directly borrowed from the large-scale classification task, and which may be expressed as in Equation 5 below.”). Wang as well as Lee are directed towards machine learning. Therefore, Wang as well as Lee are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Wang with the teachings of Lee and further modify the modified system of Wang. Lee provides as additional motivation for combination (see e.g., Para [40]-[50] – to share features among category classifiers). This motivation for combination also applies to the remaining claims which depend on this combination. With respect to dependent claim 4, the modified Wang teaches the contextual data is indicative of user input that identifies the first predictive category (see e.g., Fig. 2 and Section 1 and 3.1-3.3 – “We extend it to capture diverse cross-media interactions, named as Multi-Modality Multi-Head Attention (M3H-Att). Moreover, to well align the images’ semantics to texts’, we adopt image wordings and define two forms for that — explicit optical characters (such as “NBA Finals” in post (b)) detected from the optical character reader (OCR) and implicit image attributes (Wu et al., 2006), high-level text labels predicted to summarize the image’s semantic concepts (such as a “cat” label for post (a)).”). With respect to dependent claim 5, the modified Wang teaches the multi-headed composite model comprises a neural network (see e.g., Section 2). With respect to dependent claim 6, the modified Wang teaches the model body comprises a first plurality of attention blocks of the neural network and each model head of the plurality of model heads comprises a second plurality of attention blocks of the neural network (see e.g., Fig. 2). With respect to dependent claim 7, the modified Wang teaches each of the plurality of model heads corresponds to a particular predictive category of a plurality of predictive categories of the predictive domain (see e.g., Fig. 2 and section 3 - “first encode a text-image tweet into three modalities: text, attribute, and vision (§3.1), and propose a Multi-Modality Multi-Head Attention (M3H-Att) to capture their intricate interactions (§3.2). Then, we feed the learned multi-modality representations for either keyphrase classification or generation, followed with a tailored aggregator to combine their outputs (§3.3). Lastly, the entire framework can be jointly trained via multi-task learning (§3.4).”). With respect to dependent claim 8, the modified Wang teaches generating the output embedding comprises: generating, using the model body, an intermediate output for the text input (see e.g. Section 3 - “We represent each input as a triplet (x, I, y), where x and y are formulated as word sequences x = hx1, ..., xlx i and y = hy1, ..., yly i (lx and ly denote the number of words)”); and generating, using the first model head, the output embedding based on the intermediate output (see e.g., Section 3.2). Claim 12 is rejected for the similar reasons discussed above with respect to claim 1. Claim 15 is rejected for the similar reasons discussed above with respect to claim 4. Claim 16 is rejected for the similar reasons discussed above with respect to claim 5. Claim 17 is rejected for the similar reasons discussed above with respect to claim 6. Claim 18 is rejected for the similar reasons discussed above with respect to claim 1. Claim 19 is rejected for the similar reasons discussed above with respect to claim 5. Claim 20 is rejected for the similar reasons discussed above with respect to claim 6. Claims 2, 3, 13 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Wang in view of Jin, Lee and further in view of Lambert (“MSeg: A Composite Dataset for Multi-domain Semantic Segmentation”, 2020). With respect to dependent claim 2, Wang does not expressly show the contextual data is indicative of a third-party category for the text input and the first predictive category is based on the third-party category. However, Lambert, in the same field of endeavor, teaches similar feature (Page 1, 3 - "A computer vision professional will likely resort to multiple models, each trained on a different dataset." Each dataset specific domain interpreted as a predictive category corresponding to a respective dataset, each model explicitly trained for the respective domain). Wang as well as Lambert are directed towards machine learning. Therefore, Wang as well as Lambert are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Wang with the teachings of Lambert by using the model in Wang as the model architecture for the multiple models trained on different domains/datasets in Lambert. Lambert provides as additional motivation for combination (Page 1, 3 - "A computer vision professional will likely resort to multiple models, each trained on a different dataset."). This motivation for combination also applies to the remaining claims which depend on this combination. With respect to dependent claim 3, the modified Wang teaches the predictive category is based on a semantic mapping between a plurality of third-party categories and the plurality of predictive categories (see e.g., Page 1 – “We present MSeg, a composite dataset that unifies semantic segmentation datasets from different domains: COCO [10], ADE20K [11], Mapillary [9], IDD [13], BDD [14], Cityscapes [8], and SUN RGB-D [15]. A naive merge of the taxonomies of the seven datasets would yield more than 300 classes, with substantial internal inconsistency in definitions. Instead, we reconcile the taxonomies, merging and splitting classes to arrive at a unified taxonomy with 194 categories.”). Claim 13 is rejected for the similar reasons discussed above with respect to claim 2. Claim 14 is rejected for the similar reasons discussed above with respect to claim 3. Claims 9-11 are rejected under 35 U.S.C. 103 as being unpatentable over Wang in view Jin, Lee and further in view of Lai (“Ontology-based Interpretable Machine Learning for Textual Data”, 2020). With respect to dependent claim 9, Wang does not expressly show the predictive label is one of a plurality of predefined ontology agnostic predictive labels. However, Lai, in the same field of endeavor, teaches similar feature (Page 1, 3) Wang as well as Lambert are directed towards machine learning. Therefore, Wang as well as Lai are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Wang with the teachings of Lai by using the model in Wang as the model architecture for the multiple models trained on different domains/datasets in Lai. Lai provides as additional motivation ontology based machine learning based on textual data (Page 1, 3). This motivation for combination also applies to the remaining claims which depend on this combination. With respect to dependent claim 10, the modified Wang teaches providing the predictive label for the text input based on the output embedding comprises: generating a plurality of label probabilities based on a comparison between the output embedding and a plurality of label embeddings corresponding to the plurality of predefined ontology agnostic predictive labels; and identifying the predictive label based on the plurality of label probabilities (see e.g., Lai P3 - “To learn the local behavior of f in its vicinity (Eq. 1), we approximate L(f, g, φx) by drawing samples based on x, with the proximity indicated by φx. A sample z can be sampled as: z = ∪xi∈x,i6=k,i6=l R(xi) ∪ R({xk, xl}) (3) where R(xi) and R({xk, xl}) are probabilities randomly drawn for each word xi ∈ x(i 6= k, l) and words xk, xl ∈ x together, respectively. If R is greater than a predefined threshold, then the word(s) will be included in z.”). With respect to dependent claim 11, the modified Wang teaches each of the plurality of label probabilities are indicative of a distance between the output embedding and a respective label embedding of the plurality of label embeddings (see e.g., Lai P3 - “To learn the local behavior of f in its vicinity (Eq. 1), we approximate L(f, g, φx) by drawing samples based on x, with the proximity indicated by φx. A sample z can be sampled as: z = ∪xi∈x,i6=k,i6=l R(xi) ∪ R({xk, xl}) (3) where R(xi) and R({xk, xl}) are probabilities randomly drawn for each word xi ∈ x(i 6= k, l) and words xk, xl ∈ x together, respectively. If R is greater than a predefined threshold, then the word(s) will be included in z.”). It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. “The use of patents as references is not limited to what the patentees describe as their own inventions or to the problems with which they are concerned. They are part of the literature of the art, relevant for all they contain.” In re Heck, 699 F.2d 1331, 1332-33, 216 USPQ 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 USPQ 275, 277 (CCPA 1968)). Further, a reference may be relied upon for all that it would have reasonably suggested to one having ordinary skill the art, including nonpreferred embodiments. Merck & Co. v. Biocraft Laboratories, 874 F.2d 804, 10 USPQ2d 1843 (Fed. Cir.), cert. denied, 493 U.S. 975 (1989). See also Upsher-Smith Labs. v. Pamlab, LLC, 412 F.3d 1319, 1323, 75 USPQ2d 1213, 1215 (Fed. Cir. 2005); Celeritas Technologies Ltd. v. Rockwell International Corp., 150 F.3d 1354, 1361, 47 USPQ2d 1516, 1522-23 (Fed. Cir. 1998). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to PEIYONG WENG whose telephone number is (571)270-1660. The examiner can normally be reached on Mon.-Fri. 8 am to 5 pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Matthew Ell, can be reached on (571) 270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://portal.uspto.gov/external/portal. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). /PEI YONG WENG/Primary Examiner, Art Unit 2141
Read full office action

Prosecution Timeline

Apr 28, 2023
Application Filed
Feb 26, 2026
Non-Final Rejection mailed — §103
Apr 23, 2026
Applicant Interview (Telephonic)
Apr 23, 2026
Examiner Interview Summary
May 21, 2026
Response Filed
Jun 10, 2026
Final Rejection mailed — §103
Jul 02, 2026
Examiner Interview Summary
Jul 02, 2026
Applicant Interview (Telephonic)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699878
GENERATING IMPLICIT PLANS FOR ACCOMPLISHING GOALS IN AN ENVIRONMENT USING ATTENTION OPERATIONS OVER PLANNING EMBEDDINGS
4y 0m to grant Granted Aug 04, 2026
Patent 12699923
SYSTEM AND METHOD FOR DISTRIBUTED LEARNING OF UNIVERSAL VECTOR REPRESENTATIONS ON EDGE DEVICES
3y 10m to grant Granted Aug 04, 2026
Patent 12694299
TRAINING A CONVOLUTIONAL NEURAL NETWORK
3y 9m to grant Granted Jul 28, 2026
Patent 12675733
SYSTEMS AND METHODS FOR GENERATING UNIFORM FRAMES HAVING SENSOR AND AGENT DATA
4y 2m to grant Granted Jul 07, 2026
Patent 12670438
MACHINE LEARNING SYSTEM, METHOD, INFERENCE APPARATUS AND COMPUTER-READABLE STORAGE MEDIUM
3y 4m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
80%
Grant Probability
99%
With Interview (+23.1%)
3y 1m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 645 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month