Prosecution Insights
Last updated: August 18, 2026
Application No. 17/822,029

MACHINE LEARNING CONTEXT BASED CONFIDENCE CALIBRATION

Non-Final OA §103
Filed
Aug 24, 2022
Examiner
ROSTAMI, MOHAMMAD S
Art Unit
2154
Tech Center
2100 — Computer Architecture & Software
Assignee
Adobe Inc.
OA Round
3 (Non-Final)
67%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 67% — above average
67%
Career Allowance Rate
431 granted / 641 resolved
+12.2% vs TC avg
Strong +26% interview lift
Without
With
+26.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 9m
Avg Prosecution
26 currently pending
Career history
688
Total Applications
across all art units

Statute-Specific Performance

§101
20.1%
-19.9% vs TC avg
§103
56.8%
+16.8% vs TC avg
§102
9.8%
-30.2% vs TC avg
§112
4.6%
-35.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 641 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of Claims Claims 1-20 are pending of which claims 1, 9 and 16 are in independent form. Claims 1-20 are rejected under 35 U.S.C. 103. Response to Arguments Applicant’s arguments with respect to claim(s) 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Regarding 35 USC 101 (Abstract Idea): Applicant’s arguments, see “Remarks”, filed on 2/19/2026, with respect to 35 USC 101 have been fully considered and are persuasive. The 35 USC 101 rejection of claim 1-20 has been withdrawn. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-4, 6, 8-11, 13, 15, 16, 18 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Li; Yuguang et al. (US 20230032888 A1) [Li] in view of KURMA; Sai Sree Bhargav et al. (US 20220019734 A1) [Kurma] in view of Joshi; Siddharth Vivek et al. (US 11928558 B1) [Joshi]. Regarding claims 1, 9 and 16, Li discloses, a system comprising: a memory component; and one or more processing devices coupled to the memory component (see Fig. 3), the one or more processing devices to perform operations comprising: obtaining an image frame (analyzing the visual data of the target image ¶ [0052], obtaining the target image ¶ [0101], [0023], [0028], [0030]); generating, with a first machine learning model, a confidence score, a bounding box (scoring layer that produces a confidence score … trained neural network is used that takes as input two lists of object embedding vectors associated with determined 3D and/or 2D bounding boxes (one list for each image) and the associated directional vector orientation for each bounding box in their respective image's local camera coordinate system ¶ [0064]. Also see ¶ [0052],[0061]-[0062], [0069], [0074], [0101]), and a tensor representation of an instance embedding (generating object embedding vectors ¶ [0061], object embedding vectors ¶ [0062], [0064], generating an object embedding vector for each such object ¶ [0101]; examiner specifies that NN generated object embedding vectors constitute tensor representations of object-instance embeddings), corresponding to an object instance inferred from the image frame (wall objects ¶ [0061], matching objects ¶ [0062], object image descriptor ¶ [0069], generating an object embedding vector for each such object ¶ [0101], also see ¶ [0016], [0074], [0102]); inputting the confidence score, the bounding box, and the tensor representation [into a second machine learning model] (measure the differences between object embedding vectors ¶ [0062], object embedding vector comparison…scoring layer ¶ [0064], object embedding vector ¶ [0101]). However, Li does not explicitly facilitate into a second machine learning model. Kurma discloses, into a second machine learning model (using the first deep learning model…bounding box coordinate…confidence score …using a second deep learning model, contextual embedding ¶ [0041]-[0043], provided as input to the contextual language model reasoner ¶ [0038], [0040], contextual language model reasoner can be fed into a feed forward network ¶ [0061]; examiner specifies that these passages expressly teach a downstream second ML model, receiving and processing outputs received by an earlier ML model, including bounding box and confidence related outputs). It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the cited references because Kurma’s system would have allowed Li to facilitate into a second machine learning model. The motivation to combine is apparent in the Li’s reference, because there is a need to improve artificial intelligence, and more particularly, to a method and system for visio-linguistic understanding using contextual language model reasoners by improving computational power, memory and time. However, neither Li nor Kurma does not explicitly facilitate computing, with the second machine learning model, a calibrated confidence score for the object instance based on the tensor representation, the confidence score, and the bounding box; providing the calibrated confidence score to an application, wherein the application disregards the object instance responsive to the calibrated confidence score being less than a confidence threshold. Joshi discloses, computing, with the second machine learning model, a calibrated confidence score for the object instance based on the tensor representation, the confidence score, and the bounding box (the first ML model may have a first confidence score associated with fields of interest that correspond to the request of the user… determine whether the first confidence is trustworthy [col. 31, ll. 55-67], second confidence associated with the accuracy of the first ML model… comparing the first confidence with a second confidence that is trained via a calibration set [col. 32, ll. 23-43], the ML models may be trained to increase the confidence and accuracy of the ML models [col. 2, ll. 34-61], stateful calibrated adaptive threshold technique [col. 33, ll. 2-16], also see [col. 6, ll. 5-24], [col. 10, ll. 10-28], [col. 12, ll. 56-col. 13, ll. 5]); providing the calibrated confidence score to an application (ML models determine that the confidence score of a prediction is less than a defined confidence (e.g., threshold) [col. 2, ll. 34-61], confidence score [col. 4, ll. 30-56], [col. 5, ll. 56-col. 6, ll. 24]), wherein the application disregards the object instance responsive to the calibrated confidence score being less than a confidence threshold (ML models determine that the confidence score of a prediction is less than a defined confidence (e.g., threshold) [col. 2, ll. 34-61], the first ML model may have a first confidence score associated with fields of interest that correspond to the request of the user… determine whether the first confidence is trustworthy [col. 31, ll. 55-67], determining whether to trust the first confidence of the first ML model [col. 32, ll. 23-43], calibrated adaptive threshold technique [col. 33, ll. 2-16]). It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the cited references because Joshi’s system would have allowed Li and Kurma to facilitate computing, with the second machine learning model, a calibrated confidence score for the object instance based on the tensor representation, the confidence score, and the bounding box; providing the calibrated confidence score to an application, wherein the application disregards the object instance responsive to the calibrated confidence score being less than a confidence threshold. The motivation to combine is apparent in the Li and Kurma’s reference, because there is a need to improve ML models universalness or scalability to accept various conditional inputs when analyzing content and outputting predictions. Regarding claims 2 and 10, the combination of Li, Kurma and Joshi discloses, wherein the first machine learning model and the second machine learning model are executed via a neutral network (Kurma: Both contextual visio-linguistic reasoner and contextual language model reasoners are neural networks and neural networks do not perform well if test data varies much from the data it is trained on. This makes a contextual language model reasoner better suited as a general-purpose model ¶ [0025]. Also see ¶ [0006]-[0007]). Regarding claim 3, the combination of Li, Kurma and Joshi discloses, wherein the image frame comprises an image of at least one of text, a graphic, a video image frame, and a photograph (Kurma: receiving an image and a text corresponding to the image, wherein the image comprises one or more embedded texts and wherein the image and the text correspond to a downstream task ¶ [0006]-[0007]. Also see Fig. 8). Regarding claims 4, 11, 18 and 19, the combination of Li, Kurma and Joshi discloses, wherein the first machine learning model is trained to generate the object instance based on a first training set comprising image frame samples (Joshi: At 302, the process 300 may analyze a dataset using a ML model to train the ML model to recognize one or more field(s) of interest or item(s) within content. For example, the dataset may include various forms of content, such as documents, PDFs, images, videos, and so forth that are searchable by the ML model. The ML model may be instructed to analyze the dataset or to be trained on the dataset, or content within the dataset, for use in recognizing or searching for item(s) within content at later instances. In some instances, human reviewers may label or classify samples within the dataset (e.g., a calibration set) and the ML model may accept these as input these as inputs for training the ML model. For example, the ML model may be trained to identify certain objects within the content, such as dogs or cats. That is, utilizing the dataset and/or the labels provided by human reviews, the ML models may be trained to recognize or identify dogs or cats with presented content [col. 20, ll. 58-col. 21, ll. 8]. Also see [co. 3, ll. 6-59], [col. 4, ll. 57-col. 5, ll. 18]); and wherein the second machine learning model is trained separately from the first machine learning model using a second training set after training of the first machine learning model with the first training set is completed (Joshi: In some instances, more than one ML model(s) 126 may be utilized when carrying out requests. For example, a first ML model may identify objects within an image and a second ML model may label the objects. In some instances, each of the ML model(s) 126 may be previously trained from a specific subset of the content data 124 and/or a calibration set within the content data 124 [col. 10, ll. 10-28]. For example, as illustrated, at 810 the process 800 may determine a calibration set for the second ML model. The calibration set used to train the second ML model may include random samplings of content or content that has been identified with high confidences. In other instances, the calibration set may include content labeled by human reviewers. The calibration set may therefore be utilized to train the second ML model to identify, search, or review particular field(s) of interest or content [col. 32, ll. 14-43]. Also see [col. 2, ll. 34-col. 3, ll. 59]). Regarding claims 6 and 13, the combination of Li, Kurma and Joshi discloses, the operations further comprising: responsive to determining that a difference between the confidence score and the calibrated confidence score exceeds a first threshold, searching a set of training image samples for similar object instances based on the instance embedding, wherein the first machine learning model was trained using the set of training image samples (Joshi: If the ML models determine that the confidence score of a prediction is less than a defined confidence (e.g., threshold), the content (or a portion thereof) may be sent for human review. Alternatively, if the ML model(s) determine that the confidence score is greater than the defined confidence threshold, the content may not be sent for human review. Users may therefore define the conditions when predictions or results of the ML model(s) are sent for human review. Based on the review of the ML model(s), the ML models may be trained to increase the confidence and accuracy of the ML models [col. 2, ll. 51-61]. However, the results of the human review may be utilized to train the ML models to increase their associated accuracy. For example, if the ML model(s) are accurate, the confidence threshold for screening the results of the predicted outputs may be reduced as the outputs of the ML model(s) are accurate [col. 4, ll. 51-56]. Also see [col. 5, ll. 56-col. 6, ll. 4], [col. 28, ll. 52-col. 28, ll. 2] and [col. 30, ll. 4-15]). Regarding claims 8 and 15, the combination of Li, Kurma and Joshi discloses, generating, with the first machine learning model, a set of object instances from a first set of image samples, wherein for each object instance of the set of object instances, the first machine learning model computes a respective confidence score and a respective instance embedding; computing, with the second machine learning model, a respective calibrated confidence score for each object instance of the set of object instances; generating a second set of image samples based on one or more object instances from the first set of image samples for which a difference between the respective confidence score and the respective calibrated confidence score exceeds a threshold; and clustering object instances from the second set of image samples based on the respective instance embedding for each of the one or more object instances (Joshi: For example, the predictions may include text classification or labeling (e.g., assigning tags, categorizing text, mining text, etc.), image classification (e.g., categorizing images into classes), object detection (e.g., locating objects in images via bounding boxes), or semantic segmentation (e.g., locating objects in images with pixel-level precision) associated with the content. In some instances, when generating predictions or analyzing the content, the ML models may utilize conditions or user-defined criteria. For example, users may define confidence scores that are associated with the predicted outputs. If the ML models determine that the confidence score of a prediction is less than a defined confidence (e.g., threshold), the content (or a portion thereof) may be sent for human review. Alternatively, if the ML model(s) determine that the confidence score is greater than the defined confidence threshold, the content may not be sent for human review. Users may therefore define the conditions when predictions or results of the ML model(s) are sent for human review. Based on the review of the ML model(s), the ML models may be trained to increase the confidence and accuracy of the ML models [col. 2, ll. 34-61]. Also see [col. 4, ll. 57-col. 5, ll. 18]). Claim(s) 5, 12, 17 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Kurma in view of Joshi in view of Lin; Zhe et al. (US 20200151448 A1) [Lin]. Regarding claims 5, 12 and 20, the combination of Li, Kurma and Joshi teaches all the limitations of claims 4, 11 and 16. However, neither one of Li, Kurma or Joshi explicitly facilitate wherein the second machine learning model is trained by: computing a binary classification score for the bounding box responsive to determining that the bounding box corresponds to an annotated ground truth bounding box; and adjusting the second machine learning model based on a difference between the calibrated confidence score and the binary classification score. Lin discloses, wherein the second machine learning model is trained by: computing a binary classification score for the bounding box responsive to determining that the bounding box corresponds to an annotated ground truth bounding box; and adjusting the second machine learning model based on a difference between the calibrated confidence score and the binary classification score (In one example, a conditional detection network includes a binary classifier that assigns a positive training label to detection outputs of the conditional detection network that substantially overlap with a ground truth bounding box for the word-based concept. The binary classifier assigns a negative training label to other detection outputs of the conditional detection network that are not assigned a positive training label ¶ [0026]. Training module 154 can train any suitable network according to any suitable loss function. To keep a conditional detection network label agnostic, a binary loss can be used. For instance, in conditional detection network 300 of FIG. 3, CNN 304 includes binary classifier 322. In binary classifier 322, a binary sigmoid cross-entropy loss is used, rather than a softmax cross-entropy loss of Faster R-CNN. A binary classifier, such as binary classifier 322, can assign a positive label (e.g., a positive training label) to detection outputs of the conditional detection network that substantially overlap with a ground truth bounding box that corresponds to a given word-based concept, and a negative label (e.g., a negative training label) to other detection outputs that are not assigned a positive label. Additionally or alternatively, a binary classifier can be used to train a conditional detection network for negative classes of inputs. For instance, a negative class for a word-based concept can be provided to a conditional detection network, and a negative label can be assigned by a binary classifier to detection outputs that substantially overlap with a ground truth bounding box corresponding to the word-based concept. Training with positive and negative training labels, and with negative classes of inputs is illustrated in FIG. 4 ¶ [0086]. Also see ¶ [0044]). It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the cited references because Lin’s system would have allowed Li, Kurma and Joshi to facilitate wherein the second machine learning model is trained by: computing a binary classification score for the bounding box responsive to determining that the bounding box corresponds to an annotated ground truth bounding box; and adjusting the second machine learning model based on a difference between the calibrated confidence score and the binary classification score. The motivation to combine is apparent in the Li, Kurma and Joshi’s reference, because there is a need to improve object detectors trained using heterogeneous training datasets. Regarding claim 17, the combination of Li, Kurma, Joshi and Lin discloses, computing a training correction using ground truth images used by the another machine learning model to generate the instance embedding, the confidence score, and the bounding box for each of the one or more object instances (Lin: In one example, a conditional detection network includes a binary classifier that assigns a positive training label to detection outputs of the conditional detection network that substantially overlap with a ground truth bounding box for the word-based concept. The binary classifier assigns a negative training label to other detection outputs of the conditional detection network that are not assigned a positive training label ¶ [0026]-[0027]. Also see ¶ [0086], [0091]-[0094]). Claim(s) 7 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Kurma in view of Joshi in view of Mehra; Ashutosh et al. (US 20210133439 A1) [Mehra]. Regarding claims 7 and 14, the combination of Li, Kurma and Joshi teaches discloses, determining, using the first machine learning model, a respective confidence score for each of the similar object instances from the set of similar image samples; determining, using the second machine learning model, a respective calibrated confidence score for each of the similar object instances from the set of similar image samples (Joshi: In some instances, the ML models may be retrained or calibrated from a calibration set of data within the dataset. In some instances, the calibration set may include predicted outputs from the ML models as well as outputs provided by the reviewers. The calibration set may, in some instances, represent new content recently added to the dataset as well as old content within the dataset. For example, old content within the dataset may be periodically removed from the calibration set based on various expiration and/or sampling strategies. In some instances, content within the dataset may be randomly sampled for inclusion within the calibration set. Additionally, or alternatively, a percent or sampling of newly added content to the dataset may be randomly chosen for inclusion within the calibration set. Through the calibration set, the confidence thresholds of the ML models may be re-computed by iterating the data within the dataset and then comparing the predicted outputs with human review. The desired confidence thresholds may be influenced, in some instances, by accuracy, precision, and/or other recall configurations [col. 6, ll. 5-24]. Also see [col. 10, ll. 10-28], [col. 12, ll. 56-col. 13, ll. 5] and [col. 32, ll. 1-43], [col. 32, ll. 57-col. 33, ll. 2]) However, neither one of Li, Kurma or Joshi explicitly facilitate responsive to determining that, for a first similar object instance, a difference between the respective confidence score and the respective calibrated confidence score exceeds a second threshold, generating an indication of a potential training data annotation error. Mehra discloses, responsive to determining that, for a first similar object instance, a difference between the respective confidence score and the respective calibrated confidence score exceeds a second threshold, generating an indication of a potential training data annotation error (Training or tuning of the CNN or any machine learning model can include minimizing a loss function between the target variable or output (e.g., 0.90) and the expected output (e.g., 100%). Accordingly, it may be desirable to arrive as close to 100% confidence of a particular classification as possible so as to reduce the prediction error. This may happen overtime as more training images/documents and baseline data sets are fed into the learning models so that classification/detection can occur with higher prediction probabilities. Accordingly, in some embodiments, block 1008 represents tuning or training, which is done in various stages (e.g., a first stage and a second stage) to reduce prediction error. In these embodiments for example, a first training set can be created (e.g., a first document with content order values) and training can occur in a first stage using the first training set and then a second training set can be created (e.g., a first document with other content order values) and training can occur in a second stage using the second training set to reduce error rate or tune the model. In other embodiments, the prediction at block 1008 represents prediction on a deployed model that has already been trained ¶ [0086]. Also see ¶ [0068]). It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the cited references because Mehra’s system would have allowed Li, Kurma and Joshi to facilitate responsive to determining that, for a first similar object instance, a difference between the respective confidence score and the respective calibrated confidence score exceeds a second threshold, generating an indication of a potential training data annotation error. The motivation to combine is apparent in the Li, Kurma and Joshi’s reference, because there is a need to improve the detection of objects within documents using machine learning models. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MOHAMMAD S ROSTAMI whose telephone number is (571)270-1980. The examiner can normally be reached Mon-Fri From 9 a.m. to 5 p.m.. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Boris Gorney can be reached at (571)270-5626. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. 5/21/2026 /MOHAMMAD S ROSTAMI/Primary Examiner, Art Unit 2154
Read full office action

Prosecution Timeline

Show 5 earlier events
Feb 19, 2026
Response Filed
May 27, 2026
Final Rejection mailed — §103
Jul 01, 2026
Interview Requested
Jul 10, 2026
Applicant Interview (Telephonic)
Jul 24, 2026
Request for Continued Examination
Jul 25, 2026
Examiner Interview Summary
Jul 27, 2026
Response after Non-Final Action
Aug 10, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705292
USER PROFILE FILTERING BASED UPON SENSITIVE TOPICS
3y 1m to grant Granted Aug 11, 2026
Patent 12675493
SYSTEMS AND METHODS FOR MAPPING A TERM TO A VECTOR REPRESENTATION IN A SEMANTIC SPACE
2y 0m to grant Granted Jul 07, 2026
Patent 12670223
Search System Having Task-Based Machined-Learned Models
2y 6m to grant Granted Jun 30, 2026
Patent 12670162
UNIFIED STATISTICS COLLECTION FRAMEWORK USING A PROCESS-BASED TOP-DOWN APPROACH FOR POSTGRES-BASED DATABASE SYSTEMS
2y 9m to grant Granted Jun 30, 2026
Patent 12657174
STORAGE MANAGEMENT METHODS AND APPARATUSES FOR DISTRIBUTED DATABASE
2y 8m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
67%
Grant Probability
93%
With Interview (+26.0%)
3y 9m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 641 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month