Prosecution Insights
Last updated: August 17, 2026
Application No. 18/401,186

AI-GENERATED DATASETS FOR AI MODEL TRAINING AND VALIDATION

Non-Final OA §103
Filed
Dec 29, 2023
Examiner
JACKSON, JAKIEDA R
Art Unit
2657
Tech Center
2600 — Communications
Assignee
Microsoft Technology Licensing, LLC
OA Round
3 (Non-Final)
74%
Grant Probability
Favorable
3-4
OA Rounds
5m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
681 granted / 919 resolved
+12.1% vs TC avg
Strong +16% interview lift
Without
With
+15.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
29 currently pending
Career history
949
Total Applications
across all art units

Statute-Specific Performance

§101
27.1%
-12.9% vs TC avg
§103
42.2%
+2.2% vs TC avg
§102
20.9%
-19.1% vs TC avg
§112
2.8%
-37.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 919 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on May 26, 2026 has been entered. Response to Arguments Applicants argue that the prior art cited fails to teach the claims as amended. Applicants’ arguments are persuasive, but are moot in view of new grounds of rejection. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-7, 15-16 and 18-19 is/are rejected under 35 U.S.C. 103 as being unpatenable over Folkens et al. (PGPUB 2016/0283595), hereinafter referenced as Folkens in view of Klein et al. (PGPUB 2024/0241624), hereinafter referenced as Klein and in further view of Budurean et al. (USPN 9,934,129), hereinafter referenced as Budurean. Regarding claim 1, Folkens discloses a method, comprising: obtaining a screenshot of an active portion of application window rendered by an application (capture screen; fig. 2 with p. 0128-0129); obtaining an application window metadata (p. 0097, 0112-0115); identifying an image within the screenshot (image of interest; fig. 2 element 230 with p. 0128-0129); generating, with a multi-modal model, a caption of the image (tag; fig. 2, element 240 with p. 0128-0129); and generating a label of the screenshot based on the caption of the image and the application window metadata (tag; fig. 2, element 240 with p. 0128-0129, 0097, 0112-0115), but does not specifically teach training a machine learning model to predict human-computer interactions from screenshots by providing the machine learning model with the screenshot and the label of the screenshot, wherein another application personalizes a user experience by using the machine learning model to predict, from an individual screenshot, what a user is doing on a computing device and a tree of a user interface elements captured when the screenshot is taken. Klein discloses a method comprising training a machine learning model to predict human-computer interactions from screenshots by providing the machine learning model (training a machine learning by learning past data of user interactions; p. 0156) with the screenshot and the label of the screenshot (screenshot as well as context data in the screenshot), wherein another application personalizes a user experience by using the machine learning model to predict (predict user intent from screenshot), from an individual screenshot, what a user is doing on a computing device (p. 0107-0115), to determine, detect and predict user intent. Therefore, it would have been obvious to one of ordinary skill of the art, before the effective filing date of the claimed invention, to modify the method as described above, to assist with seamlessly performing native functions and to support multi-modal input. Budurean discloses a method comprising a tree of a user interface elements captured when the screenshot is taken (The screenshot metadata traversed by developer service module 162 may be organized as a tree structure or other type of hierarchal data structure so as to enable developer service module 162 to quickly compare the metadata between two or more screenshots 164 to determine what differences there are if any. For example, using tree comparison techniques, developer service module 162 may determine whether the metadata of two screenshots 164 is isomorphic, and if not, what the differences between the two metadata tree structures are; column 9, lines 7-25), to assist with developing data. Therefore, it would have been obvious to one of ordinary skill of the art, before the effective filing date of the claimed invention, to modify the method as described above, to more quickly and easily navigate through data. Regarding claim 2, Folkens discloses a method further comprising: identifying a region of text within the screenshot (fig. 2, element 230); and extracting text from the region of text (OCR; p. 0113). In addition, Klein discloses generating the label of the screenshot at least in part based on the extracted text (OCR; p. 0115-118). Regarding claim 3, Folkens discloses a method further comprising: a title of the application window, wherein the title of the application window is displayed in a title bar of the application window (title; figs. 2 and 3 with p. 0113, 0248). Regarding claim 4, Folkens discloses a method further comprising: automatically navigating the application to a website and causing an automated agent to interact with the website in accordance with a usage history of the website (generating results automatically; fig. 3 with p. 0130). In addition, Klein discloses a method wherein the screenshot is captured after the automated agent interacts with the website (p. 0107-0115). Regarding claim 5, it is interpreted and rejected for similar reasons as set forth above. In addition, Folkens discloses a method further comprising: generating an activity set by grouping screenshots taken while causing an automated agent to navigate to frequently visited locations of the application, validating a feature of an individual application that uses an individual machine learning model trained with the activity set (measure of confidence assigned to correctly characterized content; p. 0021, 0037, 0067, 0086-0089, 0099). Regarding claim 6, Folkens discloses a method further comprising: generating an activity set by grouping screenshots taken while causing an automated agent to navigate through a stream of locations within the application window (classification; p. 0094-0099, 0107-0108, 0032, 0068-0069); and training a machine learning model with the activity set (retag/retrain for improved processing; p. 0115-0123, 0127, 0166). Regarding claim 7, Folkens discloses a method further comprising: providing an individual screenshot to a feature of an application that uses the machine learning model; (fig. 3); and validating the feature by comparing an output of the feature by comparing an output of the feature with a label of the individual screenshot (comparing data; p. 0212). Regarding claim 15, it is interpreted and rejected for similar reasons as set forth in claim 1. In addition, Folkens teaches obtaining metadata of the application (p. 0097, 0112-0115, 0122, 0241). Regarding claim 16, Folkens discloses a medium wherein the metadata comprises a tree of properties of application windows of a desktop that includes the application (list/rank/order/relevance; p. 0097, 0241-0242, 0258-0259). Regarding claim 18, it is interpreted and rejected for similar reasons as set forth above. In addition, Klein discloses a medium wherein the metadata includes a description of an image displayed in the application, wherein the description of the image is obtained from the application and wherein the label of the screenshot is generated in part based on the description of the image (p. 0115-0118). Regarding claim 19, Folkens discloses a medium wherein the label is generated by a large language model based on a prompt that instructs the large language model to label the annotated screenshot with a particular use (large library; p. 0213). Claim(s) 9-14 is/are rejected under 35 U.S.C. 103 as being unpatenable over Folkens in view of Klein and Budurean and in further view of Spivack et al. (PGPUB 20100268720), hereinafter referenced as Spicack. Regarding claim 9, it is interpreted and rejected for similar reasons as set forth in claim 1. In addition, Folkens discloses a system comprising: a processing unit (p. 0094-0096); and a memory storing having computer-executable instructions, which, when executed by the processing unit (p. 0094-0096), cause the processing unit to: receive a source text derived from a screenshot of an application window of an application (OCR; p. 0112-0115); receive a label that annotates the screenshot (tag; fig. 2, element 240 with p. 0128-0129); determine a level of correctness of the label in relation to the source text (measure of confidence assigned to correctly characterized content; p. 0021, 0037, 0067, 0086-0089); determine that the label satisfies a quality criteria (measure of confidence assigned to correctly characterized content; p. 0021, 0037, 0067, 0086-0089); determine a label grade of the label based on the level of correctness and the determination that the label satisfies the quality criteria (measure of confidence assigned to correctly characterized content; p. 0021, 0037, 0067, 0086-0089); and determine that the label grade exceeds a defined threshold and validate a feature of an application that uses a machine learning model trained on labeled screenshots of application windows, wherein validating comprises providing the label to the machine learning model and verifying that an output of the machine learning model returns or identifies the screenshot (measure of confidence assigned to correctly characterized content; p. 0021, 0037, 0067, 0086-0089, 0099), however, the Folkens in view of Klein and Budurean fails to teach determining using a label quality classifier, that the label satisfies a quality criteria, wherein the quality criteria indicates whether the label is useful for understanding user behavior with the application that led to the screenshot. Spivack discloses a system comprising determining using a label quality classifier, that the label satisfies a quality criteria (tag), wherein the quality criteria (rating) indicates whether the label is useful for understanding user behavior with the application that led to the screenshot (tracking user behavior; p. 0088, 0096, 0099, 0156-0159, 0036, 0170-0178), to assist with customizing, enhance and optimize the experience. Therefore, it would have been obvious to one of ordinary skill of the art, before the effective filing date of the claimed invention, to modify the method as described above, to associate properties and attributes defined in one or more ontologies from behaviors and patterns. Regarding claim 10, Folkens discloses a system wherein the computer-executable instructions further cause the processing unit to: identify a portion of the source text that satisfies a usefulness criteria, wherein the level of correctness of the label is determined in relation to the identified portion of the source text (measure of confidence assigned to correctly characterized content; p. 0021, 0037, 0067, 0086-0089). Regarding claim 11, Folkens discloses a system wherein the label is one of a plurality of labels, and wherein the computer-executable instructions further cause the processing unit to: compute a diversity score of the plurality of labels, wherein the label grade is additionally based on the diversity score (weightings/uniqueness/numeric value; p. 0099-0101, 0122, 0127). Regarding claim 12, Folkens discloses a system wherein the diversity score is computed based on distances between embedding scores computed for each of the plurality of labels, and wherein the label grade (weightings) is proportional to the diversity score (uniqueness; p; 0099-0101, 0122, 0127). Regarding claim 13, Folkens discloses a system, wherein the source text is processed by a short text clustering engine before the portion of the source text is determined to satisfy the usefulness criteria (association/classification; p. 0107-0108, 0032, 0068-0069, 0094-0099). Regarding claim 14, Folkens discloses a system wherein the computer-executable instructions further cause the processing unit to: identify, across a plurality of labels applied to a plurality of screenshots, clusters of related explanations of label incorrectness or low label quality (lower weight; p. 0094-0099, 0107-0108, 0032, 0068-0069). Claim(s) 17 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Folkens in view of Klein and Budurean and in further view of Winn et al. (PGPUB 2020/0150832), hereinafter referenced as Winn. Regarding claim 17, Folkens and Klein and Budurean disclose a medium as described above, but fails to teach wherein the screenshot is of a desktop, and wherein the screenshot is cropped to the application based on a location and a size of the application window obtained from the metadata. Winn discloses a medium wherein the screenshot is of a desktop, and wherein the screenshot is cropped to the application based on a location and a size of the application window obtained from the metadata (automatically enhance, crop, reduce file size; p. 0079, 0135), to provide a high-quality output. Therefore, it would have been obvious to one of ordinary skill of the art, before the effective filing date of the claimed invention, to modify the method as described above, to refresh the user interface. Regarding claim 20, it is interpreted and rejected for similar reasons as set forth above. In addition, Folkens discloses a medium wherein the screenshot is generated by a computing device configured with a language (p. 0113). In addition, Winn teaches a medium wherein a screen resolution and a user interface theme selected to create screenshots under a diversity of computing environments (image resolution; p. 0054, 0066-0067, 0079, 0135, 0022-0023). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. This information has been detailed in the PTO 892 attached (Notice of References Cited). Barhoumeh et al. teaches a user behavior rule of taking a screenshot. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAKIEDA R JACKSON whose telephone number is (571)272-7619. The examiner can normally be reached Mon - Fri 6:30a-2:30p. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at 571.272.5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JAKIEDA R JACKSON/Primary Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Show 1 earlier event
Aug 15, 2025
Non-Final Rejection mailed — §103
Oct 23, 2025
Examiner Interview Summary
Nov 17, 2025
Response Filed
Feb 23, 2026
Final Rejection mailed — §103
May 26, 2026
Request for Continued Examination
May 28, 2026
Response after Non-Final Action
Jun 16, 2026
Non-Final Rejection mailed — §103
Aug 07, 2026
Interview Requested

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12706082
SPEECH TRANSLATION USING LATENCY BASED FILLER GENERATION
2y 5m to grant Granted Aug 11, 2026
Patent 12706090
METHOD AND APPARATUS WITH REAL-TIME TRANSLATION
2y 5m to grant Granted Aug 11, 2026
Patent 12700396
SYSTEMS AND METHODS FOR AUTOMATED SYNTHETIC VOICE PIPELINES
3y 9m to grant Granted Aug 04, 2026
Patent 12700412
CACHING SCHEME FOR VOICE RECOGNITION ENGINES
2y 1m to grant Granted Aug 04, 2026
Patent 12700414
MULTI-CHANNEL SIGNAL ENCODING METHOD, MULTI-CHANNEL SIGNAL DECODING METHOD, ENCODING DEVICE, DECODING DEVICE, AND TERMINAL DEVICE
1y 10m to grant Granted Aug 04, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
74%
Grant Probability
90%
With Interview (+15.7%)
3y 0m (~5m remaining)
Median Time to Grant
High
PTA Risk
Based on 919 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month