Prosecution Insights
Last updated: October 02, 2026
Application No. 18/626,165

METHOD AND APPARATUS FOR TRAINING IMAGE RECOGNITION MODEL, DEVICE, AND MEDIUM

Non-Final OA §103
Filed
Apr 03, 2024
Priority
May 17, 2022 — CN 202210533141.4 +3 more
Examiner
LIN, JESSICA YIFANG
Art Unit
2668
Tech Center
2600 — Communications
Assignee
Tencent Technology (Shenzhen) Company Limited
OA Round
3 (Non-Final)
82%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
78%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
9 granted / 11 resolved
+19.8% vs TC avg
Minimal -3% lift
Without
With
+-3.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
52 currently pending
Career history
67
Total Applications
across all art units

Statute-Specific Performance

§101
2.7%
-37.3% vs TC avg
§103
67.3%
+27.3% vs TC avg
§102
26.7%
-13.3% vs TC avg
§112
3.0%
-37.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 11 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. Information Disclosure Statement The information disclosure statement (IDS) submitted on 4/16/2024 and 3/9/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 7/21/2026 has been entered. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1, 6, 8, 12-13, 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sharma, Yash et al. “Cluster-to-Conquer: A Framework for End-to-End Multi-Instance Learning for Whole Slide Image Classification.” International Conference on Medical Imaging with Deep Learning (2021) in view of Pao (United States Patent Application Publication US 2025/0104450 A1) and Hohne et. al. (United States Patent Application Publication US 2025/0265832 A1). Regarding claim 1, Sharma et. al. discloses a method for training an image recognition model, performed by a computer device, the method comprising (Sharma et. al. Figure 1): obtaining a sample image and a corresponding sample label; segmenting the sample image into a sample image patch bag of sample image patches corresponding to the sample image (Sharma et. al., Fig.1(a), section 3.1 “For digital pathology classification problems, WSIs (W) of patients are available along with their disease labels.”), the sample patch bag having a bag label corresponding to the sample label of the sample image (Sharma et. al. Fig.1(b)-(c), section 3.1 “Hence, using the Otsu thresholding approach and sliding window approach, patches containing substantial tissue area (>50%) of desirable size are extracted. Given a WSI W (bag) with label y, we extract w1, w2, w3, …, wn patches (instances) from it for training.”); predicting an attention distribution of the sample image patch bag from the bag feature and the plurality of patch features by the attention layer in the image recognition model (Sharma et. al. Fig. 1(d)-(f), section 3.2 “We used the weighted-average aggregation approach proposed in Ilse et. al. (2018) for aggregating the patch-level representation to obtain WSI-level representations.”, section 3.4 “Using the aggregated representation of the WSI and representation of patches (instances), end-to-end training is performed using cross-entropy and KL-divergence loss. Along with WSI and patch cross-entropy loss, for each cluster, KL-divergence loss between the patches’ attention weight and a uniform distribution is included.”); determining a relative entropy loss corresponding to the sample patch bag based on a difference the attention distribution predicted and an expected distribution corresponding to the bag label of the sample patch bag, the attention distribution being a distribution obtained by predicting image content in the sample patch bag, and the expected distribution being a distribution of the sample patch bag indicated by the bag label (Sharma et. al. Fig. 1(d)-(f), section 3.2 “We used the weighted-average aggregation approach proposed in Ilse et. al. (2018) for aggregating the patch-level representation to obtain WSI-level representations.”, section 3.4 “Using the aggregated representation of the WSI and representation of patches (instances), end-to-end training is performed using cross-entropy and KL-divergence loss. Along with WSI and patch cross-entropy loss, for each cluster, KL-divergence loss between the patches’ attention weight and a uniform distribution is included.”); determining a first cross entropy loss corresponding to the sample patch bag based on a difference between the bag label and the corresponding bag feature analysis result of the sample patch bag; for each sample image patch of the sample image patches in the sample patch bag: performing feature analysis on the sample image patch by using the second fully connected layer in the image recognition model to obtain a patch feature analysis result of the sample image patch; determining a second cross entropy loss corresponding to the sample image patch based on a difference between a patch label of the sample image patch and the corresponding patch analysis result of the sample image patch (Sharma et. al. Fig.1 (g), section 3.4 “…the instance representation are passed through Gy’: h [Wingdings font/0xE0] y’ to obtain patches prediction probability. Instance loss is included with weak supervision assumption. Along with WSI and patch cross-entropy loss, for each cluster, KL-divergence loss between the patches’ attention weight and a uniform distribution is included.”); performing weighted fusion on the relative entropy loss, the first cross entropy loss, and the plurality of second cross entropy losses to obtain a total loss value of the image recognition model (Sharma et. al. section 3.4 “The instance representation are passed through Gy’: h[Wingdings font/0xE0]y’ to obtain patches prediction probability…Instance loss is included with weak supervision assumption”); and training the first fully connected layer, the second fully connected layer, and the attention layer in the image recognition model simultaneously based on the total loss value, the trained image recognition model being configured to recognize image content in an image (Sharma et.al. loss L(Gy, Gy’, Ga, Ge) in section 3.4). However, Sharma et. al. fails to disclose wherein the image recognition model includes a first fully connected layer, a second fully connected layer that is communicatively coupled to the first fully connected layer, and an attention layer that is communicatively coupled to both the first fully connected layer and the second fully connected layer, processing the sample image patch bag by the first fully connected layer in the image recognition model to obtain a bag feature corresponding to the sample image patch bag; processing the sample image patches in the sample image patch bag by the first fully connected layer in the image recognition model to obtain a plurality of patch features corresponding to the sample image patches in the sample image patch bag; generating a fused feature of the sample image patch bag from the attention distribution of the sample image patch bag, the bag feature and the plurality of patch features; processing the fused feature of the sample image patch bag by using the second fully connected layer to obtain a bag feature analysis result of the sample patch bag indicating a recognition result of the image content in the sample patch bag. Pao teaches wherein the image recognition model includes a first fully connected layer, a second fully connected layer that is communicatively coupled to the first fully connected layer, and an attention layer that is communicatively coupled to both the first fully connected layer and the second fully connected layer, processing the sample image patch bag by the first fully connected layer in the image recognition model to obtain a bag feature corresponding to the sample image patch bag; processing the sample image patches in the sample image patch bag by the first fully connected layer in the image recognition model to obtain a plurality of patch features corresponding to the sample image patches in the sample image patch bag (Pao Abstract, [0005], [0010], [0022], Fig 9A: attention weights). PNG media_image1.png 602 508 media_image1.png Greyscale Hohne et. al. teaches generating a fused feature of the sample image patch bag from the attention distribution of the sample image patch bag, the bag feature and the plurality of patch features; processing the fused feature of the sample image patch bag by using the second fully connected layer to obtain a bag feature analysis result of the sample patch bag indicating a recognition result of the image content in the sample patch bag (Hohne et. al. [0156]: These regional features are fused with the local features of selected patches, and optionally patches adjacent to the selected patches, to form a global embedding. [0025]-[0033]: computing a loss based on a difference between the class to which the global embedding is assigned and the class to which the training image is assigned, modifying parameters of the first machine learning model based on the computed loss). These features are important to the claimed invention because the architecture of the machine learning model via attention neural network improves the accuracy based on the three different types of entropy losses. Thus, it would have been obvious to one skilled in the art prior to the effective filing date of the claimed invention to have combined the teachings of Sharma et. al., Pao and Hohne et. al., so these features are included in the solution of the claimed invention. Regarding claim 8, which is a computer device, comprising a processor and a memory, the memory having at least one program stored therein that, when executed by the processor, causes the computer device to implement a method for training an image recognition model of claim 1, which the rejection analysis is incorporated herein. Regarding claim 15, which is a non-transitory computer-readable storage medium, having at least one program stored thereon that, when executed by a processor of a computer device, causes the computer device to implement a method for training an image recognition model of claim 1, which the rejection analysis is incorporated herein. Regarding claim 6, Sharma et. al. further discloses the method according to claim 5, wherein the method further comprises: using, in response to the sample image patches in the sample patch bag belonging to a same sample image, a sample label corresponding to the sample image as a bag label corresponding to the sample patch bag; and determining, in response to the sample image patches in the sample patch bag belonging to different sample images, the bag label corresponding to the sample patch bag based on the patch labels corresponding to the sample image patches (Sharma et. al. section 3.4 “The aggregated representation is passed through Gy: z[Wingdings font/0xE0]y to obtain WSI prediction probability). Regarding claim 12, Sharma et. al. further discloses the computer device according to claim 8, wherein the segmenting the sample image into a sample image patch bag of sample image patches corresponding to the sample image comprises: segmenting an image region of the sample image to obtain the sample image patches; and allocating sample image patches belonging to a same sample image to a same bag to obtain the sample patch bag (Sharma et. al. Fig. 1 (d)-(g), section 3.2 “weighted-average aggregation approach…to obtain WSI-level representations.” And section 3.4 “KL-divergence loss between the patches’ attention weight and a uniform distribution is included. The aggregated representation is passed through Gy: y[Wingdings font/0xE0] z to obtain WSI prediction probability”). Regarding claim 13, Sharma et. al. further discloses the computer device according to claim 12, wherein the method further comprises: using, in response to the sample image patches in the sample patch bag belonging to a same sample image, a sample label corresponding to the sample image as a bag label corresponding to the sample patch bag; and determining, in response to the sample image patches in the sample patch bag belonging to different sample images, the bag label corresponding to the sample patch bag based on the patch labels corresponding to the sample image patches (Sharma et. al. Fig. 1 (d)-(g), section 3.2 “weighted-average aggregation approach…to obtain WSI-level representations.” And section 3.4 “KL-divergence loss between the patches’ attention weight and a uniform distribution is included. The aggregated representation is passed through Gy: y[Wingdings font/0xE0] z to obtain WSI prediction probability”). Claim(s) 5 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Sharma, Yash et al. “Cluster-to-Conquer: A Framework for End-to-End Multi-Instance Learning for Whole Slide Image Classification.” International Conference on Medical Imaging with Deep Learning (2021), Pao (United States Patent Application Publication US 2025/0104450 A1) and Hohne et. al. (United States Patent Application Publication US 2025/0265832 A1) as applied to claims 1 and 15 above, in further view of Jewsbury, Robert et al. “A QuadTree Image Representation for Computational Pathology.” 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) (2021): 648-656. (Year: 2021). Regarding claim 5, Sharma et. al., Pao, and Hohne et. al. disclose the method according to claim 1. However, Sharma et. al., Pao and Hohne et. al. fail to disclose wherein the segmenting the sample image into a sample image patch bag of sample image patches corresponding to the sample image comprises: segmenting an image region of the sample image to obtain the sample image patches; and allocating sample image patches belonging to a same sample image to a same bag to obtain the sample patch bag. Jewsbury et. al. teaches wherein the segmenting the sample image into a sample image patch bag of sample image patches corresponding to the sample image comprises: segmenting an image region of the sample image to obtain the sample image patches; and allocating sample image patches belonging to a same sample image to a same bag to obtain the sample patch bag (Jewsbury et. al. 2.2 MIL framework, where generated collection of down-sampled image regions are given labels either positive or negative based on categories of features). PNG media_image2.png 533 1142 media_image2.png Greyscale Separating the sample images into different categories based on features that are extracted are critical to the claimed invention. This defines the solution to the problem. Thus, it would have been obvious to one skilled in the art prior to the effective filing date of the claimed invention to have combined the teachings of Sharma et. al., Pao, Hohne et. al. and Jewsbury so that this solution is implemented fully. Regarding claim 19, Sharma et. al., Pao, and Hohne et. al. disclose the non-transitory computer-readable storage medium according to claim 15. However, Sharma et. al., Pao, and Hohne et. al. fail to disclose wherein the segmenting the sample image into a sample image patch bag of sample image patches corresponding to the sample image comprises: segmenting an image region of the sample image to obtain the sample image patches; and allocating sample image patches belonging to a same sample image to a same bag to obtain the sample patch bag. Jewsbury et. al. teaches wherein the segmenting the sample image into a sample image patch bag of sample image patches corresponding to the sample image comprises: segmenting an image region of the sample image to obtain the sample image patches; and allocating sample image patches belonging to a same sample image to a same bag to obtain the sample patch bag (Jewsbury et. al. 2.2 MIL framework, where generated collection of down-sampled image regions are given labels either positive or negative based on categories of features). PNG media_image2.png 533 1142 media_image2.png Greyscale Separating the sample images into different categories based on features that are extracted are critical to the claimed invention. This defines the solution to the problem. Thus, it would have been obvious to one skilled in the art prior to the effective filing date of the claimed invention to have combined the teachings of Sharma et. al., Pao, Hohne et. al. and Jewsbury so that this solution is implemented fully. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JESSICA YIFANG LIN whose telephone number is (571)272-6435. The examiner can normally be reached M-F 7:00am-6:15pm, with optional day off. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be reached at 571-272-7332. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JESSICA YIFANG LIN/Examiner, Art Unit 2668 September 14, 2026 /VU LE/Supervisory Patent Examiner, Art Unit 2668
Read full office action

Prosecution Timeline

Show 2 earlier events
Mar 24, 2026
Interview Requested
Mar 30, 2026
Examiner Interview Summary
Mar 31, 2026
Response Filed
Apr 22, 2026
Final Rejection mailed — §103
Jun 18, 2026
Response after Non-Final Action
Jul 21, 2026
Request for Continued Examination
Jul 23, 2026
Response after Non-Final Action
Sep 18, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743744
CROSS-VIEW IMAGE GEO-LOCALIZATION
2y 8m to grant Granted Sep 22, 2026
Patent 12743869
CLASS BOUNDARY DETECTION APPARATUS, CONTROL METHOD, AND NON-TRANSITORY COMPUTER-READABLE MEDIUM
2y 5m to grant Granted Sep 22, 2026
Patent 12738010
OBJECT RECOGNITION DEVICE AND OBJECT RECOGNITION METHOD
2y 5m to grant Granted Sep 15, 2026
Patent 12718396
AVIATION DOCUMENT TARGET AREA EXTRACTION SYSTEM AND METHOD
2y 6m to grant Granted Aug 25, 2026
Patent 12711738
Real-time Media Alteration Using Generative Techniques
2y 9m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
82%
Grant Probability
78%
With Interview (-3.3%)
2y 5m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 11 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month