Prosecution Insights
Last updated: October 02, 2026
Application No. 18/663,730

Character recognition-based augmentation for multimodal model inputs

Final Rejection §101§103
Filed
May 14, 2024
Examiner
LI, RUIPING
Art Unit
2676
Tech Center
2600 — Communications
Assignee
Google LLC
OA Round
2 (Final)
77%
Grant Probability
Favorable
3-4
OA Rounds
4m
Est. Remaining
95%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
740 granted / 963 resolved
+14.8% vs TC avg
Strong +18% interview lift
Without
With
+18.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
28 currently pending
Career history
982
Total Applications
across all art units

Statute-Specific Performance

§101
10.8%
-29.2% vs TC avg
§103
44.7%
+4.7% vs TC avg
§102
25.3%
-14.7% vs TC avg
§112
15.8%
-24.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 963 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . 2. This is in response to the applicant response filed on 07/23/2026. In the applicant’s response, claims 1, 3, 8-10, 13, and 20 were amended. Accordingly, claims 1-20 are pending and being examined. Claims 1, 13, and 20 are independent form. Claim Rejections - 35 USC § 101 3. The claim rejections under 35 USC § 101 make in the previous office action are withdrawn in view of applicant’s amendments. Claim Rejections - 35 USC § 103 4. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 5. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 6. Claim 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al (“Convolutional Neural Networks and Multimodal Fusion for Text Aided Image Classification”, 2017) in view of Gupta et al (“Automatic Assessment of OCR Quality in Historical Documents”, 2015). Regarding claim 1, Wang discloses a method (the Convolutional Neural Networks (the CNNs) based method/system for text-aided image classification; see fig.1), comprising: receiving, by one or more processors, multimodal input comprising at least two of text, images, video, or audio (the CNNs may receive a “test image” including text data; see the “text image” of fig.1 and Sec. II-A and II-B); determining, as the multimodal input is received by the one or more processors and based on one or more criteria, the reliability of character recognition (CR) data through a first multimodal model trained to receive the multimodal input and generate a multimodal model output (the CNNs may assign the weight w t j   defined by Eq(8) to the text classification result d 2 ( j ) to determine the reliability of the textual feature-based classification result and generate the final classification result s j (or q j ) defined by Eq(6). It should be noticed that: the textual feature-based classification includes character recognition (CD), see ‘airplanes’... and ‘jet’ in fig.2); generating, by the one or more processors, the multimodal model output using the multimodal input and the CR data (the CNNs may output the final classification decision score D based on the test image and the three reliabilities w p j ,   w t j   ,   and w c j , wherein the w t j   indicates the weight of the text/character classification result d 2 ( j ) ; see the right col of fig.1, Eq(6), Eq(8), and Eq(14), and Sec. III), wherein in generating the multimodal model output the first multimodal model (see the final classification decision score D of fig.1) processes the multimodal input to generate a first output (see fig.1, from “test image”[Wingdings font/0xE0]“Pre-trained VGG-16 model”[Wingdings font/0xE0]”FCa”[Wingdings font/0xE0]”FCb”[Wingdings font/0xE0]FC1, i.e., to the first output D1), processes both the multimodal input with the CR data to generate a second output (see fig.1, wherein the “Early fusion” outputs the third output D3 based on the image information “FCb” and the text information ”FCc”), and compares the first output to the second output to select the multimodal model output (see fig.1, wherein the “Later fusion” outputs the final output D by comparing the reliabilities   w p j   a n d   w t j   in form of q j = w p j d 1 ( j ) + w t j d 2 ( j ) + w c j d 3 ( j ) ; Eq(11) and the first three lines right-up the Eq(14). In other words, the system in Wang “compares” all the three outputs and gives them three different weights); and training a second multimodal model using the first output and the second output (training, I.e., minimizing the cost by Eq(12) and taking the highest decision score by Eq(14); see III-B). As explained above, though Wang does not explicitly disclose the feature of “determining [...] whether to process the multimodal input with character recognition (CR) data” recited in the claim, Wang discloses assigning the reliability (i.e., the weight) w t j   defined by Eq(8) to the textual feature-based classification result d 2 ( j ) , see Eqs(6), (8), (11), and (12). And then, the final classification result q i j output by the CNNs is determined by comparing all three reliabilities. In other words, Wang has appreciated that OCR quality issue needs to be considered in multimodal document classification. In fact, in the same field of endeavor, Gupta clearly points out “when a document has poor quality, the OCR engine generally produces a large number of spurious bounding boxes (BBs) in addition to those that correspond to words in the document.” See Gupta, Sec. “Introduction”, Parapraph.3. To resolve this issue, Gupta, see Sec. “Method”, teaches a “pre-filtering” process prior to OCR, wherein the pre-filtering process uses three criteria (rules), that is, OCR word confidence, height-to-width rate, and area, to determine whether to process the input with character recognition (CR) data. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was made to incorporate the teachings of Gupta into the teachings of Wang and use the criterion for OCR word confidence to determine whether to perform a pre-filtering to input text images taught by Gupta. Suggestion or motivation for doing so would have been to provide a robust OCR for document images. Gupta, see Sec. “Introduction”, paragraph 3; Wang, see Sec. III-B, paragraph 2. Therefore, the claim is unpatentable over Wang in view of Gupta. Regarding claim 2, the combination of Wang and Gupta discloses the method of claim 1, wherein the CR data identifies or characterizes text in the multimodal input (Wang, see the word/text classification decision result D2 output by the CNN in fig.2/fig.1 and Sec. II-B). Regarding claim 3, 14, the combination of Wang and Gupta discloses, wherein determining whether to process the multimodal input with the CR data comprises: generating the CR data; and determining whether the generated CR data meet the one or more predetermined criteria (Gupta, see “Rule I: OCR word confidence. BBs with very low or very high confidence predominantly consist of noise, and are flagged accordingly during pre-filtering.”), and in response, generating, by the one or more processors, the multimodal model output using the multimodal input without the CR data (Gupta, performing a “pre-filtering” process to the text image prior to OCR; see “Method”, “Pre-Filtering”) Regarding claim 4, 15, the combination of Wang and Gupta discloses, wherein determining whether to process the multimodal input with CR data comprises determining whether the multimodal input or the CR data satisfy the one or more predetermined criteria, comprising one or more of whether: Regarding claim 5, the combination of Wang and Gupta discloses the method of claim 1, wherein determining whether to process the multimodal input with the CR data comprises: determining, based on the multimodal input and the one or more criteria, whether to generate the CR data from the multimodal input; and generating the CR data from the multimodal input (Gupta, determining whether a BB (a bounding box extracted from the input text image) is a noise BB or a text BB and performing OCR only to text BBs; see fig.2 and “Method”, “Pre-Filtering”). Regarding claim 6, 16, 17, 18, the combination of Wang and Gupta discloses, wherein the one or more criteria are based on at least one of: the length of the multimodal input, the quantity of images or videos in the multimodal input, the size of the images or video in the multimodal input, or the resolution or quality of components of the multimodal input (Gupta, see “Rule 1”, “Rule 2”, and “Rule 3”,). Regarding claim 7, the combination of Wang and Gupta discloses the method of claim 1, wherein the method further comprises: generating, by the one or more processors, the CR data; and formatting, by the one or more processors, the multimodal input and the CR data according to one of one or more predetermined formats (Wang, the CNNs may assign the weight w t j   defined by Eq(8) to the textual feature-based classification result d 2 ( j ) to determine the reliability of the textual feature-based classification result and then generate final classification result q i j defined by Eq(11).). Regarding claim 8, the combination of Wang and Gupta discloses the method of claim 1, wherein determining whether to process the multimodal input with the CR data comprises: training the multimodal model to: receive the multimodal input, and determine, based on the multimodal input, whether to generate a model output with the multimodal input or the multimodal input with the CR data (Gupta, wherein the labelling for noise and text BBs is trained by dataset to optimize the threshold values; see fig.5 and Sec. “Results”, “Pre-filtering”. Wang, see the CNN training, shown by fig.1 and stated Sec. II). Regarding claim 9, the combination of Wang and Gupta discloses the method of claim 8, wherein the method further comprises: training, by the one or more processors, the multimodal model on training data comprising: examples of model outputs generated with multimodal inputs, and examples of model outputs generated with the multimodal inputs and respective CR data identifying or characterizing text in each of the multimodal inputs (ibid.). Regarding claim 10, the combination of Wang and Gupta discloses the method of claim 9, wherein determining whether to generate the CR data comprises: executing the multimodal model with the multimodal input to generate a first output; executing the multimodal model with the multimodal input and the CR data to generate a second output; and outputting one of the first output and the second output based on a comparison of the first output and the second output (ibid.). Regarding claim 11, the combination of Wang and Gupta discloses the method of claim 1, further comprising: processing, by the one or more processors, the response through a machine learning model trained to generate output at least from the multimodal input (Wang, see training the CNN, shown by fig.1 and stated in Sec. II. Gupta, see training the labelling for noise and text BBs, shown by fig.5 and stated in Sec. “Results”,). Regarding claim 12, 19, the combination of Wang and Gupta discloses, wherein the CR data is optical character recognition (OCR) data generated by performing an OCR process on at least a portion of the multimodal input (Wang, see text recognition in fig.2 and Sec. II-B. Gupta, see word recognition, in “Rule I: OCR word confidence”). Regarding claim 13, 20, each of them is an inherent variation of claim 1, thus it is interpreted and rejected for the reasons set forth in the rejection of claim 1. Response to Arguments 7. Applicant’s arguments, filed on 07/23/2026, have been fully considered but they are not persuasive. On page 9, regarding claims 1, 13, and 20, applicant argues: it is respectfully submitted that there is no disclosure in the cited references from which a POSA could reasonably conclude that the references disclose "wherein in generating the multimodal model output the first multimodal model processes the multimodal input to generate a first output, processes both the multimodal input with the CR data to generate a second output, and compares the first output to the second output to select the multimodal model output; and training a second multimodal model using the first output and the second output," as recited in claims 1 and 20. Similarly, it is respectfully submitted that a POSA would also be unable to reach such a conclusion with respect to claim 13 with respect to recital of "wherein to generate the multimodal model output the first multimodal model processes the multimodal input to generate a first output and processes both the multimodal input with the CR data to generate a second output, and compares the first output to the second output to select the multimodal model output; and train a second multimodal model using the first output and the second output." The examiner respectfully disagrees with the applicant’s arguments. As explained above, Wang, see Sec. III-B, paragraph 2, clearly states “Since web resources have great reliability diversity, it may not be an optimal practice to allocate fixed weights to the visual feature-based and textual feature-based decisions [,]. In this paper, an adaptive multimodal fusion algorithm is developed, where the classification results D1, D2 and D3 are dynamically weighted to obtain the final decision.” To this end, the three different indexes w p j ,   w t j   a n d   w c j   are assigned to three different classifications, respectively, namely, visual-only based classification D1, text-only based classification D2, and visual-and-text based classification D3. After then the three weights are compared to be normalized to 1, as shown by Eqs (7)-(9). A person skilled in the art (POSA) would be able to reach such a conclusion is because, as long as he/she understands the techniques on “Convolutional Neural Networks and Multimodal Fusion for Text Aided Image Classification”, the person would understand that the three indexes w p j ,   w t j   a n d   w c j   are competed (i.e., comparison) to reach the value one. For example, as w p j = 1 , and then w t j = w c j = 0 . In addition, Wang, see the right below Eq(9), further states “where λj is a regularizer to control the reliability-dependent weights wpj, wtj , and wcj with respect to the j-th class.” In other words, the system in Wang may choose different λj values to control the weights by comparing the respective results. The arguments thus are not persuasive. Conclusion 8. THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. 9. Any inquiry concerning this communication or earlier communications from the examiner should be directed to RUIPING LI whose telephone number is (571)270-3376. The examiner can normally be reached 8:30am--5:30pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, HENOK SHIFERAW can be reached on (571)272-4637. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit https://patentcenter.uspto.gov; https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center, and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /RUIPING LI/Primary Examiner, Ph.D., Art Unit 2676
Read full office action

Prosecution Timeline

May 14, 2024
Application Filed
Mar 23, 2026
Non-Final Rejection mailed — §101, §103
Jul 23, 2026
Response Filed
Aug 10, 2026
Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12738093
DEPTH ASSISTED IMAGES REFINEMENT
4y 1m to grant Granted Sep 15, 2026
Patent 12738092
LEARNING DOMAIN AND POSE INVARIANCE FOR THERMAL-TO-VISIBLE FACE RECOGNITION
3y 5m to grant Granted Sep 15, 2026
Patent 12737958
GENERATING PERSONALIZED VIDEOS WITH CUSTOMIZED TEXT MESSAGES
2y 10m to grant Granted Sep 15, 2026
Patent 12731313
Object Detection Training Based on Artificially Generated Images
2y 12m to grant Granted Sep 08, 2026
Patent 12718528
REAL TIME FACE SWAPPING SYSTEM AND METHODS THEREOF
3y 7m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
77%
Grant Probability
95%
With Interview (+18.4%)
2y 9m (~4m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 963 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month