Prosecution Insights
Last updated: October 01, 2026
Application No. 19/376,222

AUTOMATIC IMAGE CROPPING

Final Rejection §103
Filed
Oct 31, 2025
Priority
Sep 30, 2022 — reissue of 12/300,007
Examiner
HANCE, ROBERT J
Art Unit
3992
Tech Center
3900
Assignee
Amazon Technologies Inc.
OA Round
2 (Final)
66%
Grant Probability
Favorable
3-4
OA Rounds
1y 11m
Est. Remaining
88%
With Interview

Examiner Intelligence

Grants 66% — above average
66%
Career Allowance Rate
506 granted / 761 resolved
+6.5% vs TC avg
Strong +22% interview lift
Without
With
+21.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
32 currently pending
Career history
792
Total Applications
across all art units

Statute-Specific Performance

§101
8.1%
-31.9% vs TC avg
§103
51.0%
+11.0% vs TC avg
§102
14.4%
-25.6% vs TC avg
§112
16.1%
-23.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 761 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Reissue Applications This application seeks to reissue US Patent No. 12,300,007 (“the ‘007 patent”). In a 07/15/2026 response to the 06/05/2026 non-final Office action (“NFOA”), claims 1, 2, 4, and 13 have been amended. Claims 1-20 are pending. For reissue applications filed before September 16, 2012, all references to 35 U.S.C. 251 and 37 CFR 1.172, 1.175, and 3.73 are to the law and rules in effect on September 15, 2012. Where specifically designated, these are “pre-AIA ” provisions. For reissue applications filed on or after September 16, 2012, all references to 35 U.S.C. 251 and 37 CFR 1.172, 1.175, and 3.73 are to the current provisions. Applicant is reminded of the continuing obligation under 37 CFR 1.178(b), to timely apprise the Office of any prior or concurrent proceeding in which Patent No. 12,300,007 is or was involved. These proceedings would include any trial before the Patent Trial and Appeal Board, interferences, reissues, reexaminations, supplemental examinations, and litigation. Applicant is further reminded of the continuing obligation under 37 CFR 1.56, to timely apprise the Office of any information which is material to patentability of the claims under consideration in this reissue application. These obligations rest with each individual associated with the filing and prosecution of this application for reissue. See also MPEP §§ 1404, 1442.01 and 1442.04. Applicant’s Response to the NFOA § 101 Rejections Claims 1-20 were rejected in the NFOA under § 101 as being directed to abstract ideas. NFOA at 3-7. The applicant’s arguments traversing this rejection (see Remarks at 11-16) are persuasive. In particular, the limitations that fall outside of the abstract ideas, when evaluated along with the limitations containing the judicial exception, are understood to integrate the abstract ideas into a practical application. See Remarks at 12 and MPEP 2106.04(d)(III). In other words, the “way in which the additional elements use or interact with the exception … integrate it into a practical application.” MPEP 2106.04(d)(III). This rejection is withdrawn. § 103 Rejections The independent claims were rejected in the NFOA under § 103 as obvious over Zhang in view of Gonsalves. See NFOA at 8-11. The applicant traverses, and argues that Gonsalves “compares image embeddings rather than representations of generated text.” Remarks at 17. The applicant also argues that Gonsalves “does not generate any text caption that is encoded thereafter.” Id. This is not persuasive. Gonsalves describes that “an image-encoder” generates object embeddings by and “encod[ing] images and text” into a vector. Gonsalves ¶ 20. The vector represents a “semantic embedding for the image.” Id. Therefore “images and text” are encoded to generate the embeddings. The embeddings (that is, the vectors) of two images are compared to determine a similarity between the images. Id. ¶¶ 20-22. While Gonsalves discloses that “images and text” are encoded to generate the embeddings, the reference itself does not explicitly describe how the “text” is generated. However, Gonsalves describes that the “image-encoder” is the “Oscar” model that is incorporated by reference in ¶ 20. In other words, Oscar is used “to generate object embeddings.” Id. ¶ 20. The Oscar paper describes that its model performs five understanding tasks and two generation tasks. See Oscar at 6. Because Gonsalves relies on Oscar to “generate object embeddings” (see Gonsalves ¶ 20), the POSITA would conclude that Gonsalves uses one of Oscar’s “generation tasks.” Oscar pp. 6-8 lists these tasks; the two generation tasks in this list are “Image Captioning” and “NoCaps.” Id. at 7. These tasks both describe generating captions for an input image. See id. The POSITA would therefore conclude that Gonsalves uses Oscar to generate text captions. The system will therefore “encode images and text [i.e., the text captions generated in Oscar] to a compatible vector space” that “represents a semantic embedding for the image.” Gonsalves ¶ 20. These embeddings are compared for an original image and a cropped image to determine their similarity. Id. ¶¶ 20-22. This disclosure falls within the broadest reasonable interpretation of the language of the claims1. The applicant has not addressed the content of the Oscar paper, or provided a convincing argument that Gonsalves/Oscar does not operate as described in the NFOA. The § 103 rejections are maintained. The amended claims are similar in scope to the previous claims and are addressed in the § 103 rejections below. Objection, 37 CFR 1.175 – Defective Declaration This application is objected under 37 CFR 1.175. In the applicant’s 07/15/2026 response, the claims were amended to recite determining “quantitative similarity metrics” in place of the previously-recited “cosine similarity scores.” A quantitative similarity metric encompasses cosine similarity scores along with a number of other metrics, such as “Euclidean distance, and/or some other vector similarity metrics.” See the ‘007 patent at 5:53-56. The claims have been broadened because they now cover something that the claims of the ‘007 patent did not. See MPEP 1412.03(I). For a reissue application “that seeks to enlarge the scope of the claims of the patent, the reissue oath or declaration must also identify a claim that the application seeks to broaden in the identification of the error that is relied upon to support the reissue application.” MPEP 1414(II). The declaration on file only states that this application was filed to correct inventorship. To overcome this, the applicant should submit a new declaration form PTO/AIA /05 that includes an error statement that identifies the patent claims that have been broadened in this reissue application. Please refer to MPEP 1414(II) for guidance on proper error statements in a broadening reissue application. Claim Rejections, 35 USC § 251 – Defective Declaration Claims 1-20 are rejected as being based upon a defective reissue declaration under 35 U.S.C. 251 as set forth above. See 37 CFR 1.175. The nature of the defect(s) in the declaration is set forth in the discussion above in this Office action. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 4, 11-13, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang, US 20210056663 in view of Gonsalves, US 20240054748. Claim 1: Zhang discloses a computer-implemented method, comprising: receiving first image data representing a first image (Image data representing a first, full-size image is received by a crop generation system. ¶25.); generating second image data representing a portion of the first image that is generated by cropping the first image data according to a first aspect ratio (A number of crop candidates is generated by cropping the first image data. ¶26. The cropping is performed to generate cropped images that have a particular aspect ratio. ¶¶ 26 and 42. Based on an evaluation of the cropped image to the original image, one of the crop candidates is selected as the final output image, i.e. second image data. ¶30); generating, based at least in part on a similarity score, first computer- executable instructions to cause the second image data to be displayed (Each crop candidate is evaluated based on a comparison with the source image, and how well the crop candidate preserves visual content and composition of the original image. ¶¶18 and 29. The crop candidate with the highest evaluation score is selected and is displayed. ¶30.). Zhang fails to disclose that the crop criteria relates to an aspect ratio of a first display of a target device, and displaying the second image data on the first display of the target device. In addition, while there are various disclosed methods by which the crop candidates are scored (see Zhang ¶ 30), Zhang does not determine similarity in the same way that is recited in this claim. Therefore Zhang fails to disclose the following limitations: generating, by inputting the first image data into an image captioning model, first text data representing a first description of first content of the first image data; generating, using a first text encoder, a first vector representation of the first text data; generating, by inputting the second image data into the image captioning model, second text data representing a second description of second content of the second image data; generating, using the first text encoder, a second vector representation of the second text data; determining a quantitative similarity metric between the first vector representation and the second vector representation Gonsalves discloses: cropping a source image into second image data, wherein cropping is performed according to an aspect ratio of a first display of a target device, and displaying the second image data on the first display of the target device (An image is cropped to change the aspect ratio to fit a target display device. ¶¶ 5 and 30); generating, by inputting the first image data into an image captioning model, first text data representing a first description of first content of the first image data; generating, using a first text encoder, a first vector representation of the first text data (Text captions describing the content of the image are generated. See Gonsalves ¶¶20-21, and “Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks,” by Li, X. et al., (incorporated by reference into Gonsalves), page 7 and 11. After the captions are generated, the image and text are encoded into a vector representation. Gonsalves ¶¶20-21.); generating, by inputting the second image data into the image captioning model, second text data representing a second description of second content of the second image data; generating, using the first text encoder, a second vector representation of the second text data (The above-described process to generate captions and vector encodings is also used for the cropped images. Gonsalves ¶22.); determining a quantitative similarity metric between the first vector representation and the second vector representation (The encoded vector representations (i.e., the embeddings) of the cropped image and the source image are compared to determine similarity. Gonsalves ¶¶ 21-22. Similarity is determined as a cosine metric, which is a quantitative similarity metric. Id. See also “Term-weighting Approaches in Automatic Text Retrieval” by G. Salton and C. Buckley., incorporated by reference by Gonsalves.). It would have been obvious to a skilled artisan before the effective filing date of the claimed invention to modify Zhang with teachings in Gonsalves. The POSITA would have recognized that Zhang would be improved by scoring the candidate crop images based on the routines described in Gonsalves, which would enable the system to “find important objects within an image so that cropping could be performed more efficiently.” Gonsalves ¶3 (internal omissions removed). Claim 4 is rendered obvious by the Zhang-Gonsalves combination for reasons given above in the rejection of claim 1. This combination discloses a method comprising: receiving first image data representing a first image; receiving second image data representing a second image comprising a first subset of pixels of the first image data (Zhang ¶¶ 25-26. See above.); generating, using an image captioning model executed by at least one computing device, first text data describing the first image; generating, using the image captioning model, second text data describing the second image (Gonsalves ¶¶ 20-22. See above.); generating first data comprising a first structured data representation of the first text data; generating second data comprising a second structured data representation of the second text data (Gonsalves ¶¶ 20-22. See above.); determining a third data representing a quantitative similarity metric using the first data and the second data (Gonsalves ¶¶ 21-22.); and generating first computer-executable instructions configured to cause the second image data to be displayed on a display based at least in part on the third data (Zhang ¶¶ 18 and 29-30. See above.). Claim 11: Zhang-Gonsalves renders obvious: receiving a first input describing a first aspect ratio of an output display (The desired aspect ratio of the cropped image is known, thus this ratio is received. Gonsalves ¶¶ 30 and 32.); generating a plurality of cropped images from the first image data by iterating a window of the first aspect ratio over a plurality of positions overlaying the first image data, wherein each position of the plurality of positions corresponds to one of the plurality of cropped images (A candidate pool of cropped images is produced. Zhang ¶16. In a Zhang-Gonsalves combination, this candidate pool will be of the desired output aspect ratio.); generating, for a first cropped image of the plurality of cropped images, a first score representing a first degree of similarity between the first cropped image and the first image; generating, for a second cropped image of the plurality of cropped images, a second score representing a second degree of similarity between the second cropped image and the first image; and selecting the first cropped image from among the plurality of cropped images based on the first score and the second score (The multiple crop candidates are scored for similarity to the input image, and a first image is selected based on this score. Zhang ¶¶ 18 and 29-30.). Claim 12: Zhang-Gonsalves discloses that the third data represents a similarity between first content of the first image and second content of the second image (Zhang ¶29). Claim 13: see rejection of claim 4. Zhang-Gonsalves further discloses a system comprising: at least one processor; and at least one non-transitory computer-readable memory storing instructions that, when executed by the at least one processor, are effective to program the at least one processor to perform the method of claim 4 (see e.g. Zhang Fig. 7 and its description.). Claim 17: see rejection of claim 11. Claims 2, 3, 5, 7, 9, 14, 16, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang, Gonsalves, and Anderson et al., “SPICE: Semantic Propositional Image Caption Evaluation” (arXiv:1607.08822, publicly available on arXiv on 07/29/2016). Claim 2: the Zhang-Gonsalves combination does not disclose generating first graph data representing the first text data using a semantic propositional image caption evaluation (SPICE) model, wherein a node in the first graph data represents a word of the first text data; and generating second graph data representing the second text data using SPICE. But Anderson discloses generating first graph data representing the first text data using a semantic propositional image caption evaluation (SPICE) model, wherein a node in the first graph data represents a word of the first text data (Anderson §§ 2.2-3.1, pages 4-7). It would have been obvious to a skilled artisan before the effective filing date of the claimed invention to modify Zhang-Gonsalves with teachings found in Anderson, the rationale being to produce improved image captions. When modified in this manner, the POSITA would have found it likewise obvious to have the method include generating second graph data representing the second text data using SPICE, as Zhang-Gonsalves involves comparing first and second images using first and second sets of captions (see rejection of claim 1). Claim 3: the Zhang-Gonsalves-Anderson combination discloses: generating, by the first encoder, the first vector representation at least in part by generating a first embedding (Gonsalves ¶¶20-22) for a first node in the first graph data, the first node representing a first word (Anderson pg. 4-7); generating, by the first encoder, the second vector representation at least in part by generating a second embedding for a second node in the second graph data, the second node representing a second word (Gonsalves ¶¶20-22 and Anderson pg. 4-7. See also the rejection of claim 2, which described why generating multiple graphs for the multiple images would have been the natural result of modifying Zhang-Gonsalves with Anderson); and determining the quantitative similarity metric based at least in part by determining a cosine similarity between the first embedding and the second embedding (Gonsalves ¶¶ 21-22.). Claims 5 and 14: see rejection of claim 2. Claims 7 and 16: see rejection of claim 3. Claims 9 and 19: the Zhang-Gonsalves-Anderson combination renders obvious the steps of: receiving third image data representing a third image comprising a second subset of pixels of the first image data different from the first subset; generating, using the image captioning model, third text data describing the third image data; generating, using a first encoder and the third text data, fourth data representing the third text data; determining fifth data representing a degree of similarity between the first data and the fourth data; and selecting the second image data for output based on a comparison of the third data and the fifth data (Zhang discloses comparing multiple candidate cropped images to an input image to determine saliency scores, and comparing these scores to select a candidate. See Zhang Abstract. When Zhang is modified using teachings found in Gonsalves and Anderson, the invention that is recited in claim 9 is rendered obvious for reasons discussed in the rejection of claims 1, 2, and 4, above). Claims 6 and 15 rejected under 35 U.S.C. 103 as being unpatentable over Zhang, Gonsalves, and Anderson in view of Xie, US 20230394855. Claims 6 and 15: the Zhang-Gonsalves-Anderson combination fails to disclose: determining, using the first graph data and the second graph data, a first set of words present in the first graph data that are also present in the second graph data; and generating the third data based at least in part on the first set of words. Xie discloses determining, using first graph data and second graph data, a first set of words present in the first graph data that are also present in the second graph data; and generating third data based at least in part on the first set of words (Two sets of SPICE graphs are extracted and are compared to produce a comparison score. ¶¶ 29-30 and 46.). It would have been obvious to a skilled artisan before the effective filing date of the claimed invention to modify Zhang-Gonsalves-Anderson with teachings found in Xie, the rationale being to produce improved comparison scores for the two images. See Xie ¶29. While Xie relates to determining similarity between an image and a text description, the POSITA would have been motivated by this disclosure to compare the caption graphs of Zhang-Gonsalves-Anderson using this method, as this would have produced a more accurate determination of the similarity of the images. Claims 8 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang, Gonsalves, and Anderson in view of Gray, US 20100309225. Claims 8 and 18: the Zhang-Gonsalves-Anderson combination fails to disclose: determining first keypoints in the first image data using a keypoint detection model; determining second keypoints in the second image data using the keypoint detection model; determining a ratio of a number of the second keypoints to a number of the first keypoints; and determining the third data based at least in part on the ratio. Gray discloses: determining first keypoints in first image data using a keypoint detection model; determining second keypoints in second image data using the keypoint detection model; determining a ratio of a number of the second keypoints to a number of the first keypoints; and determining third data based at least in part on the ratio (A candidate image is compared with a query image to determine similarity score based on the number of matching keypoints. ¶31.). It would have been obvious to a skilled artisan before the effective filing date of the claimed invention to modify Zhang-Gonsalves-Anderson with teachings found in Gray, the rationale being to produce improved comparison scores for the two images. Claims 10 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang-Gonsalves-Anderson in view of Gandhi, US 20230394391. Claims 10 and 20: the Zhang-Gonsalves-Anderson combination discloses that the first data and the second data are generated using a first encoder (see rejection of claim 1). Zhang-Gonsalves-Anderson combination does not disclose: generating, using a second encoder executed by the at least one computing device, fourth data representing the first text data; generating, using the second encoder and the second text data, fifth data representing the second text data; determining sixth data representing a degree of similarity between the fourth data and the fifth data; inputting the third data and the sixth data into a neural network; and generating, by the neural network, output data indicating a semantic similarity between the first image data and the second image data. However, Gandhi discloses using a first and second encoder on first and second extracted text to generate data indicative of the similarity between the first and second text (Different word embedding models are used on two sets of text to produce a similarity score. ¶¶ 13-15 and 30.). This teaching in Gandi would have suggested to the POSITA to modify Zhang-Gonsalves-Anderson to arrive at the invention that is recited in claim 10. Gandi provides clear motivation for using multiple word embedding models when determining a similarity score between two sets of text: doing so produces more accurate similarity scores. See Gandhi ¶30. When modifying Zhang-Gonsalves-Anderson combination in this manner, the POSITA would have found it obvious to produce the claimed fourth, fifth, and sixth data using a different word embedding model than was used to generate the first, second, and third data, as this would have improved accuracy. Zhang-Gonsalves-Anderson-Gandhi does not disclose inputting the third data and the sixth data into a neural network; and generating, by the neural network, output data indicating a semantic similarity between the first image data and the second image data. However, official notice is taken that this was well known in the art before the effective filing date of the claimed invention. For example, it was widely practiced to use a neural network to analyze various items of input that describe the same data, and to output a result of that analysis. Therefore it would have been obvious to the POSITA to modify Zhang-Gonsalves-Anderson-Gandhi to do this, the rationale being to produce a more accurate final similarity score relating to the two images. See also Gandhi ¶41, which would have provided further suggestion to modify Zhang-Gonsalves-Anderson-Gandhi to use a neural network in determining the semantic similarity score. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ROBERT J HANCE whose telephone number is (571)270-5319. The examiner can normally be reached M-F 11:00am-7:00pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael Fuelling can be reached at (571) 270-1367. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ROBERT J HANCE/ Reexamination Specialist, Art Unit 3992 Conferees: /CHARLES R CRAVER/ Reexamination Specialist, Art Unit 3992 /M.F/Supervisory Patent Examiner, Art Unit 3992 1 Because the images “and text” are encoded, Gonsalves describes a “text encoder” as required in the amended claims.
Read full office action

Prosecution Timeline

Oct 31, 2025
Application Filed
Jun 05, 2026
Non-Final Rejection mailed — §103
Jun 11, 2026
Interview Requested
Jun 26, 2026
Examiner Interview Summary
Jul 15, 2026
Response Filed
Aug 06, 2026
Final Rejection mailed — §103
Sep 12, 2026
Interview Requested
Sep 21, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent RE51040
COMMUNICATION TECHNIQUES INVOLVING PAIRWISE ORTHOGONALITY OF ADJACENT ROWS IN LPDC CODE
1y 4m to grant Granted Sep 22, 2026
Patent RE51038
MOTION VECTOR CALCULATION METHOD, PICTURE CODING METHOD, PICTURE DECODING METHOD, MOTION VECTOR CALCULATION APPARATUS, AND PICTURE CODING AND DECODING APPARATUS
10m to grant Granted Sep 15, 2026
Patent RE51020
SYSTEM CONTROLLER FOR SERIES HYBRID POWERTRAIN
2y 8m to grant Granted Sep 01, 2026
Patent RE51022
MOTION VECTOR CALCULATION METHOD, PICTURE CODING METHOD, PICTURE DECODING METHOD, MOTION VECTOR CALCULATION APPARATUS, AND PICTURE CODING AND DECODING APPARATUS
11m to grant Granted Sep 01, 2026
Patent RE50999
PREDICTING RESPONSE TO IMMUNOTHERAPY USING COMPUTER EXTRACTED FEATURES RELATING TO SPATIAL ARRANGEMENT OF TUMOR INFILTRATING LYMPHOCYTES IN NON-SMALL CELL LUNG CANCER
3y 6m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
66%
Grant Probability
88%
With Interview (+21.5%)
2y 10m (~1y 11m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 761 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month