Prosecution Insights
Last updated: August 15, 2026
Application No. 18/938,587

DOCUMENT RETRIEVAL USING INTRA-IMAGE RELATIONSHIPS

Final Rejection §103
Filed
Nov 06, 2024
Priority
Apr 16, 2021 — provisional 63/175,972 +1 more
Examiner
GIULIANI, GIUSEPPI J
Art Unit
2153
Tech Center
2100 — Computer Architecture & Software
Assignee
Georgetown University
OA Round
2 (Final)
58%
Grant Probability
Moderate
3-4
OA Rounds
1y 7m
Est. Remaining
64%
With Interview

Examiner Intelligence

Grants 58% of resolved cases
58%
Career Allowance Rate
167 granted / 289 resolved
+2.8% vs TC avg
Moderate +7% lift
Without
With
+6.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
12 currently pending
Career history
314
Total Applications
across all art units

Statute-Specific Performance

§101
13.0%
-27.0% vs TC avg
§103
54.8%
+14.8% vs TC avg
§102
11.7%
-28.3% vs TC avg
§112
14.0%
-26.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 289 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Remarks This action is in response to the applicant’s response filed 26 May 2026, which is in response to the USPTO office action mailed 25 February 2026. Claims 1, 7, 8 and 15 are amended. Claims 3, 14 and 17 are cancelled. Claims 1, 2, 4-13, 15, 16 and 18-20 are currently pending. Response to Arguments With respect to the 35 USC §103 rejections of claims 1-20, the applicant’s arguments are moot in view of a new grounds of rejection, as necessitated by the applicant's amendments. Terminal Disclaimer The terminal disclaimer filed on 26 May 2026 disclaiming the terminal portion of any patent granted on this application which would extend beyond the expiration date of US 12,169,518 B2 has been reviewed and is accepted. The terminal disclaimer has been recorded. Allowable Subject Matter The following is a statement of reasons for the indication of allowable subject matter: The prosecution history and applicant's reply make evident reasons for allowance, satisfying the record "record as a whole" as required by rule 37 CFR 1.104(e). In this case, the substance of applicant's remarks and the amendments herein made to the claims clarifying the claimed invention indicate the reasons claims are patentable over the prior art of record. Further search and consideration was given, however no prior art resulting from the search was found which could fairly teach or suggest the claimed invention. Accordingly, the claims 8-13 are allowed. Claim Objections Claims 6 and 20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 2, 4, 5, 7, 15, 16, 18 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Jin et al., US 20210319243 A1 (hereinafter “Jin” – as cited in the IDS filed 6 November 2024) in view of Mosseri et al., US 20210342386 A1 (hereinafter “Mosseri”) in further view of Liu et al., US 2021/0049202 A1 (hereinafter “Liu” – as cited in the IDS filed 6 November 2024). Claim 1: Jin teaches a method, performed by one or more computing devices, the method comprising: obtaining an image representation comprising two or more segmentations of an image (Jin, [0042] note a user may specify a to-be-processed image (or target image) by using a terminal device (the smart phone 101, the tablet computer 102, or the portable computer 103 shown in FIG. 1). For example, the user transmits a target image to the server 105 by using the terminal device, [0043] note a plurality of target regions… a non-salient region in the image may be weakened, and a salient region in the image may be highlighted); generating one or more representations for each of a first segmentation and a second segmentation of the two or more segmentations (Jin, [0043] note after determining the target image, the server 105 may extract a feature map of the target image. For example, the server may extract a feature map of the target image by using any convolutional layer in a convolutional neural network (CNN) model. After the feature map of the target image is extracted, the feature map may be divided into a plurality of target regions); generating first one or more feature vectors for the first segmentation and second one or more feature vectors for the second segmentation based on the one or more representations (Jin, [0061] note generating a feature vector of a target image according to weights of target regions and feature vectors of the target regions); determining that the first segmentation is a dominant segmentation in the image representation and the second segmentation is a non-dominant segmentation in the image representation (Jin, [0043] note a non-salient region in the image… a salient region in the image; i.e. salient reads on dominant and non-salient reads on non-dominant); assigning a greater weight to the first one or more feature vectors and a lower weight to the second one or more feature vectors to generate a weighted set of feature vectors (Jin, [0043] note the target regions may be weighted according to the feature vectors of the target regions in the image, so that a non-salient region in the image may be weakened, and a salient region in the image may be highlighted, thereby effectively improving accuracy and appropriateness (or quality) of the generated feature vector of the image, and improving an image processing effect, for example, improving an image retrieval effect and accuracy in image recognition); comparing the weighted set of feature vectors to one or more other feature vectors, the one or more other feature vectors generated from one or more representations based on one or more other image representations (Jin, [0066] note after the feature vector of the target image is obtained, an image matching the target image may be retrieved according to the feature vector of the target image, [0097] note image retrieval model may further include a similarity determining module, configured to determine a similarity between different images based on feature vectors of the images, to determine similar images based on the similarity); and retrieving, based on the comparing of the weighted set of feature vectors to one or more other feature vectors, at least one image representation from the one or more other image representations that has a greater similarity to the dominant segmentation than the non-dominant segmentation (Jin, [Fig. 10], [0102] note after feature vectors of images (or to-be-retrieved images) are extracted according to the technical solution of an example embodiment of the disclosure, retrieval may be performed according to the extracted feature vectors, and then retrieved images are sequentially returned in descending order based on similarity). Jin does not explicitly teach latent space; wherein the determining that the first segmentation is the dominant segmentation is based at least on determining that the second segmentation is a text element and the first segmentation is a non-text element; and based at least on the determining that the second segmentation is the text element and is the non-dominant segmentation and the first segmentation is the non-text element and is the dominant segmentation. However, Mosseri teaches wherein the determining that the first segmentation is the dominant segmentation is based at least on determining that the second segmentation is a text element and the first segmentation is a non-text element; and based at least on the determining that the second segmentation is the text element and is the non-dominant segmentation and the first segmentation is the non-text element and is the dominant segmentation (Mosseri, [0008] note searching for similarity between digital visual objects, [0018] note digital visual objects are images of trademarks, [0037] note aspects of similarity may involve one or more of the following search queries, [0039] note 2. Image/pixel similarity—This similarity aspect is responsible for catching structural similarities of the images. For example, according to such similarity aspect, the images shown in FIGS. 1A and 1B may be considered similar by the system, [Fig. 2] note image/pixel similarity 24, [0044] note performing a search query (bloc 22) for each one of the above four aspects (blocs 23-26), [Fig. 5] note 54, 55, [0067] note FIG. 5 shows a search conducted for an “Ankori” sign 50… the search results by each category are indicated as follows: results by tags 53, results by image 54, results by text 55). It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the similarity between different images of Jin with the image/pixel similarity of Mosseri according to known methods (i.e. determining the similarity between images based on an image/pixel similarity responsible for catching structural similarities of the images). Motivation for doing so is that provides an improvement in computer tools (in particular a computer based similarity search engine) that may save trademark examiners effort or time or otherwise improve their performance is required (Mosseri, [0004]). Jin and Mosseri do not explicitly teach latent space. However, Liu teaches this (Liu, [0020] note images are initially mapped to base descriptors that characterize the images as vectors in a latent space. The first layer of the graph neural network may be configured to receive base descriptors as input. The base descriptors for images in the repository and the query image may be generated by applying a machine-learned model, such as an artificial neural network model (ANN), a convolutional neural network model (CNN), or other models that are configured to map an image to a base descriptor in the latent space such that images with similar content are closer to each other in the latent space, [0051] note content retrieval system 130 identifies 510 images relevant to the query image by selecting a relevant subset of image nodes. The image descriptors for the relevant subset of image nodes have above a similarity threshold with the query descriptor. The content retrieval system 130 returns 512 the images represented by the relevant subset of image nodes as a query result to the client device). It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the feature map of Jin and Mosseri with the latent space of Liu according to known methods (i.e. characterizing images in a latent space). Motivation for doing so is that can effectively learn a new descriptor space that improves retrieval accuracy while maintaining computational efficiency (Liu, [0007]). Claim 2: Jin, Mosseri and Liu teach the method of claim 1, further comprising analyzing a visual focus of each of the first segmentation and the second segmentation within the image representation, wherein the determining that the first segmentation is the dominant segmentation is based at least on the analyzing of the visual focus of each of the first segmentation and the second segmentation within the image representation (Jin, [0108] note the division unit 1104 is configured to: divide the feature map in a predetermined region division manner, to obtain the plurality of target regions; or perform an ROI pooling operation on the feature map, to map ROIs to the feature map to obtain the plurality of target regions, [0009] note a region of interest (ROI) pooling operation). Claim 4: Jin, Mosseri and Liu teach the method of claim 1, further comprising analyzing a location of each of the first segmentation and the second segmentation within the image representation, wherein the determining that the first segmentation is the dominant segmentation is based at least on the analyzing of the location of each of the first segmentation and the second segmentation within the image representation (Jin, [0089] note the image may be divided in the three manners shown in FIG. 7, part (1) to part (3), to obtain 14 regions R1 to R14. Then a max pooling operation is performed in each region according to a coordinate position of the each region, to determine a feature vector v of the each region). Claim 5: Jin, Mosseri and Liu teach the method of claim 1, further comprising: analyzing a classification of each of the first segmentation and the second segmentation within the image representation; and annotating the first segmentation with a first classification and the second segmentation with a second classification, wherein the determining that the first segmentation is the dominant segmentation is based at least on the first classification (Jin, [0099] note FIG. 9 shows a schematic diagram of weights of regions in an image according to an embodiment of the disclosure. For ease of description of the effects of the technical solution of an example embodiment of the disclosure, weights of the regions are annotated in the image, illustratively shown in FIG. 9. A box denoted as “GT” shown in FIG. 9 represents a region in which a salient object is located in each image. It can be seen from FIG. 9 that a weight of a region including a salient object is generally relatively large, and a weight of a region including no salient object is generally relatively small. In this way, a feature of a foreground region may be strengthened, and a feature of a background region may be weakened, thereby implementing more appropriate and more accurate image feature encoding, and greatly improving image retrieval performance). Claim 7: Jin, Mosseri and Liu teach the method of claim 1, wherein the at least one image representation from the one or more other image representations comprises a set of external search results that is ranked in order of similarity to the dominant segmentation (Jin, [Fig. 10], [0102] note after feature vectors of images (or to-be-retrieved images) are extracted according to the technical solution of an example embodiment of the disclosure, retrieval may be performed according to the extracted feature vectors, and then retrieved images are sequentially returned in descending order based on similarity). Claim 15: Jin teaches one or more computing devices comprising: one or more processors; and memory having a plurality of computer-executable instructions stored thereon; wherein the computer-executable instructions are configured to, when executed by the one or more processors, cause the one or more computing devices to perform a plurality of operations, the operations comprising: obtaining an image representation comprising two or more segmentations of an image (Jin, [0042] note a user may specify a to-be-processed image (or target image) by using a terminal device (the smart phone 101, the tablet computer 102, or the portable computer 103 shown in FIG. 1). For example, the user transmits a target image to the server 105 by using the terminal device, [0043] note a plurality of target regions… a non-salient region in the image may be weakened, and a salient region in the image may be highlighted); generating one or more representations for each of a first segmentation and a second segmentation of the two or more segmentations (Jin, [0043] note after determining the target image, the server 105 may extract a feature map of the target image. For example, the server may extract a feature map of the target image by using any convolutional layer in a convolutional neural network (CNN) model. After the feature map of the target image is extracted, the feature map may be divided into a plurality of target regions); generating first one or more feature vectors for the first segmentation and second one or more feature vectors for the second segmentation based on the one or more representations (Jin, [0061] note generating a feature vector of a target image according to weights of target regions and feature vectors of the target regions); determining that the first segmentation is a dominant segmentation in the image representation and the second segmentation is a non-dominant segmentation in the image representation (Jin, [0043] note a non-salient region in the image… a salient region in the image; i.e. salient reads on dominant and non-salient reads on non-dominant); assigning a greater weight to the first one or more feature vectors and a lower weight to the second one or more feature vectors to generate a weighted set of feature vectors (Jin, [0043] note the target regions may be weighted according to the feature vectors of the target regions in the image, so that a non-salient region in the image may be weakened, and a salient region in the image may be highlighted, thereby effectively improving accuracy and appropriateness (or quality) of the generated feature vector of the image, and improving an image processing effect, for example, improving an image retrieval effect and accuracy in image recognition); comparing the weighted set of feature vectors to one or more other feature vectors, the one or more other feature vectors generated from one or more representations based on one or more other image representations (Jin, [0066] note after the feature vector of the target image is obtained, an image matching the target image may be retrieved according to the feature vector of the target image, [0097] note image retrieval model may further include a similarity determining module, configured to determine a similarity between different images based on feature vectors of the images, to determine similar images based on the similarity); and retrieving, based on the comparing of the weighted set of feature vectors to one or more other feature vectors, at least one image representation from the one or more other image representations that has a greater similarity to the dominant segmentation than the non-dominant segmentation (Jin, [Fig. 10], [0102] note after feature vectors of images (or to-be-retrieved images) are extracted according to the technical solution of an example embodiment of the disclosure, retrieval may be performed according to the extracted feature vectors, and then retrieved images are sequentially returned in descending order based on similarity). Jin does not explicitly teach latent space; wherein the determining that the first segmentation is the dominant segmentation and the second segmentation is a non-dominant segmentation in the image representation is based at least on determining that the second segmentation is a text element and the first segmentation is a non-text element; based at least on the determining that the second segmentation is the text element and is the non-dominant segmentation and the first segmentation is the non-text element and is the dominant segmentation. However, Mosseri teaches wherein the determining that the first segmentation is the dominant segmentation and the second segmentation is a non-dominant segmentation in the image representation is based at least on determining that the second segmentation is a text element and the first segmentation is a non-text element; based at least on the determining that the second segmentation is the text element and is the non-dominant segmentation and the first segmentation is the non-text element and is the dominant segmentation (Mosseri, [0008] note searching for similarity between digital visual objects, [0018] note digital visual objects are images of trademarks, [0037] note aspects of similarity may involve one or more of the following search queries, [0039] note 2. Image/pixel similarity—This similarity aspect is responsible for catching structural similarities of the images. For example, according to such similarity aspect, the images shown in FIGS. 1A and 1B may be considered similar by the system, [Fig. 2] note image/pixel similarity 24, [0044] note performing a search query (bloc 22) for each one of the above four aspects (blocs 23-26), [Fig. 5] note 54, 55, [0067] note FIG. 5 shows a search conducted for an “Ankori” sign 50… the search results by each category are indicated as follows: results by tags 53, results by image 54, results by text 55). It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the similarity between different images of Jin with the image/pixel similarity of Mosseri according to known methods (i.e. determining the similarity between images based on an image/pixel similarity responsible for catching structural similarities of the images). Motivation for doing so is that provides an improvement in computer tools (in particular a computer based similarity search engine) that may save trademark examiners effort or time or otherwise improve their performance is required (Mosseri, [0004]). Jin and Mosseri do not explicitly teach latent space. However, Liu teaches this (Liu, [0020] note images are initially mapped to base descriptors that characterize the images as vectors in a latent space. The first layer of the graph neural network may be configured to receive base descriptors as input. The base descriptors for images in the repository and the query image may be generated by applying a machine-learned model, such as an artificial neural network model (ANN), a convolutional neural network model (CNN), or other models that are configured to map an image to a base descriptor in the latent space such that images with similar content are closer to each other in the latent space, [0051] note content retrieval system 130 identifies 510 images relevant to the query image by selecting a relevant subset of image nodes. The image descriptors for the relevant subset of image nodes have above a similarity threshold with the query descriptor. The content retrieval system 130 returns 512 the images represented by the relevant subset of image nodes as a query result to the client device). It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the feature map of Jin and Mosseri with the latent space of Liu according to known methods (i.e. characterizing images in a latent space). Motivation for doing so is that can effectively learn a new descriptor space that improves retrieval accuracy while maintaining computational efficiency (Liu, [0007]). Claim 16: Jin, Mosseri and Liu teach the one or more computing devices of claim 15, wherein the operations further comprise analyzing a visual focus of each of the first segmentation and the second segmentation within the image representation, wherein the determining that the first segmentation is the dominant segmentation is based at least on the analyzing of the visual focus of each of the first segmentation and the second segmentation within the image representation (Jin, [0108] note the division unit 1104 is configured to: divide the feature map in a predetermined region division manner, to obtain the plurality of target regions; or perform an ROI pooling operation on the feature map, to map ROIs to the feature map to obtain the plurality of target regions, [0009] note a region of interest (ROI) pooling operation). Claim 18: Jin, Mosseri and Liu teach the one or more computing devices of claim 15, wherein the operations further comprise analyzing a location of each of the first segmentation and the second segmentation within the image representation, wherein the determining that the first segmentation is the dominant segmentation is based at least on the analyzing of the location of each of the first segmentation and the second segmentation within the image representation (Jin, [0089] note the image may be divided in the three manners shown in FIG. 7, part (1) to part (3), to obtain 14 regions R1 to R14. Then a max pooling operation is performed in each region according to a coordinate position of the each region, to determine a feature vector v of the each region). Claim 19: Jin, Mosseri and Liu teach the one or more computing devices of claim 15, wherein the operations further comprise: analyzing a classification of each of the first segmentation and the second segmentation within the image representation; and annotating the first segmentation with a first classification and the second segmentation with a second classification, wherein the determining that the first segmentation is the dominant segmentation is based at least on the first classification (Jin, [0099] note FIG. 9 shows a schematic diagram of weights of regions in an image according to an embodiment of the disclosure. For ease of description of the effects of the technical solution of an example embodiment of the disclosure, weights of the regions are annotated in the image, illustratively shown in FIG. 9. A box denoted as “GT” shown in FIG. 9 represents a region in which a salient object is located in each image. It can be seen from FIG. 9 that a weight of a region including a salient object is generally relatively large, and a weight of a region including no salient object is generally relatively small. In this way, a feature of a foreground region may be strengthened, and a feature of a background region may be weakened, thereby implementing more appropriate and more accurate image feature encoding, and greatly improving image retrieval performance). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Giuseppi Giuliani whose telephone number is (571)270-7128. The examiner can normally be reached Monday-Friday. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kavita Stanley can be reached at (571)272-8352. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /GIUSEPPI GIULIANI/Primary Examiner, Art Unit 2153
Read full office action

Prosecution Timeline

Nov 06, 2024
Application Filed
Feb 25, 2026
Non-Final Rejection mailed — §103
May 26, 2026
Response Filed
Jul 08, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12675463
APPARATUSES, SYSTEMS, AND METHODS FOR PROVIDING AN EVENT MANAGEMENT FRAMEWORK FOR A GEOGRAPHIC INFORMATION SYSTEM
3y 5m to grant Granted Jul 07, 2026
Patent 12675393
EFFICIENT BURST SORT BASED ON NETWORK-ATTACHED MEMORY
2y 0m to grant Granted Jul 07, 2026
Patent 12639293
SYSTEMS AND METHODS FOR PAGINATING SEARCH RESULTS RETRIEVED FROM DATABASES THAT SUPPORT CURSOR-BASED PAGINATION
3y 4m to grant Granted May 26, 2026
Patent 12632499
SYSTEMS AND METHODS FOR USING GRAPH DATA STRUCTURES
1y 8m to grant Granted May 19, 2026
Patent 12613916
SYSTEMS AND METHODS TO INCREASE VIEWERSHIP OF ONLINE CONTENT
5y 0m to grant Granted Apr 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
58%
Grant Probability
64%
With Interview (+6.6%)
3y 5m (~1y 7m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 289 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month