DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Remarks
This action is in response to the applicant’s response filed 26 May 2026, which is in response to the USPTO office action mailed 25 February 2026. Claims 1, 7, 8 and 15 are amended. Claims 3, 14 and 17 are cancelled. Claims 1, 2, 4-13, 15, 16 and 18-20 are currently pending.
Response to Arguments
With respect to the 35 USC §103 rejections of claims 1-20, the applicant’s arguments are moot in view of a new grounds of rejection, as necessitated by the applicant's amendments.
Terminal Disclaimer
The terminal disclaimer filed on 26 May 2026 disclaiming the terminal portion of any patent granted on this application which would extend beyond the expiration date of US 12,169,518 B2 has been reviewed and is accepted. The terminal disclaimer has been recorded.
Allowable Subject Matter
The following is a statement of reasons for the indication of allowable subject matter:
The prosecution history and applicant's reply make evident reasons for allowance, satisfying the record "record as a whole" as required by rule 37 CFR 1.104(e). In this case, the substance of applicant's remarks and the amendments herein made to the claims clarifying the claimed invention indicate the reasons claims are patentable over the prior art of record. Further search and consideration was given, however no prior art resulting from the search was found which could fairly teach or suggest the claimed invention. Accordingly, the claims 8-13 are allowed.
Claim Objections
Claims 6 and 20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 2, 4, 5, 7, 15, 16, 18 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Jin et al., US 20210319243 A1 (hereinafter “Jin” – as cited in the IDS filed 6 November 2024) in view of Mosseri et al., US 20210342386 A1 (hereinafter “Mosseri”) in further view of Liu et al., US 2021/0049202 A1 (hereinafter “Liu” – as cited in the IDS filed 6 November 2024).
Claim 1: Jin teaches a method, performed by one or more computing devices, the method comprising:
obtaining an image representation comprising two or more segmentations of an image (Jin, [0042] note a user may specify a to-be-processed image (or target image) by using a terminal device (the smart phone 101, the tablet computer 102, or the portable computer 103 shown in FIG. 1). For example, the user transmits a target image to the server 105 by using the terminal device, [0043] note a plurality of target regions… a non-salient region in the image may be weakened, and a salient region in the image may be highlighted);
generating one or more representations for each of a first segmentation and a second segmentation of the two or more segmentations (Jin, [0043] note after determining the target image, the server 105 may extract a feature map of the target image. For example, the server may extract a feature map of the target image by using any convolutional layer in a convolutional neural network (CNN) model. After the feature map of the target image is extracted, the feature map may be divided into a plurality of target regions);
generating first one or more feature vectors for the first segmentation and second one or more feature vectors for the second segmentation based on the one or more representations (Jin, [0061] note generating a feature vector of a target image according to weights of target regions and feature vectors of the target regions);
determining that the first segmentation is a dominant segmentation in the image representation and the second segmentation is a non-dominant segmentation in the image representation (Jin, [0043] note a non-salient region in the image… a salient region in the image; i.e. salient reads on dominant and non-salient reads on non-dominant);
assigning a greater weight to the first one or more feature vectors and a lower weight to the second one or more feature vectors to generate a weighted set of feature vectors (Jin, [0043] note the target regions may be weighted according to the feature vectors of the target regions in the image, so that a non-salient region in the image may be weakened, and a salient region in the image may be highlighted, thereby effectively improving accuracy and appropriateness (or quality) of the generated feature vector of the image, and improving an image processing effect, for example, improving an image retrieval effect and accuracy in image recognition);
comparing the weighted set of feature vectors to one or more other feature vectors, the one or more other feature vectors generated from one or more representations based on one or more other image representations (Jin, [0066] note after the feature vector of the target image is obtained, an image matching the target image may be retrieved according to the feature vector of the target image, [0097] note image retrieval model may further include a similarity determining module, configured to determine a similarity between different images based on feature vectors of the images, to determine similar images based on the similarity); and
retrieving, based on the comparing of the weighted set of feature vectors to one or more other feature vectors, at least one image representation from the one or more other image representations that has a greater similarity to the dominant segmentation than the non-dominant segmentation (Jin, [Fig. 10], [0102] note after feature vectors of images (or to-be-retrieved images) are extracted according to the technical solution of an example embodiment of the disclosure, retrieval may be performed according to the extracted feature vectors, and then retrieved images are sequentially returned in descending order based on similarity).
Jin does not explicitly teach latent space; wherein the determining that the first segmentation is the dominant segmentation is based at least on determining that the second segmentation is a text element and the first segmentation is a non-text element; and based at least on the determining that the second segmentation is the text element and is the non-dominant segmentation and the first segmentation is the non-text element and is the dominant segmentation.
However, Mosseri teaches wherein the determining that the first segmentation is the dominant segmentation is based at least on determining that the second segmentation is a text element and the first segmentation is a non-text element; and based at least on the determining that the second segmentation is the text element and is the non-dominant segmentation and the first segmentation is the non-text element and is the dominant segmentation (Mosseri, [0008] note searching for similarity between digital visual objects, [0018] note digital visual objects are images of trademarks, [0037] note aspects of similarity may involve one or more of the following search queries, [0039] note 2. Image/pixel similarity—This similarity aspect is responsible for catching structural similarities of the images. For example, according to such similarity aspect, the images shown in FIGS. 1A and 1B may be considered similar by the system, [Fig. 2] note image/pixel similarity 24, [0044] note performing a search query (bloc 22) for each one of the above four aspects (blocs 23-26), [Fig. 5] note 54, 55, [0067] note FIG. 5 shows a search conducted for an “Ankori” sign 50… the search results by each category are indicated as follows: results by tags 53, results by image 54, results by text 55).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the similarity between different images of Jin with the image/pixel similarity of Mosseri according to known methods (i.e. determining the similarity between images based on an image/pixel similarity responsible for catching structural similarities of the images). Motivation for doing so is that provides an improvement in computer tools (in particular a computer based similarity search engine) that may save trademark examiners effort or time or otherwise improve their performance is required (Mosseri, [0004]).
Jin and Mosseri do not explicitly teach latent space.
However, Liu teaches this (Liu, [0020] note images are initially mapped to base descriptors that characterize the images as vectors in a latent space. The first layer of the graph neural network may be configured to receive base descriptors as input. The base descriptors for images in the repository and the query image may be generated by applying a machine-learned model, such as an artificial neural network model (ANN), a convolutional neural network model (CNN), or other models that are configured to map an image to a base descriptor in the latent space such that images with similar content are closer to each other in the latent space, [0051] note content retrieval system 130 identifies 510 images relevant to the query image by selecting a relevant subset of image nodes. The image descriptors for the relevant subset of image nodes have above a similarity threshold with the query descriptor. The content retrieval system 130 returns 512 the images represented by the relevant subset of image nodes as a query result to the client device).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the feature map of Jin and Mosseri with the latent space of Liu according to known methods (i.e. characterizing images in a latent space). Motivation for doing so is that can effectively learn a new descriptor space that improves retrieval accuracy while maintaining computational efficiency (Liu, [0007]).
Claim 2: Jin, Mosseri and Liu teach the method of claim 1, further comprising analyzing a visual focus of each of the first segmentation and the second segmentation within the image representation, wherein the determining that the first segmentation is the dominant segmentation is based at least on the analyzing of the visual focus of each of the first segmentation and the second segmentation within the image representation (Jin, [0108] note the division unit 1104 is configured to: divide the feature map in a predetermined region division manner, to obtain the plurality of target regions; or perform an ROI pooling operation on the feature map, to map ROIs to the feature map to obtain the plurality of target regions, [0009] note a region of interest (ROI) pooling operation).
Claim 4: Jin, Mosseri and Liu teach the method of claim 1, further comprising analyzing a location of each of the first segmentation and the second segmentation within the image representation, wherein the determining that the first segmentation is the dominant segmentation is based at least on the analyzing of the location of each of the first segmentation and the second segmentation within the image representation (Jin, [0089] note the image may be divided in the three manners shown in FIG. 7, part (1) to part (3), to obtain 14 regions R1 to R14. Then a max pooling operation is performed in each region according to a coordinate position of the each region, to determine a feature vector v of the each region).
Claim 5: Jin, Mosseri and Liu teach the method of claim 1, further comprising:
analyzing a classification of each of the first segmentation and the second segmentation within the image representation; and annotating the first segmentation with a first classification and the second segmentation with a second classification, wherein the determining that the first segmentation is the dominant segmentation is based at least on the first classification (Jin, [0099] note FIG. 9 shows a schematic diagram of weights of regions in an image according to an embodiment of the disclosure. For ease of description of the effects of the technical solution of an example embodiment of the disclosure, weights of the regions are annotated in the image, illustratively shown in FIG. 9. A box denoted as “GT” shown in FIG. 9 represents a region in which a salient object is located in each image. It can be seen from FIG. 9 that a weight of a region including a salient object is generally relatively large, and a weight of a region including no salient object is generally relatively small. In this way, a feature of a foreground region may be strengthened, and a feature of a background region may be weakened, thereby implementing more appropriate and more accurate image feature encoding, and greatly improving image retrieval performance).
Claim 7: Jin, Mosseri and Liu teach the method of claim 1, wherein the at least one image representation from the one or more other image representations comprises a set of external search results that is ranked in order of similarity to the dominant segmentation (Jin, [Fig. 10], [0102] note after feature vectors of images (or to-be-retrieved images) are extracted according to the technical solution of an example embodiment of the disclosure, retrieval may be performed according to the extracted feature vectors, and then retrieved images are sequentially returned in descending order based on similarity).
Claim 15: Jin teaches one or more computing devices comprising: one or more processors; and memory having a plurality of computer-executable instructions stored thereon; wherein the computer-executable instructions are configured to, when executed by the one or more processors, cause the one or more computing devices to perform a plurality of operations, the operations comprising:
obtaining an image representation comprising two or more segmentations of an image (Jin, [0042] note a user may specify a to-be-processed image (or target image) by using a terminal device (the smart phone 101, the tablet computer 102, or the portable computer 103 shown in FIG. 1). For example, the user transmits a target image to the server 105 by using the terminal device, [0043] note a plurality of target regions… a non-salient region in the image may be weakened, and a salient region in the image may be highlighted);
generating one or more representations for each of a first segmentation and a second segmentation of the two or more segmentations (Jin, [0043] note after determining the target image, the server 105 may extract a feature map of the target image. For example, the server may extract a feature map of the target image by using any convolutional layer in a convolutional neural network (CNN) model. After the feature map of the target image is extracted, the feature map may be divided into a plurality of target regions);
generating first one or more feature vectors for the first segmentation and second one or more feature vectors for the second segmentation based on the one or more representations (Jin, [0061] note generating a feature vector of a target image according to weights of target regions and feature vectors of the target regions);
determining that the first segmentation is a dominant segmentation in the image representation and the second segmentation is a non-dominant segmentation in the image representation (Jin, [0043] note a non-salient region in the image… a salient region in the image; i.e. salient reads on dominant and non-salient reads on non-dominant);
assigning a greater weight to the first one or more feature vectors and a lower weight to the second one or more feature vectors to generate a weighted set of feature vectors (Jin, [0043] note the target regions may be weighted according to the feature vectors of the target regions in the image, so that a non-salient region in the image may be weakened, and a salient region in the image may be highlighted, thereby effectively improving accuracy and appropriateness (or quality) of the generated feature vector of the image, and improving an image processing effect, for example, improving an image retrieval effect and accuracy in image recognition);
comparing the weighted set of feature vectors to one or more other feature vectors, the one or more other feature vectors generated from one or more representations based on one or more other image representations (Jin, [0066] note after the feature vector of the target image is obtained, an image matching the target image may be retrieved according to the feature vector of the target image, [0097] note image retrieval model may further include a similarity determining module, configured to determine a similarity between different images based on feature vectors of the images, to determine similar images based on the similarity); and
retrieving, based on the comparing of the weighted set of feature vectors to one or more other feature vectors, at least one image representation from the one or more other image representations that has a greater similarity to the dominant segmentation than the non-dominant segmentation (Jin, [Fig. 10], [0102] note after feature vectors of images (or to-be-retrieved images) are extracted according to the technical solution of an example embodiment of the disclosure, retrieval may be performed according to the extracted feature vectors, and then retrieved images are sequentially returned in descending order based on similarity).
Jin does not explicitly teach latent space; wherein the determining that the first segmentation is the dominant segmentation and the second segmentation is a non-dominant segmentation in the image representation is based at least on determining that the second segmentation is a text element and the first segmentation is a non-text element; based at least on the determining that the second segmentation is the text element and is the non-dominant segmentation and the first segmentation is the non-text element and is the dominant segmentation.
However, Mosseri teaches wherein the determining that the first segmentation is the dominant segmentation and the second segmentation is a non-dominant segmentation in the image representation is based at least on determining that the second segmentation is a text element and the first segmentation is a non-text element; based at least on the determining that the second segmentation is the text element and is the non-dominant segmentation and the first segmentation is the non-text element and is the dominant segmentation (Mosseri, [0008] note searching for similarity between digital visual objects, [0018] note digital visual objects are images of trademarks, [0037] note aspects of similarity may involve one or more of the following search queries, [0039] note 2. Image/pixel similarity—This similarity aspect is responsible for catching structural similarities of the images. For example, according to such similarity aspect, the images shown in FIGS. 1A and 1B may be considered similar by the system, [Fig. 2] note image/pixel similarity 24, [0044] note performing a search query (bloc 22) for each one of the above four aspects (blocs 23-26), [Fig. 5] note 54, 55, [0067] note FIG. 5 shows a search conducted for an “Ankori” sign 50… the search results by each category are indicated as follows: results by tags 53, results by image 54, results by text 55).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the similarity between different images of Jin with the image/pixel similarity of Mosseri according to known methods (i.e. determining the similarity between images based on an image/pixel similarity responsible for catching structural similarities of the images). Motivation for doing so is that provides an improvement in computer tools (in particular a computer based similarity search engine) that may save trademark examiners effort or time or otherwise improve their performance is required (Mosseri, [0004]).
Jin and Mosseri do not explicitly teach latent space.
However, Liu teaches this (Liu, [0020] note images are initially mapped to base descriptors that characterize the images as vectors in a latent space. The first layer of the graph neural network may be configured to receive base descriptors as input. The base descriptors for images in the repository and the query image may be generated by applying a machine-learned model, such as an artificial neural network model (ANN), a convolutional neural network model (CNN), or other models that are configured to map an image to a base descriptor in the latent space such that images with similar content are closer to each other in the latent space, [0051] note content retrieval system 130 identifies 510 images relevant to the query image by selecting a relevant subset of image nodes. The image descriptors for the relevant subset of image nodes have above a similarity threshold with the query descriptor. The content retrieval system 130 returns 512 the images represented by the relevant subset of image nodes as a query result to the client device).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the feature map of Jin and Mosseri with the latent space of Liu according to known methods (i.e. characterizing images in a latent space). Motivation for doing so is that can effectively learn a new descriptor space that improves retrieval accuracy while maintaining computational efficiency (Liu, [0007]).
Claim 16: Jin, Mosseri and Liu teach the one or more computing devices of claim 15, wherein the operations further comprise analyzing a visual focus of each of the first segmentation and the second segmentation within the image representation, wherein the determining that the first segmentation is the dominant segmentation is based at least on the analyzing of the visual focus of each of the first segmentation and the second segmentation within the image representation (Jin, [0108] note the division unit 1104 is configured to: divide the feature map in a predetermined region division manner, to obtain the plurality of target regions; or perform an ROI pooling operation on the feature map, to map ROIs to the feature map to obtain the plurality of target regions, [0009] note a region of interest (ROI) pooling operation).
Claim 18: Jin, Mosseri and Liu teach the one or more computing devices of claim 15, wherein the operations further comprise analyzing a location of each of the first segmentation and the second segmentation within the image representation, wherein the determining that the first segmentation is the dominant segmentation is based at least on the analyzing of the location of each of the first segmentation and the second segmentation within the image representation (Jin, [0089] note the image may be divided in the three manners shown in FIG. 7, part (1) to part (3), to obtain 14 regions R1 to R14. Then a max pooling operation is performed in each region according to a coordinate position of the each region, to determine a feature vector v of the each region).
Claim 19: Jin, Mosseri and Liu teach the one or more computing devices of claim 15, wherein the operations further comprise: analyzing a classification of each of the first segmentation and the second segmentation within the image representation; and annotating the first segmentation with a first classification and the second segmentation with a second classification, wherein the determining that the first segmentation is the dominant segmentation is based at least on the first classification (Jin, [0099] note FIG. 9 shows a schematic diagram of weights of regions in an image according to an embodiment of the disclosure. For ease of description of the effects of the technical solution of an example embodiment of the disclosure, weights of the regions are annotated in the image, illustratively shown in FIG. 9. A box denoted as “GT” shown in FIG. 9 represents a region in which a salient object is located in each image. It can be seen from FIG. 9 that a weight of a region including a salient object is generally relatively large, and a weight of a region including no salient object is generally relatively small. In this way, a feature of a foreground region may be strengthened, and a feature of a background region may be weakened, thereby implementing more appropriate and more accurate image feature encoding, and greatly improving image retrieval performance).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Giuseppi Giuliani whose telephone number is (571)270-7128. The examiner can normally be reached Monday-Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kavita Stanley can be reached at (571)272-8352. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/GIUSEPPI GIULIANI/Primary Examiner, Art Unit 2153