DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to s without significantly more.
With regard to claim 1:
Step 2A, Prong One:
The claim recites the following limitations which are drawn towards an abstract idea:
determine,
filter the image data based on the request to identify a filtered set of the image data (recites mental process steps of comparison and selection of information);
obtain a subset of the image embeddings corresponding to the filtered set of the image data (recites mental process steps of selection/judgment of content);
determine, based on a comparison of the search embeddings and the subset of the image embeddings, recommended image data (recites mental process steps of analyzing and comparing information to form a decision/judgement);
As seen from above, the identified limitations recite concepts associated with an abstract idea and thus the respective claim recites a judicial exception (see 2106.04(a)) and thus requires further analysis as discussed below.
Step 2A, Prong Two:
The following limitations have been identified as being additional elements as discussed below.
A system, comprising: a database (recites apply-it type limitations by reciting generic hardware elements to implement the abstract idea, see MPEP 2106.05(f)) including image data and image embeddings (recites field of use limitations describing the type of data that is being stored, see MPEP 2106.05(h)); a processor; and a non-transitory memory storing instructions, that when executed, cause the processor to:
receive, from a user device, a request for an image (recites insignificant extrasolution activity of receiving information, see MPEP 2106.05(g); with the user device reciting generic computer elements to implement the abstract idea, see MPEP 2106.05(f));
“…using a machine-learning model,…” (recites using the computer components of a computer as a tool to implement the abstract idea, see MPEP 2106.05(f)),
cause a presentation of the recommended image data at the user device (recites insignificant extrasolution activity of transmitting information such as displaying information, see MPEP 2106.05(g));
and in response to a selection of a recommended image from the recommended image data, provide the recommended image to the user device (recites insignificant extrasolution activity of transmitting information such as retrieving and displaying selected information, see MPEP 2106.05(g)).
As seen from the above discussion, the identified limitations did not integrate the judicial exception into a practical application (see MPEP 2106.04(d)). This judicial exception is not integrated into a practical application because the additional elements recite generic computer components to be used as a tool to implement the abstract idea including insignificant extrasolution activity functionality such as receiving and transmitting information.
Step 2B:
Below is the analysis of the claims:
A system, comprising: a database (recites apply-it type limitations by reciting generic hardware elements to implement the abstract idea, see MPEP 2106.05(f)) including image data and image embeddings (recites field of use limitations describing the type of data that is being stored, see MPEP 2106.05(h)); a processor; and a non-transitory memory storing instructions, that when executed, cause the processor to:
receive, from a user device, a request for an image (recites insignificant extrasolution activity of receiving information, see MPEP 2106.05(g); with the user device reciting generic computer elements to implement the abstract idea, see MPEP 2106.05(f));
“…using a machine-learning model,…” (recites using the computer components of a computer as a tool to implement the abstract idea, see MPEP 2106.05(f)),
cause a presentation of the recommended image data at the user device (recites insignificant extrasolution activity of transmitting information such as displaying information, see MPEP 2106.05(g));
and in response to a selection of a recommended image from the recommended image data, provide the recommended image to the user device (recites insignificant extrasolution activity of transmitting information such as retrieving and displaying selected information, see MPEP 2106.05(g)).
As seen from above, the respective claim elements taken individually do not amount to significantly more than the judicial exception. When taken as a whole (in combination), the claim also does not amount to significantly more than the abstract idea because the additional elements recite generic computer components to be used as a tool to implement the abstract idea including well-understood, routine, and conventional activity such as receiving and transmitting information.
With regard to claim 2, this claim recites wherein the request comprises at least one of: a text portion, an image portion, or a campaign related portion (recites field of use limitations describing the data type of that is received, see MPEP 2106.05(h)).
With regard to claim 3, this claim recites wherein the instructions, when executed, cause the processor to determine the search embeddings based at least by: generating, using a text query encoder of the machine-learning model, at least one text query embedding based on the text portion; generating, using an image query encoder of the machine-learning model, at least one image query embedding based on the image portion; and determining the search embeddings based on the at least one text query embedding and the at least one image query embedding (recites mental process steps of converting textual content to another format/vector and also converting an image to a vector format and then determining/generating a vector that can be the combination of the two other vectors/embeddings such as concatenating vectors together).
With regard to claim 4, this claim recites wherein the instructions, when executed, further cause the processor to train the machine-learning model based at least by: training the machine-learning model using a first set of tasks having a first complexity (recites apply-it type limitations of using the computer as a tool to implement the abstract idea, in particular the setup of a computer tool/computer model recited at a high level of generality, see MPEP 2106.05(f));
and re-training the machine-learning model using a second set of tasks having a second complexity greater than the first complexity, after completion of the first set of tasks (recites apply-it type limitations of using the computer as a tool to implement the abstract idea, in particular the setup of a computer tool/computer model recited at a high level of generality; see applicant’s specification at paragraph 59 to indicate second set of tasks relate to category classification task, see MPEP 2106.05(f)).
With regard to claim 5, this claim recites wherein: the first set of tasks includes training data related to one or more image-text pairs each of which is formed by an image and a text corresponding to the image; the machine-learning model is trained using the first set of tasks based on a cross entropy loss of image embeddings and text embeddings (recites field of use limitations describing the particular data formats being used as well as the type of loss function being used at a high-level of generality, see MPEP 2106.05(h));
the second set of tasks includes training data related to one or more image-category pairs each of which is formed by an image of a corresponding product and a category of the corresponding product; and the machine-learning model is re-trained using the second set of tasks based on a cross entropy loss of image embeddings and category embeddings (recites field of use limitations describing the particular data formats being used as well as the type of loss function being used at a high-level of generality, see MPEP 2106.05(h)).
With regard to claim 6, this claim recites wherein the instructions, when executed, cause the processor to filter the image data based by at least one of: identifying and excluding circular images from the filtered set of the image data using a circular filter based on the request, wherein each of the circular images has a circular shape; identifying and excluding text-heavy images from the filtered set of the image data using a text filter based on the request, wherein each text-heavy image of the text-heavy images has a text portion occupying more than half of the text-heavy image; or identifying and excluding duplicated images from the filtered set of the image data using a duplication filter based on the request, wherein each of the duplicated images has a hash vector based similarity score higher than a predetermined threshold with respect to an existing image in the filtered set of the image data (recites insignificant extrasolution activity of sorting information which amounts to well-understood, routine, and conventional activity of sorting/filtering information, see MPEP 2106.05(d)).
With regard to claim 7, this claim recites wherein the instructions, when executed, cause the processor to determine the recommended image data based at least by: comparing the search embeddings with the subset of the image embeddings to compute cosine similarity distances; generate ranking scores for the subset of the image embeddings based on the cosine similarity distances (recites mental process steps of evaluation/comparisons and mathematical vector calculations to generate a value/result);
selecting one or more image embeddings having highest ranking scores among the subset of the image embeddings, and determining the recommended image data corresponding to the one or more image embeddings (recites mental process steps of evaluation and decision/judgement making).
With regard to claims 8-14, these claims are substantially similar to claims 1-7 respectively and are rejected for similar reasons as discussed above.
With regard to claims 15-20, these claims are substantially similar to claims 1 and 3-7 respectively and are rejected for similar reasons as discussed above.
Claims 15-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter because the claims are directed towards signals per se. The applicant’s specification at paragraph [0037] indicates that a non-transitory computer-readable medium can be anything from an open-list of examples including “any other suitable memory” where it is unclear what the metes and bounds of “suitable memory” encompasses and could encompass non-statutory embodiments such as signals. Therefore, the claims are being rejected for being directed towards signals per se.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-4, 7-11, 14-17, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Kotov et al [US 2022/0138247 A1] in view of Zhang et al [US 2023/0245418 A1] and Huynh et al [US 2015/0269231 A1].
With regard to claim 1, Kotov teaches a system, comprising:
a processor; and a non-transitory memory storing instructions, that when executed, cause the processor to (see Figure 7):
receive, from a user device, a request for an image (see paragraph [0073]; the system allows for a user to provide a search input including an image input; “The method 400, at block 410 includes generating a first image embedding of an image received as a search input.”;
determine, using a machine-learning model, search embeddings based on the request (see paragraphs [0029]-[0030] and [0073]-[0075]; the system has means to generate search embeddings based on the received search request input);
filter the image data based on the request to identify a filtered set of the image data; obtain a subset of the image embeddings corresponding to the filtered set of the image data (see paragraphs [0047]-[0049] and [0076]; the system can utilize the request to filter the respective content to identify an initial result set of image embeddings;
“The visual similarity component 152 performs a visual similarity analysis between an image embedding of the search image and image embeddings of other images within the embedding storage 130. The result of the visual similarity analysis is a visual similarity score that quantifies a similarity between the search image and a second image.”, para 47);
determine, based on a comparison of the search embeddings and the subset of the image embeddings, recommended image data (see paragraphs [0051]-[0052]; the system can utilize comparisons of the embeddings to determine recommended content including via scoring and ranking the results;
“In one embodiment, an aggregate similarity score is generated for each image in the first plurality of images. In one embodiment, the aggregate similarity score is the sum of a weighted visual similarity score and a weighted textual similarity score.”, para [0051];
“The results shown are selected and ordered according to the aggregate similarity score assigned to the result.”, para [0052]);
cause a presentation of the recommended image data at the user device (see Figures 3A and 3B and paragraph [0052]; the system can present the result or recommended images back to the user).
Kotov does not appear to explicitly teach:
a database including image data and image embeddings;
and in response to a selection of a recommended image from the recommended image data, provide the recommended image to the user device.
Zhang teaches a database including image data and image embeddings (see paragraphs [0063], [0065], and [0098]; the system can have a database that can store image information including image embeddings in advance of a query,
“The product database 144 includes information of the product, such as title, description, main image, and optionally other text or images of the product.”, para 65;
“As a result, the hidden features from the transformers for each product can be extracted offline and stored. ”, para 98).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify the search system of Kotov by storing image data and their respective embeddings in a single database as taught by Zhang in order to consolidate the information locally in the server system so that the server can reduce network transmissions by having to transmit information from the query to get a result before providing a response to the user.
Kotov in view of Zhang teach search results (see Kotov, paragraphs [0027] and [0044]-[0045]; see Zhang, paragraph [0096]) but do not appear to explicitly teach:
in response to a selection of a recommended image from the recommended image data, provide the recommended image to the user device.
Huynh teaches in response to a selection of a recommended image from the recommended image data, provide the recommended image to the user device (see paragraphs [0035] and [0040]; the results can be thumbnail or smaller version of the image where the user can interact with the results to receive a larger version of the image).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify the search result presentation process of Kotov in view of Zhang by presenting thumbnails of images as taught by Huynh in order to reduce transmission time and bandwidth by presenting reduced size results so that the user can select and retrieve only the larger size (higher resolution) images/results that the user desires.
With regard to claim 2, Kotov in view of Zhang and Huynh teach wherein the request comprises at least one of: a text portion, an image portion, or a campaign related portion (see Kotov, paragraph [0073]; an image can be received as the search input/request).
With regard to claim 3, Kotov in view of Zhang and Huynh teach wherein the wherein the instructions, when executed, cause the processor to determine the search embeddings based at least by: generating, using a text query encoder of the machine-learning model, at least one text query embedding based on the text portion; generating, using an image query encoder of the machine-learning model, at least one image query embedding based on the image portion; and determining the search embeddings based on the at least one text query embedding and the at least one image query embedding (see Kotov, paragraphs [0073]-[0075]; see Zhang, paragraph [0104]; the system can generate a text embedding and image embedding based on the input query request; “When both query text and query image exist, the query embeddings are combination of query text embeddings and query image embeddings.”, Zhang, para 104).
With regard to claim 4, Kotov in view of Zhang and Huynh teach wherein the instructions, when executed, further cause the processor to train the machine-learning model based at least by: training the machine-learning model using a first set of tasks having a first complexity (see Kotov, paragraphs [0030] and [0065]; the system can train the model using particular tasks to achieve a particular output);
and re-training the machine-learning model using a second set of tasks having a second complexity greater than the first complexity, after completion of the first set of tasks (see Zhang, paragraph [0084]; the system can be trained to match images to categories).
With regard to claim 7, Kotov in view of Zhang and Huynh teach wherein the instructions, when executed, cause the processor to determine the recommended image data based at least by: comparing the search embeddings with the subset of the image embeddings to compute cosine similarity distances (see Kotov, paragraphs [0020] and [0047]; the system can compare embeddings via cosine similarity);
generate ranking scores for the subset of the image embeddings based on the cosine similarity distances; selecting one or more image embeddings having highest ranking scores among the subset of the image embeddings, and determining the recommended image data corresponding to the one or more image embeddings (see Kotov, paragraphs [0044]-[0049]; the system can determine rankings based on cosine similarity and be able to select particular images via their embeddings and scores to determine the result to provide to the user).
With regard to claims 8-11 and 14, these claims are substantially similar to claims 1-4 and 7 and are rejected for similar reasons as discussed above.
With regard to claims 15-17 and 20, these claims are substantially similar to claims 1, 3, 4 and 7 respectively and are rejected for similar reasons as discussed above.
Claims 5, 12, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Kotov et al [US 2022/0138247 A1] in view of Zhang et al [US 2023/0245418 A1] and Huynh et al [US 2015/0269231 A1] in further view of Hendler et al [US 12,135,742].
With regard to claim 5, Kotov in view of Zhang and Huynh teach all the claim limitations of claims 1 and 4 as discussed above.
Kotov in view of Zhang and Huynh teach the second set of tasks includes training data related to one or more image-category pairs each of which is formed by an image of a corresponding product and a category of the corresponding product (see Zhang, paragraph [0084]; see Kotov, paragraph [0065]; the system can train the models including using image classification with training data have correct labels).
Kotov in view of Zhang and Huynh teach training and cross-entropy (see Zhang, paragraph [0095]) but do not appear to explicitly teach:
wherein: the first set of tasks includes training data related to one or more image-text pairs each of which is formed by an image and a text corresponding to the image; the machine-learning model is trained using the first set of tasks based on a cross entropy loss of image embeddings and text embeddings;
and the machine-learning model is re-trained using the second set of tasks based on a cross entropy loss of image embeddings and category embeddings.
Hendler teaches wherein: the first set of tasks includes training data related to one or more image-text pairs each of which is formed by an image and a text corresponding to the image; the machine-learning model is trained using the first set of tasks based on a cross entropy loss of image embeddings and text embeddings (see col 5, lines 7-54; the system can utilize image-text pairs when training the model including using cross-entropy loss).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify the machine learning process of the search system of Kotov in view of Zhang and Huynh by including means to do loss function evaluation using cross-entropy between text and image embeddings as taught by Hendler in order to form a more accurate model by being able to minimize the loss for both image to text evaluation as well as text to image evaluation thus helping to improve the robustness of the system by being able to handle different inputs to find desired information while still maintaining the ability to accurately find respective results when using multi-modal embeddings.
Kotov in view of Zhang and Huynh in further view of Hendler teach the machine-learning model is re-trained using the second set of tasks based on a cross entropy loss of image embeddings and category embeddings (see Hendler, col 5, lines 32-36; see Zhang, paragraph [0084]; see Kotov, paragraph [0065]; the system can train the models including using image classification with training data have correct labels and use cross-entropy loss).
With regard to claims 12 and 18, these claims are substantially similar to claim 5 and are rejected for similar reasons as discussed above.
Claims 6, 13, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Kotov et al [US 2022/0138247 A1] in view of Zhang et al [US 2023/0245418 A1] and Huynh et al [US 2015/0269231 A1] in further view of Stoop et al [US 2018/0101540 A1].
With regard to claim 6, Kotov in view of Zhang and Huynh teach all the claim limitations of claim 1 as discussed above.
Kotov in view of Zhang and Huynh do not appear to explicitly teach:
wherein the instructions, when executed, cause the processor to filter the image data based by at least one of:
identifying and excluding circular images from the filtered set of the image data using a circular filter based on the request, wherein each of the circular images has a circular shape;
identifying and excluding text-heavy images from the filtered set of the image data using a text filter based on the request, wherein each text-heavy image of the text-heavy images has a text portion occupying more than half of the text-heavy image;
or identifying and excluding duplicated images from the filtered set of the image data using a duplication filter based on the request, wherein each of the duplicated images has a hash vector based similarity score higher than a predetermined threshold with respect to an existing image in the filtered set of the image data.
Stoop teaches wherein the instructions, when executed, cause the processor to filter the image data based by at least one of: identifying and excluding circular images from the filtered set of the image data using a circular filter based on the request, wherein each of the circular images has a circular shape; identifying and excluding text-heavy images from the filtered set of the image data using a text filter based on the request, wherein each text-heavy image of the text-heavy images has a text portion occupying more than half of the text-heavy image; or identifying and excluding duplicated images from the filtered set of the image data using a duplication filter based on the request, wherein each of the duplicated images has a hash vector based similarity score higher than a predetermined threshold with respect to an existing image in the filtered set of the image data (see paragraphs [0050], [0006], and [0069], and [0056]-[0057]; the system can identify duplicates in results including using a threshold to determine how similar the visual content).
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to modify the search result presentation process of Kotov in view of Zhang and Huynh by removing/excluding duplicates as taught by Stoop in order to reduce the amount similar/duplicate information to the user so that results can be relevant but distinct so that the user isn’t shown a result page with results that are all duplicates of one another thereby providing a wider amount of unique and relevant results while not cluttering the search result page.
With regard to claims 13 and 19, these claims are substantially similar to claim 6 and are rejected for similar reasons as discussed above.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Zhao et al [US 2023/0376828 A1] teaches at paragraph [0019] cross-modal retrieval of an image in response via text query or vice versa.
Du et al [US 10,043,109] teaches at Figures 4B and 5 the ability to input an image search query and receive recommend image results including image evaluation, similarity comparisons and classifiers.
Baltescu et al [US 2023/0252550 A1] teaches at Figure 2B creating product embeddings based on image and text information.
Al Jadda et al [US 2021/0073891] teaches at Figures 1 and 2 and paragraphs 41-43 the ability to select a product and have similarity comparisons done between respective text embeddings and visual/image feature vectors.
Aggarwal et al [US 2020/0380027] teaches at paragraph 29 teaches encoders and having text and digital image embeddings and be able to compare text embeddings to image embeddings.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MARC S SOMERS whose telephone number is (571)270-3567. The examiner can normally be reached M-F 11-8 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ann Lo can be reached at 5712729767. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MARC S SOMERS/Primary Examiner, Art Unit 2159 7/10/2026