DETAILED ACTION
Status of Claims
The following is a non-final, first office action in response to the communication filed 8/8/2023
Claims 1-20 are currently pending and have been examined.
Priority
The applicant’s claim for benefit of Provisional Patent Application Serial No. 63/377104, filed 9/26/2022 has been received and acknowledged.
Information Disclosure Statement
Information Disclosure Statements received 9/23/2024 and 4/25/2024 have been reviewed and considered.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception without significantly more.
Step 1 of the Subject Matter Eligibility Test entails considering whether the claimed subject matter falls within the four statutory categories of patentable subject matter identified by 35 U.S.C. 101: Process, machine, manufacture, or composition of matter.
Claims 1-20 are directed to a method (process), a system (machine or manufacture), and a non-transitory medium (manufacture), respectively. As such, the claims are directed to statutory categories of invention.
If the claim recites a statutory category of invention, the claim requires further analysis in Step 2A. Step 2A of the Subject Matter Eligibility Test is a two-prong inquiry. In Prong One, examiners evaluate whether the claim recites a judicial exception.
Claim 1 recites abstract limitations, including those identified in bold below.
1. A computer-implemented method for generating images that represent design alternatives for three-dimensional (3D) objects, the method comprising: generating a first keyword prompt based on design intent text that describes at least a first 3D object; executing a first machine learning model on the first keyword prompt to generate a first plurality of keywords; generating a rephrase prompt based on a set of keywords that includes at least one keyword included in the first plurality of keywords; executing the first machine learning model on the rephrase prompt to generate a final text prompt; and executing a second machine learning model on the final text prompt to generate a plurality of images.
These limitations, as drafted, are a process that, under its broadest reasonable interpretation, cover performance of the limitations in the mind, or by a human using pen and paper, and therefore recite mental processes. More specifically, other than reciting that the claim is computer-implemented, nothing in the claim element precludes the aforementioned steps from practically being performed in the human mind, or by a human using pen and paper. The mere recitation of a generic computer does not take the claim out of the mental process grouping. Thus, the claim recites an abstract idea.
Claims 11 and 20 recite limitations analogous to those presented above with respect to claim 1.
If the claim recites a judicial exception in step 2A Prong One , the claim requires further analysis in step 2A Prong Two. In step 2A Prong Two, examiners evaluate whether the claim recites additional elements that integrate the exception into a practical application of that exception.
Claim 1 recites additional elements, including those identified with underlining below:
1. A computer-implemented method for generating images that represent design alternatives for three-dimensional (3D) objects, the method comprising: generating a first keyword prompt based on design intent text that describes at least a first 3D object; executing a first machine learning model on the first keyword prompt to generate a first plurality of keywords; generating a rephrase prompt based on a set of keywords that includes at least one keyword included in the first plurality of keywords; executing the first machine learning model on the rephrase prompt to generate a final text prompt; and executing a second machine learning model on the final text prompt to generate a plurality of images.
Claim 11 recites further additional elements, including those identified with underlining below:
11. One or more non-transitory computer readable media including instructions that, when executed by one or more processors, cause the one or more processors to…
Claim 20 recites further additional elements, including those identified with underlining below:
20. A system comprising: one or more memories storing instructions; and one or more processors coupled to the one or more memories that, when executing the instructions, perform the steps of:…
The additional elements recited above (i.e., computer/system and its respective components such as memory and processor, machine learning model) are recited at a high level of generality, and, as applied, are a tool used in their ordinary capacity to perform the abstract idea, and therefore amount to “apply it.”
Accordingly, in combination, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
If the additional elements do not integrate the exception into a practical application in step 2A Prong Two, then the claim is directed to the recited judicial exception, and requires further analysis under Step 2B to determine whether they provide an inventive concept (i.e., whether the additional elements amount to significantly more than the exception itself).
As discussed above, the additional elements amount to mere instructions to apply the exception. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Use of a computer or other machinery in its ordinary capacity for economic or other tasks (e.g., to receive, store, or transmit data) or simply adding a general purpose computer or computer components after the fact to an abstract idea does not provide significantly more. See Affinity Labs v. DirecTV, 838 F.3d 1253, 1262, 120 USPQ2d 1201, 1207 (Fed. Cir. 2016) (cellular telephone); TLI Communications LLC v. AV Auto, LLC, 823 F.3d 607, 613, 118 USPQ2d 1744, 1748 (Fed. Cir. 2016) (computer server and telephone unit).
Thus, even when viewed as an ordered combination, nothing in the claims add significantly more (i.e. an inventive concept) to the abstract idea.
The various limitations of claims 2, 7, 12, and 17 merely narrow the previously recited abstract idea limitations (e.g., further characterizing the prompt, the keywords) without introducing any further additional elements. For the reasons described above with respect to claims 1 and 11, this judicial exception is not meaningfully integrated into a practical application, or significantly more than the abstract idea
Claims 3-4, 13-14 recites limitations that further characterize previously identified abstract concepts (adding/grouping keywords, generation of images), and further recites the display of selectable options on a user interface and the receipt of a user input, which when recited at this level of breadth amount to extra-solution activity. MPEP 2106.05(d)(II), and the cases cited therein, including in Trading Techs. Int’l v. IBG LLC, 921 F.3d 1084, 1093 (Fed. Cir. 2019), and Intellectual Ventures I LLC v. Erie Indemnity Co., 850 F.3d 1315, 1331 (Fed. Cir. 2017), for example, indicated that the mere displaying of data is a well understood, routine, and conventional function. In addition, the Symantec, TLI, OIP Techs. and buySAFE court decisions cited in MPEP 2106.05(d)(II) indicate that mere collection or receipt of data via an interface is a well‐understood, routine, conventional function when it is claimed in a merely generic manner (as it is here). Thus, even when viewed as an ordered combination, nothing in the claims integrate the abstract idea into a practical application or amount to significantly more than the abstract idea itself.
Claims 5-6, and 15-16 recites limitations that further characterize previously identified abstract concepts (calculating a score based on inputs, displaying information based on a score), and further recites (1) the display of selectable options on a user interface and the receipt of a user input, which when recited at this level of breadth amount to extra-solution activity and (2) the application of a machine learning model, which at this level of breadth amounts to applying the abstract idea using generic computing components. MPEP 2106.05(d)(II), and the cases cited therein, including in Trading Techs. Int’l v. IBG LLC, 921 F.3d 1084, 1093 (Fed. Cir. 2019), and Intellectual Ventures I LLC v. Erie Indemnity Co., 850 F.3d 1315, 1331 (Fed. Cir. 2017), for example, indicated that the mere displaying of data is a well understood, routine, and conventional function. In addition, the Symantec, TLI, OIP Techs. and buySAFE court decisions cited in MPEP 2106.05(d)(II) indicate that mere collection or receipt of data via an interface is a well‐understood, routine, conventional function when it is claimed in a merely generic manner (as it is here). Thus, even when viewed as an ordered combination, nothing in the claims integrate the abstract idea into a practical application or amount to significantly more than the abstract idea itself.
Claims 8-9, and 18-19 recite limitations that further characterize previously identified abstract concepts (generating images based on image prompts, and capturing a prompt from a displayed image), and re-recites the application of a machine learning model, which at this level of breadth amounts to applying the abstract idea using generic computing components, and the display of information of a graphical user interface, which amounts to extra-solution activity. MPEP 2106.05(d)(II), and the cases cited therein, including in Trading Techs. Int’l v. IBG LLC, 921 F.3d 1084, 1093 (Fed. Cir. 2019), and Intellectual Ventures I LLC v. Erie Indemnity Co., 850 F.3d 1315, 1331 (Fed. Cir. 2017), for example, indicated that the mere displaying of data is a well understood, routine, and conventional function. Thus, even when viewed as an ordered combination, nothing in the claims integrate the abstract idea into a practical application or amount to significantly more than the abstract idea itself.
Claim 10 recites limitations that further characterize the machine learning models (which apply the abstract idea). Thus, even when viewed as an ordered combination, nothing in the claims integrate the abstract idea into a practical application or amount to significantly more than the abstract idea itself.
The specification demonstrates the well-understood, routine, conventional nature of additional elements as it describes the additional elements as well-understood or routine or conventional (or an equivalent term), as a commercially available product, or in a manner that indicates that the additional elements are sufficiently well-known that the specification does not need to describe the particulars of such additional elements to satisfy 35 U.S.C. §112(a).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-5 and 7-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang et al. (US 20100205202) in view of Jacobs (US 20180129188) and further in view of Sethi et al. (20220382763 A1)
Regarding claim 1, Yang discloses:
A computer-implemented method for generating images that represent design alternatives for … objects, the method comprising:
Yang [0017] In illustrated architecture 100, user 102 may access search engine 106 for the purpose of conducting a search for content on one or more content providers 110(1), . . . , 110(N). Content providers 110(1)-(N) may comprise websites, databases, or any other entity that includes content 112 that search engine 106 may search in response to receiving a user query.
Yang [0018] As illustrated, search engine 106 includes one or more processors 118 and memory 120. Memory 120 stores or has access to a search module 122, a keyword-suggestion module 124, an image-suggestion module 126 and a search-refinement module 128.
Yang, Fig. 4-5
generating a first keyword prompt based on design intent text that describes at least a first … object;
Yang [0020] For instance, envision that user 102 requests to receive images that are associated with the search query "Apple."
Yang [0031] User interface 200 includes a text box 202 in which user 102 inputted the original query "Apple."
executing a first … model on the first keyword prompt to generate a first plurality of keywords;
Yang [0022] To determine which of multiple keyword candidates to suggest as keywords, keyword-suggestion module 124 includes a relatedness calculator 130 and an informativeness calculator 132. …
Yang [0023] Informativeness calculator 132, meanwhile, attempts to find keywords that are each informative enough (when coupled with the original query) to reflect a different aspect of the original query. Returning to the example of the search query "Apple," calculator 132 may determine that the words "Computer," "Fruit" and "Smartphone" each reflect diverse aspects of the query "Apple." To make this determination, calculator may determine that images associated with the query "Apple Computer" make up a very different set of images than sets of images associated with the queries "Apple Fruit" and "Apple Smartphone," respectively.
Yang [0024] In some instances, keyword-suggestion module 124 combines the input from relatedness calculator 130 with the input from informativeness calculator 132 to determine a set of keywords associated with the originally-inputted query. By doing so, keyword-suggestion module 124 determines a set of keywords that are sufficiently related to the original query and that sufficiently represent varying aspects of the original query. In some instances, module 124 sets a predefined number of keywords (e.g., one, three, ten, etc.). In other instances, however, module 124 may set a predefined threshold score that each keyword candidate should score in order to be deemed a keyword, which may result in varying numbers of keywords for different queries.
generating a rephrase prompt based on a set of keywords that includes at least one keyword included in the first plurality of keywords;
Yang [0027] Once keyword-suggestion module 124 determines a set of keywords associated with a received query and image-suggestion module determines representative images of clusters therein, search engine 106 may suggest to user 102 to refine the search request based on selection of a keyword and a representative image. For instance, search engine 106 may output a user interface 138 to user 102 that allows user 102 to select a keyword (e.g., "Fruit") and an image associated with a cluster of that keyword (e.g., an image of a red apple).
Yang [0031] User interface 200 also includes one or more keywords 204 and one or images 206 that user 102 may select to refine the search. Keywords 204 include a keyword 208 entitled "Computer," a keyword 210 entitled "Fruit," and a keyword 212 entitled "Smartphone." Each of keywords 208, 210, and 212 is associated with a set of images 214, 216, and 218, respectively. While FIG. 2 illustrates three keywords and three images, other implementations may employ any number of one or more keywords and any number of one or more images.
Yang [0033] In response to receiving UI 200 at device 104, user 102 may select a keyword and one or more images on which to refine the image search.
executing the first … model on the rephrase prompt to generate a final text prompt; and
Yang [0028] Upon receiving a selected keyword and image, search-refinement module 128 may refine the user's search. First, a keyword-similarity module 140 may search for or determine the images that are associated with the query and the keyword ("Apple Fruit").
Yang [0033] By doing so, user 102 would be choosing to refine the search to include images associated with the query "Apple Fruit."
Yang [0039] Once a user chooses a keyword-image suggestion, the selected keyword may be appended to the initial query, which results in a composite or combined query. The techniques may then first search for images that are associated with the composite textual query.
executing a second … model on the final text prompt to generate a plurality of images.
Yang [0007] FIG. 3 depicts an illustrative UI that the search engine may serve to the client computing device in response to receiving a user selection of a keyword "computer" and a particular image from the UI of FIG. 2. As illustrated, this UI includes multiple images that are associated with the query "Apple Computer" and with the image selected by the user.
Yang [0035] FIG. 3 represents a user interface (UI) 300 that search engine 106 may serve to device 104 after user 102 selects a particular image associated with keyword 208 ("Computer") of UI 200. As illustrated in text box 202, in response to receiving the user selection, search engine 106 ran a search of content providers 110(1)-(N) for images associated with the refined query "Apple Computer." Furthermore, search engine 106 may have compared these images with the user-selected image in order to determine a similarity there between. Search engine 106 may have then ranked these images based on the determined similarities to present a set of images 302 in a manner based at least in part on the ranking. For instance, those images determined to be most similar to the selected image may be at or near the top of UI 300 when compared to less similar images.
Yang [0039] The final results are then presented to the user. In most instances, these final results more adequately conform to the intent of the searching user than when compared with traditional image-search techniques.
Yang, as shown above, discloses the generation of a images that represent design alternatives for objects that are 3D (e.g., an apple computer), but the generated set of objects are not explicitly recited to be three-dimensional (3D). Jacobs is directed to a natural language interface for computer-aided design systems. Jacobs discloses that said objects can be two-dimensional objects or three-dimensional (3D) objects.
Jacobs [0034] For example, when information needed to respond to a query resides outside of the CAD system and natural language program server module 122, automated searching of appropriate databases is initiated, the databases being supplied as resource provider server modules 132 in order to provide information from suppliers, marketplaces, and other external services. In some examples, resource provider server module 132 is an external service supplier marketplace database, either a source of information or standalone entity that will perform calculations.
Jacobs [0091] Alternatively or additionally, words included in the additional information may be added to query without replacing one or more words. For instance, CAD context database 500 may link one or more words of user voice input to categories or geometries; as a non-limiting example, CAD context database 500 may link the word “cup” to words describing upwardly opening recesses, cylindrical forms, or the like. Additional information may result in return of “best match” queries that, as indicated by CAD context database 500, may describe some but not all features of the object to which user voice input refers.
Jacobs [0094] Still referring to FIG. 9, at least a modified graphical model may be displayed to a user. This may be performed using GUI 136 as described above. Where at least a modified graphical model includes a plurality of graphical models, the plurality of graphical models may be displayed to the user; plurality may be ranked, for instance according to degree of match with query, and displayed in rank-order.
Jacobs [0100] Still viewing FIG. 12, at least an image may include a three-dimensional graphical model of object. Three-dimensional graphical model may be any three-dimensional graphical model as described above, including without limitation a CAD model of object. … In another embodiment, the at least an image includes at least a two-dimensional image, such as a photograph or two-dimensional CAD model illustrating a view of object; view may be a straightaway side view, an isometric view, a perspective view, or the like.
Since each individual element and its function are shown in the prior art, albeit shown in separate references, the difference between the claimed subject matter and the prior art rests not on any individual element or function but in the very combination itself. That is in the substitution of the three-dimensional object image of Jacobs for the two-dimensional object image of Yang. Thus, the simple substitution of one known element for another producing a predictable result renders the claim obvious.
Yang, as shown above, also discloses a first and second model, but does not explicitly disclose that the models are machine-learning models. Sethi is directed to populating search results with intent and context-based images.
Sethi discloses that the first model may be a machine learning model.
Sethi [0034] NLP service manager 250 may be configured to analyze the search query to determine the intent of the user and the context data associated with the search query. Analysis of the search query may include tokenization, sentence segmentation, part of speech tagging, lemmatization, etc. NLP service manager 325 may include at least one component to perform one or more of the above functions. For example, NLP service manager 325 may include a tokenizer, a part-of-speech tagger, a lemmatization engine, etc. NLP service manager 325 may include OpenNLP™, natural language toolkit, Stanford CoreNLP™, etc. The above components may be included as part of an NLP processor 255.
Sethi [0055] At stage D, NLP service manager 250 may predict the user's intent and subsequent action. The prediction may be converted into keywords or key attributes and passed as input to machine learning engine 215 at stage E.
Sethi also discloses that the second model may be a machine learning model.
Sethi [0038] In one embodiment, machine learning engine 215 may use a convolutional neural network to predict a set of images with features that match one or more key attributes of the search query from images stored in data store 210.
Sethi [0055] At stage F, machine learning engine 215 may perform a search term to context mapping, filtering of semantic keywords, and generating a set of images for the search results. At stage G, the set of images may be ranked and displayed at the user interface for the user.
One of ordinary skill in the art at the time of filing would have recognized that applying the machine learning models of Sethi to the combined system of Yang and Jacobs would have yielded predictable results and resulted in an improved system that is more capable of presenting a desired result in response to a user search.
Regarding claim 2, the combination of Yang, Jacobs and Sethi disclose the limitations of claim 1, and further discloses:
The computer-implemented method of claim 1, wherein the first keyword prompt comprises a request to list at least one of a design, a style, or a part associated with the design intent text.
Yang [0020] For instance, envision that user 102 requests to receive images that are associated with the search query "Apple."
Yang [0022] To determine which of multiple keyword candidates to suggest as keywords, keyword-suggestion module 124 includes a relatedness calculator 130 and an informativeness calculator 132. …
Yang [0023] Informativeness calculator 132, meanwhile, attempts to find keywords that are each informative enough (when coupled with the original query) to reflect a different aspect of the original query. Returning to the example of the search query "Apple," calculator 132 may determine that the words "Computer," "Fruit" and "Smartphone" each reflect diverse aspects of the query "Apple." To make this determination, calculator may determine that images associated with the query "Apple Computer" make up a very different set of images than sets of images associated with the queries "Apple Fruit" and "Apple Smartphone," respectively.
In addition, see: Jacobs Fig. 8 , [0072]-[0078] disclosing an exemplary model descriptor database; As a non-limiting example, at least a descriptive word 816 may include one or more words, phrases, or other terms used by a user to describe an object represented by an object model descriptor of the at least an object model descriptor 800; for instance, at without limitation, an object model descriptor associated with a drinking vessel may also be linked in model descriptor database 148 to a first descriptive word of “cup,” a second descriptive word of “glass,” and a third descriptive word of “mug,” with the result that a query containing any one of those three words may occasion the retrieval of the object model descriptor. See claim 1 for rationale to combine.
Regarding claim 3, the combination of Yang, Jacobs and Sethi disclose the limitations of claim 1, and further discloses:
The computer-implemented method of claim 1, further comprising:
displaying within a graphical user interface a selectable version of a first keyword included in the first plurality of keywords;
Yang [0006] FIG. 2 depicts an illustrative user interface (UI) that the search engine may serve to a client computing device in response to receiving a search query. As illustrated, this UI suggests that the user refine the search request for images associated with the query "Apple" by selecting an additional keyword and an image.
Yang [0023] Informativeness calculator 132, meanwhile, attempts to find keywords that are each informative enough (when coupled with the original query) to reflect a different aspect of the original query. Returning to the example of the search query "Apple," calculator 132 may determine that the words "Computer," "Fruit" and "Smartphone" each reflect diverse aspects of the query "Apple." To make this determination, calculator may determine that images associated with the query "Apple Computer" make up a very different set of images than sets of images associated with the queries "Apple Fruit" and "Apple Smartphone," respectively.
determining that the first keyword has been selected based on user input received via the graphical user interface; and adding the first keyword to the first set of keywords.
Yang [0039] Once a user chooses a keyword-image suggestion, the selected keyword may be appended to the initial query, which results in a composite or combined query. The techniques may then first search for images that are associated with the composite textual query.
Regarding claim 4, the combination of Yang, Jacobs and Sethi disclose the limitations of claim 1, and further discloses:
The computer-implemented method of claim 1, further comprising:
displaying within a graphical user interface a selectable version of a first keyword associated with … design;
Yang [0006] FIG. 2 depicts an illustrative user interface (UI) that the search engine may serve to a client computing device in response to receiving a search query. As illustrated, this UI suggests that the user refine the search request for images associated with the query "Apple" by selecting an additional keyword and an image.
Yang [0023] Informativeness calculator 132, meanwhile, attempts to find keywords that are each informative enough (when coupled with the original query) to reflect a different aspect of the original query. Returning to the example of the search query "Apple," calculator 132 may determine that the words "Computer," "Fruit" and "Smartphone" each reflect diverse aspects of the query "Apple." To make this determination, calculator may determine that images associated with the query "Apple Computer" make up a very different set of images than sets of images associated with the queries "Apple Fruit" and "Apple Smartphone," respectively.
Jacobs, as shown above, discloses that said objects can be 2D designs or 3D designs. See claim 1 for rationale to combine.
Jacobs [0094] Still referring to FIG. 9, at least a modified graphical model may be displayed to a user. This may be performed using GUI 136 as described above. Where at least a modified graphical model includes a plurality of graphical models, the plurality of graphical models may be displayed to the user; plurality may be ranked, for instance according to degree of match with query, and displayed in rank-order.
Jacobs [0100] Still viewing FIG. 12, at least an image may include a three-dimensional graphical model of object. See also - Jacobs [0034], [0091]
determining that the first keyword has been selected based on user input received via the graphical user interface; and adding the first keyword to the first set of keywords.
Yang [0039] Once a user chooses a keyword-image suggestion, the selected keyword may be appended to the initial query, which results in a composite or combined query. The techniques may then first search for images that are associated with the composite textual query.
Regarding claim 5, the combination of Yang, Jacobs and Sethi disclose the limitations of claim 1, and further discloses:
The computer-implemented method of claim 1, further comprising:
executing a multimodal … model on an image prompt and a first keyword included in the first plurality of keywords to compute a first score; and
Yang [0037] FIGS. 4-5 depict an illustrative process 400 for suggesting to a user to refine a search based on selection of both a keyword and an image.
Yang [0038] Generally, process 400 describes techniques for providing both keyword and image suggestions in response to receiving a search query in order to help users express the search intent of the user more clearly.
Yang [0039] Once a user chooses a keyword-image suggestion, the selected keyword may be appended to the initial query, which results in a composite or combined query.
Yang [0042] In some instances, the described techniques implement a two-step approach to generating the keyword-image suggestions. First, a statistical method is proposed to suggest keywords (e.g., tags or text surrounding initial search results) that can reduce the ambiguity of the initial query. After that, for each keyword suggestion, the techniques may collect the images associated with both the initial query and the suggested keyword and cluster these images into several groups or clusters. Each cluster represents a different aspect of the combined query, and the techniques select the most representative images from the clusters to form the image suggestions.
Yang [0043] Finally, the techniques refine the text-based results by using visual information of the selected suggested image. That is, the techniques compare the selected suggested image to the text-based image search results to determine a similarity there between. In some instances, the techniques employ content-based image retrieval (CBIR) to compare these images based on one or more visual modalities, such as color, texture, and shape. The techniques may then rank and reorder the search results based on the determined visual similarities
Yang [0063] In response, user 102 may select a keyword and an image from the rendered user interface. In some instances, the user selects both the keyword and the image by selecting, from the UI, an image that is associated with a particular keyword (and, in some instances, associated with a particular cluster of the keyword). At operation 414, search engine 106 receives the selection and, in response, attempts to rank and return images according to the selection. That is, the search engine analyzes the images that associated the combined query "Q+q.sub.i" (e.g., "Apple Fruit") and then ranks the images associated with the combine query according to each image's similarity to the selected image. In some instances, this comparison is made on the basis of color, texture, and/or shape similarity.
See also: Yang [0046], [0048], [0055]- [0057], [0060]
Sethi, as shown above, discloses that the applied models may be machine learning models. See claim 1 for rationale to combine.
Sethi [0055] At stage D, NLP service manager 250 may predict the user's intent and subsequent action. The prediction may be converted into keywords or key attributes and passed as input to machine learning engine 215 at stage E.
Sethi [0055] At stage F, machine learning engine 215 may perform a search term to context mapping, filtering of semantic keywords, and generating a set of images for the search results. At stage G, the set of images may be ranked and displayed at the user interface for the user.
displaying a selectable version of the first keyword within a graphical user interface, wherein at least one visual characteristic of the selectable version of the first keyword is based on the first score.
Yang [0062] Next, operation 412 represents that search engine 106 may return, to the client computing device 104 of user 102, the suggested keywords and images for selection in order to refine the image search. For instance, search engine 106 may return the keywords "Computer," "Fruit," and "Smart Phone," along with representative images of clusters therein.
Yang [0024] In some instances, keyword-suggestion module 124 combines the input from relatedness calculator 130 with the input from informativeness calculator 132 to determine a set of keywords associated with the originally-inputted query. By doing so, keyword-suggestion module 124 determines a set of keywords that are sufficiently related to the original query and that sufficiently represent varying aspects of the original query. In some instances, module 124 sets a predefined number of keywords (e.g., one, three, ten, etc.). In other instances, however, module 124 may set a predefined threshold score that each keyword candidate should score in order to be deemed a keyword, which may result in varying numbers of keywords for different queries.
Yang, Fig. 2, disclosing display of selectable keywords and associated images
See claim 1 for rationale to combine.
Regarding claim 7, the combination of Yang, Jacobs and Sethi disclose the limitations of claim 1, and further discloses:
The computer-implemented method of claim 1, wherein generating the rephrase prompt comprises constructing a request to combine every keyword included in the set of keywords.
Yang [0037] FIGS. 4-5 depict an illustrative process 400 for suggesting to a user to refine a search based on selection of both a keyword and an image.
Yang [0038] Generally, process 400 describes techniques for providing both keyword and image suggestions in response to receiving a search query in order to help users express the search intent of the user more clearly.
Yang [0039] Once a user chooses a keyword-image suggestion, the selected keyword may be appended to the initial query, which results in a composite or combined query.
Regarding claim 8, the combination of Yang, Jacobs and Sethi disclose the limitations of claim 1, and further discloses:
The computer-implemented method of claim 1, further comprising executing the second … model on one or more image prompts when generating the plurality of images.
Yang [0023] Informativeness calculator 132, meanwhile, attempts to find keywords that are each informative enough (when coupled with the original query) to reflect a different aspect of the original query. Returning to the example of the search query "Apple," calculator 132 may determine that the words "Computer," "Fruit" and "Smartphone" each reflect diverse aspects of the query "Apple." To make this determination, calculator may determine that images associated with the query "Apple Computer" make up a very different set of images than sets of images associated with the queries "Apple Fruit" and "Apple Smartphone," respectively.
Yang [0031] User interface 200 also includes one or more keywords 204 and one or images 206 that user 102 may select to refine the search. Keywords 204 include a keyword 208 entitled "Computer," a keyword 210 entitled "Fruit," and a keyword 212 entitled "Smartphone." Each of keywords 208, 210, and 212 is associated with a set of images 214, 216, and 218, respectively. While FIG. 2 illustrates three keywords and three images, other implementations may employ any number of one or more keywords and any number of one or more images.
Yang [0007] FIG. 3 depicts an illustrative UI that the search engine may serve to the client computing device in response to receiving a user selection of a keyword "computer" and a particular image from the UI of FIG. 2. As illustrated, this UI includes multiple images that are associated with the query "Apple Computer" and with the image selected by the user.
Yang [0035] FIG. 3 represents a user interface (UI) 300 that search engine 106 may serve to device 104 after user 102 selects a particular image associated with keyword 208 ("Computer") of UI 200. As illustrated in text box 202, in response to receiving the user selection, search engine 106 ran a search of content providers 110(1)-(N) for images associated with the refined query "Apple Computer." Furthermore, search engine 106 may have compared these images with the user-selected image in order to determine a similarity there between. Search engine 106 may have then ranked these images based on the determined similarities to present a set of images 302 in a manner based at least in part on the ranking. For instance, those images determined to be most similar to the selected image may be at or near the top of UI 300 when compared to less similar images.
Yang [0039] The final results are then presented to the user. In most instances, these final results more adequately conform to the intent of the searching user than when compared with traditional image-search techniques.
Sethi, as shown above, discloses that the second model may be a machine learning model. See claim 1 for rationale to combine.
Sethi [0038] In one embodiment, machine learning engine 215 may use a convolutional neural network to predict a set of images with features that match one or more key attributes of the search query from images stored in data store 210.
Sethi [0055] At stage F, machine learning engine 215 may perform a search term to context mapping, filtering of semantic keywords, and generating a set of images for the search results. At stage G, the set of images may be ranked and displayed at the user interface for the user.
Sethi [0064] Machine learning engine 460 may use one or more images in product images store 470 to determine a set of images corresponding to the search query. In addition, machine learning engine 460 may use images in product images store to determine and/or combine images resulting in images 455.
Regarding claim 9, the combination of Yang, Jacobs and Sethi disclose the limitations of claim 8, and further discloses:
The computer-implemented method of claim 8, further comprising capturing a first image prompt included in the one or more image prompts from a … model displayed within a graphical user interface.
Yang [0007] FIG. 3 depicts an illustrative UI that the search engine may serve to the client computing device in response to receiving a user selection of a keyword "computer" and a particular image from the UI of FIG. 2. As illustrated, this UI includes multiple images that are associated with the query "Apple Computer" and with the image selected by the user.
Yang [0012] The tools may then suggest both the keywords and a representative image for each cluster associated with the combined query comprising the original query and the selected keyword. …. In response to receiving a selection of this image, the tools may rank the images associated with "Apple Fruit" based on similarity to the selected image. The tools then output, to the user, images associated with "Apple Fruit" in a manner based at least in part on the ranking of the images (e.g., in descending order, beginning with the highest rank). By doing so, the tools allow for better understanding of the user's intent in submitting a request for images and, hence, allow for better service to the user.
Yang [0023] Informativeness calculator 132, meanwhile, attempts to find keywords that are each informative enough (when coupled with the original query) to reflect a different aspect of the original query. Returning to the example of the search query "Apple," calculator 132 may determine that the words "Computer," "Fruit" and "Smartphone" each reflect diverse aspects of the query "Apple." To make this determination, calculator may determine that images associated with the query "Apple Computer" make up a very different set of images than sets of images associated with the queries "Apple Fruit" and "Apple Smartphone," respectively.
Yang [0031] User interface 200 also includes one or more keywords 204 and one or images 206 that user 102 may select to refine the search. Keywords 204 include a keyword 208 entitled "Computer," a keyword 210 entitled "Fruit," and a keyword 212 entitled "Smartphone." Each of keywords 208, 210, and 212 is associated with a set of images 214, 216, and 218, respectively. While FIG. 2 illustrates three keywords and three images, other implementations may employ any number of one or more keywords and any number of one or more images.
Jacobs, as shown above, discloses that said objects can be 2D designs or 3D designs. See claim 1 for rationale to combine.
Jacobs [0094] Still referring to FIG. 9, at least a modified graphical model may be displayed to a user. This may be performed using GUI 136 as described above. Where at least a modified graphical model includes a plurality of graphical models, the plurality of graphical models may be displayed to the user; plurality may be ranked, for instance according to degree of match with query, and displayed in rank-order.
Jacobs [0100] Still viewing FIG. 12, at least an image may include a three-dimensional graphical model of object. See also - Jacobs [0034], [0091]
Regarding claim 10 the combination of Yang, Jacobs and Sethi disclose the limitations of claim 1, and further discloses:
The computer-implemented method of claim 1, wherein the first machine learning model comprises a generative prompt-to-text … model, and the second machine learning model comprises a generative prompt-to-image … model.
Yang [0018] As illustrated, search engine 106 includes one or more processors 118 and memory 120. Memory 120 stores or has access to a search module 122, a keyword-suggestion module 124, an image-suggestion module 126 and a search-refinement module 128.
Sethi discloses that the first model may be a machine learning model. See claim 1 for rationale to combine.
Sethi [0034] NLP service manager 250 may be configured to analyze the search query to determine the intent of the user and the context data associated with the search query. Analysis of the search query may include tokenization, sentence segmentation, part of speech tagging, lemmatization, etc. NLP service manager 325 may include at least one component to perform one or more of the above functions. For example, NLP service manager 325 may include a tokenizer, a part-of-speech tagger, a lemmatization engine, etc. NLP service manager 325 may include OpenNLP™, natural language toolkit, Stanford CoreNLP™, etc. The above components may be included as part of an NLP processor 255.
Sethi [0055] At stage D, NLP service manager 250 may predict the user's intent and subsequent action. The prediction may be converted into keywords or key attributes and passed as input to machine learning engine 215 at stage E.
Sethi also discloses that the second model may be a machine learning model. See claim 1 for rationale to combine.
Sethi [0038] In one embodiment, machine learning engine 215 may use a convolutional neural network to predict a set of images with features that match one or more key attributes of the search query from images stored in data store 210.
Sethi [0055] At stage F, machine learning engine 215 may perform a search term to context mapping, filtering of semantic keywords, and generating a set of images for the search results. At stage G, the set of images may be ranked and displayed at the user interface for the user.
Though not explicitly recited in the claims, Examiner notes that Jacobs further discloses queries may include voice input (see abstract), and one of ordinary skill in the art of the time of filing would understand that the inclusion of different types of prompts (text, voice, etc.) would improve the interface between users and the search engine, thus facilitating a more seamless search experience.
Regarding claim 11, Yang discloses:
One or more non-transitory computer readable media including instructions that, when executed by one or more processors, cause the one or more processors to …Yang [0018] As illustrated, search engine 106 includes one or more processors 118 and memory 120. Memory 120 stores or has access to a search module 122, a keyword-suggestion module 124, an image-suggestion module 126 and a search-refinement module 128.
In addition, claim 11 recites limitations analogous to the limitations of claim 1. See rejection of claim 1 above.
Regarding claim 12, the combination of Yang, Jacobs and Sethi disclose the limitations of claim 11. In addition, claim 12 recites limitations analogous to the limitations of claim 2. See rejection of claim 2 above.
Regarding claim 13, the combination of Yang, Jacobs and Sethi disclose the limitations of claim 11. In addition, claim 13 recites limitations analogous to the limitations of claim 3. See rejection of claim 3 above.
Regarding claim 14, the combination of Yang, Jacobs and Sethi disclose the limitations of claim 11 and further disclose:
The one or more non-transitory computer readable media of claim 11, further comprising: designating a first word or a first phrase as a first user keyword based on user input received via a graphical user interface; and adding the first user keyword to the first set of keywords.
Yang [0006] FIG. 2 depicts an illustrative user interface (UI) that the search engine may serve to a client computing device in response to receiving a search query. As illustrated, this UI suggests that the user refine the search request for images associated with the query "Apple" by selecting an additional keyword and an image.
Yang [0007] FIG. 3 depicts an illustrative UI that the search engine may serve to the client computing device in response to receiving a user selection of a keyword "computer" and a particular image from the UI of FIG. 2. As illustrated, this UI includes multiple images that are associated with the query "Apple Computer" and with the image selected by the user.
Yang [0037] FIGS. 4-5 depict an illustrative process 400 for suggesting to a user to refine a search based on selection of both a keyword and an image.
Yang [0038] Generally, process 400 describes techniques for providing both keyword and image suggestions in response to receiving a search query in order to help users express the search intent of the user more clearly.
Yang [0039] Once a user chooses a keyword-image suggestion, the selected keyword may be appended to the initial query, which results in a composite or combined query.
Regarding claim 15, the combination of Yang, Jacobs and Sethi disclose the limitations of claim 11. In addition, claim 15 recites limitations analogous to the limitations of claim 5. See rejection of claim 5 above.
Regarding claim 16, the combination of Yang, Jacobs and Sethi disclose the limitations of claim 15 and further disclose:
The one or more non-transitory computer readable media of claim 15, wherein the first score estimates a similarity between the image prompt and the first keyword.
Yang [0037] FIGS. 4-5 depict an illustrative process 400 for suggesting to a user to refine a search based on selection of both a keyword and an image.
Yang [0038] Generally, process 400 describes techniques for providing both keyword and image suggestions in response to receiving a search query in order to help users express the search intent of the user more clearly.
Yang [0039] Once a user chooses a keyword-image suggestion, the selected keyword may be appended to the initial query, which results in a composite or combined query.
Yang [0042] In some instances, the described techniques implement a two-step approach to generating the keyword-image suggestions. First, a statistical method is proposed to suggest keywords (e.g., tags or text surrounding initial search results) that can reduce the ambiguity of the initial query. After that, for each keyword suggestion, the techniques may collect the images associated with both the initial query and the suggested keyword and cluster these images into several groups or clusters. Each cluster represents a different aspect of the combined query, and the techniques select the most representative images from the clusters to form the image suggestions.
Yang [0043] Finally, the techniques refine the text-based results by using visual information of the selected suggested image. That is, the techniques compare the selected suggested image to the text-based image search results to determine a similarity there between. In some instances, the techniques employ content-based image retrieval (CBIR) to compare these images based on one or more visual modalities, such as color, texture, and shape. The techniques may then rank and reorder the search results based on the determined visual similarities
Yang [0063] In response, user 102 may select a keyword and an image from the rendered user interface. In some instances, the user selects both the keyword and the image by selecting, from the UI, an image that is associated with a particular keyword (and, in some instances, associated with a particular cluster of the keyword). At operation 414, search engine 106 receives the selection and, in response, attempts to rank and return images according to the selection. That is, the search engine analyzes the images that associated the combined query "Q+q.sub.i" (e.g., "Apple Fruit") and then ranks the images associated with the combine query according to each image's similarity to the selected image. In some instances, this comparison is made on the basis of color, texture, and/or shape similarity.
See also: Yang [0046], [0048], [0055]- [0057], [0060]
Regarding claim 17, the combination of Yang, Jacobs and Sethi disclose the limitations of claim 11. In addition, claim 17 recites limitations analogous to the limitations of claim 7. See rejection of claim 7 above.
Regarding claim 18, the combination of Yang, Jacobs and Sethi disclose the limitations of claim 11. In addition, claim 18 recites limitations analogous to the limitations of claim 8. See rejection of claim 8 above.
Regarding claim 19, the combination of Yang, Jacobs and Sethi disclose the limitations of claim 18 and further disclose:
The one or more non-transitory computer readable media of claim 18, further comprising setting a first image prompt included in the one or more image prompts equal to at least a portion of an image displayed within a graphical user interface to recursively generate the plurality of images.
Yang [0035] FIG. 3 represents a user interface (UI) 300 that search engine 106 may serve to device 104 after user 102 selects a particular image associated with keyword 208 ("Computer") of UI 200. As illustrated in text box 202, in response to receiving the user selection, search engine 106 ran a search of content providers 110(1)-(N) for images associated with the refined query "Apple Computer." Furthermore, search engine 106 may have compared these images with the user-selected image in order to determine a similarity there between. Search engine 106 may have then ranked these images based on the determined similarities to present a set of images 302 in a manner based at least in part on the ranking. For instance, those images determined to be most similar to the selected image may be at or near the top of UI 300 when compared to less similar images.
Yang [0037] FIGS. 4-5 depict an illustrative process 400 for suggesting to a user to refine a search based on selection of both a keyword and an image.
Yang [0038] Generally, process 400 describes techniques for providing both keyword and image suggestions in response to receiving a search query in order to help users express the search intent of the user more clearly.
Yang [0039] Once a user chooses a keyword-image suggestion, the selected keyword may be appended to the initial query, which results in a composite or combined query.
Yang Fig. 2-3 (Apple – Computer to Apple Computer)
Regarding claim 20, Yang discloses:
One or more non-transitory computer readable media including instructions that, when executed by one or more processors, cause the one or more processors to …Yang [0018] As illustrated, search engine 106 includes one or more processors 118 and memory 120. Memory 120 stores or has access to a search module 122, a keyword-suggestion module 124, an image-suggestion module 126 and a search-refinement module 128.
In addition, claim 20 recites limitations analogous to the limitations of claim 1. See rejection of claim 1 above.
Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang et al. (US 20100205202) in view of Jacobs (US 20180129188) in view of Sethi et al. (20220382763 A1) and further in view of Chundi et al. (US 20210303568).
Regarding claim 6, the combination of Yang, Jacobs and Sethi disclose the limitations of claim 5, and further discloses:
The computer-implemented method of claim 5, wherein the at least one visual characteristic comprises at least one of a color, an intensity, an opacity, a size, or a position.
Yang [0024] In some instances, keyword-suggestion module 124 combines the input from relatedness calculator 130 with the input from informativeness calculator 132 to determine a set of keywords associated with the originally-inputted query. By doing so, keyword-suggestion module 124 determines a set of keywords that are sufficiently related to the original query and that sufficiently represent varying aspects of the original query. In some instances, module 124 sets a predefined number of keywords (e.g., one, three, ten, etc.). In other instances, however, module 124 may set a predefined threshold score that each keyword candidate should score in order to be deemed a keyword, which may result in varying numbers of keywords for different queries.
Yang [0062] Next, operation 412 represents that search engine 106 may return, to the client computing device 104 of user 102, the suggested keywords and images for selection in order to refine the image search. For instance, search engine 106 may return the keywords "Computer," "Fruit," and "Smart Phone," along with representative images of clusters therein. Yang recites the concept of scoring keywords and presenting keywords based on said scoring, which strongly suggests that the keywords ranked highest are presented first. Chundi discloses methods and systems for search queries. Chundi more explicitly discloses that visual characteristics may comprise position (Chundi [0069] the list of keywords 536 being displayed with the highest weighted retrieved keyword at the top of the list 536.)
One of ordinary skill in the art at the time of filing would have recognized that applying the machine learning models of Sethi to the combined system of Yang, Jacobs and Sethi would have yielded predictable results and resulted in an improved system that is more capable of optimized presentation of information that is perceived as most relevant to the inquiry.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Volum et al. (US 20230125036 ), disclosing a natural language interface for virtual environment generation.
Fong et al. (US 20090058860), disclosing a method for transforming language into a visual form.
Harp et al. (US 20150186418), disclosing methods and systems for use of a database of 3D object models for search queries.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ABBY J FLYNN whose telephone number is (571)272-9855. The examiner can normally be reached Monday - Friday 8:30-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, James Trammell can be reached at 571-272-6712. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ABBY J FLYNN/Examiner, Art Unit 3663