DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The disclosure is objected to because paragraph 33 provides that the “Natural language recognition (NER) model” is “a machine learning model that is able to identify and classify named entities in natural language texts, sentences, and strings”. The acronym NER typically stands for Named Entity Recognition, a type of Natural Language Processing (NLP) task, but the specification describes NER with the language “natural language recognition”, which typically corresponds to NLR, not NER. However, the description in paragraph 33 is more akin to a description of named entity recognition (NER) than it is a description of NLR. NER and NLR are both NLP tasks, but NLR is generally considered a broader term that may include more capabilities than NER. For example, an NER task would be to determine which spans of text are named entities and what type of entities they are, whereas NLR could derive semantic meaning based on the context of two named entities appearing in the same passage of text.
The specification mentions the phrase “natural language recognition model” in paragraphs 6, 15, and 39. Applicant is encouraged to review the entire specification and clarify any instance related to the above subject matter to clarify which type of processing is intended to be referenced.
Appropriate correction is required. Depending upon the particular changes made to the specification, step S451 in Figure 4 may need to be updated with the corresponding changes made to the description of step S451 and Figure 4.
Claim Objections
Claims 1, 2, 4, 5, 7-9, 11, 12 and 14 are objected to because of the following informalities: the uncommon phrases “to-be-searched image” and “to-be-searched images” should be changed to “stored image” or “stored images” or the like to promote clarity and readability of the claims. Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claims 1-14 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Claims 1 and 8 use the terms “whether” and “while”. The term “whether” can be interchangeable with the term “if”. The term “while” is also a conditional type of term and implies periods of time or durations where a matching condition is true or false. The claims do not sufficiently explain what constitutes a “match” such that a reader of claims 1 and 8 would understand the meaning of phrases like “while the search string matches” and “while the search string does not match” for the same classification tag, which appears contradictory. The “one” classification tag either matches the search string or it does not. The term “while” in this context is confusing because it is conditioned on “whether the search string” does or does not match the “one” classification tag. If a determination has been made that a search string matches an image tag, then the unspecified duration of the matching has already ended, which makes “while the search string matches” or “not match” so confusing. The claims do not explain a period of time which necessarily corresponds to a match or lack thereof.
Based on an example provided in the specification, “while” could be interpreted as being more akin to “if”, “when” or “in response to”. For example, paragraph 18 discloses, “In step S215 of FIG. 2, the processor 110 determines whether the to-be-searched image has a corresponding geographical location. If step S215 in FIG. 2 is YES, then step S220 in FIG. 2 is carried out.” (emphasis added). However, the claims do not require that interpretation. The term “while”, notably, only occurs in the specification in paragraphs 4 and 5 of the Summary section of the specification. Accordingly, for purposes of applying prior art, the term “whether” as it appears in any pending claim is considered equivalent to “if” and “while” is interpreted to refer to the event of a match of the search string being determined or a lack thereof without any reference to a particular period or duration of time, i.e., if the classification tag matches the query string, then generate the first search result, and if the classification tag does not match the query string, then obtain the second search result. Dependent claims 2-7 and 9-14 are rejected for inheriting and not curing the deficiencies of their respective parent claims.
Claim 1 recites, “a plurality of to-be-searched images, each of the to-be-searched images corresponds to a first comparison vector, and at least one classification tag and first location information respectively correspond to a part of the to-be-searched images” (emphasis added). The phrase “the to-be-searched images” could refer to any of the “plurality of to-be-searched images”. Therefore, “the to-be-searched images” lacks antecedent basis. Claim 8 recites similar language as claim 1, “a plurality of to-be-searched images; and a display device comprising a user interface, wherein the processor obtains a search string, wherein the search string is provided to search for one or more of the plurality of to-be-searched images, each of the to-be-searched image corresponds to a first comparison vector, and at least one classification tag and first location information respectively correspond to a part of the to-be-searched image” (emphasis added). The phrase “the to-be-searched image” lacks a proper antecedent basis due to the earlier-recited plurality of to-be-searched images”. For purposes of applying prior art, “the to-be-searched images” and “the to-be-searched image”, as those phrases appear in the claims, are interpreted as “the plurality of to-be-searched images” and “the plurality of to-be-searched images” respectively. Dependent claims 2-7 and 9-14 are rejected for inheriting and not curing the deficiencies of their respective parent claims.
Claim 1 recites, “each of the to-be-searched images corresponds to a first comparison vector ... determining a correlation degree between the second comparison vector and the at least one first comparison vector corresponding to the to-be-searched image to generate a second search result,” (emphasis added). There may be a plurality of “to-be-searched images”, but there is only one recited “first comparison vector” that corresponds to the plurality, meaning there could be one or any number of first comparison vectors. Whether there is one, a plurality, or the exact same number of first comparison vectors as “to-be-searched images”, it remains unclear what the antecedent basis of “the at least one first comparison vector” would be. For purposes of applying prior art, “the at least one first comparison vector” is interpreted as “the image. Claim 8 recites substantially similar limitations, is rejected for the same reasons above, and is interpreted in the same manner. Dependent claims 2-7 and 9-14 are rejected for inheriting and not curing the deficiencies of their respective parent claims.
Claim 1 recites, “at least one classification tag and first location information respectively correspond to a part of the to-be-searched images” and “wherein the second search result comprises a part of the to-be-searched images” (emphasis added). The classification tag and first location are reasonably grouped together as types of metadata. It is unclear how this metadata corresponds to the to-be-searched images because it is also unclear what “a part” of such images would be. A part could be a region within a single image, such as a segmented object region, or a part could be a subset of the images as in one or more whole images amongst the plurality of images. Classification tags and location information can be actual pieces of a metadata file that is part of the same data structure as an image or they could represent a semantic match, meaning the object or place is represented in an image. The second search result is based on a vector similarity, but it remains unclear what “part” of the “to-be-searched images” said result corresponds and what that correspondence is. For purposes of applying prior art, the classification and first location information are interpreted to correspond to an image tag matching process by which a classification tag and location tag are matched to images having the same tags. Additionally, a “part” is interpreted to be a subset of whole images amongst the “to-be-searched images”. Claim 8 recites substantially similar limitations, is rejected for the same reasons above, and is interpreted in the same manner. Dependent claims 2-7 and 9-14 are rejected for inheriting and not curing the deficiencies of their respective parent claims.
Claim 3 recites, “wherein the multi-modal AI model is a multi-modal AI model for images and text” (emphasis added). Claim 1 already provides that the multi-modal AI model is used to obtain a vector from the search string for searching images, but does not explicitly recite “text”. However, a search string in the context of computer processing is a data type that represents text. A string can be empty, but the “search string” in the context of claim 1 must be non-empty, otherwise there would be nothing to match. Thus, the multi-modal AI model of claim 1 is already “for images and text”. Furthermore, the phrase “for images and text” is an intended use of the multi-modal AI model and does not specify any particular features of the model itself, such as being trained to find similarities between text and image modalities. Accordingly, it is unclear how claim 3 changes the scope of claim 1. If claim 3 is intended to mean that the model is used and trained with images and text, then claim 1 would necessarily be broader, meaning the model of claim 1 could have no training or particular features related to images or text, which begs the question of what else it would be designed for. The specification does not appear to contemplate such a model. For purposes of applying prior art, claim 3 is interpreted as not further limiting claim 1, and the model of claim 1 is interpreted as being able to process image and text data. Claim 10 recites substantially similar limitations, is rejected for the same reasons above, and is interpreted in the same manner.
The following is a quotation of 35 U.S.C. 112(d):
(d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
Claims 3 and 10 are rejected under 35 U.S.C. 112(d) as being of improper dependent form for failing to further limit the subject matter of the claim upon which they depend, or for failing to include all the limitations of the claim upon which they depend. Based on the interpretation made for limitations in claims 3 and 10 in the corresponding rejection under 35 U.S.C. 112(b) above, claims 3 and 10 are interpreted as reciting features already included in their parent claims, and therefore, do not further limit the parent claims.
Applicant may cancel the claims, amend the claims to place the claims in proper dependent form, rewrite the claims in independent form, or present a sufficient showing that the dependent claims comply with the statutory requirements.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-14 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
[Claim 1] A method for image search, comprising:
(a) obtaining a search string, wherein the search string is provided to search for one or more of a plurality of to-be-searched images, each of the to-be-searched images corresponds to a first comparison vector, and at least one classification tag and first location information respectively correspond to a part of the to-be-searched images;
(b) determining whether the search string matches one of the at least one classification tag, and while the search string matches the one of the at least one classification tag, generating a first search result based on the to-be-searched image corresponding to the one of the at least one classification tag, and presenting the first search result in a user interface;
(c) wherein while the search string does not match the one of the at least one classification tag, obtaining a second comparison vector corresponding to the search string according to the search string based on a multi-modal artificial intelligence (AI) model,
(d) determining a correlation degree between the second comparison vector and the at least one first comparison vector corresponding to the to-be-searched image to generate a second search result, wherein the second search result comprises a part of the to-be-searched images; and
(e) identifying whether the search string has second location information, so that a third search result is generated based on the second search result and the second location information, and the third search result is presented in the user interface.
[Claim 8] An electronic device, comprising:
(f) a processor;
(g) a storage device provided to store a plurality of to-be-searched images; and
(h) a display device comprising a user interface, wherein the processor:
(performs) limitations (a-e).
Claim Interpretation
Under the broadest reasonable interpretation, the terms of the claims are presumed to have their plain meaning consistent with the specification as it would be interpreted by one of ordinary skill in the art. See MPEP 2111. The preamble of claim 1 specifies that the claim is to a method for searching images. The disclosure, e.g., par. 30, gives an example of an embodiment of a broadly-described electronic device having a processor, storage device and display device for displaying a user interface. Examples of electronic devices include camera-equipped consumer electronic devices such as a smartphone, tablet computer, and notebook computer. See specification at par. 14. The preamble of claim 8 specifies that the claim is to an electronic device, as in the above examples.
Regarding claim element (a), the domain of potential matches to the search string, the plurality of “to-be-searched images”, is a collection of an open-ended number of images that is not required by the claims to be exhaustively searched. Claim 8 includes a storage device in claim element (g) for storing the plurality of to-be-searched images, but the only added concept compared to claim 1 is that the images are stored, which is not implicit from claim 1 because “to-be-searched images” could refer to future images that have not yet been acquired.
Regarding claim element (b), the claims do not place any limits on the type of matching or specifics of the data structure of a tag used in said matching, except for claims 4 and 11, which specify that “a natural language recognition model” is used to determine a match. The claims do not put any limits on how the “first search result” is presented in the user interface.
Regarding claim element (c), the claims do not place any limits on what constitutes “not matching”. Vectors can be correlated in many different ways. Matching can be absolute or based on confidences or probabilities, for example. The claims do not place any limits on the unspecified “multi-modal artificial intelligence (AI) model” or how the “second comparison vector” is obtained “based on” or “according to” such a model other than by the processor as further specified in claim 8. A multi-modal model, in general, is based on at least two modes and the claims do not specify any mode of the model, though text or tag-based matching is implied through context. The disclosure provides a CNN as an example of a “classification AI model”, which is little more than a general link to deep learning methods and not an example of any specific model, architecture or classification scheme. See specification at par. 21. Paragraph 25 provides an example of a “multi-modal AI” that “may be implemented through a Contrastive Language-Image Pre-Training (CLIP) model or a CLIP algorithm”. The same paragraph also describes a “multi-modal AI model” that “consists of two parts: an image model and a text model” (a text encoder and an image encoder). Paragraph 33 mentions Bidirectional Encoder Representations from Transformers (BERT), which does not appear in the claims and the disclosure thereof amounts to a reference to a well-known transformer model that could be integrated into the search string matching in some manner, but no specific details on its implementation would be any different than using a completely different encoder model. None of these specific features are explicitly included in or required by claims 1 and 8.
Regarding claim element (d), similarity or “correlation degree” by vector comparison is broadly recited without any particular distance or similarity metric. Claims 5 and 12 specify a cosine similarity.
Claim element (e) describes initiating a second search based on the first search using “second location information” without putting any limits on what that information is or how it is different from the first location information other than being used in the event the first location does not match the search string. The claims do not specify any further use of the “correlation degree” after it is introduced.
Therefore, the broadest reasonable interpretation of claims 1 and 8 is a method and generic computer having a display to implement a method of conditional matching steps where a text-based query is used to conduct a first search for corresponding tagged images, displaying any results, and if no results are found, performing a second search of the same collection of potentially matching images based on a vector similarity, and performing a third search based on location information in the second search results. Put more simply, the claims describe repeated text and/or vector similarity-based image matching using an implicit computer as in claim 1 or a explicitly recited, but generically described, computer as in claim 8, the computer having a user interface and the ability to receive a text-based query.
Step 1: do the claims fall within any statutory category?
Claim 1 recites a series of steps and therefore, is a process. Claim 8 recites an electronic device and therefore, is a machine. See MPEP 2106.03. (Step 1: YES).
Step 2A, Prong One: do the claims recite a judicial exception?
Overall, the subject matter of the independent claims recite limitations of operating a user interface to search for tagged images based on a text query, where matches are determined by tags associated with each candidate image match and/or a vector-based similarity. The matching determinations in claims 1 and 9, as drafted, are a process that, under their broadest reasonable interpretation, cover performance of the limitations in the mind but for the recitation of generic computer components, namely claim elements (f), (g) and (h) in claim 9, the user interface in each claim, and the vague reference to a multi-modal artificial intelligence (AI) model that is related to obtaining a comparison vector, which could simply be a retrieval operation without any complex processing. For example, but for the “processor” language in claim 9, claims 1 and 9 encompasses the user of the user interface manually constructing an image search based on a text string as input and deciding whether to make subsequent searches based on the results (or lack thereof) of a current search. The search is at least based on tag matching, but a tag is merely data and could be as simple as retrieving photos from a user’s digital library from user-named folders that group their images by some tag or common feature, or performing a simple text search based on the existing metadata that is typically embedded in images produces by consumer electronics like smartphones. The mere nominal recitation of a generic processor does not take the claim limitation out of the mental processes grouping.
Although vector-based similarity is included in claims 1 and 8, the comparison is so broadly recited that it could easily be performed in the mind. For example, a vector is simply an array of data. A user may have one folder for dog pictures and another for cat pictures. If searching with “cat” as the search string and the to-be-searched images are the dog and cat folders, then a vector [dog, cat] of the input string is [0, 1], which has a dot product of 0 ([0,1]*[1,0] = (0×1) + (1×0) = 0) with the dog folder images ([1, 0]) and a dot product of 1 ([0,1]*[0,1] = (0×0) + (1×1) = 1) with the cat folder images ([0, 1]). Thus, claims 1 and 8 recite a mental process. See MPEP 2106.04(a)(2), subsection III. (Step 2A, Prong One: YES).
Step 2A, Prong Two: Do the claims as a whole integrate the recited judicial exception into a practical application of the exception?
Each of claims 1 and 8 is viewed as a whole amounts to more than merely searching for images matching a text search string using a generic computer with a display, as they include up to three searches that are conditioned on the result of each search. However, this amounts to the normal process of searching: sometimes a match is found on the first try and sometimes no match is found which requires broadening the initial search parameters and/or using a different method.
An additional element is found in claim 8 and not claim 1: a processor which performs the initial match as well as the determination of the correlation degree between the first and second comparison vectors of the to-be-searched images. However, the role of the processor is broadly described and could reasonably amount to merely providing the user interface for the user to mark or provide input to enter a text query and navigate through any displayed search results, which is a normal function of a notebook computer, for example. Even if claim 1 or claim 9 were interpreted as requiring a processor to perform vector similarity calculations on its own volition, the claims still fail to recite how a processor-centric solution is accomplished and merely invoke computers as a tool to perform an existing process, for example using the Microsoft Windows-provided calculator to calculate a dot product as a vector similarity. See MPEP 2106.05(f). Here, there are no details how the “multi-modal artificial intelligence (AI) model” functions to provide the user with image searching functionality beyond what their computer is already capable of without any deep learning model or artificial intelligence. The claims omit any significant details as to how the multi-modal AI model solves a technical problem, and instead recite only the idea of a solution or outcome as somehow involving such a model. Therefore, this additional limitation represents no more than mere instructions to apply the judicial exception on a computer. It can also be viewed as nothing more than an attempt to generally link the use of the judicial exception to the technological environment of computers capable of locally running AI models and/or being capable of remotely communicating with such models.
Claims 2 and 9 merely describe how tags are applied to images for subsequent matching and further adds “geographical information” as a type of matching data, but the claims merely amount to further practicing the abstract idea while still only tangentially referring to the multi-modal AI model, with the processor in claim 9 performing standard computer functions of retrieving and storing data.
Claims 3 and 13 include an additional element of finally specifying modalities of the multi-modal AI model: images and text, though those could be implied from context within claims 1 and 8, though claims 1 and 8 do not explicitly require such modalities. Thus claims 3 and 13 merely confine the technological environment of the user interface to that of text-based and image-based searching by generally linking the use of the judicial exception to text and image encoders without specifying any specific use beyond their general functions and being used in a sequence of searches.
Claims 4, 6, 11 and 13 include an additional element of : analyzing the search string based on a natural language model to determine if the search strong contains second location information. Natural Language Processing is considered to be a subfield of AI, not a type of AI in itself, and its development predates modern connotations of “artificial intelligence” systems. NLP is not one technique or process, it is an entire broad field of study and practice. These claims therefore generally link the use of the judicial exception to NLP algorithms without specifying any particular algorithm.
Claims 5 and 12 include an additional element of : the vector similarity is a cosine similarity, which is just the normalized dot product of two vectors and can easily be performed in the mind given the breadth of the claims and the lack of specificity in how the commonly-used cosine similarity is used beyond its normal function of computing the dot product of two vectors divided by their magnitudes to obtain a value, which is then compared to a user-specified threshold to distinguish “similar” from “not similar”.
Claims 7 and 14 specify that the first result from the to-be-searched images is matched via the “classification tag”, which is functionally, no different than any other recited tag because a tag is just data added to other data and serves the same purpose in terms of tag-based matching whether the tag or label is about a classification or a geographic area. Thus claims 7 and 14 merely further practice the abstract idea without reciting additional elements.
The additional elements describe generic computer components (e.g., using a pre-built PC, continuously searching through tagged photos using the built-in search feature) or broadly refer to entire subfields of artificial intelligence or machine learning recited at a high level of generality and that do not amount to any of the relevant considerations for evaluating whether additional limitations integrate a judicial exception into a practical application provided in MPEP 2106.04(d), subsection I.
The additional elements amount to merely including instructions to implement the abstract idea on a computer or merely using a computer as a tool to perform the abstract idea and perform an existing process: use a standard computer that stores or has access to a collection of tagged images to continuously search for matches to a text query. See MPEP 2106.05(f)(2).
The additional elements of “a multi-modal AI model”, “natural language recognition” and the multi-modal AI model implicitly including text and image modalities merely confine the use of the abstract idea to a particular technological environment (ML/AI image searching) and thus fail to add an inventive concept to the claims. See MPEP 2106.05(h).
It should be noted that because the courts have made it clear that mere physicality or tangibility of an additional element or elements is not a relevant consideration in the eligibility analysis, the physical nature of any underlying hardware in the additional elements, e.g., “a multi-modal AI model” and “natural language recognition”, does not affect this analysis. See MPEP 2106.05(I).
The additional elements do not improve the functioning of a computer. See MPEP 2106.04(d)(1). The specification sets forth a technical problem “a technology for searching images, allowing users to use the text description of the image ... and combine with a plurality of artificial intelligence (AI) models (such as classification AI models, image and text multi-modal AI model, natural language recognition model, etc.) and corresponding databases to more accurately find the required image from a large number of images.” (par. 15). Figures 4 and 5 provide a conditional processing flowchart and schematic diagram for a multi-modal AI model respectively. However, the interrelationships between the two is not featured in the claims. Furthermore, the diagram in Figure 5 suggests than “a multi-modal AI model” is merely one which uses a combination of a text encoder and an image encoder and a cosine similarity between their outputs, thereby amounting to a multi-model image search engine, where two of the modalities can be text and images, and the search engine is merely a platform to plug in well-known algorithms and strategies,
Whether evaluated individually or in combination, the additional elements do not integrate the recited judicial exception into a practical application and the independent claims, therefore, are directed to the judicial exception. (Step 2A, Prong Two: NO).
Step 2B: do the claims as a whole amount to significantly more than the judicial exception?
As explained with respect to Step 2A Prong Two, the additional elements of the pending claims amount to performing the abstract idea using a computer as a tool to perform an existing process (retrieve data and manipulate data at a user’s direction), which cannot provide an inventive concept. See MPEP 2106.05(f). Based on the high-level of specify of the technical aspects of the independent and dependent claims as compared to subject matter in the specification, including the drawings, that is in certain aspects more specific and rooted in the technical problems being solved in the additional elements, the additional elements do not constitute an improvement to the functioning of a computer or to another technology because they represent what is well-understood, routine, conventional activity. See MPEP 2106.04(d)(1). For example, compare Figure 1A of U.S. Pat. No. 11,922,550 and Figure 3 of U.S. Pat. Appl. Pub. No. 20200380027 to applicant’s Figure 5.
The additional elements in combination with the judicial exception do not provide an improvement to the functioning of a computer or any other technology or technical field. (Step 2A, Prong Two: NO). Even considering each claim as a whole, the claims do not amount to significantly more than the recited judicial exception and fail to encompass an inventive concept (Step 2B: NO). Claims 1-14, therefore, are not eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-14 are rejected under 35 U.S.C. 103 as being unpatentable over LifeSeeker 4.0: An Interactive Lifelog Search Engine for LSC’22 (Published 27 June 2022) to Nguyen et al (hereinafter “Nguyen”) in view of U.S. Pat. Appl. Pub. No. 20230359651 (filed 18 July 2023) to Mei et al. (hereinafter “Mei”).
Regarding claim 1, Nguyen teaches a method for image search (Nguyen, abstract, “In LifeSeeker 4.0, we focus on enhancing our previous system to allow users who have little to no knowledge of underlying system functioning and lifelog data to use it with ease by not only enhancing the text parser but also employing a Contrastive Language-Image Pre-training (CLIP) model as an extra search mechanism.”), comprising:
obtaining a search string (Nguyen, Figure 1, “eating egg at home from 4pm to 7pm”), wherein the search string is provided to search for one or more of a plurality of to-be-searched images (Nguyen, section 3.1, “Newly released this year, the LSC’22 dataset 1 is a novel multimodal data collection of one lifelogger over the period of 18 months in 2019 and 2020. This dataset contains over 725K egocentric photos in which identifiable faces are redacted and sensitive texts removed to protect personal privacy. Along with those images, visual and textual annotations were extracted by utilising Microsoft Computer Vision API and Google Cloud Vision API respectively. The organisers also provide various metadata, including location, time, biometrics, and music listening history to provide some contextual evidence for the egocentric images.”), each of the to-be-searched images corresponds to a first comparison vector (Each image is represented by features in the shared CLIP space, including visual concepts that are encoded by the CLIP Text Encoder. See Nguyen at section 3.3. The CLIP Text Encoder produces first comparison feature vectors for comparison to the feature vectors produced by the CLIP Image Encoder. See Nguyen at Figure 1.), and at least one classification tag (Nguyen, section 3.2, “visual concepts indexed”) and first location information (Location information, such as GPS coordinates indicating the location of a person’s home, are parsed from the input and used to filter the top ranked matches. See Nguyen at section 3.1 and Figure 2.) respectively correspond to a part of the to-be-searched images (Images matching the parsed concepts and filtered by location and/or time. See Nguyen at Figure 2);
determining whether (i.e., if) the search string matches one of the at least one classification tag (Visual concepts extracted by the text parser are matched to visual concept labels indexed from the LSC’22 dataset. See Nguyen at sections 2, 3.1 and 3.2. Matching similarity between image and text embeddings is determined from the cosine similarity. See Nguyen at section 3.3.), and while (i.e., if) the search string matches the one of the at least one classification tag, generating a first search result based on the to-be-searched image corresponding to the one of the at least one classification tag, and presenting the first search result in a user interface (Nguyen - The example user search in Figure 1 includes the text string “egg” and matching images depicting eggs are presented.), and
wherein while (i.e., if) the search string does not match the one of the at least one classification tag (If users do not already know or happen to guess the actual indexed visual concepts in their text search, the result is missed matches that are semantically equivalent, so the enhanced text parser is provided to account for that no-match scenario, which is one of the main improvements over the prior LifeSeeker version. See Nguyen at section 3.2. Keywords that are not initially matched to the location and time vocabularies become concepts. See id. So, if a user enters a non-indexed, i.e., not directly matched by concept, it is treated as a concept anyway.), obtaining a second comparison vector (The image embeddings are second comparison feature vectors of the indexed images generated by the CLIP Image Encoder. See Nguyen at Figure 2. Nguyen, section 3.2), determining a correlation degree between the second comparison vector and the at least one first comparison vector corresponding to the to-be-searched image (Nguyen, Figure 1, “Cosine Similarity”) to generate a second search result (Nguyen, Figure 2, “Inverted Indices” in the Retrieval stage.), wherein the second search result comprises a part of the to-be-searched images (Before the indices are inverted, the top matches, i.e., a part of the whole collection of possible matches or a second search result, are subsequently filtered by the location and time keyword collections/vocabularies. See Nguyen at section 3.2); and
identifying whether (i.e., if) the search string has second location information (Location keywords are parsed from the query and if matched to a location name collection (vocabulary), they are used for filtering the top matches, thereby producing a third search result. See Nguyen at section 3.2.), so that a third search result is generated based on the second search result and the second location information (Nguyen - The inverted indices are generated based on the second search result’s top matches filtered by the second location data extracted from the query and matched to the location vocabulary.), and the third search result is presented in the user interface (Nguyen, Figure 2, “Inverted Indices”. See also Figure 1.), but does not teach that which is explicitly taught by Mei.
Mei teaches obtaining a second comparison vector (Mei, par. 74, “similar image feature vectors are retrieved from the image feature library by the text feature vector.”, par. 98, “The search text query may be inputted to the text encoder to output the text feature vector, then the similar image feature vectors are retrieved from the image feature vector retrieval set based on the text feature vector, and the corresponding image set is recalled based on the similar image feature vectors.”) corresponding to a search string according to the search string (Mei, pars. 70-71, “S301: Acquire the semantic feature of the first modality data. In an embodiment, the semantic feature of the first modality data may be obtained by processing with a cross-modal search model, specifically, the cross-modal search model includes a first modality processing network ... The first modality processing network is a processing network for the first modality data. Exemplarily, in response to that the first modality data is a text, the first modality processing network may be a text processing network, and the text processing network may be a bidirectional encoder representation from transformers (BERT) model, and may also be other natural language processing (NLP) models. As shown in FIG. 4 a , a schematic diagram of text encoder processing is shown. The text is used as an input, and a text encoder may output a text feature vector.” BERT is considered an AI model.) based on a multi-modal artificial intelligence (AI) model (Mei, par. 29, “Multi-modal learning: the multi-modal learning refers to that data of two different modes are mapped to a same feature space (such as a semantic space) to enable the data of two different modes to be correlated, the modality data with similar semantics has similar features in the feature space, and the data of two different modes, for example, may be an image and a text., par. 85, “The feature extraction network may be a depth model for image processing, such as a convolutional neural network (CNN) model, or a vision transformer (VIT) model for feature extraction. The feature extraction network is a backbone network in the second modality processing network, which is mainly configured to extract the initial feature of the second modality data to for the subsequent network to use.” Vision Transformers (ViTs) and BERT are AI models of different modalities. Together, they form a multi-modal AI model for text and image modalities.).
Nguyen discloses a multi-modal search system implemented on a server that communicates with a user interface to receive a textual query and display matching image results (moments). The matching is determined from a cosine similarity between a feature vector of the textual query and image-based feature vectors associated with archived images, the different modalities of feature vector generated by a pair of CLIP encoders. Matching images are provided through a user interface as a ranked list in a vertically-scrollable panel. The weighting also uses a visual concept dictionary of object labels (classification tags). Nguyen identifies a limitation of their approach: retrieval performance is dependent on the user’s choice of words and the available indexed visual concepts. As such, Nguyen views semantic searching as a substitute for concept-based searching. See Nguyen at section 2, page 15. Thus, Nguyen shows that it was known in the art before the effective filing date of the claimed invention to leverage multiple modalities to improve matching and specifically to prioritize semantic location information over classification labels (tags), which is analogous to the claimed invention in that it is pertinent to the problem being solved by the claimed invention, making image collections easier to search. Mei discloses generating text-modality comparison feature vectors and image-modality feature vectors from a multi-modal AI model ((BERT and NLP) + ViT) and matching input text provided by a user of a computer based on a cross-modality match degree. See Mei at pars. 44, 71, 85, and 94. Mei also discloses a solution that does not depend on fixed category labels and can support more complex text queries by cross-modality matching. See Mei at par. 98. The backbone network is also capable of generating category labels in addition to image feature vectors. See Mei at FIG. 4b. Thus, Mei provides an improvement responsive to the limitation acknowledged by Nguyen and shows that it was known in the art before the effective filing date of the claimed invention to implement cross-modality searches using a multi-modal AI model, to use semantic searching when text/object based searching reaches its limit, and to generate category labels using the same image encoder, which is analogous to the claimed invention in that it is pertinent to the problem being solved by the claimed invention, making image collections easier to search.
A person of ordinary skill in the art would have been motivated to replace the text-encoding network and image-encoding network of Nguyen with a BERT and NLP-based network for encoding text features and a ViT network for encoding image features and generating category labels as disclosed by Mei, and to hard-code the switch to semantic matching in the event an exact match text-based fails as suggested by both Nguyen and Mei, to thereby generate the feature vectors of both modalities using AI models as part of a multi-modal AI mode during training, testing and re-training. Based on the foregoing, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have made such modification according to known methods to yield the predictable results to have the benefit of leveraging NLP processing to extract additional semantic information to improve the results of cross-modal search.
Regarding claim 2, Nguyen in view of Mei teaches the method according to claim 1, further comprising:
querying and obtaining corresponding first geographical information based on geographical location information corresponding to the to-be-searched image (Nguyen, section 3.1, “With the use of the new dataset this year, we re-define some additional metadata following the similar data enhancement process in Lifeseeker 3.0. From the given GPS coordinates, we initiate some supplemental labels that include address, city, and country be-fore clustering those geographic points into 32 primary categories.” Nguyen discloses receiving the LSC’22 dataset and accompanying metadata, including location metadata. See Nguyen at section 3.1), and storing the first geographical information and a corresponding relationship between the to-be-searched image and the first geographical information in a location database (The multimodal LSC’22 dataset images are indexed using the provided visual and textual labels/annotations, and the metadata is searched for related location-based types of metadata to store as a vocabulary/database in association with the indexed images thereby encouraging presented matches to include images with the correct location data specified in the text query. See Nguyen at section 3.1);
generating (By a multi-modal AI model. See Mei at pars. 44, 71, 85, and 94.) the at least one classification tag corresponding to the to-be-searched image based on the to-be-searched image (The visual concept dictionary is built from the multimodal LSC’22 dataset. See Nguyen at section 3.1.), and storing the at least one classification tag and a corresponding relationship between the to-be-searched image and the at least one classification tag in a category database (The visual concept dictionary is a visual object category database. See Nguyen at section 3.1 The visual concepts, by being placed in the concepts database.); and
obtaining the at least one first comparison vector corresponding to each of the to-be-searched images based on each of the to-be-searched images according to the multi-modal AI model (See Mei at pars. 44, 71, 85, and 94), and storing the at least one first comparison vector and a corresponding relationship between the to-be-searched image and the at least one first comparison vector in an image vector database (Mei, Figure 4b, “Category label C” and “image feature vector”).
The rationale for obviousness is the same as provided for claim 1.
Regarding claim 3, Nguyen in view of Mei teaches the method according to claim 1, wherein the multi-modal AI model is a multi-modal AI model for images and text (See Mei at pars. 44, 71, 85, and 94).
The rationale for obviousness is the same as provided for claim 1.
Regarding claim 4, Nguyen in view of Mei teaches the method according to claim 1, wherein the step of identifying whether the search string has the second location information so that the third search result is generated based on the second search result and the second location information comprises:
analyzing the search string based on a natural language recognition model (Mei, par. 74, “NLP”) to determine whether the search string comprises the second location information (Nguyen, section 3.1, “We also construct three different vocabularies in terms of concept, location, and time, used as a filter to refine the search results.” The top k matches that are retained after filtering are an even stronger match.),
determining whether the second location information matches the first location information (Cross-modal similarity. See Nguyen at Figure 2);
wherein when the second location information does not match the first location information (The hard coded-mode switch per the combination in claim 1.), using the second search result as the third search result (Generating the next-best thing to an exact match: a cross-modal semantic match. See Nguyen at Figure 2); and
wherein when the second location information matches the first location information (i.e., an exact match), using the to-be-searched image in the second search result that matches the first location information as the third search result (Displaying the exact match per the workflow in Figure 2 of Nguyen).
The rationale for obviousness is the same as provided for claim 1.
Regarding claim 5, Nguyen in view of Mei teaches the method according to claim 1, wherein the step of determining the correlation degree between the second comparison vector and the at least one first comparison vector corresponding to the to-be-searched image comprises:
calculating a cosine similarity between the second comparison vector and the at least one first comparison vector corresponding to the to-be-searched image (Nguyen, Figure 2, “Cosine Similarity”); and
comparing the cosine similarity with a preset threshold (Implicit by “Top Matches”. See Nguyen at Figure 2) to determine the correlation degree between the second comparison vector and the at least one first comparison vector corresponding to the to-be-searched image.
Regarding claim 6, Nguyen in view of Mei teaches the method according to claim 1, wherein a natural language recognition model (Mei, par. 74, “NLP”) is utilized to identify whether the search string has the second location information .
The rationale for obviousness is the same as provided for claim 1.
Regarding claim 7, Nguyen in view of Mei teaches the method according to claim 1, wherein the first search result is the to-be-searched image corresponding to the one of the at least one classification tag that the search string matches (i.e., an exact concept match with cosine similarity sufficient to be a top match. See Nguyen at Figure 2).
Claims 8-14 substantially correspond to claims 1-7 by reciting an electronic device, comprising: a processor (A computer including a display, a processor and a storage, is required to implement the user interface and underlying processing of LifeSeeker 4.0. See Nguyen at Figure 1), a storage device provided to store a plurality of to-be-searched images (The images must be stored to later be retrieved. See Nguyen at Figure 1), and a display device (Nguyen, section 4, “display continuously”) comprising a user interface (Nguyen, Figure 1, “The user Interface of LifeSeeker 4.0”), wherein the processor implements the steps of the methods of claims 1-7. The rationale for obviousness of each of claims 1-7 is the same as provided or indicated for each corresponding claim 8-14.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RYAN P POTTS whose telephone number is (571)272-6351. The examiner can normally be reached M-F, 9am-5pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sumati Lefkowitz can be reached at 571-272-3638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RYAN P POTTS/Examiner, Art Unit 2672