DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The title of the invention: “INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND RECORDING MEDIUM” is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
Claim Objections
Claim 7 is objected to because of the following informalities: 1) the seventh line of the claim has improper spacing where the claim limitation “a process of evaluating” begins; this claim limitation should begin on the eighth line following other limitations in the text group obtaining process.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-10 are rejected are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., an abstract idea) without integration into a practical application or recitation of significantly more.
In the analysis below, the method of independent claim 1 is considered representative of independent claims 9 and 10 since all of the independent claims recite identical steps despite being directed to different statutory matter. Furthermore, independent claims 1 and 8-10 are directed to one of the four statutory categories of eligible subject matter (an apparatus for independent claim 1; an apparatus for independent claim 8; a process for independent claim 9; a non-transitory computer readable medium for claim 10); thus, the claims pass Step 1 of the Subject Matter Eligibility Test (See flowchart in MPEP 2106).
Step 2A, prong 1 analysis:
Independent claims 1 and 9-10 are directed to a text group obtaining process of obtaining a visually expressing text group that includes a plurality of texts which visually express a detection target, with reference to input data that specifies the detection target; a prompt generating process of generating a prompt with reference to the visually expressing text group; and a providing process of providing the prompt that has been generated in the prompt generating process, to detect, from the image, a detection target that is specified by the prompt.
Each of the above steps can be performed mentally. In particular, a human comes up with a target object (basketball) they want to detect and identify from a library of images from internet searching as input data to the search engine; they research the object and come up with different visual descriptors of the target object in text and notes words such as “spherical”, “bright orange”, “curved rib channels” etc. and then uses those visual terms to come up with a prompt (i.e. “please show me a ball that is spherical, bright orange, has curved rib channels, and can bounce”) for searching to look through the internet images and to identify the target object in images that correspond to the prompt inputted into the search engine; therefore, this process can all be done mentally.
Independent claim 8 recites a similar claim to independent claim 1 except it specifically obtaining input data that specifies a detection target which is already done by a human in the process above regarding the targe object; therefore, this process can all be done mentally.
As such, the description in independent claims 1 and 8-10 is an abstract idea – namely, a mental process. Accordingly, the analysis under prong one of step 2A of the Subject Matter Eligibility Test does not result in a conclusion of eligibility (See flowchart in MPEP 2106).
Additional elements:
The additional element recited in independent claims 1 and 8-10 are a detection model being a model into which a prompt and an image are inputted.
Step 2A, prong 2 analysis:
The above-identified additional elements do not integrate the judicial exception into a practical application. A detection model being a model into which a prompt and an image are inputted is a generic computer that simply automates what a human has capabilities to do with their own human vision and pen and paper; using different written words as a prompt, a human easily can relate visual descriptor words to identify objects in an image.
Each of the other additional elements (a detection model being a model into which a prompt and an image are inputted) amounts to merely using different devices as tools to perform the claimed mental process. Implementing an abstract idea on a computer or using known generic devices does not integrate a judicial exception into a practical application (See MPEP 2106.05(f)).
Moreover, the additional elements of the claims do not recite an improvement in the functioning of a computer or other technology or technical field, the claimed steps are not performed using a particular machine, the claimed steps do not effect a transformation, and the claims do not apply the judicial exception in any meaningful way beyond generically linking the use of the judicial exception to a particular technological environment (See MPEP 2106.04(d)). Therefore, the analysis under prong two of step 2A of the Subject Matter Eligibility Test does not result in a conclusion of eligibility (See flowchart in MPEP 2106).
Step 2B:
Finally, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Each of the other additional elements (a detection model being a model into which a prompt and an image are inputted) are generic computer features which perform generic computer functions that are well-understood, routine, and conventional and do not amount to more than implementing the abstract idea with a computerized system. Thus, taken alone, the additional elements do not amount to significantly more than the above-identified judicial exception (the abstract idea).
Looking at the limitations as an ordered combination adds nothing that is not already present when looking at the elements taken individually. There is no indication that the combination of elements improves the functioning of a computer or improves any other technology. Their collective functions merely provide conventional computer implementation, and mere implementation on a generic computer does not add significantly more to the claims. Accordingly, the analysis under step 2B of the Subject Matter Eligibility Test does not result in a conclusion of eligibility (See flowchart in MPEP 2106).
For all of the foregoing reasons, independent claims 1 and 8-10 do not recite eligible subject matter under 35 USC 101.
Claim 2 recites wherein the prompt generating process includes an evaluating process of evaluating appropriateness of at least any text included in the visually expressing text group. The human evaluates all visual descriptor words about the target object to see if they are correct or not depending on their research on the object (for example: “green” would not do with at target object being a basketball typically) therefore, this process can all be done mentally.
Claim 3 recites wherein the prompt generating process includes a selecting process of selecting, from the visually expressing text group, one or more texts to be used to generate the prompt, with reference to a result of the evaluating process. The human has numerous descriptor words found during research of the target object and uses main visual identifiers to look at images and determine if the object is in the image (terms “orange” and “black lines” and “ball” would be selected to recognize basketballs in images by a human); therefore, this process can all be done mentally.
Claim 4 recites wherein the prompt generating process includes a searching process of searching for a text other than the one or more texts that have been selected in the selecting process, as an additional text to be used to generate the prompt. The human has numerous descriptor words found during research of the target object and uses main visual identifiers to look at images and determine if the object is in the image (terms “orange” and “black lines” and “ball” would be selected to recognize basketballs in images by a human); however, the human additionally uses an additional term “sphere” that will guide evaluating the images with the prompt even more optimally; ); therefore, this process can all be done mentally.
Claim 5 recites wherein the prompt generating process includes a process of generating the prompt from the one or more texts that have been selected in the selecting process, in a case where the additional text is not found in the searching process. The human has numerous descriptor words found during research of the target object and uses only main visual identifiers to look at images and determine if the object is in the image (terms “orange” and “black lines” and “ball” would be selected to recognize basketballs in images by a human while terms like “curved rib channels”, while true visual descriptors of a basketball, are ignored) while ignoring visual identifiers that are do not help looking for the object in images with human vision; therefore, this process can all be done mentally.
Claim 6 recites wherein: the text group obtaining process includes a process of generating a plurality of visually expressing text groups with use of a plurality of generation models that differ from each other; and the evaluating process includes a process of evaluating the plurality of visually expressing text groups with use of a plurality of evaluation models that differ from each other. Each “generation model” is a generic computer and is akin to simply using multiple humans for the task that can detect objects in images with their eyes and determine language prompts to aid in determining if an image has the target object multiple humans are different evaluators since they have different visual acuity; therefore, this process can all be done mentally.
Claim 7 recites wherein the text group obtaining process includes a process of generating a first text group with use of a first generation model, a process of generating a second text group with use of a second generation model, a process of evaluating the second text group with use of a first evaluation model that includes the first generation model, and a process of evaluating the first text group with use of a second evaluation model that includes the second generation model. Each “generation model” is a generic computer and is akin to simply using multiple humans for the task that can detect objects in images with their eyes and determine language prompts to aid in determining if an image has the target object multiple humans are different evaluators since they have different visual acuity; once the visual descriptor text groups are generated by two humans, the first human evaluates the second human’s researched visual descriptor text groups and vice versa; the humans verify each other’s research on their text prompts for helping find the target object in images; therefore, this process can all be done mentally.
Therefore, dependent claims 2-7 recite the same abstract idea of a mental process which can be performed in the mind with the aid of pen and paper, and are therefore also rejected under 35 U.S.C. 101.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-6 and 8-10 are rejected under 35 U.S.C. 103 as being unpatentable over International Patent Application Publication No.: WO 2025101175 A1 (Avinash et al.) (hereinafter Avinash), in view of Chinese Patent Application Publication No.: CN 117437209 A (Wang et al.) (hereinafter Wang).
Regarding claim 1, Avinash teaches an information processing apparatus comprising at least one processor, the at least one processor executing: (Avinash, para. [0243]; FIG. 17: “Computing device 50 can include one or more processors 51 and a memory 52. Processor(s) 51 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. Memory 52 can include one or more non-transitory computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. Memory 52 can store data 53 and instructions 54 which can be executed by processor(s) 51 to cause computing device 50 to perform operations. The operations can implement any one or multiple features described herein. The operations can implement example methods and techniques described herein.”;
PNG
media_image1.png
289
265
media_image1.png
Greyscale
)
a text group obtaining process of obtaining a visually expressing text group that includes a plurality of texts which visually express a detection target, with reference to input data that specifies the detection target; a prompt generating process of generating a prompt with reference to the visually expressing text group (Avinash, para. [0073]; para. [0078]; para. [0046] FIG. 2: “Figure 2 depicts one or more language models 220 receiving an input concept 102 and sending queries 224, which can be based on the input concept 102, to one or more vision language models 222. The vision language model(s) 222 can then generate one or more responses 226 to the queries 224 based on one or more images 106. The language model(s) 220 can then generate one or more image labels 110 based on the responses 226.”; “Figure 2 depicts the language models 220 generating one or more queries 224 to send to one or more vision language models 222. In some instances, generating the queries 224 can comprise extracting one or more queries 224 from the input concept 102. For example, in some instances where an input concept 102 is obtained via a multi-tum conversation with a user, a user can be asked to provide examples of image types (e.g. “bluefin nigiri”) that are associated with the input concept 102 (or not associated with the input concept 102); in such instances, the user’s examples may be extracted and used as queries. In some instances, generating the queries 224 can comprise generating a textual output (e.g. token or sequence of tokens) based on the input concept 102 (e.g. via autoregressive token generation) … In some instances, generating the queries 224 can comprise determining one or more in-scope image attributes (i.e. image attributes that would make an image likely to belong to a user- defined category) or one or more out-of-scope image attributes. In some instances, generating the queries 224 can comprise generating questions based on one or more in-scope or out-of- scope image attributes (e.g. “is there a table with four legs in this image?”).”; “For example, a large language model can generate a set of queries to prompt one or more vision language models.”
PNG
media_image2.png
626
858
media_image2.png
Greyscale
;
Avinash teaches that an input concept 102 that specifies some sort of detection target (ex: furniture) is inputted (referred to in Avinash as “user-defined category”) into a language model (LM) or large language model (LLM) 220 that outputs queries 224 which, in the example discussed above in para. [0078] for a input concept of “furniture”, could yield the visual expressing categories of table and four legs that is then expressed in a prompt question format “is there a table with four legs in this image?”; table and four legs are two visual expressive categories referring back to the original input concept of “furniture”’; and the prompt refers back to the two visual expressing categories so the steps of obtaining a visually expressing text group and generating a prompt are simultaneously done by the language model 220); and
a providing process of providing, to a model, the prompt that has been generated in the prompt generating process, the model being a model into which a prompt and an image are inputted (Avinash, para. [0081]; para. [0083]: “Figure 2 depicts vision language models 222. The vision language models 222 can comprise, for example, any machine-learned models configured to process (e.g. input or output) both language data (e.g. textual data) and image data. In some instances, the vision language models 222 can comprise a vision language model configured for visual question answering (e.g. answering a question about an input image). In some instances, the vision language models 222 can comprise a vision language model configured for visual captioning (e.g. producing a verbose description of an input image) … In some instances, the vision language models 222 can comprise one or more models that process text and images using the same module. In some instances, a vision language model 220 can leverage an attention mechanism such as self-attention. For example, a vision language model 220 can comprise one or more multi-headed self-attention models (e.g. encoder-only, encoder-decoder, or decoder-only transformer language model).”; Figure 2 depicts the vision language models 222 generating the responses 226. In some instances, generating the responses can comprise prompting the vision language models 222 with the queries 224 and an image 106, and receiving a textual response 226 based on the prompt.”; as seen in FIG. 2 above, the query/prompt 224 and an image 106 are inputted into the vision-language model 222 that outputs text stating whether the object discussed in the prompt is in the image or not).
Avinash fails to teach
a providing process of providing, to a detection model, the prompt that has been generated in the prompt generating process, the detection model being a model into which a prompt and an image are inputted and which detects, from the image, a detection target that is specified by the prompt.
Wang teaches
a providing process of providing, to a detection model, the prompt that has been generated in the prompt generating process, the detection model being a model into which a prompt and an image are inputted and which detects, from the image, a detection target that is specified by the prompt (Wang, page 7, para. 3-5; page 8, para. 1-4; page 17, para. 4; FIG. 3: “Step 110, inputting the image to be detected and the prompt text corresponding to the image to be detected into a focusing network to obtain at least one candidate image block in the image to be detected output by the focusing network, wherein the prompt text comprises information related to a target to be detected, and the focusing network is obtained after training an initial focusing network based on an image sample and a prompt text sample corresponding to the image sample. Specifically, the prompt text may include any information related to the object to be detected, for example, may be information indicating a position of the object to be detected in the image, may be information indicating a shape, a color, a texture, a size, or the like of the object to be detected, and may be scene information related to the object to be detected … the target prompt text is determined based on the type of the image to be detected, and the target prompt text is determined to be the prompt text corresponding to the image to be detected … The focusing network may be a neural network model or algorithm obtained after training the initial focusing network based on the image sample and the prompt text sample corresponding to the image sample. Because the focusing network is based on the image sample and the prompt text sample corresponding to the image sample during training, the initial focusing network for training can learn the text information corresponding to the image information besides the information contained in the image sample … The candidate image block may be an image area determined from the image to be detected, and optionally, the image area determined from the image to be detected is subjected to image segmentation to obtain at least one candidate image block. And 120, detecting targets to be detected in each candidate image block to obtain target detection results. Specifically, after at least one candidate image block is obtained, object detection may be performed based on each candidate image block to detect an object to be detected.”; “Fig. 3 is a schematic structural diagram of an object detection device according to an embodiment of the present invention”;
PNG
media_image3.png
1076
502
media_image3.png
Greyscale
).
It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the providing process of providing, to a model, the prompt that has been generated in the prompt generating process, the model being a model into which a prompt and an image are inputted, as taught by Avinash, to include providing, to a detection model, the prompt that has been generated in the prompt generating process, the detection model being a model into which a prompt and an image are inputted and which detects, from the image, a detection target that is specified by the prompt, as taught by Wang.
The suggestion/motivation for doing so would have been that “input of the prompt text is beneficial to introducing priori knowledge which is useful for specific application scenes in the target detection task, and the prompt text can be used for effectively guiding the focusing network to focus the target to be detected, so that the target can be conveniently identified or extracted” (Wang, page 7, para. 5).
Therefore, it would have been obvious to combine Avinash, with Wang, to obtain the invention as specified in claim 1.
Regarding claim 2, Avinash, in view of Wang, teaches the information processing apparatus as set forth in claim 1, wherein the prompt generating process includes an evaluating process of evaluating appropriateness of at least any text included in the visually expressing text group (Avinash, para. [0092]; FIG. 3: “In some instances, generating one or more positive queries 332 and one or more negative queries 334 can include a query expansion step or query mutation step. For example, in some instances a first plurality of positive queries 332 and first plurality of negative queries 334 can be created (e.g. based on an input context, feedback from a user, etc.). One or more language models 330 can then generate a second plurality of positive queries 332 or negative queries 334 based on the first plurality of positive queries 332 and negative queries 334. In some instances, a query expansion or mutation step can include prompting a language model 330 with one or more instructions to generate additional queries 332, 334 that are different from existing queries 332, 334 in specified ways (e.g. “expand the query to be much more general,” “narrow the query down to a specific subcase,” “completely change the [first few/last few/middle] words of the query,” “make sure to create a query with an almost entirely different meaning; for example, the query does not need to share any words with the original and doesn’t have to be directly related.”). In some instances, multiple query expansion or query mutation steps may be performed recursively until a satisfactory set of queries 332, 334 has been generated (e.g. a second query mutation step can generate a third plurality of queries 332, 334 based on a second plurality of queries 332, 334, etc.). In some instances, a set of generated queries 332, 334 can be satisfactory when a desired number of queries 332, 334 (e.g. 50, 100, 200, 500, etc.) has been generated. In some instances, a set of queries 332, 334 can be satisfactory when a desired statistical measure is achieved (e.g. a measure of semantic diversity of the generated queries 332, 334 in relation to a desired size of a training dataset). In some example experiments according to the present disclosure, a query expansion step can be used to generate 100 additional queries 332, 334 based on an initial query set of 13 queries 332, 334. In some examples, the language model 330 or other query expansion tool may not necessarily have knowledge or receive an indication of whether the queries being expanded are “positive” or “negative” queries, but such knowledge may be applied by the larger system after expansion.”;
PNG
media_image4.png
579
725
media_image4.png
Greyscale
).
Regarding claim 3, Avinash, in view of Wang, teaches the information processing apparatus as set forth in claim 2, wherein the prompt generating process includes a selecting process of selecting, from the visually expressing text group, one or more texts to be used to generate the prompt, with reference to a result of the evaluating process (Avinash, para. [0161]-[0173]; FIG. 12: “Figure 12 is a block diagram of an example implementation of an example machine-learned model configured to process sequences of information … For example, some example sequence processing models in the text domain are referred to as “Large Language Models,” or LLMs. In general, sequence processing model(s) 4 can obtain input sequence 5 using data from input(s) 2 … Sequence processing model(s) 4 can ingest the data from input(s) 2 and parse the data into a sequence of elements to obtain input sequence 5. For example, a portion of input data from input(s) 2 can be broken down into pieces that collectively represent the content of the portion of the input data … For example, for textual input source(s), the elements can correspond to groups of one or more words or sub-word components, such as sets of one or more characters. Prediction layer(s) 6 can predict one or more output elements 7-1, 7-2, . . . , 7- N based on the input elements. Prediction layer(s) 6 can include one or more machine-learned model architectures, such as one or more layers of learned parameters that manipulate and transform the input(s) to extract higher-order meaning from, and relationships between, input element(s) 5-1, 5-2, . . . , 5-M. In this manner, for instance, example prediction layer(s) 6 can predict new output element(s) in view of the context provided by input sequence 5. Prediction layer(s) 6 can evaluate associations between portions of input sequence 5 and a particular output element. These associations can inform a prediction of the likelihood that a particular output follows the input context. For example, consider the textual snippet, “The carpenter’s toolbox was small and heavy. It was full of .” Example prediction layer(s) 6 can identify that “It” refers back to “toolbox” by determining a relationship between the respective embeddings. Example prediction layer(s) 6 can also link “It” to the attributes of the toolbox, such as “small” and “heavy.” Based on these associations, prediction layer(s) 6 can, for instance, assign a higher probability to the word “nails” than to the word “sawdust.” Output sequence 7 can include or otherwise represent the same or different data types as input sequence 5. For instance, input sequence 5 can represent textual data, and output sequence 7 can represent textual data … For instance, for some applications, an output of one or more prediction layer(s) 6 can be passed through one or more output layers (e.g., softmax layer) to obtain a probability distribution over an output vocabulary (e.g., a textual or symbolic vocabulary) conditioned on a set of input elements in a context window.
PNG
media_image5.png
782
680
media_image5.png
Greyscale
;
Avinash teaches using a language model that both generates the visual expressing groups and generates the prompts; while creating the visual expressing groups that describe the object described by the input concept, the language model also weights identified visual expressive terms associated with the concept differently depending on how close the visual terms are associated with the main concept object; these steps are I the language model that can also be incorporated into the query expansion step discussed above in the rejection of claim 2, where each prompt output is reevaluated and resubmitted again into the language model which then optimizes the visual expressive terms (descriptors of the object in the input concept); this builds in an evaluation process of the visual terms used in the prompt and after iterations of the query expansion step an optimized prompt having the most accurate visually descriptive words of the object named in the original input concept can be input into the visual-language model (VLM) to output an accurate indication that the prompt question identifies the object in the model).
Regarding claim 4, Avinash, in view of Wang, teaches the information processing apparatus as set forth in claim 3, wherein the prompt generating process includes a searching process of searching for a text other than the one or more texts that have been selected in the selecting process, as an additional text to be used to generate the prompt (Avinash, para. [0161]-[0173]; FIG. 12; Avinash, para. [0161]-[0173]; see rejection of claims 2-3 above; Avinash discusses an process of inputting an input concept of an object (words) into a language model, outputting a prompt with different weighted visual expressing terms that visually describe the object in the input concept, and doing an iterative query expansion of the generated prompt to further optimize the terms used in the prompt; this includes keeping visual expressive terms previously generated and/or optimizing the prompt multiple times so there are new additional visual expressing terms added to the prompt that explain how the object in the original input concept is described; this process optimizes generating a prompt for inputting into a vision-language model that can classify the object from the input concept based on the prompt).
Regarding claim 5, Avinash, in view of Wang, teaches the information processing apparatus as set forth in claim 4, wherein the prompt generating process includes a process of generating the prompt from the one or more texts that have been selected in the selecting process, in a case where the additional text is not found in the searching process (Avinash, para. [0161]-[0173]; FIG. 12; Avinash, para. [0161]-[0173]; see rejection of claims 2-3 above; Avinash discusses an process of inputting an input concept of an object (words) into a language model, outputting a prompt with different weighted visual expressing terms that visually describe the object in the input concept, and doing an iterative query expansion of the generated prompt to further optimize the terms used in the prompt; this includes keeping visual expressive terms previously generated and/or optimizing the prompt multiple times so there are new additional visual expressing terms added to the prompt that explain how the object in the original input concept is described; this process optimizes generating a prompt for inputting into a vision-language model that can classify the object from the input concept based on the prompt; in the situation where there are no more additional descriptive words of the input concept object, the query expansion has found the most optimized prompt, without any additional visual text found in the searching (query expansion) process); the more specific an object that is described in the main concept the more likely the searching process will not add any additional visual expressing text groups beyond what the language model initially outputs with the prompt).
Regarding claim 6, Avinash, in view of Wang, teaches the information processing apparatus as set forth in claim 2, wherein: the text group obtaining process includes a process of generating a plurality of visually expressing text groups with use of a plurality of generation models that differ from each other; and the evaluating process includes a process of evaluating the plurality of visually expressing text groups with use of a plurality of evaluation models that differ from each other (Avinash, (Avinash, para. [0092]; see rejection of claim 2 above; para. [0075]: “In some instances, the language models 220 can comprise two or more language models. In some instances, the language models 220 can comprise a supervisor language model and a worker language model. In such instances, the worker language model can have an architecture that can be the same as, different from, based on, or the basis for an architecture of the supervisor language model … According to some example experiments, the supervisor language model and worker language model can each be a decoder-only transformer model adapted to dialogue applications (e.g. conversation). In some instances, a supervisor language model 220 can be configured to correct one or more outputs of a worker model 220. In some instances, a supervisor language model 220 may be configured (e.g. prompted) to detect whether a worker language model 220 has failed to follow one or more prompt instructions, or has generated output with one or more errors (e.g. factual errors). In such instances, a supervisor language model 220 can in some instances be configured to generate a corrected output (e.g. a corrected query 224).”; two different language models can be used in place of single language model 220 shown in FIG. 2; a worker language model outputs visual expressing text groups in prompts and the supervisor language model outputs a corrected output for finetuning/correcting the queries/prompts output from the worker language model; they are both programmed/trained to have different semantic levels of accuracy for outputting different prompts with different visually expressing text groups; the two language models also operate as two different evaluation models in the query expansion described in the rejection of claim 2 above when the iterative process of evaluating the prompts multiple times to refine the prompts and the visual expressing text groups used in the prompts to be more optimally describing the visual nature of the object in the input concept).
Regarding claim 8, Avinash teaches an information processing apparatus comprising at least one processor, the at least one processor executing: (Avinash, para. [0243]; FIG. 17: “Computing device 50 can include one or more processors 51 and a memory 52. Processor(s) 51 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. Memory 52 can include one or more non-transitory computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. Memory 52 can store data 53 and instructions 54 which can be executed by processor(s) 51 to cause computing device 50 to perform operations. The operations can implement any one or multiple features described herein. The operations can implement example methods and techniques described herein.”;
PNG
media_image1.png
289
265
media_image1.png
Greyscale
)
an obtaining process of obtaining input data that specifies a detection target; and the prompt that is provided by the at least one processor in the providing process being generated in a process including: a text group obtaining process of obtaining a visually expressing text group that includes a plurality of texts which visually express the detection target; and a prompt generating process of generating a prompt with reference to the visually expressing text group (Avinash, para. [0073]; para. [0078]; para. [0046] FIG. 2: “Figure 2 depicts one or more language models 220 receiving an input concept 102 and sending queries 224, which can be based on the input concept 102, to one or more vision language models 222. The vision language model(s) 222 can then generate one or more responses 226 to the queries 224 based on one or more images 106. The language model(s) 220 can then generate one or more image labels 110 based on the responses 226.”; “Figure 2 depicts the language models 220 generating one or more queries 224 to send to one or more vision language models 222. In some instances, generating the queries 224 can comprise extracting one or more queries 224 from the input concept 102. For example, in some instances where an input concept 102 is obtained via a multi-tum conversation with a user, a user can be asked to provide examples of image types (e.g. “bluefin nigiri”) that are associated with the input concept 102 (or not associated with the input concept 102); in such instances, the user’s examples may be extracted and used as queries. In some instances, generating the queries 224 can comprise generating a textual output (e.g. token or sequence of tokens) based on the input concept 102 (e.g. via autoregressive token generation) … In some instances, generating the queries 224 can comprise determining one or more in-scope image attributes (i.e. image attributes that would make an image likely to belong to a user- defined category) or one or more out-of-scope image attributes. In some instances, generating the queries 224 can comprise generating questions based on one or more in-scope or out-of- scope image attributes (e.g. “is there a table with four legs in this image?”).”; “For example, a large language model can generate a set of queries to prompt one or more vision language models.”
PNG
media_image2.png
626
858
media_image2.png
Greyscale
;
Avinash teaches that an input concept 102 that specifies some sort of detection target (ex: furniture) is inputted (referred to in Avinash as “user-defined category”) into a language model (LM) or large language model (LLM) 220 that outputs queries 224 which, in the example discussed above in para. [0078] for a input concept of “furniture”, could yield the visual expressing categories of table and four legs that is then expressed in a prompt question format “is there a table with four legs in this image?”; table and four legs are two visual expressive categories referring back to the original input concept of “furniture”’; and the prompt refers back to the two visual expressing categories so the steps of obtaining a visually expressing text group and generating a prompt are simultaneously done by the language model 220); and
a providing process of providing, to a model, a prompt that is obtained with reference to the input data, the model being a model into which a prompt and an image are inputted (Avinash, para. [0081]; para. [0083]: “Figure 2 depicts vision language models 222. The vision language models 222 can comprise, for example, any machine-learned models configured to process (e.g. input or output) both language data (e.g. textual data) and image data. In some instances, the vision language models 222 can comprise a vision language model configured for visual question answering (e.g. answering a question about an input image). In some instances, the vision language models 222 can comprise a vision language model configured for visual captioning (e.g. producing a verbose description of an input image) … In some instances, the vision language models 222 can comprise one or more models that process text and images using the same module. In some instances, a vision language model 220 can leverage an attention mechanism such as self-attention. For example, a vision language model 220 can comprise one or more multi-headed self-attention models (e.g. encoder-only, encoder-decoder, or decoder-only transformer language model).”; Figure 2 depicts the vision language models 222 generating the responses 226. In some instances, generating the responses can comprise prompting the vision language models 222 with the queries 224 and an image 106, and receiving a textual response 226 based on the prompt.”; as seen in FIG. 2 above, the query/prompt 224 and an image 106 are inputted into the vision-language model 222 that outputs text stating whether the object discussed in the prompt is in the image or not).
Avinash fails to teach
a providing process of providing, to a detection model, a prompt that is obtained with reference to the input data, the detection model being a model into which a prompt and an image are inputted and which detects, from the image, a detection target that is specified by the prompt.
Wang teaches
a providing process of providing, to a detection model, a prompt that is obtained with reference to the input data, the detection model being a model into which a prompt and an image are inputted and which detects, from the image, a detection target that is specified by the prompt (Wang, page 7, para. 3-5; page 8, para. 1-4; page 17, para. 4; FIG. 3: “Step 110, inputting the image to be detected and the prompt text corresponding to the image to be detected into a focusing network to obtain at least one candidate image block in the image to be detected output by the focusing network, wherein the prompt text comprises information related to a target to be detected, and the focusing network is obtained after training an initial focusing network based on an image sample and a prompt text sample corresponding to the image sample. Specifically, the prompt text may include any information related to the object to be detected, for example, may be information indicating a position of the object to be detected in the image, may be information indicating a shape, a color, a texture, a size, or the like of the object to be detected, and may be scene information related to the object to be detected … the target prompt text is determined based on the type of the image to be detected, and the target prompt text is determined to be the prompt text corresponding to the image to be detected … The focusing network may be a neural network model or algorithm obtained after training the initial focusing network based on the image sample and the prompt text sample corresponding to the image sample. Because the focusing network is based on the image sample and the prompt text sample corresponding to the image sample during training, the initial focusing network for training can learn the text information corresponding to the image information besides the information contained in the image sample … The candidate image block may be an image area determined from the image to be detected, and optionally, the image area determined from the image to be detected is subjected to image segmentation to obtain at least one candidate image block. And 120, detecting targets to be detected in each candidate image block to obtain target detection results. Specifically, after at least one candidate image block is obtained, object detection may be performed based on each candidate image block to detect an object to be detected.”; “Fig. 3 is a schematic structural diagram of an object detection device according to an embodiment of the present invention”;
PNG
media_image3.png
1076
502
media_image3.png
Greyscale
).
It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the a providing process of providing, to a model, a prompt that is obtained with reference to the input data, the model being a model into which a prompt and an image are inputted, as taught by Avinash, to include providing, to a detection model, a prompt that is obtained with reference to the input data, the detection model being a model into which a prompt and an image are inputted and which detects, from the image, a detection target that is specified by the prompt, as taught by Wang.
The suggestion/motivation for doing so would have been that “input of the prompt text is beneficial to introducing priori knowledge which is useful for specific application scenes in the target detection task, and the prompt text can be used for effectively guiding the focusing network to focus the target to be detected, so that the target can be conveniently identified or extracted” (Wang, page 7, para. 5).
Therefore, it would have been obvious to combine Avinash, with Wang, to obtain the invention as specified in claim 8.
With regards to independent claim 9, it recites the functions of the apparatus of independent claim 1, as a process. Thus, the analysis in rejecting claim 1 is equally applicable to claim 8.
Regarding claim 10, Avinash teaches a non-transitory recording medium in which a program for causing a computer (Avinash, para. [0243]: “Computing device 50 can include one or more processors 51 and a memory 52. Processor(s) 51 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. Memory 52 can include one or more non-transitory computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. Memory 52 can store data 53 and instructions 54 which can be executed by processor(s) 51 to cause computing device 50 to perform operations. The operations can implement any one or multiple features described herein. The operations can implement example methods and techniques described herein.”).
With regards to the remaining limitations of claim 10, they recite the functions of the apparatus of claim 1, as a non-transitory computer readable medium storing a program. Thus, the analysis in rejecting claim 1 is equally applicable to the remaining limitations of claim 10.
Conclusion
Claim 7 was rejected under 35 U.S.C. 101 but not rejected under prior art 35 U.S.C. 102 and 103.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: U.S. Patent Application Publication No.: 2024/0311652 (Kulkarni et al.) that teaches refining text prompts to be more descriptive of objects.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL ADAM SHARIFF whose telephone number is 571-272-9741. The examiner can normally be reached M-F 8:30-5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sumati Lefkowitz can be reached on 571-272-3638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL ADAM SHARIFF/
Examiner, Art Unit 2672