DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
Claims 1-4, 6-9, 11-14, 16-17, and 19-24 are pending. Claims 5, 10, 15, and 18 are cancelled.
Response to Arguments
Applicant’s arguments, see p. 1, filed 2/13/26, with respect to the abstract have been fully considered and are persuasive. The abstract objections of 11/24/25 have been withdrawn.
Applicant’s arguments, see p. 1-2, filed 2/13/26, with respect to the title of the invention have been fully considered and are persuasive. The objection to the title of the invention of 11/24/25 has been withdrawn.
Applicant’s arguments, see p. 2, filed 2/13/26, with respect to claims 1, 8, 10-11, and 18-22 have been fully considered and are persuasive. The claim objections of 11/24/25 have been withdrawn.
Applicant’s arguments, see p. 2, filed 2/13/26, with respect to claims 19-22 have been fully considered and are persuasive. The 35 U.S.C. 112(f) claim interpretation of 11/24/25 has been withdrawn.
Applicant’s arguments, see p. 2-3, filed 2/13/26, with respect to claims 4, 6-7, 14, and 16-18 have been fully considered and are persuasive. The 35 U.S.C. 112(b) rejections of 11/24/25 have been withdrawn.
Applicant's arguments filed 2/13/26 with respect to the 35 U.S.C. 101 rejections and 35 U.S.C. 103 rejections have been fully considered but they are not persuasive.
First, Applicant argues, in p. 3-4 of the remarks filed 2/13/26, that the prior art of record does not disclose the following limitations in claim 1: “generating a caption in a case where there is an image in the candidate image group to which no caption is given; wherein the analyzing of the captions includes analyzing the caption during generation of the caption by the generating.” The Examiner respectfully disagrees. The prior art of record Krishnan teaches uploading an image, generating a caption for the image, and storing the image and the caption (Fig. II) in which the caption is stored along with the image ID and relevant tags (Fig. IV). Under broadest reasonable interpretation of the claim, the Examiner interprets the initial input image as an image without a caption since a caption is generated for each input image. Additionally, the prior art of record Vinyals teaches using BeamSearch when generating a sentence/caption for an image in which a set of the k best sentences up to time t are iteratively considered as candidates to generate sentences of size t + 1, and only the resulting best k of them are kept (Pg. 3159). Applicant does not recite specifics of the analyzing of the captions, therefore, under broadest reasonable interpretation of the claim, the Examiner interprets the BeamSearch iterative process of keeping only the k best sentences generated as analyzing the caption during generation of it.
Second, in response to applicant’s argument that there is no teaching, suggestion, or motivation to combine the references, the examiner recognizes that obviousness may be established by combining or modifying the teachings of the prior art to produce the claimed invention where there is some teaching, suggestion, or motivation to do so found either in the references themselves or in the knowledge generally available to one of ordinary skill in the art. See In re Fine, 837 F.2d 1071, 5 USPQ2d 1596 (Fed. Cir. 1988), In re Jones, 958 F.2d 347, 21 USPQ2d 1941 (Fed. Cir. 1992), and KSR International Co. v. Teleflex, Inc., 550 U.S. 398, 82 USPQ2d 1385 (2007). In this case, it would have been obvious, before the effective filing date of the claim invention, to one of ordinary skill in the art to combine generating and analyzing captions of images and selecting an image based on the captions as taught by Krishnan with the method of Iguchi in order to find the most similar caption and corresponding image while saving valuable time (Krishnan, Abstract). Also, It would have been obvious, before the effective filing date of the claim invention, to one of ordinary skill in the art to combine analyzing the caption during generation of the caption as taught by Vinyals with the combined method of Iguchi and Krishnan in order to automatically describe the content of an image (Vinyals, Abstract) and provide a better approximation (Vinyals, Pg. 3159).
Third, Applicant argues, in p. 4-5 of the remarks filed 2/13/26, that the claims are not directed to an abstract idea and constitute a practical application. The Examiner respectfully disagrees. The limitations in the independent claims, for example in claim 1, recite concepts that fall under the grouping of abstract idea mental processes, i.e., a concept performed in the human mind, evaluation, judgment, and/or opinion of a human. That is, a person can select a preference for the type of image they want, evaluate the contents in the images, create captions for images, evaluate/analyze the captions for the images, and select an image based on their selected preference and their evaluation/analysis of the images and captions. Additionally, in claim 1, for example, the “obtaining” step is just data gathering/image gathering, and the image processing apparatus, at least one memory, and at least one processor are just generic computer components. These limitations are regarded as adding routine and conventional elements to perform the judicial exception, and do not apply into a practical application. Therefore, contrary to Applicant’s remarks, the additional elements/limitations do not integrate the abstract idea into a practical application, and the claims do not recite non-conventional particular solutions to problems and/or non-conventional particular ways to achieve a desired outcome.
Examiner’s Note
Please note that the claims originally filed on 12/12/22 did not contain “1wherein” in claim 12. It appears that Applicant mistakenly added “1wherein” to claim 12 and subsequently amended it.
Information Disclosure Statement
The information disclosure statement (IDS) submitted 11/21/25 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claim 2 is objected to because of the following informalities: In line 1, “an image processing apparatus” should read –the image processing apparatus–. Appropriate correction is required.
Claim 3 is objected to because of the following informalities: In line 1, “an image processing apparatus” should read –the image processing apparatus–. Appropriate correction is required.
Claim 4 is objected to because of the following informalities: In line 1, “an image processing apparatus” should read –the image processing apparatus–. Appropriate correction is required.
Claim 6 is objected to because of the following informalities: In line 1, “an image processing apparatus” should read –the image processing apparatus–. Appropriate correction is required.
Claim 7 is objected to because of the following informalities: In line 1, “an image processing apparatus” should read –the image processing apparatus–. Appropriate correction is required.
Additionally, Applicant is advised that should claim 1 be found allowable, claim 7 will be objected to under 37 CFR 1.75 as being a substantial duplicate thereof. When two claims in an application are duplicates or else are so close in content that they both cover the same thing, despite a slight difference in wording, it is proper after allowing one claim to object to the other as being a substantial duplicate of the allowed claim. See MPEP § 608.01(m). In this case, both claim 1 and claim 7 recite analyzing captions in which the generated caption is analyzed.
Claim 8 is objected to because of the following informalities: In line 1, “an image processing apparatus” should read –the image processing apparatus–. Appropriate correction is required.
Claim 9 is objected to because of the following informalities: In line 1, “an image processing apparatus” should read –the image processing apparatus–. Appropriate correction is required.
Claim 12 is objected to because of the following informalities: In line 1, “an image processing apparatus” should read –the image processing apparatus–. Appropriate correction is required.
Claim 13 is objected to because of the following informalities: In line 1, “an image processing apparatus” should read –the image processing apparatus–. Appropriate correction is required.
Claim 14 is objected to because of the following informalities: In line 1, “an image processing apparatus” should read –the image processing apparatus–. Appropriate correction is required.
Claim 16 is objected to because of the following informalities: In line 1, “an image processing apparatus” should read –the image processing apparatus–. Appropriate correction is required.
Claim 17 is objected to because of the following informalities: In line 1, “an image processing apparatus” should read –the image processing apparatus–. Appropriate correction is required.
Additionally, Applicant is advised that should claim 11 be found allowable, claim 17 will be objected to under 37 CFR 1.75 as being a substantial duplicate thereof. When two claims in an application are duplicates or else are so close in content that they both cover the same thing, despite a slight difference in wording, it is proper after allowing one claim to object to the other as being a substantial duplicate of the allowed claim. See MPEP § 608.01(m). In this case, both claim 11 and claim 17 recite analyzing captions in which the generated caption is analyzed.
Claim 23 is objected to because of the following informalities:
In line 1, “an image processing apparatus” should read –the image processing apparatus–.
In line 3, “analyzing of the caption” should read –analyzing of the captions–.
Appropriate correction is required.
Claim 24 is objected to because of the following informalities:
In line 1, “an image processing apparatus” should read –the image processing apparatus–.
In line 3, “analyzing of the caption” should read –analyzing of the captions–.
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-4, 6-9, 11-14, 16-17, and 19-24 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite methods (claims 1 and 11), apparatuses (claims 19 and 21), and non-transitory computer readable storage media (claims 20 and 22) for selecting an image for image captioning. With respect to the analysis of Claims 1 and 11:
Step 1:
With regard to Step 1, claims 1 and 11 are directed to a method; and therefore, the claims are directed to one of the statutory categories of inventions.
Step 2A, Prong One:
With regard to Step 2A, Prong One, the limitations in claim 1 “(2) determining a specific condition for preferentially selecting an image from the candidate image group; (3) analyzing the plurality of images in the candidate image group; (4) generating a caption in a case where there is an image in the candidate image group to which no caption is given; (5) analyzing captions attached to the plurality of images in the candidate image group or caption generated by the generating; and (6) selecting a specific image from the candidate image group based on results of (a) the determining of the specific condition, (b) the analyzing of the plurality of images, and (c) the analyzing of the captions, wherein the analyzing of the captions includes analyzing the caption during generation of the caption by the generating” as drafted, recite an abstract idea, such as a process that, under its broadest reasonable interpretation, covers performance of the limitation manually or in the mind by a human. That is, a person can select a preference for the type of image they want, evaluate the contents in the images, create captions for images, evaluate/analyze the captions for the images, and select an image based on their selected preference and their evaluation/analysis of the images and captions. These are concepts that fall under the grouping of abstract idea mental processes, i.e., a concept performed in the human mind, evaluation, judgment, and/or opinion of a human.
With regard to Step 2A, Prong One, the limitations in claim 11 “(2) determining a specific condition for preferentially selecting an image from the candidate image group; (3) generating a caption in a case where there is an image in the candidate image group to which no caption is given; (4) analyzing captions attached to the plurality of images in the candidate image group; and (5) selecting a specific image from the candidate image group based on results of (a) the determining of the specific condition and (b) the analyzing of the captions, wherein the analyzing of the captions includes analyzing the caption during generation of the caption by the generating” as drafted, recite an abstract idea, such as a process that, under its broadest reasonable interpretation, covers performance of the limitation manually or in the mind by a human. That is, a person can select a preference for the type of image they want, create captions for images, evaluate/analyze the captions for the images, and select an image based on their selected preference and their evaluation/analysis of the captions. These are concepts that fall under the grouping of abstract idea mental processes, i.e., a concept performed in the human mind, evaluation, judgment, and/or opinion of a human.
Step 2A, Prong Two:
The 2019 PEG defines the phrase “integration into a practical application” to require an additional element or a combination of additional elements in the claim to apply, rely on, or use the judicial exception. In the instant case, there are no additional steps/elements/limitations in the claims, with the exception of the following in the claims: “image processing apparatus”, “at least one processor”, “at least one memory”, and “obtaining a candidate image group including a plurality of images” in claims 1 and 11, “at least one memory”, “at least one processor”, “an obtaining unit configured to obtain a candidate image group including a plurality of images”, “a determining unit”, “an image analyzing unit”, “caption generating unit”, “caption analyzing unit”, and a “selecting unit” in claim 19, “one or more processors”, “image processing apparatus”, “an obtaining unit configured to obtain a candidate image group including a plurality of images”, “a determining unit”, “an image analyzing unit”, “caption generating unit”, “caption analyzing unit”, and “a selecting unit” in claim 20, “at least one memory”, “at least one processor”, “an obtaining unit configured to obtain a candidate image group including a plurality of images”, “a determining unit”, “caption generating unit”, “caption analyzing unit”, and a “selecting unit” in claim 21, and “one or more processors”, “image processing apparatus”, “an obtaining unit configured to obtain a candidate image group including a plurality of images”, “a determining unit”, “caption generating unit”, “caption analyzing unit”, and “a selecting unit” in claim 22. The “obtaining” step is just data gathering/image gathering. The image processing apparatus, at least one memory, at least one processor/one or more processors, obtaining unit, determining unit, image analyzing unit, caption generating unit, caption analyzing unit, and selecting unit are just generic computer components. These limitations are regarded as adding routine and conventional elements to perform the judicial exception, and do not apply into a practical application. Accordingly, the above-mentioned additional elements/limitations do not integrate the abstract idea into a practical application; and therefore, the claims recite an abstract idea.
Step 2B:
Because the claims fail under Step 2A, the claims are further evaluated under Step 2B. The claims herein do not include additional elements that are sufficient to amount to significantly more than the judicial exception, because as discussed above with respect to integration of the abstract idea into practical application, the additional elements/limitations to perform the steps, amount to no more than insignificant routine and conventional elements. Mere instructions to apply an exception using generic components cannot provide an inventive concept. Therefore, claims 1, 11, 19, 20, 21, and 22 are not patent eligible.
Furthermore, with regard to claims 2-4, 6-9, 12-14, 16-17, and 23-24 viewed individually, these additional steps, under their broadest reasonable interpretation, provide extra-solution activities to cover performance of the limitations as an abstract idea, and do not provide meaningful limitations to transform the abstract idea into a patent eligible application of the abstract idea such that the claims amount to significantly more than the abstract idea itself. Accordingly, they are not patent eligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 6-9, 11-12, 16-17, and 19-24 are rejected under 35 U.S.C. 103 as being unpatentable over Iguchi et al. (US 2018/0165538 A1; hereinafter “Iguchi”) in view of Text-based Image Retrieval Using Captioning by Krishnan et al. (hereinafter “Krishnan”) and further in view of “Show and Tell: A Neural Image Caption Generator” by Vinyals et al. (hereinafter “Vinyals”).
Regarding claim 1, Iguchi teaches, A method of controlling an image processing apparatus, the image processing apparatus including at least one processor and at least one memory, the method comprising (Iguchi, Para. [0028]: operating an album creation application on an image processing apparatus and generating an automatic layout; Iguchi, Para. [0030]: CPU generally controls the image processing apparatus and loads a program stored in a ROM or RAM and executes the program; Iguchi, Fig. 1: CPU 101, ROM 102, RAM 103):
by the at least one processor (Iguchi, Fig. 1: CPU 101; Iguchi, Para. [0030]):
(1) obtaining a candidate image group including a plurality of images (Iguchi, Para. [0037]: an image obtaining unit obtains an image data group of still or moving images);
(2) determining a specific condition for preferentially selecting an image from the candidate image group (Iguchi, Para. [0037]: a priority mode selection unit inputs information that designates whether to preferentially select a person or pet image for an album to be created; Iguchi, Para. [0053]);
(3) analyzing the plurality of images in the candidate image group (Iguchi, Para. [0037]; Iguchi, Para. [0038]: the image analysis unit executes the processes of feature amount obtaining, face detection, expression recognition, and personal recognition);
and (6) selecting a specific image from the candidate image group based on results of (a) the determining of the specific condition, (b) the analyzing of the plurality of imagesIguchi, Para. [0037]: a priority mode selection unit inputs information that designates whether to preferentially select a person or pet image for an album to be created; Iguchi, Paras. [0038]-[0039]: image analysis is performed, and an image scoring unit gives a score to the image data. The image scoring unit can give a higher score to image data including an object in a case in which the object other than persons is recognized based on the priority mode designated by the condition designation unit; Iguchi, Para. [0040]: based on the score, an image selection unit selects an image from the part of the image data group assigned to each double page spread),
.
Iguchi does not expressly disclose the following limitations: (4) generating a caption in a case where there is an image in the candidate image group to which no caption is given; (5) analyzing captions attached to the plurality of images in the candidate image group or captions generated by the generating; and (c) the analyzing of the captions, wherein the analyzing of the captions includes analyzing the caption during generation of the caption by the generating.
However, Krishnan teaches, (4) generating a caption in a case where there is an image in the candidate image group to which no caption is given (Krishnan, Fig. II: an image is uploaded, a caption is generated for the image, and the image and caption are stored; Krishnan: Pgs. 3-4 of article, section A; Krishnan, Fig. IV: captions are stored along with the image ID and relevant tags; Note: since a caption is generated for each input image, the Examiner interprets the initial input image as an image without a caption);
(5) analyzing captions attached to the plurality of images in the candidate image group or captions generated by the generating (Krishnan, Fig. II: a caption is generated for an image and is stored. Then an image is searched for by comparing the input query with each of the captions; Krishnan, Pg. 4 of article, section B: the captions are ranked based on their similarity to the input search query; Krishnan, Pg. 4 of article, IV. Results);
and (c) the analyzing of the captions (Krishnan, Fig. II: a caption is generated for an image and is stored. Then an image is searched for by comparing the input query with each of the captions, and an image corresponding to the closest matching captions are retrieved; Krishnan, Pg. 4 of article, section B; Krishnan, Pg. 4 of article, IV. Results),
.
It would have been obvious, before the effective filing date of the claim invention, to one of ordinary skill in the art to combine generating and analyzing captions of images and selecting an image based on the captions as taught by Krishnan with the method of Iguchi in order to find the most similar caption and corresponding image while saving valuable time (Krishnan, Abstract). Therefore, one of ordinary skill in the art would be capable to have combined the elements as claimed by known methods, and that in combination, each element merely performs the same function as it does separately.
The combination of Iguchi and Krishnan does not expressly disclose the following limitation: wherein the analyzing of the captions includes analyzing the caption during generation of the caption by the generating.
However, Vinyals teaches, wherein the analyzing of the captions includes analyzing the caption during generation of the caption by the generating (Vinyals, Pg. 3159: BeamSearch was used when generating a sentence/caption for an image. During BeamSearch, a set of the k best sentences up to time t are iteratively considered as candidates to generate sentences of size t + 1, and only the resulting best k of them are kept; Note: the Examiner interprets the BeamSearch iterative process of keeping only the k best sentences generated as analyzing the caption during generation of it).
It would have been obvious, before the effective filing date of the claim invention, to one of ordinary skill in the art to combine analyzing the caption during generation of the caption as taught by Vinyals with the combined method of Iguchi and Krishnan in order to automatically describe the content of an image (Vinyals, Abstract) and provide a better approximation (Vinyals, Pg. 3159). Therefore, one of ordinary skill in the art would be capable to have combined the elements as claimed by known methods, and that in combination, each element merely performs the same function as it does separately. It is for at least the aforementioned that the Examiner has reached a conclusion of obviousness with respect to claim 1.
Regarding claim 2, the combination of Iguchi, Krishnan, and Vinyals teaches the limitations as explained above in claim 1.
The combination of Iguchi, Krishnan, and Vinyals further teaches, The method of controlling an image processing apparatus according to claim 1 (see claim 1 above), wherein the specific condition includes a setting of a preferred photographic subject for preferentially selecting the specific image (Iguchi, Para. [0037]: a priority mode selection unit inputs information that designates whether to preferentially select a person or pet image for an album to be created; Iguchi, Para. [0053]).
Regarding claim 6, the combination of Iguchi, Krishnan, and Vinyals teaches the limitations as explained above in claim 1.
The combination of Iguchi, Krishnan, and Vinyals further teaches, The method of controlling an image processing apparatus according to claim 1, wherein, in the generating of the caption, the caption is generated by using a Show and Tell model (Vinyals, Pg. 3156 and Figure 1: the model consists of a vision CNN followed by a language generating RNN and captions/complete sentences are generated; Note: the “show” part of the model is the vision deep CNN and the “tell” part of the model is the language generating RNN).
The proposed combination as well as the motivation for combining the Iguchi, Krishnan, and Vinyals references presented in the rejection of claim 1 apply to claim 6 and are incorporated herein by reference. Therefore, method recited in claim 6 is met by Iguchi, Krishnan, and Vinyals.
Regarding claim 7, the combination of Iguchi, Krishnan, and Vinyals teaches the limitations as explained above in claim 6.
The combination of Iguchi, Krishnan, and Vinyals further teaches, The method of controlling an image processing apparatus according to claim 6 (see claim 6 above), wherein, in the analyzing of the captions, the caption generated in the generating of the caption is also analyzed (Krishnan, Fig. II: a caption is generated for an image and is stored. Then an image is searched for by comparing the input query with each of the captions; Krishnan, Pg. 4 of article, section B: the captions are ranked based on their similarity to the input search query).
The proposed combination as well as the motivation for combining the Iguchi, Krishnan, and Vinyals references presented in the rejection of claim 6 apply to claim 7 and are incorporated herein by reference. Therefore, method recited in claim 7 is met by Iguchi, Krishnan, and Vinyals.
Regarding claim 8, the combination of Iguchi, Krishnan, and Vinyals teaches the limitations as explained above in claim 1.
The combination of Iguchi, Krishnan, and Vinyals further teaches, The method of controlling an image processing apparatus according to claim 1 (see claim 1 above), wherein, in the analyzing of the plurality of images, estimation of a degree of in-focus, face detection, personal recognition, or object determination of each of the plurality of images is performed (Iguchi, Para. [0037]; Iguchi, Para. [0038]: the image analysis unit executes the processes of feature amount obtaining, face detection, expression recognition, and personal recognition; Note: the Examiner considers the face detection and personal recognition limitations).
Regarding claim 9, the combination of Iguchi, Krishnan, and Vinyals teaches the limitations as explained above in claim 1.
The combination of Iguchi, Krishnan, and Vinyals further teaches, The method of controlling an image processing apparatus according to claim 1 (see claim 1 above), wherein the specific condition includes a degree of in-focus, the number of faces, or an object type (Iguchi, Para. [0036]: a person priority or pet priority is set to the object of an image to be employed in an album; Iguchi, Para. [0037]: a priority mode selection unit inputs information that designates whether to preferentially select a person or pet image for an album to be created; Iguchi, Para. [0053]; Note: the Examiner considers the object type limitation).
Regarding claim 11, Iguchi teaches, A method of controlling an image processing apparatus, the image processing apparatus including at least one processor and at least one memory, the method comprising (Iguchi, Para. [0028]: operating an album creation application on an image processing apparatus and generating an automatic layout; Iguchi, Para. [0030]: CPU generally controls the image processing apparatus and loads a program stored in a ROM or RAM and executes the program; Iguchi, Fig. 1: CPU 101, ROM 102, RAM 103):
by the at least one processor (Iguchi, Fig. 1: CPU 101; Iguchi, Para. [0030]):
(1) obtaining a candidate image group including a plurality of images (Iguchi, Para. [0037]: an image obtaining unit obtains an image data group of still or moving images);
(2) determining a specific condition for preferentially selecting an image from the candidate image group (Iguchi, Para. [0037]: a priority mode selection unit inputs information that designates whether to preferentially select a person or pet image for an album to be created; Iguchi, Para. [0053]);
and (5) selecting a specific image from the candidate image group based on results of (a) the determining of the specific condition Iguchi, Para. [0037]: a priority mode selection unit inputs information that designates whether to preferentially select a person or pet image for an album to be created; Iguchi, Paras. [0038]-[0039]: image analysis is performed, and an image scoring unit gives a score to the image data. The image scoring unit can give a higher score to image data including an object in a case in which the object other than persons is recognized based on the priority mode designated by the condition designation unit; Iguchi, Para. [0040]: based on the score, an image selection unit selects an image from the part of the image data group assigned to each double page spread),
.
Iguchi does not expressly disclose the following limitations: (3) generating a caption in a case where there is an image in the candidate image group to which no caption is given; (4) analyzing captions attached to the plurality of images in the candidate image group; and (b) the analyzing of the captions, wherein the analyzing of the captions includes analyzing the caption during generation of the caption by the generating.
However, Krishnan teaches, (3) generating a caption in a case where there is an image in the candidate image group to which no caption is given (Krishnan, Fig. II: an image is uploaded, a caption is generated for the image, and the image and caption are stored; Krishnan: Pgs. 3-4 of article, section A; Krishnan, Fig. IV: captions are stored along with the image ID and relevant tags; Note: since a caption is generated for each input image, the Examiner interprets the initial input image as an image without a caption);
(4) analyzing captions attached to the plurality of images in the candidate image group (Krishnan, Fig. II: a caption is generated for an image and is stored. Then an image is searched for by comparing the input query with each of the captions; Krishnan, Pg. 4 of article, section B: the captions are ranked based on their similarity to the input search query; Krishnan, Pg. 4 of article, IV. Results);
and (b) the analyzing of the captions (Krishnan, Fig. II: a caption is generated for an image and is stored. Then an image is searched for by comparing the input query with each of the captions, and an image corresponding to the closest matching captions are retrieved; Krishnan, Pg. 4 of article, section B; Krishnan, Pg. 4 of article, IV. Results),
.
It would have been obvious, before the effective filing date of the claim invention, to one of ordinary skill in the art to combine generating and analyzing captions of images and selecting an image based on the captions as taught by Krishnan with the method of Iguchi in order to find the most similar caption and corresponding image while saving valuable time (Krishnan, Abstract). Therefore, one of ordinary skill in the art would be capable to have combined the elements as claimed by known methods, and that in combination, each element merely performs the same function as it does separately.
The combination of Iguchi and Krishnan does not expressly disclose the following limitation: wherein the analyzing of the captions includes analyzing the caption during generation of the caption by the generating.
However, Vinyals teaches, wherein the analyzing of the captions includes analyzing the caption during generation of the caption by the generating (Vinyals, Pg. 3159: BeamSearch was used when generating a sentence/caption for an image. During BeamSearch, a set of the k best sentences up to time t are iteratively considered as candidates to generate sentences of size t + 1, and only the resulting best k of them are kept; Note: the Examiner interprets the BeamSearch iterative process of keeping only the k best sentences generated as analyzing the caption during generation of it).
It would have been obvious, before the effective filing date of the claim invention, to one of ordinary skill in the art to combine analyzing the caption during generation of the caption as taught by Vinyals with the combined method of Iguchi and Krishnan in order to automatically describe the content of an image (Vinyals, Abstract) and provide a better approximation (Vinyals, Pg. 3159). Therefore, one of ordinary skill in the art would be capable to have combined the elements as claimed by known methods, and that in combination, each element merely performs the same function as it does separately. It is for at least the aforementioned that the Examiner has reached a conclusion of obviousness with respect to claim 11.
Regarding claim 12, the combination of Iguchi, Krishnan, and Vinyals teaches the limitations as explained above in claim 11.
The combination of Iguchi, Krishnan, and Vinyals further teaches, The method of controlling an image processing apparatus according to claim 11 (see claim 11 above), wherein the specific condition includes a setting of a preferred photographic subject for preferentially selecting the specific image (Iguchi, Para. [0037]: a priority mode selection unit inputs information that designates whether to preferentially select a person or pet image for an album to be created; Iguchi, Para. [0053]).
Regarding claim 16, the combination of Iguchi, Krishnan, and Vinyals teaches the limitations as explained above in claim 11.
The combination of Iguchi, Krishnan, and Vinyals further teaches, The method of controlling an image processing apparatus according to claim 11 (see claim 11 above), wherein, in the generating of the caption, the caption is generated by using a Show and Tell model (Vinyals, Pg. 3156 and Figure 1: the model consists of a vision CNN followed by a language generating RNN and captions/complete sentences are generated; Note: the “show” part of the model is the vision deep CNN and the “tell” part of the model is the language generating RNN).
The proposed combination as well as the motivation for combining the Iguchi, Krishnan, and Vinyals references presented in the rejection of claim 11 apply to claim 16 and are incorporated herein by reference. Therefore, method recited in claim 16 is met by Iguchi, Krishnan, and Vinyals.
Regarding claim 17, the combination of Iguchi, Krishnan, and Vinyals teaches the limitations as explained above in claim 16.
The combination of Iguchi, Krishnan, and Vinyals further teaches, The method of controlling an image processing apparatus according to claim 16 (see claim 16 above), wherein, in the analyzing of the captions, the caption generated in the generating of the caption is also analyzed (Krishnan, Fig. II: a caption is generated for an image and is stored. Then an image is searched for by comparing the input query with each of the captions; Krishnan, Pg. 4 of article, section B: the captions are ranked based on their similarity to the input search query).
The proposed combination as well as the motivation for combining the Iguchi, Krishnan, and Vinyals references presented in the rejection of claim 16 apply to claim 17 and are incorporated herein by reference. Therefore, method recited in claim 17 is met by Iguchi, Krishnan, and Vinyals.
Regarding claim 19, Iguchi teaches, An image processing apparatus comprising (Iguchi, Para. [0030] and Fig. 1: image processing apparatus 100):
at least one memory and at least one processor configured to function as a plurality of units comprising (Iguchi, Para. [0028]: operating an album creation application on an image processing apparatus and generating an automatic layout; Iguchi, Para. [0030]: CPU generally controls the image processing apparatus and loads a program stored in a ROM or RAM and executes the program; Iguchi, Fig. 1: CPU 101, ROM 102, RAM 103):
(1) an obtaining unit configured to obtain a candidate image group including a plurality of images (Iguchi, Fig. 2: image obtaining unit 202; Iguchi, Para. [0037]: “image obtaining unit 202 obtains an image data group of still or moving images”);
(2) a determining unit configured to determine a specific condition for preferentially selecting an image from the candidate image group (Iguchi, Fig. 2: priority mode selection unit 203; Iguchi, Para. [0037]: “priority mode selection unit 203 inputs information that designates whether to preferentially select a person image or a pet image for an album to be created”; Iguchi, Para. [0053]);
(3) an image analyzing unit configured to analyze the plurality of images in the candidate image group (Iguchi, Fig. 2: image analysis unit 205; Iguchi, Para. [0037]; Iguchi, Para. [0038]: “the image analysis unit 205 executes the processes of feature amount obtaining, face detection, expression recognition, and personal recognition”);
and (6) a selecting unit configured to select a specific image from the candidate image group based on results of (a) the determining unit, (b) the image analyzing unitIguchi, Fig. 2: image selection unit 211; Iguchi, Para. [0037]: “priority mode selection unit 203 inputs information that designates whether to preferentially select a person image or a pet image for an album to be created”; Iguchi, Paras. [0038]-[0039]: image analysis is performed in which the image scoring unit 208 performs scoring using information from the image analysis unit 205. The image scoring unit 208 can give a higher score to image data including an object in a case in which the object other than persons is recognized based on the priority mode designated by the condition designation unit 201; Iguchi, Para. [0040]: based on the score, an image selection unit 211 selects an image from the part of the image data group assigned to each double page spread; Iguchi: As shown in Fig. 2, the units (i.e., condition designation unit 201, image obtaining unit 202, priority mode selection unit 203, image analysis unit 205, image scoring unit 208, image selection unit 211) are connected/interact with each other. Output from the units are also input into the image selection unit 211),
.
Iguchi does not expressly disclose the following limitations: (4) a caption generating unit configured to generate a caption in a case where there is an image in the candidate image group to which no caption is given; (5) a caption analyzing unit configured to analyze captions attached to the plurality of images in the candidate image group or captions generated by the caption generating unit; and (c) the caption analyzing unit, wherein the caption analyzing unit includes a unit configured to analyze the caption during generation of the caption by the caption generating unit.
However, Krishnan teaches, (4) a caption generating unit configured to generate a caption in a case where there is an image in the candidate image group to which no caption is given (Krishnan, Pg. 1 of article, I. Introduction: computer is used in image captioning; Krishnan, Fig. II: an image is uploaded, a caption is generated for the image, and the image and caption are stored. Each block of the methodology in Fig. II is interpreted as a unit/software unit; Krishnan: Pgs. 3-4 of article, section A; Krishnan, Fig. IV: captions are stored along with the image ID and relevant tags; Note: since a caption is generated for each input image, the Examiner interprets the initial input image as an image without a caption);
(5) a caption analyzing unit configured to analyze captions attached to the plurality of images in the candidate image group or captions generated by the caption generating unit; (Krishnan, Pg. 1 of article, I. Introduction: computer is used in image captioning; Krishnan, Fig. II: a caption is generated for an image and is stored. Then an image is searched for by comparing the input query with each of the captions. Each block of the methodology in Fig. II is interpreted as a unit/software unit; Krishnan, Pg. 4 of article, section B: the captions are ranked based on their similarity to the input search query; Krishnan, Pg. 4 of article, IV. Results);
and (c) the caption analyzing unit (Krishnan, Pg. 1 of article, I. Introduction: computer is used in image captioning; Krishnan, Fig. II: a caption is generated for an image and is stored. Then an image is searched for by comparing the input query with each of the captions, and an image corresponding to the closest matching captions are retrieved. Each block of the methodology in Fig. II is interpreted as a unit/software unit; Krishnan, Pg. 4 of article, section B; Krishnan, Pg. 4 of article, IV. Results),
.
It would have been obvious, before the effective filing date of the claim invention, to one of ordinary skill in the art to combine generating and analyzing captions of images and selecting an image based on the captions as taught by Krishnan with the method of Iguchi in order to find the most similar caption and corresponding image while saving valuable time (Krishnan, Abstract). Therefore, one of ordinary skill in the art would be capable to have combined the elements as claimed by known methods, and that in combination, each element merely performs the same function as it does separately.
The combination of Iguchi and Krishnan does not expressly disclose the following limitation: wherein the caption analyzing unit includes a unit configured to analyze the caption during generation of the caption by the caption generating unit.
However, Vinyals teaches, wherein the caption analyzing unit includes a unit configured to analyze the caption during generation of the caption by the caption generating unit (Vinyals, Abstract: “we present a generative model based on a deep recurrent architecture that combines recent advances in computer vision and machine translation and that can be used to generate natural sentences describing an image”; Vinyals, Pg. 3159: BeamSearch was used when generating a sentence/caption for an image. During BeamSearch, a set of the k best sentences up to time t are iteratively considered as candidates to generate sentences of size t + 1, and only the resulting best k of them are kept; Note: the Examiner interprets the BeamSearch iterative process of keeping only the k best sentences generated as analyzing the caption during generation of it).
It would have been obvious, before the effective filing date of the claim invention, to one of ordinary skill in the art to combine analyzing the caption during generation of the caption as taught by Vinyals with the combined method of Iguchi and Krishnan in order to automatically describe the content of an image (Vinyals, Abstract) and provide a better approximation (Vinyals, Pg. 3159). Therefore, one of ordinary skill in the art would be capable to have combined the elements as claimed by known methods, and that in combination, each element merely performs the same function as it does separately. It is for at least the aforementioned that the Examiner has reached a conclusion of obviousness with respect to claim 19.
Regarding claim 20, Iguchi teaches, A non-transitory computer-readable storage medium storing program instructions that, when executed by one or more processors of an image processing apparatus, cause the one or more processors to function as a plurality of units comprising (Iguchi, Fig. 1: CPU 101, ROM 102, RAM 103; Iguchi, Para. [0030]: CPU generally controls the image processing apparatus and loads a program stored in a ROM or RAM and executes the program; Iguchi, Para. [0138]: a computer of a system or apparatus reads out and executes computer executable instructions recorded on a storage medium/non-transitory computer-readable storage medium):
an obtaining unit configured to obtain a candidate image group including a plurality of images (Iguchi, Fig. 2: image obtaining unit 202; Iguchi, Para. [0037]: “image obtaining unit 202 obtains an image data group of still or moving images”);
a determining unit configured to determine a specific condition for preferentially selecting an image from the candidate image group (Iguchi, Fig. 2: priority mode selection unit 203; Iguchi, Para. [0037]: “priority mode selection unit 203 inputs information that designates whether to preferentially select a person image or a pet image for an album to be created”; Iguchi, Para. [0053]);
an image analyzing unit configured to analyze the plurality of images in the candidate image group (Iguchi, Fig. 2: image analysis unit 205; Iguchi, Para. [0037]; Iguchi, Para. [0038]: “the image analysis unit 205 executes the processes of feature amount obtaining, face detection, expression recognition, and personal recognition”);
and a selecting unit configured to select a specific image from the candidate image group based on results of (a) the determining unit, (b) the image analyzing unitIguchi, Fig. 2: image selection unit 211; Iguchi, Para. [0037]: “priority mode selection unit 203 inputs information that designates whether to preferentially select a person image or a pet image for an album to be created”; Iguchi, Paras. [0038]-[0039]: image analysis is performed in which the image scoring unit 208 performs scoring using information from the image analysis unit 205. The image scoring unit 208 can give a higher score to image data including an object in a case in which the object other than persons is recognized based on the priority mode designated by the condition designation unit 201; Iguchi, Para. [0040]: based on the score, an image selection unit 211 selects an image from the part of the image data group assigned to each double page spread; Iguchi: As shown in Fig. 2, the units (i.e., condition designation unit 201, image obtaining unit 202, priority mode selection unit 203, image analysis unit 205, image scoring unit 208, image selection unit 211) are connected/interact with each other. Output from the units are also input into the image selection unit 211),
.
Iguchi does not expressly disclose the following limitations: a caption generating unit configured to generate a caption in a case where there is an image in the candidate image group to which no caption is given; a caption analyzing unit configured to analyze captions attached to the plurality of images in the candidate image group or captions generated by the caption generating unit; and (c) the caption analyzing unit, wherein the caption analyzing unit includes a unit configured to analyze the caption during generation of the caption by the caption generating unit.
However, Krishnan teaches, a caption generating unit configured to generate a caption in a case where there is an image in the candidate image group to which no caption is given (Krishnan, Pg. 1 of article, I. Introduction: computer is used in image captioning; Krishnan, Fig. II: an image is uploaded, a caption is generated for the image, and the image and caption are stored. Each block of the methodology in Fig. II is interpreted as a unit/software unit; Krishnan: Pgs. 3-4 of article, section A; Krishnan, Fig. IV: captions are stored along with the image ID and relevant tags; Note: since a caption is generated for each input image, the Examiner interprets the initial input image as an image without a caption);
a caption analyzing unit configured to analyze captions attached to the plurality of images in the candidate image group or captions generated by the caption generating unit; (Krishnan, Pg. 1 of article, I. Introduction: computer is used in image captioning; Krishnan, Fig. II: a caption is generated for an image and is stored. Then an image is searched for by comparing the input query with each of the captions. Each block of the methodology in Fig. II is interpreted as a unit/software unit; Krishnan, Pg. 4 of article, section B: the captions are ranked based on their similarity to the input search query; Krishnan, Pg. 4 of article, IV. Results);
and (c) the caption analyzing unit (Krishnan, Pg. 1 of article, I. Introduction: computer is used in image captioning; Krishnan, Fig. II: a caption is generated for an image and is stored. Then an image is searched for by comparing the input query with each of the captions, and an image corresponding to the closest matching captions are retrieved. Each block of the methodology in Fig. II is interpreted as a unit/software unit; Krishnan, Pg. 4 of article, section B; Krishnan, Pg. 4 of article, IV. Results),
.
It would have been obvious, before the effective filing date of the claim invention, to one of ordinary skill in the art to combine generating and analyzing captions of images and selecting an image based on the captions as taught by Krishnan with the method of Iguchi in order to find the most similar caption and corresponding image while saving valuable time (Krishnan, Abstract). Therefore, one of ordinary skill in the art would be capable to have combined the elements as claimed by known methods, and that in combination, each element merely performs the same function as it does separately.
The combination of Iguchi and Krishnan does not expressly disclose the following limitation: wherein the caption analyzing unit includes a unit configured to analyze the caption during generation of the caption by the caption generating unit.
However, Vinyals teaches, wherein the caption analyzing unit includes a unit configured to analyze the caption during generation of the caption by the caption generating unit (Vinyals, Abstract: “we present a generative model based on a deep recurrent architecture that combines recent advances in computer vision and machine translation and that can be used to generate natural sentences describing an image”; Vinyals, Pg. 3159: BeamSearch was used when generating a sentence/caption for an image. During BeamSearch, a set of the k best sentences up to time t are iteratively considered as candidates to generate sentences of size t + 1, and only the resulting best k of them are kept; Note: the Examiner interprets the BeamSearch iterative process of keeping only the k best sentences generated as analyzing the caption during generation of it).
It would have been obvious, before the effective filing date of the claim invention, to one of ordinary skill in the art to combine analyzing the caption during generation of the caption as taught by Vinyals with the combined method of Iguchi and Krishnan in order to automatically describe the content of an image (Vinyals, Abstract) and provide a better approximation (Vinyals, Pg. 3159). Therefore, one of ordinary skill in the art would be capable to have combined the elements as claimed by known methods, and that in combination, each element merely performs the same function as it does separately. It is for at least the aforementioned that the Examiner has reached a conclusion of obviousness with respect to claim 20.
Regarding claim 21, Iguchi teaches, An image processing apparatus comprising (Iguchi, Para. [0030] and Fig. 1: image processing apparatus 100):
at least one memory and at least one processor configured to function as a plurality of units comprising (Iguchi, Para. [0028]: operating an album creation application on an image processing apparatus and generating an automatic layout; Iguchi, Para. [0030]: CPU generally controls the image processing apparatus and loads a program stored in a ROM or RAM and executes the program; Iguchi, Fig. 1: CPU 101, ROM 102, RAM 103):
(1) an obtaining unit configured to obtain a candidate image group including a plurality of images (Iguchi, Fig. 2: image obtaining unit 202; Iguchi, Para. [0037]: “image obtaining unit 202 obtains an image data group of still or moving images”);
(2) a determining unit configured to determine a specific condition for preferentially selecting an image from the candidate image group (Iguchi, Fig. 2: priority mode selection unit 203; Iguchi, Para. [0037]: “priority mode selection unit 203 inputs information that designates whether to preferentially select a person image or a pet image for an album to be created”; Iguchi, Para. [0053]);
and (5) a selecting unit configured to select a specific image from the candidate image group based on results of (a) the determining unit Iguchi, Fig. 2: image selection unit 211; Iguchi, Para. [0037]: “priority mode selection unit 203 inputs information that designates whether to preferentially select a person image or a pet image for an album to be created”; Iguchi, Paras. [0038]-[0039]: image analysis is performed in which the image scoring unit 208 performs scoring using information from the image analysis unit 205. The image scoring unit 208 can give a higher score to image data including an object in a case in which the object other than persons is recognized based on the priority mode designated by the condition designation unit 201; Iguchi, Para. [0040]: based on the score, an image selection unit 211 selects an image from the part of the image data group assigned to each double page spread; Iguchi: As shown in Fig. 2, the units (i.e., condition designation unit 201, image obtaining unit 202, priority mode selection unit 203, image analysis unit 205, image scoring unit 208, image selection unit 211) are connected/interact with each other. Output from the units are also input into the image selection unit 211),
.
Iguchi does not expressly disclose the following limitations: (3) a caption generating unit configured to generate a caption in a case where there is an image in the candidate image group to which no caption is given; (4) a caption analyzing unit configured to analyze captions attached to the plurality of images in the candidate image group or captions generated by the caption generating unit; and (b) the caption analyzing unit, wherein the caption analyzing unit includes a unit configured to analyze the caption during generation of the caption by the caption generating unit.
However, Krishnan teaches, (3) a caption generating unit configured to generate a caption in a case where there is an image in the candidate image group to which no caption is given (Krishnan, Pg. 1 of article, I. Introduction: computer is used in image captioning; Krishnan, Fig. II: an image is uploaded, a caption is generated for the image, and the image and caption are stored. Each block of the methodology in Fig. II is interpreted as a unit/software unit; Krishnan: Pgs. 3-4 of article, section A; Krishnan, Fig. IV: captions are stored along with the image ID and relevant tags; Note: since a caption is generated for each input image, the Examiner interprets the initial input image as an image without a caption);
(4) a caption analyzing unit configured to analyze captions attached to the plurality of images in the candidate image group or captions generated by the caption generating unit; (Krishnan, Pg. 1 of article, I. Introduction: computer is used in image captioning; Krishnan, Fig. II: a caption is generated for an image and is stored. Then an image is searched for by comparing the input query with each of the captions. Each block of the methodology in Fig. II is interpreted as a unit/software unit; Krishnan, Pg. 4 of article, section B: the captions are ranked based on their similarity to the input search query; Krishnan, Pg. 4 of article, IV. Results);
and (b) the caption analyzing unit (Krishnan, Pg. 1 of article, I. Introduction: computer is used in image captioning; Krishnan, Fig. II: a caption is generated for an image and is stored. Then an image is searched for by comparing the input query with each of the captions, and an image corresponding to the closest matching captions are retrieved. Each block of the methodology in Fig. II is interpreted as a unit/software unit; Krishnan, Pg. 4 of article, section B; Krishnan, Pg. 4 of article, IV. Results),
.
It would have been obvious, before the effective filing date of the claim invention, to one of ordinary skill in the art to combine generating and analyzing captions of images and selecting an image based on the captions as taught by Krishnan with the method of Iguchi in order to find the most similar caption and corresponding image while saving valuable time (Krishnan, Abstract). Therefore, one of ordinary skill in the art would be capable to have combined the elements as claimed by known methods, and that in combination, each element merely performs the same function as it does separately.
The combination of Iguchi and Krishnan does not expressly disclose the following limitation: wherein the caption analyzing unit includes a unit configured to analyze the caption during generation of the caption by the caption generating unit.
However, Vinyals teaches, wherein the caption analyzing unit includes a unit configured to analyze the caption during generation of the caption by the caption generating unit (Vinyals, Abstract: “we present a generative model based on a deep recurrent architecture that combines recent advances in computer vision and machine translation and that can be used to generate natural sentences describing an image”; Vinyals, Pg. 3159: BeamSearch was used when generating a sentence/caption for an image. During BeamSearch, a set of the k best sentences up to time t are iteratively considered as candidates to generate sentences of size t + 1, and only the resulting best k of them are kept; Note: the Examiner interprets the BeamSearch iterative process of keeping only the k best sentences generated as analyzing the caption during generation of it).
It would have been obvious, before the effective filing date of the claim invention, to one of ordinary skill in the art to combine analyzing the caption during generation of the caption as taught by Vinyals with the combined method of Iguchi and Krishnan in order to automatically describe the content of an image (Vinyals, Abstract) and provide a better approximation (Vinyals, Pg. 3159). Therefore, one of ordinary skill in the art would be capable to have combined the elements as claimed by known methods, and that in combination, each element merely performs the same function as it does separately. It is for at least the aforementioned that the Examiner has reached a conclusion of obviousness with respect to claim 21.
Regarding claim 22, Iguchi teaches, A non-transitory computer-readable storage medium storing program instructions that, when executed by one or more processors of an image processing apparatus, cause the one or more processors to function as a plurality of units comprising (Iguchi, Fig. 1: CPU 101, ROM 102, RAM 103; Iguchi, Para. [0030]: CPU generally controls the image processing apparatus and loads a program stored in a ROM or RAM and executes the program; Iguchi, Para. [0138]: a computer of a system or apparatus reads out and executes computer executable instructions recorded on a storage medium/non-transitory computer-readable storage medium):
an obtaining unit configured to obtain a candidate image group including a plurality of images (Iguchi, Fig. 2: image obtaining unit 202; Iguchi, Para. [0037]: “image obtaining unit 202 obtains an image data group of still or moving images”);
a determining unit configured to determine a specific condition for preferentially selecting an image from the candidate image group (Iguchi, Fig. 2: priority mode selection unit 203; Iguchi, Para. [0037]: “priority mode selection unit 203 inputs information that designates whether to preferentially select a person image or a pet image for an album to be created”; Iguchi, Para. [0053]);
and a selecting unit configured to select a specific image from the candidate image group based on results of (a) the determining unit Iguchi, Fig. 2: image selection unit 211; Iguchi, Para. [0037]: “priority mode selection unit 203 inputs information that designates whether to preferentially select a person image or a pet image for an album to be created”; Iguchi, Paras. [0038]-[0039]: image analysis is performed in which the image scoring unit 208 performs scoring using information from the image analysis unit 205. The image scoring unit 208 can give a higher score to image data including an object in a case in which the object other than persons is recognized based on the priority mode designated by the condition designation unit 201; Iguchi, Para. [0040]: based on the score, an image selection unit 211 selects an image from the part of the image data group assigned to each double page spread; Iguchi: As shown in Fig. 2, the units (i.e., condition designation unit 201, image obtaining unit 202, priority mode selection unit 203, image analysis unit 205, image scoring unit 208, image selection unit 211) are connected/interact with each other. Output from the units are also input into the image selection unit 211),
.
Iguchi does not expressly disclose the following limitations: a caption generating unit configured to generate a caption in a case where there is an image in the candidate image group to which no caption is given; a caption analyzing unit configured to analyze captions attached to the plurality of images in the candidate image group or captions generated by the caption generating unit; and (b) the caption analyzing unit, wherein the caption analyzing unit includes a unit configured to analyze the caption during generation of the caption by the caption generating unit.
However, Krishnan teaches, a caption generating unit configured to generate a caption in a case where there is an image in the candidate image group to which no caption is given (Krishnan, Pg. 1 of article, I. Introduction: computer is used in image captioning; Krishnan, Fig. II: an image is uploaded, a caption is generated for the image, and the image and caption are stored. Each block of the methodology in Fig. II is interpreted as a unit/software unit; Krishnan: Pgs. 3-4 of article, section A; Krishnan, Fig. IV: captions are stored along with the image ID and relevant tags; Note: since a caption is generated for each input image, the Examiner interprets the initial input image as an image without a caption);
a caption analyzing unit configured to analyze captions attached to the plurality of images in the candidate image group or captions generated by the caption generating unit (Krishnan, Pg. 1 of article, I. Introduction: computer is used in image captioning; Krishnan, Fig. II: a caption is generated for an image and is stored. Then an image is searched for by comparing the input query with each of the captions. Each block of the methodology in Fig. II is interpreted as a unit/software unit; Krishnan, Pg. 4 of article, section B: the captions are ranked based on their similarity to the input search query; Krishnan, Pg. 4 of article, IV. Results);
and (b) the caption analyzing unit (Krishnan, Pg. 1 of article, I. Introduction: computer is used in image captioning; Krishnan, Fig. II: a caption is generated for an image and is stored. Then an image is searched for by comparing the input query with each of the captions, and an image corresponding to the closest matching captions are retrieved. Each block of the methodology in Fig. II is interpreted as a unit/software unit; Krishnan, Pg. 4 of article, section B; Krishnan, Pg. 4 of article, IV. Results),
.
It would have been obvious, before the effective filing date of the claim invention, to one of ordinary skill in the art to combine generating and analyzing captions of images and selecting an image based on the captions as taught by Krishnan with the method of Iguchi in order to find the most similar caption and corresponding image while saving valuable time (Krishnan, Abstract). Therefore, one of ordinary skill in the art would be capable to have combined the elements as claimed by known methods, and that in combination, each element merely performs the same function as it does separately.
The combination of Iguchi and Krishnan does not expressly disclose the following limitation: wherein the caption analyzing unit includes a unit configured to analyze the caption during generation of the caption by the caption generating unit.
However, Vinyals teaches, wherein the caption analyzing unit includes a unit configured to analyze the caption during generation of the caption by the caption generating unit (Vinyals, Abstract: “we present a generative model based on a deep recurrent architecture that combines recent advances in computer vision and machine translation and that can be used to generate natural sentences describing an image”; Vinyals, Pg. 3159: BeamSearch was used when generating a sentence/caption for an image. During BeamSearch, a set of the k best sentences up to time t are iteratively considered as candidates to generate sentences of size t + 1, and only the resulting best k of them are kept; Note: the Examiner interprets the BeamSearch iterative process of keeping only the k best sentences generated as analyzing the caption during generation of it).
It would have been obvious, before the effective filing date of the claim invention, to one of ordinary skill in the art to combine analyzing the caption during generation of the caption as taught by Vinyals with the combined method of Iguchi and Krishnan in order to automatically describe the content of an image (Vinyals, Abstract) and provide a better approximation (Vinyals, Pg. 3159). Therefore, one of ordinary skill in the art would be capable to have combined the elements as claimed by known methods, and that in combination, each element merely performs the same function as it does separately. It is for at least the aforementioned that the Examiner has reached a conclusion of obviousness with respect to claim 22.
Regarding claim 23, the combination of Iguchi, Krishnan, and Vinyals teaches the limitations as explained above in claim 1.
The combination of Iguchi, Krishnan, and Vinyals further teaches, The method of controlling an image processing apparatus according to claim 1 (see claim 1 above), wherein the generating of the caption also terminates generation-in-progress of the caption in a case where the analyzing of the caption is completed (Vinyals, Pg. 3159: There is a special start word and stop word which designates the start and end of the sentence describing the image. The stop word signals that a complete sentence has been generated. BeamSearch was used when generating a sentence/caption for an image. During BeamSearch, a set of the k best sentences up to time t are iteratively considered as candidates to generate sentences of size t + 1, and only the resulting best k of them are kept; Note: the Examiner interprets the signaling of the stop word that a complete sentence has been generated as terminating generation-in-progress of the caption when caption analysis is completed).
The proposed combination as well as the motivation for combining the Iguchi, Krishnan, and Vinyals references presented in the rejection of claim 1 apply to claim 23 and are incorporated herein by reference. Therefore, method recited in claim 23 is met by Iguchi, Krishnan, and Vinyals.
Regarding claim 24, the combination of Iguchi, Krishnan, and Vinyals teaches the limitations as explained above in claim 11.
The combination of Iguchi, Krishnan, and Vinyals further teaches, The method of controlling an image processing apparatus according to claim 11 (see claim 11 above), wherein the generating of the caption also terminates generation-in-progress of the caption in a case where the analyzing of the caption is completed (Vinyals, Pg. 3159: There is a special start word and stop word which designates the start and end of the sentence describing the image. The stop word signals that a complete sentence has been generated. BeamSearch was used when generating a sentence/caption for an image. During BeamSearch, a set of the k best sentences up to time t are iteratively considered as candidates to generate sentences of size t + 1, and only the resulting best k of them are kept; Note: the Examiner interprets the signaling of the stop word that a complete sentence has been generated as terminating generation-in-progress of the caption when caption analysis is completed).
The proposed combination as well as the motivation for combining the Iguchi, Krishnan, and Vinyals references presented in the rejection of claim 11 apply to claim 24 and are incorporated herein by reference. Therefore, method recited in claim 24 is met by Iguchi, Krishnan, and Vinyals.
Claims 3-4 and 13-14 are rejected under 35 U.S.C. 103 as being unpatentable over Iguchi et al. (US 2018/0165538 A1; hereinafter “Iguchi”) in view of Text-based Image Retrieval Using Captioning by Krishnan et al. (hereinafter “Krishnan”) and further in view of “Show and Tell: A Neural Image Caption Generator” by Vinyals et al. (hereinafter “Vinyals”) and Endo et al. (JP 2017028412 A, see previously provided machine translation; hereinafter “Endo”).
Regarding claim 3, the combination of Iguchi, Krishnan, and Vinyals teaches the limitations as explained above in claim 2.
The combination of Iguchi, Krishnan, and Vinyals does not expressly disclose the following limitation: wherein, in the analyzing of the captions, (1) each of the captions is segmented into words and (2) a subject of the image in the candidate image group is determined.
However, Endo teaches, wherein, in the analyzing of the captions, (1) each of the captions is segmented into words and (2) a subject of the image in the candidate image group is determined (Endo, Para. [0036]: words are extracted (i.e., through parsing) from the descriptions/captions in which the main word is related to the subject of the image; Endo, Paras. [0042]-[0047]: the selected main word determines whether the phrase refers to the particular subject or not (i.e., a person, flowers, room, etc.); Endo, Para. [0058]: the subject of the photographic image is determined from the extracted word).
It would have been obvious, before the effective filing date of the claim invention, to one of ordinary skill in the art to combine segmenting/extracting captions into words and then determining the subject of the image as taught by Endo with the combined method of Iguchi, Krishnan, and Vinyals in order to improve accuracy of automatic image processing (Endo, Abstract). Therefore, one of ordinary skill in the art would be capable to have combined the elements as claimed by known methods, and that in combination, each element merely performs the same function as it does separately. It is for at least the aforementioned that the Examiner has reached a conclusion of obviousness with respect to claim 3.
Regarding claim 4, the combination of Iguchi, Krishnan, Vinyals, and Endo teaches the limitations as explained above in claim 3.
The combination of Iguchi, Krishnan, Vinyals, and Endo further teaches, The method of controlling an image processing apparatus according to claim 3 (see claim 3 above), wherein, in the selecting of the specific image, in the case where the subject of the image in the candidate image group determined in the analyzing of the captions matches the preferred photographic subject, the image in the candidate image group is preferentially selected (Endo, [Para. 0036]: words are extracted (i.e., through parsing) from the descriptions/captions in which the main word is related to the subject of the image; Endo, Paras. [0042]-[0047]: the selected main word determines whether the phrase refers to the particular subject or not (i.e., a person, flowers, room, etc.); Endo, Para. [0058]: the subject of the photographic image is determined from the extracted word; Krishnan, Fig. II: a caption is generated for an image and is stored. Then an image is searched for by comparing the input query with each of the captions, and an image corresponding to the closest matching captions are retrieved; Krishnan, Pg. 4 of article, section B: for retrieving images, captions are ranked based on their similarity to the input search query; Note: the Examiner interprets the image that matches closest to the captions as the image preferentially selected).
The proposed combination as well as the motivation for combining the Iguchi, Krishnan, Vinyals, and Endo references presented in the rejection of claim 3 apply to claim 4 and are incorporated herein by reference. Therefore, method recited in claim 4 is met by Iguchi, Krishnan, Vinyals, and Endo.
Regarding claim 13, the combination of Iguchi, Krishnan, and Vinyals teaches the limitations as explained above in claim 12.
The combination of Iguchi, Krishnan, and Vinyals does not expressly disclose the following limitation: wherein, in the analyzing of the captions, (1) each of the captions is segmented into words and (2) a subject of the image in the candidate image group is determined.
However, Endo teaches, wherein, in the analyzing of the captions, (1) each of the captions is segmented into words and (2) a subject of the image in the candidate image group is determined (Endo, Para. [0036]: words are extracted (i.e., through parsing) from the descriptions/captions in which the main word is related to the subject of the image; Endo, Paras. [0042]-[0047]: the selected main word determines whether the phrase refers to the particular subject or not (i.e., a person, flowers, room, etc.); Endo, Para. [0058]: the subject of the photographic image is determined from the extracted word).
It would have been obvious, before the effective filing date of the claim invention, to one of ordinary skill in the art to combine segmenting/extracting captions into words and then determining the subject of the image as taught by Endo with the combined method of Iguchi Krishnan, and Vinyals in order to improve accuracy of automatic image processing (Endo, Abstract). Therefore, one of ordinary skill in the art would be capable to have combined the elements as claimed by known methods, and that in combination, each element merely performs the same function as it does separately. It is for at least the aforementioned that the Examiner has reached a conclusion of obviousness with respect to claim 13.
Regarding claim 14, the combination of Iguchi, Krishnan, Vinyals, and Endo teaches the limitations as explained above in claim 13.
The combination of Iguchi, Krishnan, Vinyals, and Endo further teaches, The method of controlling an image processing apparatus according to claim 13 (see claim 13 above), wherein, in the selecting of the specific image, in the case where the subject of the image in the candidate image group determined in the analyzing of the captions matches the preferred photographic subject, the image in the candidate image group is preferentially selected (Endo, Para. [0036]: words are extracted (i.e., through parsing) from the descriptions/captions in which the main word is related to the subject of the image; Endo, Paras. [0042]-[0047]: the selected main word determines whether the phrase refers to the particular subject or not (i.e., a person, flowers, room, etc.); Endo, Para. [0058]: the subject of the photographic image is determined from the extracted word; Krishnan, Fig. II: a caption is generated for an image and is stored. Then an image is searched for by comparing the input query with each of the captions, and an image corresponding to the closest matching captions are retrieved; Krishnan, Pg. 4 of article, section B: for retrieving images, captions are ranked based on their similarity to the input search query; Note: the Examiner interprets the image that matches closest to the captions as the image preferentially selected).
The proposed combination as well as the motivation for combining the Iguchi, Krishnan, Vinyals, and Endo references presented in the rejection of claim 13 apply to claim 14 and are incorporated herein by reference. Therefore, method recited in claim 14 is met by Iguchi, Krishnan, Vinyals, and Endo.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
“Image Retrieval based on Similarity of Image Features and Captions” by Hirokazu Ito et al.
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Daniella M. DiGuglielmo whose telephone number is (571)272-0183. The examiner can normally be reached Monday - Friday 8:00 AM - 4:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emily Terrell can be reached at (571)270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Daniella M. DiGuglielmo/Examiner, Art Unit 2666
/EMILY C TERRELL/Supervisory Patent Examiner, Art Unit 2666