DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged that application is a National Stage application of PCT PCT/US2022/038839. Priority to PCT/US2022/038839 with a priority date of 7/29/2022 is acknowledged under 35 USC 119(e) and 37 CFR 1.78.
Claim Interpretation
Under MPEP 2143.03, "All words in a claim must be considered in judging the patentability of that claim against the prior art." In re Wilson, 424 F.2d 1382, 1385, 165 USPQ 494, 496 (CCPA 1970). As a general matter, the grammar and ordinary meaning of terms as understood by one having ordinary skill in the art used in a claim will dictate whether, and to what extent, the language limits the claim scope. Language that suggests or makes a feature or step optional but does not require that feature or step does not limit the scope of a claim under the broadest reasonable claim interpretation. In addition, when a claim requires selection of an element from a list of alternatives, the prior art teaches the element if one of the alternatives is taught by the prior art. See, e.g., Fresenius USA, Inc. v. Baxter Int’l, Inc., 582 F.3d 1288, 1298, 92 USPQ2d 1163, 1171 (Fed. Cir. 2009).
Claim 10 recites “at least one” then listing “average vividness, agglomerate size, agglomerate splotchiness, and agglomerate location.” Since “at least one” is disjunctive, any one of the elements found in the prior art is sufficient to reject the claim. While citations have been provided for completeness and rapid prosecution, only one element is required. Because, on balance, it appears the disjunctive interpretation enjoys the most specification support and for that reason the disjunctive interpretation (one of A, B OR C) is being adopted for the purposes of this Office Action. Applicant’s comments and/or amendments relating to this issue are invited to clarify the claim language and the prosecution history.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title.
Claims 1, 12, and 17 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. All of the claims are method claims (1 and 12) or apparatus/machine claim (17) under (Step 1), but under Step 2A all of these claims recite abstract ideas and specifically mental processes—concepts performed in the human mind including observation, evaluation, judgement and opinion which are generally described as a human visually observing a label to judge the locations and dimensions of empty regions in order to insert content into these empty regions; furthermore these mental processes are more particularly:
Recited in claims 1, 12, and 17 as:
performing vividness scoring for a plurality of pixels of the digital image;
determining one or more candidate pixels based on the vividness scoring for the plurality of pixels;
determining at least one subject of the digital image;
It is noted that the above analysis is according to the 2019 Revised Patent Subject Matter Eligibility Guidance published in the Federal Register (84 FR 50) on January 7, 2019 and MPEP 2106.04(a)(2)(III).
Consider also that “If a claim recites a limitation that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper, the limitation falls within the mental processes grouping, and the claim recites an abstract idea” as per MPEP 2106.04(a)(2)(III)(B). See also footnotes 14 and 15 of the Federal Register Notice. As detailed above, the steps of performing vividness scoring, determining candidate pixels, and determining a subject, etc. may be practically performed in the human mind with the use of a physical aid such as a pen and paper (marking the label on the package with a pen).
Under Step 2B, this judicial exception is not integrated into a practical application because claims 1, 12, and 17 do not recite additional elements that integrate the exception into a practical application. The only additional elements processor and computer readable medium are recited at a high level of generality and merely equate to “apply it” or otherwise merely uses a generic computer as a tool to perform an abstract which are not indicative of integration into a practical application as per MPEP 2106.05(f). See also MPEP 2106.04(a)(2)(III) with respect to Mental Processes: “Nor do the courts distinguish between claims that recite mental processes performed by humans and claims that recite mental processes performed on a computer”. See also MPEP 2106.04(a)(2)(III)(C)(3) Using a computer as tool to perform a mental process and MPEP 2106.04(a)(2)(III)(D) as well as the case law cited therein.
In other words, the additional elements and/or are recited at a high level of generality that does not amount to significantly more and/ such that they could practically be performed in the human mind.
For all of the above reasons, taken alone or in combination, claims 1, 12, and 17 recite a non-statutory mental process.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claim 11 recites the limitation "wherein agglomerate splotchiness…". Because Claim 10 recites a Markush group in the alternative ("at least one characteristic selected from the group..."), agglomerate splotchiness is an optional element that may not be present in every embodiment of Claim 10. By further limiting an element that is not guaranteed to exist in the independent scope, claim 11 becomes conditional and indefinite.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 2, 3, 4, 10, 12, 13, 14, 16, 17, 18, and 19 are rejected under 35 U.S.C. 102(a)(1) and (a)(2) as being anticipated by “Personal Photo Enhancement via Saliency Driven Color Transfer,” (Gao et al.).
Claim 1
Regarding claim 1, Gao et al. disclose a method for modifying a digital image, the method comprising: performing vividness scoring for a plurality of pixels of the digital image; ("calculates the saliency value of each pixel sx;y as the Euclidean distance between its color and the average color of the Gaussian blurred image in L*a*b* space," pg. 2, sec. 2.2) determining one or more candidate pixels ("we select the salient super-pixels from the given image," pg. 2, sec. 2.3) based on the vividness scoring for the plurality of pixels; ("we further define two thresholds Tg = 2T and Tl = T to measure global saliency and local saliency of super-pixels," pg. 2, sec. 2.3) agglomerating the one or more candidate pixels into one or more suggested agglomerates; ("clusters pixels to super-pixels based on their similarity in L*a*b* space and spatial proximity," pg. 2, sec. 2.1) determining at least one subject of the digital image; ("super-pixels representation helps to annotate the subject(s) with user interaction," pg. 2, sec. 2.1) removing at least one agglomerate from the one or more
PNG
media_image1.png
480
666
media_image1.png
Greyscale
suggested agglomerates based on at least one of the at least one subject of the digital image or one or more characteristics of the at least one agglomerate; ("For each super-pixel selected as a part of distractive object, [AltContent: textbox (Figure 3 shows the modified image output (e).)]we reduce its attraction by saliency-driven color transfer," pg. 3, sec. 2.3) generating a modified digital image with the one or more suggested agglomerates modified; and outputting the modified digital image.
Claim 2
Regarding claim 2, Gao et al. disclose the method of claim 1, wherein performing vividness scoring for the plurality of pixels includes accessing a mapping between an identified value for each of the plurality of pixels and an associated colorfulness value ("calculates the saliency value of each pixel sx;y as the Euclidean distance between its color and the average color of the Gaussian blurred image in L*a*b* space," pg. 2, sec. 2.2).
Claim 3
Regarding claim 3, Gao et al. disclose the method of claim 1, wherein determining the candidate pixels comprises: determining a set of mask pixels ("we select the salient super-pixels from the given image," pg. 2, sec. 2.3) based on the vividness scoring for the plurality of pixels, wherein mask pixels are pixels with a vividness score above a first threshold and are pixels connected to pixels with a vividness score above a seed threshold; ("we further define two thresholds Tg = 2T and Tl = T to measure global saliency and local saliency of super-pixels," pg. 2, sec. 2.3) and outputting the set of mask pixels as the one or more candidate pixels ("All the super-pixels satisfy the following requirement are selected as the parts of distractive objects: … where Ω is the set of super-pixels annotated as parts of the subjects," pg. 3, sec. 2.3).
Claim 4
Regarding claim 4, Gao et al. disclose the method of claim 1, wherein agglomerating the one or more candidate pixels further comprises: clustering the one or more candidate pixels into one or more super-pixels based on a similar appearance of a subset of the one or more candidate pixels; ("clusters pixels to super-pixels based on their similarity in L*a*b* space and spatial proximity," pg. 2, sec. 2.1) and outputting the one or more super-pixels as the one or more candidate pixels for agglomeration ("SLIC searches the optimal matching pixels from a square neighborhood around each initial seed, and generates super-pixel representation after iterations," pg. 2, sec. 2.1).
Claim 10
Regarding claim 10, Gao et al. disclose the method of claim 1, wherein the one or more characteristics includes at least one characteristic selected from the group of characteristics consisting of average vividness, agglomerate size, agglomerate splotchiness, and agglomerate location ("reduce its attraction by transferring its color to its surrounding super-pixels in L*a*b* space. In color transfer, we keep the color value on L channel of each pixel to avoid artifacts and change the color values on a and b channels of each pixel as follows: [see formulas 7 and 8] where ax,y and bx,y are the color values on a and b channels of pixel px,y; ai and bi are the average color values on a and b channels of all the super-pixels surrounding spi," pg. 3, sec. 2.3)
Claim 12
Regarding claim 12, Gao et al. disclose a non-transitory, computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform a process, the process comprising: ("The proposed method is implemented in Matlab. All of the experiments are carried out on a desktop computer," pg. 3, sec. 3.1) performing vividness scoring for a plurality of pixels pixel of a digital image; ("calculates the saliency value of each pixel sx;y as the Euclidean distance between its color and the average color of the Gaussian blurred image in L*a*b* space," pg. 2, sec. 2.2) determining one or more candidate pixels ("we select the salient super-pixels from the given image," pg. 2, sec. 2.3) based on the vividness scoring for the plurality of pixels; ("we further define two thresholds Tg = 2T and Tl = T to measure global saliency and local saliency of super-pixels," pg. 2, sec. 2.3) agglomerating the one or more candidate pixels into one or more suggested agglomerates; ("clusters pixels to super-pixels based on their similarity in L*a*b* space and spatial proximity," pg. 2, sec. 2.1) determining at least one subject of the digital image; ("super-pixels representation helps to annotate the subject(s) with user interaction," pg. 2, sec. 2.1) removing at least one agglomerate from the one or more suggested agglomerates based on at least one of the at least one subject of the digital image or one or more characteristics of the at least one agglomerate; ("For each super-pixel selected as a part of distractive object, we reduce its attraction by saliency-driven color transfer," pg. 3, sec. 2.3) generating a modified digital image with the one or more suggested agglomerates modified; and outputting the modified digital image (See figure 3).
Claim 13
Regarding claim 13, Gao et al. disclose the non-transitory, computer-readable medium of claim 12, the process further comprising: determining a set of mask pixels ("we select the salient super-pixels from the given image," pg. 2, sec. 2.3) based on the vividness scoring for the plurality of pixels, wherein mask pixels are pixels with a vividness score above a first threshold and are pixels connected to pixels with a vividness score above a seed threshold; ("we further define two thresholds Tg = 2T and Tl = T to measure global saliency and local saliency of super-pixels," pg. 2, sec. 2.3) and outputting the set of mask pixels as the one or more candidate pixels ("All the super-pixels satisfy the following requirement are selected as the parts of distractive objects: … where Ω is the set of super-pixels annotated as parts of the subjects," pg. 3, sec. 2.3).
Claim 14
Regarding claim 14, Gao et al. disclose the non-transitory, computer-readable medium of claim 12, the process further comprising: clustering the one or more candidate pixels into one or more super-pixels based on a similar appearance of a subset of the one or more candidate pixels; ("clusters pixels to super-pixels based on their similarity in L*a*b* space and spatial proximity," pg. 2, sec. 2.1) and outputting the one or more super-pixels as the one or more candidate pixels for agglomeration ("SLIC searches the optimal matching pixels from a square neighborhood around each initial seed, and generates super-pixel representation after iterations," pg. 2, sec. 2.1).
Claim 16
Regarding claim 16, Gao et al. disclose the non-transitory, computer-readable medium of claim 12, wherein the one or more characteristics includes at least one characteristic selected from the group of characteristics consisting of average vividness, agglomerate size, agglomerate splotchiness, and agglomerate location ("reduce its attraction by transferring its color to its surrounding super-pixels in L*a*b* space. In color transfer, we keep the color value on L channel of each pixel to avoid artifacts and change the color values on a and b channels of each pixel as follows: [see formulas 7 and 8] where ax,y and bx,y are the color values on a and b channels of pixel px,y; ai and bi are the average color values on a and b channels of all the super-pixels surrounding spi," pg. 3, sec. 2.3).
Claim 17
Regarding claim 17, Gao et al. disclose a computing system for modifying a digital image, the computing system comprising: one or more processors; and a non-transitory, computer-readable memory comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform a process, the process comprising: ("The proposed method is implemented in Matlab. All of the experiments are carried out on a desktop computer," pg. 3, sec. 3.1) performing vividness scoring for a plurality of pixels of the digital image; ("calculates the saliency value of each pixel sx;y as the Euclidean distance between its color and the average color of the Gaussian blurred image in L*a*b* space," pg. 2, sec. 2.2) determining one or more candidate pixels ("we select the salient super-pixels from the given image," pg. 2, sec. 2.3) based on the vividness scoring for the plurality of pixels; ("we further define two thresholds Tg = 2T and Tl = T to measure global saliency and local saliency of super-pixels," pg. 2, sec. 2.3) agglomerating the one or more candidate pixels into one or more suggested agglomerates; ("clusters pixels to super-pixels based on their similarity in L*a*b* space and spatial proximity," pg. 2, sec. 2.1) determining at least one subject of the digital image; ("super-pixels representation helps to annotate the subject(s) with user interaction," pg. 2, sec. 2.1) removing at least one agglomerate from the one or more suggested agglomerates based on at least one of the at least one subject of the digital image or one or more characteristics of the at least one agglomerate; ("For each super-pixel selected as a part of distractive object, we reduce its attraction by saliency-driven color transfer," pg. 3, sec. 2.3) generating a modified digital image with the one or more suggested agglomerates modified; and outputting the modified digital image (See figure 3).
Claim 18
Regarding claim 18, Gao et al. disclose the computing system of claim 17, the process further comprising: determining a set of mask pixels ("we select the salient super-pixels from the given image," pg. 2, sec. 2.3) based on the vividness scoring for the plurality of pixels, wherein mask pixels are pixels with a vividness score above a first threshold and are pixels connected to pixels with a vividness score above a seed threshold; ("we further define two thresholds Tg = 2T and Tl = T to measure global saliency and local saliency of super-pixels," pg. 2, sec. 2.3) and outputting the set of mask pixels as the one or more candidate pixels ("All the super-pixels satisfy the following requirement are selected as the parts of distractive objects: … where Ω is the set of super-pixels annotated as parts of the subjects," pg. 3, sec. 2.3).
Claim 19
Regarding claim 19, Gao et al. disclose the computing system of claim 17, the process further comprising: clustering the one or more candidate pixels into one or more super-pixels based on a similar appearance of a subset of the one or more candidate pixels; ("clusters pixels to super-pixels based on their similarity in L*a*b* space and spatial proximity," pg. 2, sec. 2.1) and outputting the one or more super-pixels as the one or more candidate pixels for agglomeration ("SLIC searches the optimal matching pixels from a square neighborhood around each initial seed, and generates super-pixel representation after iterations," pg. 2, sec. 2.1).
1st Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 5, 15, and 20 are rejected under 35 U.S.C. 103 as obvious over “Personal Photo Enhancement via Saliency Driven Color Transfer,” (Gao et al.) in view of US Patent Publication 2013 0128021 A1, (Lee et al.).
Claim 5
Regarding claim 5, Gao et al. teach method of claim 1 as noted above.
Gao et al. do not explicitly teach all of wherein determining at least one subject of the digital image includes determining if the digital image includes at least one human subject.
However, Lee et al. teach wherein determining at least one subject of the digital image includes determining if the digital image includes at least one human subject ("The human figure detecting module 32 determines whether one scene image of the number of consecutive scene images captured by the image capturing device 10 includes a human figure," par. 10).
Therefore, taking the teachings of Gao et al. and Lee et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the distractor detection and removal system as taught by Gao et al. to use human subject detection as taught by Lee et al. The suggestion/motivation for doing so would have been that, “The human figure detecting module 32 determines whether one scene image of the number of consecutive scene images captured by the image capturing device 10 includes a human figure based on a human figure detecting technology. The human figure detecting technology is well known” as noted by the Lee et al. disclosure in paragraph [0010], which also motivates combination because the combination would predictably have an additional utility as there is a reasonable expectation that identifying a human figure within the image allows the system to accurately determine the specific type or category of the image based on its contents, thereby ensuring that a human captured in the image is properly recognized as the primary subject of interest rather than being erroneously classified and processed as an unwanted distractor; and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
The rejection of method claim 5 above applies mutatis mutandis to the corresponding limitations of computer-readable medium claim 15 and system claim 20 while noting that the rejection above cites to both device and method disclosures. Claims 15 and 20 are mapped below for clarity of the record and to specify any new limitations not included in claim 5.
Claim 15
Regarding claim 15, Gao et al. teach the non-transitory, computer-readable medium of claim 12 as noted above.
Gao et al. do not explicitly teach all of wherein determining at least one subject of the digital image includes determining if the digital image includes at least one human subject.
However, Lee et al. teach wherein determining at least one subject of the digital image includes determining if the digital image includes at least one human subject ("The human figure detecting module 32 determines whether one scene image of the number of consecutive scene images captured by the image capturing device 10 includes a human figure," par. 10).
Gao et al. and Lee et al. are combined as per claim 1.
Claim 20
Regarding claim 20, Gao et al. teach the computing system of claim 17 as noted above.
Gao et al. do not explicitly teach all of wherein determining at least one subject of the digital image includes determining if the digital image includes at least one human subject.
However, Lee et al. teach wherein determining at least one subject of the digital image includes determining if the digital image includes at least one human subject ("The human figure detecting module 32 determines whether one scene image of the number of consecutive scene images captured by the image capturing device 10 includes a human figure," par. 10).
Gao et al. and Lee et al. are combined as per claim 1.
2nd Claim Rejections - 35 USC § 103
Claims 8 and 9 are rejected under 35 U.S.C. 103 as obvious over “Personal Photo Enhancement via Saliency Driven Color Transfer,” (Gao et al.) and US Patent Publication 2013 0128021 A1, (Lee et al.) in view of “Deep Saliency Prior for Reducing Visual Distraction,” (Aberman et al.) and “Finding Distractors In Images,” (Fried et al.)
Claim 8
Regarding claim 8, Gao et al. and Lee et al. teach method of claim 5 as noted above.
Gao et al. do not explicitly teach all of collecting gaze data associated with the digital image; performing filtering on the gaze data, wherein filtering the gaze data includes at least one of determining early gaze data for the digital image and determining dense gaze data for the digital image; and determining the subject of the digital image based on the gaze data.
However, Aberman et al. teach collecting gaze data associated with the digital image; ("tracks with high accuracy the eye fixation of 20 subjects," pg. 8, sec. 4) and performing filtering on the gaze data, wherein filtering the gaze data includes at least one of determining early gaze data for the digital image and determining dense gaze data for the digital image; ("To ensure that our data is statistically significant we ran a statistical test for three features: (i) gaze saliency within the mask, (ii) consecutive gaze duration within the mask, and (iii) first time gaze stays within the mask for more than 50 ms," pg. 8, sec. 4).
Therefore, taking the teachings of Gao et al., Lee et al., and Aberman et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the distractor detection and removal system as taught by Gao et al. and human subject detection as taught by Lee et al. to use generating saliency maps from gaze data as taught by Aberman et al. The suggestion/motivation for doing so would have been that, “while the research community has so far focused on developing models for predicting where people look, almost no attention has been given to utilizing the knowledge embedded in such recent, deep saliency models to actually drive and direct editing of images and videos, so as to tweak the attention drawn to different regions in them” as noted by the Aberman et al. disclosure on pg. 1, sec. 1, which also motivates combination because the combination would predictably have a greater efficiency as there is a reasonable expectation that generating saliency maps from gaze data to direct the editing of images and videos will yield more accurate, automated, and human-centric identification and suppression of distracting regions than conventional heuristic-based detection systems alone; and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Additionally, Fried et al. teach determining the subject of the digital image based on the gaze data ("Our datasets and annotations give us clues about the properties of distractors and how to detect them. The distractors are, by definition, salient, but not all salient regions are distractors. Thus, previous features used for saliency prediction are good candidate features for our predictor (but not sufficient). We also detect features that distinguish main subjects from salient distractors that might be less important, such as objects near the image boundary," pg. 4, sec. 4.2).
Therefore, taking the teachings of Gao et al., Lee et al., Aberman et al., and Fried et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the distractor detection and removal system as taught by Gao et al., human subject detection as taught by Lee et al., and generating saliency maps from gaze data as taught by Aberman et al. to use subject determination based on gaze data as taught by Fried et al. The suggestion/motivation for doing so would have been that, “The distractors are, by definition, salient, but not all salient regions are distractors. Thus, previous features used for saliency prediction are good candidate features for our predictor (but not sufficient). We also detect features that distinguish main subjects from salient distractors” as noted by the Fried et al. disclosure on pg. 4, sec. 4.2, which also motivates combination because the combination would predictably have a higher accuracy as there is a reasonable expectation that utilizing gaze data provides a direct behavioral indicator of human attention, thereby allowing the system to accurately differentiate between mere salient background distractors and the actual intended main subject of the image; and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Claim 9
Regarding claim 9, Gao et al. and Lee et al. teach method of claim 5 as noted above.
Gao et al. do not explicitly teach all of wherein determining the subject of the digital image based on the gaze data includes: inputting the digital image into a saliency-based artificial intelligence model, the saliency-based artificial intelligence model being trained with images with filtered ground-truth gaze points; receiving an output from the saliency-based artificial intelligence model, the output including a heatmap representing a probability of one or more portions of the digital image containing the subject of the image; and determining the subject of the digital image based on the heatmap.
[AltContent: textbox (Figure 3 shows the saliency heat-map output.)]
PNG
media_image2.png
224
486
media_image2.png
Greyscale
However, Aberman et al. teach inputting the digital image into a saliency-based artificial intelligence model, ("We introduced a novel framework that utilizes the power of a saliency model trained to predict human eye-gaze, to guide a range of editing effects (e.g., recoloring, inpainting, camouflage, semantic object and attribute editing) that result in meaningful changes to visual attention in images," pg. 10, sec. 5) the saliency-based artificial intelligence model being trained with images with filtered ground-truth gaze points; ("It requires a single saliency model that was pre-trained on high-quality eye-gaze tracking data," pg. 3, sec. 2.2) and receiving an output from the saliency-based artificial intelligence model, the output including a heatmap representing a probability of one or more portions of the digital image containing the subject of the image ("a saliency model S(·) that predicts a spatial map (per-pixel value in the range of [0; 1])," pg. 4, sec. 3) .
[AltContent: textbox (Figure 5 shows the segmented image heat-map.)]
PNG
media_image3.png
468
408
media_image3.png
Greyscale
Additionally, Fried et al. teach determining the subject of the digital image based on the heatmap ("Our datasets and annotations give us clues about the properties of distractors and how to detect them. The distractors are, by definition, salient, but not all salient regions are distractors. Thus, previous features used for saliency prediction are good candidate features for our predictor (but not sufficient). We also detect features that distinguish main subjects from salient distractors that might be less important, such as objects near the image boundary," pg. 4, sec. 4.2).
3rd Claim Rejections - 35 USC § 103
Claim 11 is rejected under 35 U.S.C. 103 as obvious over “Personal Photo Enhancement via Saliency Driven Color Transfer,” (Gao et al.) in view of US Patent 7454067 B1, (Pati).
Claim 11
Regarding claim 11, Gao et al. teach method of claim 10 as noted above.
Gao et al. do not explicitly teach all of wherein agglomerate splotchiness is a metric describing the regularity or irregularity of a shape of the agglomerate.
However, Pati teach wherein agglomerate splotchiness is a metric describing the regularity or irregularity of a shape of the agglomerate ("determining a relative arrangement of ON pixels in the cluster, where the determined relative arrangement specifies a shape for the cluster. Assigning one or more scores to the difference image can include assigning a score to each cluster based on the shape of the cluster," col. 2, line 17).
Therefore, taking the teachings of Gao et al. and Pati as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the distractor detection and removal system as taught by Gao et al. to use measuring cluster shape as taught by Pati. The suggestion/motivation for doing so would have been that, “For example, the difference evaluator 114 can assign scores that depend on the shape of clusters in the difference image. The classifier 110 uses the characteristic values assigned to difference images to classify the corresponding image elements. By using characteristic values that depend on cluster shapes, the classifier 110 can recognize that a large but disperse cluster represents a small difference and classify the corresponding image element in the right class” as noted by the Pati disclosure in paragraph [0014], which also motivates combination because the combination would predictably have a higher efficiency as there is a reasonable expectation that measuring and utilizing cluster shapes during difference evaluation improves classification accuracy, thereby reducing false positives and accelerating the identification and removal of distractor elements within the image data; and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Allowable Subject Matter
Claims 6 and 7 are objected to as being dependent upon a rejected base claim but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Reasons for Indicating Allowable Subject Matter
The following is an examiner’s statement of reasons for allowance: The prior art of record does not teach certain distinguishing features as described below in reference to claim 6.
Regarding claim 6, the prior art Lee et al. teaches identifying human subjects in images.
However none teaches: wherein determining at least one subject of the digital image comprises, when the image contains no human subjects, identifying a center of the digital image as the subject of the digital image.
Further, none of the reference teaches or fairly suggests the combination of claimed elements. The Examiner finds no reason or motivation to combine the above references in an obviousness rejection thus placing the claims in condition for allowance.
Claim 7 is objected as depending on an objected claim.
Reference Cited
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure.
US Patent 9870617 A1 to Piekniewski et al. discloses detecting and tracking salient objects in digital images by analyzing pixel and color variations.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KARSTEN F LANTZ whose telephone number is (571) 272-4564. The examiner can normally be reached Monday-Friday 8:00-4:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ms. Jennifer Mehmood can be reached on 571-272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Karsten F. Lantz/Examiner, Art Unit 2664
Date: 9/16/2026
/JENNIFER MEHMOOD/Supervisory Patent Examiner, Art Unit 2664