DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS), filed 11/25/2024, is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Specification
The disclosure is objected to because of the following informalities:
¶0043 of the specification recites “As shown bin Fig. 2, …”. The examiner believes this was meant to recite “As shown in Fig. 2, …”
Appropriate correction is required.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are:
Claim 18: “one or more devices configured to:
receive user input that indicates…
generate a probability map…
identify an object of interest….”
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
The following limitations are computer-implemented means-plus-function limitations, wherein the corresponding structure and associated algorithm as disclosed in the specification for each limitation are being interpreted (see MPEP § 2181(II)(B)).
“one or more devices” (device 400 of Fig. 4, process 500 of Fig. 5, process 600 of Fig. 6, process 700 of Fig. 7; ¶0017-18, 25, 43-45, 68-76, 86-89, 97-100)
The associated structure of the one or more devices is device 400 of Fig. 4, including processor 420 for executed the associated algorithm(s).
The associated algorithm for performing the claimed function is as follows: the device 400 is described as being configured to perform any one or more of the processes 500, 600, or 700. In process 500, the device is configured to receive one or more reference images with an indication of an object of interest (block 510), generating a probability map associated with a target based on soft matching scores (block 520), and then identifying the object of interest in the target image based on the probability map (block 530). More specifically this soft matching is performed using frozen feature extraction backbone with probabilistic feature correspondence derived via Optimal Transport utilizing a quadratic cosine similarity matrix as a cost matrix. A pretrained segmentation model for translating course user inputs into reference masks. Reference images are featurized utilizing either a convolutional neural network or vision transformer (ViT) for soft matching. Process 600 is performed similarly, but instead with block 610 specifying receiving an indication of a reference object of interest in one or more reference images. Process 700 functions similarly, but with block 710 specifying that a user indicates a reference object of interest in the one or more reference images.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 U.S.C. § 112(b)
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 4 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 4 describe an algorithm for generating a probability map has a computational complexity “less than or equal to a product of a size in pixels of the target image and respective sizes in pixels of the reference objects.” The term “size in pixels” is unclear in the context of the specification [¶0014; 43] when discussing the use of optimal transport approach for soft feature matching. The examiner believes the applicant is referencing the size of an object in a reference image or the size of the target image in terms of the number of pixels.
35 USC § 101 – Directed to Statutory Subject Matter
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Concerning claim 12, typically, the broadest reasonable interpretation of a “computer readable storage media” would be considered to include transitory media in the form signals or carrier waves, which would render the claim non-statutory under 35 U.S.C. § 101. That being said, in ¶0049 of the originally filed specification, the applicant defines “storage media” to also be referred to as “mediums”, which is later defined “to not be construed as storage in the form of transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, light pulses passing through a fiber optic cable, electrical signals communicated through a wire, and/or other transmission media”. This clarifies that the claim, in light of the specification, cannot encompass a transitory.
Therefore, claim 12 is considered to pertain to statutory subject matter, and is patent eligible under 35 U.S.C. § 101.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1, 3-9 & 12-16 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Liu et al (“Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching”, 2024, ArXiv), hereinafter referred to as “Liu”.
Regarding claim 1, Liu teach A method (the method of Matcher [Sec 3: Method; outlined in Fig. 1]) comprising:
receiving one or more reference images respectively having an indication of a reference object of interest within the one or more reference images (a reference image (xr) is provided with a mask (mr) indicating an object to be segmented in one shot [Sec 3: Method - ¶01; Figs. 1 & 2]);
generating a probability map associated with a target image based at least in part on matching scores that are associated with the reference object (patch-level features (zr, zt) of both reference (xr) and target (xt) images are extracted to compute a similarity between matching regions in the form of a correspondence matrix (S) [Sec 3.1: Correspondence Matrix Extraction; Fig. 1]– the examiner notes that that the disclosed correspondence matrix is analogous to the probabilistic feature matching as outlined in ¶0017-18 of the specification of the instant application); and
identifying an object of interest within the target image based at least in part on the probability map (Matcher then segments objects or parts of the target image using the correspondence matrix to generate a course segmentation mask over the object to be identified in the target image [Sec 3.2: Prompts Generation - ¶01; Figs. 1 & 2]).
Considering claim 3, Liu teach The method of claim 1 (as described above), further comprising:
providing, after identifying the object of interest, an indication of the object of interest that is identified within the target image (a segmentation mask is overlayed over the identified object [Sec 3.2: Prompts Generation - ¶01; Figs. 1 & 2]).
As for claim 4, Liu teach The method of claim 1 (as described previously), wherein an algorithm associated with generation of the probability map has a computation complexity that is less than or equal to a product of a size in pixels of the target image and respective sizes in pixels of the reference objects (Matcher leverages an optimal transport approach for dense semantic feature matching when determining mask relevance based on the correspondence matrix [Sec 3.3: Controllable Masks Generation - ¶01-03 & Sec: Appendix A - ¶01-02; Fig. 2] – the examiner notes that the optimal transport approach described in Liu is analogous to the optimal transport approach employed by the instant application [¶0021-22; 43] for computationally efficient matching).
With respect to claim 5, Liu teach The method of claim 1 (as described previously), wherein identifying the object of interest comprises:
sampling candidate locations of the target image and identifying the matching scores to identify the object of interest (image patches are compared between features of the reference (xr) and target (xt) images to discovery the best matching regions defined by the cosine similarity of the correspondence matrix (S) [Sec 3.1: Correspondence Matrix Extraction - ¶01; Fig. 1]).
Concerning claim 6, Liu teach The method of claim 1 as described previously), further comprising obtaining one or more segmentation masks associated with the one or more reference images based at least in part on the indication of the object of interest (a reference image (xr) is provided with a mask (mr) indicating an object is obtained for one-shot segmentation [Sec 3: Method - ¶01; Figs. 1 & 2]),
wherein the matching scores are based at least in part on application of the one or more segmentation masks to the target image (the correspondence matrix (S) is utilized to define the mask outputs of the object of interest [Sec 3.2: Prompt Generation & Sec 3.3: Controllable Masks Generation; Fig. 1]).
Turning to claim 7, The method of claim 6 (described above), wherein obtaining the one or more segmentation masks comprises:
applying a pretrained segmentation model (a segment anything model (SAM) with ViT-H is used for segmentation, utilizing 1B masks and 11M images for pretraining [Sec 4.1: Experiments Setting - ¶01]).
As for claim 8, Liu teach The method of claim 6 (as previously described), further comprising:
receiving an input that indicates a requested modification of the one or more segmentation masks (during instance-level matching, the correspondence matrix is utilized to select high-quality masks from the mask proposals [Sec 3.3: Controllable Masks Generation - ¶01; Fig. 1] – here, the examiner is interpreting the input from the correspondence matrix for selecting masks from the mask proposals as a form of a modification request); and
modifying the one or more segmentation masks based at least in part on the input (the selected high-quality masks are then utilized and merged to obtain a final target mask [Sec 3.3: Controllable Masks Generation - ¶01; Fig. 1]).
Considering claim 9, Liu teach The method of claim 1 (as described previously), further comprising:
receiving the indication of the reference object of interest after receiving the one or more reference images (a mask (mr) indicates the object to be segmented in the reference image (xr) [Sec 3: Method - ¶01; Figs. 1 & 2]).
Regarding claim 12, Liu teach A computer program product (Matcher, a computer vision framework for one shot semantic segmentation [Sec 1: Introduction - ¶03-04]) comprising:
one or more computer readable storage media (GPU cluster [Sec 5: Conclusion – Subsec: Acknowledgements] – a GPU contains one or more processors or memories for storing and implementing computer-readable instructions), and program instructions collectively stored on the one or more computer readable storage media (GPU cluster [Sec 5: Conclusion – Subsec: Acknowledgements]), the program instructions comprising:
program instructions to receive an indication of a reference object of interest within one or more reference images (a reference image (xr) is provided with a mask (mr) indicating an object to be segmented in one shot [Sec 3: Method - ¶01; Figs. 1 & 2]);
program instructions to generate a probability map associated with a target image based at least in part on matching scores that are associated with the reference object (patch-level features (zr, zt) of both reference (xr) and target (xt) images are extracted to compute a similarity between matching regions in the form of a correspondence matrix (S) [Sec 3.1: Correspondence Matrix Extraction; Fig. 1]– the examiner notes that that the disclosed correspondence matrix is analogous to the probabilistic feature matching as outlined in ¶0017-18 of the specification of the instant application); and
program instructions to identify an object of interest within the target image based at least in part on the probability map (Matcher then segments objects or parts of the target image using the correspondence matrix to generate a course segmentation mask over the object to be identified in the target image [Sec 3.2: Prompts Generation - ¶01; Figs. 1 & 2]).
Considering claim 13, Liu teach The computer program product of claim 12 (described above), wherein the program instructions comprise:
program instructions to provide, after identifying the object of interest, an indication of the object of interest that is identified within the target image (a segmentation mask is overlayed over the identified object [Sec 3.2: Prompts Generation - ¶01; Figs. 1 & 2]).
As for claim 14, Liu teach The computer program product of claim 12 (as described previously), wherein to identify the object of interest, the program instructions comprise:
program instructions to sample candidate locations of the target image and identify the matching scores to identify the object of interest (image patches are compared between features of the reference (xr) and target (xt) images to discovery the best matching regions defined by the cosine similarity of the correspondence matrix (S) [Sec 3.1: Correspondence Matrix Extraction - ¶01; Fig. 1]).
With respect to claim 15, Liu teach The computer program product of claim 12 (as described previously), wherein the program instructions comprise program instructions to obtain one or more segmentation masks associated with the one or more reference images based at least in part on the indication of the object of interest (a reference image (xr) is provided with a mask (mr) indicating an object is obtained for one-shot segmentation [Sec 3: Method - ¶01; Figs. 1 & 2]),
wherein the matching scores are based at least in part on application of the one or more segmentation masks to the target image (the correspondence matrix (S) is utilized to define the mask outputs of the object of interest [Sec 3.2: Prompt Generation & Sec 3.3: Controllable Masks Generation; Fig. 1]).
Turning to claim 16, Liu teach The computer program product of claim 15, wherein the program instructions comprise:
program instructions to receive an input that indicates a requested modification of the one or more segmentation masks (during instance-level matching, the correspondence matrix is utilized to select high-quality masks from the mask proposals [Sec 3.3: Controllable Masks Generation - ¶01; Fig. 1] – here, the examiner is interpreting the input from the correspondence matrix for selecting masks from the mask proposals as a form of a modification request); and
program instructions to modify the one or more segmentation masks based at least in part on the input (the selected high-quality masks are then utilized and merged to obtain a final target mask [Sec 3.3: Controllable Masks Generation - ¶01; Fig. 1]).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 2 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al (“Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching”, 2024, ArXiv), hereinafter referred to as “Liu” , in view of Xie et al (US 2026/0024278 A1), hereinafter referred to as “Xie”.
Regarding claim 2, Liu teach The method of claim 1 (as described previously), however, Lie fails to teach modifying the target image.
Xie, on the other hand, is analogous art pertinent to the field of endeavor and disclose a method for generating different object view after obtaining a reference object image via a user-provided text prompt. Xie teach further comprising:
receiving, after identifying the object of interest, input requesting a modification of the target image (Xie: a user 115 provides an input image 505 depicting an object of interest as well as a transformation instruction such as rotating the object [¶0026-27, 75-76; Figs. 1 & 5]); and
performing the modification of the target image (Xie: the image processing apparatus 100 then generates a synthetic image 520 depicting the object via an image generation model 515 after applying the transformation [¶0026-27, 75-76; Figs. 1 & 5]).
Xie further explains a motivation behind the implementation of their invention, with the synthesis of new image views of an object of interest being particularly useful when generating 3D models of a particular object. Their system provides a consistent and accurate way to create these new synthetic views of an object [¶0020-22]. One of ordinary skill would recognize the benefit of being able to utilize Xie’s text-prompted image modification system with the Liu’s disclosed Matcher object segmentation process to not just segment objects in a target image, but also make synthetic image predictions of the object under different viewing conditions.
Claim(s) 10-11 & 17-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al (“Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching”, 2024, ArXiv), hereinafter referred to as “Liu”, in view of Yamamoto; Takuma (US 2021/0366145 A1), hereinafter referred to as “Yamamoto”.
Regarding claim 10, Liu teach The method of claim 9 (as described previously), however fails to recite any particular user input for indicating an object of interest.
Yamamoto, per contra, disclose an apparatus for acquiring information of an object in an image via user-defined course inputs in the form of scribbles. Yamamoto teach wherein receiving the indication of the reference object of interest comprises:
receiving the indication of the reference object via user input (Yamamoto: a user may draw a curved line (scribbles S1-S4) along an image to specify an image area for a particular object of interest [¶0032; Fig. 1]).
Yamamoto further describe that allowing a user to input a free-form curve (a scribble) to specify a target object, the designation of the target object has a relatively low degree of indefiniteness for what regions of the image would correspond to the object [¶0026-27]. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to utilize the user-defined scribble as an input for determining the reference frame mask mr so a user may be able properly specify a particular object of interest in a reference image. This provides a particular advantage when a reference image may contain one or more potential objects and allows a user to specifically designate which object to segment in an image.
As for claim 11, Liu in view of Yamamoto The method of claim 10 (as described above), wherein receiving the indication of the reference object via user input comprises:
receiving the user input as coarse user input (Yamamoto: a user may draw a curved line (scribbles S1-S4) along an image to specify an image area for a particular object of interest [¶0032; Fig. 1]).
Yamamoto further describe that allowing a user to input a free-form curve (a scribble) to specify a target object, the designation of the target object has a relatively low degree of indefiniteness for what regions of the image would correspond to the object [¶0026-27]. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to utilize the user-defined scribble as an input for determining the reference frame mask mr so a user may be able properly specify a particular object of interest in a reference image. This provides a particular advantage when a reference image may contain one or more potential objects and allows a user to specifically designate which object to segment in an image.
Considering claim 17, Liu teach The computer program product of claim 16 (as previously described), but fails to recite any particular user input for indicating an object of interest.
Yamamoto, on the other hand, teach wherein, to receive the indication of the reference object of interest, the program instructions comprise:
program instructions to receive the indication of the reference object via coarse user input (Yamamoto: a user may draw a curved line (scribbles S1-S4) along an image to specify an image area for a particular object of interest [¶0032; Fig. 1]).
Yamamoto further describe that allowing a user to input a free-form curve (a scribble) to specify a target object, the designation of the target object has a relatively low degree of indefiniteness for what regions of the image would correspond to the object [¶0026-27]. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to utilize the user-defined scribble as an input for determining the reference frame mask mr so a user may be able properly specify a particular object of interest in a reference image. This provides a particular advantage when a reference image may contain one or more potential objects and allows a user to specifically designate which object to segment in an image.
Regarding claim 18, Liu teach A system (Liu: Matcher, a computer vision framework for one shot semantic segmentation [Sec 1: Introduction - ¶03-04]) comprising:
one or more devices† (†the examiner notes that this limitation is being mapped to its associated claim interpretation under 35 U.S.C. § 112(f) – Liu: GPU cluster [Sec 5: Conclusion – Subsec: Acknowledgements] – a GPU contains one or more processors or memories for storing and implementing computer-readable instructions, a reference image (xr) is provided with a mask (mr) indicating an object to be segmented in one shot [Sec 3: Method - ¶01; Figs. 1 & 2], references images features and patches are encoded using a ViT [Sec 4.1: Experiment Settings - ¶01], patch-level features (zr, zt) of both reference (xr) and target (xt) images are extracted to compute a similarity between matching regions in the form of a correspondence matrix (S) [Sec 3.1: Correspondence Matrix Extraction; Fig. 1]– the examiner notes that that the disclosed correspondence matrix is analogous to the probabilistic feature matching as outlined in ¶0017-18 of the specification of the instant application, Matcher leverages an optimal transport approach for dense semantic feature matching when determining mask relevance based on the correspondence matrix [Sec 3.3: Controllable Masks Generation - ¶01-03 & Sec: Appendix A - ¶01-02; Fig. 2] – the examiner notes that the optimal transport approach described in Liu is analogous to the optimal transport approach employed by the instant application [¶0021-22; 43] for computationally efficient matching, Matcher then segments objects or parts of the target image using the correspondence matrix to generate a course segmentation mask over the object to be identified in the target image [Sec 3.2: Prompts Generation - ¶01; Figs. 1 & 2]) configured to:
receive user input that indicates a reference object of interest within one or more reference images (Liu: a reference image (xr) is provided with a mask (mr) indicating an object to be segmented in one shot [Sec 3: Method - ¶01; Figs. 1 & 2]);
generate a probability map associated with a target image based at least in part on matching scores that are associated with the reference object (Liu: patch-level features (zr, zt) of both reference (xr) and target (xt) images are extracted to compute a similarity between matching regions in the form of a correspondence matrix (S) [Sec 3.1: Correspondence Matrix Extraction; Fig. 1]– the examiner notes that that the disclosed correspondence matrix is analogous to the probabilistic feature matching as outlined in ¶0017-18 of the specification of the instant application); and
identify an object of interest within the target image based at least in part on the probability map (Liu: Matcher then segments objects or parts of the target image using the correspondence matrix to generate a course segmentation mask over the object to be identified in the target image [Sec 3.2: Prompts Generation - ¶01; Figs. 1 & 2]). Liu, however, fails to disclose a device† in accordance with the claim interpretation under 35 U.S.C. § 112(f) that allows a user to specify a course input to select an object of interest in an image.
Yamamoto, on the other hand, disclose an apparatus for acquiring information of an object in an image via user-defined course inputs in the form of scribbles. Yamamoto teach one or more devices† (Yamamoto: CPU 121 [¶0080; Fig. 14], a user may draw a curved line (scribbles S1-S4) along an image to specify an image area for a particular object of interest [¶0032; Fig. 1]).
Yamamoto further describe that allowing a user to input a free-form curve (a scribble) to specify a target object, the designation of the target object has a relatively low degree of indefiniteness for what regions of the image would correspond to the object [¶0026-27]. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to utilize the user-defined scribble as an input for determining the reference frame mask mr so a user may be able properly specify a particular object of interest in a reference image. This provides a particular advantage when a reference image may contain one or more potential objects and allows a user to specifically designate which object to segment in an image.
Considering claim 19, Liu in view of Yamamoto teach The system of claim 18 (described above), wherein the one or more devices are configured to:
provide, after identifying the object of interest, an indication of the object of interest that is identified within the target image (a segmentation mask is overlayed over the identified object [Sec 3.2: Prompts Generation - ¶01; Figs. 1 & 2]).
With respect to claim 20, Liu in view of Yamamoto teach The system of claim 18 (described previously), wherein, to identify the object of interest, the one or more devices are configured to:
sample candidate locations of the target image and identify the matching scores to identify the object of interest (image patches are compared between features of the reference (xr) and target (xt) images to discovery the best matching regions defined by the cosine similarity of the correspondence matrix (S) [Sec 3.1: Correspondence Matrix Extraction - ¶01; Fig. 1]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Chakraborty et al (US 11,971,955 B1) disclose an example-based image annotation system wherein a machine learning model obtains a query image of an object of interest, and the model then annotates the objects of interest in target images.
Lee et al (US 2019/0311202 A1) describe a video object segmentation method utilizing reference frames with an annotated ground truth to predict a mask in a target frame for propagating a mask across frames, with some embodiments allowing a user to designate an object using a reference scribble.
Bar et al (“Visual Prompting via Image Inpainting”, 2022, arXiv) outline a computer vision system wherein an input image can be provided to provide a visual prompt for segmenting matching objects in a query image.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael M. Sofroniou whose telephone number is (571)272-0287. The examiner can normally be reached M-F: 8:30 AM - 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, John M. Villecco can be reached at (571) 272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL M SOFRONIOU/Examiner, Art Unit 2661
/AARON W CARTER/Primary Examiner, Art Unit 2661