DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The abstract of the disclosure is objected to because it contains grammatical and idiomatic errors and therefore does not provide a sufficiently clear and concise narrative of the technical disclosure: (i) “The embodiments of the disclosure provides --” should be corrected to “The embodiments of the disclosure provide --”; (ii) “performing target segmentation on the target object at location of a viewpoint” should be clarified, for example, as “performing target segmentation on the target object at a location indicated by the current viewpoint.”; (iii) The phrase “in response to a segmentation end operation … taking the current segmentation result” lacks a finite verb and should be revised, for example, to “in response to a segmentation end operation … the method takes the current segmentation result.”; (iv) “any target may be real-time segmented” should be revised to “any target may be segmented in real time.”.
The abstract recites method steps in claim-like language (“determining...; performing...; taking...") rather than a concise narrative summary.
A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b).
The disclosure is objected to because of the following informalities:
[0002]: “and, more particular, to” should be “and, more particularly, to --”
[0003]: “a segmentation mask image mask” is redundant; “is currently urgent needed” should be “is currently urgently needed.”
[0004]: “to real-time segment any target” should be “to segment any target in real time.”
[0017]: “the target which the target user currently want to segment” should use “wants.”; “current viewpoint point” is redundant.
[0027]: “the implement method” should be revised, such as to “the implementation method.”; “omit the steps shown by performing” is grammatically unclear.
[0037]: “the sight line is viewed to the target location which wants to be segmented currently” should be rewritten to identify that the user directs the sight line toward the target to be segmented.
[0043]: “obtained by preforming pre-training” should be corrected to “obtained by performing pre-training” or “obtained by pre-training.”; “a target mask image which has a same size with the target object” should be “a target mask image having the same size as the target object.”; “mask image mask” is redundant.
[0044]: “the pre-trained obtained visual foundation model” should be revised, such as to “the pre-trained visual foundation model.”
[0050]: “is used to be a specified eye action” and “is used to be a specified gesture action” should be grammatically corrected.
[0053]: “sample mask image mask” is redundant.
[0058]: “in details” should be revised to “in detail.”
[0069]: inconsistent terminology, "the cached historical segmentation result corresponding to the target audience" should read "target object" to match usage throughout the rest of the specification; this is a clerical drafting error that should be corrected for clarity.
Appropriate correction is required.
Claim Objections
Claims 1, 10, and 19 are objected to because of the following informalities: claims 1, 10, and 19 inconsistently recite “a current viewpoint” and subsequently “a viewpoint” in the claims. The language should be corrected for grammatical and terminological consistency, for example: “at a location indicated by the current viewpoint”. Appropriate correction is required.
Claims 6, 9, 15, and 18 are objected to because of the following informalities: claims 6, 9, 15, and 18 use the plural verb “comprise” with the singular subject “performing”; “comprise” should be corrected to “comprises”. Appropriate correction is required.
Claims 8 and 17 are objected to because of the following informalities: claims 8 and 17 repeatedly recite: “less than or equals to” and “greater than or equals to”. These phrases should be corrected to: “less than or equal to” and “greater than or equal to”. Appropriate correction is required.
Claims 8 and 17 recite both “a score of a segmentation quality” and “a predetermined score of a segmentation quality”. For consistency and clarity, applicant should consider using the defined terminology “segmentation-quality score” throughout.
Claim Rejections - 35 USC § 112(b)
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 5, 8, 14, and 17 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 5 and 14 recite the limitation “before in response to the segmentation end operation triggered by the target user for a current segmentation result, in response to a re-segmentation operation…” in claims. The phrase “before in response to” does not grammatically or logically identify the required temporal relationship among the segmentation-end operation, the re-segmentation operation, re-obtaining the viewpoint-location information, and performing target re-segmentation. It is unclear whether: (i) Re-segmentation occurs before the segmentation-end operation; (ii) Re-segmentation occurs in response to the segmentation-end operation; (iii) Re-segmentation occurs in response to the separate re-segmentation operation but before the current result is accepted as final; or (iv) Both operations must be triggered before re-segmentation occurs.
The specification suggests that re-segmentation occurs when the user is dissatisfied with the current result and before the user triggers the ending operation. However, that intended sequence cannot be imported into the claims to repair the malformed language. The Examiner must explain why the language prevents a skilled artisan from determining the claim boundaries.
Claims 8 and 17 recite the limitation “a variation amount of a current scenario is less than or equals to a first predetermined variation amount” in claims. The phrase “a variation amount of a current scenario” fails to identify the reference scenario, image, video frame, time, or state against which the variation of the current scenario is determined. Consequently, it is unclear what quantity must be less than or equal to the first predetermined variation amount. The specification states that the variation may be based on the states of a historical scenario and a current scenario and that the scenario state may be represented using a hash value of an image or video. Accordingly, the specification appears to contemplate a comparison between the historical and current scenarios, but the claims recite only “a variation amount of a current scenario”.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1–20 are rejected under 35 U.S.C. § 101 because the claimed invention is directed to a judicial exception without significantly more.
This rejection has been made in accordance with the current USPTO subject matter eligibility framework, including MPEP §§ 2103–2106.07, the 2019 Revised Patent Subject Matter Eligibility Guidance, the October 2019 Patent Eligibility Guidance Update, the 2024 Guidance Update on Patent Subject Matter Eligibility, Including on Artificial Intelligence, the July 2024 AI Subject Matter Eligibility Examples, the August 4, 2025 USPTO memorandum titled "Reminders on evaluating subject matter eligibility of claims under 35 U.S.C. § 101," and the USPTO's guidance concerning Ex parte Desjardins, Appeal No. 2024-000567. The claims have been evaluated under the broadest reasonable interpretation, and the claims have been considered as a whole.
Step 1: Statutory category
Independent claim 1 is directed to a method of target segmentation and therefore falls within the statutory category of a process. Independent claims 10 and 19 are directed to an electronic device and a storage medium, respectively, and therefore fall within the statutory categories of a machine and an article of manufacture. Accordingly, the analysis proceeds to Step 2A.
Step 2A, Prong One (Judicial exception)
Independent claim 1 recites: determining location information of a current viewpoint when a target user watches a target object; performing target segmentation on the target object at location of a viewpoint based on location information of a current viewpoint and a visual foundation model, to determine and present a current segmentation result; and in response to a segmentation end operation triggered by the target user for the current segmentation result, taking the current segmentation result as a target segmentation result corresponding to the target object.
These limitations recite an abstract idea, namely a mental process and mathematical concepts. Specifically, identifying an object a person is looking at (watching a target object to determine a viewpoint), outlining or segmenting that object conceptually or mathematically, and deciding to accept that segmentation based on human observation are mental processes historically performed by humans. Furthermore, using a "visual foundation model" to perform "target segmentation" based on input coordinates ("location information") relies on mathematical neural networks and algorithms to group and classify data points.
The claim recites information collection (determining viewpoint location), mathematical data analysis (processing the location via a visual foundation model), and outputting a result (taking the current segmentation result). The claim does NOT recite a specific physical improvement to the way images are captured, digitized, or displayed, nor does it improve the internal physical operation of the computer or eye-tracking hardware itself. Rather, the claim uses generic sensors to obtain coordinate data and mathematically processes that data through an AI model.
The claim is similar in character to claims that courts have found abstract where the focus is collecting information, analyzing the information mathematically, and presenting or acting on the results of the analysis (e.g., Electric Power Group, LLC v. Alstom S.A.). The present claims similarly collect gaze location data, mathematically analyze the data to segment an image using an AI model, and output a classification mask.
Dependent claims 2-9 and 11-18 further limit the method and apparatus to generic data gathering (e.g., "obtaining current eye movement information"), generic display techniques ("labeling the current segmentation result"), basic UI triggers ("predetermined eye action or a predetermined gesture action"), caching data ("obtaining a cached historical segmentation result"), applying conditional logic ("detecting that a condition of continued segmentation is currently satisfied"), comparing data to thresholds (e.g., checking if variations in scenarios, eye movements, or quality scores are less than/greater than predefined thresholds), and generic mathematical alignment ("time alignment processing"). These limitations recite pure mathematical concepts, logic rules, and data manipulation steps.
Accordingly, claims 1–20 recite an abstract idea under Step 2A, Prong One.
Step 2A, Prong Two (Practical Application)
The additional elements, considered individually and in combination, do not integrate the abstract idea into a practical application.
The recited "visual foundation model", "location information", "segmentation result", and "variation amount" amount to data objects, mathematical models, and generic computer implementations of the abstract image segmentation concept.
The claims do not recite a particular improvement to computer or imaging technology. They do not improve how a digital image is physically captured, sensed, or stored. The claims do not recite a particular improvement to graphical processing hardware or eye-tracking hardware. Rather, the claims use standard location coordinates as input data for mathematical algorithms.
The claims further do not recite a particular improvement to artificial intelligence or machine learning technology itself. The claims do not train a model in a novel way that improves the computer's operation, update model parameters to reduce processing load, or modify model architecture to save hardware resources. The "visual foundation model" is recited purely functionally as a black-box mathematical tool for segmenting images based on coordinates.
This analysis is consistent with the USPTO's 2024 AI subject matter eligibility guidance and AI examples, which emphasize that AI-related claims may be eligible when they recite a specific technological improvement or otherwise integrate a judicial exception into a practical application. The present claims do NOT recite such a specific technological improvement. Instead, the claims use generic computer operations to collect location information, mathematically segment an image, and accept the result. Applying AI segmentation to a new field of use (e.g., gaze-based UI) is a field-of-use limitation, not an integration of the abstract idea into a practical application.
Accordingly, the claims do not integrate the judicial exception into a practical application under Step 2A, Prong Two.
Step 2B: (Inventive Concept)
The additional elements, considered both individually and as an ordered combination, do not amount to significantly more than the abstract idea.
The claims use generic computer components to perform ordinary computer functions, including obtaining eye tracking data, running mathematical models, comparing values to thresholds, and caching historical data. These are conventional data-processing and mathematical operations performed using generic computer technology.
The ordered combination also does not provide an inventive concept. The ordered combination follows the abstract idea itself: determine where a user is looking, mathematically segment the object at that location, and save the result if the user triggers an end operation. This is no more than the abstract idea implemented on generic computer components and displays.
The dependent claims recite additional steps of obtaining eye movement via a wearable device, matching variations against thresholds, detecting specific gestures, and caching results. These limitations merely specify the generic hardware tools, known logic formulas, and standard UI gathering mechanisms used in the abstract evaluation and do not add significantly more.
The independent electronic device and storage medium claims recite generic counterparts using basic computing logic to perform substantially the same operations. The recitation of generic processors and memory does not transform the abstract idea into patent-eligible subject matter.
Accordingly, claims 1–20 are directed to a judicial exception without significantly more and are therefore rejected under 35 U.S.C. § 101.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 10 and 19 are rejected under 35 U.S.C. §102(a)(1) as being anticipated by Wang (Wang et al, GazeSAM: What You See is What You Segment. arXiv.Org, April 26th 2023).
Regarding claim 1, Wang teaches a method of target segmentation, comprising:
determining location information of a current viewpoint when a target user watches a target object;
( Abstract; §§2.1–2.2; Figs. 1–2: Wang discloses tracking a user’s eye movement while the user looks at a region of interest, collecting eye-gaze data as location coordinates on the screen after calibration, and transforming the screen-coordinate gaze points into image-coordinate points for use in segmentation, thereby determining location information corresponding to the user’s current viewpoint while the user watches the target object. )
performing target segmentation on the target object at location of a viewpoint based on location information of a current viewpoint and a visual foundation model, to determine and present a current segmentation result; and
( Abstract; §§2.2; §3; Figs. 1–3: Wang discloses using the eye-gaze data (point or sequence of eye-gaze points, transformed into the image-coordinate space) as the input prompt for the Segment Anything Model (SAM) [visual foundation model], where SAM’s prompt encoder accepts point prompts, and, given the prompt transformed from eye-gaze data, SAM generates a segmentation mask in near real-time, with the resulting mask shown on the interface in real time. )
in response to a segmentation end operation triggered by the target user for the current segmentation result, taking the current segmentation result as a target segmentation result corresponding to the target object.
( §3; Fig. 3(a): Wang discloses that once segmentation is generated and displayed, users can save the eye-captured segmentation mask displayed on the interface at any point during the process using the “Save Mask” function, which teaches taking the then-current segmentation result as the segmentation result for the target object in response to user action. )
Regarding claims 10 and 19, the rationale provided in the rejection of claim 1 is incorporated herein. In addition, Wang teaches a human-computer interaction system (GazeSAM) including a computing device running the SAM model having a “user interface that incorporates multiple functions”, computer components required to encode prompts and generate masks in real time [§§2.1-2.2; §3; Fig. 3]. Accordingly, the method of claim 1 corresponds to the electronic device of claim 10, as well as the non-transitory storage medium of claim 19 and performs the steps disclosed herein. Therefore, the claims are all rejected.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 2–4, 11–13 and 20 are rejected under 35 U.S.C. §103 as being unpatentable over Wang in view of Spencer (Spencer et al, US 2023/0110964 A1, April 13th 2023).
Regarding claim 2, Wang teaches the method of target segmentation according to claim 1, wherein determining location information of the current viewpoint when the target user watches the target object comprises:
obtaining current eye movement information or current head movement information when the target user watches the target object via a screen-based eye tracker;
( §2.1; §3; Figs. 1 & 3(b)]: Wang discloses that GazeSAM utilizes a screen-based eye-tracker, specifically a Tobii Pro Nano, for eye-gaze data collection, and that, after calibration, eye-gaze data can be collected in the form of location coordinates on the screen while the user looks at the target region. )
determining, based on the current eye movement information or the current head movement information, location information of a current viewpoint of the target user.
( §2.1; §2.2; Eq. (1); Figs. 1–2: Wang discloses that the collected eye-gaze point coordinates are transformed from screen coordinate space into image coordinate space for use as point prompts for SAM, thereby determining, based on the current eye movement information, location information corresponding to the user’s current viewpoint. )
Wang, however, fails to disclose obtaining the eye-movement or head-movement information through a wearable device, where Spencer teaches:
obtaining current eye movement information or current head movement information when the target user watches the target object via a wearable device;
( [0010], [0025–0027], [Figs. 1A–1C]: Spencer teaches a wearable computing device with a gaze tracking device that tracks eye gaze direction and movement while the user views displayed image content, and further teaches position/ orientation sensors for the wearable device, thereby teaching obtaining current eye movement information or current head movement information via a wearable device. )
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Wang’s screen-based gaze-tracking arrangement by incorporating Spencer’s known wearable gaze-tracking device. Wang and Spencer both obtain gaze information for identifying the location at which a user is looking. Spencer’s wearable implementation would predictably permit Wang’s gaze-directed segmentation system to obtain the same viewpoint information in a portable, head-mounted, and hands-free manner, while preserving Wang’s established use of the gaze location as a segmentation prompt.
Regarding claim 3, Wang [as modified by Spencer] teaches the method of target segmentation according to claim 2, wherein presenting the current segmentation result comprises:
labeling the current segmentation result in the target object presented by the wearable device.
( Wang teaches presenting the selected image and the corresponding current segmentation mask through a user interface, wherein the generated segmentation mask is shown in green and automatically updated based on the location of the user’s eye gaze, thereby visually labeling the segmented target object with the current segmentation result [§3; Figs. 3(a) & 4]. Spencer teaches a wearable computing device having a display device configured to present image content viewed by the user [0025–0027] & [Figs. 1A–1C]. )
Regarding claim 4, Wang teaches the method of target segmentation according to claim 1, wherein the segmentation end operation is triggered by performing a conventional UI input by the target user.
( §3; Fig. 3(a): Wang teaches a graphical user interface having a function panel in which each function may be activated by a corresponding keyboard shortcut or by clicking a displayed button. Wang further teaches that the target user may activate the “Save Mask” function at any point during the segmentation process to save the segmentation mask currently displayed on the interface. )
Wang, however, fails to disclose triggering the segmentation-end operation through a predetermined eye action or predetermined gesture action, where Spencer teaches:
the segmentation end operation is triggered by performing a predetermined eye action or a predetermined gesture action by the target user.
( [0053–0054], [Figs. 2E–2F]: Spencer teaches verifying a detected object selection through a predetermined eye action, including maintaining a fixation gaze for longer than a previously set time threshold or performing an eye gesture such as a blink. Spencer further teaches that a user selection may be triggered through a gaze input or a gesture input. Accordingly, Spencer teaches triggering confirmation of a selected result through a predetermined eye action or predetermined gesture action performed by the user. )
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Wang’s keyboard- or button-triggered “Save Mask” operation to use Spencer’s known eye-action or gesture-based confirmation input. Wang and Spencer both concern user confirmation of a visually selected result, and Spencer’s fixation, blink, or gesture input would have predictably provided a hands-free and more intuitive mechanism for accepting Wang’s currently displayed segmentation mask, particularly in a gaze-controlled segmentation system.
Regarding claims 11–13 and 20, the rationale provided in the rejection of claim 2–4 is incorporated herein. In addition, Wang [as modified by Spencer] teaches a computer-program product tangibly embodied in a computer system, including instructions configured to cause one or more data processors to perform the disclosed method (Wang: [§§2.1-2.2; §3; Fig. 3]; Spencer: [0074]). Accordingly, the method of claim 2 corresponds to the electronic device of claim 11, as well as the non-transitory storage medium of claim 20; the method of claims 3–4 corresponds to the electronic device of claim 12–13; and performs the steps disclosed herein. Therefore, the claims are all rejected.
Claims 5 and 14 are rejected under 35 U.S.C. §103 as being unpatentable over Wang.
Regarding claim 5, Wang teaches the method of target segmentation according to claim 1, further comprising: before in response to the segmentation end operation triggered by the target user for a current segmentation result,
in response to a re-segmentation operation triggered by the target user for the current segmentation result, re-obtaining location information of the current viewpoint of the target user; and performing target re-segmentation based on the re-obtained location information of the current viewpoint.
( §2.1; §2.2; §3; Eq. (1); Figs. 1–3: Wang teaches that SAM may initially generate an imperfect segmentation mask, particularly at boundary regions, and that the user may refine the current mask by continuing to look at additional desired regions. In response to the user’s continued gaze interaction, the eye tracker collects updated gaze-point coordinates, which are transformed from the screen-coordinate space into corresponding locations in the image-coordinate space and supplied as additional point prompts to SAM. Wang further teaches that the segmentation mask is automatically updated based on the newly obtained gaze location, thereby dynamically adjusting the segmented target or iteratively refining the current segmentation result before the user activates the “Save Mask” function. )
Although different embodiments of Wang have been referred to (Wang does not expressly label the user’s continued gaze-based refinement as a separate re-segmentation operation/ module), it would have been exceedingly obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement Wang’s disclosed continued gaze-based mask refinement as a user-triggered re-segmentation operation. Wang already teaches collecting updated gaze locations and automatically refining the current mask based on those locations; expressly initiating that same refinement process in response to the user’s continued gaze interaction would merely formalize Wang’s existing iterative workflow and predictably improve user control over correction of an unsatisfactory segmentation result.
Regarding claim 14, the rationale provided in the rejection of claim 5 is incorporated herein. In addition, Wang teaches a human-computer interaction system (GazeSAM) including a computing device running the SAM model having a “user interface that incorporates multiple functions”, computer components required to encode prompts and generate masks in real time [§§2.1-2.2; §3; Fig. 3]. Accordingly, the method of claim 5 corresponds to the electronic device of claim 14, and performs the steps disclosed herein. Therefore, the claims are all rejected.
Claims 6–7, 9, 15–16, and 18 are rejected under 35 U.S.C. §103 as being unpatentable over Wang in view of Yang (Yang et al, Track Anything: Segment Anything Meets Videos. arXiv.Org., April 28th 2023).
Regarding claim 6, Wang teaches the method of target segmentation according to claim 1, wherein performing target segmentation on the target object at location of a viewpoint based on location information of the current viewpoint and the visual foundation model, to determine the current segmentation result comprise:
Wang teaches gaze-location prompting of SAM and generation of a current segmentation result and an iterative process where a user can add more gaze points to fix a bad segmentation mask (the "AllPoints" mode), but fails to disclose where Yang teaches:
obtaining a cached historical segmentation result corresponding to the target object; and
( §3.1; §3.2, Step 2; Fig. 1: Yang teaches that, given an initialized mask describing the target object, XMem tracks the target object and generates corresponding predicted masks in subsequent frames using temporal and spatial correspondence. When the quality of a predicted mask is unsatisfactory, Yang saves the XMem prediction and corresponding intermediate parameters, segmentation masks are then propagated and maintained over time for the same object, thereby obtaining and caching a historical segmentation result corresponding to the target object. )
performing, based on location information of the current viewpoint, the historical segmentation result, and the visual foundation model, target segmentation on the target object at location of the viewpoint and segmentation result superposition processing, to determine a current segmentation result after superposition.
( §3.2, Steps 2–3; Fig. 1: Yang teaches using the XMem-predicted mask as a mask prompt for SAM and projecting the corresponding probes and affinities as point prompts for SAM. SAM jointly processes the image, point prompts, and historical predicted-mask prompt to generate a refined current segmentation mask. In Wang’s system as modified by Yang, Wang’s current gaze location provides the current viewpoint point prompt, while Yang’s cached XMem prediction provides the historical segmentation-result mask prompt; SAM processes the current viewpoint prompt and historical mask prompt together, corresponding to the recited segmentation-result superposition processing, to produce the refined current segmentation result after superposition. Yang further adds the refined mask to XMem’s temporal correspondence for use in subsequent segmentation. )
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Wang’s gaze-prompted SAM segmentation system by incorporating Yang’s known XMem-based historical-mask prompting and refinement technique. Wang and Yang both use interactive prompts with SAM to generate and refine target-object segmentation masks, and Yang expressly identifies SAM’s lack of temporal correspondence as a deficiency and addresses it by supplying an XMem-predicted historical mask to SAM as a mask prompt. The modification would predictably preserve prior segmentation information, improve temporal consistency, reduce repeated user input, and enable Wang’s current gaze prompt and Yang’s historical mask prompt to be jointly processed by SAM to produce a refined current segmentation result.
Regarding claim 7, Wang [as modified by Yang] teaches the method of target segmentation according to claim 6, wherein obtaining the cached historical segmentation result corresponding to the target object comprises:
in response to a continued segmentation operation triggered by the target user for the historical segmentation result, obtaining a cached historical segmentation result corresponding to the target object; or
( §3.1; §3.2, Steps 2 and 4; Fig. 1: Yang further teaches that, during continued segmentation, the user may compulsively pause the TAM process and correct the segmentation mask of the current frame using positive and negative clicks. Therefore, in response to the user-triggered continued-segmentation or correction operation directed to the previously generated segmentation result, Yang obtains and uses the available XMem-predicted segmentation mask corresponding to the target object for continued segmentation and correction. )
in accordance with detecting that a condition of continued segmentation is currently satisfied, obtaining the cached historical segmentation result corresponding to the target object.
( §3.2, Steps 2–3; Fig. 1: Yang teaches that XMem performs video-object segmentation on subsequent frames based on an initialized target-object mask and outputs predicted masks during continued tracking. Yang detects whether the quality of an XMem-predicted mask is satisfactory. When the mask quality is not satisfactory, corresponding to detecting that a condition for continued segmentation and refinement is satisfied, Yang saves the XMem prediction and its corresponding intermediate parameters, including probes and affinities, and proceeds to SAM refinement. Accordingly, Yang obtains and caches the historical segmentation result corresponding to the target object upon detecting satisfaction of the continued-segmentation condition. )
Regarding claim 9, Wang [as modified by Yang] teaches the method of target segmentation according to claim 6, wherein performing, based on location information of the current viewpoint, the historical segmentation result, and the visual foundation model, target segmentation on the target object at location of the viewpoint and segmentation result superposition processing, to determine the current segmentation result after superposition comprise:
performing time alignment processing on the historical segmentation result, to obtain an aligned historical segmentation result at the current moment;
( §3.1; §3.2, Step 2; Fig. 1: Yang teaches that, given an initialized mask describing the target object in an earlier frame, XMem tracks the target object and generates a corresponding predicted mask in a subsequent or current frame based on temporal and spatial correspondence. Accordingly, XMem propagates the historical target-object segmentation information into the current frame, thereby performing time-alignment processing on the historical segmentation result to obtain a predicted historical-mask result aligned with the current frame or moment. )
performing, by inputting the target object, location information of the current viewpoint, and an aligned historical segmentation result into the visual foundation model, target segmentation at location of the viewpoint and segmentation result superposition processing; and
( Wang, [§§2.1–2.2; Eq. (1); Figs. 1–2]; Yang, [§3.2, Steps 2–3; Fig. 1]: Wang teaches transforming the user’s current eye-gaze location into image-coordinate information and supplying that current viewpoint location as a point prompt to SAM. Yang teaches supplying the current-frame image to SAM, projecting XMem’s probes and affinities as point prompts, and supplying the XMem-predicted mask aligned with the current frame as a mask prompt to SAM. In Wang [as modified by Yang]’s system, SAM therefore receives the target-object image, Wang’s current viewpoint-location prompt, and Yang’s aligned historical segmentation-mask prompt. SAM jointly processes the current viewpoint prompt and historical-mask prompt, corresponding to the recited segmentation-result superposition processing, to perform target segmentation at the current viewpoint location. )
obtaining, based on output of the visual foundation model, the current segmentation result after superposition.
( §3.2, Step 3; Fig. 1: Yang teaches that, using the projected point prompts and the XMem-predicted mask prompt, SAM produces a refined segmentation mask for the current frame. The refined mask is the current segmentation result obtained from the visual foundation model after the combined processing of the current prompt information and aligned historical segmentation-mask information. Yang further adds the refined mask to XMem’s temporal correspondence for use in refining subsequent target-object segmentation. )
Regarding claims 15–16 and 18, the rationale provided in the rejection of claims 6–7 and 9 is incorporated herein. In addition, Wang [as modified by Yang] teaches a computer-program product tangibly embodied in a computer system, including instructions configured to cause one or more data processors to perform the disclosed method (Wang: [§§2.1-2.2; §3; Fig. 3]). Accordingly, the method of claims 6–7 and 9 corresponds to the electronic device of claims 15–16 and 18, and performs the steps disclosed herein. Therefore, the claims are all rejected.
Claim 8, and 17 are rejected under 35 U.S.C. §103 as being unpatentable over Wang [as modified by Yang] in view of Kirillov (Kirillov et al, Segment Anything. arXiv:2304.02643 [Cs], April 5th 2023).
Regarding claim 8, Wang [as modified by Yang] teaches the method of target segmentation according to claim 7, wherein the condition of continued segmentation comprises at least one of:
Wang [as modified by Yang] teaches obtaining and using historical segmentation results during continued target segmentation, but fails to expressly disclose where Kirillov teaches:
a variation amount of a current scenario is less than or equals to a first predetermined variation amount;
a variation amount of a current eye movement is less than or equals to a second predetermined variation amount;
a variation amount of a current head movement is less than or equals to a third predetermined variation amount;
a score of a segmentation quality corresponding to a historical segmentation result is greater than or equals to a predetermined score of a segmentation quality.
( §3, “Segment Anything Model,” Fig. 3; Supplementary Material, §C, “Automatic Mask Generation Details,” “Filtering”: Kirillov teaches that SAM predicts a confidence score, namely an estimated intersection-over-union (IoU) score, for each predicted segmentation mask to evaluate and rank mask quality. Kirillov further teaches filtering the predicted masks using a predetermined predicted-IoU threshold of 88.0, such that only masks having a predicted IoU score greater than or equal to the predetermined threshold are retained as confident masks. In Wang [as modified by Yang]’s system, applying Kirillov’s predicted-IoU quality score and threshold to Yang’s cached historical segmentation mask teaches determining that the segmentation-quality score corresponding to the historical segmentation result is greater than or equal to a predetermined segmentation-quality score. )
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to further modify Wang [as modified by Yang] by applying Kirillov’s known predicted-IoU quality score and predetermined threshold to Yang’s cached historical segmentation results. Yang already evaluates whether a predicted historical mask has satisfactory quality before continued segmentation or refinement, while Kirillov provides an express, known technique for quantifying mask quality and retaining masks that satisfy a predetermined confidence threshold. Applying Kirillov’s threshold-based quality assessment to Yang’s historical masks would have been a routine use of a known segmentation-quality technique to predictably prevent propagation of unreliable masks and improve the accuracy and temporal consistency of continued target segmentation.
Regarding claim 17, the rationale provided in the rejection of claim 8 is incorporated herein. In addition, Wang [as modified by Yang and Kirillov] teaches a computer-program product tangibly embodied in a computer system, including instructions configured to cause one or more data processors to perform the disclosed method (Wang: [§§2.1-2.2; §3; Fig. 3]; Kirillov: [§3; Appendix, Training recipe]). Accordingly, the method of claim 8 corresponds to the electronic device of claim 17, and performs the steps disclosed herein. Therefore, the claims are all rejected.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEN KUDO whose telephone number is (571)272-4498. The examiner can normally be reached M-F 8am - 5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent Rudolph can be reached at 571-272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
KEN KUDO
Examiner
Art Unit 2671
/KEN KUDO/Examiner, Art Unit 2671
/VINCENT RUDOLPH/Supervisory Patent Examiner, Art Unit 2671