DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Notice to Applicant
This office action is in response to application filed on 01/18/2025.
Limitations appearing inside of {} are intended to indicate the limitations not taught by said prior art(s)/combinations.
Claims 1-15 are pending in the application.
Information Disclosure Statement
Information Disclosure Statement(s) (IDS) filed on 01/18/2025 has been considered.
Claim Objections
Claim 5 is objected to because of the following informalities:
Claim 5 recites “an user input”. Consider revising to “a user input”.
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 5-7 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being incomplete for omitting essential structural cooperative relationships of elements, such omission amounting to a gap between the necessary structural connections. See MPEP § 2172.01. The omitted structural cooperative relationships are:
Regarding claim 5, is “the user input regarding a target object” related to “the at least one object to be recognized”? There does not appear to be a connection between these two terms in the claim language. The specification does not appear to refer to user input regarding “a target object”. The portion of the specification that appears to be the closest to this limitation is ¶[0082]-[0083]. For the purpose of examination, the limitation will be interpreted accord to the noted specification portion.
Regarding Claim 6, is “obtain the mask” of claim 6 referring to the limitation of claim 1? Claim 1 obtains “a mask”, and then Claim 6 obtains “a plurality of masks” and “the mask based on…the plurality of masks”. As a result, the two “the mask” limitations claimed in 1 and 6 appear unrelated. Perhaps the antecedent basis “the mask” of claim 6 would become clear if claim 6 recited “wherein the obtaining of the mask comprises” followed by “obtain a plurality of masks” and “obtain the mask”. For the purpose of examination, the limitations of claim 6 are interpreted as further limiting “the mask” of claim 1.
Regarding Claim 7, is “obtain the second video” different than limitation of claim 1? Claim 1 obtains “a second video”, and then Claim 7 obtains “the second video using the plurality of second image frames”. As a result, the two “second video” limitations of claims 1 and 7 appear unrelated. Perhaps the antecedent basis “the second video” of claim 7would become clear if claim 7 recited “wherein the obtaining of the second video comprises” followed by “obtain a plurality of second image frames” and “obtain the second video”. For the purpose of examination, the limitations of claim 6 are interpreted as further limiting “the second video” of claim 1.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 11 and 15 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by “Shin” (SHIN YUN JAE et al., KR 20220169312 A), as cited in the IDS (01/18/2025).
Regarding claim 11, Shin teaches a method of operating an electronic device, the method comprising:
obtaining a first video including a plurality of first image frames from a user device connected to the electronic device (a video taken by a user on a smartphone or the like; Shin, ¶[0020]; In step S1100, the image processing device obtains a first image frame and at least one second image frame; Shin, ¶[0068]);
obtaining a first point set and a first bounding box set for at least one object included in the first video using an object recognition model trained to generate points and bounding boxes for objects in a video (the boundary feature points of the deleted area in the first image frame and the searched insertion in the second image frame. Calculating value coordinate values, calculating and storing median values of bounding boxes of objects existing in the second image frame; Shin, ¶[0073]) using the object recognition model (using a learning-based algorithm. Detect all existing objects; Shin, ¶[0068]);
identifying the at least one object based on the first point set and the first bounding box set the tracking of the selected deletion request object may include calculating a median coordinate value of a bounding box of the deletion request object in the first image frame including the deletion request object, the second image frame calculating and storing median values of bounding boxes of objects existing in , comparing the calculated median coordinate values with median coordinate values of the deletion request object, and an object having a minimum distance difference as a result of the comparison Tracking to the deletion request object; Shin ¶[0009];
obtaining a mask for segmenting the identified at least one object in the first video (object masking step (S100); Shin, ¶[0060]);
obtaining a second video by using the mask to remove regions other than the at least one object in the first video (searches for a background area to be inserted into the area where the deletion request object is deleted in units of frames; Shin, ¶[0041]; third image frame (i.e., interpreted as the second frame of the instant application) is created by inserting; Shin, [0069]); and
transmitting the second video to the user device such that the user device outputs the second video(The author accesses the object service authoring device locally or through a network to edit the moving picture; Shin, ¶[0035]) such that the user device outputs the second video (and the output unit 1240 encodes and outputs the adjusted image; Shin, ¶[0077]).
Claim 15 is similarly analyzed as analogous claim 11.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claims 1-3,6, 9-10, and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Shin in view of “Chen”, (Chen et al., US 20160379371 A1).
1. Shin teaches an electronic device comprising:
a communication device (The author accesses the object service authoring device locally or through a network to edit the moving picture; Shin, ¶[0035]);
a storage device configured to store an object recognition model trained to generate points and bounding boxes for at least one object included in a video (implemented as a program and stored in various non-transitory computer readable media; Shin, ¶[0095]); and
at least one processor (image processing device; Shin, ¶[0041]),
wherein the at least one processor is configured to:
obtain, via the communication device, a first video including a plurality of first image frames from a user device connected to the electronic device (a video taken by a user on a smartphone or the like; Shin, ¶[0020]; In step S1100, the image processing device obtains a first image frame and at least one second image frame; Shin, ¶[0068]);
obtain a point set and a bounding box for at least one object included in the first video (the boundary feature points of the deleted area in the first image frame and the searched insertion in the second image frame. Calculating value coordinate values, calculating and storing median values of bounding boxes of objects existing in the second image frame; Shin, ¶[0073]) using the object recognition model (using a learning-based algorithm. Detect all existing objects; Shin, ¶[0068]);
identify {a contour of the} at least one object based on the point set and the bounding box set (the tracking of the selected deletion request object may include calculating a median coordinate value of a bounding box of the deletion request object in the first image frame including the deletion request object, the second image frame calculating and storing median values of bounding boxes of objects existing in , comparing the calculated median coordinate values with median coordinate values of the deletion request object, and an object having a minimum distance difference as a result of the comparison Tracking to the deletion request object; Shin, ¶[0009]);
obtain a mask for segmenting the at least one object based on the contour of the at least one object identified from the first video (object masking step (S100); Shin, ¶[0060])
obtain a second video by using the mask to remove regions {other than the at least one object} in the first video (searches for a background area to be inserted into the area where the deletion request object is deleted in units of frames; Shin, ¶[0041]; third image frame (i.e., interpreted as the second frame of the instant application) is created by inserting; Shin, [0069]); and
transmit the second video to the user device via the communication device (The author accesses the object service authoring device locally or through a network to edit the moving picture; Shin, ¶[0035]) such that the user device outputs the second video (and the output unit 1240 encodes and outputs the adjusted image; Shin, ¶[0077]).
Shi does not explicitly disclose identify a contour of the at least one object based on the point set and the bounding box set.
However, Chen, a similar field of endeavor of extracting objects from video, in addition to teaching some limitations as taught by Shin, also teaches missing limitations, as follows:
obtain a point set (select a local maxima of the heat map as the additional foreground seed point; Chen, ¶[0049]) and a bounding box set (obtain a series of candidate object bounding boxes; Chen, ¶[0049]) for at least one object included in the first video (detect the t-the frame of the input video; Chen, ¶[0049]) using the object recognition model (bounding box detector; Chen, ¶[0049]);
identify a contour of the at least one object based on the point set and the bounding box set (the seed point indicates an optimal pixel region to be detected by the object contour detector; Chen, ¶[0048]);
obtain a mask for segmenting the at least one object based on the contour of the at least one object identified from the first video (object segmentation in videos tagged with semantic labels (i.e., segmentation mask) … detect each frame of the input video with an object contour detector based on constrained parametric min-cuts, to obtain a candidate object contour set for each frame of the input video, thus roughly determining a location of the object of a given semantic category, and improving accuracy of the subsequent video object segmentation; Chen, ¶[0050]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include segmentation based on contours as taught by Chen to the invention of Shin. The motivation to do so would be to improve the accuracy of segmentation and overcome wrong results occurring from fuzzy sample classification.
2. The combination of Chen and Shin teaches the electronic device of claim 1.
Shin further teaches, wherein the at least one processor is configured to
generate image frame groups by grouping the plurality of first image frames included in the first video into the image frame groups based on a predetermined reference unit and obtain the point set and the bounding box set by tracking the at least one object in each of the image frame groups (tracks and deletes the deletion request object in all frames, and searches for a background area to be inserted into the area where the deletion request object is deleted in units of frames; Shin, ¶[0045]; Units of frames are interpreted as “image frame groups” of the instant application)
using the object recognition model (deletion request object tracking unit; Shin, ¶[0015]).
3. The combination of Chen and Shin teaches the electronic device of claim 2.
Shin further teaches wherein the at least one processor is configured to
recognize scene change in the first video using the object recognition model and
generate the image frame groups by grouping the plurality of first image frames based on the scene change. (In the video … the background and foreground move over time, …, by simultaneously tracking the foreground and the background in an image frame obtained through a camera, a background area corresponding to an object to be erased can be found even when the camera is moving. Shin, ¶[0035]; searches for a background area to be inserted into the area where the deletion request object is deleted in units of frames. and search for the optimal background area; Shin, ¶[0045]).
6. The combination of Chen and Shin teaches the electronic device of claim 1.
wherein the at least one processor is configured to:
obtain a plurality of masks for segmenting the at least one object (all objects 105a and 105b included in each image frame are detected by using a Mask R convolutional neural network (CNN) method among deep learning networks, and the coordinates of the detected objects are stored; Shin, ¶[0041]; tracks and deletes the deletion request object in all frames, and searches for a background area to be inserted into the area where the deletion request object is deleted in units of frames; Shin, ¶[0045]) {based on the contour of the at least one object} ; and
{obtain the mask based on the reliability of the plurality of masks}.
Chan further teaches obtain the mask based on the reliability of the plurality of masks (optimizing each sequence containing the object in sequence with graph cut algorithm in combination with the shape probability distribution (i.e., reliability) of the object, to obtain an optimal segmenting sequence corresponding to the object in the input video; Chen, ¶[0115]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include mask reliability as taught by Chen to the invention of Shin. The motivation to do so would be to generate a more accurate and robust estimated contour.
9. The combination of Chen and Shin teaches the electronic device of claim 1.
Shin further teaches wherein the at least one processor is configured to:
obtain an user input for mask generation from the user device (generates a first image frame 905a when using Mask R CNN in the object masking step (S100); Shin, ¶[0060]; in step S1110, the image processing device stores the coordinate information of the detected objects, and when a user selects a deletion request object; Shin, ¶[0068]); and
obtain the mask based on the user input for mask generation {and the contour} of the at least one object (generates a first image frame 905a when using Mask R CNN in the object masking step (S100); Shin, ¶[0060]).
Chen further teaches obtain the mask based on the user input for mask generation and the contour of the at least one object (the method for object segmentation in videos tagged with semantic labels …, the Step 102, …, building a joint assignment model containing the candidate object bounding box set and the candidate object contour set, to solve an initial segmenting sequence corresponding to the object in the input video; Chen, ¶[0051]).
The motivation to combine the teachings of Shin and Chen are provided with claim 1.
10. The combination of Chen and Shin teaches the electronic device of claim 9.
Shin further teaches wherein the user input for mask generation includes at least one of text input regarding the at least one object among the plurality of objects in the first video, point selection for the at least one object, bounding box selection, or masking region selection. (Reference number 110 in FIG. 2 shows that the deletion request object 110a is selected by the user.; Shin, ¶[0041]. See Fig 2 shown below:
PNG
media_image1.png
594
411
media_image1.png
Greyscale
).
Claim 12 is similarly analyzed as analogous claim 2.
Claims 4-5, and 13 are rejected under pre-AIA 35 U.S.C. 103(a) as being unpatentable over Shin in view of Chen and further in view of Holdsworth” (Holdsworth et al., KR 20220049389 A), as cited in the IDS (01/18/2025).
4. The combination of Chen and Shin teaches the electronic device of claim 1.
Shin further teaches wherein the object recognition model is configured to
{generate the point set by extracting skeleton data of the at least one object included in each of the plurality of first image frames, and}
generate the bounding box set for the at least one object included in each of the plurality of first image frames (Calculating value coordinate values, calculating and storing median values of bounding boxes of objects existing in the second image frame; Shin, ¶[0073]).
The combination does not explicitly disclose generate the point set by extracting skeleton data of the at least one object included in each of the plurality of first image frames.
However, Holdsworth, a similar field of endeavor of removing unnecessary dynamic objects by utilizing the patio-temporal characteristics of video images captured in a dynamic environment, teaches generate the point set by extracting skeleton data of the at least one object included in each of the plurality of first image frames (receives skeleton information and object tracking through joint and scale analysis for a specific object identified in the frame each time a frame constituting an image is received, …, based on the skeleton information of the object, the predicted movement position and the scale change of the object expected for the key point; Holdsworth, ¶[0107]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include generating a point set from skeleton data as taught by Holdsworth to the combined invention of Shin and Chen. The motivation to do so would be to detect the same object by comparing it with the prediction information for each object generated for the change of the joint point and scale of the object at the position predicted in the next frame.
5. The combination of Chen, Shin, and Holdsworth teaches the electronic device of claim 4.
Shin further teaches wherein the at least one processor is configured to: obtain an user input regarding a target object through the user device; and determine the at least one object to be recognized by the object recognition model based on the user input (An object detection unit that detects all objects existing in the first and second image frames using an algorithm and stores coordinate information of the detected objects, and an object to be deleted that sets an object requested by the user as a deletion request object; Shin, ¶[0014]).
Claim 13 is similarly analyzed as analogous claim 4.
Claim 7 is rejected under pre-AIA 35 U.S.C. 103(a) as being unpatentable over Shin in view of Chen and further in view of “Liu” (Liu; Zhidong et al., US 20240031517 A1).
7. The combination of Shin and Chen teaches the electronic device of claim 1.
Shin further teaches wherein the at least one processor is configured to:
obtain a plurality of second image frames by using the mask to remove regions other than the at least one object (tracks and deletes the deletion request object in all frames, and searches for a background area to be inserted into the area where the deletion request object is deleted in units of frames; Shin, ¶[0045]); and
{obtain the second video using the plurality of second image frames and audio information of the first video}.
The combination does not explicitly disclose obtain the second video using the plurality of second image frames and audio information of the first video.
However, Liu, a similar field of endeavor of video segmentation, teaches obtain the second video using the plurality of second image frames and audio information of the first video (Audio segment extractor 410 inputs the extracted audio segment Ausubos into video reconstructor 420, and video reconstructor 420 may acquire foreground fusion information based on the audio segment subset, the plurality of image frames in video segment 160, and the corresponding plurality of mask maps; Liu, ¶[0043]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include audio information as taught by Liu to the combined invention of Shin and Chen. The motivation to do so would be to provide a video reconstruction with greatly improved audio-video synchronization.
Claims 8 and 14 are rejected under pre-AIA 35 U.S.C. 103(a) as being unpatentable over Shin in view of Chen and further in view of “Gupta” (Gupta, Meenal, et al. "Image Segmentation based Background Removal and Replacement." Proceedings of the 2021 Thirteenth International Conference on Contemporary Computing. 2021.).
8. The combination of Shin and Chen teaches the electronic device of claim 1.
The combination does not explicitly disclose wherein the at least one object is a foreground of the first video, and the second video is a video in which the background, excluding the foreground, has been removed from the first video (Shin teaches removing a foreground object and replacing with a background object; Shin, ¶[0035].).
However, Gupta, a similar field of endeavor of background removal in video segmentation, teaches wherein the at least one object is a foreground of the first video, and the second video is a video in which the background, excluding the foreground, has been removed from the first video (Background subtraction is intended specifically for use in a video, rather than single images, and aims to separate the foreground object from the background by observing change across consecutive frames; Gupta; [§3, page 79, col 2, ¶3]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include background subtraction as taught by Gupta to the invention of Shin and Chen. The motivation to do so would be to provide some privacy during an online meeting our main objective.
Claim 14 is similarly analyzed as analogous claim 8.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. See PTO-892 Notice of References Cited for full citations.
Wang et al., (US 20120213432 A1) teaches automatic video segmentation. A segmentation shape prediction and a segmentation color model are determined for a current image of a video sequence based on existing segmentation information for at least one previous image of the video sequence. A segmentation of the current image is automatically generated based on a weighted combination of the segmentation shape prediction and the segmentation color model. The segmentation of the current image is stored in a memory medium. Wang would have been relied upon for teaching mask reliability.
Liba et al., (US 20230118361 A1) is the US Patent Application Publication for WO 2023069445 A1, cited in the IDS (01/18/2025)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHANDHANA PEDAPATI whose telephone number is 571-272-5325. The examiner can normally be reached M-F 8:30am-6pm (ET).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chan Park can be reached at 571-272-7409. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CHANDHANA PEDAPATI/Examiner, Art Unit 2669 /CHAN S PARK/Supervisory Patent Examiner, Art Unit 2669