Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 06/28/2024 and 2/13/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Response to Arguments
Applicant has amended claims 1, 3, 10, 14, 15, and 20. Claims 2 and 19 have been cancelled. Claims 1, 3-16 and 20-24 are currently being considered. Applicant’s arguments, filed 7/8/2026, with respect to the rejection(s) of claim(s) 1-2, 7-16 and 19 under 35 U.S.C 102 and claims 3-6 and 20 under 35 U.S.C 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of El-Khamy et. al. (United States Patent Application Publication US 2019/0057507 A1).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1, 7-16, 22 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wen et. al. (International Publication WO 2021/175252 A1) in view of El-Khamy et. al. (United States Patent Application Publication US 2019/0057507 A1).
Regarding claim 1, Wen et. al. discloses a video processing method, comprising: determining a target image to be processed in a video (Wen et. al., figure 8, page 14, lines 15-20); performing semantic segmentation on the target image through a convolutional neural network to obtain a first feature map, wherein the first feature map comprises a feature map corresponding to at least one semantic class (Wen et. al. figure 8, page 15, lines 1-27: categories of pixels, i.e. the semantic classes); determining a target image region corresponding to the at least one semantic class in the target image according to the first feature map having the same size of the target image (Wen et. al., page 16, lines 10-23, bounding boxes correspond to the target image region); wherein the at least one semantic class comprises an object-in-hand for any object held by hand (Wen et. al., page 15, lines 5-7 “define four categories of pixels, i.e., hand, product in hand, product on table, and background”), and a training image adopted by the convolutional neural network in a training process is marked with an image region corresponding to the at least one semantic class (Wen et. al., page 15, lines 12-22: according to the output result, it can be inferred that an image region corresponding to the semantic class is implicitly disclosed).
However, Wen et. al. fails to disclose wherein a plurality of semantic classes are provided, and different semantic classes have contextual relations, the determining a target image region corresponding to the at least one semantic class in the target image according to the first feature map, comprises: determining the target image region according to a first sub-feature map and a plurality of second sub-feature maps; wherein the first sub-feature map is a feature map corresponding to an image background in the first feature map, each second sub-feature map in the plurality of second sub-feature maps is a feature map corresponding to a semantic class in the first feature map, and different second sub-feature maps correspond to different semantic classes.
El-Khamy et. al. teaches wherein a plurality of semantic classes are provided, and different semantic classes have contextual relations, the determining a target image region corresponding to the at least one semantic class in the target image according to the first feature map, comprises: determining the target image region according to a first sub-feature map and a plurality of second sub-feature maps; wherein the first sub-feature map is a feature map corresponding to an image background in the first feature map, each second sub-feature map in the plurality of second sub-feature maps is a feature map corresponding to a semantic class in the first feature map, and different second sub-feature maps correspond to different semantic classes (El-Khamy et. al. Figure 1B, [0003]: An image recognition system provides a computer application that detects and identifies an object or multiple objects from a digital image or a video frame, [0005]: a method for detecting instances of objects in an input image includes: extracting a plurality of core instance features from the input image; calculating a plurality of feature maps at multiscale resolutions for core instance features. [0056]: The multi-resolution feature maps and the bounding boxes are supplied to a segmentation mask prediction network or segmentation mask head, which generates predicts a segmentation mask for each object class at each of the resolutions of the feature maps. The segmentation mask head is further configured to provide a pixel-level classification score for each class (e.g., each class of object to be detected by the instance semantic segmentation system, where the classes may include, for example, humans, dogs, cats, cars, debris, furniture and the like), and can also provide a pixel-level score of falling inside or outside a mask. [0057]: Generally, the later stage feature maps (e.g. the third and fifth feature maps) have a larger receptive field and also have higher resolutions than earlier stage feature maps (e.g. first feature map).
This is important to the claimed invention because this further classifies the object of interest and assigns a semantic class. Thus, it would have been obvious to one skilled in the art prior to the effective filing date of the claimed invention to have combined the teachings of Wen et. al. and El-Khamy et. al. so that these features are included in the solution of the claimed invention.
Regarding claim 14, which is a video processing device claim, corresponding to the method of claim 1, which the rejection analysis is incorporated herein.
Regarding claim 15, which is an electronic device claim comprising: at least one processor and a memory; the memory stores computer-executed instructions; the at least one processor executes the computer-executed instructions stored in the memory, causing the at least one processor to execute a video processing method, corresponding to the method of claim 1, which the rejection analysis is incorporated herein.
Regarding claim 16, which is a computer-readable storage medium with computer-executable instructions stored therein, which corresponds to the method of claim 1, which the rejection analysis is incorporated herein.
Regarding claim 7, Wen et. al. and El-Khamy et. al. disclose the video processing method according to claim 1, and Wen et. al. further discloses wherein the determining a target image to be processed in a video, comprises: determining the target image in the video according to a preset interval frame number (Wen et. al. page 5). The target image is determined in the video according to a preset number of spaced frames, which is commonly used in the art.
Regarding claim 8 and claim 22, Wen et. al. and El-Khamy et. al. disclose the video processing method according to claim 1 and the electronic device according to claim 15, and Wen et. al. further discloses after determining a target image region corresponding to the at least one semantic class in the target image according to the first feature map, further comprising: generating a target box corresponding to the object-in-hand in the target image according to the target image region; tracking the object-in-hand appearing in the target image according to the target box (Wen et. al. page 27, line 18, page 28, line 9).
Regarding claim 9, Wen et. al. further discloses the video processing method according to claim 8, wherein the tracking the object-in-hand appearing in the target image according to the target box, comprises: acquiring a tracking box, wherein the tracking box is from a previous image frame of the target image, different tracking boxes correspond to different object IDs, and each object ID is used for uniquely identifying a corresponding object-in-hand; matching the tracking box with the target box to obtain a matching result; updating the tracking box in the target image according to the matching result (Wen et. al. page 27, line 18, page 28, line 9).
Regarding claim 10, Wen et. al. further discloses the video processing method according to claim 9, wherein the updating the tracking box in the target image according to the matching result, comprises: for each target box, in response to a first tracking box matching the target box, updating the first tracking box matching the target box into the target box; in response to no tracking box matching the target box, determining that the target box is a new tracking box, and assigning an object ID to the new tracking box; for any tracking box, in response to no target box matching the tracking box, deleting the tracking box (Wen et. al. page 5, the step of tracking the product is performed using tracking-by-detection and greedy search).
Regarding claim 11, Wen et. al. further discloses the video processing method according to claim 9, wherein at least one object-in-hand appears in the video, and at least one tracking box and at least one target box are provided, the matching the tracking box with the target box to obtain a matching result, comprises: for each tracking box in the at least one tracking box and each target box in the at least one target box, determining an overlapping degree between the tracking box and the target box to obtain at least one overlapping degree; for each target box, determining a tracking box matched with the target box according to the at least one overlapping degree (Wen et. al. page 27, line 18 to page 28, line 9: matching can be carried out on the basis of the position of the bounding box, the size of the bounding box, the pixel label of the pixels in the bounding box and the depth information).
Regarding claim 12, Wen et. al. and El-Khamy et. al. disclose the video processing method according to claim 1, and Wen et. al. further discloses after determining a target image region corresponding to the semantic class in the target image according to the first feature map, further comprising: in response to a target object leaving a hand appearing in the video, tracking the target object based on an image region where the target object is located in the target image region (Wen et. al. page 5). Tracking of the target object based on the image region of the target image region in which the exiting target object is located is well within the ordinary skill in the art for a person skilled in the art.
Regarding claim 13, Wen et. al. and El-Khamy et. al. disclose the video processing method according to claim 1, and Wen et. al. further discloses wherein the performing semantic segmentation on the target image through a convolutional neural network to obtain a first feature map, comprises: performing semantic segmentation on the target image and a preset number of images in the video located in front of the target image through the convolutional neural network to obtain the first feature map, wherein the preset number of images located in front of the target image are used for assisting the semantic segmentation of the target image (Wen et. al. page 5). It is well known to those skilled in the art that semantic segmentation of a pre-set plurality of images preceding the target image assists the semantic segmentation result of the current target image, which is commonly used in the art.
Claim(s) 3-6, and 20-21, 23-24 are rejected under 35 U.S.C. 103 as being unpatentable over Wen et. al. (International Publication WO2021/175252 A1) and El-Khamy et. al. (United States Patent Application Publication US 2019/0057507 A1) in view of Cheng et. al. (Chinese Patent CN 112866797B).
Regarding claim 3, Wen et. al. and El-Khamy et. al. disclose the video processing method according to claim 2. However, Wen et. al. and El-Khamy et. al. fail to disclose wherein a pixel value in the second sub-feature map is a weight that a corresponding pixel in the target image belongs to the semantic class corresponding to the second sub-feature map, the determining the target image region according to a first sub-feature map and a plurality of second sub-feature maps, comprises: for each pixel in a plurality of pixels in the target image, determining a maximum value of the pixel among corresponding pixel values in the first sub-feature map and the plurality of second sub-feature maps, in response to a feature map where the maximum value is located being the first sub-feature map, determining that the pixel belongs to the image background, in response to the feature map where the maximum value is located not being the first sub-feature map, determining that the pixel belongs to a semantic class corresponding to the feature map where the maximum value is located; determining the target image region according to a semantic class to which each of the plurality of pixels belong.
Cheng et. al. discloses wherein a pixel value in the second sub-feature map is a weight that a corresponding pixel in the target image belongs to the semantic class corresponding to the second sub-feature map, the determining the target image region according to a first sub-feature map and a plurality of second sub-feature maps, comprises: for each pixel in a plurality of pixels in the target image, determining a maximum value of the pixel among corresponding pixel values in the first sub-feature map and the plurality of second sub-feature maps, in response to a feature map where the maximum value is located being the first sub-feature map, determining that the pixel belongs to the image background, in response to the feature map where the maximum value is located not being the first sub-feature map, determining that the pixel belongs to a semantic class corresponding to the feature map where the maximum value is located; determining the target image region according to a semantic class to which each of the plurality of pixels belong (Cheng et. al. see specification [0004]-[0274], the first semantic segmentation result of the target location region of the first image can be represented by a first probability value that a first pixel point within the target location region of the first image belongs to a target class, the target class may be at least one class pre-set, such as a foreground category that the user is focused on and a background category that the user is not focused on may be included in a frame of an image, here the target class may comprise a foreground class and a background class, and the acquired first probability value that the first pixel point belongs to the target class may comprise a first probability value of belonging to the foreground class and a first probability value of belonging to the background class. Exemplarily, target classes may include a plurality, like the foreground and background categories mentioned above, which target class the first pixel point specifically belongs to, may be determined by a first probability value that the first pixel point belongs to the target class, say the first probability value that the first pixel point belongs to the foreground class is greater than the first probability value that the first pixel point belongs to the background class, consider the first pixel point to belong to the foreground class, upon determining a target class to which the first pixel point and the second pixel point belong based on the probability values, it is also possible to further determine a confidence that a first pixel point belongs to the target class and a confidence that a second pixel point belongs to the target class, in accordance with a first sub-feature map and a plurality of second sub-feature maps, determining the target image area, comprising, for a plurality of pixel points on the target image, determining a maximum value of the pixel point among corresponding pixel values in the first sub-feature map and the second sub-feature map, determining that the pixel point belongs to an image background if the feature map in which the maximum is located is the first sub-feature map, otherwise determining that the pixel point belongs to a semantic class to which the feature map in which the maximum is located corresponds; determining the target image area according to the semantic class to which the plurality of pixel points belongs). The above technical features play the same role in the present application and Cheng et. al., as they play in the claims, all based on the probability big value to determine the corresponding semantic class. Thus, it would have been obvious to one skilled in the art prior to the claimed invention prior to the effective filing date of the claimed invention to have combined the features of Wen et. al., El-Khamy et. al. and Cheng et. al. so that these solutions are incorporated to the problems discussed in the claimed invention.
Regarding claim 20, the electronic device according to claim 15, which corresponds to the method of claim 3, which the rejection analysis is incorporated herein.
Regarding claim 4 and claim 21, Wen et. al. and El-Khamy et. al. disclose the video processing method according to claim 1 and the electronic device according to claim 15. However, Wen et. al. and El-Khamy et. al. fails to disclose wherein the determining a target image region corresponding to the at least one semantic class in the target image according to the first feature map, comprises: acquiring a second feature map, wherein the second feature map is a feature map obtained by performing semantic segmentation on a reference image, and the reference image is an image located in front of the target image in the video and at least one frame apart; fusing the first feature map and the second feature map to obtain a fused feature map; determining the target image region according to the fused feature map
Cheng et. al. teaches wherein the determining a target image region corresponding to the at least one semantic class in the target image according to the first feature map, comprises: acquiring a second feature map, wherein the second feature map is a feature map obtained by performing semantic segmentation on a reference image, and the reference image is an image located in front of the target image in the video and at least one frame apart; fusing the first feature map and the second feature map to obtain a fused feature map; determining the target image region according to the fused feature map (Cheng et. al. [0004]-[0274], In a first aspect, an embodiment of the present disclosure provides a video processing method comprising: obtaining a video segment; a first image of a current frame and a second image of a previous frame are included in the video footage; determining a first semantic segmentation result of a target location area of the first image and first feature information of the target location area of the first image; obtaining a second semantic segmentation result of the second image and second feature information of the second image; based on the first semantic segmentation result and the first feature information, the second semantic segmentation result and the second feature information, determining a second semantic segmentation result of the first image). The above technical features are important to the claimed invention as they play in the claims, all based on the fusion result of two-frame feature maps for the target area division. Thus, it would have been obvious to one skilled in the art prior to the effective filing date of the claimed invention to have combined the features of Wen et. al., El-Khamy et. al. and Cheng et. al. so that these solutions are incorporated in the method of Wen et. al.
Regarding claim 5 and claim 23, the combination of Wen et. al., El-Khamy et. al. and Cheng et. al. discloses the video processing method according to claim 4 and the electronic device according to claim 21, and Cheng et. al. further discloses wherein the fusing the first feature map and the second feature map to obtain a fused feature map, comprises: fusing a feature map corresponding to an image background in the first feature map and a feature map corresponding to the image background in the second feature map to obtain a first fused sub-feature map; fusing a feature map corresponding to a semantic class in the first feature map with a feature map corresponding to a same semantic class in the second feature map to obtain a second fused sub-feature map (Cheng et. al. see specification [0004]-[0274]). It is within the ordinary skill in the art to fuse the feature maps corresponding to the background in the front and rear two frame images while simultaneously fusing the feature maps corresponding to the same foreground semantic class. Thus, it would have been obvious to one skilled in the art prior to the effective filing date of the claimed invention to have combined the features of Wen et. al., El-Khamy et. al. and Cheng et. al. so that these solutions are incorporated in the method of Wen et. al.
Regarding claim 6 and claim 24, Wen et. al., El-Khamy et. al. and Cheng et. al. disclose the video processing method according to claim 4 and the electronic device according to claim 21, and Wen et. al. further discloses before fusing the first feature map and the second feature map to obtain a fused feature map, further comprising: determining an optical flow between the target image and the reference image; adjusting the second feature map according to the optical flow (Wen et. al. Fig. 1 is a flowchart of a method of positioning a target object in an image according to one embodiment of the present application. As shown in Figure 1, the method comprises the following steps: S1 1: a sequence of input images, said sequence of images comprising at least two frames of images; note that a sequence of images refers to a series of images sequentially acquired at different locations at different times for a target. S12: processing each image of each frame separately using a first neural network to obtain a confidence image of each image of each frame in which each pixel distinguishes between different objects with different markers; s13: performing an optical flow calculation on the sequence of images using a second neural network, resulting in an optical flow vector map; it is noted that optical flow is a concept in object motion detection in a field of view. Regarding optical flow detection of moving objects: Each pixel point in the image is given a velocity vector, which forms a motion vector field. Depending on the velocity vector characteristics of the individual pixel points, the image can be dynamically analyzed. If there is no moving object in the image, the optical flow vector varies continuously over the entire image area, here the loss of reduction if there is no object movement is 0. When there is a target object in the image, there is relative motion of the target and the background. The velocity vector formed by the target object is necessarily different from the velocity vector of the background, so that the position of the target object can be calculated. However, object detection by optical flow is not possible to detect objects with a particularly large moving distance). It is within the ordinary skill of the art to adapt the second feature map according to the optical flow. Thus, it would have been obvious to one skilled in the art prior to the effective filing date of the claimed invention to have combined the features of Wen et. al., El-Khamy et. al. and Cheng et. al. so that these solutions are incorporated in the method of Wen et. al.
Conclusion
Response to Amendment
Examiner has carefully considered the amended claims and reconsidered the prior arts of record along with an updated search. The amendment to claim 10 is sufficient to overcome the previous claim objection. New prior art was found to reject all claims 1, 3-16 and 20-24.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JESSICA YIFANG LIN whose telephone number is (571)272-6435. The examiner can normally be reached M-F 7:00am-6:15pm, with optional day off.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be reached at 571-272-7332. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JESSICA YIFANG LIN/ Examiner, Art Unit 2668 August 14, 2026
/VU LE/ Supervisory Patent Examiner, Art Unit 2668