DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Response to Preliminary Amendment
This Office Action is responsive to the Preliminary Amendment submitted on 10/18/2024 and is entered to amend paragraphs 1 and 12 of the specification. Applicant has amended Claim 10. Claims 7-9 have been canceled. Claims 11-15 have been added. Applicant has amended the abstract. Applicant submits no new matter has been added by way of these amendments. As such, the Preliminary Amendment has been entered and an Office Action on the merits follows here below.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 10/18/2024, 08/07/2025 and 06/05/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claims 2 and 11 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 2 recites: “… the label of the sample image pair comprises a first label and a second label, the first label is a segmentation result that is of the RGB image or the depth image and that is formed manually in advance, and the second label is obtained after Gaussian blur processing is performed on the first label...”
Claim 11 also recites: “…wherein the label of the sample image pair comprises a first label and a second label, the first label is a segmentation result that is of the RGB image or the depth image and that is formed manually in advance, and the second label is obtained after Gaussian blur processing is performed on the first label;”
With respect to both claims that are similarly mirrored in language, it is unclear how the segmentation results can be formed manually in advance when the first, second and third neural network/network model must be presented via a processing device. The “labeling”, according to one of ordinary skill in the art would reasonably use a computing device partnered with executable code in order to calculate segmentation results. Appropriate correction/clarification is required.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 6 and 10 are rejected under 35 U.S.C. 102(a)(1) and/or (a)(2) as being anticipated by Wang (US 20200272888 A1).
Regarding Claim 1: (Original) Wang discloses a segmentation model training method (Refer to para [002]; “Segmentation in images involves determining which pixels are assigned to figures, such as in pictures that include human bodies.”) wherein the segmentation model comprises a first network model (“a plurality of first head neural networks”) a second network model (“a plurality of second head neural networks”) and a third network model (“a plurality of third head neural networks”) and the method comprises: obtaining a sample image pair (Refer to para [023]; “One or more image sensors 28 may be provided to capture input images 27.”) wherein the sample image pair comprises an RGB image and a depth image that are obtained by photographing the same visual range (Refer to para [023]; “The computing device 12 may be configured to receive visible light image data, for example in RGB format, from the visible light camera 29 for processing. The computing device 12 may also be configured to receive depth maps from depth camera 30 and active brightness (AB) maps from an IR camera such as IR camera 34. Computing system 10 may be equipped to process, via image sensors 28, input images 27 that are real-time images captured and received in real-time.”)
inputting the depth image into the first network model (Refer to para [038]; “In this implementation, visible light and/or IR images and/or depth images may also be received as input.”) to obtain a first depth feature extraction result output by the first network model (Refer to para [044]; “The method 200 at 204 may include receiving, at the backbone network 42, an input image 27 as input and outputting feature maps extracted from the input image 27. As the CCPN learns body part localization and association, feature maps may be extracted from input images. Feature maps from the intermediate layers C.sub.1 to C.sub.5 may be concatenated as input to downstream convolutional layers, as shown in FIG. 3.”)
inputting a combined image of the depth image and the RGB image into the second network model (Refer to para [037]; “Turning now to FIGS. 8A-8C, the input image 27 may include overlapping bodies, where one or more parts of at least one body in the input image 27 are occluded by at least one additional body. FIG. 8A is an example input RGB image for human pose estimation that includes overlapping bodies.”) to obtain an edge feature that is of a target object and that is output by the second network model (Refer to para [028]; “As shown in FIG. 4, one implementation may include 18 keypoints 46 (19 keypoints when the background is included); thus, 18 keypoint heatmaps 52 corresponding to each respective keypoint 46 may be generated by the first head neural networks 44. As shown in FIG. 2, the keypoint heatmaps 52 may be fed to a program executing an edge detector, which may execute edge thinning using a technique such as non-maximum suppression (NMS). This additional processing may output a set of peaks for the keypoints 46 that identify locations of keypoints 46 with labels and associated probabilities for each of the keypoints 46 to be a body part.”)
inputting the edge feature of the target object and the first depth feature extraction result into the third network model, to obtain a segmentation result that is of the target object and that is output by the third network model (Refer to para [038]; “Turning now to FIG. 9, the processor 14 may be configured to execute a plurality of third head neural networks 62 having been trained to determine a probability that each pixel in the input image 27 belongs to one of a plurality of segments 68. FIG. 9 depicts the computing system 10 similarly to FIG. 2 but with additional third head neural networks 62 for body part segmentation. In this implementation, visible light and/or IR images and/or depth images may also be received as input. At the third head neural networks 62, the feature maps may be processed using each of the third head neural networks 62 to output corresponding instance segmentation maps 70 for each segment 68 indicating the probability that each pixel in the input image 27 belongs to a corresponding one of the plurality of segments 68.”) and performing a parameter adjustment on the first network model, the second network model, and the third network model based on a label of the sample image pair and the segmentation result of the target object (Refer to para [061]; “Non-limiting examples of training procedures for adjusting trainable parameters include supervised training (e.g., using gradient descent or any other suitable optimization method), zero-shot, few-shot, unsupervised learning methods (e.g., classification based on classes derived from unsupervised clustering methods), reinforcement learning (e.g., deep Q learning based on feedback) and/or generative adversarial neural network training methods, belief propagation, RANSAC (random sample consensus), contextual bandit methods, maximum likelihood methods, and/or expectation maximization. In some examples, a plurality of methods, processes, and/or components of systems described herein may be trained simultaneously with regard to an objective function measuring performance of collective functioning of the plurality of components (e.g., with regard to reinforcement feedback and/or with regard to labelled training data). Simultaneously training the plurality of methods, processes, and/or components may improve such collective functioning. In some examples, one or more methods, processes, and/or components may be trained independently of other components (e.g., offline training on historical data).”).
Regarding Claim 6: (Original) Wang discloses an image recognition method (Refer to para [043]; “FIG. 13 shows a flowchart of a method 200 for use with a computing device of the computing system 10. The following description of method 200 is provided with reference to the computing system 10, implementations of which are described above and shown in FIGS. 2 and 8. It will be appreciated that method 200 may also be performed in other contexts using other suitable components.”) comprising: obtaining a to-be-processed RGB image and a to-be-processed depth image that are obtained by photographing the same visual range (Refer to para [025]; “The input image 27 may include one or more of visible light image data, depth data, and active brightness data, which may include input received from a visible light camera, a depth camera, or an infrared camera. The input may be real-time input as well. Thus, to name a few examples, the input image 27 may include color image data such as RGB data captured from a visible light camera, depth data (referred to herein as a depth map), and/or active brightness (AB) data from raw infrared (IR) images (referred to herein as an AB map) from an infrared camera.”) inputting the to-be-processed depth image into a first network model (Refer to para [038]; “In this implementation, visible light and/or IR images and/or depth images may also be received as input.”) to obtain a depth feature extraction result output by the first network model (Refer to para [044]; “The method 200 at 204 may include receiving, at the backbone network 42, an input image 27 as input and outputting feature maps extracted from the input image 27. As the CCPN learns body part localization and association, feature maps may be extracted from input images. Feature maps from the intermediate layers C.sub.1 to C.sub.5 may be concatenated as input to downstream convolutional layers, as shown in FIG. 3.”) inputting a combined image of the to-be-processed depth image and the to-be-processed RGB image into a second network model (Refer to para [037]; “Turning now to FIGS. 8A-8C, the input image 27 may include overlapping bodies, where one or more parts of at least one body in the input image 27 are occluded by at least one additional body. FIG. 8A is an example input RGB image for human pose estimation that includes overlapping bodies.”) to obtain an edge feature that is of a target object and that is output by the second network model (Refer to para [028]; “As shown in FIG. 4, one implementation may include 18 keypoints 46 (19 keypoints when the background is included); thus, 18 keypoint heatmaps 52 corresponding to each respective keypoint 46 may be generated by the first head neural networks 44. As shown in FIG. 2, the keypoint heatmaps 52 may be fed to a program executing an edge detector, which may execute edge thinning using a technique such as non-maximum suppression (NMS). This additional processing may output a set of peaks for the keypoints 46 that identify locations of keypoints 46 with labels and associated probabilities for each of the keypoints 46 to be a body part.”)
and inputting the edge feature of the target object and the depth feature extraction result (Refer to para [028]; “As shown in FIG. 4, one implementation may include 18 keypoints 46 (19 keypoints when the background is included); thus, 18 keypoint heatmaps 52 corresponding to each respective keypoint 46 may be generated by the first head neural networks 44. As shown in FIG. 2, the keypoint heatmaps 52 may be fed to a program executing an edge detector, which may execute edge thinning using a technique such as non-maximum suppression (NMS). This additional processing may output a set of peaks for the keypoints 46 that identify locations of keypoints 46 with labels and associated probabilities for each of the keypoints 46 to be a body part.”) into a third network model (“a plurality of third head neural networks”) to obtain a segmentation result that is of the target object and that is output by the third network model (Refer to para [061]; “Non-limiting examples of training procedures for adjusting trainable parameters include supervised training (e.g., using gradient descent or any other suitable optimization method), zero-shot, few-shot, unsupervised learning methods (e.g., classification based on classes derived from unsupervised clustering methods), reinforcement learning (e.g., deep Q learning based on feedback) and/or generative adversarial neural network training methods, belief propagation, RANSAC (random sample consensus), contextual bandit methods, maximum likelihood methods, and/or expectation maximization. In some examples, a plurality of methods, processes, and/or components of systems described herein may be trained simultaneously with regard to an objective function measuring performance of collective functioning of the plurality of components (e.g., with regard to reinforcement feedback and/or with regard to labelled training data). Simultaneously training the plurality of methods, processes, and/or components may improve such collective functioning. In some examples, one or more methods, processes, and/or components may be trained independently of other components (e.g., offline training on historical data).”).
Regarding Claim 10: (Currently Amended) Wang discloses a computing device (“computing device 12”) comprising a memory (“volatile memory 16 and non-volatile memory 18 in a computing device 12.”) and a processor (“the processor 14 may be configured to execute one or more programs stored in the memory to execute a CNN 22 that has been trained using a training data set 60.”) wherein the memory stores executable code, and when the processor executes the executable code (Refer to para [022]; “The processor 14 may be a GPU, CPU, FPGA, ASIC, and so on, and may be configured to execute one or more programs stored in the non-volatile memory 18. These programs and the data they utilize may include image pre-processing programs 20, convolutional neural network (CNN) 22, image algorithms 24 for statistical analysis such as non-maximum suppression (NMS), image rendering programs 26 for the output of image processing by the CNN 22.”) implemented the processor is caused to implement a segmentation model training method (Refer to para [002]; “Segmentation in images involves determining which pixels are assigned to figures, such as in pictures that include human bodies.”) wherein the segmentation model comprises a first network model (“a plurality of first head neural networks”) a second network model (“a plurality of second head neural networks”) and a third network model (“a plurality of third head neural networks”) the method comprising:
obtaining a sample image pair (Refer to para [023]; “One or more image sensors 28 may be provided to capture input images 27.”) wherein the sample image pair comprises an RGB image and a depth image that are obtained by photographing the same visual range (Refer to para [023]; “The computing device 12 may be configured to receive visible light image data, for example in RGB format, from the visible light camera 29 for processing. The computing device 12 may also be configured to receive depth maps from depth camera 30 and active brightness (AB) maps from an IR camera such as IR camera 34. Computing system 10 may be equipped to process, via image sensors 28, input images 27 that are real-time images captured and received in real-time.”) inputting the depth image into the first network model (Refer to para [038]; “In this implementation, visible light and/or IR images and/or depth images may also be received as input.”) to obtain a first depth feature extraction result output by the first network model (Refer to para [044]; “The method 200 at 204 may include receiving, at the backbone network 42, an input image 27 as input and outputting feature maps extracted from the input image 27. As the CCPN learns body part localization and association, feature maps may be extracted from input images. Feature maps from the intermediate layers C.sub.1 to C.sub.5 may be concatenated as input to downstream convolutional layers, as shown in FIG. 3.”) inputting a combined image of the depth image and the RGB image into the second network model (Refer to para [037]; “Turning now to FIGS. 8A-8C, the input image 27 may include overlapping bodies, where one or more parts of at least one body in the input image 27 are occluded by at least one additional body. FIG. 8A is an example input RGB image for human pose estimation that includes overlapping bodies.”) to obtain an edge feature that is of a target object and that is output by the second network model (Refer to para [028]; “As shown in FIG. 4, one implementation may include 18 keypoints 46 (19 keypoints when the background is included); thus, 18 keypoint heatmaps 52 corresponding to each respective keypoint 46 may be generated by the first head neural networks 44. As shown in FIG. 2, the keypoint heatmaps 52 may be fed to a program executing an edge detector, which may execute edge thinning using a technique such as non-maximum suppression (NMS). This additional processing may output a set of peaks for the keypoints 46 that identify locations of keypoints 46 with labels and associated probabilities for each of the keypoints 46 to be a body part.”) inputting the edge feature of the target object and the first depth feature extraction result into the third network model (Refer to para [028]; “As shown in FIG. 4, one implementation may include 18 keypoints 46 (19 keypoints when the background is included); thus, 18 keypoint heatmaps 52 corresponding to each respective keypoint 46 may be generated by the first head neural networks 44. As shown in FIG. 2, the keypoint heatmaps 52 may be fed to a program executing an edge detector, which may execute edge thinning using a technique such as non-maximum suppression (NMS). This additional processing may output a set of peaks for the keypoints 46 that identify locations of keypoints 46 with labels and associated probabilities for each of the keypoints 46 to be a body part.”) to obtain a segmentation result that is of the target object and that is output by the third network model (Refer to para [038]; “Turning now to FIG. 9, the processor 14 may be configured to execute a plurality of third head neural networks 62 having been trained to determine a probability that each pixel in the input image 27 belongs to one of a plurality of segments 68. FIG. 9 depicts the computing system 10 similarly to FIG. 2 but with additional third head neural networks 62 for body part segmentation. In this implementation, visible light and/or IR images and/or depth images may also be received as input. At the third head neural networks 62, the feature maps may be processed using each of the third head neural networks 62 to output corresponding instance segmentation maps 70 for each segment 68 indicating the probability that each pixel in the input image 27 belongs to a corresponding one of the plurality of segments 68.”) and performing a parameter adjustment on the first network model, the second network model, and the third network model based on a label of the sample image pair and the segmentation result of the target object (Refer to para [061]; “Non-limiting examples of training procedures for adjusting trainable parameters include supervised training (e.g., using gradient descent or any other suitable optimization method), zero-shot, few-shot, unsupervised learning methods (e.g., classification based on classes derived from unsupervised clustering methods), reinforcement learning (e.g., deep Q learning based on feedback) and/or generative adversarial neural network training methods, belief propagation, RANSAC (random sample consensus), contextual bandit methods, maximum likelihood methods, and/or expectation maximization. In some examples, a plurality of methods, processes, and/or components of systems described herein may be trained simultaneously with regard to an objective function measuring performance of collective functioning of the plurality of components (e.g., with regard to reinforcement feedback and/or with regard to labelled training data). Simultaneously training the plurality of methods, processes, and/or components may improve such collective functioning. In some examples, one or more methods, processes, and/or components may be trained independently of other components (e.g., offline training on historical data).”).
Allowable Subject Matter
Claims 3-5, 12-15 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claims 2 and 11 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 112(b) set forth in this Office action and to include all of the limitations of the base claim and any intervening claims.
The prior art either singly or in combination does not teach, disclose or suggest at least the following claim limitation(s): “… performing feature extraction on the combined image through each layer of front-end neural network comprised in the second network model, to obtain a primary edge feature; and processing the primary edge feature and the second depth feature extraction result though each layer of back-end neural network comprised in the second network model, to obtain and output the edge feature of the target object.”
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Tsai (US 11276177 B1) discloses “… obtaining a first image of a target and a second image of the target, the first image having a first field-of-view (FOV) and the second image having a second FOV that is different than the first FOV; determining, based on the first image, a first segmentation map comprising foreground prediction values associated with a first estimated foreground region in the first image; determining, based on the second image, a second segmentation map comprising foreground prediction values associated with a second estimated foreground region in the second image; generating a third segmentation map based on the first segmentation map and the second segmentation map; and generating, using the second segmentation map and the third segmentation map, a refined segmentation mask that identifies at least a portion of the target as a foreground region of at least one of the first image and the second image.”
Related Application(s): Hu (US 20230411574 A1)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MIA M THOMAS whose telephone number is (571)270-1583. The examiner can normally be reached M-Th 8:30am-4:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Stephen (Steve) Koziol can be reached at (408) 918-7630. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
MIA M. THOMAS
Primary Examiner
Art Unit 2665
/MIA M THOMAS/Primary Examiner
Art Unit 2665