DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged that application claims priority to foreign application
with application number JP 2023-203957 dated 12/01/2023. Copies of certified
papers required by 37 CFR 1.55 have been retrieved.
Information Disclosure Statement
The information disclosure statement (“IDS”) filed on 11/26/2024, 02/14/2025, and 04/09/2025 has been reviewed and the listed references have been considered.
Drawings
The 22-page drawings have been considered and placed on record in the file.
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
Status of Claims
Claims 1-19 are pending.
Claim Objections
Claims 14 is objected to because of the following informality:
Claim 14 recites "joint point likelihood map of one object in two objects in a vertical direction to calculate" should be "joint point likelihood map of one object of the plurality of objects in a vertical direction to calculate".
Appropriate corrections are required.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpretated under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Because the claim limitations use a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are, "…person detection unit…" and "…fall determination unit…" in claims 1-19, "…removal unit…" in claim 4, "…similarity calculation unit…" in claims 7-8, 11, 13-17, "…acquisition unit…" in claim 8, "…selection unit…" in claims 9-10, "…joint point detection unit…" in claims 11, 13-16, "…human body ratio calculation unit…" in claim 16, and "…generation unit…" in claim 17.
Because of these claim limitations being interpretated under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 19 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter as follows. Claim 18 is directed to a “storage medium”. Applicant’s specification states “storage medium (which may also be referred to more fully as a 'non-transitory computer-readable storage medium')”; however, the specification does not exclude use of transitory propagating signals such as carrier waves. The broadest reasonable interpretation of a claim drawn to a storage medium typically covers forms of non-transitory tangible media and transitory propagating signals per se in view of the ordinary and customary meaning of computer readable media. See Subject Matter Eligibility of Computer Readable Media, 1351 OG 212 (26 Jan 2010). See MPEP 2111.01. Signals are nothing but the physical characteristics of a form of energy, and as such is nonstatutory natural phenomena. See, e.g., In re Nuitjen, 500 F. 3d 1346, 1357 (Fed. Cir. 2007)(slip. op. at 18)("A transitory, propagating signal like Nuitjen's is not a process, machine, manufacture, or composition of matter.' … Thus, such a signal cannot be patentable subject matter."). Thus, claim 19 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-2, 4-5, 7-9, 11-13, and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Ng et al. (US 2020/0211154 A1) in view of Balavalikar Krishnamurthy et al. (US 2023/0351727 A1).
Regarding claim 1, Ng teaches “An image processing apparatus (Ng paragraph [0053] "embedded fall-detection system 100 includes a fall-detection engine 101 and a camera 102") comprising:
a person detection unit configured to detect an object representing a person from an image (Ng paragraph [0081] "Referring back to FIG. 1, note that pose-estimation module 106 is coupled to action-recognition module 108, which is configured to receive the outputs from the pose estimation module. In some embodiments, the outputs from pose-estimation module 106 can include detected human keypoints 122, the associated skeleton diagram (also referred to as the "stick figure diagram" throughout), and a two-dimensional (2-D) image 132 of the detected person cropped out from original video image 104 based on the detected keypoints 122 (also referred to as "cropped image 132" of the detected person)"); and
a fall determination unit configured to determine, (Ng paragraph [0054] "Embedded fall-detection system 100 can use camera 102 to monitor human activities within a given space such as a room, a house, a lobby, or a hallway, and to capture video images and/or still images which can be used for fall analysis and prediction. In some embodiments, when embedded fall-detection system 100 is active, camera 102 generates and outputs video images 104 which can includes video images of one or multiple persons present in the monitored space. Fall-detection engine 101 receives video images 104 as input and subsequently processes input video images 104 and makes fall/non-fall predictions/decisions based on the processed video images 104").“
However, Ng is not relied on to teach “on similarity between a plurality of objects detected by the person detection unit and an image in each area of the plurality of objects”.
Balavalikar Krishnamurthy teaches “on similarity between a plurality of objects detected by the person detection unit and an image in each area of the plurality of objects (Balavalikar Krishnamurthy paragraph [0035] "The number of a similarity measure is a measure of the degree to which the pair of the sub-images (106) match each other. Thus, for example, a higher similarity measure indicates a higher probability that a pair of instances of a selected object type match each other. In a specific example, if the first sub-image (110) and the second sub-image (112) both are in the set of selected object types (116), (e.g., both the first sub-image (110) and the second sub-image (112) are "heads"), then the similarity measures (122) indicates how closely the first sub-image (110) and the second sub-image (112) match each other (e.g., whether the first sub-image (110) and the second sub-image (112) represent both a physical head and a reflection of that physical head)")”.
It would have been obvious to a person having ordinary skill in the art before
effective filing date of the claimed invention of the instant application to combine an
imaging processing system for detecting whether a person has fell or not as taught by Ng to include analysis of whether an object is an reflection of an object as taught by Balavalikar Krishnamurthy.
The suggestion/motivation for doing so would have been that there is a need in the field of fall detection to accurately determine whether a person has fallen and reduce false detections. Once with ordinary skill in the art understands the shadows and reflections can cause false detection and unexpected behavior "Video conferencing systems may use detection and tracking and detection software to identify sub-images of objects shown in an image or a video stream. However, the tracking and detection and detection software may undesirably detect a sub-image of a reflection of a person as a sub-image of a real person. Thus, for example, if a camera is capturing an image or a video stream of a conference room having a glass wall, glass window, or any reflective surface, then the tracking and detection and detection software undesirably may treat images of person's reflection in the glass as images of a real person" as disclosed by Balavalikar Krishnamurthy in paragraph 1.
Therefore, it would have been obvious to combine the disclosure of Ng with the Balavalikar Krishnamurthy disclosure to obtain the invention as specified in claim 1 as there is a reasonable expectation of success and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Claim 18 recites a method with steps corresponding to the apparatus with elements recited in claim 1. Therefore, the recited steps of this claim are mapped to the
proposed combination in the same manner as the corresponding elements of device
claim 1. Additionally, the rationale and motivation to combine the Ng and Balavalikar Krishnamurthy references, presented in rejection of claim 1 apply to this claim.
Claim 19 recites a storage medium including computer executable instructions corresponding to the elements of the apparatus recited in claim 1. Therefore, the recited instructions of the storage of claim 15 are mapped to the proposed combination in the same manner as the corresponding elements of the apparatus claim 1. Additionally, the rationale and motivation to combine Ng and Balavalikar Krishnamurthy presented in rejection of claim 1, apply to this claim.
Regarding claim 2, the combination of Ng and Balavalikar Krishnamurthy teaches “The image processing apparatus according to claim 1, wherein, in a case where there is a second object of which similarity to a first object is greater than or equal to a threshold value (Balavalikar Krishnamurthy paragraph [0038] "if the first sub-image (110) and the second sub-image (112) together have a similarity measure of 0.99, and if the similarity threshold value (124) is 0.85, then the first sub-image (110) is determined to match the second sub-image (112). In a more specific example, the first sub-image (110) and the second sub-image (112) are both heads. As a pair, the first sub-image (110) and the second sub-image (112) have a similarity measure of0.99. Thus, in this example, a determination is made that the first sub image (110) and the second sub-image (112) are either matching heads (e.g., twins are present in the room) or that one of the first sub-image (110) and the second sub-image (112) is a sub-image of a physical person's head and the other is a sub-image of a reflection of the physical person's head"), the fall determination unit does not determine a fall state of the first object or the second object (Ng paragraph [0120] "Process 600 next determines if a fall has been detected based on the fall/non-fall decision (step 610). For example, using state transition diagram 500, step 610 determines that, after processing the multiple consecutive video images, whether the current state of the system is in red state 508 of state transition diagram 500 or not. If so, process 600 generates a fall alarm/notification (step 612). Otherwise, process 600 can return to step 608 to use the most recent").”
The proposed combination as well as the motivation for combining Ng and Balavalikar Krishnamurthy references presented in the rejection of claim 1, applies to claim 2. Finally the apparatus recited in claim 2 is met by Ng and Balavalikar Krishnamurthy.
Regarding claim 4, the combination of Ng and Balavalikar Krishnamurthy teaches “The image processing apparatus according to claim 1, further comprising a removal unit configured to remove (Balavalikar Krishnamurthy paragraph [0015] "The second filter compares detected sub-images with each other, and discards one or more detected sub images when two or more detected sub-images are sufficiently similar"), in a case where a first object of which a fall state is determined by the fall determination unit satisfies a predetermined condition (Balavalikar Krishnamurthy paragraph [0038] "if the first sub-image (110) and the second sub-image (112) together have a similarity measure of 0.99, and if the similarity threshold value (124) is 0.85, then the first sub-image (110) is determined to match the second sub-image (112). In a more specific example, the first sub-image (110) and the second sub-image (112) are both heads. As a pair, the first sub-image (110) and the second sub-image (112) have a similarity measure of0.99. Thus, in this example, a determination is made that the first sub image (110) and the second sub-image (112) are either matching heads (e.g., twins are present in the room) or that one of the first sub-image (110) and the second sub-image (112) is a sub-image of a physical person's head and the other is a sub-image of a reflection of the physical person's head"), the first object from a detection target (Balavalikar Krishnamurthy paragraph [0064] "Step 208 includes removing, responsive to the similarity measure exceeding a similarity threshold value and the first confidence score exceeding the second confidence score, the second sub-image from the set of sub images").”
The proposed combination as well as the motivation for combining Ng and Balavalikar Krishnamurthy references presented in the rejection of claim 1, applies to claim 4. Finally the apparatus recited in claim 4 is met by Ng and Balavalikar Krishnamurthy.
Regarding claim 5, the combination of Ng and Balavalikar Krishnamurthy teaches “The image processing apparatus according to claim 4, wherein the predetermined condition is a condition regarding a fall state of the first object and a second object of which similarity to the first object is greater than or equal to a threshold value (Balavalikar Krishnamurthy paragraph [0038] "if the first sub-image (110) and the second sub-image (112) together have a similarity measure of 0.99, and if the similarity threshold value (124) is 0.85, then the first sub-image (110) is determined to match the second sub-image (112). In a more specific example, the first sub-image (110) and the second sub-image (112) are both heads. As a pair, the first sub-image (110) and the second sub-image (112) have a similarity measure of0.99. Thus, in this example, a determination is made that the first sub image (110) and the second sub-image (112) are either matching heads (e.g., twins are present in the room) or that one of the first sub-image (110) and the second sub-image (112) is a sub-image of a physical person's head and the other is a sub-image of a reflection of the physical person's head").”
The proposed combination as well as the motivation for combining Ng and Balavalikar Krishnamurthy references presented in the rejection of claim 1, applies to claim 5. Finally the apparatus recited in claim 5 is met by Ng and Balavalikar Krishnamurthy.
Regarding claim 7, the combination of Ng and Balavalikar Krishnamurthy teaches “The image processing apparatus according to claim 1, wherein the person detection unit detects an object (Ng paragraph [0081] "Referring back to FIG. 1, note that pose-estimation module 106 is coupled to action-recognition module 108, which is configured to receive the outputs from the pose estimation module. In some embodiments, the outputs from pose-estimation module 106 can include detected human keypoints 122, the associated skeleton diagram (also referred to as the "stick figure diagram" throughout), and a two-dimensional (2-D) image 132 of the detected person cropped out from original video image 104 based on the detected keypoints 122 (also referred to as "cropped image 132" of the detected person)") using a first learning model (Ng paragraph [0052] "embedded fall-detection system are based on implementing various deep-learning-based fast neural networks while combining various optimization techniques, such as network pruning, quantization, and depth-wise convolution. As a result, the disclosed embedded fall-detection system can perform a multitude of deep-learning-based functionalities such as real-time deep-learning-based pose estimation, action recognition, fall detection, face detection, and face recognition") and acquires a feature map of the object from the first learning model (Ng paragraph [0071] "These pose-estimation techniques first use a strong CNN-based feature extractor to extract visual features from an input image, and then use a two-branch multi-stage CNN to detect various human keypoints within the input image), and
wherein the image processing apparatus further comprises a similarity calculation unit (Balavalikar Krishnamurthy paragraph [0015] "Similarity software assigns a similarity measure to the two detected sub-images, as compared to each other. If the similarity measure is above a similarity threshold, then the detected sub-image with the lower confidence score is removed before further processing of the video stream or image") configured to calculate similarity between the plurality of objects using the feature map acquired by the person detection unit (Balavalikar Krishnamurthy paragraph [0034] "The similarity measures (122) are numbers assigned to pairs of the sub-images (106) that match one of the selected object types (116) to within a confidence threshold value (120). There are various methods to compute the similarity measures (122). One of the methods is to compute the L2-distance (Euclidean) distance between features extracted from the detections. The smaller the distance, the larger the match. Computing the Cosine similarity match is another method for computing the one or more similarity measures (122). Computing an image hash is yet another such method").“
The proposed combination as well as the motivation for combining Ng and Balavalikar Krishnamurthy references presented in the rejection of claim 1, applies to claim 7. Finally the apparatus recited in claim 7 is met by Ng and Balavalikar Krishnamurthy.
Regarding claim 8, the combination of Ng and Balavalikar Krishnamurthy teaches “The image processing apparatus according to claim 1, further comprising: an acquisition unit configured to input the image of each area of the plurality of objects to a second learning model (Ng paragraph [0052] "embedded fall-detection system are based on implementing various deep-learning-based fast neural networks while combining various optimization techniques, such as network pruning, quantization, and depth-wise convolution. As a result, the disclosed embedded fall-detection system can perform a multitude of deep-learning-based functionalities such as real-time deep-learning-based pose estimation, action recognition, fall detection, face detection, and face recognition") for determining a fall state of a person to acquire a feature map corresponding to each of the plurality of objects (Ng paragraph [0081] "Referring back to FIG. 1, note that pose-estimation module 106 is coupled to action-recognition module 108, which is configured to receive the outputs from the pose estimation module. In some embodiments, the outputs from pose-estimation module 106 can include detected human keypoints 122, the associated skeleton diagram (also referred to as the "stick figure diagram" throughout), and a two-dimensional (2-D) image 132 of the detected person cropped out from original video image 104 based on the detected keypoints 122 (also referred to as "cropped image 132" of the detected person). Action-recognition module 108 is further configured to predict, based on the outputs from pose-estimation module 106, what type of action or activity the detected person is associated with"); and
a similarity calculation unit (Balavalikar Krishnamurthy paragraph [0015] "Similarity software assigns a similarity measure to the two detected sub-images, as compared to each other. If the similarity measure is above a similarity threshold, then the detected sub-image with the lower confidence score is removed before further processing of the video stream or image") configured to calculate similarity between the plurality of objects using the feature map acquired by the acquisition unit (Balavalikar Krishnamurthy paragraph [0034] "The similarity measures (122) are numbers assigned to pairs of the sub-images (106) that match one of the selected object types (116) to within a confidence threshold value (120). There are various methods to compute the similarity measures (122). One of the methods is to compute the L2-distance (Euclidean) distance between features extracted from the detections. The smaller the distance, the larger the match. Computing the Cosine similarity match is another method for computing the one or more similarity measures (122). Computing an image hash is yet another such method").”
The proposed combination as well as the motivation for combining Ng and Balavalikar Krishnamurthy references presented in the rejection of claim 1, applies to claim 8. Finally the apparatus recited in claim 8 is met by Ng and Balavalikar Krishnamurthy.
Regarding claim 9, the combination of Ng and Balavalikar Krishnamurthy teaches “The image processing apparatus according to claim 1, further comprising: a selection unit configured to select a pair that is highly likely to be a combination of a person and a reflected image of the person in the plurality of objects detected by the person detection unit (Balavalikar Krishnamurthy paragraph [0090] "Step 708 includes applying a second filter. The second filter may be the second filter (138) of FIG. 1. The second filter finds pairs of heads that are similar to each other, and removes sub-images of heads from the pairs. The removed sub-images are those sub-images that have lower confidence scores"); and
a similarity calculation unit configured to calculate similarity between objects corresponding to the pair selected by the selection unit (Balavalikar Krishnamurthy paragraph [0015] "Similarity software assigns a similarity measure to the two detected sub-images, as compared to each other. If the similarity measure is above a similarity threshold, then the detected sub-image with the lower confidence score is removed before further processing of the video stream or image").”
The proposed combination as well as the motivation for combining Ng and Balavalikar Krishnamurthy references presented in the rejection of claim 1, applies to claim 9. Finally the apparatus recited in claim 9 is met by Ng and Balavalikar Krishnamurthy.
Regarding claim 11, the combination of Ng and Balavalikar Krishnamurthy teaches “The image processing apparatus according to claim 1, further comprising: a joint point detection unit configured to detect a joint point of a person from an image of each area of the plurality of objects (Ng paragraph [0070] "pose-estimation module 106 can first identify a set of human keypoints 122 ( or simply "human keypoints 122" or "keypoints 122") for the detected person within an input video image 104, and then represent a pose of the detected person using the configuration and/or localization of the set of keypoints, wherein the set of keypoints 122 can include, but are not limited to: the eyes, the nose, the ears, the chest, the shoulders, the elbows, the wrists, the knees, the hip joints, and the ankles of the person. In some embodiments, instead of using a full set of keypoints, a simplified set of keypoints 122 can include just the head, the shoulders, the arms, and the legs of the detected person. A person of ordinary skill in the art can easily appreciate that a different pose of the detected person can be represented by a different geometric configuration of the set of keypoints 122"); and
a similarity calculation unit configured to calculate similarity between the plurality of objects based on a detection result by the joint point detection unit (Balavalikar Krishnamurthy paragraph [0015] "Similarity software assigns a similarity measure to the two detected sub-images, as compared to each other. If the similarity measure is above a similarity threshold, then the detected sub-image with the lower confidence score is removed before further processing of the video stream or image").”
The proposed combination as well as the motivation for combining Ng and Balavalikar Krishnamurthy references presented in the rejection of claim 1, applies to claim 11. Finally the apparatus recited in claim 11 is met by Ng and Balavalikar Krishnamurthy.
Regarding claim 12, the combination of Ng and Balavalikar Krishnamurthy teaches “The image processing apparatus according to claim 11, wherein the fall determination unit determines fall states of the plurality of persons corresponding to the plurality of objects (Ng paragraph [0120] "Process 600 next determines if a fall has been detected based on the fall/non-fall decision (step 610). For example, using state transition diagram 500, step 610 determines that, after processing the multiple consecutive video images, whether the current state of the system is in red state 508 of state transition diagram 500 or not. If so, process 600 generates a fall alarm/notification (step 612). Otherwise, process 600 can return to step 608 to use the most recent") based on the detection result by the joint point detection unit (Ng paragraph [0117] "for a given video image in the sequence of video images, process 600 detects each person in the video image, and subsequently estimates a pose for each detected person and generates a cropped image for the detected person (step 604). For example, process 600 can first identify a set of human keypoints for each detected person and then generate a skeleton diagram/stick figure of the detected person by connecting neighboring keypoints with straight lines. In various embodiments, step 604 can be performed by the disclosed pose-estimation module 106 of embedded fall detection system 100").”
The proposed combination as well as the motivation for combining Ng and Balavalikar Krishnamurthy references presented in the rejection of claim 1, applies to claim 12. Finally the apparatus recited in claim 12 is met by Ng and Balavalikar Krishnamurthy.
Regarding claim 13, the combination of Ng and Balavalikar Krishnamurthy teaches “The image processing apparatus according to claim 11, wherein the joint point detection unit detects a joint point using a third learning model (Ng paragraph [0052] "embedded fall-detection system are based on implementing various deep-learning-based fast neural networks while combining various optimization techniques, such as network pruning, quantization, and depth-wise convolution. As a result, the disclosed embedded fall-detection system can perform a multitude of deep-learning-based functionalities such as real-time deep-learning-based pose estimation, action recognition, fall detection, face detection, and face recognition") and acquires a feature map of the object from the third learning model (Ng paragraph [0070] "pose-estimation module 106 can first identify a set of human keypoints 122 ( or simply "human keypoints 122" or "keypoints 122") for the detected person within an input video image 104, and then represent a pose of the detected person using the configuration and/or localization of the set of keypoints, wherein the set of keypoints 122 can include, but are not limited to: the eyes, the nose, the ears, the chest, the shoulders, the elbows, the wrists, the knees, the hip joints, and the ankles of the person. In some embodiments, instead of using a full set of keypoints, a simplified set of keypoints 122 can include just the head, the shoulders, the arms, and the legs of the detected person. A person of ordinary skill in the art can easily appreciate that a different pose of the detected person can be represented by a different geometric configuration of the set of keypoints 122"), and
wherein the image processing apparatus further comprises a similarity calculation unit (Balavalikar Krishnamurthy paragraph [0015] "Similarity software assigns a similarity measure to the two detected sub-images, as compared to each other. If the similarity measure is above a similarity threshold, then the detected sub-image with the lower confidence score is removed before further processing of the video stream or image") configured to calculate similarity between the plurality of objects using the feature map acquired by the joint point detection unit (Balavalikar Krishnamurthy paragraph [0034] "The similarity measures (122) are numbers assigned to pairs of the sub-images (106) that match one of the selected object types (116) to within a confidence threshold value (120). There are various methods to compute the similarity measures (122). One of the methods is to compute the L2-distance (Euclidean) distance between features extracted from the detections. The smaller the distance, the larger the match. Computing the Cosine similarity match is another method for computing the one or more similarity measures (122). Computing an image hash is yet another such method").”
The proposed combination as well as the motivation for combining Ng and Balavalikar Krishnamurthy references presented in the rejection of claim 1, applies to claim 13. Finally the apparatus recited in claim 13 is met by Ng and Balavalikar Krishnamurthy.
Claims 3, 6, and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Ng and Balavalikar Krishnamurthy in view of Daehee et al. ("Identifying Reflected Images From Object Detector in Indoor Environment Utilizing Depth Information" - From IDS).
Regarding claim 3, the combination of Ng and Balavalikar Krishnamurthy teaches “The image processing apparatus according to claim 1, wherein, in a case where there is a second object of which similarity to a first object is greater than or equal to a threshold value (Balavalikar Krishnamurthy paragraph [0038] "if the first sub-image (110) and the second sub-image (112) together have a similarity measure of 0.99, and if the similarity threshold value (124) is 0.85, then the first sub-image (110) is determined to match the second sub-image (112). In a more specific example, the first sub-image (110) and the second sub-image (112) are both heads. As a pair, the first sub-image (110) and the second sub-image (112) have a similarity measure of0.99. Thus, in this example, a determination is made that the first sub image (110) and the second sub-image (112) are either matching heads (e.g., twins are present in the room) or that one of the first sub-image (110) and the second sub-image (112) is a sub-image of a physical person's head and the other is a sub-image of a reflection of the physical person's head"), based on a (Ng paragraph [0120] "Process 600 next determines if a fall has been detected based on the fall/non-fall decision (step 610). For example, using state transition diagram 500, step 610 determines that, after processing the multiple consecutive video images, whether the current state of the system is in red state 508 of state transition diagram 500 or not. If so, process 600 generates a fall alarm/notification (step 612). Otherwise, process 600 can return to step 608 to use the most recent").”
However, the combination of Ng and Balavalikar Krishnamurthy is not relied on to teach “positional relationship between the first object and the second object”.
Daehee teaches “on a positional relationship between the first object and the second object (Daehee page 3 right hand column paragraph 2 "The proposed algorithm is composed of the following orders. First, bounding boxes of objects (person in our experiment) are detected from color images by a conventional detector. Coordinates of bounding boxes are then converted from color image coordinates into depth camera world coordinates provided with depth information. It is then used to compare positional relationship with reference surface in world coordinate").”
It would have been obvious to a person having ordinary skill in the art before
effective filing date of the claimed invention of the instant application to combine an
imaging processing system for detecting whether a person has fell or not as taught by Ng and Balavalikar Krishnamurthy to include positional relationship between objects as taught Daehee.
The suggestion/motivation for doing so would have been that there is a need in the field of fall detection to accurately determine whether a person has fallen and reduce false detections. Once with ordinary skill in the art understands the shadows and reflections can cause false detection and unexpected behavior “From the perspective of a conventional computer vision system that does not consider reflection, there is no difference between reflected object image and real object image. However, from the perspective of robot services, the reflected image is obviously not a real object. In other words, errors by reflection can cause serious performance degradation to overall robot services" as disclosed by Daehee in page 1 left hand column paragraph 1.
Therefore, it would have been obvious to combine the disclosure of Ng and Balavalikar Krishnamurthy with the Daehee disclosure to obtain the invention as specified in claim 3 as there is a reasonable expectation of success and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Regarding claim 6, the combination of Ng, Balavalikar Krishnamurthy, and Daehee teaches “The image processing apparatus according to claim 1, wherein the predetermined condition is a condition regarding a positional relationship between the first object and a second object (Daehee page 3 right hand column paragraph 2 "The proposed algorithm is composed of the following orders. First, bounding boxes of objects (person in our experiment) are detected from color images by a conventional detector. Coordinates of bounding boxes are then converted from color image coordinates into depth camera world coordinates provided with depth information. It is then used to compare positional relationship with reference surface in world coordinate") of which similarity to the first object is greater than or equal to a threshold value (Balavalikar Krishnamurthy paragraph [0038] "if the first sub-image (110) and the second sub-image (112) together have a similarity measure of 0.99, and if the similarity threshold value (124) is 0.85, then the first sub-image (110) is determined to match the second sub-image (112). In a more specific example, the first sub-image (110) and the second sub-image (112) are both heads. As a pair, the first sub-image (110) and the second sub-image (112) have a similarity measure of 0.99. Thus, in this example, a determination is made that the first sub image (110) and the second sub-image (112) are either matching heads (e.g., twins are present in the room) or that one of the first sub-image (110) and the second sub-image (112) is a sub-image of a physical person's head and the other is a sub-image of a reflection of the physical person's head").”
The proposed combination as well as the motivation for combining Ng, Balavalikar Krishnamurthy, and Daehee references presented in the rejection of claim 3, applies to claim 6. Finally the apparatus recited in claim 6 is met by Ng, Balavalikar Krishnamurthy, and Daehee.
Regarding claim 10, the combination of Ng, Balavalikar Krishnamurthy, and Daehee teaches “The image processing apparatus according to claim 9, wherein the selection unit selects the pair based on a positional relationship between the plurality of objects (Daehee page 3 right hand column paragraph 2 "The proposed algorithm is composed of the following orders. First, bounding boxes of objects (person in our experiment) are detected from color images by a conventional detector. Coordinates of bounding boxes are then converted from color image coordinates into depth camera world coordinates provided with depth information. It is then used to compare positional relationship with reference surface in world coordinate").”
The proposed combination as well as the motivation for combining Ng, Balavalikar Krishnamurthy, and Daehee references presented in the rejection of claim 3, applies to claim 10. Finally the apparatus recited in claim 10 is met by Ng, Balavalikar Krishnamurthy, and Daehee.
Claims 14 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Ng and Balavalikar Krishnamurthy in view of Wikipedia ("Point Reflection" - Version dated November 2023).
Regarding claim 14, the combination of Ng and Balavalikar Krishnamurthy teaches “The image processing apparatus according to claim 13, wherein the joint point detection unit acquires a joint point likelihood map of the object from the third learning model (Ng paragraph [0070] "pose-estimation module 106 can first identify a set of human keypoints 122 ( or simply "human keypoints 122" or "keypoints 122") for the detected person within an input video image 104, and then represent a pose of the detected person using the configuration and/or localization of the set of keypoints, wherein the set of keypoints 122 can include, but are not limited to: the eyes, the nose, the ears, the chest, the shoulders, the elbows, the wrists, the knees, the hip joints, and the ankles of the person. In some embodiments, instead of using a full set of keypoints, a simplified set of keypoints 122 can include just the head, the shoulders, the arms, and the legs of the detected person. A person of ordinary skill in the art can easily appreciate that a different pose of the detected person can be represented by a different geometric configuration of the set of keypoints 122"), and
wherein (Balavalikar Krishnamurthy paragraph [0034] "The similarity measures (122) are numbers assigned to pairs of the sub-images (106) that match one of the selected object types (116) to within a confidence threshold value (120). There are various methods to compute the similarity measures (122). One of the methods is to compute the L2-distance (Euclidean) distance between features extracted from the detections. The smaller the distance, the larger the match. Computing the Cosine similarity match is another method for computing the one or more similarity measures (122). Computing an image hash is yet another such method").”
However, the combination of Ng and Balavalikar Krishnamurthy is not relied on to teach “the similarity calculation unit inverts a joint point likelihood map of one object in two objects in a vertical direction”.
Wikipedia teaches “similarity calculation unit inverts a joint point likelihood map of one object in two objects in a vertical direction (Wikipedia page 1 paragraph 1 "In geometry, a point reflection (also called a point inversion or central inversion) is a transformation of affine space in which every point is reflected across a specific fixed point. When dealing with crystal structures and in the physical sciences the terms inversion symmetry, inversion center or centrosymmetric are more commonly used").”
It would have been obvious to a person having ordinary skill in the art before
effective filing date of the claimed invention of the instant application to combine an
imaging processing system for detecting whether a person has fell or not as taught by Ng and Balavalikar Krishnamurthy to include conversion of object coordinates into a reflected coordinate as taught Wikipedia.
The suggestion/motivation for doing so would have been that there is a need in the field of fall detection to accurately determine whether a person has fallen and reduce false detections. Once with ordinary skill in the art understands the shadows and reflections can cause false detection and unexpected behavior, and that inverting an object would provide a reflected image “In Euclidean space, a point reflection is an isometry (preserves distance). In the Euclidean lane, a point reflection is the same as a half-turn rotation (180° or n: radians); a point reflection through the object's centroid is the same as a half-turn spin " as disclosed by Wikipedia in page 1 paragraph 4.
Therefore, it would have been obvious to combine the disclosure of Ng and Balavalikar Krishnamurthy with the Wikipedia disclosure to obtain the invention as specified in claim 14 as there is a reasonable expectation of success and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Regarding claim 15, the combination of Ng, Balavalikar Krishnamurthy, and Wikipedia teaches “The image processing apparatus according to claim 13, wherein the joint point detection unit calculates joint point coordinates (Balavalikar Krishnamurthy paragraph [0021] "The sub-images (106) may be referred to as detection bounding boxes and represented by their { x,y} coordinates on a pre-determined or generated coordinate system, along the width and height of the detection bounding boxes") from a joint point likelihood map acquired from the third learning model (Ng paragraph [0070] "pose-estimation module 106 can first identify a set of human keypoints 122 ( or simply "human keypoints 122" or "keypoints 122") for the detected person within an input video image 104, and then represent a pose of the detected person using the configuration and/or localization of the set of keypoints, wherein the set of keypoints 122 can include, but are not limited to: the eyes, the nose, the ears, the chest, the shoulders, the elbows, the wrists, the knees, the hip joints, and the ankles of the person. In some embodiments, instead of using a full set of keypoints, a simplified set of keypoints 122 can include just the head, the shoulders, the arms, and the legs of the detected person. A person of ordinary skill in the art can easily appreciate that a different pose of the detected person can be represented by a different geometric configuration of the set of keypoints 122"), and
wherein the similarity calculation unit inverts a joint point position of one object in two objects in a vertical direction (Wikipedia page 1 paragraph 1 "In geometry, a point reflection (also called a point inversion or central inversion) is a transformation of affine space in which every point is reflected across a specific fixed point. When dealing with crystal structures and in the physical sciences the terms inversion symmetry, inversion center or centrosymmetric are more commonly used") to calculate a difference from a joint point position of another object and calculates similarity between the two objects based on the difference calculated (Balavalikar Krishnamurthy paragraph [0034] "The similarity measures (122) are numbers assigned to pairs of the sub-images (106) that match one of the selected object types (116) to within a confidence threshold value (120). There are various methods to compute the similarity measures (122). One of the methods is to compute the L2-distance (Euclidean) distance between features extracted from the detections. The smaller the distance, the larger the match. Computing the Cosine similarity match is another method for computing the one or more similarity measures (122). Computing an image hash is yet another such method").”
The proposed combination as well as the motivation for combining Ng, Balavalikar Krishnamurthy, and Wikipedia references presented in the rejection of claim 14, applies to claim 15. Finally the apparatus recited in claim 15 is met by Ng, Balavalikar Krishnamurthy, and Wikipedia.
Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Ng and Balavalikar Krishnamurthy in view of Yokoi et al. (US 2015/0071529 A1).
Regarding claim 16, the combination of Ng and Balavalikar Krishnamurthy teaches the apparatus of claim 11. However, the combination of Ng and Balavalikar Krishnamurthy is not relied on to teach “a human body ratio calculation unit configured to calculate a human body ratio representing a ratio of a size of each region in a human body based on the detection result by the joint point detection unit wherein the similarity calculation unit calculates the similarity based on the human body ratio”.
Yokoi teaches “The image processing apparatus according to claim 11, further comprising a human body ratio calculation unit configured to calculate a human body ratio representing a ratio of a size of each region in a human body based on the detection result by the joint point detection unit (Yokoi paragraph [0036] "extraction unit has extracted a plurality of candidate areas, a user selects an image necessary for learning (Step S13). In the case where a person is desired to be detected, the user selects only an area including the person. In the case where a vehicle is desired to be detected, for example, the user may select only an area including the vehicle. Also in the case where an object other than the person and the vehicle is desired to be detected, the same processing may be performed with the object being a target"), wherein the similarity calculation unit calculates the similarity based on the human body ratio (Yokoi paragraph [0039] "width_A represents the length of a candidate area A in a horizontal direction, height_A represents the length of the candidate area A in a vertical direction, width_B represents the length of the target object area designated in advance in a horizontal direction, and height_B represents the length of the target object area in a vertical direction. Moreover, it is also possible to calculate the degree of similarity S2 based on an aspect ratio by the following formula 2").”
It would have been obvious to a person having ordinary skill in the art before
effective filing date of the claimed invention of the instant application to combine an
imaging processing system for detecting whether a person has fell or not as taught by Ng and Balavalikar Krishnamurthy to a comparison of object size as taught Yokoi.
The suggestion/motivation for doing so would have been that “In order to efficiently collect images including a target object with different environmental conditions, in images with different environmental conditions as shown in FIG. 3, an individual condition such as (a) comparing aspect ratios of rectangles and (b) a change in position on an image is used" as disclosed by Yokoi in paragraph 35.
Therefore, it would have been obvious to combine the disclosure of Ng and Balavalikar Krishnamurthy with the Yokoi disclosure to obtain the invention as specified in claim 16 as there is a reasonable expectation of success and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Ng, Balavalikar Krishnamurthy, and Wikipedia in view of Chen et al. (US 2022/0392325 A1).
Regarding claim 17, the combination of Ng, Balavalikar Krishnamurthy, and Wikipedia teaches “The image processing apparatus according to claim 1, further comprising: a generation unit configured to estimate three-dimensional models of the plurality of persons corresponding to the plurality of objects from an image of each area of the plurality of and generate a prediction image that reproduces a reflected image of the person reflected on another object based on the estimated three-dimensional model (Wikipedia page 1 paragraph 1 "In geometry, a point reflection (also called a point inversion or central inversion) is a transformation of affine space in which every point is reflected across a specific fixed point. When dealing with crystal structures and in the physical sciences the terms inversion symmetry, inversion center or centrosymmetric are more commonly used"); and
a similarity calculation unit configured to calculate similarity between the plurality of objects based on the image of each area of the plurality of objects and the prediction image corresponding to each of the plurality of objects (Balavalikar Krishnamurthy paragraph [0034] "The similarity measures (122) are numbers assigned to pairs of the sub-images (106) that match one of the selected object types (116) to within a confidence threshold value (120). There are various methods to compute the similarity measures (122). One of the methods is to compute the L2-distance (Euclidean) distance between features extracted from the detections. The smaller the distance, the larger the match. Computing the Cosine similarity match is another method for computing the one or more similarity measures (122). Computing an image hash is yet another such method").“
However, the combination of Ng, Balavalikar Krishnamurthy, and Wikipedia is not relied on to teach “a generation unit configured to estimate three-dimensional models of the plurality of persons corresponding to the plurality of objects from an image of each area of the plurality of objects”.
Chen teaches “a generation unit configured to estimate three-dimensional models of the plurality of persons corresponding to the plurality of objects from an image of each area of the plurality of objects (Chen paragraph [0013] "The fall detection system 100 of the embodiment may include a data generator 12 configured to generate a point cloud according to the reflected radio waves (step 22). The point cloud includes a plurality of three-dimensional (3D) data representing the 3D shape of the person 10 under detection. Generally speaking, a point cloud refers a set of data points in space representing a 3D shape of an object, and each data point has 3D coordinates").”
It would have been obvious to a person having ordinary skill in the art before
effective filing date of the claimed invention of the instant application to combine an
imaging processing system for detecting whether a person has fell or not as taught by Ng, Balavalikar Krishnamurthy, and Wikipedia to include three dimensional models of objects as taught Chen.
The suggestion/motivation for doing so would have been that “A need has thus arisen to propose a novel detection scheme to overcome drawbacks of the conventional fall detection" as disclosed by Chen in paragraph 5.
Therefore, it would have been obvious to combine the disclosure of Ng, Balavalikar Krishnamurthy, Wikipedia with the Chen disclosure to obtain the invention as specified in claim 17 as there is a reasonable expectation of success and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JASPREET KAUR whose telephone number is (571)272-5534. The examiner can normally be reached Monday - Friday 7:30 am - 4:00 PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached at (571)272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JASPREET KAUR/Examiner, Art Unit 2662
/AMANDEEP SAINI/Supervisory Patent Examiner, Art Unit 2662