Prosecution Insights
Last updated: October 02, 2026
Application No. 18/739,269

MULTI-SENSOR COORDINATION METHOD, PROCESSING DEVICE, AND INFORMATION DISPLAY SYSTEM

Non-Final OA §102§103
Filed
Jun 10, 2024
Priority
Feb 21, 2024 — TW 113106215
Examiner
FATIMA, UROOJ
Art Unit
2676
Tech Center
2600 — Communications
Assignee
Industrial Technology Research Institute
OA Round
1 (Non-Final)
75%
Grant Probability
Favorable
1-2
OA Rounds
4m
Est. Remaining
75%
With Interview

Examiner Intelligence

Grants 75% — above average
75%
Career Allowance Rate
6 granted / 8 resolved
+13.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
23 currently pending
Career history
29
Total Applications
across all art units

Statute-Specific Performance

§101
13.3%
-26.7% vs TC avg
§103
60.8%
+20.8% vs TC avg
§102
7.7%
-32.3% vs TC avg
§112
14.7%
-25.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 8 resolved cases

Office Action

§102 §103
CTNF 18/739,269 CTNF 101309 DETAILED ACTION Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Priority 02-27 AIA Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. TW113106215 , filed on 02/21/2024 . Information Disclosure Statement The information disclosure statement (IDS) submitted on 06/10/2024 and 07/22/2025 has been considered by the examiner. Positive Statement regarding 35 U.S.C. 101 : Claims 1-20 are determined to be eligible under 35 U.S.C. 101. The claim 1, for example, at lines 7-10 recites “ performing an information fusion processing on the plurality of object information according to the fusion weight of each of the image sensors to obtain an object fusion information of the real scene object; and determining display content of a display according to the object fusion information.” It is given the weight of the description in the specification paragraph [0053] that determining the display content of the display based on the object fusion information causes the content to not jitter significantly thereby improving the viewing comfort of the displayed content . Because the claims recite specific, claimed steps and structural elements that produce tangible technical results, they are not directed to an abstract idea. Therefore, the claims are determined to be eligible under 35 U.S.C. 101. Claim Rejections - 35 USC § 102 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-07-aia AIA 07-07 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – 07-12-aia AIA (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. 07-15-03-aia AIA Claim s 1, 14, 19, and 20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Xu (US 12,217,545 B2) . Regarding claim 1, Xu discloses a multi-sensor coordination method, adapted to an information display system comprising a display and a plurality of image sensors (Column 3 [lines 18-25] “data processing environment 100 having one or more servers 102 communicatively coupled to one or more client devices 104, in accordance with some embodiments. The one or more client devices 104 may be, for example, desktop computers 104A, tablet computers 104B, mobile phones 104C, augmented reality (AR) glasses 150, or intelligent, multi-sensing, network-connected home devices (e.g., a camera).” , and comprising: obtaining a plurality of images captured by the image sensors (Column 10 [lines 51-61] “The first electronic device has a first camera 504, and the second electronic device has a second camera 506. Examples of these electronic devices having cameras include…The first camera 504 captures one or more first images including a first perspective (i.e., a first viewing angle) of a hand. The second camera 506 captures one or more second images including a second perspective (i.e., a second viewing angle) of the hand”) ; performing an object recognition processing on the images (Column 16 [lines 49-51] "an image backbone network 1006 is applied to identify a plurality of visual features from the input image 1004”) or a stitching image of the images to obtain a plurality of object information of a real scene object (Column 16 [lines 51-53] “…a hand detection module 1008 processes the plurality of visual features to determine a location of a hand in the input image, e.g., using a CNN.") ; determining a fusion weight of each of the image sensors (Column 18 [lines 24-34] "the final hand gesture is determined from the different hand gestures 1104 associated with the different cameras 1106 according to a weighting scheme. Each camera 1106 is assigned with a weight 1108 that accounts for an impact of quality of images captured by the respective camera 1106. For example, a camera 1106B that is closest to the hand has a greatest weight among the cameras 1106. The weights for the cameras 1106 are optionally normalized into a range of [0, 1]. A weighted sum of confidence scores is calculated for each type of hand gesture 1104 determined from images captured by the cameras) ; performing an information fusion processing on the plurality of object information according to the fusion weight of each of the image sensors to obtain an object fusion information of the real scene object (Column 9 [lines 22-28] “a weight w′ associated with each link 412 is applied to the node output. Likewise, the one or more node inputs are combined based on corresponding weights w.sub.1, w.sub.2, w.sub.3, and w.sub.4 according to the propagation function. In an example, the propagation function is a product of a non-linear activation function and a linear weighted combination of the one or more node inputs. “; Column 18 [lines 24-50] "the final hand gesture is determined from the different hand gestures 1104 associated with the different cameras 1106 according to a weighting scheme. Each camera 1106 is assigned with a weight 1108 that accounts for an impact of quality of images captured by the respective camera 1106. For example, a camera 1106B that is closest to the hand has a greatest weight among the cameras 1106. The weights for the cameras 1106 are optionally normalized into a range of [0, 1]. A weighted sum of confidence scores is calculated for each type of hand gesture 1104 determined from images captured by the cameras… the second weighted sum of the hand gesture 900 is greater than the first weighted sum of the hand gesture 800, and the hand gesture 900 is determined as the final hand gesture.") ; and determining display content of a display according to the object fusion information (Column 13 [lines 11-16] "after the first and second hand gestures are generated, the first and second hand gestures are provided (706) to the TV device 510 by the HMD 150, client device 104, or both of them. The TV device 510 generates the final hand gesture from the first and second hand gestures and is controlled by the final hand gesture.]”) . Regarding claim 14, which claim 1 is incorporated, Xu discloses wherein the image sensors comprise a first image sensor (first camera in (Column 10 [lines 51-61] equates to first image sensor) and a second image sensor (second camera in (Column 10 [lines 51-61] equates to second image sensor) (Column 10 [lines 51-61] “The first electronic device has a first camera 504, and the second electronic device has a second camera 506. Examples of these electronic devices having cameras include…The first camera 504 captures one or more first images including a first perspective (i.e., a first viewing angle) of a hand. The second camera 506 captures one or more second images including a second perspective (i.e., a second viewing angle) of the hand”) , and the step of determining the fusion weight of each of the plurality of image sensors comprises (Column 18 [lines 24-34] "the final hand gesture is determined from the different hand gestures 1104 associated with the different cameras 1106 according to a weighting scheme. Each camera 1106 is assigned with a weight 1108 that accounts for an impact of quality of images captured by the respective camera 1106. For example, a camera 1106B that is closest to the hand has a greatest weight among the cameras 1106. The weights for the cameras 1106 are optionally normalized into a range of [0, 1]. A weighted sum of confidence scores is calculated for each type of hand gesture 1104 determined from images captured by the cameras”) : obtaining a first object detection rate of the first image sensor within a preset period; obtaining a second object detection rate of the second image sensor within the preset period (Column 16 [lines 20-28] "The hand gesture 800 or 900 is formed and changes at a low rate, e.g., does not change within 0.5-1 second. Given such a low changing rate of the hand gesture 800 or 900, two images are captured concurrently by the first and second cameras 504 and 506 when the two images are captured within a time window having a predefined width, e.g., within 1 second. It can be assumed that these two images are captured concurrently and correspond to the same hand gesture 800 or 900.") ; and determining the fusion weight of the first image sensor and the fusion weight of the second image sensor according to the first object detection rate and the second object detection rate (Column 18 [lines 31 - 50] "A weighted sum of confidence scores is calculated for each type of hand gesture 1104 determined from images captured by the cameras 1106, and the hand gesture 1104 having the largest weighted sum of confidence scores is determined to be the final hand gesture. Referring to FIG. 11, five cameras 1106A, 1106B, 1106C, 1106D, and 1106E are used to capture images associated with hand gestures in the scene, and each camera 1106 is associated with a respective weight 1108. Two types of hand gestures (hand gesture 800 and hand gesture 900) are determined from the images captured by these five cameras. A first weighted sum of confidence scores is equal to 0.445 for the hand gesture 800, and a second weighted sum of confidence scores is equal to 0.8 for the hand gesture 900. Although the gesture 800 has been determined from images captured by three cameras, the second weighted sum of the hand gesture 900 is greater than the first weighted sum of the hand gesture 800, and the hand gesture 900 is determined as the final hand gesture.") . Regarding claim 19 , Xu discloses an information display system, comprising: a display (Column 5 [lines 12-18] “a pair of AR glasses 150 (also called a head-mounted display (HMD)) that can be communicatively coupled in a data processing environment 100, in accordance with some embodiments. The AR glasses 150 includes a camera, a microphone, a speaker and a display. The camera and microphone are configured to capture video and audio data from a scene of the AR glasses 150.”) ; a plurality of image sensors (Column 5 [lines 54-60] “the client device 104 includes one or more cameras, scanners, or photo sensor units for capturing images, for example, of graphic serial codes printed on the electronic devices. The data processing system 200 also includes one or more output devices 212 that enable presentation of user interfaces and display content”) ; and a processing apparatus, connecting to the display and the image sensors (Column 5 [lines 12-20] “a pair of AR glasses 150 (also called a head-mounted display (HMD)) that can be communicatively coupled in a data processing environment 100, in accordance with some embodiments. The AR glasses 150 includes a camera, a microphone, a speaker and a display. The camera and microphone are configured to capture video and audio data from a scene of the AR glasses 150.”) , and configured to: obtain a plurality of images captured by the image sensors (Column 10 [lines 51-61] “The first electronic device has a first camera 504, and the second electronic device has a second camera 506. Examples of these electronic devices having cameras include…The first camera 504 captures one or more first images including a first perspective (i.e., a first viewing angle) of a hand. The second camera 506 captures one or more second images including a second perspective (i.e., a second viewing angle) of the hand”) ; perform an object recognition processing on the images (Column 16 [lines 49-51] "an image backbone network 1006 is applied to identify a plurality of visual features from the input image 1004” or a stitching image of the images to obtain a plurality of object information of a real scene object (Column 16 [lines 51-53] “…a hand detection module 1008 processes the plurality of visual features to determine a location of a hand in the input image, e.g., using a CNN."); determine a fusion weight of each of the image sensors (Column 18 [lines 24-34] "the final hand gesture is determined from the different hand gestures 1104 associated with the different cameras 1106 according to a weighting scheme. Each camera 1106 is assigned with a weight 1108 that accounts for an impact of quality of images captured by the respective camera 1106. For example, a camera 1106B that is closest to the hand has a greatest weight among the cameras 1106. The weights for the cameras 1106 are optionally normalized into a range of [0, 1]. A weighted sum of confidence scores is calculated for each type of hand gesture 1104 determined from images captured by the cameras ”); perform an information fusion processing on the plurality of object information according to the fusion weight of each of the image sensors to obtain an object fusion information of the real scene object (Column 9 [lines 22-28] “a weight w′ associated with each link 412 is applied to the node output. Likewise, the one or more node inputs are combined based on corresponding weights w.sub.1, w.sub.2, w.sub.3, and w.sub.4 according to the propagation function. In an example, the propagation function is a product of a non-linear activation function and a linear weighted combination of the one or more node inputs. “; Column 18 [lines 24-50] "the final hand gesture is determined from the different hand gestures 1104 associated with the different cameras 1106 according to a weighting scheme. Each camera 1106 is assigned with a weight 1108 that accounts for an impact of quality of images captured by the respective camera 1106. For example, a camera 1106B that is closest to the hand has a greatest weight among the cameras 1106. The weights for the cameras 1106 are optionally normalized into a range of [0, 1]. A weighted sum of confidence scores is calculated for each type of hand gesture 1104 determined from images captured by the cameras… the second weighted sum of the hand gesture 900 is greater than the first weighted sum of the hand gesture 800, and the hand gesture 900 is determined as the final hand gesture.") ; and determine display content of a display according to the object fusion information (Column 13 [lines 11-16] "after the first and second hand gestures are generated, the first and second hand gestures are provided (706) to the TV device 510 by the HMD 150, client device 104, or both of them. The TV device 510 generates the final hand gesture from the first and second hand gestures and is controlled by the final hand gesture.]”) . Regarding claim 20, Xu discloses a processing apparatus, comprising: a memory configured to store data (Column 5 [lines 38-45] “The data processing system 200 includes a server 102, a client device 104 (e.g., AR glasses 150 in FIG. 1B), a storage 106, or a combination thereof. The data processing system 200, typically, includes one or more processing units (CPUs) 202, one or more network interfaces 204, memory 206, and one or more communication buses 208 for interconnecting these components (sometimes called a chipset).”) ; and a processor connected to the memory (Column 6 [lines 4-6] “Memory 206, optionally, includes one or more storage devices remotely located from one or more processing units 202.”) and configured to: obtain a plurality of images captured by a plurality of image sensors (Column 10 [lines 51-61] “The first electronic device has a first camera 504, and the second electronic device has a second camera 506. Examples of these electronic devices having cameras include…The first camera 504 captures one or more first images including a first perspective (i.e., a first viewing angle) of a hand. The second camera 506 captures one or more second images including a second perspective (i.e., a second viewing angle) of the hand”) ; perform an object recognition processing on the images (Column 16 [lines 49-51] "an image backbone network 1006 is applied to identify a plurality of visual features from the input image 1004”) or a stitching image of the images to obtain a plurality of object information of a real scene object (Column 16 [lines 51-53] “…a hand detection module 1008 processes the plurality of visual features to determine a location of a hand in the input image, e.g., using a CNN.") ; determine a fusion weight of each of the image sensors (Column 18 [lines 24-34] "the final hand gesture is determined from the different hand gestures 1104 associated with the different cameras 1106 according to a weighting scheme. Each camera 1106 is assigned with a weight 1108 that accounts for an impact of quality of images captured by the respective camera 1106. For example, a camera 1106B that is closest to the hand has a greatest weight among the cameras 1106. The weights for the cameras 1106 are optionally normalized into a range of [0, 1]. A weighted sum of confidence scores is calculated for each type of hand gesture 1104 determined from images captured by the cameras) ; perform an information fusion processing on the plurality of object information according to the fusion weight of each of the image sensors to obtain an object fusion information of the real scene object (Column 9 [lines 22-28] “a weight w′ associated with each link 412 is applied to the node output. Likewise, the one or more node inputs are combined based on corresponding weights w.sub.1, w.sub.2, w.sub.3, and w.sub.4 according to the propagation function. In an example, the propagation function is a product of a non-linear activation function and a linear weighted combination of the one or more node inputs. “; Column 18 [lines 24-50] "the final hand gesture is determined from the different hand gestures 1104 associated with the different cameras 1106 according to a weighting scheme. Each camera 1106 is assigned with a weight 1108 that accounts for an impact of quality of images captured by the respective camera 1106. For example, a camera 1106B that is closest to the hand has a greatest weight among the cameras 1106. The weights for the cameras 1106 are optionally normalized into a range of [0, 1]. A weighted sum of confidence scores is calculated for each type of hand gesture 1104 determined from images captured by the cameras… the second weighted sum of the hand gesture 900 is greater than the first weighted sum of the hand gesture 800, and the hand gesture 900 is determined as the final hand gesture.") ; and determine display content of a display according to the object fusion information (Column 13 [lines 11-16] "after the first and second hand gestures are generated, the first and second hand gestures are provided (706) to the TV device 510 by the HMD 150, client device 104, or both of them. The TV device 510 generates the final hand gesture from the first and second hand gestures and is controlled by the final hand gesture.]”) . Claim Rejections - 35 USC § 103 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-21-aia AIA Claim s 2 and 3 are rejected under 35 U.S.C. 103 as being unpatentable over Xu (US 12,217,545 B2) in view of Kalayeh (US 8,600,193) . Regarding claim 2, which claim 1 is incorporated, Xu discloses obtaining a capture angle of each of the image sensors (Column 10 [lines 57-62] "The first camera 504 captures one or more first images including a first perspective (i.e., a first viewing angle) of a hand. The second camera 506 captures one or more second images including a second perspective (i.e., a second viewing angle) of the hand, concurrently with the first camera 504 capturing the one or more first images"; Column 17 [lines 28-34] “a table 1100 showing confidence scores 1102 of hand gestures 1104 identified from images captured by different cameras 1106 from different perspectives (e.g., different viewing angles)…The cameras 1106 are disposed in different locations of a scene where a hand is located and configured to capture the images including a hand with the different perspectives.”) . However, Xu fails to teach performing an image stitching processing according to the capture angle of each of the image sensors to generate the stitching image. Kalayeh teaches performing an image stitching processing [according to the capture angle] of each of the image sensors to generate the stitching image (Column 7 [lines 40-45] "Image stitching preprocessing includes a selection of windows from multiple images (typically, but not necessarily two sequentially obtained images). Later, features can be selected from each of the windows for further analysis. The windows are selected from an overlapping region of each of two images to be stitched together.") . Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Xu’s reference to include performing an image stitching processing [according to the capture angle] of each of the image sensors to generate the stitching image taught by Kalayeh’s reference. The motivation for doing so would have been to seamlessly overlap two images that have an overlap region as suggested by Kalayeh (see Kalayeh, Column 5 [lines 35-40]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Kalayeh with Xu to obtain the invention specified in claim 2. Regarding claim 3, which claim 2 is incorporated, Xu discloses wherein the image sensors comprise a first image sensor (first camera in (Column 10 [lines 51-61] equates to first image sensor) and a second image sensor (second camera in (Column 10 [lines 51-61] equates to second image sensor) , the images comprise a first image of the first image sensor and a second image of the second image sensor, and the capture angle of each of the image sensors (cameras 1106 from different perspectives (e.g., different viewing angles) in Column 17 [lines 28-34] equates to angle of each of the image sensors) (Column 10 [lines 51-61] “The first electronic device has a first camera 504, and the second electronic device has a second camera 506. Examples of these electronic devices having cameras include…The first camera 504 captures one or more first images including a first perspective (i.e., a first viewing angle) of a hand. The second camera 506 captures one or more second images including a second perspective (i.e., a second viewing angle) of the hand”; Column 17 [lines 28-34] “a table 1100 showing confidence scores 1102 of hand gestures 1104 identified from images captured by different cameras 1106 from different perspectives (e.g., different viewing angles)…The cameras 1106 are disposed in different locations of a scene where a hand is located and configured to capture the images including a hand with the different perspectives.”) . However, Xu fails to teach the step of performing the image stitching processing [according to the capture angle] of each of the image sensors to generate the stitching image comprises: performing an image geometric correction processing on the first image or the second image [according to the capture angle] of the first image sensor and the [capture angle] of the second image sensor to generate at least one corrected image; and performed the image stitching processing according to the at least one corrected image to generate the stitching image. Kalayeh’s teaches the step of performing the image stitching processing [according to the capture angle] of each of the image sensors to generate the stitching image comprises (Column 7 [lines 40-45] "Image stitching preprocessing includes a selection of windows from multiple images (typically, but not necessarily two sequentially obtained images). Later, features can be selected from each of the windows for further analysis. The windows are selected from an overlapping region of each of two images to be stitched together.") : performing an image geometric correction processing on the first image or the second image [according to the capture angle] of the first image sensor and the [capture angle of the ] second image sensor to generate at least one corrected image (Column 6 [lines 60-64] " The image stitching algorithm according to this embodiment can include the following steps: Image Motion Preprocessing, Image (fix figures) Motion Correction, Image Stitching Preprocessing"; Column 7 [lines 1-14] "This image motion correction uses the two selected sub-images for estimation of a possible translation and rotation of the second image relative to the first one and then can compensate or correct the second image motion relative to the first one. The second motion corrected sub-image therefore has zero motion components relative to the first image. In the image stitching preprocessing step, the size and shape of the overlapped region in the first image, second image, and the template are determined. ") ; and performed the image stitching processing according to the at least one corrected image to generate the stitching image (Column 7 [lines 14-24] "The step of feature extraction extracts and selects multi-resolution radon gray level, gradient, or range based features from the overlap region. The template matching algorithm can use 2D and 1D normalized cross correlation, normalized city block distance, and/or city block distance to estimate the stitching location. The blending algorithm combines the overlapped region to make the stitching seam substantially invisible. Finally, the performance evaluation algorithm quantitatively measures the stitching error or correlation in the joined region.") . Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Xu’s reference to include the step of performing the image stitching processing [according to the capture angle] of each of the image sensors to generate the stitching image comprises: performing an image geometric correction processing on the first image or the second image [according to the capture angle] of the first image sensor and the [capture angle of the] second image sensor to generate at least one corrected image; and performed the image stitching processing according to the at least one corrected image to generate the stitching image taught by Kalayeh’s reference. The motivation for doing so would have been to correct the second image relative to the first one as suggested by Kalayeh (see Kalayeh, Column 7 [lines 8-9]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Kalayeh with Xu to obtain the invention specified in claim 3 . 07-21-aia AIA Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Xu (US 12,217,545 B2) in view of Dong et al. (CN 117,495,697 A) (hereinafter, “Dong”) . Regarding claim 4, which claim 1 is incorporated, Xu discloses wherein the image sensors comprise a first image sensor and a second image sensor, the images comprise a first image of the first image sensor and a second image of the second image sensor (Column 10 [lines 51-61] “The first electronic device has a first camera 504, and the second electronic device has a second camera 506. Examples of these electronic devices having cameras include…The first camera 504 captures one or more first images including a first perspective (i.e., a first viewing angle) of a hand. The second camera 506 captures one or more second images including a second perspective (i.e., a second viewing angle) of the hand”) , wherein the step of performing the information fusion processing on the plurality of object information according to the fusion weight of each of the image sensors to obtain the object fusion information of the real scene object (Column 9 [lines 22-28] “a weight w′ associated with each link 412 is applied to the node output. Likewise, the one or more node inputs are combined based on corresponding weights w.sub.1, w.sub.2, w.sub.3, and w.sub.4 according to the propagation function. In an example, the propagation function is a product of a non-linear activation function and a linear weighted combination of the one or more node inputs. “; Column 18 [lines 24-50] "the final hand gesture is determined from the different hand gestures 1104 associated with the different cameras 1106 according to a weighting scheme. Each camera 1106 is assigned with a weight 1108 that accounts for an impact of quality of images captured by the respective camera 1106. For example, a camera 1106B that is closest to the hand has a greatest weight among the cameras 1106. The weights for the cameras 1106 are optionally normalized into a range of [0, 1]. A weighted sum of confidence scores is calculated for each type of hand gesture 1104 determined from images captured by the cameras… the second weighted sum of the hand gesture 900 is greater than the first weighted sum of the hand gesture 800, and the hand gesture 900 is determined as the final hand gesture.") : However, Xu fails to teach the plurality of object information of the real scene object comprise a first spatial coordinate associated with the first image sensor and a second spatial coordinate associated with the second image sensor, and performing a weighted operation on the first spatial coordinate and the second spatial coordinate according to the fusion weight of the first image sensor and the fusion weight of the second image sensor to obtain a fusion spatial coordinate of the real scene object. Dong teaches the plurality of object information of the real scene object comprise a first spatial coordinate (L10 (x,y,t) in Paragraph [0031] equates to first spatial coordinates) associated with the first image sensor and a second spatial coordinate (L20 (x,y,t) in Paragraph [0031] equates to second spatial coordinates) associated with the second image sensor (Paragraph [031] “Two images can be obtained, namely the first image corresponding to the first image information and the second image corresponding to the second image information, so that the format size of the first and second images after necessary alignment, scaling, etc. Completely consistent, assuming that the current frame to be processed is the t-th frame, read the first and second image brightness signals L10 (x,y,t) and L20 (x,y,t) whose spatial coordinates are (x,y) respectively. As the first image information and the second image information;”) , and performing a weighted operation on the first spatial coordinate and the second spatial coordinate according to the fusion weight of the first image sensor and the fusion weight of the second image sensor to obtain a fusion spatial coordinate of the real scene object (Paragraphs [033-034] “According to the first fusion weight and the second fusion weight, obtain a spatio-temporal domain weighting number; according to the spatio-temporal domain weighting number, perform spatio-temporal domain filtering on the first image information and the second image information. The output result of the spatiotemporal domain filtering is used as the calculation basis for image fusion calculation.”. Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Xu’s reference to include the plurality of object information of the real scene object comprise a first spatial coordinate associated with the first image sensor and a second spatial coordinate associated with the second image sensor, and performing a weighted operation on the first spatial coordinate and the second spatial coordinate according to the fusion weight of the first image sensor and the fusion weight of the second image sensor to obtain a fusion spatial coordinate of the real scene object taught by Dong’s reference. The motivation for doing so would have been to obtain comprehensive fused motion information as suggested by Dong (see Dong, Paragraph [088]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Dong with Xu to obtain the invention specified in claim 4 . 07-21-aia AIA Claim s 5-7 are rejected under 35 U.S.C. 103 as being unpatentable over Xu (US 12,217,545 B2) in view of Kalayeh (US 8600193), and further in view of Dong et al. (CN 117,495,697 A) (hereinafter, “Dong”) . Regarding claim 5, which claim 1 is incorporated, Xu discloses wherein the image sensors comprise a first image sensor and a second image sensor, the images comprise a first image of the first image sensor and a second image of the second image sensor (Column 10 [lines 51-61] “The first electronic device has a first camera 504, and the second electronic device has a second camera 506. Examples of these electronic devices having cameras include…The first camera 504 captures one or more first images including a first perspective (i.e., a first viewing angle) of a hand. The second camera 506 captures one or more second images including a second perspective (i.e., a second viewing angle) of the hand”) , and the step of performing the object recognition (Column 16 [lines 49-51] "an image backbone network 1006 is applied to identify a plurality of visual features from the input image 1004”) processing on the images or the stitching image of the images to obtain the plurality of object information of the real scene object comprises (Column 16 [lines 51-53] “…a hand detection module 1008 processes the plurality of visual features to determine a location of a hand in the input image, e.g., using a CNN.") : performing the object recognition processing on the first image of the first image sensor and the second image of the second image sensor respectively (Column 16 [lines 49-51] "an image backbone network 1006 is applied to identify a plurality of visual features from the input image 1004”) However, Xu fails to teach to obtain a first region of interest (ROI) and a second region of interest; and performing a coordinate conversion on the first ROI of the first image and the second ROI of the second image respectively to obtain a first spatial coordinate of the first ROI and a second spatial coordinate of the second ROI. Kalayeh teaches to obtain a first region of interest (ROI) and a second region of interest (Column 4 [lines 14-17) “the autostitching algorithm is configured to read the first digital image and the second digital image, determine a first Region of Interest (ROI) in the first digital image and a second ROI in the second digital image”). Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Xu’s reference to include to obtain a first region of interest (ROI) and a second region of interest taught by Kalayeh’s reference. The motivation for doing so would have been to later extract features from the windows containing the objects of interest as suggested by Kalayeh (see Kalayeh, Column 7 [lines 47-50]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. However, Xu and Kalayeh fail to teach performing a coordinate conversion on the [first ROI of the first image and the second ROI of the second image] respectively to obtain a first spatial coordinate of the [first ROI] and a second spatial coordinate of the [second ROI]. Dong teaches performing a coordinate conversion on the [first ROI of the first image and the second ROI of the second image] respectively to obtain a first spatial coordinate of the [first ROI] and a second spatial coordinate of the [second ROI] (Paragraph [078] “Read the first and second image brightness signals with spatial coordinates (x, y) respectively .”) Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Xu in view of Kalayeh to include performing a coordinate conversion on the [first ROI of the first image and the second ROI of the second image] respectively to obtain a first spatial coordinate of the [first ROI] and a second spatial coordinate of the [second ROI] taught by Dong’s reference. The motivation for doing so would have been to obtain image information after necessary alignment to use as a basis for fusion weight calculation as suggested by Dong (see Dong, Paragraph [031] and Paragraph [034]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Kalayeh and Dong with Xu to obtain the invention specified in claim 5. Regarding claim 6, which claim 5 is incorporated, Xu discloses wherein the step of performing the object recognition processing on the images (Column 16 [lines 49-51] "an image backbone network 1006 is applied to identify a plurality of visual features from the input image 1004”) or the stitching image of the images to obtain the plurality of object information of the real scene object (Column 16 [lines 51-53] “…a hand detection module 1008 processes the plurality of visual features to determine a location of a hand in the input image, e.g., using a CNN.") further comprises: an object classification category of the first [ROI] , and an object classification category of the second [ROI] (Column 18 [lines 37-43] “five cameras 1106A, 1106B, 1106C, 1106D, and 1106E are used to capture images associated with hand gestures in the scene, and each camera 1106 is associated with a respective weight 1108. Two types of hand gestures (hand gesture 800 and hand gesture 900) are determined from the images captured by these five cameras.”). However, Xu fails to teach obtaining an area intersection information of the first ROI and the second ROI according to the first spatial coordinate of the first ROI and the second spatial coordinate of the second ROI; and determining that the first spatial coordinate and the second spatial coordinate correspond to the real scene object based on the area intersection information. Kalayeh teaches obtaining an area intersection information of the first ROI and the second ROI [according to the first spatial coordinate] of the first ROI and the [second spatial coordinate] of the second ROI (Column 7 [lines 1-15] "selects the size of two sub-images based on the size of suggested or required minimum and maximum overlapped regions of two typically sequentially acquired images, whether vertically or horizontally displaced or a combination thereof. This image motion correction uses the two selected sub-images for estimation of a possible translation and rotation of the second image relative to the first one and then can compensate or correct the second image motion relative to the first one. The second motion corrected sub-image therefore has zero motion components relative to the first image. In the image stitching preprocessing step, the size and shape of the overlapped region in the first image, second image, and the template are determined.") . Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Xu’s reference to include to obtaining an area intersection information of the first ROI and the second ROI [according to the first spatial coordinate] of the first ROI and the [second spatial coordinate] of the second ROI taught by Kalayeh’s reference. The motivation for doing so would have been to later extract features from the windows containing the objects of interest and seamlessly combine the overlapping regions as suggested by Kalayeh (see Kalayeh, Column 7 [lines 20-22] and [lines 47-50]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. However, Xu and Kalayeh fail to teach determining that the first spatial coordinate and the second spatial coordinate correspond to the real scene object based on [the area intersection information]. Dong teaches determining that the first spatial coordinate and the second spatial coordinate correspond to the real scene object based on [the area intersection information] (Paragraph [078] “the first image information and the second image information originate from the same frame in multiple images of the same shooting target taken by different sensors in the same time period…Two images can be obtained, namely the first image corresponding to the first image information and the second image corresponding to the second image information…assuming that the current frame to be processed is the t-th frame, read the first and second image brightness signals L10 (x,y,t) and L20 (x,y,t) whose spatial coordinates are (x,y) ”) . Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Xu in view of Kalayeh to include determining that the first spatial coordinate and the second spatial coordinate correspond to the real scene object based on [the area intersection information] taught by Dong’s reference. The motivation for doing so would have been to obtain image information after necessary alignment to use as a basis for fusion weight calculation as suggested by Dong (see Dong, Paragraph [031] and Paragraph [034]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Kalayeh and Dong with Xu to obtain the invention specified in claim 6. Regarding claim 7, which claim 6 is incorporated, Xu teaches wherein the object recognition processing is performed by using a convolutional neural network model, and the step of performing the object recognition processing on the images or the stitching image of the images to obtain the plurality of object information of the real scene object comprises (Column 9 line 66 continued to column 10 line 3 “The pre-processed video or image data is abstracted by each layer of the CNN to a respective feature map. By these means, video and image data can be processed by the CNN for video and image recognition, classification, analysis, imprinting, or synthesis.”) : performing a weighted operation on a plurality of classification confidence values of the first [ROI] (confidence scores is calculated for each type of hand gesture in Column 18 [lines 24-50] classification confidence values of the first [ROI]) and a plurality of classification confidence values of the second [ROI] (confidence scores is calculated for each type of hand gesture in Column 18 [lines 24-50] classification confidence values of the second [ROI]) according to the fusion weight of the first image sensor and the fusion weight of the second image sensor, to obtain a plurality of fusion classification confidence values (Column 18 [lines 24-50] "In some embodiments, the final hand gesture is determined from the different hand gestures 1104 associated with the different cameras 1106 according to a weighting scheme. Each camera 1106 is assigned with a weight 1108 that accounts for an impact of quality of images captured by the respective camera 1106. For example, a camera 1106B that is closest to the hand has a greatest weight among the cameras 1106. The weights for the cameras 1106 are optionally normalized into a range of [0, 1]. A weighted sum of confidence scores is calculated for each type of hand gesture 1104 determined from images captured by the cameras 1106, and the hand gesture 1104 having the largest weighted sum of confidence scores is determined to be the final hand gesture.”) ; and determining a target classification category of the real scene object according to the highest one among the fusion classification confidence values (the hand gesture 1104 having the largest weighted sum of confidence scores is determined to be the final hand gesture. Referring to FIG. 11, five cameras 1106A, 1106B, 1106C, 1106D, and 1106E are used to capture images associated with hand gestures in the scene, and each camera 1106 is associated with a respective weight 1108. Two types of hand gestures (hand gesture 800 and hand gesture 900) are determined from the images captured by these five cameras. A first weighted sum of confidence scores is equal to 0.445 for the hand gesture 800, and a second weighted sum of confidence scores is equal to 0.8 for the hand gesture 900. Although the gesture 800 has been determined from images captured by three cameras, the second weighted sum of the hand gesture 900 is greater than the first weighted sum of the hand gesture 800, and the hand gesture 900 is determined as the final hand gesture.") . However, Xu fails to teach first and second ROI. Kalayeh teaches a first and second ROI (Column 4 [lines 14-17) “the autostitching algorithm is configured to read the first digital image and the second digital image, determine a first Region of Interest (ROI) in the first digital image and a second ROI in the second digital image”). Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Xu’s reference to include to a first and second ROI taught by Kalayeh’s reference. The motivation for doing so would have been to later extract features from the windows containing the objects of interest as suggested by Kalayeh (see Kalayeh, Column 7 [lines 47-50]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Kalayeh with Xu to obtain the invention specified in claim 7 . 07-21-aia AIA Claim s 8 and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Xu (US 12,217,545 B2) in view of Lin (US 2022/0165031 A1) . Regarding claim 8, which claim 1 is incorporated, Xi discloses wherein the image sensors comprise a first image sensor and a second image sensor, the images comprise a first image of the first image sensor and a second image of the second image sensor (Column 10 [lines 51-61] “The first electronic device has a first camera 504, and the second electronic device has a second camera 506. Examples of these electronic devices having cameras include…The first camera 504 captures one or more first images including a first perspective (i.e., a first viewing angle) of a hand. The second camera 506 captures one or more second images including a second perspective (i.e., a second viewing angle) of the hand”) , wherein the step of performing the information fusion processing on the plurality of object information according to the fusion weight of each of the image sensors to obtain the object fusion information of the real scene object (Column 18 [lines 24-34] "the final hand gesture is determined from the different hand gestures 1104 associated with the different cameras 1106 according to a weighting scheme. Each camera 1106 is assigned with a weight 1108 that accounts for an impact of quality of images captured by the respective camera 1106. For example, a camera 1106B that is closest to the hand has a greatest weight among the cameras 1106. The weights for the cameras 1106 are optionally normalized into a range of [0, 1]. A weighted sum of confidence scores is calculated for each type of hand gesture 1104 determined from images captured by the cameras). However Xu fails to teach the real scene object is a face object, and the plurality of object information of the real scene object comprise a first facial orientation angle of a first face ROI sensed by the first image sensor and a second facial orientation angle of a second face ROI sensed by the second image sensor, and performing a weighted operation on the first facial orientation angle of the first face ROI and the second facial orientation angle of the second face ROI according to the fusion weight of the first image sensor and the fusion weight of the second image sensor, to obtain a fusion facial orientation angle of the face object. Lin teaches the real scene object is a face object, and the plurality of object information of the real scene object comprise a first facial orientation angle of a first face ROI (front face image in Paragraph [0081] equates to first facial orientation angle of a first face ROI) sensed by the first image sensor and a second facial orientation angle of a second face ROI (right and left face image in Paragraph [0081] equates to second facial orientation angle of a second face ROI) sensed by the second image sensor (Paragraph [0041] “The terminal prompts a user to photograph a face from different angles, so as to obtain at least two initial images recorded with the face photographed from different angles.”] , and performing a weighted operation on the first facial orientation angle of the first face ROI and the second facial orientation angle of the second face ROI according to the fusion weight of the first image sensor and the fusion weight of the second image sensor, to obtain a fusion facial orientation angle of the face object (Paragraph [0081] “during point fusion, different weights need to be assigned to points from different images. In some embodiments, weight values may be assigned to first points according to at least one of a shooting angle, an image noise value, or a normal direction of the initial image in which the first point is located…a weight assigned to the first point of the front face image is 60%, and weights assigned to the first points of the left face image and the right face image are respectively 20%.”; Paragraph [0089] “effective point clouds in each initial image are projected to a reference coordinate system (a camera coordinate system in which the first frame is located), and then weighted fusion is performed on points in an overlapping area to obtain the second point cloud information more precisely, so that a more precise 3D model is constructed.”). Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Xu’s reference to include the real scene object is a face object, and the plurality of object information of the real scene object comprise a first facial orientation angle of a first face ROI sensed by the first image sensor and a second facial orientation angle of a second face ROI sensed by the second image sensor, and performing a weighted operation on the first facial orientation angle of the first face ROI and the second facial orientation angle of the second face ROI according to the fusion weight of the first image sensor and the fusion weight of the second image sensor, to obtain a fusion facial orientation angle of the face object taught by Lin’s reference. The motivation for doing so would have been to more accurately fuse images based on their different weights as suggested by Lin (see Lin, Paragraph [0083]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Lin with Xu to obtain the invention specified in claim 8. Regarding claim 9 , which claim 1 is incorporated, Xu discloses wherein the image sensors comprise a first image sensor and a second image sensor, the images comprise a first image of the first image sensor and a second image of the second image sensor ( Column 10 [lines 51-61] “The first electronic device has a first camera 504, and the second electronic device has a second camera 506. Examples of these electronic devices having cameras include…The first camera 504 captures one or more first images including a first perspective (i.e., a first viewing angle) of a hand. The second camera 506 captures one or more second images including a second perspective (i.e., a second viewing angle) of the hand”) , and the step of determining the fusion weight of each of the image sensors comprises (Column 18 [lines 24-34] "the final hand gesture is determined from the different hand gestures 1104 associated with the different cameras 1106 according to a weighting scheme. Each camera 1106 is assigned with a weight 1108 that accounts for an impact of quality of images captured by the respective camera 1106. For example, a camera 1106B that is closest to the hand has a greatest weight among the cameras 1106. The weights for the cameras 1106 are optionally normalized into a range of [0, 1]. A weighted sum of confidence scores is calculated for each type of hand gesture 1104 determined from images captured by the cameras) . However, Xu fails to teach obtaining a first image object position of the first ROI in the first image and a second image object position of the second ROI in the second image; determining the fusion weight of the first image sensor according to the first image object position; and determining the fusion weight of the second image sensor according to the second image object position. Lin teaches obtaining a first image object position of the first ROI in the first image and a second image object position of the second ROI in the second image (Paragraph [0078] " in the front face image, there is a 3D point A1 at a nose tip of a user, and a coordinate value is (x1, y1, z1); in the left face image, there is a 3D point B1 at the nose tip of the user, and a coordinate value is (x2, y2, z2); and in the right face image, there is a 3D point C1 at the nose tip of the user, and a coordinate value is (x3, y3, z3). After the left face image and the right face image are moved according to step 306 above, three 3D points A1, B1 and C1 overlap with each other. In this case, fusion on A1, B1, and C1 is performed to obtain a 3D point D1 (x4, y4, z4), which is a 3D point at the nose tip of the user in the second point cloud information."); determining the fusion weight of the first image sensor according to the first image object position ( Paragraph [0081] “a front face image, a left face image, and a right face image. Each of the three images includes a point used for representing a nose tip of a user (for example, the nose tip is a first point). In this case, during point fusion, different weights need to be assigned to points from different images.”) ; and determining the fusion weight of the second image sensor according to the second image object position (Paragraph [0081] " weight values may be assigned to first points according to at least one of a shooting angle, an image noise value, or a normal direction of the initial image in which the first point is located. For example, a shooting angle of the front face image is the most positive, and has a higher precision among the nose tip points. Therefore, a weight assigned to the first point of the front face image is 60%, and weights assigned to the first points of the left face image and the right face image are respectively 20%. ") . Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Xu’s reference to include obtaining a first image object position of the first ROI in the first image and a second image object position of the second ROI in the second image; determining the fusion weight of the first image sensor according to the first image object position; and determining the fusion weight of the second image sensor according to the second image object position taught by Lin’s reference. The motivation for doing so would have been to more accurately fuse images based on their different weights as suggested by Lin (see Lin, Paragraph [0083]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Lin with Xu to obtain the invention specified in claim 9 . 07-21-aia AIA Claim s 10-13 are rejected under 35 U.S.C. 103 as being unpatentable over Xu (US 12,217,545 B2) in view of Lin (US 2022/0165031 A1), and further in view Gupta et al. (US 2023/0360566 A1) (hereinafter, “Gupta”) . Regarding claim 10, which claim 9 is incorporated, Xu fails to teach wherein the step of determining the fusion weight of the first image sensor according to the first image object position comprises: obtaining a target distance between the first image object position and an image boundary of the first image; and determining the fusion weight of the first image sensor according to the target distance. Lin teaches wherein the step of determining the fusion weight of the first image sensor according to the first image object position (Paragraph [0078] " in the front face image, there is a 3D point A1 at a nose tip of a user, and a coordinate value is (x1, y1, z1); in the left face image, there is a 3D point B1 at the nose tip of the user, and a coordinate value is (x2, y2, z2); and in the right face image, there is a 3D point C1 at the nose tip of the user, and a coordinate value is (x3, y3, z3). After the left face image and the right face image are moved according to step 306 above, three 3D points A1, B1 and C1 overlap with each other. In this case, fusion on A1, B1, and C1 is performed to obtain a 3D point D1 (x4, y4, z4), which is a 3D point at the nose tip of the user in the second point cloud information.; Paragraph [0081] " weight values may be assigned to first points according to at least one of a shooting angle, an image noise value, or a normal direction of the initial image in which the first point is located. For example, a shooting angle of the front face image is the most positive, and has a higher precision among the nose tip points. Therefore, a weight assigned to the first point of the front face image is 60%, and weights assigned to the first points of the left face image and the right face image are respectively 20%.”). Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Xu’s reference to include the step of determining the fusion weight of the first image sensor according to the first image object position taught by Lin’s reference. The motivation for doing so would have been to more accurately fuse images based on their different weights as suggested by Lin (see Lin, Paragraph [0083]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. However, Xu and Lin fail to teach obtaining a target distance between the first image object position and an image boundary of the first image; and determining the fusion weight of the first image sensor according to the target distance. Gupta teaches obtaining a target distance between the first image object position and an image boundary of the first image; and determining the fusion weight of the first image sensor according to the target distance (closest distance in Paragraph [0073] equates to target distance) (Paragraph [0073] "When cropped ROI may not be fully encompassed by the blue overlapped FOV, the content of the first view may be 100% selected. On the flip side, if cropped ROI fully resides inside the inner blue region, i.e., overlapped FOV + buffering region for image fusion, the second/narrow view which normally carries the higher IQ may be 100% selected. Anywhere in between the image fusion algorithm may be called whenever any edge of the cropped ROI (in red) locates inside the blue shaded area and stops when it exits. The fusion weights are dynamically adjusted based on the closest distance between red edge and blue edges.) . Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Xu in view of Feng to include obtaining a target distance between the first image object position and an image boundary of the first image; and determining the fusion weight of the first image sensor according to the target distance taught by Gupta’s reference. The motivation for doing so would have been to seamlessly transition views when viewing regions of interest from close objects to far objects as suggested by Gupta (see Gupta, Paragraph [0073]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Lin and Gupta with Xu to obtain the invention specified in claim 10. Regarding claim 11, which claim 10 is incorporated, Xu and Lin fail to teach wherein the step of obtaining the target distance between the first image object position and the image boundary of the first image comprises: calculating a first reference distance between the first image object position and a first reference image boundary of the first image; calculating a second reference distance between the first image object position and a second reference image boundary of the first image, wherein the first reference image boundary is parallel to the second reference image boundary; and determining the smaller one between the first reference distance and the second reference distance as the target distance. Gupta teaches wherein the step of obtaining the target distance between the first image object position and the image boundary of the first image comprises: calculating a first reference distance between the first image object position and a first reference image boundary (distance from red edge Paragraph [0073 equates to first reference image boundary) of the first image; calculating a second reference distance between the first image object position and a second reference image boundary (distance from blue edge Paragraph [0073 equates to second reference image boundary) of the first image, wherein the first reference image boundary is parallel to the second reference image boundary; and determining the smaller one between the first reference distance and the second reference distance as the target distance (closest distance in Paragraph [0073] equates to target distance) (Paragraph [0073] "When cropped ROI may not be fully encompassed by the blue overlapped FOV, the content of the first view may be 100% selected. On the flip side, if cropped ROI fully resides inside the inner blue region, i.e., overlapped FOV + buffering region for image fusion, the second/narrow view which normally carries the higher IQ may be 100% selected. Anywhere in between the image fusion algorithm may be called whenever any edge of the cropped ROI (in red) locates inside the blue shaded area and stops when it exits. The fusion weights are dynamically adjusted based on the closest distance between red edge and blue edges.) . Figure 17 PNG media_image1.png 524 914 media_image1.png Greyscale Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Xu in view of Feng to include wherein the step of obtaining the target distance between the first image object position and the image boundary of the first image comprises: calculating a first reference distance between the first image object position and a first reference image boundary of the first image; calculating a second reference distance between the first image object position and a second reference image boundary of the first image, wherein the first reference image boundary is parallel to the second reference image boundary; and determining the smaller one between the first reference distance and the second reference distance as the target distance taught by Gupta’s reference. The motivation for doing so would have been to seamlessly transition views when viewing regions of interest from close objects to far objects as suggested by Gupta (see Gupta, Paragraph [0073]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Lin and Gupta with Xu to obtain the invention specified in claim 11. Regarding claim 12, which claim 11 is incorporated, Xu and Lin fail to teach wherein the step of obtaining the target distance between the first image object position and the image boundary of the first image further comprises: determining that the first reference image boundary and the second reference image boundary are vertical image boundaries or horizontal image boundaries according to arrangement of the first image sensor and the second image sensor. Gupta teaches wherein the step of obtaining the target distance between the first image object position and the image boundary of the first image (Paragraph [0073] "When cropped ROI may not be fully encompassed by the blue overlapped FOV, the content of the first view may be 100% selected. On the flip side, if cropped ROI fully resides inside the inner blue region, i.e., overlapped FOV + buffering region for image fusion, the second/narrow view which normally carries the higher IQ may be 100% selected. Anywhere in between the image fusion algorithm may be called whenever any edge of the cropped ROI (in red) locates inside the blue shaded area and stops when it exits. The fusion weights are dynamically adjusted based on the closest distance between red edge and blue edges. This strategy provides an economic way to seamlessly transit from view0 to view1 when viewing interest regions shifts from close/middle-field objects to far-field objects and vice versa. FIG. 17 illustraties the position weight from relative position of crop ROI and overlapped FOV.") further comprises: determining that the first reference image boundary and the second reference image boundary are vertical image boundaries or horizontal image boundaries according to arrangement of the first image sensor and the second image sensor (Paragraph [0069] "[0069] FIG. 11 illustrates another exemplary class diagram of image fusion with the incorporation of the weights needed for fusion. There are two types of weights calculated in this image fusion. The first type or a blended weight may be derived from the position of intermediate view with respect to wide view as the pivot. At the beginning of the transition where the point of view may be close to the wide camera, a is close to 1.0. As the intermediate view transits to the right/narrow view") . Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Xu in view of Feng to include wherein the step of obtaining the target distance between the first image object position and the image boundary of the first image further comprises: determining that the first reference image boundary and the second reference image boundary are vertical image boundaries or horizontal image boundaries according to arrangement of the first image sensor and the second image sensor taught by Gupta’s reference. The motivation for doing so would have been to seamlessly transition views when viewing regions of interest from close objects to far objects as suggested by Gupta (see Gupta, Paragraph [0073]). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Lin and Gupta with Xu to obtain the invention specified in claim 12 . 07-21-aia AIA Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Xu (US 12,217,545 B2) in view of Lin (US 2022/0165031 A1), further in view Gupta et al. (US 2023/0360566 A1) (hereinafter, “Gupta”), and further in view of Feng et al. (US 2022/0392036 A1) (hereinafter “Feng”) Regarding claim 13, which claim 10 is incorporated, Xu, Lin, and Gupta fail to teach wherein the fusion weight of the first image sensor is the Nth power of the target distance, and N is greater than 0. Feng teaches wherein the fusion weight of the first image sensor is the Nth power of the target distance, and N is greater than 0 (Examiner interprets “fusion weight…is the Nth power of the target distance, and N is greater than 0” to cover any positive relationship between the weight and the distance, wherein when n=1, the weighting factor corresponds to a linear relationship.”) (Paragraph [0114] “determine the values of the distance weighting table based at least in part on the diagonal length and pixel distance values. For example, camera processor(s) 14 may determine pixel distance values from the center (C) of composite frame 2002 (e.g., pixel_distance). Camera processor(s) 14 may, for each pixel location, determine W_dist values by dividing pixel_distance by half of the diagonal length (L) to determine W_dist values for a distance weighting table…the transition area of composite frame 2002 corresponds to the region shown by ramp region 1408. The weight value correlations may be linear”) . Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Xu’s reference to include wherein the fusion weight of the first image sensor is the Nth power of the target distance, and N is greater than 0 taught by Feng’s reference. The motivation for doing so would have been to generate a composite frame in which regions at different distances are in focus as suggested by Feng (see Feng, Abstract). Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Feng with Xu, Lin, and Gupta to obtain the invention specified in claim 13 . Allowable Subject Matter 12-151-08 AIA 07-43 12-51-08 Claim 15-18 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claims 15-18 contain subject matter that is not disclosed or made obvious in the cited art: In regard to claim 15 , when considering claim 15 as a whole, prior art of record fails to disclose or render obvious, alone or in combination: “[…] wherein the step of performing the information fusion processing on the plurality of object information according to the fusion weight of each of the image sensors to obtain the object fusion information of the real scene object comprises: performing a weighted operation on the first finger bending information and the second finger bending information according to the fusion weight of the first image sensor and the fusion weight of the second image sensor, to obtain a fusion finger bending information of the hand object determining a target control gesture according to the fusion finger bending information of the hand object and a plurality of bending thresholds corresponding to a plurality of fingers.”. In regard to claims 16-18, claims 16-18 depend on objected claim 15. Therefore, by virtue of their dependency, claims 16-18 are also indicated as objected subject matter. Conclusion 07-96 AIA The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Henningsson et al. (US 20180367789 A1) discloses a method to combine video frames from multiple sensors with overlapping views. Zhong et al. (US 2022/0044356 A1) discloses method for real-time image stitching using different camera calibrations and homography to identify overlapping regions. Zhang et al. (US 2023/0260155 A1) discloses a wrist-mounted system that uses a plurality of cameras mounted on a band to capture images of a hand. The images are used to track poses and gestures, with the camera positioned close to the wrist for optimal field of view. Any inquiry concerning this communication or earlier communications from the examiner should be directed to UROOJ FATIMA whose telephone number is (571)272-2096. The examiner can normally be reached M-F 8:00-5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Henok Shiferaw can be reached at (571) 272-4637. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /UROOJ FATIMA/Examiner, Art Unit 2676 /Henok Shiferaw/Supervisory Patent Examiner, Art Unit 2676 Application/Control Number: 18/739,269 Page 2 Art Unit: 2676 Application/Control Number: 18/739,269 Page 3 Art Unit: 2676 Application/Control Number: 18/739,269 Page 4 Art Unit: 2676 Application/Control Number: 18/739,269 Page 5 Art Unit: 2676 Application/Control Number: 18/739,269 Page 6 Art Unit: 2676 Application/Control Number: 18/739,269 Page 7 Art Unit: 2676 Application/Control Number: 18/739,269 Page 8 Art Unit: 2676 Application/Control Number: 18/739,269 Page 9 Art Unit: 2676 Application/Control Number: 18/739,269 Page 10 Art Unit: 2676 Application/Control Number: 18/739,269 Page 11 Art Unit: 2676 Application/Control Number: 18/739,269 Page 12 Art Unit: 2676 Application/Control Number: 18/739,269 Page 13 Art Unit: 2676 Application/Control Number: 18/739,269 Page 14 Art Unit: 2676 Application/Control Number: 18/739,269 Page 15 Art Unit: 2676 Application/Control Number: 18/739,269 Page 16 Art Unit: 2676 Application/Control Number: 18/739,269 Page 17 Art Unit: 2676 Application/Control Number: 18/739,269 Page 18 Art Unit: 2676 Application/Control Number: 18/739,269 Page 19 Art Unit: 2676 Application/Control Number: 18/739,269 Page 20 Art Unit: 2676 Application/Control Number: 18/739,269 Page 21 Art Unit: 2676 Application/Control Number: 18/739,269 Page 22 Art Unit: 2676 Application/Control Number: 18/739,269 Page 23 Art Unit: 2676 Application/Control Number: 18/739,269 Page 24 Art Unit: 2676 Application/Control Number: 18/739,269 Page 25 Art Unit: 2676 Application/Control Number: 18/739,269 Page 26 Art Unit: 2676 Application/Control Number: 18/739,269 Page 27 Art Unit: 2676 Application/Control Number: 18/739,269 Page 28 Art Unit: 2676 Application/Control Number: 18/739,269 Page 29 Art Unit: 2676 Application/Control Number: 18/739,269 Page 30 Art Unit: 2676 Application/Control Number: 18/739,269 Page 31 Art Unit: 2676 Application/Control Number: 18/739,269 Page 32 Art Unit: 2676 Application/Control Number: 18/739,269 Page 33 Art Unit: 2676 Application/Control Number: 18/739,269 Page 34 Art Unit: 2676 Application/Control Number: 18/739,269 Page 35 Art Unit: 2676 Application/Control Number: 18/739,269 Page 36 Art Unit: 2676 Application/Control Number: 18/739,269 Page 37 Art Unit: 2676
Read full office action

Prosecution Timeline

Jun 10, 2024
Application Filed
Apr 15, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705860
COMPUTER-IMPLEMENTED OBJECT DETECTION METHOD, OBJECT DETECTION APPARATUS, AND COMPUTER-READABLE MEDIUM
2y 8m to grant Granted Aug 11, 2026
Patent 12693409
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND NON-TRANSITORY STORAGE MEDIUM
2y 8m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 2 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
75%
Grant Probability
75%
With Interview (+0.0%)
2y 8m (~4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 8 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month