DETAILED ACTION
Claims 1-20 are pending in this application. Claims 1-20 have been given the effective filing date of 10/27/2022 in accordance with the applicant’s claim for foreign priority. Claims 1-3, 11-13, 15, 18 and 20 have been amended in this present application.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority, certified copies have been received.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 03/31/2023 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 03/11/2025 has been entered.
Response to Arguments
35 U.S.C. 112(f)
Applicant’s arguments (see Remarks filed 03/11/2026) have been fully considered by the examiner and are not persuasive as discussed in the previously issued Final Rejection mailed on 01/14/2026. As previously discussed, taking “first sub model” as example, applicant states in claim 3, that “a first sub-model” is “configured to estimate the illuminant map from the input image set”, which constitutes a placeholder (First sub-model), and functional language to modify the placeholder a first sub-model” which meets Prong A and B of the three prongs for invoking 35 U.S.C. 112(f) (see MPEP section 2181, section 1, subsections a and b). Further, the “first sub-model” is not modified by sufficient structure, material or acts for which to perform the specified function, which meets Prong C of the three prongs for invoking 35 U.S.C. 112(f) (see MPEP section 2181, section 1, subsection c). Additionally, “second sub-model”, in claims 3, 4, 13, 14, and 20, “First and second feature extraction model” in claims 4, 5, 14 and 15 and “confidence score estimation model” in claims 4, 5, 14 and 15 all follow the same logic as described above for meeting the three prongs for invoking 35 U.S.C. 112(f). Therefore, for at least these reasons, the examiner respectfully maintains the claim interpretations under 35 U.S.C. 112(f)
35 U.S.C. 103
Applicant’s arguments (see Remarks filed 03/11/2026) have been fully considered by the examiner and are persuasive. However, in view of the newly added limitations to claims 1-3, 11-13, 15, 18 and 20, a new grounds of rejection is presented and fully discussed below in view of Du.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are:
First and second sub-model in claims 3, 4, 13, 14, and 20.
First and second feature extraction model in claims 4, 5, 14 and 15.
And confidence score estimation model in claims 4, 5, 14 and 15.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Du (CN 114782298 A).
Regarding claim 1 Du discloses; A method comprising:
obtaining a visible light image comprising first local areas and an infrared image comprising second local areas (Du, [0017] an infrared and an a visible light image are obtained, [0039] multiple salient regions of the two images are extracted);
PNG
media_image1.png
82
728
media_image1.png
Greyscale
(Du, [0017])
PNG
media_image2.png
274
746
media_image2.png
Greyscale
(Du, [0039])
estimating an illuminant map based on visible light and the infrared image wherein the illuminant map comprises an estimate of an effect of an illumination source on a color of an object in the visible light image (Du,[0019]-[0021] the visible light and infrared images are input into an encoder after being converted to YCbCr color space to generate feature maps (FR and FV), [0072] a global fusion feature map FF_0 (illuminant map) is then generated using FR and FV (feature maps of illuminant effects) since the images are converted to YCbCr space, this color space provides information about the luminance and chrominance values on the images, which are effects of illumination in the image);
PNG
media_image3.png
400
730
media_image3.png
Greyscale
(Du, [0019]-[0021])
PNG
media_image4.png
112
732
media_image4.png
Greyscale
(Du, [0072])
generating, using a neural network model (Du, [0007] the system uses a fusion network, which is a type of neural network), a confidence score map representing a correlation between the first local areas of the visible light image and the second local areas of the infrared image by performing a cross- attention operation between the visible light image and the infrared image to obtain an attention map (Du, Claim 5, Figure 3 and [0033] Feature of a visible light image (FV) and a feature of an infrared image (FR) are input into the model to obtain an attention feature map (attention feature map map) for both images this is then used to generate the fusion attention feature map MRV (confidence score map), [0045] this is done for multiple regions of each image, where a feature map F is input for the region and an attention map M is output for the region, [0074] the generation of the fusion attention feature map (confidence score map) is done using a weighting of the features in the feature map, since a fusion attention feature map is created using weighted features from both the IR and visible light images this indicates that the two images are correlated),
PNG
media_image5.png
112
734
media_image5.png
Greyscale
(Du, claim 5)
PNG
media_image6.png
180
728
media_image6.png
Greyscale
(Du, [0033])
PNG
media_image7.png
552
738
media_image7.png
Greyscale
(Du, [0074])
PNG
media_image8.png
272
444
media_image8.png
Greyscale
(Du, figure 3)
wherein the confidence score map is based on the attention map (Du, [0033] Feature of a visible light image (FV) and a feature of an infrared image (FR) are input into the model to obtain an attention feature map (attention feature map) for both images this is then used to generate the fusion attention feature map MRV (confidence score map)),
and wherein the confidence score map comprises weights for each corresponding pairs of the first local area and the second local areas depending on the correlation (Du, [0073]-[0074] in generating the fusion attention feature map (confidence score map) a global average pooling function is performed on the feature maps to generate a vector, which is then weighted using a fully connected layer and an activation layer, where the weight corresponds to the importance of the feature determined for the pair of inputs, therefore the weights are determined based on correlation between the inputs),
and each weight of the weights indicates a likelihood that the estimate of the effect of the illumination source on the color of the object in the visible light image is accurate (Du, [0074] the feature maps are weighted to ensure that the regions with the greatest effect on the illumination have a bigger impact, which impacts the accuracy of the illumination mapping as described in applicant’s specification paragraphs [0030], [0042]-[0043], then the weighted feature maps are convolved to generate an attention feature map (confidence score map) which is based on the weighted features/weight vectors);
determining illuminant information of the visible light image based on the illuminant map and the confidence score map by obtaining a weighted sum between the illuminant map and the confidence score map based on the weights (Du, [0076] a final fused feature map is generated FF (illuminant information) which contains the illuminant information and features of both the visible light and infrared images, this final fused feature map is generated by summing the global fusion feature map FF_0 and the attention feature map MRV as shown in figure 3, these maps are summed for each of the pixels at each location, where [0073]-[0075] notes that the attention feature map (confidence score map) is generated from weight vectors for the features of the images),
PNG
media_image9.png
250
732
media_image9.png
Greyscale
(Du, [0075]-[0076])
PNG
media_image10.png
270
430
media_image10.png
Greyscale
(Du, figure 3 emphasis added to show the generation of the illuminant information (FF ) from the confidence score map (MRV) and the illuminant map (FF_0))
wherein a first area of the first local areas has a greater influence on the illuminant information than a second area of the second local areas based on the confidence score map (Du [0073]-[0076] the weights of the attention feature map (confidence score map) are determined based on the region’s importance/impact on the image, further [0007] states that regions in the visible light image (first areas) which are well exposed may be preserved in the final fused image by increasing their intensity, this is done via the weighting described in [0073]-[0076]);
PNG
media_image11.png
308
736
media_image11.png
Greyscale
PNG
media_image12.png
100
730
media_image12.png
Greyscale
(Du, [0007])
PNG
media_image13.png
274
756
media_image13.png
Greyscale
(Du, [0039])
PNG
media_image14.png
228
740
media_image14.png
Greyscale
(Du, [0095])
and processing the visible light image using the illuminant information to obtain an output image (Du, [0076]-[0077] and figure 1, the final fused feature map (illuminant information) is input into the decoder to generate the fused image (output image) OF, [0095] the final fused images are shown in figure 5, they capture both the visible and infrared features in a single enhanced image).
PNG
media_image15.png
226
696
media_image15.png
Greyscale
(Du, [0076]-[0077])
PNG
media_image16.png
238
736
media_image16.png
Greyscale
(Du, [0095])
PNG
media_image17.png
240
586
media_image17.png
Greyscale
(Du, figure 1)
PNG
media_image18.png
442
584
media_image18.png
Greyscale
(Du, figure 5)
Regarding claim 2 Du discloses; The method of claim 1, wherein:
an illuminant map represents the illuminant configuration for each local area (Du,[0019]-[0021] the visible light and infrared images are input into an encoder after being converted to YCbCr color space to generate feature maps (FR and FV), [0072] a global fusion feature map FF_0 (illuminant map) is then generated using FR and FV (feature maps of illuminant effects) since the images are converted to YCbCr space, this color space provides information about the luminance and chrominance values on the images, which are effects of illumination in the image),
and the confidence score map represents the correlation between the visible light image and the infrared image for the each local area (Du, Claim 5, Figure 3 and [0033] Feature of a visible light image (FV) and a feature of an infrared image (FR) are input into the model to obtain an attention feature map (attention feature map) for both images this is then used to generate the fusion attention feature map MRV (confidence score map), [0045] this is done for multiple regions of each image, where a feature map F is input for the region and an attention map M is output for the region, [0074] the generation of the fusion attention feature map (confidence score map) is done using a weighting of the features in the feature map, since a fusion attention feature map is created using weighted features from both the IR and visible light images this indicates that the two images are correlated).
Regarding claim 3 Du discloses; The method of claim 1, wherein the neural network model comprises:
a first sub-model configured to estimate the illuminant map from an input image set (Du,[0019]-[0021] the visible light and infrared images are input into an encoder after being converted to YCbCr color space to generate feature maps (FR and FV), [0072] a global fusion feature map FF_0 (illuminant map) is then generated using FR and FV (feature maps of illuminant effects) since the images are converted to YCbCr space, this color space provides information about the luminance and chrominance values on the images, which are effects of illumination in the image, figure 3 shows the set of convolutional layers/sub model portion which generates the illuminant map FF_0);
and a second sub-model configured to estimate the confidence score map from the input image set (Du, Claim 5, Figure 3 and [0033] Feature of a visible light image (FV) and a feature of an infrared image (FR) are input into the model to obtain an attention feature map (attention feature map) for both images this is then used to generate the fusion attention feature map MRV (confidence score map), [0045] this is done for multiple regions of each image, where a feature map F is input for the region and an attention map M is output for the region, [0074] the generation of the fusion attention feature map (confidence score map) is done using a weighting of the features in the feature map, since a fusion attention feature map is created using weighted features from both the IR and visible light images this indicates that the two images are correlated, figure 3 shows the two Regional attention modules (RABs) which are the second submodel parts to generate the fusion attention feature map MRV).
PNG
media_image19.png
272
472
media_image19.png
Greyscale
(Du, figure 3 emphasis added)
Regarding claim 4 Du discloses; The method of claim 3, wherein the second sub-model comprises (Du, figure 3 shows a pair of Regional Attention Modules RABs (second sub-model) used to generate the fused attention feature map (confidence score map):
a first feature extraction model configured to extract a visible light feature map from the visible light image of the input image set (Du, [0066] the visible light image Iv is input into a first encoder (first feature extraction model) to obtain a visible light feature map Fv as shown in figure 1 below);
a second feature extraction model configured to extract an infrared feature map from the infrared image of the input image set (Du, [0066] the infrared light image IR is input into a second encoder (second feature extraction model) to obtain an infrared light feature map FR as shown in figure 1 below);
PNG
media_image20.png
128
744
media_image20.png
Greyscale
(Du, [0066])
PNG
media_image21.png
254
580
media_image21.png
Greyscale
(Du, figure 1, emphasis added to show the first and second feature extraction models (encoders) of the second submodel)
and a confidence score estimation model configured to estimate the confidence score map based on a correlation between the visible light feature map and the infrared feature map (Du, [0067] the model contains two paths (two submodels), where the first one is a set of convolutional layers to obtain a fused feature map (illuminant map) and the second (second submodel) obtains the attention feature map MRV (confidence score map), where attention feature map MRV (confidence score map) is generated using the Regional Attention Modules RAB (confidence score estimation model), [0073]-[0074] the two feature maps are input into the RAB modules to generate the fusion attention feature map (confidence score map), in generating the fusion attention feature map (confidence score map) a global average pooling function is performed on the feature maps to generate a vector, which is then weighted using a fully connected layer and an activation layer, where the weight corresponds to the importance of the feature determined for the pair of inputs, therefore the weights are determined based on correlation between the inputs).
PNG
media_image22.png
296
426
media_image22.png
Greyscale
(Du, Figure 3, emphasis added)
PNG
media_image23.png
314
732
media_image23.png
Greyscale
(Du, [0073])
Regarding claim 5 Du discloses; The method of claim 4, wherein the cross-attention operation comprises: determining first data corresponding to one of a visible light feature map and an infrared feature map and second data corresponding to another one (Du, [0066] two feature maps FR and FV corresponding to the input visible light and infrared images respectively are generated to be used as inputs);
generating the attention map based on query data according to the first data and key data according to the second data (Du, [0073] – [0074] an attention map is generated from the two input feature maps corresponding to each of the visible light (MV) and infrared images (MR) (first/query data and second/key data respectively));
and estimating the confidence score map based on value data according to one of the first data and the second data and the attention map (Du, [0084]-[0075] the attention map (MR , MV), which is based on the feature maps (first and second data) is input into a convolutional attention layer in feature space and used to generate the fused attention feature map M RV(confidence score map) as shown in figure 3 below).
PNG
media_image22.png
296
426
media_image22.png
Greyscale
(Du, Figure 3 emphasis added)
Regarding claim 6 Du discloses; The method of claim 1, wherein the infrared image comprises a multi-band (Du, [0035] notes that the infrared images may contain sub-images which are captured from the public TNO dataset, the TNO dataset is a multiband image dataset therefore the images may be multiband).
Regarding claim 7 Du discloses; The method of claim 1, wherein the determining of the illuminant information comprises:
obtaining a weighted sum between the illuminant map and the confidence score map (Du, [0076] a final fused feature map is generated FF (illuminant information) which contains the illuminant information and features of both the visible light and infrared images, this final fused feature map is generated by summing the global fusion feature map FF_0 and the attention feature map MRV as shown in figure 3, these maps are summed for each of the pixels at each location, where [0073]-[0075] notes that the attention feature map (confidence score map) is generated from weight vectors for the features of the images);
and determining the illuminant information according to the weighted sum (Du, [0076] a final fused feature map is generated FF (illuminant information) which contains the illuminant information and features of both the visible light and infrared images, this final fused feature map is generated by summing the global fusion feature map FF_0 and the attention feature map MRV as shown in figure 3, these maps are summed for each of the pixels at each location, where [0073]-[0075] notes that the attention feature map (confidence score map) is generated from weight vectors for the features of the images, further, [0054] notes that the final fused feature map is input into the decoder to obtain the luminance channel values for the final fused image, indicating this final fused feature map contains the illumination information for the images).
Regarding claim 8 Du discloses; The method of claim 7, wherein the obtaining of the weighted sum comprises: obtaining the weighted sum by summing vectors for each local area according to the illuminant map using a weight according to the confidence score map (Du, [0074] the vectors of the weights for the attention feature map are summed with the feature map inputs to generate the attention feature map (confidence score maps) and then the values for the attention feature map (confidence score map) is summed with the pixel values for each local area of the global feature map (illuminant map) using a convolution layer, indicating this is a vector operation).
Regarding claim 9 Du discloses; The method of claim 8, wherein: local areas of the visible light image and local areas of the infrared image form corresponding pairs (Du, [0063], the visible light and infrared images are acquired in pairs, where [0039] multiple salient regions of the two images are extracted (local areas)),
the confidence score map comprises weights for each of the corresponding pairs (Du, Claim 5, Figure 3 and [0033] Feature of a visible light image (FV) and a feature of an infrared image (FR) are input into the model to obtain an attention feature map (attention feature map map) for both images this is then used to generate the fusion attention feature map MRV (confidence score map), [0045] this is done for multiple regions of each image, where a feature map F is input for the region and an attention map M is output for the region),
and a weight of a corresponding pair corresponds to a correlation between visible light data and infrared data of the corresponding pair (Du, Claim 5, Figure 3 and [0033] Feature of a visible light image (FV) and a feature of an infrared image (FR) are input into the model to obtain an attention feature map (attention feature map map) for both images this is then used to generate the fusion attention feature map MRV (confidence score map), [0045] this is done for multiple regions of each image, where a feature map F is input for the region and an attention map M is output for the region, [0074] the generation of the fusion attention feature map (confidence score map) is done using a weighting of the features in the feature map, since a fusion attention feature map is created using weighted features from both the IR and visible light images this indicates that the two images are correlated).
Regarding claim 10 Du discloses; A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method of claim 1 (Du, [0060] the system was trained and implemented on an NVIDIA Titan V GPU (processor)).
Regarding claim 11, Du discloses; An apparatus comprising: a processor (Du, [0060] the system was trained and implemented on an NVIDIA Titan V GPU (processor));
and a memory configured to store instructions executable by the processor, wherein, in response to the instructions being executed by the processor, the processor is configured to (Du, [0060] the system was trained and implemented on an NVIDIA Titan V GPU (processor), this processor has both memory to store the code and images as well as the capacity of a processor to execute the program):
obtain a visible light image comprising first local areas and an infrared image comprising second local areas (Du, [0017] an infrared and an a visible light image are obtained, [0039] multiple salient regions of the two images are extracted);
estimate an illuminant map based on visible light and the infrared image wherein the illuminant map comprises an estimate of an effect of an illumination source on a color of an object in the visible light image (Du,[0019]-[0021] the visible light and infrared images are input into an encoder after being converted to YCbCr color space to generate feature maps (FR and FV), [0072] a global fusion feature map FF_0 (illuminant map) is then generated using FR and FV (feature maps of illuminant effects) since the images are converted to YCbCr space, this color space provides information about the luminance and chrominance values on the images, which are effects of illumination in the image);
generate, using a neural network model (Du, [0007] the system uses a fusion network, which is a type of neural network), a confidence score map representing a correlation between the first local areas of the visible light image and the second local areas of the infrared image by performing a cross- attention operation between the visible light image and the infrared image to obtain an attention map (Du, Claim 5, Figure 3 and [0033] Feature of a visible light image (FV) and a feature of an infrared image (FR) are input into the model to obtain an attention feature map (attention feature map map) for both images this is then used to generate the fusion attention feature map MRV (confidence score map), [0045] this is done for multiple regions of each image, where a feature map F is input for the region and an attention map M is output for the region, [0074] the generation of the fusion attention feature map (confidence score map) is done using a weighting of the features in the feature map, since a fusion attention feature map is created using weighted features from both the IR and visible light images this indicates that the two images are correlated),
wherein the confidence score map is based on the attention map (Du, [0033] Feature of a visible light image (FV) and a feature of an infrared image (FR) are input into the model to obtain an attention feature map (attention feature map) for both images this is then used to generate the fusion attention feature map MRV (confidence score map)),
and wherein the confidence score map comprises weights for each corresponding pairs of the first local area and the second local areas depending on the correlation (Du, [0073]-[0074] in generating the fusion attention feature map (confidence score map) a global average pooling function is performed on the feature maps to generate a vector, which is then weighted using a fully connected layer and an activation layer, where the weight corresponds to the importance of the feature determined for the pair of inputs, therefore the weights are determined based on correlation between the inputs),
and each weight of the weights indicates a likelihood that the estimate of the effect of the illumination source on the color of the object in the visible light image is accurate (Du, [0074] the feature maps are weighted to ensure that the regions with the greatest effect on the illumination have a bigger impact, which impacts the accuracy of the illumination mapping as described in applicant’s specification paragraphs [0030], [0042]-[0043], then the weighted feature maps are convolved to generate an attention feature map (confidence score map) which is based on the weighted features/weight vectors);
determine illuminant information of the visible light image based on the illuminant map and the confidence score map by obtaining a weighted sum between the illuminant map and the confidence score map based on the weights (Du, [0076] a final fused feature map is generated FF (illuminant information) which contains the illuminant information and features of both the visible light and infrared images, this final fused feature map is generated by summing the global fusion feature map FF_0 and the attention feature map MRV as shown in figure 3, these maps are summed for each of the pixels at each location, where [0073]-[0075] notes that the attention feature map (confidence score map) is generated from weight vectors for the features of the images),
wherein a first area of the first local areas has a greater influence on the illuminant information than a second area of the second local areas based on the confidence score map (Du [0073]-[0076] the weights of the attention feature map (confidence score map) are determined based on the region’s importance/impact on the image, further [0007] states that regions in the visible light image (first areas) which are well exposed may be preserved in the final fused image by increasing their intensity, this is done via the weighting described in [0073]-[0076]);
and process the visible light image using the illuminant information to obtain an output image (Du, [0076]-[0077] and figure 1, the final fused feature map (illuminant information) is input into the decoder to generate the fused image (output image) OF, [0095] the final fused images are shown in figure 5, they capture both the visible and infrared features in a single enhanced image).
Regarding claim 12, Du discloses; The apparatus of claim 11, wherein: an illuminant map represents the illuminant configuration for each local area (Du,[0019]-[0021] the visible light and infrared images are input into an encoder after being converted to YCbCr color space to generate feature maps (FR and FV), [0072] a global fusion feature map FF_0 (illuminant map) is then generated using FR and FV (feature maps of illuminant effects) since the images are converted to YCbCr space, this color space provides information about the luminance and chrominance values on the images, which are effects of illumination in the image),
and the confidence score map represents the correlation between the visible light image and the infrared image for the each local area (Du, Claim 5, Figure 3 and [0033] Feature of a visible light image (FV) and a feature of an infrared image (FR) are input into the model to obtain an attention feature map (attention feature map) for both images this is then used to generate the fusion attention feature map MRV (confidence score map), [0045] this is done for multiple regions of each image, where a feature map F is input for the region and an attention map M is output for the region, [0074] the generation of the fusion attention feature map (confidence score map) is done using a weighting of the features in the feature map, since a fusion attention feature map is created using weighted features from both the IR and visible light images this indicates that the two images are correlated).
Regarding claim 13 Du discloses; The apparatus of claim 11, wherein the neural network model comprises:
a first sub-model configured to estimate the illuminant map from an input image set (Du,[0019]-[0021] the visible light and infrared images are input into an encoder after being converted to YCbCr color space to generate feature maps (FR and FV), [0072] a global fusion feature map FF_0 (illuminant map) is then generated using FR and FV (feature maps of illuminant effects) since the images are converted to YCbCr space, this color space provides information about the luminance and chrominance values on the images, which are effects of illumination in the image, figure 3 shows the set of convolutional layers/sub model portion which generates the illuminant map FF_0);
and a second sub-model configured to estimate the confidence score map from the input image set (Du, Claim 5, Figure 3 and [0033] Feature of a visible light image (FV) and a feature of an infrared image (FR) are input into the model to obtain an attention feature map (attention feature map) for both images this is then used to generate the fusion attention feature map MRV (confidence score map), [0045] this is done for multiple regions of each image, where a feature map F is input for the region and an attention map M is output for the region, [0074] the generation of the fusion attention feature map (confidence score map) is done using a weighting of the features in the feature map, since a fusion attention feature map is created using weighted features from both the IR and visible light images this indicates that the two images are correlated, figure 3 shows the two Regional attention modules (RABs) which are the second submodel parts to generate the fusion attention feature map MRV).
Regarding claim 14, Du discloses; The apparatus of claim 13, wherein the second sub-model comprises(Du, figure 3 shows a pair of Regional Attention Modules RABs (second sub-model) used to generate the fused attention feature map (confidence score map):
a first feature extraction model configured to extract a visible light feature map from the visible light image of the input image set (Du, [0066] the visible light image Iv is input into a first encoder (first feature extraction model) to obtain a visible light feature map Fv as shown in figure 1 below);
a second feature extraction model configured to extract an infrared feature map from the infrared image of the input image set (Du, [0066] the infrared light image IR is input into a second encoder (second feature extraction model) to obtain an infrared light feature map FR as shown in figure 1 below);
and a confidence score estimation model configured to estimate the confidence score map based on a correlation between the visible light feature map and the infrared feature map (Du, [0067] the model contains two paths (two submodels), where the first one is a set of convolutional layers to obtain a fused feature map (illuminant map) and the second (second submodel) obtains the attention feature map MRV (confidence score map), where attention feature map MRV (confidence score map) is generated using the Regional Attention Modules RAB (confidence score estimation model), [0073]-[0074] the two feature maps are input into the RAB modules to generate the fusion attention feature map (confidence score map), in generating the fusion attention feature map (confidence score map) a global average pooling function is performed on the feature maps to generate a vector, which is then weighted using a fully connected layer and an activation layer, where the weight corresponds to the importance of the feature determined for the pair of inputs, therefore the weights are determined based on correlation between the inputs).
Regarding claim 15, Du discloses; The apparatus of claim 14, wherein the cross- attention operation comprises: determining first data corresponding to one of a visible light feature map and an infrared feature map and second data corresponding to another one (Du, [0066] two feature maps FR and FV corresponding to the input visible light and infrared images respectively are generated to be used as inputs);
generating the attention map based on query data according to the first data and key data according to the second data (Du, [0073] – [0074] an attention map is generated from the two input feature maps corresponding to each of the visible light (MV) and infrared images (MR) (first/query data and second/key data respectively));
and estimating the confidence score map based on value data according to one of the first data and the second data and the attention map (Du, [0084]-[0075] the attention map (MR , MV), which is based on the feature maps (first and second data) is input into a convolutional attention layer in feature space and used to generate the fused attention feature map M RV(confidence score map) as shown in figure 3 below).
Regarding claim 16, Du discloses; The apparatus of claim 11, wherein the infrared image comprises a multi-band (Du, [0035] notes that the infrared images may contain sub-images which are captured from the public TNO dataset, the TNO dataset is a multiband image dataset therefore the images may be multiband).
Regarding claim 17, Du discloses; The apparatus of claim 11, wherein, to determine the illuminant information, the processor is configured to:
obtain a weighted sum between the illuminant map and the confidence score map (Du, [0076] a final fused feature map is generated FF (illuminant information) which contains the illuminant information and features of both the visible light and infrared images, this final fused feature map is generated by summing the global fusion feature map FF_0 and the attention feature map MRV as shown in figure 3, these maps are summed for each of the pixels at each location, where [0073]-[0075] notes that the attention feature map (confidence score map) is generated from weight vectors for the features of the images);
and determine the illuminant information according to the weighted sum (Du, [0076] a final fused feature map is generated FF (illuminant information) which contains the illuminant information and features of both the visible light and infrared images, this final fused feature map is generated by summing the global fusion feature map FF_0 and the attention feature map MRV as shown in figure 3, these maps are summed for each of the pixels at each location, where [0073]-[0075] notes that the attention feature map (confidence score map) is generated from weight vectors for the features of the images, further, [0054] notes that the final fused feature map is input into the decoder to obtain the luminance channel values for the final fused image, indicating this final fused feature map contains the illumination information for the images).
Regarding claim 18, Du discloses; An electronic device comprising: a visible light camera configured to generate a visible light image (Du, [0002] – [0003] multiple image sensors can be used to capture visible and infrared images);
an infrared camera configured to generate an infrared image (Du, [0002] – [0003] multiple image sensors can be used to capture visible and infrared images);
and a processor configured to (Du, [0060] the system was trained and implemented on an NVIDIA Titan V GPU (processor)):
obtain a visible light image comprising first local areas and an infrared image comprising second local areas (Du, [0017] an infrared and an a visible light image are obtained, [0039] multiple salient regions of the two images are extracted);
estimate an illuminant map based on visible light and the infrared image wherein the illuminant map comprises an estimate of an effect of an illumination source on a color of an object in the visible light image (Du,[0019]-[0021] the visible light and infrared images are input into an encoder after being converted to YCbCr color space to generate feature maps (FR and FV), [0072] a global fusion feature map FF_0 (illuminant map) is then generated using FR and FV (feature maps of illuminant effects) since the images are converted to YCbCr space, this color space provides information about the luminance and chrominance values on the images, which are effects of illumination in the image);
generate, using a neural network model (Du, [0007] the system uses a fusion network, which is a type of neural network), a confidence score map representing a correlation between the first local areas of the visible light image and the second local areas of the infrared image by performing a cross- attention operation between the visible light image and the infrared image to obtain an attention map (Du, Claim 5, Figure 3 and [0033] Feature of a visible light image (FV) and a feature of an infrared image (FR) are input into the model to obtain an attention feature map (attention feature map map) for both images this is then used to generate the fusion attention feature map MRV (confidence score map), [0045] this is done for multiple regions of each image, where a feature map F is input for the region and an attention map M is output for the region, [0074] the generation of the fusion attention feature map (confidence score map) is done using a weighting of the features in the feature map, since a fusion attention feature map is created using weighted features from both the IR and visible light images this indicates that the two images are correlated),
wherein the confidence score map is based on the attention map (Du, [0033] Feature of a visible light image (FV) and a feature of an infrared image (FR) are input into the model to obtain an attention feature map (attention feature map) for both images this is then used to generate the fusion attention feature map MRV (confidence score map)),
and wherein the confidence score map comprises weights for each corresponding pairs of the first local area and the second local areas depending on the correlation (Du, [0073]-[0074] in generating the fusion attention feature map (confidence score map) a global average pooling function is performed on the feature maps to generate a vector, which is then weighted using a fully connected layer and an activation layer, where the weight corresponds to the importance of the feature determined for the pair of inputs, therefore the weights are determined based on correlation between the inputs),
and each weight of the weights indicates a likelihood that the estimate of the effect of the illumination source on the color of the object in the visible light image is accurate (Du, [0074] the feature maps are weighted to ensure that the regions with the greatest effect on the illumination have a bigger impact, which impacts the accuracy of the illumination mapping as described in applicant’s specification paragraphs [0030], [0042]-[0043], then the weighted feature maps are convolved to generate an attention feature map (confidence score map) which is based on the weighted features/weight vectors);
determine illuminant information of the visible light image based on the illuminant map and the confidence score map by obtaining a weighted sum between the illuminant map and the confidence score map based on the weights (Du, [0076] a final fused feature map is generated FF (illuminant information) which contains the illuminant information and features of both the visible light and infrared images, this final fused feature map is generated by summing the global fusion feature map FF_0 and the attention feature map MRV as shown in figure 3, these maps are summed for each of the pixels at each location, where [0073]-[0075] notes that the attention feature map (confidence score map) is generated from weight vectors for the features of the images),
wherein a first area of the first local areas has a greater influence on the illuminant information than a second area of the second local areas based on the confidence score map (Du [0073]-[0076] the weights of the attention feature map (confidence score map) are determined based on the region’s importance/impact on the image, further [0007] states that regions in the visible light image (first areas) which are well exposed may be preserved in the final fused image by increasing their intensity, this is done via the weighting described in [0073]-[0076]);
and process the visible light image using the illuminant information to obtain an output image (Du, [0076]-[0077] and figure 1, the final fused feature map (illuminant information) is input into the decoder to generate the fused image (output image) OF, [0095] the final fused images are shown in figure 5, they capture both the visible and infrared features in a single enhanced image).
Regarding claim 19, Du discloses; The electronic device of claim 18, wherein; the illuminant map represents the illuminant configuration for each local area (Du,[0019]-[0021] the visible light and infrared images are input into an encoder after being converted to YCbCr color space to generate feature maps (FR and FV), [0072] a global fusion feature map FF_0 (illuminant map) is then generated using FR and FV (feature maps of illuminant effects) since the images are converted to YCbCr space, this color space provides information about the luminance and chrominance values on the images, which are effects of illumination in the image),
and the confidence score map represents the correlation between the visible light image and the infrared image for the each local area (Du, Claim 5, Figure 3 and [0033] Feature of a visible light image (FV) and a feature of an infrared image (FR) are input into the model to obtain an attention feature map (attention feature map) for both images this is then used to generate the fusion attention feature map MRV (confidence score map), [0045] this is done for multiple regions of each image, where a feature map F is input for the region and an attention map M is output for the region, [0074] the generation of the fusion attention feature map (confidence score map) is done using a weighting of the features in the feature map, since a fusion attention feature map is created using weighted features from both the IR and visible light images this indicates that the two images are correlated).
Regarding claim 20, Du discloses; The electronic device of claim 18, wherein the neural network model comprises: a first sub-model configured to estimate the illuminant map from an input image set (Du,[0019]-[0021] the visible light and infrared images are input into an encoder after being converted to YCbCr color space to generate feature maps (FR and FV), [0072] a global fusion feature map FF_0 (illuminant map) is then generated using FR and FV (feature maps of illuminant effects) since the images are converted to YCbCr space, this color space provides information about the luminance and chrominance values on the images, which are effects of illumination in the image, figure 3 shows the set of convolutional layers/sub model portion which generates the illuminant map FF_0);
and a second sub-model configured to estimate the confidence score map from the input image set (Du, Claim 5, Figure 3 and [0033] Feature of a visible light image (FV) and a feature of an infrared image (FR) are input into the model to obtain an attention feature map (attention feature map) for both images this is then used to generate the fusion attention feature map MRV (confidence score map), [0045] this is done for multiple regions of each image, where a feature map F is input for the region and an attention map M is output for the region, [0074] the generation of the fusion attention feature map (confidence score map) is done using a weighting of the features in the feature map, since a fusion attention feature map is created using weighted features from both the IR and visible light images this indicates that the two images are correlated, figure 3 shows the two Regional attention modules (RABs) which are the second submodel parts to generate the fusion attention feature map MRV).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. For a listing of analogous art as determined by the examiner please see the attached PTO-892 Notice of References cited form.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JORDAN M ELLIOTT whose telephone number is (703)756-5463. The examiner can normally be reached M-F 8AM-5PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emily Terrell can be reached at (571) 270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.M.E./Examiner, Art Unit 2666
/EMILY C TERRELL/Supervisory Patent Examiner, Art Unit 2666