DETAILED ACTION
Claims 1-20 are pending in this application. Claims 1-20 have been given the effective filing date of 10/27/2022 in accordance with the applicant’s claim for foreign priority. Claims 1-2, 5, 9, 11-12, 15, and 18-19 have been amended in this present application.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority, certified copies have been received.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 03/31/2023 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Response to Arguments
35 U.S.C. 112(f)
Applicant’s arguments (see Remark’s filed 10/30/2025) have been fully considered by the examiner and are not persuasive. Taking “first sub model” as example, applicant states in claim 3, that “a first sub-model” is “configured to estimate the illuminant map from the input image set”, which constitutes a placeholder (First sub-model), and functional language to modify the placeholder a first sub-model” which meets Prong A and B of the three prongs for invoking 35 U.S.C. 112(f) (see MPEP section 2181, section 1, subsections a and b). Further, the “first sub-model” is not modified by sufficient structure, material or acts for which to perform the specified function, which meets Prong C of the three prongs for invoking 35 U.S.C. 112(f) (see MPEP section 2181, section 1, subsection c). Additionally, “second sub-model”, in claims 3, 4, 13, 14, and 20, “First and second feature extraction model” in claims 4, 5, 14 and 15 and “confidence score estimation model” in claims 4, 5, 14 and 15 all follow the same logic as described above for meeting the three prongs for invoking 35 U.S.C. 112(f). Therefore, for at least these reasons, the examiner respectfully maintains the claim interpretations under 35 U.S.C. 112(f)
35 U.S.C. 103
Applicant’s arguments (see Remark’s filed 10/30/2025) have been fully considered by the examiner and are persuasive. However, in view of the newly added limitations to claims 1, 5, 11, 15, and 18, a new grounds of rejection is presented and fully discussed below in view of Li, Peake and Zhou.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are:
First and second sub-model in claims 3, 4, 13, 14, and 20.
First and second feature extraction model in claims 4, 5, 14 and 15.
And confidence score estimation model in claims 4, 5, 14 and 15.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
1. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Li (CN 112184604A, Published 01/05/2021) in view of Peake (US 20220044044 A1) and in further view of Zhou (US 20220392201 A1).
Regarding claim 1 Li discloses; A method comprising:
forming an input image set by combining a visible light image and an infrared image (Li, [0049] the method provided is based on the fusion of a visible light image and a near infrared image, in paragraph [0038] of applicant’s specification the visible light image is stated as using light in the visible range of 400mn-700nm and the infrared image is stated as using a near infrared (NIR) band of 700-1000 nm range);
PNG
media_image1.png
94
580
media_image1.png
Greyscale
(Li, [0049])
estimating an illuminant map representing an illuminant configuration of the input image set (Li, [0050] a weighted least squares filter is used as step (1) to obtain filtered base layer images [0067] a base layer is obtained in step (1) of the method which removes the high frequency information, and leaving the illumination information and color information)
[ generating, using a neural network model, a confidence score map representing a correlation between the visible light image and the infrared image by performing a cross- attention operation between the visible light image and the infrared image to obtain an attention map, wherein the confidence score map is based on the attention map,
(Li [0042] and [0043] the fused basic layer (illuminant map) and the fused detail layer (confidence score map) are added to obtain a final result which takes into account the illumination information (from the base layers) and the confidence/correlation information (illuminant information), this is used to obtain a modified/fused image, Figure 15 of applicant’s specification details that the illuminant information is what is used to generate the modified image, so the combination of the two fused layers as in Li would be the equivalent to this information);
[and processing the visible light image using the illuminant information to obtain an output image.]
Li fails to teach; and processing the visible light image using the illuminant information to obtain an output image.
In the same field of endeavor, Peake teaches; and processing the visible light image using the illuminant information to obtain an output image (Peake, [0083]- [0086] channels in the stereo image pair are processed to yield and output image, which is denoted as the “augmented spectral image” and is functionally equivalent to an output image).
The combination of Li and Peake would be obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. Given that both Li and Peake teach different methods and motivation for visible and infrared light image fusion and mapping of this information, the addition of the use of a neural network as taught in Peake would be an improvement to the method of Li because the use of a neural network would improve speed and accuracy of the image fusion operations (See Peake, [0083]- [0090])
Both Li and Peake fail to teach;
generating, using a neural network model, a confidence score map representing a correlation between the visible light image and the infrared image by performing a cross- attention operation between the visible light image and the infrared image to obtain an attention map, wherein the confidence score map is based on the attention map,
However, in the same field of endeavor of image fusion, Zhou teaches; generating, using a neural network model (Zhou, [0033] neural networks are used to perform feature extraction operations and mappings), a confidence score map representing a correlation between the visible light image and the infrared image by performing a cross- attention operation between the visible light image and the infrared image to obtain an attention map (Zhou, Figure 4, S21-S26 two images which are infrared and visible light are acquired, features are extracted from both to create two feature maps (feature maps) and then per S26 a matching result map is generated based on the two feature/attention maps (attention map), [0079] the feature matching and aggregation are performed using a cross-attention operation, which generates the attention map, [0081] a matching confidence matrix (confidence map) is generated based on the match result from the feature map matching operation (attention map)), wherein the confidence score map is based on the attention map (Zhou, [0081] a matching confidence matrix (confidence map) is generated based on the match result from the feature map matching operation (attention map)),
The combination of Li, Peake and Zhou would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The combination of Li and Peake teach a system which performs fusion of two images, one infrared and one visible light, however it does not teach the use of a cross-attention operation to perform this. Zhou teaches this deficiency, where the use of a cross-attention operation would reasonably improve the system of Li and Peake by allowing for matching features from both inputs to generate matched global features, which would be effective for fusion two images of different modalities. (Zhou, [0044]- [0046])
Regarding claim 2 the combination of Li, Peake and Zhou teaches; The method of claim 1, wherein:
the illuminant map represents the illuminant configuration for each local area (Li, [0017]-[0020] the basic layers of the images (illuminant map) are generated by filtering the visible and infrared images using a weighted least squares formula, since this is being computed over the whole image it would also be computed for all local areas or pixels), and the confidence score map represents the correlation between the visible light image and the infrared image for the each local area (Li, [0035]-[0036] the generation of the fused detail layer (confidence map) includes calculation of local energies of the pixels (local areas) for each image and then correlating the two).
Regarding claim 3 the combination of Li, Peake and Zhou teaches; The method of claim 1, wherein the neural network model comprises:
a first sub-model configured to estimate the illuminant map from the input image set (Peake, [0083]-[0085]] a machine learning model may be trained to obtain information from the stereo image pair including RGB and infrared information as well as stereo information for identification of objects in images, the applicant notes that illuminant information may have RGB information, [0086] the machine learning model of Peake has multiple neural units (first and second sub models) that are used to generate the analysis outputs);
and a second sub-model configured to estimate the confidence score map from the input image set (Peake, [0078] the confidence between the features of the pairs of images (visible and infrared) are validate for confidence for each feature label using the neural network, [0086] the machine learning model of Peake has multiple neural units (first and second sub models) that are used to generate the analysis outputs).
The combination of Li and Peake would be obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. Given that both Li and Peake teach different methods and motivation for visible and infrared light image fusion and mapping of this information, the addition of the use of a neural network as taught in Peake would be an improvement to the method of Li because the use of a neural network would improve speed and accuracy of the image fusion operations (See Peake, [0083]- [0090]).
Regarding claim 4 the combination of Li, Peake and Zhou teaches; The method of claim 3, wherein the second sub-model comprises:
a first feature extraction model configured to extract a visible light feature map from the visible light image of the input image set (Peake, [0084] the NN may access the color feature data in the color (visible light) image to help estimate color feature data of the stereo image pair, [0086] the machine learning model of Peake has multiple neural units (multiple sub models) that are used to generate the analysis outputs);
a second feature extraction model configured to extract an infrared feature map from the infrared image of the input image set (Peake, [0085] the neural network may access the multispectral image color information and use the NN to general an augmented/estimated set of color data (feature map), [0086] the machine learning model of Peake has multiple neural units (multiple sub models) that are used to generate the analysis outputs);
and a confidence score estimation model configured to estimate the confidence score map based on a correlation between the visible light feature map and the infrared feature map (Peake, [0084] the NN can determine color information in the pair of images based on the color and infrared feature data previously collected, [0086] the machine learning model of Peake has multiple neural units (multiple sub models) that are used to generate the analysis outputs).
The combination of Li and Peake would be obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. Given that both Li and Peake teach different methods and motivation for visible and infrared light image fusion and mapping of this information. The use of a neural network and its neural units to map color features of each of the images as taught in Peake would have improved the method of Li because the ability to map the features of the corresponding images using a neural network would allow for faster estimation of this data. (Peake, [0083]- [0090])
Regarding claim 5 the combination of Li, Peake and Zhou teaches; The method of claim 4, wherein the cross-attention operation comprises (Zhou, Figure 4, S21-S26 two images which are infrared and visible light are acquired, features are extracted from both to create two feature maps (feature maps) and then per S26 a matching result map is generated based on the two feature/attention maps (attention map), [0079] the feature matching and aggregation are performed using a cross-attention operation, which generates the attention map, [0081] a matching confidence matrix (confidence map) is generated based on the match result from the feature map matching operation (attention map))):
determining first data corresponding to one of a visible light feature map and an infrared feature map and second data corresponding to another one (Li, [0074] and [0075] the visible and infrared images are put through a Laplacian pyramid to generate a feature vector for each pixel in each image (first and second data respectively));
generating the attention map based on query data according to the first data and key data according to the second data (Li, [0074] and [0075], the feature vectors for each of the images (First and second data) are transformed using a transformation matrix (reshaping), the layers of the Laplacian pyramid (which include the reshaped feature vectors, which are the first and second data/query and key data respectively) are reconstructed to obtain fused layers, this method by definition uses the dot product of the layers, therefore the reshaped feature vectors are combined using a dot product operation to create a fused layer with feature information which is functionally equivalent to the attention map as taught by [0050] of applicant’s specification. [0050] of the applicant’s spec, notes that the query and key data are generated by putting the first and second data, respectively, through an operation later where the operation layer includes one of a set of operations including reshaping or convolutions, further the dot product of the key and query data generates the attention map);
and estimating the confidence score map based on value data according to one of the first data and the second data and the attention map (Li, [0023]- [0031] Gaussian and Laplace pyramids are constructed for the two images to get the base layer information, which is then used to generate the detail layer (confidence map), claim 3 of Li further illustrates this concept. The confidence map (detail layer) formation is based off of the fused base layer, and the formation of this is dependent on feature data from the infrared (value data) and visible light images that have been fused (formation of the attention map), Further, [0051] of applicant’s specification details that the value data is determined from the infrared feature map through an operation layer group).
The combination of Li, Peake and Zhou would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The combination of Li and Peake teach a system which performs fusion of two images, one infrared and one visible light, however it does not teach the use of a cross-attention operation to perform this. Zhou teaches this deficiency, where the use of a cross-attention operation would reasonably improve the system of Li and Peake by allowing for matching features from both inputs to generate matched global features, which would be effective for fusion two images of different modalities. (Zhou, [0044]- [0046])
Regarding claim 6 the combination of Li, Peake and Zhou teaches; The method of claim 1, wherein the infrared image comprises a multi-band (LI, [0032] the infrared image can be split into 4 dimensions of information (RGB and infrared), indicating that multiple frequencies/bands exist within the image).
Regarding claim 7 the combination of Li, Peake and Zhou teaches; The method of claim 1, wherein the determining of the illuminant information comprises:
obtaining a weighted sum between the illuminant map and the confidence score map (Li, [0040] – [0042] the detail layer (confidence map) is weighted and added to the basic layer (illuminant map));
and determining the illuminant information according to the weighted sum (Li, [0042] a final result (illuminant information) is obtained by summing the detail layer and the basic layer, where the detail layer is weighted (Li (0041])).
Regarding claim 8 the combination of Li, Peake and Zhou teaches; The method of claim 7, wherein the obtaining of the weighted sum comprises
obtaining the weighted sum by summing vectors for each local area according to the illuminant map using a weight according to the confidence score map (Li, [0040] –[0042] the detail layer (confidence map) is weighted and added to the basic layer (illuminant map), [0075]-[0079] the fused base layer image (Illuminant map) involves getting a feature vector for each pixel and mapping it to the constraint image, therefore a weighted summing of vectors is involved in the generation of the basic layer (illuminant map), further, weights for each pixel are determined in [0079]-[0081] of Li).
Regarding claim 9 the combination of Li, Peake and Zhou teaches; The method of claim 8, wherein: local areas of the visible light image and local areas of the infrared image form corresponding pairs (LI, Figure 2 shows the pairs of images, where the two images are the same scene and therefore all local areas correspond),
the confidence score map comprises weights for each of the corresponding pairs (Li, [0054] the detail layers of the infrared and visible light images are weighted prior to fusion to create one detail layer (Confidence map)),
and a weight of a corresponding pair corresponds to a correlation between visible light data and infrared data of the corresponding pair (Li, [0054] the detail layers of the infrared and visible light images are weighted prior to fusion to create one detail layer (Confidence map), [0036]-[0041] the detail layer is determined by determining the correlation between pixels between the visible and infrared light images and weighting the detail layers of the two images prior to fusion to create the fused detail layer (confidence map)).
Regarding claim 10 the combination of Li, Peake and Zhou teaches; A non-transitory computer-readable storage medium storing instructions that (Peake, [0107 the program is stored in non-transitory computer readable medium), when executed by a processor (Peake, [0091] the system has a processor for performing the method), cause the processor to perform the method of claim 1 (Peake, [0091] the system has a processor for performing the method).
The combination of Li, Peake and Zhou would be obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. Li, Peake and Zhou each teach similar methods of multi-spectral image fusion for image enhancement or correction. Li does not teach hardware for executing this method, however in the same field of endeavor Peake teaches a processor and non-transitory computer readable media for execution of an image fusion method which one of ordinary skill in the art would appreciate as being necessary for translation of the method into practical application. (Peake, [0091] and [0107])
Regarding claim 11, the combination of Li, Peake and Zhou teaches; An apparatus comprising: a processor (Peak, [0091] the system has a processor for performing the method);
and a memory configured to store instructions executable by the processor (Peake, [0099] the processor is coupled to a memory), wherein, in response to the instructions being executed by the processor, the processor is configured to (Peak, [0091] the system has a processor for performing the method):
form an input image set by combining a visible light image and an infrared image (Li, [0049] the method provided is based on the fusion of a visible light image and a near infrared image, in paragraph [0038] of applicant’s specification the visible light image is stated as using light in the visible range of 400mn-700nm and the infrared image is stated as using a near infrared (NIR) band of 700-1000 nm range);
PNG
media_image1.png
94
580
media_image1.png
Greyscale
(Li, [0049])
estimate an illuminant map representing an illuminant configuration of the input image set (Li, [0050] a weighted least squares filter is used as step (1) to obtain filtered base layer images [0067] a base layer is obtained in step (1) of the method which removes the high frequency information, and leaving the illumination information and color information)
generating, using a neural network model, (Zhou, [0033] neural networks are used to perform feature extraction operations and mappings), a confidence score map representing a correlation between the visible light image and the infrared image by performing a cross- attention operation between the visible light image and the infrared image to obtain an attention map (Zhou, Figure 4, S21-S26 two images which are infrared and visible light are acquired, features are extracted from both to create two feature maps (feature maps) and then per S26 a matching result map is generated based on the two feature/attention maps (attention map), [0079] the feature matching and aggregation are performed using a cross-attention operation, which generates the attention map, [0081] a matching confidence matrix (confidence map) is generated based on the match result from the feature map matching operation (attention map)), wherein the confidence score map is based on the attention map (Zhou, [0081] a matching confidence matrix (confidence map) is generated based on the match result from the feature map matching operation (attention map)),
and determine illuminant information of the visible light image based on the illuminant map and the confidence score map (Li [0042] and [0043] the fused basic layer (illuminant map) and the fused detail layer (confidence score map) are added to obtain a final result which takes into account the illumination information (from the base layers) and the confidence/correlation information (illuminant information), this is used to obtain a modified/fused image, Figure 15 of applicant’s specification details that the illuminant information is what is used to generate the modified image, so the combination of the two fused layers as in Li would be the equivalent to this information);
and process the visible light image using the illuminant information to obtain an output image (Peake, [0083]- [0086] channels in the stereo image pair are processed to yield and output image, which is denoted as the “augmented spectral image” and is functionally equivalent to an output image).
The combination of Li and Peake would be obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. Given that both Li and Peake teach different methods and motivation for visible and infrared light image fusion and mapping of this information, the addition of the use of a neural network as taught in Peake would be an improvement to the method of Li because the use of a neural network would improve speed and accuracy of the image fusion operations (See Peake, [0083]- [0090]). Further, the combination of Li, Peake and Zhou would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The combination of Li and Peake teach a system which performs fusion of two images, one infrared and one visible light, however it does not teach the use of a cross-attention operation to perform this. Zhou teaches this deficiency, where the use of a cross-attention operation would reasonably improve the system of Li and Peake by allowing for matching features from both inputs to generate matched global features, which would be effective for fusion two images of different modalities. (Zhou, [0044]- [0046])
Regarding claim 12, the combination of Li, Peake and Zhou teaches; The apparatus of claim 11, wherein: the illuminant map represents the illuminant configuration for each local area (Li, [0017]-[0020] the basic layers of the images (illuminant map) are generated by filtering the visible and infrared images using a weighted least squares formula, since this is being computed over the whole image it would also be computed for all local areas or pixels), and the confidence score map represents the correlation between the visible light image and the infrared image for the each local area (Li, [0035]-[0036] the generation of the fused detail layer (confidence map) includes calculation of local energies of the pixels (local areas) for each image and then correlating the two).
Regarding claim 13 the combination of Li, Peake and Zhou teaches; The apparatus of claim 11, wherein the neural network model comprises:
a first sub-model configured to estimate the illuminant map from the input image set (Peake, [0083]-[0085]] a machine learning model may be trained to obtain information from the stereo image pair including RGB and infrared information as well as stereo information for identification of objects in images, the applicant notes that illuminant information may have RGB information, [0086] the machine learning model of Peake has multiple neural units (first and second sub models) that are used to generate the analysis outputs);
and a second sub-model configured to estimate the confidence score map from the input image set (Peake, [0078] the confidence between the features of the pairs of images (visible and infrared) are validate for confidence for each feature label using the neural network, [0086] the machine learning model of Peake has multiple neural units (first and second sub models) that are used to generate the analysis outputs).
The combination of Li, Peake and Zhou would be obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. Given that Li, Peake and Zhou teach different methods and motivation for visible and infrared light image fusion and mapping of this information, the addition of the use of a neural network as taught in Peake would be an improvement to the method of Li because the use of a neural network would improve speed and accuracy of the image fusion operations (See Peake, [0083]- [0090]).
Regarding claim 14, the combination of Li, Peake and Zhou teaches; The apparatus of claim 13, wherein the second sub-model comprises: a first feature extraction model configured to extract a visible light feature map from the visible light image of the input image set (Peake, [0084] the NN may access the color feature data in the color (visible light) image to help estimate color feature data of the stereo image pair, [0086] the machine learning model of Peake has multiple neural units (multiple sub models) that are used to generate the analysis outputs);
a second feature extraction model configured to extract an infrared feature map from the infrared image of the input image set (Peake, [0085] the neural network may access the multispectral image color information and use the NN to general an augmented/estimated set of color data (feature map), [0086] the machine learning model of Peake has multiple neural units (multiple sub models) that are used to generate the analysis outputs);
and a confidence score estimation model configured to estimate the confidence score map based on a correlation between the visible light feature map and the infrared feature map (Peake, [0084] the NN can determine color information in the pair of images based on the color and infrared feature data previously collected, [0086] the machine learning model of Peake has multiple neural units (multiple sub models) that are used to generate the analysis outputs).
The combination of Li, Peake and Zhou would be obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. Given that Li, Peake and Zhou teach different methods and motivation for visible and infrared light image fusion and mapping of this information. The use of a neural network and its neural units to map color features of each of the images as taught in Peake would have improved the method of Li because the ability to map the features of the corresponding images using a neural network would allow for faster estimation of this data. (Peake, [0083]- [0090])
Regarding claim 15, the combination of Li, Peake and Zhou teaches; The apparatus of claim 14, wherein the cross- attention operation comprises (Zhou, Figure 4, S21-S26 two images which are infrared and visible light are acquired, features are extracted from both to create two feature maps (feature maps) and then per S26 a matching result map is generated based on the two feature/attention maps (attention map), [0079] the feature matching and aggregation are performed using a cross-attention operation, which generates the attention map, [0081] a matching confidence matrix (confidence map) is generated based on the match result from the feature map matching operation (attention map))):
determining first data corresponding to one of the visible light feature map and the infrared feature map and second data corresponding to another one (Li, [0074] and [0075] the visible and infrared images are put through a Laplacian pyramid to generate a feature vector for each pixel in each image (first and second data respectively));
generating the attention map based on query data according to the first data and key data according to the second data (Li, [0074] and [0075], the feature vectors for each of the images (First and second data) are transformed using a transformation matrix (reshaping), the layers of the Laplacian pyramid (which include the reshaped feature vectors, which are the first and second data/query and key data respectively) are reconstructed to obtain fused layers, this method by definition uses the dot product of the layers, therefore the reshaped feature vectors are combined using a dot product operation to create a fused layer with feature information which is functionally equivalent to the attention map as taught by [0050] of applicant’s specification. [0050] of the applicant’s spec, notes that the query and key data are generated by putting the first and second data, respectively, through an operation later where the operation layer includes one of a set of operations including reshaping or convolutions, further the dot product of the key and query data generates the attention map);
and estimating the confidence score map based on value data according to one of the first data and the second data and the attention map (Li, [0023] - [0031] Gaussian and Laplace pyramids are constructed for the two images to get the base layer information, which is then used to generate the detail layer (confidence map), claim 3 of Li further illustrates this concept. The confidence map (detail layer) formation is based off of the fused base layer, and the formation of this is dependent on feature data from the infrared (value data) and visible light images that have been fused (formation of the attention map), Further, [0051] of applicant’s specification details that the value data is determined from the infrared feature map through an operation layer group).
The combination of Li, Peake and Zhou would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The combination of Li and Peake teach a system which performs fusion of two images, one infrared and one visible light, however it does not teach the use of a cross-attention operation to perform this. Zhou teaches this deficiency, where the use of a cross-attention operation would reasonably improve the system of Li and Peake by allowing for matching features from both inputs to generate matched global features, which would be effective for fusion two images of different modalities. (Zhou, [0044]- [0046])
Regarding claim 16, the combination of Li, Peake and Zhou teaches; The apparatus of claim 11, wherein the infrared image comprises a multi-band (Li, [0032] the infrared image can be split into 4 dimensions of information (RGB and infrared), indicating that multiple frequencies/bands exist within the image).
Regarding claim 17, the combination of Li, Peake and Zhou teaches; The apparatus of claim 11, wherein, to determine the illuminant information, the processor is configured to:
obtain a weighted sum by summing vectors for each local area according to the illuminant map using a weight according to the confidence score map (Li, [0040] –[0042] the detail layer (confidence map) is weighted and added to the basic layer (illuminant map), [0075]-[0079] the fused base layer image (Illuminant map) involves getting a feature vector for each pixel and mapping it to the constraint image, therefore a weighted summing of vectors is involved in the generation of the basic layer (illuminant map), further, weights for each pixel are determined in [0079]-[0081] of Li);
and determining the illuminant information according to the weighted sum (Li, [0042] a final result (illuminant information) is obtained by summing the detail layer and the basic layer, where the detail layer is weighted (Li (0041]))
Regarding claim 18, the combination of Li, Peake and Zhou teaches; An electronic device comprising:
a visible light camera configured to generate a visible light image (Peake, [0039] the system includes a color camera);
an infrared camera configured to generate an infrared image (Peake, [0039] the system includes an IR camera;
and a processor configured to (Peak, [0091] the system has a processor for performing the method):
form an input image set by combining a visible light image and an infrared image (Li, [0049] the method provided is based on the fusion of a visible light image and a near infrared image, in paragraph [0038] of applicant’s specification the visible light image is stated as using light in the visible range of 400mn-700nm and the infrared image is stated as using a near infrared (NIR) band of 700-1000 nm range);
PNG
media_image1.png
94
580
media_image1.png
Greyscale
(Li, [0049])
estimate an illuminant map representing an illuminant configuration of the input image set (Li, [0050] a weighted least squares filter is used as step (1) to obtain filtered base layer images [0067] a base layer is obtained in step (1) of the method which removes the high frequency information, and leaving the illumination information and color information)
generating, using a neural network model (Zhou, [0033] neural networks are used to perform feature extraction operations and mappings), a confidence score map representing a correlation between the visible light image and the infrared image by performing a cross- attention operation between the visible light image and the infrared image to obtain an attention map (Zhou, Figure 4, S21-S26 two images which are infrared and visible light are acquired, features are extracted from both to create two feature maps (feature maps) and then per S26 a matching result map is generated based on the two feature/attention maps (attention map), [0079] the feature matching and aggregation are performed using a cross-attention operation, which generates the attention map, [0081] a matching confidence matrix (confidence map) is generated based on the match result from the feature map matching operation (attention map)), wherein the confidence score map is based on the attention map (Zhou, [0081] a matching confidence matrix (confidence map) is generated based on the match result from the feature map matching operation (attention map)),
and(Li [0042] and [0043] the fused basic layer (illuminant map) and the fused detail layer (confidence score map) are added to obtain a final result which takes into account the illumination information (from the base layers) and the confidence/correlation information (illuminant information), this is used to obtain a modified/fused image, Figure 15 of applicant’s specification details that the illuminant information is what is used to generate the modified image, so the combination of the two fused layers as in Li would be the equivalent to this information);
and process the visible light image using the illuminant information to obtain an output image (Peake, [0083]- [0086] channels in the stereo image pair are processed to yield and output image, which is denoted as the “augmented spectral image” and is functionally equivalent to an output image).
The combination of Li and Peake would be obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. Given that both Li and Peake teach different methods and motivation for visible and infrared light image fusion and mapping of this information, the addition of the use of a neural network as taught in Peake would be an improvement to the method of Li because the use of a neural network would improve speed and accuracy of the image fusion operations (See Peake, [0083]- [0090]). Further, the combination of Li, Peake and Zhou would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The combination of Li and Peake teach a system which performs fusion of two images, one infrared and one visible light, however it does not teach the use of a cross-attention operation to perform this. Zhou teaches this deficiency, where the use of a cross-attention operation would reasonably improve the system of Li and Peake by allowing for matching features from both inputs to generate matched global features, which would be effective for fusion two images of different modalities. (Zhou, [0044]- [0046])
Regarding claim 19, the combination of Li, Peake and Zhou teaches; The electronic device of claim 18, wherein; the illuminant map represents the illuminant configuration for each local area (Li, [0017]-[0020] the basic layers of the images (illuminant map) are generated by filtering the visible and infrared images using a weighted least squares formula, since this is being computed over the whole image it would also be computed for all local areas or pixels), and the confidence score map represents the correlation between the visible light image and the infrared image for the each local area (Li, [0035]-[0036] the generation of the fused detail layer (confidence map) includes calculation of local energies of the pixels (local areas) for each image and then correlating the two).
Regarding claim 20, the combination of Li, Peake and Zhou teaches; The electronic device of claim 18, wherein the neural network model comprises: a first sub-model configured to estimate the illuminant map from the input image set (Peake, [0083]-[0085]] a machine learning model may be trained to obtain information from the stereo image pair including RGB and infrared information as well as stereo information for identification of objects in images, the applicant notes that illuminant information may have RGB information, [0086] the machine learning model of Peake has multiple neural units (first and second sub models) that are used to generate the analysis outputs);
and a second sub-model configured to estimate the confidence score map from the input image set (Peake, [0078] the confidence between the features of the pairs of images (visible and infrared) are validate for confidence for each feature label using the neural network, [0086] the machine learning model of Peake has multiple neural units (first and second sub models) that are used to generate the analysis outputs).
The combination of Li, Peake and Zhou would be obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. Given that Li, Peake, and Zhou teach different methods and motivation for visible and infrared light image fusion and mapping of this information, the addition of the use of a neural network as taught in Peake would be an improvement to the method of Li because the use of a neural network would improve speed and accuracy of the image fusion operations (See Peake, [0083]- [0090]).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. For a listing of analogous art as determined by the examiner please see the attached PTO-892 Notice of References cited form.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JORDAN M ELLIOTT whose telephone number is (703)756-5463. The examiner can normally be reached M-F 8AM-5PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emily Terrell can be reached at (571) 270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.M.E./Examiner, Art Unit 2666
/EMILY C TERRELL/Supervisory Patent Examiner, Art Unit 2666