Prosecution Insights
Last updated: October 01, 2026
Application No. 19/002,223

TRAINING METHOD FOR IMAGE PROCESSING NETWORK, AND IMAGE PROCESSING METHOD AND APPARATUS

Non-Final OA §103
Filed
Dec 26, 2024
Priority
Jun 30, 2022 — CN 202210772860.1 +1 more
Examiner
YANG, WEI WEN
Art Unit
Tech Center
Assignee
Honda Motor Co., Ltd.
OA Round
1 (Non-Final)
82%
Grant Probability
Favorable
1-2
OA Rounds
8m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
560 granted / 684 resolved
+21.9% vs TC avg
Moderate +12% lift
Without
With
+11.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
31 currently pending
Career history
705
Total Applications
across all art units

Statute-Specific Performance

§101
7.8%
-32.2% vs TC avg
§103
75.0%
+35.0% vs TC avg
§102
9.3%
-30.7% vs TC avg
§112
7.8%
-32.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 684 resolved cases

Office Action

§103
DETAILED ACTION Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 8-13, 15, 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over GUPTA (US 20220383663 A1), in view of SARGENT (EP 3614308 A). Re Claim 1, GUPTA discloses a method for training an image processing network performed by an electronic device (see GUPTA: e.g., --[0017] FIG. 7 shows the process of training the neural network in the identification operation--, and, --[0022] FIG. 12 shows the process of training the neural network used in the liveliness-detection operation--, and, --[0029] FIG. 19 presents a flowchart of the training of the neural network used in the cropping operation--, in [0017]-[0029]; and, --[0065] The neural network, according to some embodiments, is trained and/or otherwise adapted to be able to distinguish, through processing of the image, between those portions of the image that contain the region of interest and other portions of the image that do not contain the region of interest. This can be achieved in variety of ways and is thus not to be understood in a limiting way. That the neural network “distinguishes” that region comprising the ROI from another region is to be understood as the ability of the neural network to provide an output that distinguishes between the ROI and other regions of the image or makes it possible to distinguish between those regions. For example, the output could be an indication of pixels comprising the ROI but not other regions of the ROI. In any case, the outcome of the processing of the image by the neural network at least comprises that a first portion of the image comprising the region of interest is identified as different from another portion of the image. In this regard, it is noted that the specific size, shape of position of the region of interest is found out by the neural network during processing of the image and it is not preset.--, in [0065]) and comprising: determining a reference pixel based on a training image annotated with a truth value (see Gupta: e.g., -- reference will usually be made to “the image” or “the input image” or the “original image”. In view of the foregoing, it is clear that this does not only comprise the full image or the original image obtained by the optical sensor but also any realization of the pre-processing, including using, instead of the full image, only a part of the image or using only one or more images comprising one colour value or being restricted to brightness values for the respective pixels. Any of these pre-processings and any other pre-processing that can be thought of will thus be considered included when the further processing of the original image is described.--, in [0465] {herein reference of “one colour value or being restricted to brightness values for the respective pixels” of the “original image” read on claimed limitation of “truth value”}; and, --[0543] The resulting matrix can be considered “black and white” image where the entries in the matrix having a value x=0 might be considered to be white and the entries and the resulting matrix having values x=1 may be considered black. The other way around is also possible and the reference to the “black and white” picture is only for exemplary purpose. [0544] Due to the processing of the original image by the neural network, this will result in the region of interest being visible in the output matrix or output decoded image as having a specific shape for example an elliptical shape. This is because, due to the learned neural network and the processing of the input image, the ROI either corresponds to the values x=1 or x=0. The rest of the image will be faded out (corresponding to have the other value x, respectively) which then allows to distinguish between the regions of interest and other portions or parts of the image. --, in [0543]-[0544] {herein “the ROI, such as “pixel of fingertip”, either corresponds to the values x=1 or x=0” read on reference pixel of “truth value”}); determining, with the reference pixel as a starting point, and cropping probabilities of the image processing network processing the training image (see Gupta: e.g., -- [0436] The method of FIG. 12 begins with the provision of training data 1201. The training data may be constituted by a plurality of images of real objects as well as images of spoofs of real objects. For example, the images may comprise a number of images of real hands or fingers or the like and images of images (i.e. spoofs) of those objects. The images may be multiplied by using, from the same image, rotated, highlighted, darkened, enlarged or otherwise modified copies that are introduced as training data. In some embodiments, modifications involving image flips, image rotation and translation, shears, crops, multiplication to increase brightness and Gaussian blurs may be used to obtain a larger number of training images. Arbitrary combinations of the mentioned techniques may also be used. The values σ.sub.1 and σ.sub.2 provided according to item 1208 are the values indicating the “correct” output of the first node and second node of the last layer in the neural network that provide the probability of the image showing a spoof of an object or a real object. These values are provided for each image in the training data.--, in [0436]; and, -- specifically extracting a region of interest from the original image obtained and processed during the steps described above and before data comprising the biometric characteristic is sent to the third party computing device. In fact, with the embodiment now described, one embodiment of obtaining image data that only comprises the region of interest, like the fingertip carrying the fingerprint of interest, is described. [0448] This embodiment focuses on the extraction of a region of interest (ROI) from an image taken from an object of the user, where the image comprises a biometric characteristic that could be used to identify the user, but could also be used for any other purpose on the third party computing device. Such an object can be anything like a fingertip of one or more fingers of a hand of the user, the palm creases of a hand of the user or the face of the user or even the eye or the foot of the user. Each of these entities is known to carry biometric characteristics that can be used for identifying a user. [0449]… as a “cropping step” which separates a part of an image from another part of the image. This method is performed on the image that was already obtained during step 1 of the method according to FIG. 1. The step may occur after the identification step 2 and after the liveliness-detection and comparison steps 3 and 4. However, it can also occur in parallel to steps 3 and 4 or immediately after step 2. In some embodiments, it does not occur before step 2 since, in step 2 of FIG. 1, it is just determined whether there is an object carrying a biometric characteristic in the image at all. Processing the image in line with the embodiment to be described now would thus not be reasonable and potentially a waste of resources if it were performed before step 2 in FIG. 1.--, in [0447]-[0449], [0453]-[0457]; and, -- [0327] As explained above, it is assumed that the neural network is already perfectly learned for identifying a specific object, such as a fingertip. This involves that the neural network is able to identify a specific pattern of pixels that are most likely representing a fingertip. This might refer to specific patterns of color values or other characteristics like the brightness of those spots. It is, however, clear that the image 300 may arbitrarily show a fingertip which might not correspond in size and arrangement to a fingertip that was used for learning the neural network. [0328] With the help of the bounding boxes and the grid, however, it is possible for the neural network to identify the specific bounding box that will most likely comprise the fingertip. In order to identify this specific bounding box, the neural network (or an associated component that processes the image 300) compares the values of the pixels within each bounding box of each grid cell to a pattern of pixels that corresponds to a fingertip as was previously learned by the neural network. In this first stage, it is most unlikely that a perfect match will be found but there will be bounding boxes that are already more likely to contain at least a portion of a fingertip than other bounding boxes. [0329]…depicted in FIG. 6, for example, the bounding box 341 centered around the point M in grid cell 313 includes a portion of the fingertip of the hand 350. In contrast to this, none of the grid cells 310 and 311 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 341 and potentially the bounding box 340, the process can determine that the bounding box 341 includes even more of a pattern that corresponds to a fingertip than the bounding box 340. [0330] In view of this, the method can conclude that none of the bounding boxes 331 and 332 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0331] As both bounding boxes 340 and 341 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0332] While the smaller grid cell 340 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 341 may be regarded by the process to include a pattern that corresponds to a fingertip. [0333] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 341 and 340 to a pattern obtained from learning which indeed corresponds to a fingertip. [0334] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 341 is used as the starting point and its position and shapes modified or the smaller bounding box 340 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern….[0337] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0338] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 300 that contain the respective fingertip.--, in [0329]-[0338], and, --depicted in FIG. 18, for example, the bounding box 1841 centered around the point M in grid cell 1813 includes a portion of the fingertip of the hand 1850. In contrast to this, none of the grid cells 1810 and 1811 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 1841 and potentially the bounding box 1840, the process can determine that the bounding box 1841 includes even more of a pattern that corresponds to a fingertip than the bounding box 1840. [0565] In view of this, the method can conclude that none of the bounding boxes 1831 and 1832 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0566] As both bounding boxes 1840 and 1841 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0567] While the smaller grid cell 1840 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 1841 may be regarded by the process to include a pattern that corresponds to a fingertip. [0568] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 1841 and 1840 to a pattern obtained from learning which indeed corresponds to a fingertip. [0569] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 1841 is used as the starting point and its position and shapes modified or the smaller bounding box 1840 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern. [0570] This minimizing process can firstly comprise modifying the position of the bounding box (in the following, it will be assumed that the bounding box 1841 is used for the further calculations) by moving it a small amount into orthogonal directions first along the x-axis and then along the y-axis (or vice versa) as depicted in FIG. 18 around the center point M of the respective grid cell. The movement will be along the positive and the negative x-axis and y-axis and at each position, a comparison will be made to determine a difference function between the pattern obtained from the learning and the actual pattern identified in the image. This allows for calculating a two-dimensional function that represents the difference d(x,y) depending on the coordinates…..[0572] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0573] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 1800 that contain the respective fingertip or object carrying the biometric characteristic.--, in [0564]-[0573]; also see: -- [0071] It may be that the cropping operation comprises processing the image or the part of the image by a neural network, wherein processing the image or the part of the image by the neural network comprises processing the image by an encoder to obtain an encoded image and, after that, processing the encoded image by a decoder to obtain a decoded output image. [0072] Furthermore, the image or the part of the image provided to the neural network in the cropping operation for processing may comprise N×M pixels and the encoded image comprises n×m pixels, where n<N,m<M and the decoded output image comprises N×M pixels. [0073] Reducing the number of pixels when encoding the image results in a loss of information. When enlarging the image once again during the decoding, however, the most relevant information to distinguish the ROI from other portions of the image can be more easily discernable as not relevant information or very detailed information that is not necessary for identifying ROI is faded out with this procedure. [0074] In this regard, it can also be provided that distinguishing a portion of the image or the part of the image comprises distinguishing a portion of the decoded output image from another portion of the decoded output image. The distinguishing of the portions in the decoded image can be much easier compared to distinguishing the portion comprising the ROI from another portion of the original image. Thus, the processing power required for distinguishing a portion in the decoded output image from another portion in the decoded output image are reduced significantly compared to processing the original input image. [0075] More specifically, extracting the portion comprising the biometric characteristic can comprise identifying pixel in the decoded output image that are within the distinguished portion and, after that, identifying the pixels in the decoded output image that are in the distinguished portion with corresponding pixels in the original image or the part of the image and extracting, from the original image or the part of the image, the corresponding pixels, the extracted corresponding pixels constituting the portion of the image or the part of the image comprising the biometric characteristic.--, in [0071]-[0075]; and, -- [0101] Separating the obtained image into grid cells with predefined bounding boxes allows for properly displaying and providing feedback on objects identified by using the bounding boxes in the final result to mark the location of the object and the object itself. [0102] In a more specific realization of this embodiment, the position of the bounding box is calculated relative to a center of the grid cell in two dimensions and the geometrical characteristics of the bounding box comprise a height and a width of the bounding box, wherein, further, a probability of the object being within the bounding box is associated with each bounding box. [0103] Associating the bounding boxes with corresponding probabilities allows for providing a matrix or vector that represents the bounding box and can be handled by graphical processing units with accurate efficiency when having to combine this with other objects that are represented in the form of a matrix or vector. Thereby, the required computer resources are reduced even further…. [0105] The resulting tensor can be processed by graphic processing units in a highly efficient manner. Additionally, providing the identification result in the form of such a tensor allows for easily deducing the results having the greatest probability for identifying a specific object. [0106] Moreover, outputting the output may comprise displaying the image and the resulting bounding boxes in each grid cell that have the highest probability among the bounding boxes in the grid cell.--, in [0102]-[0107]); GUPTA however does not explicitly disclose {about “to identify this specific bounding box” based on the “identifying pixels that are most likely representing a fingertip” as of “cropping probabilities”}, based on a Markov chain of the training image, SARGENT discloses {determining cropping probabilities in “to identify this specific bounding box”} based on a Markov chain of the training image (see SARGENT: e.g., -- Determining the object convolutional positions may include finding a moment bounding box of each object, the moment bounding box being a minimum bounding rectangle surrounding an object and aligned with the orientation of the major axis of the object, the object convolutional positions being determined in dependence on properties of the bounding box. Properties of the moment bounding box may be used to determine, for an object, a single first object convolutional position which should be used to represent the object when being processed by the first neural network, and further to select, for the object, a plurality of second object convolutional positions distributed across the bounding box which should be used to represent the object when being processed by the second neural network…. Figure 13 displays joint deep learning with joint distribution modelling (a) through iterative process for pixel-level cover (LC) and patch-based land use (LU) extraction and decision-making (b); and Figure 14 is a flow diagram of the training of the JDL model. --, in [0019]-[0020], and, [0026]; and, -- The image of land is passed into a first machine learning network which is used to calculate LC probabilities on a per-pixel basis across the image. The LC probabilities are then fed into a second machine learning network to calculate LU probabilities for each identified object. The LU probability is then fed back into the first machine learning network, along with the original image, and the process using the first and second machine learning networks is iterated over. This iteration process produces a highly accurate LC classification for each pixel in the image, and a LU classification for each identified object. In more detail, the described method simultaneously determines land cover (LC) and land use (LU) classifications from remotely sensed images using two machine learning networks, a multilayer perceptron (MLP) and a convolutional neural network (CNN), which together form a joint deep learning (JDL) model. The JDL model is implemented via a Markov process involving iterating between the two machine learning networks, with the output of one of the networks being fed as an input for the next iteration of the other of the networks, and vice versa. The CNN is trained with labelled image patches for the LU classes.--, in [0027]-[0028]; and, -- The assumption of the LC - LU joint deep learning (LC-LU JDL) model is that both LC and LU are manifested over same geographical space and are nested with each other in a hierarchical manner. The LC and LU representations are considered as two random variables, where the probabilistic relationship between them can be modelled through a joint probability distribution. In this way, the conditional dependencies between these two random variables are captured via an undirected graph through iteration (i.e. formulating a Markov process). The joint distribution is, thus, factorised as a product of the individual density functions,--, in [0044]; and, -- Figure 3 illustrates the general workflow of the described LC and LU joint deep learning (LC-LU JDL) model, with key components including the JDL inputs, the Markov Process to learn the joint distribution, and the classification outputs of LC and LU at each iteration. Detailed explanation is given as follows. JDL input involves LC samples with pixel locations and the corresponding land cover labels, LU samples with image patches representing specific land use categories, together with the remotely sensed imagery, and the object-based segmentation results with unique identity for each segment. These four elements were used to infer the hierarchical relationships between LC and LU, and to obtain LC and LU classification results through iteration. Markov Process models the joint probability distribution between LC and LU through iteration--, in [0049]-[0051]); GUPTA and SARGENT are combinable as they are in the same field of endeavor: training image processing neural network for object recognition, segmentation, and bounding box identification. Therefore it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify GUPTA’s method using SARGENT’s teachings by including {determining cropping probabilities in “to identify this specific bounding box of object”} based on a Markov chain of the training image to GUPTA’s identifying and segmentation, bounding box of object in order to (see SARGENT: e.g. in [0019]-[0020], [0026], [0044], and [0049]-[0051]); GUPTA as modified by SARGENT further disclose adjusting, based on an output result obtained by the image processing network processing a training cropped area and based on the truth value, a network parameter value and the cropping probabilities of the image processing network to obtain a trained image processing network, wherein the training cropped area is obtained by cropping the training image based on the cropping probabilities (see Gupta: e.g., --[0029] FIG. 19 presents a flowchart of the training of the neural network used in the cropping operation--, in [0029]; and, --[0162] The kernel allows for properly weighing information obtained from adjacent pixels in introduced matrix while not losing any information, thereby increasing the efficiency with which consecutive layers in the neural network can support the processing in order to determine a spoof or real object. For this, the kernel comprises entries that correspond to specific weights or parameters that are obtained prior to receiving the image, i.e. during training of the neural network. [0163] It is a finding of the present invention that, in case this training is performed before the mobile device is actually equipped with an application or other program that can perform the respective method according to the above embodiments, the required computer resources can be advantageously reduced on the mobile device.--, in [0162]-[0163]; and, -- [0436] The method of FIG. 12 begins with the provision of training data 1201. The training data may be constituted by a plurality of images of real objects as well as images of spoofs of real objects. For example, the images may comprise a number of images of real hands or fingers or the like and images of images (i.e. spoofs) of those objects. The images may be multiplied by using, from the same image, rotated, highlighted, darkened, enlarged or otherwise modified copies that are introduced as training data. In some embodiments, modifications involving image flips, image rotation and translation, shears, crops, multiplication to increase brightness and Gaussian blurs may be used to obtain a larger number of training images. Arbitrary combinations of the mentioned techniques may also be used. The values σ.sub.1 and σ.sub.2 provided according to item 1208 are the values indicating the “correct” output of the first node and second node of the last layer in the neural network that provide the probability of the image showing a spoof of an object or a real object. These values are provided for each image in the training data.--, in [0436]; and, -- specifically extracting a region of interest from the original image obtained and processed during the steps described above and before data comprising the biometric characteristic is sent to the third party computing device. In fact, with the embodiment now described, one embodiment of obtaining image data that only comprises the region of interest, like the fingertip carrying the fingerprint of interest, is described. [0448] This embodiment focuses on the extraction of a region of interest (ROI) from an image taken from an object of the user, where the image comprises a biometric characteristic that could be used to identify the user, but could also be used for any other purpose on the third party computing device. Such an object can be anything like a fingertip of one or more fingers of a hand of the user, the palm creases of a hand of the user or the face of the user or even the eye or the foot of the user. Each of these entities is known to carry biometric characteristics that can be used for identifying a user. [0449]… as a “cropping step” which separates a part of an image from another part of the image. This method is performed on the image that was already obtained during step 1 of the method according to FIG. 1. The step may occur after the identification step 2 and after the liveliness-detection and comparison steps 3 and 4. However, it can also occur in parallel to steps 3 and 4 or immediately after step 2. In some embodiments, it does not occur before step 2 since, in step 2 of FIG. 1, it is just determined whether there is an object carrying a biometric characteristic in the image at all. Processing the image in line with the embodiment to be described now would thus not be reasonable and potentially a waste of resources if it were performed before step 2 in FIG. 1.--, in [0447]-[0449], [0453]-[0457]; and, -- [0327] As explained above, it is assumed that the neural network is already perfectly learned for identifying a specific object, such as a fingertip. This involves that the neural network is able to identify a specific pattern of pixels that are most likely representing a fingertip. This might refer to specific patterns of color values or other characteristics like the brightness of those spots. It is, however, clear that the image 300 may arbitrarily show a fingertip which might not correspond in size and arrangement to a fingertip that was used for learning the neural network. [0328] With the help of the bounding boxes and the grid, however, it is possible for the neural network to identify the specific bounding box that will most likely comprise the fingertip. In order to identify this specific bounding box, the neural network (or an associated component that processes the image 300) compares the values of the pixels within each bounding box of each grid cell to a pattern of pixels that corresponds to a fingertip as was previously learned by the neural network. In this first stage, it is most unlikely that a perfect match will be found but there will be bounding boxes that are already more likely to contain at least a portion of a fingertip than other bounding boxes. [0329]…depicted in FIG. 6, for example, the bounding box 341 centered around the point M in grid cell 313 includes a portion of the fingertip of the hand 350. In contrast to this, none of the grid cells 310 and 311 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 341 and potentially the bounding box 340, the process can determine that the bounding box 341 includes even more of a pattern that corresponds to a fingertip than the bounding box 340. [0330] In view of this, the method can conclude that none of the bounding boxes 331 and 332 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0331] As both bounding boxes 340 and 341 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0332] While the smaller grid cell 340 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 341 may be regarded by the process to include a pattern that corresponds to a fingertip. [0333] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 341 and 340 to a pattern obtained from learning which indeed corresponds to a fingertip. [0334] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 341 is used as the starting point and its position and shapes modified or the smaller bounding box 340 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern….[0337] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0338] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 300 that contain the respective fingertip.--, in [0329]-[0338], and, --depicted in FIG. 18, for example, the bounding box 1841 centered around the point M in grid cell 1813 includes a portion of the fingertip of the hand 1850. In contrast to this, none of the grid cells 1810 and 1811 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 1841 and potentially the bounding box 1840, the process can determine that the bounding box 1841 includes even more of a pattern that corresponds to a fingertip than the bounding box 1840. [0565] In view of this, the method can conclude that none of the bounding boxes 1831 and 1832 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0566] As both bounding boxes 1840 and 1841 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0567] While the smaller grid cell 1840 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 1841 may be regarded by the process to include a pattern that corresponds to a fingertip. [0568] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 1841 and 1840 to a pattern obtained from learning which indeed corresponds to a fingertip. [0569] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 1841 is used as the starting point and its position and shapes modified or the smaller bounding box 1840 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern. [0570] This minimizing process can firstly comprise modifying the position of the bounding box (in the following, it will be assumed that the bounding box 1841 is used for the further calculations) by moving it a small amount into orthogonal directions first along the x-axis and then along the y-axis (or vice versa) as depicted in FIG. 18 around the center point M of the respective grid cell. The movement will be along the positive and the negative x-axis and y-axis and at each position, a comparison will be made to determine a difference function between the pattern obtained from the learning and the actual pattern identified in the image. This allows for calculating a two-dimensional function that represents the difference d(x,y) depending on the coordinates…..[0572] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0573] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 1800 that contain the respective fingertip or object carrying the biometric characteristic.--, in [0564]-[0573]; also see: -- [0071] It may be that the cropping operation comprises processing the image or the part of the image by a neural network, wherein processing the image or the part of the image by the neural network comprises processing the image by an encoder to obtain an encoded image and, after that, processing the encoded image by a decoder to obtain a decoded output image. [0072] Furthermore, the image or the part of the image provided to the neural network in the cropping operation for processing may comprise N×M pixels and the encoded image comprises n×m pixels, where n<N,m<M and the decoded output image comprises N×M pixels. [0073] Reducing the number of pixels when encoding the image results in a loss of information. When enlarging the image once again during the decoding, however, the most relevant information to distinguish the ROI from other portions of the image can be more easily discernable as not relevant information or very detailed information that is not necessary for identifying ROI is faded out with this procedure. [0074] In this regard, it can also be provided that distinguishing a portion of the image or the part of the image comprises distinguishing a portion of the decoded output image from another portion of the decoded output image. The distinguishing of the portions in the decoded image can be much easier compared to distinguishing the portion comprising the ROI from another portion of the original image. Thus, the processing power required for distinguishing a portion in the decoded output image from another portion in the decoded output image are reduced significantly compared to processing the original input image. [0075] More specifically, extracting the portion comprising the biometric characteristic can comprise identifying pixel in the decoded output image that are within the distinguished portion and, after that, identifying the pixels in the decoded output image that are in the distinguished portion with corresponding pixels in the original image or the part of the image and extracting, from the original image or the part of the image, the corresponding pixels, the extracted corresponding pixels constituting the portion of the image or the part of the image comprising the biometric characteristic.--, in [0071]-[0075]; and, -- [0101] Separating the obtained image into grid cells with predefined bounding boxes allows for properly displaying and providing feedback on objects identified by using the bounding boxes in the final result to mark the location of the object and the object itself. [0102] In a more specific realization of this embodiment, the position of the bounding box is calculated relative to a center of the grid cell in two dimensions and the geometrical characteristics of the bounding box comprise a height and a width of the bounding box, wherein, further, a probability of the object being within the bounding box is associated with each bounding box. [0103] Associating the bounding boxes with corresponding probabilities allows for providing a matrix or vector that represents the bounding box and can be handled by graphical processing units with accurate efficiency when having to combine this with other objects that are represented in the form of a matrix or vector. Thereby, the required computer resources are reduced even further…. [0105] The resulting tensor can be processed by graphic processing units in a highly efficient manner. Additionally, providing the identification result in the form of such a tensor allows for easily deducing the results having the greatest probability for identifying a specific object. [0106] Moreover, outputting the output may comprise displaying the image and the resulting bounding boxes in each grid cell that have the highest probability among the bounding boxes in the grid cell.--, in [0102]-[0107]). Re Claim 8, GUPTA as modified by SARGENT and ZHANG further disclose wherein after the training cropped area is obtained by cropping the training image based on the cropping probabilities, the method further comprises: determining, in the training image, a line comprising a center pixel of the training image (see Gupta: e.g., -- [0436] The method of FIG. 12 begins with the provision of training data 1201. The training data may be constituted by a plurality of images of real objects as well as images of spoofs of real objects. For example, the images may comprise a number of images of real hands or fingers or the like and images of images (i.e. spoofs) of those objects. The images may be multiplied by using, from the same image, rotated, highlighted, darkened, enlarged or otherwise modified copies that are introduced as training data. In some embodiments, modifications involving image flips, image rotation and translation, shears, crops, multiplication to increase brightness and Gaussian blurs may be used to obtain a larger number of training images. Arbitrary combinations of the mentioned techniques may also be used. The values σ.sub.1 and σ.sub.2 provided according to item 1208 are the values indicating the “correct” output of the first node and second node of the last layer in the neural network that provide the probability of the image showing a spoof of an object or a real object. These values are provided for each image in the training data.--, in [0436]; and, -- specifically extracting a region of interest from the original image obtained and processed during the steps described above and before data comprising the biometric characteristic is sent to the third party computing device. In fact, with the embodiment now described, one embodiment of obtaining image data that only comprises the region of interest, like the fingertip carrying the fingerprint of interest, is described. [0448] This embodiment focuses on the extraction of a region of interest (ROI) from an image taken from an object of the user, where the image comprises a biometric characteristic that could be used to identify the user, but could also be used for any other purpose on the third party computing device. Such an object can be anything like a fingertip of one or more fingers of a hand of the user, the palm creases of a hand of the user or the face of the user or even the eye or the foot of the user. Each of these entities is known to carry biometric characteristics that can be used for identifying a user. [0449]… as a “cropping step” which separates a part of an image from another part of the image. This method is performed on the image that was already obtained during step 1 of the method according to FIG. 1. The step may occur after the identification step 2 and after the liveliness-detection and comparison steps 3 and 4. However, it can also occur in parallel to steps 3 and 4 or immediately after step 2. In some embodiments, it does not occur before step 2 since, in step 2 of FIG. 1, it is just determined whether there is an object carrying a biometric characteristic in the image at all. Processing the image in line with the embodiment to be described now would thus not be reasonable and potentially a waste of resources if it were performed before step 2 in FIG. 1.--, in [0447]-[0449], [0453]-[0457]; and, -- [0327] As explained above, it is assumed that the neural network is already perfectly learned for identifying a specific object, such as a fingertip. This involves that the neural network is able to identify a specific pattern of pixels that are most likely representing a fingertip. This might refer to specific patterns of color values or other characteristics like the brightness of those spots. It is, however, clear that the image 300 may arbitrarily show a fingertip which might not correspond in size and arrangement to a fingertip that was used for learning the neural network. [0328] With the help of the bounding boxes and the grid, however, it is possible for the neural network to identify the specific bounding box that will most likely comprise the fingertip. In order to identify this specific bounding box, the neural network (or an associated component that processes the image 300) compares the values of the pixels within each bounding box of each grid cell to a pattern of pixels that corresponds to a fingertip as was previously learned by the neural network. In this first stage, it is most unlikely that a perfect match will be found but there will be bounding boxes that are already more likely to contain at least a portion of a fingertip than other bounding boxes. [0329]…depicted in FIG. 6, for example, the bounding box 341 centered around the point M in grid cell 313 includes a portion of the fingertip of the hand 350. In contrast to this, none of the grid cells 310 and 311 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 341 and potentially the bounding box 340, the process can determine that the bounding box 341 includes even more of a pattern that corresponds to a fingertip than the bounding box 340. [0330] In view of this, the method can conclude that none of the bounding boxes 331 and 332 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0331] As both bounding boxes 340 and 341 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0332] While the smaller grid cell 340 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 341 may be regarded by the process to include a pattern that corresponds to a fingertip. [0333] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 341 and 340 to a pattern obtained from learning which indeed corresponds to a fingertip. [0334] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 341 is used as the starting point and its position and shapes modified or the smaller bounding box 340 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern….[0337] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0338] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 300 that contain the respective fingertip.--, in [0329]-[0338], and, --depicted in FIG. 18, for example, the bounding box 1841 centered around the point M in grid cell 1813 includes a portion of the fingertip of the hand 1850. In contrast to this, none of the grid cells 1810 and 1811 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 1841 and potentially the bounding box 1840, the process can determine that the bounding box 1841 includes even more of a pattern that corresponds to a fingertip than the bounding box 1840. [0565] In view of this, the method can conclude that none of the bounding boxes 1831 and 1832 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0566] As both bounding boxes 1840 and 1841 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0567] While the smaller grid cell 1840 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 1841 may be regarded by the process to include a pattern that corresponds to a fingertip. [0568] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 1841 and 1840 to a pattern obtained from learning which indeed corresponds to a fingertip. [0569] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 1841 is used as the starting point and its position and shapes modified or the smaller bounding box 1840 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern. [0570] This minimizing process can firstly comprise modifying the position of the bounding box (in the following, it will be assumed that the bounding box 1841 is used for the further calculations) by moving it a small amount into orthogonal directions first along the x-axis and then along the y-axis (or vice versa) as depicted in FIG. 18 around the center point M of the respective grid cell. The movement will be along the positive and the negative x-axis and y-axis and at each position, a comparison will be made to determine a difference function between the pattern obtained from the learning and the actual pattern identified in the image. This allows for calculating a two-dimensional function that represents the difference d(x,y) depending on the coordinates…..[0572] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0573] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 1800 that contain the respective fingertip or object carrying the biometric characteristic.--, in [0564]-[0573]; also see: -- [0071] It may be that the cropping operation comprises processing the image or the part of the image by a neural network, wherein processing the image or the part of the image by the neural network comprises processing the image by an encoder to obtain an encoded image and, after that, processing the encoded image by a decoder to obtain a decoded output image. [0072] Furthermore, the image or the part of the image provided to the neural network in the cropping operation for processing may comprise N×M pixels and the encoded image comprises n×m pixels, where n<N,m<M and the decoded output image comprises N×M pixels. [0073] Reducing the number of pixels when encoding the image results in a loss of information. When enlarging the image once again during the decoding, however, the most relevant information to distinguish the ROI from other portions of the image can be more easily discernable as not relevant information or very detailed information that is not necessary for identifying ROI is faded out with this procedure. [0074] In this regard, it can also be provided that distinguishing a portion of the image or the part of the image comprises distinguishing a portion of the decoded output image from another portion of the decoded output image. The distinguishing of the portions in the decoded image can be much easier compared to distinguishing the portion comprising the ROI from another portion of the original image. Thus, the processing power required for distinguishing a portion in the decoded output image from another portion in the decoded output image are reduced significantly compared to processing the original input image. [0075] More specifically, extracting the portion comprising the biometric characteristic can comprise identifying pixel in the decoded output image that are within the distinguished portion and, after that, identifying the pixels in the decoded output image that are in the distinguished portion with corresponding pixels in the original image or the part of the image and extracting, from the original image or the part of the image, the corresponding pixels, the extracted corresponding pixels constituting the portion of the image or the part of the image comprising the biometric characteristic.--, in [0071]-[0075]); and correcting the training cropped area into an axisymmetric area with the line as a symmetry axis (see Gupta: e.g., -- [0436] The method of FIG. 12 begins with the provision of training data 1201. The training data may be constituted by a plurality of images of real objects as well as images of spoofs of real objects. For example, the images may comprise a number of images of real hands or fingers or the like and images of images (i.e. spoofs) of those objects. The images may be multiplied by using, from the same image, rotated, highlighted, darkened, enlarged or otherwise modified copies that are introduced as training data. In some embodiments, modifications involving image flips, image rotation and translation, shears, crops, multiplication to increase brightness and Gaussian blurs may be used to obtain a larger number of training images. Arbitrary combinations of the mentioned techniques may also be used. The values σ.sub.1 and σ.sub.2 provided according to item 1208 are the values indicating the “correct” output of the first node and second node of the last layer in the neural network that provide the probability of the image showing a spoof of an object or a real object. These values are provided for each image in the training data.--, in [0436]; and, -- specifically extracting a region of interest from the original image obtained and processed during the steps described above and before data comprising the biometric characteristic is sent to the third party computing device. In fact, with the embodiment now described, one embodiment of obtaining image data that only comprises the region of interest, like the fingertip carrying the fingerprint of interest, is described. [0448] This embodiment focuses on the extraction of a region of interest (ROI) from an image taken from an object of the user, where the image comprises a biometric characteristic that could be used to identify the user, but could also be used for any other purpose on the third party computing device. Such an object can be anything like a fingertip of one or more fingers of a hand of the user, the palm creases of a hand of the user or the face of the user or even the eye or the foot of the user. Each of these entities is known to carry biometric characteristics that can be used for identifying a user. [0449]… as a “cropping step” which separates a part of an image from another part of the image. This method is performed on the image that was already obtained during step 1 of the method according to FIG. 1. The step may occur after the identification step 2 and after the liveliness-detection and comparison steps 3 and 4. However, it can also occur in parallel to steps 3 and 4 or immediately after step 2. In some embodiments, it does not occur before step 2 since, in step 2 of FIG. 1, it is just determined whether there is an object carrying a biometric characteristic in the image at all. Processing the image in line with the embodiment to be described now would thus not be reasonable and potentially a waste of resources if it were performed before step 2 in FIG. 1.--, in [0447]-[0449], [0453]-[0457]; and, -- [0327] As explained above, it is assumed that the neural network is already perfectly learned for identifying a specific object, such as a fingertip. This involves that the neural network is able to identify a specific pattern of pixels that are most likely representing a fingertip. This might refer to specific patterns of color values or other characteristics like the brightness of those spots. It is, however, clear that the image 300 may arbitrarily show a fingertip which might not correspond in size and arrangement to a fingertip that was used for learning the neural network. [0328] With the help of the bounding boxes and the grid, however, it is possible for the neural network to identify the specific bounding box that will most likely comprise the fingertip. In order to identify this specific bounding box, the neural network (or an associated component that processes the image 300) compares the values of the pixels within each bounding box of each grid cell to a pattern of pixels that corresponds to a fingertip as was previously learned by the neural network. In this first stage, it is most unlikely that a perfect match will be found but there will be bounding boxes that are already more likely to contain at least a portion of a fingertip than other bounding boxes. [0329]…depicted in FIG. 6, for example, the bounding box 341 centered around the point M in grid cell 313 includes a portion of the fingertip of the hand 350. In contrast to this, none of the grid cells 310 and 311 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 341 and potentially the bounding box 340, the process can determine that the bounding box 341 includes even more of a pattern that corresponds to a fingertip than the bounding box 340. [0330] In view of this, the method can conclude that none of the bounding boxes 331 and 332 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0331] As both bounding boxes 340 and 341 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0332] While the smaller grid cell 340 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 341 may be regarded by the process to include a pattern that corresponds to a fingertip. [0333] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 341 and 340 to a pattern obtained from learning which indeed corresponds to a fingertip. [0334] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 341 is used as the starting point and its position and shapes modified or the smaller bounding box 340 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern….[0337] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0338] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 300 that contain the respective fingertip.--, in [0329]-[0338], and, --depicted in FIG. 18, for example, the bounding box 1841 centered around the point M in grid cell 1813 includes a portion of the fingertip of the hand 1850. In contrast to this, none of the grid cells 1810 and 1811 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 1841 and potentially the bounding box 1840, the process can determine that the bounding box 1841 includes even more of a pattern that corresponds to a fingertip than the bounding box 1840. [0565] In view of this, the method can conclude that none of the bounding boxes 1831 and 1832 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0566] As both bounding boxes 1840 and 1841 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0567] While the smaller grid cell 1840 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 1841 may be regarded by the process to include a pattern that corresponds to a fingertip. [0568] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 1841 and 1840 to a pattern obtained from learning which indeed corresponds to a fingertip. [0569] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 1841 is used as the starting point and its position and shapes modified or the smaller bounding box 1840 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern. [0570] This minimizing process can firstly comprise modifying the position of the bounding box (in the following, it will be assumed that the bounding box 1841 is used for the further calculations) by moving it a small amount into orthogonal directions first along the x-axis and then along the y-axis (or vice versa) as depicted in FIG. 18 around the center point M of the respective grid cell. The movement will be along the positive and the negative x-axis and y-axis and at each position, a comparison will be made to determine a difference function between the pattern obtained from the learning and the actual pattern identified in the image. This allows for calculating a two-dimensional function that represents the difference d(x,y) depending on the coordinates…..[0572] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0573] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 1800 that contain the respective fingertip or object carrying the biometric characteristic.--, in [0564]-[0573]; also see: -- [0071] It may be that the cropping operation comprises processing the image or the part of the image by a neural network, wherein processing the image or the part of the image by the neural network comprises processing the image by an encoder to obtain an encoded image and, after that, processing the encoded image by a decoder to obtain a decoded output image. [0072] Furthermore, the image or the part of the image provided to the neural network in the cropping operation for processing may comprise N×M pixels and the encoded image comprises n×m pixels, where n<N,m<M and the decoded output image comprises N×M pixels. [0073] Reducing the number of pixels when encoding the image results in a loss of information. When enlarging the image once again during the decoding, however, the most relevant information to distinguish the ROI from other portions of the image can be more easily discernable as not relevant information or very detailed information that is not necessary for identifying ROI is faded out with this procedure. [0074] In this regard, it can also be provided that distinguishing a portion of the image or the part of the image comprises distinguishing a portion of the decoded output image from another portion of the decoded output image. The distinguishing of the portions in the decoded image can be much easier compared to distinguishing the portion comprising the ROI from another portion of the original image. Thus, the processing power required for distinguishing a portion in the decoded output image from another portion in the decoded output image are reduced significantly compared to processing the original input image. [0075] More specifically, extracting the portion comprising the biometric characteristic can comprise identifying pixel in the decoded output image that are within the distinguished portion and, after that, identifying the pixels in the decoded output image that are in the distinguished portion with corresponding pixels in the original image or the part of the image and extracting, from the original image or the part of the image, the corresponding pixels, the extracted corresponding pixels constituting the portion of the image or the part of the image comprising the biometric characteristic.--, in [0071]-[0075]). Re Claim 9, GUPTA as modified by SARGENT further disclose wherein after the training cropped area is obtained by cropping the training image based on the cropping probabilities, the method further comprises: correcting the training cropped area into a centrosymmetric area with the center pixel as a symmetric center (see GUPTA: e.g., --[0350] The method of FIG. 7 begins with the provision of training data 401 and preset bounding boxes 408. The training data may be constituted by a plurality of images of, for example, fingertips or a plurality of fingers depicted in one image together with other objects. The images may be multiplied by using, from the same image, rotated, highlighted, darkened, enlarged or otherwise modified copies that are introduced as training data. The bounding boxes provided according to item 408 are bounding boxes corresponding to their respective image in the training data where those bounding boxes are the bounding boxes that are correctly associated with the object to be identified, i.e. have the correct size and the correct position and a corresponding probability value as explained with respect to FIG. 6. Such bounding boxes are provided for each and every image in the training data.--, in [0350]; and --. [0464] Referring to the example of FIG. 13, the finger 1311 with the fingertip is arranged almost in the middle of the image taken. Therefore, the pre-processing step 1302 could comprise cutting of the border regions of the image 1310 and only processing further a smaller part of the original image that still comprises the fingertip 1312 with the biometric feature. This is identical to extracting, from the original image, only the center portion (for example in the form of a rectangle) comprising the fingertip. [0465] In the following, reference will usually be made to “the image” or “the input image” or the “original image”. In view of the foregoing, it is clear that this does not only comprise the full image or the original image obtained by the optical sensor but also any realization of the pre-processing, including using, instead of the full image, only a part of the image or using only one or more images comprising one colour value or being restricted to brightness values for the respective pixels.--, in [0464]-[0465]). Re Claim 10, GUPTA as modified by SARGENT further disclose wherein adjusting, based on the output result obtained by the image processing network processing the training cropped area and based on the truth value, the network parameter value and the cropping probabilities of the image processing network to obtain the trained image processing network comprises: determining a value of an objective function based on the output result and the truth value (see Gupta: e.g., -- [0436] The method of FIG. 12 begins with the provision of training data 1201. The training data may be constituted by a plurality of images of real objects as well as images of spoofs of real objects. For example, the images may comprise a number of images of real hands or fingers or the like and images of images (i.e. spoofs) of those objects. The images may be multiplied by using, from the same image, rotated, highlighted, darkened, enlarged or otherwise modified copies that are introduced as training data. In some embodiments, modifications involving image flips, image rotation and translation, shears, crops, multiplication to increase brightness and Gaussian blurs may be used to obtain a larger number of training images. Arbitrary combinations of the mentioned techniques may also be used. The values σ.sub.1 and σ.sub.2 provided according to item 1208 are the values indicating the “correct” output of the first node and second node of the last layer in the neural network that provide the probability of the image showing a spoof of an object or a real object. These values are provided for each image in the training data.--, in [0436]; and, -- specifically extracting a region of interest from the original image obtained and processed during the steps described above and before data comprising the biometric characteristic is sent to the third party computing device. In fact, with the embodiment now described, one embodiment of obtaining image data that only comprises the region of interest, like the fingertip carrying the fingerprint of interest, is described. [0448] This embodiment focuses on the extraction of a region of interest (ROI) from an image taken from an object of the user, where the image comprises a biometric characteristic that could be used to identify the user, but could also be used for any other purpose on the third party computing device. Such an object can be anything like a fingertip of one or more fingers of a hand of the user, the palm creases of a hand of the user or the face of the user or even the eye or the foot of the user. Each of these entities is known to carry biometric characteristics that can be used for identifying a user. [0449]… as a “cropping step” which separates a part of an image from another part of the image. This method is performed on the image that was already obtained during step 1 of the method according to FIG. 1. The step may occur after the identification step 2 and after the liveliness-detection and comparison steps 3 and 4. However, it can also occur in parallel to steps 3 and 4 or immediately after step 2. In some embodiments, it does not occur before step 2 since, in step 2 of FIG. 1, it is just determined whether there is an object carrying a biometric characteristic in the image at all. Processing the image in line with the embodiment to be described now would thus not be reasonable and potentially a waste of resources if it were performed before step 2 in FIG. 1.--, in [0447]-[0449], [0453]-[0457]; and, -- [0327] As explained above, it is assumed that the neural network is already perfectly learned for identifying a specific object, such as a fingertip. This involves that the neural network is able to identify a specific pattern of pixels that are most likely representing a fingertip. This might refer to specific patterns of color values or other characteristics like the brightness of those spots. It is, however, clear that the image 300 may arbitrarily show a fingertip which might not correspond in size and arrangement to a fingertip that was used for learning the neural network. [0328] With the help of the bounding boxes and the grid, however, it is possible for the neural network to identify the specific bounding box that will most likely comprise the fingertip. In order to identify this specific bounding box, the neural network (or an associated component that processes the image 300) compares the values of the pixels within each bounding box of each grid cell to a pattern of pixels that corresponds to a fingertip as was previously learned by the neural network. In this first stage, it is most unlikely that a perfect match will be found but there will be bounding boxes that are already more likely to contain at least a portion of a fingertip than other bounding boxes. [0329]…depicted in FIG. 6, for example, the bounding box 341 centered around the point M in grid cell 313 includes a portion of the fingertip of the hand 350. In contrast to this, none of the grid cells 310 and 311 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 341 and potentially the bounding box 340, the process can determine that the bounding box 341 includes even more of a pattern that corresponds to a fingertip than the bounding box 340. [0330] In view of this, the method can conclude that none of the bounding boxes 331 and 332 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0331] As both bounding boxes 340 and 341 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0332] While the smaller grid cell 340 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 341 may be regarded by the process to include a pattern that corresponds to a fingertip. [0333] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 341 and 340 to a pattern obtained from learning which indeed corresponds to a fingertip. [0334] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 341 is used as the starting point and its position and shapes modified or the smaller bounding box 340 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern….[0337] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0338] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 300 that contain the respective fingertip.--, in [0329]-[0338], and, --depicted in FIG. 18, for example, the bounding box 1841 centered around the point M in grid cell 1813 includes a portion of the fingertip of the hand 1850. In contrast to this, none of the grid cells 1810 and 1811 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 1841 and potentially the bounding box 1840, the process can determine that the bounding box 1841 includes even more of a pattern that corresponds to a fingertip than the bounding box 1840. [0565] In view of this, the method can conclude that none of the bounding boxes 1831 and 1832 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0566] As both bounding boxes 1840 and 1841 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0567] While the smaller grid cell 1840 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 1841 may be regarded by the process to include a pattern that corresponds to a fingertip. [0568] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 1841 and 1840 to a pattern obtained from learning which indeed corresponds to a fingertip. [0569] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 1841 is used as the starting point and its position and shapes modified or the smaller bounding box 1840 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern. [0570] This minimizing process can firstly comprise modifying the position of the bounding box (in the following, it will be assumed that the bounding box 1841 is used for the further calculations) by moving it a small amount into orthogonal directions first along the x-axis and then along the y-axis (or vice versa) as depicted in FIG. 18 around the center point M of the respective grid cell. The movement will be along the positive and the negative x-axis and y-axis and at each position, a comparison will be made to determine a difference function between the pattern obtained from the learning and the actual pattern identified in the image. This allows for calculating a two-dimensional function that represents the difference d(x,y) depending on the coordinates…..[0572] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0573] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 1800 that contain the respective fingertip or object carrying the biometric characteristic.--, in [0564]-[0573]; also see: -- [0071] It may be that the cropping operation comprises processing the image or the part of the image by a neural network, wherein processing the image or the part of the image by the neural network comprises processing the image by an encoder to obtain an encoded image and, after that, processing the encoded image by a decoder to obtain a decoded output image. [0072] Furthermore, the image or the part of the image provided to the neural network in the cropping operation for processing may comprise N×M pixels and the encoded image comprises n×m pixels, where n<N,m<M and the decoded output image comprises N×M pixels. [0073] Reducing the number of pixels when encoding the image results in a loss of information. When enlarging the image once again during the decoding, however, the most relevant information to distinguish the ROI from other portions of the image can be more easily discernable as not relevant information or very detailed information that is not necessary for identifying ROI is faded out with this procedure. [0074] In this regard, it can also be provided that distinguishing a portion of the image or the part of the image comprises distinguishing a portion of the decoded output image from another portion of the decoded output image. The distinguishing of the portions in the decoded image can be much easier compared to distinguishing the portion comprising the ROI from another portion of the original image. Thus, the processing power required for distinguishing a portion in the decoded output image from another portion in the decoded output image are reduced significantly compared to processing the original input image. [0075] More specifically, extracting the portion comprising the biometric characteristic can comprise identifying pixel in the decoded output image that are within the distinguished portion and, after that, identifying the pixels in the decoded output image that are in the distinguished portion with corresponding pixels in the original image or the part of the image and extracting, from the original image or the part of the image, the corresponding pixels, the extracted corresponding pixels constituting the portion of the image or the part of the image comprising the biometric characteristic.--, in [0071]-[0075]); and adjusting, based on the value of the objective function and a computation quantity loss of the image processing network, the network parameter value and the cropping probabilities to obtain the trained image processing network (see Gupta: e.g., -- [0436] The method of FIG. 12 begins with the provision of training data 1201. The training data may be constituted by a plurality of images of real objects as well as images of spoofs of real objects. For example, the images may comprise a number of images of real hands or fingers or the like and images of images (i.e. spoofs) of those objects. The images may be multiplied by using, from the same image, rotated, highlighted, darkened, enlarged or otherwise modified copies that are introduced as training data. In some embodiments, modifications involving image flips, image rotation and translation, shears, crops, multiplication to increase brightness and Gaussian blurs may be used to obtain a larger number of training images. Arbitrary combinations of the mentioned techniques may also be used. The values σ.sub.1 and σ.sub.2 provided according to item 1208 are the values indicating the “correct” output of the first node and second node of the last layer in the neural network that provide the probability of the image showing a spoof of an object or a real object. These values are provided for each image in the training data.--, in [0436]; and, -- specifically extracting a region of interest from the original image obtained and processed during the steps described above and before data comprising the biometric characteristic is sent to the third party computing device. In fact, with the embodiment now described, one embodiment of obtaining image data that only comprises the region of interest, like the fingertip carrying the fingerprint of interest, is described. [0448] This embodiment focuses on the extraction of a region of interest (ROI) from an image taken from an object of the user, where the image comprises a biometric characteristic that could be used to identify the user, but could also be used for any other purpose on the third party computing device. Such an object can be anything like a fingertip of one or more fingers of a hand of the user, the palm creases of a hand of the user or the face of the user or even the eye or the foot of the user. Each of these entities is known to carry biometric characteristics that can be used for identifying a user. [0449]… as a “cropping step” which separates a part of an image from another part of the image. This method is performed on the image that was already obtained during step 1 of the method according to FIG. 1. The step may occur after the identification step 2 and after the liveliness-detection and comparison steps 3 and 4. However, it can also occur in parallel to steps 3 and 4 or immediately after step 2. In some embodiments, it does not occur before step 2 since, in step 2 of FIG. 1, it is just determined whether there is an object carrying a biometric characteristic in the image at all. Processing the image in line with the embodiment to be described now would thus not be reasonable and potentially a waste of resources if it were performed before step 2 in FIG. 1.--, in [0447]-[0449], [0453]-[0457]; and, -- [0327] As explained above, it is assumed that the neural network is already perfectly learned for identifying a specific object, such as a fingertip. This involves that the neural network is able to identify a specific pattern of pixels that are most likely representing a fingertip. This might refer to specific patterns of color values or other characteristics like the brightness of those spots. It is, however, clear that the image 300 may arbitrarily show a fingertip which might not correspond in size and arrangement to a fingertip that was used for learning the neural network. [0328] With the help of the bounding boxes and the grid, however, it is possible for the neural network to identify the specific bounding box that will most likely comprise the fingertip. In order to identify this specific bounding box, the neural network (or an associated component that processes the image 300) compares the values of the pixels within each bounding box of each grid cell to a pattern of pixels that corresponds to a fingertip as was previously learned by the neural network. In this first stage, it is most unlikely that a perfect match will be found but there will be bounding boxes that are already more likely to contain at least a portion of a fingertip than other bounding boxes. [0329]…depicted in FIG. 6, for example, the bounding box 341 centered around the point M in grid cell 313 includes a portion of the fingertip of the hand 350. In contrast to this, none of the grid cells 310 and 311 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 341 and potentially the bounding box 340, the process can determine that the bounding box 341 includes even more of a pattern that corresponds to a fingertip than the bounding box 340. [0330] In view of this, the method can conclude that none of the bounding boxes 331 and 332 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0331] As both bounding boxes 340 and 341 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0332] While the smaller grid cell 340 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 341 may be regarded by the process to include a pattern that corresponds to a fingertip. [0333] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 341 and 340 to a pattern obtained from learning which indeed corresponds to a fingertip. [0334] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 341 is used as the starting point and its position and shapes modified or the smaller bounding box 340 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern….[0337] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0338] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 300 that contain the respective fingertip.--, in [0329]-[0338], and, --depicted in FIG. 18, for example, the bounding box 1841 centered around the point M in grid cell 1813 includes a portion of the fingertip of the hand 1850. In contrast to this, none of the grid cells 1810 and 1811 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 1841 and potentially the bounding box 1840, the process can determine that the bounding box 1841 includes even more of a pattern that corresponds to a fingertip than the bounding box 1840. [0565] In view of this, the method can conclude that none of the bounding boxes 1831 and 1832 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0566] As both bounding boxes 1840 and 1841 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0567] While the smaller grid cell 1840 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 1841 may be regarded by the process to include a pattern that corresponds to a fingertip. [0568] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 1841 and 1840 to a pattern obtained from learning which indeed corresponds to a fingertip. [0569] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 1841 is used as the starting point and its position and shapes modified or the smaller bounding box 1840 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern. [0570] This minimizing process can firstly comprise modifying the position of the bounding box (in the following, it will be assumed that the bounding box 1841 is used for the further calculations) by moving it a small amount into orthogonal directions first along the x-axis and then along the y-axis (or vice versa) as depicted in FIG. 18 around the center point M of the respective grid cell. The movement will be along the positive and the negative x-axis and y-axis and at each position, a comparison will be made to determine a difference function between the pattern obtained from the learning and the actual pattern identified in the image. This allows for calculating a two-dimensional function that represents the difference d(x,y) depending on the coordinates…..[0572] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0573] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 1800 that contain the respective fingertip or object carrying the biometric characteristic.--, in [0564]-[0573]; also see: -- [0071] It may be that the cropping operation comprises processing the image or the part of the image by a neural network, wherein processing the image or the part of the image by the neural network comprises processing the image by an encoder to obtain an encoded image and, after that, processing the encoded image by a decoder to obtain a decoded output image. [0072] Furthermore, the image or the part of the image provided to the neural network in the cropping operation for processing may comprise N×M pixels and the encoded image comprises n×m pixels, where n<N,m<M and the decoded output image comprises N×M pixels. [0073] Reducing the number of pixels when encoding the image results in a loss of information. When enlarging the image once again during the decoding, however, the most relevant information to distinguish the ROI from other portions of the image can be more easily discernable as not relevant information or very detailed information that is not necessary for identifying ROI is faded out with this procedure. [0074] In this regard, it can also be provided that distinguishing a portion of the image or the part of the image comprises distinguishing a portion of the decoded output image from another portion of the decoded output image. The distinguishing of the portions in the decoded image can be much easier compared to distinguishing the portion comprising the ROI from another portion of the original image. Thus, the processing power required for distinguishing a portion in the decoded output image from another portion in the decoded output image are reduced significantly compared to processing the original input image. [0075] More specifically, extracting the portion comprising the biometric characteristic can comprise identifying pixel in the decoded output image that are within the distinguished portion and, after that, identifying the pixels in the decoded output image that are in the distinguished portion with corresponding pixels in the original image or the part of the image and extracting, from the original image or the part of the image, the corresponding pixels, the extracted corresponding pixels constituting the portion of the image or the part of the image comprising the biometric characteristic.--, in [0071]-[0075]). Re Claim 11, GUPTA as modified by SARGENT further disclose wherein determining the value of the objective function based on the output result and the truth value comprises: fusing the output result with the truth value to obtain a first fusion result of the training cropped area (see SARGENT: e.g., --fused deep CNN networks with the pixel-based multilayer perceptron (MLP) method to solve LC classification with spatial feature representation and pixel-level differentiation; Zhang et al. (2018b) proposed a regional fusion decision strategy based on rough set theory to model the uncertainties in LC classification of the CNN, and further guide data integration with other algorithms for targeted adjustment; Pan and Zhao, (2017) developed a central-point-enhanced CNN network to enhance the weight of the central pixels within image patches to strengthen the LC classification with precise land-cover boundaries.--; in [0006], and, --The predetermined classification fusion rules may include, if the land use classification data from both the first and second convolutional neural networks match for a particular object, assigning that matching land use classification to the object. Or, if the land use classification data from both the first and second convolutional neural networks do not match, selecting one of the land use classifications for the particular object in accordance with one or more predetermined criteria.--, in [0017]); and, fusing a result of the image processing network processing the training image with the truth value, to obtain a second fusion result of the training image (see SARGENT: e.g., --fused deep CNN networks with the pixel-based multilayer perceptron (MLP) method to solve LC classification with spatial feature representation and pixel-level differentiation; Zhang et al. (2018b) proposed a regional fusion decision strategy based on rough set theory to model the uncertainties in LC classification of the CNN, and further guide data integration with other algorithms for targeted adjustment; Pan and Zhao, (2017) developed a central-point-enhanced CNN network to enhance the weight of the central pixels within image patches to strengthen the LC classification with precise land-cover boundaries.--; in [0006], and, --The predetermined classification fusion rules may include, if the land use classification data from both the first and second convolutional neural networks match for a particular object, assigning that matching land use classification to the object. Or, if the land use classification data from both the first and second convolutional neural networks do not match, selecting one of the land use classifications for the particular object in accordance with one or more predetermined criteria.--, in [0017]); and obtaining, based on a ratio of the first fusion result to the second fusion result, the value of the objective function (see SARGENT: e.g., --fused deep CNN networks with the pixel-based multilayer perceptron (MLP) method to solve LC classification with spatial feature representation and pixel-level differentiation; Zhang et al. (2018b) proposed a regional fusion decision strategy based on rough set theory to model the uncertainties in LC classification of the CNN, and further guide data integration with other algorithms for targeted adjustment; Pan and Zhao, (2017) developed a central-point-enhanced CNN network to enhance the weight of the central pixels within image patches to strengthen the LC classification with precise land-cover boundaries.--; in [0006], and, --The predetermined classification fusion rules may include, if the land use classification data from both the first and second convolutional neural networks match for a particular object, assigning that matching land use classification to the object. Or, if the land use classification data from both the first and second convolutional neural networks do not match, selecting one of the land use classifications for the particular object in accordance with one or more predetermined criteria.--, in [0017]; and, --A multilayer perceptron (MLP) is a network that maps from input data to output representations through a feedforward manner. The fundamental component of a MLP involves a set of computational nodes with weights and biases at multiple layers (input, hidden, and output layers) that are fully connected. The weights and biases within the network are learned through backpropagation to approximate the complex relationship between the input features and the output characteristics. The learning objective is to minimise the difference between the predictions and the desired outputs by using a specific cost function. As one of the most representative deep neural networks, convolutional neural network (CNN) is designed to process and analyse large scale sensory data or images in consideration of their stationary characteristics at local and global scales. Within the CNN network, convolutional layers and pooling layers are connected alternatively to generalise the features towards deep and abstract representations. Typically, the convolutional layers are composed of weights and biases that are learnt through a set of image patches across the image. Those weights are shared by different feature maps, in which multiple features are learnt with a reduced amount of parameters, and an activation function (e.g. rectified linear units) is followed to strengthen the non-linearity of the convolutional operations. The pooling layer involves max-pooling or average-pooling, where the summary statistics of local regions are derived to further enhance the generalisation capability. An object-based CNN (OCNN) was proposed recently for the urban LU classification using remotely sensed imagery. The OCNN is trained as for the standard CNN model with labelled image patches, whereas the model prediction labels each segmented object derived from image segmentation. For each image object (polygon), a minimum moment bounding box was constructed by anisotropy with major and minor axes. The centre point intersected with the polygon and the bisector of the major axis was used to approximate the central location of each image patch, where the convolutional process is implemented once per object.. The size of the image patch was tuned empirically to be sufficiently large, so that the object and spatial context were captured jointly by the CNN network. The OCNN was trained on the LU classes, in which the semantic information of LU was learnt through the deep network, while the boundaries of the objects were retained through the process of segmentation. The CNN model prediction was recorded as the predicted label of the image object to formulate a LU thematic map. Here, the predictions of each object are assigned to all of its pixels.--, in [0041]-[0044]). Re Claim 12, GUPTA as modified by SARGENT further disclose wherein the network parameter value of the image processing network comprises: a weight and a pruning probability of a channel to be pruned, and adjusting, based on the value of the objective function and the computation quantity loss of the image processing network, the network parameter value and the cropping probabilities to obtain the trained image processing network comprises: obtaining a transition loss based on the value of the objective function and the computation quantity loss (see Gupta: e.g., -- [0436] The method of FIG. 12 begins with the provision of training data 1201. The training data may be constituted by a plurality of images of real objects as well as images of spoofs of real objects. For example, the images may comprise a number of images of real hands or fingers or the like and images of images (i.e. spoofs) of those objects. The images may be multiplied by using, from the same image, rotated, highlighted, darkened, enlarged or otherwise modified copies that are introduced as training data. In some embodiments, modifications involving image flips, image rotation and translation, shears, crops, multiplication to increase brightness and Gaussian blurs may be used to obtain a larger number of training images. Arbitrary combinations of the mentioned techniques may also be used. The values σ.sub.1 and σ.sub.2 provided according to item 1208 are the values indicating the “correct” output of the first node and second node of the last layer in the neural network that provide the probability of the image showing a spoof of an object or a real object. These values are provided for each image in the training data.--, in [0436]; and, -- specifically extracting a region of interest from the original image obtained and processed during the steps described above and before data comprising the biometric characteristic is sent to the third party computing device. In fact, with the embodiment now described, one embodiment of obtaining image data that only comprises the region of interest, like the fingertip carrying the fingerprint of interest, is described. [0448] This embodiment focuses on the extraction of a region of interest (ROI) from an image taken from an object of the user, where the image comprises a biometric characteristic that could be used to identify the user, but could also be used for any other purpose on the third party computing device. Such an object can be anything like a fingertip of one or more fingers of a hand of the user, the palm creases of a hand of the user or the face of the user or even the eye or the foot of the user. Each of these entities is known to carry biometric characteristics that can be used for identifying a user. [0449]… as a “cropping step” which separates a part of an image from another part of the image. This method is performed on the image that was already obtained during step 1 of the method according to FIG. 1. The step may occur after the identification step 2 and after the liveliness-detection and comparison steps 3 and 4. However, it can also occur in parallel to steps 3 and 4 or immediately after step 2. In some embodiments, it does not occur before step 2 since, in step 2 of FIG. 1, it is just determined whether there is an object carrying a biometric characteristic in the image at all. Processing the image in line with the embodiment to be described now would thus not be reasonable and potentially a waste of resources if it were performed before step 2 in FIG. 1.--, in [0447]-[0449], [0453]-[0457]; and, -- [0327] As explained above, it is assumed that the neural network is already perfectly learned for identifying a specific object, such as a fingertip. This involves that the neural network is able to identify a specific pattern of pixels that are most likely representing a fingertip. This might refer to specific patterns of color values or other characteristics like the brightness of those spots. It is, however, clear that the image 300 may arbitrarily show a fingertip which might not correspond in size and arrangement to a fingertip that was used for learning the neural network. [0328] With the help of the bounding boxes and the grid, however, it is possible for the neural network to identify the specific bounding box that will most likely comprise the fingertip. In order to identify this specific bounding box, the neural network (or an associated component that processes the image 300) compares the values of the pixels within each bounding box of each grid cell to a pattern of pixels that corresponds to a fingertip as was previously learned by the neural network. In this first stage, it is most unlikely that a perfect match will be found but there will be bounding boxes that are already more likely to contain at least a portion of a fingertip than other bounding boxes. [0329]…depicted in FIG. 6, for example, the bounding box 341 centered around the point M in grid cell 313 includes a portion of the fingertip of the hand 350. In contrast to this, none of the grid cells 310 and 311 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 341 and potentially the bounding box 340, the process can determine that the bounding box 341 includes even more of a pattern that corresponds to a fingertip than the bounding box 340. [0330] In view of this, the method can conclude that none of the bounding boxes 331 and 332 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0331] As both bounding boxes 340 and 341 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0332] While the smaller grid cell 340 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 341 may be regarded by the process to include a pattern that corresponds to a fingertip. [0333] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 341 and 340 to a pattern obtained from learning which indeed corresponds to a fingertip. [0334] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 341 is used as the starting point and its position and shapes modified or the smaller bounding box 340 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern….[0337] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0338] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 300 that contain the respective fingertip.--, in [0329]-[0338], and, --depicted in FIG. 18, for example, the bounding box 1841 centered around the point M in grid cell 1813 includes a portion of the fingertip of the hand 1850. In contrast to this, none of the grid cells 1810 and 1811 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 1841 and potentially the bounding box 1840, the process can determine that the bounding box 1841 includes even more of a pattern that corresponds to a fingertip than the bounding box 1840. [0565] In view of this, the method can conclude that none of the bounding boxes 1831 and 1832 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0566] As both bounding boxes 1840 and 1841 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0567] While the smaller grid cell 1840 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 1841 may be regarded by the process to include a pattern that corresponds to a fingertip. [0568] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 1841 and 1840 to a pattern obtained from learning which indeed corresponds to a fingertip. [0569] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 1841 is used as the starting point and its position and shapes modified or the smaller bounding box 1840 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern. [0570] This minimizing process can firstly comprise modifying the position of the bounding box (in the following, it will be assumed that the bounding box 1841 is used for the further calculations) by moving it a small amount into orthogonal directions first along the x-axis and then along the y-axis (or vice versa) as depicted in FIG. 18 around the center point M of the respective grid cell. The movement will be along the positive and the negative x-axis and y-axis and at each position, a comparison will be made to determine a difference function between the pattern obtained from the learning and the actual pattern identified in the image. This allows for calculating a two-dimensional function that represents the difference d(x,y) depending on the coordinates…..[0572] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0573] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 1800 that contain the respective fingertip or object carrying the biometric characteristic.--, in [0564]-[0573]; also see: -- [0071] It may be that the cropping operation comprises processing the image or the part of the image by a neural network, wherein processing the image or the part of the image by the neural network comprises processing the image by an encoder to obtain an encoded image and, after that, processing the encoded image by a decoder to obtain a decoded output image. [0072] Furthermore, the image or the part of the image provided to the neural network in the cropping operation for processing may comprise N×M pixels and the encoded image comprises n×m pixels, where n<N,m<M and the decoded output image comprises N×M pixels. [0073] Reducing the number of pixels when encoding the image results in a loss of information. When enlarging the image once again during the decoding, however, the most relevant information to distinguish the ROI from other portions of the image can be more easily discernable as not relevant information or very detailed information that is not necessary for identifying ROI is faded out with this procedure. [0074] In this regard, it can also be provided that distinguishing a portion of the image or the part of the image comprises distinguishing a portion of the decoded output image from another portion of the decoded output image. The distinguishing of the portions in the decoded image can be much easier compared to distinguishing the portion comprising the ROI from another portion of the original image. Thus, the processing power required for distinguishing a portion in the decoded output image from another portion in the decoded output image are reduced significantly compared to processing the original input image. [0075] More specifically, extracting the portion comprising the biometric characteristic can comprise identifying pixel in the decoded output image that are within the distinguished portion and, after that, identifying the pixels in the decoded output image that are in the distinguished portion with corresponding pixels in the original image or the part of the image and extracting, from the original image or the part of the image, the corresponding pixels, the extracted corresponding pixels constituting the portion of the image or the part of the image comprising the biometric characteristic.--, in [0071]-[0075]); and adjusting, based on the value of the objective function, the weight and the pruning probability of the channel to be pruned, and adjusting the cropping probabilities based on the transition loss to obtain the trained image processing network (see Gupta: e.g., -- [0436] The method of FIG. 12 begins with the provision of training data 1201. The training data may be constituted by a plurality of images of real objects as well as images of spoofs of real objects. For example, the images may comprise a number of images of real hands or fingers or the like and images of images (i.e. spoofs) of those objects. The images may be multiplied by using, from the same image, rotated, highlighted, darkened, enlarged or otherwise modified copies that are introduced as training data. In some embodiments, modifications involving image flips, image rotation and translation, shears, crops, multiplication to increase brightness and Gaussian blurs may be used to obtain a larger number of training images. Arbitrary combinations of the mentioned techniques may also be used. The values σ.sub.1 and σ.sub.2 provided according to item 1208 are the values indicating the “correct” output of the first node and second node of the last layer in the neural network that provide the probability of the image showing a spoof of an object or a real object. These values are provided for each image in the training data.--, in [0436]; and, -- specifically extracting a region of interest from the original image obtained and processed during the steps described above and before data comprising the biometric characteristic is sent to the third party computing device. In fact, with the embodiment now described, one embodiment of obtaining image data that only comprises the region of interest, like the fingertip carrying the fingerprint of interest, is described. [0448] This embodiment focuses on the extraction of a region of interest (ROI) from an image taken from an object of the user, where the image comprises a biometric characteristic that could be used to identify the user, but could also be used for any other purpose on the third party computing device. Such an object can be anything like a fingertip of one or more fingers of a hand of the user, the palm creases of a hand of the user or the face of the user or even the eye or the foot of the user. Each of these entities is known to carry biometric characteristics that can be used for identifying a user. [0449]… as a “cropping step” which separates a part of an image from another part of the image. This method is performed on the image that was already obtained during step 1 of the method according to FIG. 1. The step may occur after the identification step 2 and after the liveliness-detection and comparison steps 3 and 4. However, it can also occur in parallel to steps 3 and 4 or immediately after step 2. In some embodiments, it does not occur before step 2 since, in step 2 of FIG. 1, it is just determined whether there is an object carrying a biometric characteristic in the image at all. Processing the image in line with the embodiment to be described now would thus not be reasonable and potentially a waste of resources if it were performed before step 2 in FIG. 1.--, in [0447]-[0449], [0453]-[0457]; and, -- [0327] As explained above, it is assumed that the neural network is already perfectly learned for identifying a specific object, such as a fingertip. This involves that the neural network is able to identify a specific pattern of pixels that are most likely representing a fingertip. This might refer to specific patterns of color values or other characteristics like the brightness of those spots. It is, however, clear that the image 300 may arbitrarily show a fingertip which might not correspond in size and arrangement to a fingertip that was used for learning the neural network. [0328] With the help of the bounding boxes and the grid, however, it is possible for the neural network to identify the specific bounding box that will most likely comprise the fingertip. In order to identify this specific bounding box, the neural network (or an associated component that processes the image 300) compares the values of the pixels within each bounding box of each grid cell to a pattern of pixels that corresponds to a fingertip as was previously learned by the neural network. In this first stage, it is most unlikely that a perfect match will be found but there will be bounding boxes that are already more likely to contain at least a portion of a fingertip than other bounding boxes. [0329]…depicted in FIG. 6, for example, the bounding box 341 centered around the point M in grid cell 313 includes a portion of the fingertip of the hand 350. In contrast to this, none of the grid cells 310 and 311 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 341 and potentially the bounding box 340, the process can determine that the bounding box 341 includes even more of a pattern that corresponds to a fingertip than the bounding box 340. [0330] In view of this, the method can conclude that none of the bounding boxes 331 and 332 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0331] As both bounding boxes 340 and 341 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0332] While the smaller grid cell 340 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 341 may be regarded by the process to include a pattern that corresponds to a fingertip. [0333] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 341 and 340 to a pattern obtained from learning which indeed corresponds to a fingertip. [0334] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 341 is used as the starting point and its position and shapes modified or the smaller bounding box 340 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern….[0337] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0338] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 300 that contain the respective fingertip.--, in [0329]-[0338], and, --depicted in FIG. 18, for example, the bounding box 1841 centered around the point M in grid cell 1813 includes a portion of the fingertip of the hand 1850. In contrast to this, none of the grid cells 1810 and 1811 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 1841 and potentially the bounding box 1840, the process can determine that the bounding box 1841 includes even more of a pattern that corresponds to a fingertip than the bounding box 1840. [0565] In view of this, the method can conclude that none of the bounding boxes 1831 and 1832 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0566] As both bounding boxes 1840 and 1841 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0567] While the smaller grid cell 1840 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 1841 may be regarded by the process to include a pattern that corresponds to a fingertip. [0568] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 1841 and 1840 to a pattern obtained from learning which indeed corresponds to a fingertip. [0569] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 1841 is used as the starting point and its position and shapes modified or the smaller bounding box 1840 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern. [0570] This minimizing process can firstly comprise modifying the position of the bounding box (in the following, it will be assumed that the bounding box 1841 is used for the further calculations) by moving it a small amount into orthogonal directions first along the x-axis and then along the y-axis (or vice versa) as depicted in FIG. 18 around the center point M of the respective grid cell. The movement will be along the positive and the negative x-axis and y-axis and at each position, a comparison will be made to determine a difference function between the pattern obtained from the learning and the actual pattern identified in the image. This allows for calculating a two-dimensional function that represents the difference d(x,y) depending on the coordinates…..[0572] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0573] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 1800 that contain the respective fingertip or object carrying the biometric characteristic.--, in [0564]-[0573]; also see: -- [0071] It may be that the cropping operation comprises processing the image or the part of the image by a neural network, wherein processing the image or the part of the image by the neural network comprises processing the image by an encoder to obtain an encoded image and, after that, processing the encoded image by a decoder to obtain a decoded output image. [0072] Furthermore, the image or the part of the image provided to the neural network in the cropping operation for processing may comprise N×M pixels and the encoded image comprises n×m pixels, where n<N,m<M and the decoded output image comprises N×M pixels. [0073] Reducing the number of pixels when encoding the image results in a loss of information. When enlarging the image once again during the decoding, however, the most relevant information to distinguish the ROI from other portions of the image can be more easily discernable as not relevant information or very detailed information that is not necessary for identifying ROI is faded out with this procedure. [0074] In this regard, it can also be provided that distinguishing a portion of the image or the part of the image comprises distinguishing a portion of the decoded output image from another portion of the decoded output image. The distinguishing of the portions in the decoded image can be much easier compared to distinguishing the portion comprising the ROI from another portion of the original image. Thus, the processing power required for distinguishing a portion in the decoded output image from another portion in the decoded output image are reduced significantly compared to processing the original input image. [0075] More specifically, extracting the portion comprising the biometric characteristic can comprise identifying pixel in the decoded output image that are within the distinguished portion and, after that, identifying the pixels in the decoded output image that are in the distinguished portion with corresponding pixels in the original image or the part of the image and extracting, from the original image or the part of the image, the corresponding pixels, the extracted corresponding pixels constituting the portion of the image or the part of the image comprising the biometric characteristic.--, in [0071]-[0075]). Re Claim 13, GUPTA as modified by SARGENT further disclose an image processing method, comprising: acquiring an image to be processed (see GUPTA: e.g., --[0017] FIG. 7 shows the process of training the neural network in the identification operation--, and, --[0022] FIG. 12 shows the process of training the neural network used in the liveliness-detection operation--, and, --[0029] FIG. 19 presents a flowchart of the training of the neural network used in the cropping operation--, in [0017]-[0029]; and, --[0065] The neural network, according to some embodiments, is trained and/or otherwise adapted to be able to distinguish, through processing of the image, between those portions of the image that contain the region of interest and other portions of the image that do not contain the region of interest. This can be achieved in variety of ways and is thus not to be understood in a limiting way. That the neural network “distinguishes” that region comprising the ROI from another region is to be understood as the ability of the neural network to provide an output that distinguishes between the ROI and other regions of the image or makes it possible to distinguish between those regions. For example, the output could be an indication of pixels comprising the ROI but not other regions of the ROI. In any case, the outcome of the processing of the image by the neural network at least comprises that a first portion of the image comprising the region of interest is identified as different from another portion of the image. In this regard, it is noted that the specific size, shape of position of the region of interest is found out by the neural network during processing of the image and it is not preset.--, in [0065]); performing, based on cropping probabilities of a trained image processing network, pixel cropping on the image to be processed, to obtain a cropped area to be processed, wherein the trained image processing network is trained based on the method of claim 1 (see Gupta: e.g., -- [0436] The method of FIG. 12 begins with the provision of training data 1201. The training data may be constituted by a plurality of images of real objects as well as images of spoofs of real objects. For example, the images may comprise a number of images of real hands or fingers or the like and images of images (i.e. spoofs) of those objects. The images may be multiplied by using, from the same image, rotated, highlighted, darkened, enlarged or otherwise modified copies that are introduced as training data. In some embodiments, modifications involving image flips, image rotation and translation, shears, crops, multiplication to increase brightness and Gaussian blurs may be used to obtain a larger number of training images. Arbitrary combinations of the mentioned techniques may also be used. The values σ.sub.1 and σ.sub.2 provided according to item 1208 are the values indicating the “correct” output of the first node and second node of the last layer in the neural network that provide the probability of the image showing a spoof of an object or a real object. These values are provided for each image in the training data.--, in [0436]; and, -- specifically extracting a region of interest from the original image obtained and processed during the steps described above and before data comprising the biometric characteristic is sent to the third party computing device. In fact, with the embodiment now described, one embodiment of obtaining image data that only comprises the region of interest, like the fingertip carrying the fingerprint of interest, is described. [0448] This embodiment focuses on the extraction of a region of interest (ROI) from an image taken from an object of the user, where the image comprises a biometric characteristic that could be used to identify the user, but could also be used for any other purpose on the third party computing device. Such an object can be anything like a fingertip of one or more fingers of a hand of the user, the palm creases of a hand of the user or the face of the user or even the eye or the foot of the user. Each of these entities is known to carry biometric characteristics that can be used for identifying a user. [0449]… as a “cropping step” which separates a part of an image from another part of the image. This method is performed on the image that was already obtained during step 1 of the method according to FIG. 1. The step may occur after the identification step 2 and after the liveliness-detection and comparison steps 3 and 4. However, it can also occur in parallel to steps 3 and 4 or immediately after step 2. In some embodiments, it does not occur before step 2 since, in step 2 of FIG. 1, it is just determined whether there is an object carrying a biometric characteristic in the image at all. Processing the image in line with the embodiment to be described now would thus not be reasonable and potentially a waste of resources if it were performed before step 2 in FIG. 1.--, in [0447]-[0449], [0453]-[0457]; and, -- [0327] As explained above, it is assumed that the neural network is already perfectly learned for identifying a specific object, such as a fingertip. This involves that the neural network is able to identify a specific pattern of pixels that are most likely representing a fingertip. This might refer to specific patterns of color values or other characteristics like the brightness of those spots. It is, however, clear that the image 300 may arbitrarily show a fingertip which might not correspond in size and arrangement to a fingertip that was used for learning the neural network. [0328] With the help of the bounding boxes and the grid, however, it is possible for the neural network to identify the specific bounding box that will most likely comprise the fingertip. In order to identify this specific bounding box, the neural network (or an associated component that processes the image 300) compares the values of the pixels within each bounding box of each grid cell to a pattern of pixels that corresponds to a fingertip as was previously learned by the neural network. In this first stage, it is most unlikely that a perfect match will be found but there will be bounding boxes that are already more likely to contain at least a portion of a fingertip than other bounding boxes. [0329]…depicted in FIG. 6, for example, the bounding box 341 centered around the point M in grid cell 313 includes a portion of the fingertip of the hand 350. In contrast to this, none of the grid cells 310 and 311 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 341 and potentially the bounding box 340, the process can determine that the bounding box 341 includes even more of a pattern that corresponds to a fingertip than the bounding box 340. [0330] In view of this, the method can conclude that none of the bounding boxes 331 and 332 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0331] As both bounding boxes 340 and 341 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0332] While the smaller grid cell 340 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 341 may be regarded by the process to include a pattern that corresponds to a fingertip. [0333] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 341 and 340 to a pattern obtained from learning which indeed corresponds to a fingertip. [0334] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 341 is used as the starting point and its position and shapes modified or the smaller bounding box 340 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern….[0337] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0338] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 300 that contain the respective fingertip.--, in [0329]-[0338], and, --depicted in FIG. 18, for example, the bounding box 1841 centered around the point M in grid cell 1813 includes a portion of the fingertip of the hand 1850. In contrast to this, none of the grid cells 1810 and 1811 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 1841 and potentially the bounding box 1840, the process can determine that the bounding box 1841 includes even more of a pattern that corresponds to a fingertip than the bounding box 1840. [0565] In view of this, the method can conclude that none of the bounding boxes 1831 and 1832 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0566] As both bounding boxes 1840 and 1841 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0567] While the smaller grid cell 1840 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 1841 may be regarded by the process to include a pattern that corresponds to a fingertip. [0568] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 1841 and 1840 to a pattern obtained from learning which indeed corresponds to a fingertip. [0569] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 1841 is used as the starting point and its position and shapes modified or the smaller bounding box 1840 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern. [0570] This minimizing process can firstly comprise modifying the position of the bounding box (in the following, it will be assumed that the bounding box 1841 is used for the further calculations) by moving it a small amount into orthogonal directions first along the x-axis and then along the y-axis (or vice versa) as depicted in FIG. 18 around the center point M of the respective grid cell. The movement will be along the positive and the negative x-axis and y-axis and at each position, a comparison will be made to determine a difference function between the pattern obtained from the learning and the actual pattern identified in the image. This allows for calculating a two-dimensional function that represents the difference d(x,y) depending on the coordinates…..[0572] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0573] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 1800 that contain the respective fingertip or object carrying the biometric characteristic.--, in [0564]-[0573]; also see: -- [0071] It may be that the cropping operation comprises processing the image or the part of the image by a neural network, wherein processing the image or the part of the image by the neural network comprises processing the image by an encoder to obtain an encoded image and, after that, processing the encoded image by a decoder to obtain a decoded output image. [0072] Furthermore, the image or the part of the image provided to the neural network in the cropping operation for processing may comprise N×M pixels and the encoded image comprises n×m pixels, where n<N,m<M and the decoded output image comprises N×M pixels. [0073] Reducing the number of pixels when encoding the image results in a loss of information. When enlarging the image once again during the decoding, however, the most relevant information to distinguish the ROI from other portions of the image can be more easily discernable as not relevant information or very detailed information that is not necessary for identifying ROI is faded out with this procedure. [0074] In this regard, it can also be provided that distinguishing a portion of the image or the part of the image comprises distinguishing a portion of the decoded output image from another portion of the decoded output image. The distinguishing of the portions in the decoded image can be much easier compared to distinguishing the portion comprising the ROI from another portion of the original image. Thus, the processing power required for distinguishing a portion in the decoded output image from another portion in the decoded output image are reduced significantly compared to processing the original input image. [0075] More specifically, extracting the portion comprising the biometric characteristic can comprise identifying pixel in the decoded output image that are within the distinguished portion and, after that, identifying the pixels in the decoded output image that are in the distinguished portion with corresponding pixels in the original image or the part of the image and extracting, from the original image or the part of the image, the corresponding pixels, the extracted corresponding pixels constituting the portion of the image or the part of the image comprising the biometric characteristic.--, in [0071]-[0075]; and, -- [0101] Separating the obtained image into grid cells with predefined bounding boxes allows for properly displaying and providing feedback on objects identified by using the bounding boxes in the final result to mark the location of the object and the object itself. [0102] In a more specific realization of this embodiment, the position of the bounding box is calculated relative to a center of the grid cell in two dimensions and the geometrical characteristics of the bounding box comprise a height and a width of the bounding box, wherein, further, a probability of the object being within the bounding box is associated with each bounding box. [0103] Associating the bounding boxes with corresponding probabilities allows for providing a matrix or vector that represents the bounding box and can be handled by graphical processing units with accurate efficiency when having to combine this with other objects that are represented in the form of a matrix or vector. Thereby, the required computer resources are reduced even further…. [0105] The resulting tensor can be processed by graphic processing units in a highly efficient manner. Additionally, providing the identification result in the form of such a tensor allows for easily deducing the results having the greatest probability for identifying a specific object. [0106] Moreover, outputting the output may comprise displaying the image and the resulting bounding boxes in each grid cell that have the highest probability among the bounding boxes in the grid cell.--, in [0102]-[0107]); and processing the cropped area to be processed by using the trained image processing network, to obtain a processing result for the image to be processed (see Gupta: e.g., -- [0436] The method of FIG. 12 begins with the provision of training data 1201. The training data may be constituted by a plurality of images of real objects as well as images of spoofs of real objects. For example, the images may comprise a number of images of real hands or fingers or the like and images of images (i.e. spoofs) of those objects. The images may be multiplied by using, from the same image, rotated, highlighted, darkened, enlarged or otherwise modified copies that are introduced as training data. In some embodiments, modifications involving image flips, image rotation and translation, shears, crops, multiplication to increase brightness and Gaussian blurs may be used to obtain a larger number of training images. Arbitrary combinations of the mentioned techniques may also be used. The values σ.sub.1 and σ.sub.2 provided according to item 1208 are the values indicating the “correct” output of the first node and second node of the last layer in the neural network that provide the probability of the image showing a spoof of an object or a real object. These values are provided for each image in the training data.--, in [0436]; and, -- specifically extracting a region of interest from the original image obtained and processed during the steps described above and before data comprising the biometric characteristic is sent to the third party computing device. In fact, with the embodiment now described, one embodiment of obtaining image data that only comprises the region of interest, like the fingertip carrying the fingerprint of interest, is described. [0448] This embodiment focuses on the extraction of a region of interest (ROI) from an image taken from an object of the user, where the image comprises a biometric characteristic that could be used to identify the user, but could also be used for any other purpose on the third party computing device. Such an object can be anything like a fingertip of one or more fingers of a hand of the user, the palm creases of a hand of the user or the face of the user or even the eye or the foot of the user. Each of these entities is known to carry biometric characteristics that can be used for identifying a user. [0449]… as a “cropping step” which separates a part of an image from another part of the image. This method is performed on the image that was already obtained during step 1 of the method according to FIG. 1. The step may occur after the identification step 2 and after the liveliness-detection and comparison steps 3 and 4. However, it can also occur in parallel to steps 3 and 4 or immediately after step 2. In some embodiments, it does not occur before step 2 since, in step 2 of FIG. 1, it is just determined whether there is an object carrying a biometric characteristic in the image at all. Processing the image in line with the embodiment to be described now would thus not be reasonable and potentially a waste of resources if it were performed before step 2 in FIG. 1.--, in [0447]-[0449], [0453]-[0457]; and, -- [0327] As explained above, it is assumed that the neural network is already perfectly learned for identifying a specific object, such as a fingertip. This involves that the neural network is able to identify a specific pattern of pixels that are most likely representing a fingertip. This might refer to specific patterns of color values or other characteristics like the brightness of those spots. It is, however, clear that the image 300 may arbitrarily show a fingertip which might not correspond in size and arrangement to a fingertip that was used for learning the neural network. [0328] With the help of the bounding boxes and the grid, however, it is possible for the neural network to identify the specific bounding box that will most likely comprise the fingertip. In order to identify this specific bounding box, the neural network (or an associated component that processes the image 300) compares the values of the pixels within each bounding box of each grid cell to a pattern of pixels that corresponds to a fingertip as was previously learned by the neural network. In this first stage, it is most unlikely that a perfect match will be found but there will be bounding boxes that are already more likely to contain at least a portion of a fingertip than other bounding boxes. [0329]…depicted in FIG. 6, for example, the bounding box 341 centered around the point M in grid cell 313 includes a portion of the fingertip of the hand 350. In contrast to this, none of the grid cells 310 and 311 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 341 and potentially the bounding box 340, the process can determine that the bounding box 341 includes even more of a pattern that corresponds to a fingertip than the bounding box 340. [0330] In view of this, the method can conclude that none of the bounding boxes 331 and 332 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0331] As both bounding boxes 340 and 341 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0332] While the smaller grid cell 340 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 341 may be regarded by the process to include a pattern that corresponds to a fingertip. [0333] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 341 and 340 to a pattern obtained from learning which indeed corresponds to a fingertip. [0334] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 341 is used as the starting point and its position and shapes modified or the smaller bounding box 340 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern….[0337] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0338] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 300 that contain the respective fingertip.--, in [0329]-[0338], and, --depicted in FIG. 18, for example, the bounding box 1841 centered around the point M in grid cell 1813 includes a portion of the fingertip of the hand 1850. In contrast to this, none of the grid cells 1810 and 1811 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 1841 and potentially the bounding box 1840, the process can determine that the bounding box 1841 includes even more of a pattern that corresponds to a fingertip than the bounding box 1840. [0565] In view of this, the method can conclude that none of the bounding boxes 1831 and 1832 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0566] As both bounding boxes 1840 and 1841 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0567] While the smaller grid cell 1840 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 1841 may be regarded by the process to include a pattern that corresponds to a fingertip. [0568] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 1841 and 1840 to a pattern obtained from learning which indeed corresponds to a fingertip. [0569] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 1841 is used as the starting point and its position and shapes modified or the smaller bounding box 1840 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern. [0570] This minimizing process can firstly comprise modifying the position of the bounding box (in the following, it will be assumed that the bounding box 1841 is used for the further calculations) by moving it a small amount into orthogonal directions first along the x-axis and then along the y-axis (or vice versa) as depicted in FIG. 18 around the center point M of the respective grid cell. The movement will be along the positive and the negative x-axis and y-axis and at each position, a comparison will be made to determine a difference function between the pattern obtained from the learning and the actual pattern identified in the image. This allows for calculating a two-dimensional function that represents the difference d(x,y) depending on the coordinates…..[0572] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0573] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 1800 that contain the respective fingertip or object carrying the biometric characteristic.--, in [0564]-[0573]; also see: -- [0071] It may be that the cropping operation comprises processing the image or the part of the image by a neural network, wherein processing the image or the part of the image by the neural network comprises processing the image by an encoder to obtain an encoded image and, after that, processing the encoded image by a decoder to obtain a decoded output image. [0072] Furthermore, the image or the part of the image provided to the neural network in the cropping operation for processing may comprise N×M pixels and the encoded image comprises n×m pixels, where n<N,m<M and the decoded output image comprises N×M pixels. [0073] Reducing the number of pixels when encoding the image results in a loss of information. When enlarging the image once again during the decoding, however, the most relevant information to distinguish the ROI from other portions of the image can be more easily discernable as not relevant information or very detailed information that is not necessary for identifying ROI is faded out with this procedure. [0074] In this regard, it can also be provided that distinguishing a portion of the image or the part of the image comprises distinguishing a portion of the decoded output image from another portion of the decoded output image. The distinguishing of the portions in the decoded image can be much easier compared to distinguishing the portion comprising the ROI from another portion of the original image. Thus, the processing power required for distinguishing a portion in the decoded output image from another portion in the decoded output image are reduced significantly compared to processing the original input image. [0075] More specifically, extracting the portion comprising the biometric characteristic can comprise identifying pixel in the decoded output image that are within the distinguished portion and, after that, identifying the pixels in the decoded output image that are in the distinguished portion with corresponding pixels in the original image or the part of the image and extracting, from the original image or the part of the image, the corresponding pixels, the extracted corresponding pixels constituting the portion of the image or the part of the image comprising the biometric characteristic.--, in [0071]-[0075]; and, -- [0101] Separating the obtained image into grid cells with predefined bounding boxes allows for properly displaying and providing feedback on objects identified by using the bounding boxes in the final result to mark the location of the object and the object itself. [0102] In a more specific realization of this embodiment, the position of the bounding box is calculated relative to a center of the grid cell in two dimensions and the geometrical characteristics of the bounding box comprise a height and a width of the bounding box, wherein, further, a probability of the object being within the bounding box is associated with each bounding box. [0103] Associating the bounding boxes with corresponding probabilities allows for providing a matrix or vector that represents the bounding box and can be handled by graphical processing units with accurate efficiency when having to combine this with other objects that are represented in the form of a matrix or vector. Thereby, the required computer resources are reduced even further…. [0105] The resulting tensor can be processed by graphic processing units in a highly efficient manner. Additionally, providing the identification result in the form of such a tensor allows for easily deducing the results having the greatest probability for identifying a specific object. [0106] Moreover, outputting the output may comprise displaying the image and the resulting bounding boxes in each grid cell that have the highest probability among the bounding boxes in the grid cell.--, in [0102]-[0107]). Re Claims 15 and 17, claims 15 and 17 are the corresponding device claims to claims 1 and 10 respectively. Thus, claims 15, and 17 are rejected for the similar reasons as for claims 1, and 10. Furthermore, GUPTA as modified by SARGENT further disclose a device for training an image processing network, comprising a memory and a processor, wherein the memory is stored with a computer program executable on the processor, and the processor is configured to execute the computer program to the method (see GUPTA: e.g., --[0017] FIG. 7 shows the process of training the neural network in the identification operation--, and, --[0022] FIG. 12 shows the process of training the neural network used in the liveliness-detection operation--, and, --[0029] FIG. 19 presents a flowchart of the training of the neural network used in the cropping operation--, in [0017]-[0029]; and, --[0065] The neural network, according to some embodiments, is trained and/or otherwise adapted to be able to distinguish, through processing of the image, between those portions of the image that contain the region of interest and other portions of the image that do not contain the region of interest. This can be achieved in variety of ways and is thus not to be understood in a limiting way. That the neural network “distinguishes” that region comprising the ROI from another region is to be understood as the ability of the neural network to provide an output that distinguishes between the ROI and other regions of the image or makes it possible to distinguish between those regions. For example, the output could be an indication of pixels comprising the ROI but not other regions of the ROI. In any case, the outcome of the processing of the image by the neural network at least comprises that a first portion of the image comprising the region of interest is identified as different from another portion of the image. In this regard, it is noted that the specific size, shape of position of the region of interest is found out by the neural network during processing of the image and it is not preset.--, in [0065]). Re Claim 18, claim 18 the corresponding device claims to claim 13 respectively. Thus, claim 18 is rejected for the similar reasons as for claim 13. Furthermore, GUPTA as modified by SARGENT further disclose a device for training an image processing network, comprising a memory and a processor, wherein the memory is stored with a computer program executable on the processor, and the processor is configured to execute the computer program to the method (see GUPTA: e.g., --[0017] FIG. 7 shows the process of training the neural network in the identification operation--, and, --[0022] FIG. 12 shows the process of training the neural network used in the liveliness-detection operation--, and, --[0029] FIG. 19 presents a flowchart of the training of the neural network used in the cropping operation--, in [0017]-[0029]; and, --[0065] The neural network, according to some embodiments, is trained and/or otherwise adapted to be able to distinguish, through processing of the image, between those portions of the image that contain the region of interest and other portions of the image that do not contain the region of interest. This can be achieved in variety of ways and is thus not to be understood in a limiting way. That the neural network “distinguishes” that region comprising the ROI from another region is to be understood as the ability of the neural network to provide an output that distinguishes between the ROI and other regions of the image or makes it possible to distinguish between those regions. For example, the output could be an indication of pixels comprising the ROI but not other regions of the ROI. In any case, the outcome of the processing of the image by the neural network at least comprises that a first portion of the image comprising the region of interest is identified as different from another portion of the image. In this regard, it is noted that the specific size, shape of position of the region of interest is found out by the neural network during processing of the image and it is not preset.--, in [0065]). Re Claims 19-20, claims 19-20 are the corresponding medium claims to claims 1, and 13 respectively. Thus, claim 19-20 are rejected for the similar reasons as for claim, 1 and 13. Furthermore, GUPTA as modified by SARGENT and further disclose non-transitory computer-readable storage medium having stored thereon a computer program that, when executed by a processor, causes the processor to implement a method for training an image processing network (see GUPTA: e.g., --[0017] FIG. 7 shows the process of training the neural network in the identification operation--, and, --[0022] FIG. 12 shows the process of training the neural network used in the liveliness-detection operation--, and, --[0029] FIG. 19 presents a flowchart of the training of the neural network used in the cropping operation--, in [0017]-[0029]; and, --[0065] The neural network, according to some embodiments, is trained and/or otherwise adapted to be able to distinguish, through processing of the image, between those portions of the image that contain the region of interest and other portions of the image that do not contain the region of interest. This can be achieved in variety of ways and is thus not to be understood in a limiting way. That the neural network “distinguishes” that region comprising the ROI from another region is to be understood as the ability of the neural network to provide an output that distinguishes between the ROI and other regions of the image or makes it possible to distinguish between those regions. For example, the output could be an indication of pixels comprising the ROI but not other regions of the ROI. In any case, the outcome of the processing of the image by the neural network at least comprises that a first portion of the image comprising the region of interest is identified as different from another portion of the image. In this regard, it is noted that the specific size, shape of position of the region of interest is found out by the neural network during processing of the image and it is not preset.--, in [0065], and, -- this training is performed before the mobile device is actually equipped with an application or other program that can perform the respective method according to the above embodiments,--, in [0113]-[0114]). Claims 2-7, 14, and 16 is rejected under 35 U.S.C. 103 as being unpatentable over GUPTA as modified by SARGENT, and further in view of ZHANG (US 11681364 B1). Re Claim 2, GUPTA as modified by SARGENT however do not explicitly disclose the training image comprises a face image, and the truth value is annotated gaze information in the face image; Zhang discloses the training image comprises a face image, and the truth value is annotated gaze information in the face image (see Zhang: e.g., -- An image processing system may receive image data from a camera of a user device and perform gaze prediction processing of the image data to predict one or more gaze patterns. The gaze prediction processing may include processing the image data using a neural network to detect faces and/or objects and generate an image feature map. The gaze prediction processing may include performing gaze direction prediction operations using the feature map and detected faces and/or objects to determine gaze direction probability data. The gaze prediction processing may include predicting a gaze pattern based on the gaze direction probability data and the image feature map. The gaze pattern may be short-term (e.g., atomic-level) or long-term (e.g., event-level).--, in abstract, and, --(19)… a user's gaze behavior may indicate attention towards a person or an object, or an attempt to attract a person's attention to the user or to an object. Gaze behavior information may help a system understand or otherwise determine what the user is looking at, and how the user feels, or what the user is referring to. A user interface system, such as a multi-modal user interface system that processes both speech and human gaze behavior may perform better when interacting with multiple users, who may also be interacting with each other. (20) Gaze behaviors may be represented by, or categorized into, gaze patterns. The user interface system may include one or more models trained to recognize various gaze patterns, and predict whether a person's gaze behavior as exhibited in image data corresponds to one of the learned gaze patterns. Gaze patterns may be categorized based on duration. “Atomic-level gaze patterns” may be brief in duration (e.g., lasting one or several image frames up to several seconds). Some atomic-level gaze patterns may include: Single: individual gaze behavior (e.g., directed toward the device or a person/object) Mutual: two people look into eyes of each other Avert: one person shifts away the gaze from another person's gaze Refer: one person introduces one target for another person using gaze Follow: one person accepts the gaze of another person and looking at the target referred by that person Share: two people are looking at the same target (21) “Event-level gaze patterns” may be longer in duration (e.g., longer than a second or several seconds), and may include one atomic-level gaze patterns followed by another. Some event-level gaze patterns may include: Non-communicative: Single gaze held for a period of time Mutual Gaze: Mutual gaze held for a period of time Gaze Aversion: starts with Mutual, follows by Avert, and finally Single Gaze Following: Follow and Share without Mutual Joint Attention: Starts with Mutual, then Refer, Follow and Share, and ends with Mutual again to confirm the Share event.--, in line 32, col. 2 through line 12, col. 3; and, -- The decoder portion of the network may take the collection of feature vectors and up-sample them using position information from the combined feature map 245 to produce a segmented output that represents a semantic segmentation of the input. In this case, the encoder-decoder network 250 can be trained to produce a gaze direction probability map 134. The gaze direction probability map 134 can include segments corresponding to bounding boxes around entities; for example, the first region 234a corresponding to the person and the second region 234b corresponding to the object. Values corresponding to those segments can represent a probability that the POI is gazing in the direction of that physical entity. In the example shown in FIGS. 1 and 2, the gaze direction probability map 134 may reflect a higher probability (e.g., closer to 1) that the first user 5a has a gaze directed toward the object 10, relative to a lower probability (e.g., closer to 0) that the first use 5a has a gaze directed toward the person (e.g., the second user 5b).--, in lines 16-40, col. 8; and, -- (47) The NN 131, the LSTM 380 and the LSTM 480 may be trained using annotated data sets. In some implementations, the models may be trained using the publicly available VACATION (Video gAze CommunicATION) dataset. The VACATION dataset includes 300 videos cropped from TV shows and movies from YouTube. The videos are stored in MPEG4 format with 640×360 spatial resolution. The length of individual videos ranges from 2.2 to 74.5 seconds, and average 13 seconds. The dataset includes a total of 96,993 frames and is 3,880 seconds long. The dataset includes annotations for, for example, human face and object bounding boxes, the attention of each person in the image (that is, to which person/object it directs its gaze) and both atomic-level and event-level gaze labels for each person in each frame. The NN 131 model may be trained using an Adam optimizer with a learning rate of 0.0001. Multi-class entropy may be chosen as the training loss. The network may learn over 10 epochs, taking 10 hours using a commercially available graphics card. The best model is chosen based on accuracy in processing a validation set. The LSTM 380 may be trained using information from 10 successive frames. During training of the second LSTM 480, more successive frames (e.g., 30, 60, or more) may be used to account for the relatively longer time period over which event-level gaze patterns may occur.--, in lines 1-25, col. 12); GUPTA (as modified by SARGENT) and ZHANG are combinable as they are in the same field of endeavor: training image processing neural network for classification and prediction including object recognition, segmentation, and bounding box identification. Therefore it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify GUPTA (as modified by SARGENT)’s method using ZHANG’s teachings by including the training image comprises a face image, and the truth value is annotated gaze information in the face image to GUPTA (as modified by SARGENT)’s training neural network, artificial intelligence model and training data in order to recognize various gaze patterns, and predict whether a person's gaze behavior as exhibited in image data corresponds to one of the learned gaze patterns (see ZHANG: e.g. in line 32, col. 2 through line 12, col. 3, and in lines 1-25, col. 12). Re Claim 3, GUPTA as modified by SARGENT and ZHANG further disclose wherein the annotated gaze information comprises at least one of: a pitch angle of a gaze, a yaw angle of the gaze, or a roll angle of the gaze (see Zhang: e.g., -- An image processing system may receive image data from a camera of a user device and perform gaze prediction processing of the image data to predict one or more gaze patterns. The gaze prediction processing may include processing the image data using a neural network to detect faces and/or objects and generate an image feature map. The gaze prediction processing may include performing gaze direction prediction operations using the feature map and detected faces and/or objects to determine gaze direction probability data. The gaze prediction processing may include predicting a gaze pattern based on the gaze direction probability data and the image feature map. The gaze pattern may be short-term (e.g., atomic-level) or long-term (e.g., event-level).--, in abstract, and, --(19)… a user's gaze behavior may indicate attention towards a person or an object, or an attempt to attract a person's attention to the user or to an object. Gaze behavior information may help a system understand or otherwise determine what the user is looking at, and how the user feels, or what the user is referring to. A user interface system, such as a multi-modal user interface system that processes both speech and human gaze behavior may perform better when interacting with multiple users, who may also be interacting with each other. (20) Gaze behaviors may be represented by, or categorized into, gaze patterns. The user interface system may include one or more models trained to recognize various gaze patterns, and predict whether a person's gaze behavior as exhibited in image data corresponds to one of the learned gaze patterns. Gaze patterns may be categorized based on duration. “Atomic-level gaze patterns” may be brief in duration (e.g., lasting one or several image frames up to several seconds). Some atomic-level gaze patterns may include: Single: individual gaze behavior (e.g., directed toward the device or a person/object) Mutual: two people look into eyes of each other Avert: one person shifts away the gaze from another person's gaze Refer: one person introduces one target for another person using gaze Follow: one person accepts the gaze of another person and looking at the target referred by that person Share: two people are looking at the same target (21) “Event-level gaze patterns” may be longer in duration (e.g., longer than a second or several seconds), and may include one atomic-level gaze patterns followed by another. Some event-level gaze patterns may include: Non-communicative: Single gaze held for a period of time Mutual Gaze: Mutual gaze held for a period of time Gaze Aversion: starts with Mutual, follows by Avert, and finally Single Gaze Following: Follow and Share without Mutual Joint Attention: Starts with Mutual, then Refer, Follow and Share, and ends with Mutual again to confirm the Share event.--, in line 32, col. 2 through line 12, col. 3; and, -- The decoder portion of the network may take the collection of feature vectors and up-sample them using position information from the combined feature map 245 to produce a segmented output that represents a semantic segmentation of the input. In this case, the encoder-decoder network 250 can be trained to produce a gaze direction probability map 134. The gaze direction probability map 134 can include segments corresponding to bounding boxes around entities; for example, the first region 234a corresponding to the person and the second region 234b corresponding to the object. Values corresponding to those segments can represent a probability that the POI is gazing in the direction of that physical entity. In the example shown in FIGS. 1 and 2, the gaze direction probability map 134 may reflect a higher probability (e.g., closer to 1) that the first user 5a has a gaze directed toward the object 10, relative to a lower probability (e.g., closer to 0) that the first use 5a has a gaze directed toward the person (e.g., the second user 5b).--, in lines 16-40, col. 8; and, -- (47) The NN 131, the LSTM 380 and the LSTM 480 may be trained using annotated data sets. In some implementations, the models may be trained using the publicly available VACATION (Video gAze CommunicATION) dataset. The VACATION dataset includes 300 videos cropped from TV shows and movies from YouTube. The videos are stored in MPEG4 format with 640×360 spatial resolution. The length of individual videos ranges from 2.2 to 74.5 seconds, and average 13 seconds. The dataset includes a total of 96,993 frames and is 3,880 seconds long. The dataset includes annotations for, for example, human face and object bounding boxes, the attention of each person in the image (that is, to which person/object it directs its gaze) and both atomic-level and event-level gaze labels for each person in each frame. The NN 131 model may be trained using an Adam optimizer with a learning rate of 0.0001. Multi-class entropy may be chosen as the training loss. The network may learn over 10 epochs, taking 10 hours using a commercially available graphics card. The best model is chosen based on accuracy in processing a validation set. The LSTM 380 may be trained using information from 10 successive frames. During training of the second LSTM 480, more successive frames (e.g., 30, 60, or more) may be used to account for the relatively longer time period over which event-level gaze patterns may occur.--, in lines 1-25, col. 12). Re Claim 4, GUPTA as modified by SARGENT and ZHANG further disclose wherein determining the reference pixel based on the training image annotated with the truth value comprises: determining a center pixel of the training image as the reference pixel; and wherein determining, with the reference pixel as the starting point and based on the Markov chain of the training image (see Gupta: e.g., -- [0327] As explained above, it is assumed that the neural network is already perfectly learned for identifying a specific object, such as a fingertip. This involves that the neural network is able to identify a specific pattern of pixels that are most likely representing a fingertip. This might refer to specific patterns of color values or other characteristics like the brightness of those spots. It is, however, clear that the image 300 may arbitrarily show a fingertip which might not correspond in size and arrangement to a fingertip that was used for learning the neural network. [0328] With the help of the bounding boxes and the grid, however, it is possible for the neural network to identify the specific bounding box that will most likely comprise the fingertip. In order to identify this specific bounding box, the neural network (or an associated component that processes the image 300) compares the values of the pixels within each bounding box of each grid cell to a pattern of pixels that corresponds to a fingertip as was previously learned by the neural network. In this first stage, it is most unlikely that a perfect match will be found but there will be bounding boxes that are already more likely to contain at least a portion of a fingertip than other bounding boxes. [0329]…depicted in FIG. 6, for example, the bounding box 341 centered around the point M in grid cell 313 includes a portion of the fingertip of the hand 350. In contrast to this, none of the grid cells 310 and 311 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 341 and potentially the bounding box 340, the process can determine that the bounding box 341 includes even more of a pattern that corresponds to a fingertip than the bounding box 340. [0330] In view of this, the method can conclude that none of the bounding boxes 331 and 332 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0331] As both bounding boxes 340 and 341 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0332] While the smaller grid cell 340 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 341 may be regarded by the process to include a pattern that corresponds to a fingertip. [0333] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 341 and 340 to a pattern obtained from learning which indeed corresponds to a fingertip. [0334] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 341 is used as the starting point and its position and shapes modified or the smaller bounding box 340 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern….[0337] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0338] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 300 that contain the respective fingertip.--, in [0329]-[0338], and, --depicted in FIG. 18, for example, the bounding box 1841 centered around the point M in grid cell 1813 includes a portion of the fingertip of the hand 1850. In contrast to this, none of the grid cells 1810 and 1811 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 1841 and potentially the bounding box 1840, the process can determine that the bounding box 1841 includes even more of a pattern that corresponds to a fingertip than the bounding box 1840. [0565] In view of this, the method can conclude that none of the bounding boxes 1831 and 1832 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0566] As both bounding boxes 1840 and 1841 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0567] While the smaller grid cell 1840 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 1841 may be regarded by the process to include a pattern that corresponds to a fingertip. [0568] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 1841 and 1840 to a pattern obtained from learning which indeed corresponds to a fingertip. [0569] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 1841 is used as the starting point and its position and shapes modified or the smaller bounding box 1840 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern. [0570] This minimizing process can firstly comprise modifying the position of the bounding box (in the following, it will be assumed that the bounding box 1841 is used for the further calculations) by moving it a small amount into orthogonal directions first along the x-axis and then along the y-axis (or vice versa) as depicted in FIG. 18 around the center point M of the respective grid cell. The movement will be along the positive and the negative x-axis and y-axis and at each position, a comparison will be made to determine a difference function between the pattern obtained from the learning and the actual pattern identified in the image. This allows for calculating a two-dimensional function that represents the difference d(x,y) depending on the coordinates…..[0572] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0573] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 1800 that contain the respective fingertip or object carrying the biometric characteristic.--, in [0564]-[0573]; also see: -- [0071] It may be that the cropping operation comprises processing the image or the part of the image by a neural network, wherein processing the image or the part of the image by the neural network comprises processing the image by an encoder to obtain an encoded image and, after that, processing the encoded image by a decoder to obtain a decoded output image. [0072] Furthermore, the image or the part of the image provided to the neural network in the cropping operation for processing may comprise N×M pixels and the encoded image comprises n×m pixels, where n<N,m<M and the decoded output image comprises N×M pixels. [0073] Reducing the number of pixels when encoding the image results in a loss of information. When enlarging the image once again during the decoding, however, the most relevant information to distinguish the ROI from other portions of the image can be more easily discernable as not relevant information or very detailed information that is not necessary for identifying ROI is faded out with this procedure. [0074] In this regard, it can also be provided that distinguishing a portion of the image or the part of the image comprises distinguishing a portion of the decoded output image from another portion of the decoded output image. The distinguishing of the portions in the decoded image can be much easier compared to distinguishing the portion comprising the ROI from another portion of the original image. Thus, the processing power required for distinguishing a portion in the decoded output image from another portion in the decoded output image are reduced significantly compared to processing the original input image. [0075] More specifically, extracting the portion comprising the biometric characteristic can comprise identifying pixel in the decoded output image that are within the distinguished portion and, after that, identifying the pixels in the decoded output image that are in the distinguished portion with corresponding pixels in the original image or the part of the image and extracting, from the original image or the part of the image, the corresponding pixels, the extracted corresponding pixels constituting the portion of the image or the part of the image comprising the biometric characteristic.--, in [0071]-[0075]; and, -- [0101] Separating the obtained image into grid cells with predefined bounding boxes allows for properly displaying and providing feedback on objects identified by using the bounding boxes in the final result to mark the location of the object and the object itself. [0102] In a more specific realization of this embodiment, the position of the bounding box is calculated relative to a center of the grid cell in two dimensions and the geometrical characteristics of the bounding box comprise a height and a width of the bounding box, wherein, further, a probability of the object being within the bounding box is associated with each bounding box. [0103] Associating the bounding boxes with corresponding probabilities allows for providing a matrix or vector that represents the bounding box and can be handled by graphical processing units with accurate efficiency when having to combine this with other objects that are represented in the form of a matrix or vector. Thereby, the required computer resources are reduced even further…. [0105] The resulting tensor can be processed by graphic processing units in a highly efficient manner. Additionally, providing the identification result in the form of such a tensor allows for easily deducing the results having the greatest probability for identifying a specific object. [0106] Moreover, outputting the output may comprise displaying the image and the resulting bounding boxes in each grid cell that have the highest probability among the bounding boxes in the grid cell.--, in [0102]-[0107]; see SARGENT: e.g., -- Determining the object convolutional positions may include finding a moment bounding box of each object, the moment bounding box being a minimum bounding rectangle surrounding an object and aligned with the orientation of the major axis of the object, the object convolutional positions being determined in dependence on properties of the bounding box. Properties of the moment bounding box may be used to determine, for an object, a single first object convolutional position which should be used to represent the object when being processed by the first neural network, and further to select, for the object, a plurality of second object convolutional positions distributed across the bounding box which should be used to represent the object when being processed by the second neural network…. Figure 13 displays joint deep learning with joint distribution modelling (a) through iterative process for pixel-level cover (LC) and patch-based land use (LU) extraction and decision-making (b); and Figure 14 is a flow diagram of the training of the JDL model. --, in [0019]-[0020], and, [0026]; and, -- The image of land is passed into a first machine learning network which is used to calculate LC probabilities on a per-pixel basis across the image. The LC probabilities are then fed into a second machine learning network to calculate LU probabilities for each identified object. The LU probability is then fed back into the first machine learning network, along with the original image, and the process using the first and second machine learning networks is iterated over. This iteration process produces a highly accurate LC classification for each pixel in the image, and a LU classification for each identified object. In more detail, the described method simultaneously determines land cover (LC) and land use (LU) classifications from remotely sensed images using two machine learning networks, a multilayer perceptron (MLP) and a convolutional neural network (CNN), which together form a joint deep learning (JDL) model. The JDL model is implemented via a Markov process involving iterating between the two machine learning networks, with the output of one of the networks being fed as an input for the next iteration of the other of the networks, and vice versa. The CNN is trained with labelled image patches for the LU classes.--, in [0027]-[0028]; and, -- The assumption of the LC - LU joint deep learning (LC-LU JDL) model is that both LC and LU are manifested over same geographical space and are nested with each other in a hierarchical manner. The LC and LU representations are considered as two random variables, where the probabilistic relationship between them can be modelled through a joint probability distribution. In this way, the conditional dependencies between these two random variables are captured via an undirected graph through iteration (i.e. formulating a Markov process). The joint distribution is, thus, factorised as a product of the individual density functions,--, in [0044]; and, -- Figure 3 illustrates the general workflow of the described LC and LU joint deep learning (LC-LU JDL) model, with key components including the JDL inputs, the Markov Process to learn the joint distribution, and the classification outputs of LC and LU at each iteration. Detailed explanation is given as follows. JDL input involves LC samples with pixel locations and the corresponding land cover labels, LU samples with image patches representing specific land use categories, together with the remotely sensed imagery, and the object-based segmentation results with unique identity for each segment. These four elements were used to infer the hierarchical relationships between LC and LU, and to obtain LC and LU classification results through iteration. Markov Process models the joint probability distribution between LC and LU through iteration--, in [0049]-[0051]). , the cropping probabilities of the image processing network processing the training image comprises: determining, with the center pixel as the starting point and based on the Markov chain, a cropping probability for each pixel in the training image (see SARGENT: e.g., -- Determining the object convolutional positions may include finding a moment bounding box of each object, the moment bounding box being a minimum bounding rectangle surrounding an object and aligned with the orientation of the major axis of the object, the object convolutional positions being determined in dependence on properties of the bounding box. Properties of the moment bounding box may be used to determine, for an object, a single first object convolutional position which should be used to represent the object when being processed by the first neural network, and further to select, for the object, a plurality of second object convolutional positions distributed across the bounding box which should be used to represent the object when being processed by the second neural network…. Figure 13 displays joint deep learning with joint distribution modelling (a) through iterative process for pixel-level cover (LC) and patch-based land use (LU) extraction and decision-making (b); and Figure 14 is a flow diagram of the training of the JDL model. --, in [0019]-[0020], and, [0026]; and, -- The image of land is passed into a first machine learning network which is used to calculate LC probabilities on a per-pixel basis across the image. The LC probabilities are then fed into a second machine learning network to calculate LU probabilities for each identified object. The LU probability is then fed back into the first machine learning network, along with the original image, and the process using the first and second machine learning networks is iterated over. This iteration process produces a highly accurate LC classification for each pixel in the image, and a LU classification for each identified object. In more detail, the described method simultaneously determines land cover (LC) and land use (LU) classifications from remotely sensed images using two machine learning networks, a multilayer perceptron (MLP) and a convolutional neural network (CNN), which together form a joint deep learning (JDL) model. The JDL model is implemented via a Markov process involving iterating between the two machine learning networks, with the output of one of the networks being fed as an input for the next iteration of the other of the networks, and vice versa. The CNN is trained with labelled image patches for the LU classes.--, in [0027]-[0028]; and, -- The assumption of the LC - LU joint deep learning (LC-LU JDL) model is that both LC and LU are manifested over same geographical space and are nested with each other in a hierarchical manner. The LC and LU representations are considered as two random variables, where the probabilistic relationship between them can be modelled through a joint probability distribution. In this way, the conditional dependencies between these two random variables are captured via an undirected graph through iteration (i.e. formulating a Markov process). The joint distribution is, thus, factorised as a product of the individual density functions,--, in [0044]; and, -- Figure 3 illustrates the general workflow of the described LC and LU joint deep learning (LC-LU JDL) model, with key components including the JDL inputs, the Markov Process to learn the joint distribution, and the classification outputs of LC and LU at each iteration. Detailed explanation is given as follows. JDL input involves LC samples with pixel locations and the corresponding land cover labels, LU samples with image patches representing specific land use categories, together with the remotely sensed imagery, and the object-based segmentation results with unique identity for each segment. These four elements were used to infer the hierarchical relationships between LC and LU, and to obtain LC and LU classification results through iteration. Markov Process models the joint probability distribution between LC and LU through iteration--, in [0049]-[0051]). Re Claim 5, GUPTA as modified by SARGENT and ZHANG further disclose wherein determining, with the center pixel as the starting point and based on the Markov chain, the cropping probability for each pixel in the training image comprises: determining a transition probability of a next pixel from the center pixel in the Markov chain (see Gupta: e.g., -- [0436] The method of FIG. 12 begins with the provision of training data 1201. The training data may be constituted by a plurality of images of real objects as well as images of spoofs of real objects. For example, the images may comprise a number of images of real hands or fingers or the like and images of images (i.e. spoofs) of those objects. The images may be multiplied by using, from the same image, rotated, highlighted, darkened, enlarged or otherwise modified copies that are introduced as training data. In some embodiments, modifications involving image flips, image rotation and translation, shears, crops, multiplication to increase brightness and Gaussian blurs may be used to obtain a larger number of training images. Arbitrary combinations of the mentioned techniques may also be used. The values σ.sub.1 and σ.sub.2 provided according to item 1208 are the values indicating the “correct” output of the first node and second node of the last layer in the neural network that provide the probability of the image showing a spoof of an object or a real object. These values are provided for each image in the training data.--, in [0436]; and, -- specifically extracting a region of interest from the original image obtained and processed during the steps described above and before data comprising the biometric characteristic is sent to the third party computing device. In fact, with the embodiment now described, one embodiment of obtaining image data that only comprises the region of interest, like the fingertip carrying the fingerprint of interest, is described. [0448] This embodiment focuses on the extraction of a region of interest (ROI) from an image taken from an object of the user, where the image comprises a biometric characteristic that could be used to identify the user, but could also be used for any other purpose on the third party computing device. Such an object can be anything like a fingertip of one or more fingers of a hand of the user, the palm creases of a hand of the user or the face of the user or even the eye or the foot of the user. Each of these entities is known to carry biometric characteristics that can be used for identifying a user. [0449]… as a “cropping step” which separates a part of an image from another part of the image. This method is performed on the image that was already obtained during step 1 of the method according to FIG. 1. The step may occur after the identification step 2 and after the liveliness-detection and comparison steps 3 and 4. However, it can also occur in parallel to steps 3 and 4 or immediately after step 2. In some embodiments, it does not occur before step 2 since, in step 2 of FIG. 1, it is just determined whether there is an object carrying a biometric characteristic in the image at all. Processing the image in line with the embodiment to be described now would thus not be reasonable and potentially a waste of resources if it were performed before step 2 in FIG. 1.--, in [0447]-[0449], [0453]-[0457]; and, -- [0327] As explained above, it is assumed that the neural network is already perfectly learned for identifying a specific object, such as a fingertip. This involves that the neural network is able to identify a specific pattern of pixels that are most likely representing a fingertip. This might refer to specific patterns of color values or other characteristics like the brightness of those spots. It is, however, clear that the image 300 may arbitrarily show a fingertip which might not correspond in size and arrangement to a fingertip that was used for learning the neural network. [0328] With the help of the bounding boxes and the grid, however, it is possible for the neural network to identify the specific bounding box that will most likely comprise the fingertip. In order to identify this specific bounding box, the neural network (or an associated component that processes the image 300) compares the values of the pixels within each bounding box of each grid cell to a pattern of pixels that corresponds to a fingertip as was previously learned by the neural network. In this first stage, it is most unlikely that a perfect match will be found but there will be bounding boxes that are already more likely to contain at least a portion of a fingertip than other bounding boxes. [0329]…depicted in FIG. 6, for example, the bounding box 341 centered around the point M in grid cell 313 includes a portion of the fingertip of the hand 350. In contrast to this, none of the grid cells 310 and 311 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 341 and potentially the bounding box 340, the process can determine that the bounding box 341 includes even more of a pattern that corresponds to a fingertip than the bounding box 340. [0330] In view of this, the method can conclude that none of the bounding boxes 331 and 332 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0331] As both bounding boxes 340 and 341 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0332] While the smaller grid cell 340 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 341 may be regarded by the process to include a pattern that corresponds to a fingertip. [0333] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 341 and 340 to a pattern obtained from learning which indeed corresponds to a fingertip. [0334] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 341 is used as the starting point and its position and shapes modified or the smaller bounding box 340 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern….[0337] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0338] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 300 that contain the respective fingertip.--, in [0329]-[0338], and, --depicted in FIG. 18, for example, the bounding box 1841 centered around the point M in grid cell 1813 includes a portion of the fingertip of the hand 1850. In contrast to this, none of the grid cells 1810 and 1811 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 1841 and potentially the bounding box 1840, the process can determine that the bounding box 1841 includes even more of a pattern that corresponds to a fingertip than the bounding box 1840. [0565] In view of this, the method can conclude that none of the bounding boxes 1831 and 1832 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0566] As both bounding boxes 1840 and 1841 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0567] While the smaller grid cell 1840 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 1841 may be regarded by the process to include a pattern that corresponds to a fingertip. [0568] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 1841 and 1840 to a pattern obtained from learning which indeed corresponds to a fingertip. [0569] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 1841 is used as the starting point and its position and shapes modified or the smaller bounding box 1840 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern. [0570] This minimizing process can firstly comprise modifying the position of the bounding box (in the following, it will be assumed that the bounding box 1841 is used for the further calculations) by moving it a small amount into orthogonal directions first along the x-axis and then along the y-axis (or vice versa) as depicted in FIG. 18 around the center point M of the respective grid cell. The movement will be along the positive and the negative x-axis and y-axis and at each position, a comparison will be made to determine a difference function between the pattern obtained from the learning and the actual pattern identified in the image. This allows for calculating a two-dimensional function that represents the difference d(x,y) depending on the coordinates…..[0572] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0573] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 1800 that contain the respective fingertip or object carrying the biometric characteristic.--, in [0564]-[0573]; also see: -- [0071] It may be that the cropping operation comprises processing the image or the part of the image by a neural network, wherein processing the image or the part of the image by the neural network comprises processing the image by an encoder to obtain an encoded image and, after that, processing the encoded image by a decoder to obtain a decoded output image. [0072] Furthermore, the image or the part of the image provided to the neural network in the cropping operation for processing may comprise N×M pixels and the encoded image comprises n×m pixels, where n<N,m<M and the decoded output image comprises N×M pixels. [0073] Reducing the number of pixels when encoding the image results in a loss of information. When enlarging the image once again during the decoding, however, the most relevant information to distinguish the ROI from other portions of the image can be more easily discernable as not relevant information or very detailed information that is not necessary for identifying ROI is faded out with this procedure. [0074] In this regard, it can also be provided that distinguishing a portion of the image or the part of the image comprises distinguishing a portion of the decoded output image from another portion of the decoded output image. The distinguishing of the portions in the decoded image can be much easier compared to distinguishing the portion comprising the ROI from another portion of the original image. Thus, the processing power required for distinguishing a portion in the decoded output image from another portion in the decoded output image are reduced significantly compared to processing the original input image. [0075] More specifically, extracting the portion comprising the biometric characteristic can comprise identifying pixel in the decoded output image that are within the distinguished portion and, after that, identifying the pixels in the decoded output image that are in the distinguished portion with corresponding pixels in the original image or the part of the image and extracting, from the original image or the part of the image, the corresponding pixels, the extracted corresponding pixels constituting the portion of the image or the part of the image comprising the biometric characteristic.--, in [0071]-[0075]); and determining, based on the transition probability of the next pixel and transition probabilities of a plurality of pixels prior to the next pixel, a cropping probability for the next pixel (see Gupta: e.g., -- [0436] The method of FIG. 12 begins with the provision of training data 1201. The training data may be constituted by a plurality of images of real objects as well as images of spoofs of real objects. For example, the images may comprise a number of images of real hands or fingers or the like and images of images (i.e. spoofs) of those objects. The images may be multiplied by using, from the same image, rotated, highlighted, darkened, enlarged or otherwise modified copies that are introduced as training data. In some embodiments, modifications involving image flips, image rotation and translation, shears, crops, multiplication to increase brightness and Gaussian blurs may be used to obtain a larger number of training images. Arbitrary combinations of the mentioned techniques may also be used. The values σ.sub.1 and σ.sub.2 provided according to item 1208 are the values indicating the “correct” output of the first node and second node of the last layer in the neural network that provide the probability of the image showing a spoof of an object or a real object. These values are provided for each image in the training data.--, in [0436]; and, -- specifically extracting a region of interest from the original image obtained and processed during the steps described above and before data comprising the biometric characteristic is sent to the third party computing device. In fact, with the embodiment now described, one embodiment of obtaining image data that only comprises the region of interest, like the fingertip carrying the fingerprint of interest, is described. [0448] This embodiment focuses on the extraction of a region of interest (ROI) from an image taken from an object of the user, where the image comprises a biometric characteristic that could be used to identify the user, but could also be used for any other purpose on the third party computing device. Such an object can be anything like a fingertip of one or more fingers of a hand of the user, the palm creases of a hand of the user or the face of the user or even the eye or the foot of the user. Each of these entities is known to carry biometric characteristics that can be used for identifying a user. [0449]… as a “cropping step” which separates a part of an image from another part of the image. This method is performed on the image that was already obtained during step 1 of the method according to FIG. 1. The step may occur after the identification step 2 and after the liveliness-detection and comparison steps 3 and 4. However, it can also occur in parallel to steps 3 and 4 or immediately after step 2. In some embodiments, it does not occur before step 2 since, in step 2 of FIG. 1, it is just determined whether there is an object carrying a biometric characteristic in the image at all. Processing the image in line with the embodiment to be described now would thus not be reasonable and potentially a waste of resources if it were performed before step 2 in FIG. 1.--, in [0447]-[0449], [0453]-[0457]; and, -- [0327] As explained above, it is assumed that the neural network is already perfectly learned for identifying a specific object, such as a fingertip. This involves that the neural network is able to identify a specific pattern of pixels that are most likely representing a fingertip. This might refer to specific patterns of color values or other characteristics like the brightness of those spots. It is, however, clear that the image 300 may arbitrarily show a fingertip which might not correspond in size and arrangement to a fingertip that was used for learning the neural network. [0328] With the help of the bounding boxes and the grid, however, it is possible for the neural network to identify the specific bounding box that will most likely comprise the fingertip. In order to identify this specific bounding box, the neural network (or an associated component that processes the image 300) compares the values of the pixels within each bounding box of each grid cell to a pattern of pixels that corresponds to a fingertip as was previously learned by the neural network. In this first stage, it is most unlikely that a perfect match will be found but there will be bounding boxes that are already more likely to contain at least a portion of a fingertip than other bounding boxes. [0329]…depicted in FIG. 6, for example, the bounding box 341 centered around the point M in grid cell 313 includes a portion of the fingertip of the hand 350. In contrast to this, none of the grid cells 310 and 311 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 341 and potentially the bounding box 340, the process can determine that the bounding box 341 includes even more of a pattern that corresponds to a fingertip than the bounding box 340. [0330] In view of this, the method can conclude that none of the bounding boxes 331 and 332 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0331] As both bounding boxes 340 and 341 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0332] While the smaller grid cell 340 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 341 may be regarded by the process to include a pattern that corresponds to a fingertip. [0333] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 341 and 340 to a pattern obtained from learning which indeed corresponds to a fingertip. [0334] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 341 is used as the starting point and its position and shapes modified or the smaller bounding box 340 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern….[0337] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0338] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 300 that contain the respective fingertip.--, in [0329]-[0338], and, --depicted in FIG. 18, for example, the bounding box 1841 centered around the point M in grid cell 1813 includes a portion of the fingertip of the hand 1850. In contrast to this, none of the grid cells 1810 and 1811 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 1841 and potentially the bounding box 1840, the process can determine that the bounding box 1841 includes even more of a pattern that corresponds to a fingertip than the bounding box 1840. [0565] In view of this, the method can conclude that none of the bounding boxes 1831 and 1832 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0566] As both bounding boxes 1840 and 1841 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0567] While the smaller grid cell 1840 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 1841 may be regarded by the process to include a pattern that corresponds to a fingertip. [0568] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 1841 and 1840 to a pattern obtained from learning which indeed corresponds to a fingertip. [0569] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 1841 is used as the starting point and its position and shapes modified or the smaller bounding box 1840 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern. [0570] This minimizing process can firstly comprise modifying the position of the bounding box (in the following, it will be assumed that the bounding box 1841 is used for the further calculations) by moving it a small amount into orthogonal directions first along the x-axis and then along the y-axis (or vice versa) as depicted in FIG. 18 around the center point M of the respective grid cell. The movement will be along the positive and the negative x-axis and y-axis and at each position, a comparison will be made to determine a difference function between the pattern obtained from the learning and the actual pattern identified in the image. This allows for calculating a two-dimensional function that represents the difference d(x,y) depending on the coordinates…..[0572] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0573] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 1800 that contain the respective fingertip or object carrying the biometric characteristic.--, in [0564]-[0573]; also see: -- [0071] It may be that the cropping operation comprises processing the image or the part of the image by a neural network, wherein processing the image or the part of the image by the neural network comprises processing the image by an encoder to obtain an encoded image and, after that, processing the encoded image by a decoder to obtain a decoded output image. [0072] Furthermore, the image or the part of the image provided to the neural network in the cropping operation for processing may comprise N×M pixels and the encoded image comprises n×m pixels, where n<N,m<M and the decoded output image comprises N×M pixels. [0073] Reducing the number of pixels when encoding the image results in a loss of information. When enlarging the image once again during the decoding, however, the most relevant information to distinguish the ROI from other portions of the image can be more easily discernable as not relevant information or very detailed information that is not necessary for identifying ROI is faded out with this procedure. [0074] In this regard, it can also be provided that distinguishing a portion of the image or the part of the image comprises distinguishing a portion of the decoded output image from another portion of the decoded output image. The distinguishing of the portions in the decoded image can be much easier compared to distinguishing the portion comprising the ROI from another portion of the original image. Thus, the processing power required for distinguishing a portion in the decoded output image from another portion in the decoded output image are reduced significantly compared to processing the original input image. [0075] More specifically, extracting the portion comprising the biometric characteristic can comprise identifying pixel in the decoded output image that are within the distinguished portion and, after that, identifying the pixels in the decoded output image that are in the distinguished portion with corresponding pixels in the original image or the part of the image and extracting, from the original image or the part of the image, the corresponding pixels, the extracted corresponding pixels constituting the portion of the image or the part of the image comprising the biometric characteristic.--, in [0071]-[0075]). Re Claim 6, GUPTA as modified by SARGENT and ZHANG further disclose wherein determining, with the center pixel as the starting point and based on the Markov chain, the cropping probability for each pixel in the training image comprises: isotropically setting, based on a Markov chain in at least one direction starting from the center pixel, cropping probabilities of all pixels in the at least one direction starting from the center pixel (see Gupta: e.g., -- [0436] The method of FIG. 12 begins with the provision of training data 1201. The training data may be constituted by a plurality of images of real objects as well as images of spoofs of real objects. For example, the images may comprise a number of images of real hands or fingers or the like and images of images (i.e. spoofs) of those objects. The images may be multiplied by using, from the same image, rotated, highlighted, darkened, enlarged or otherwise modified copies that are introduced as training data. In some embodiments, modifications involving image flips, image rotation and translation, shears, crops, multiplication to increase brightness and Gaussian blurs may be used to obtain a larger number of training images. Arbitrary combinations of the mentioned techniques may also be used. The values σ.sub.1 and σ.sub.2 provided according to item 1208 are the values indicating the “correct” output of the first node and second node of the last layer in the neural network that provide the probability of the image showing a spoof of an object or a real object. These values are provided for each image in the training data.--, in [0436]; and, -- specifically extracting a region of interest from the original image obtained and processed during the steps described above and before data comprising the biometric characteristic is sent to the third party computing device. In fact, with the embodiment now described, one embodiment of obtaining image data that only comprises the region of interest, like the fingertip carrying the fingerprint of interest, is described. [0448] This embodiment focuses on the extraction of a region of interest (ROI) from an image taken from an object of the user, where the image comprises a biometric characteristic that could be used to identify the user, but could also be used for any other purpose on the third party computing device. Such an object can be anything like a fingertip of one or more fingers of a hand of the user, the palm creases of a hand of the user or the face of the user or even the eye or the foot of the user. Each of these entities is known to carry biometric characteristics that can be used for identifying a user. [0449]… as a “cropping step” which separates a part of an image from another part of the image. This method is performed on the image that was already obtained during step 1 of the method according to FIG. 1. The step may occur after the identification step 2 and after the liveliness-detection and comparison steps 3 and 4. However, it can also occur in parallel to steps 3 and 4 or immediately after step 2. In some embodiments, it does not occur before step 2 since, in step 2 of FIG. 1, it is just determined whether there is an object carrying a biometric characteristic in the image at all. Processing the image in line with the embodiment to be described now would thus not be reasonable and potentially a waste of resources if it were performed before step 2 in FIG. 1.--, in [0447]-[0449], [0453]-[0457]; and, -- [0327] As explained above, it is assumed that the neural network is already perfectly learned for identifying a specific object, such as a fingertip. This involves that the neural network is able to identify a specific pattern of pixels that are most likely representing a fingertip. This might refer to specific patterns of color values or other characteristics like the brightness of those spots. It is, however, clear that the image 300 may arbitrarily show a fingertip which might not correspond in size and arrangement to a fingertip that was used for learning the neural network. [0328] With the help of the bounding boxes and the grid, however, it is possible for the neural network to identify the specific bounding box that will most likely comprise the fingertip. In order to identify this specific bounding box, the neural network (or an associated component that processes the image 300) compares the values of the pixels within each bounding box of each grid cell to a pattern of pixels that corresponds to a fingertip as was previously learned by the neural network. In this first stage, it is most unlikely that a perfect match will be found but there will be bounding boxes that are already more likely to contain at least a portion of a fingertip than other bounding boxes. [0329]…depicted in FIG. 6, for example, the bounding box 341 centered around the point M in grid cell 313 includes a portion of the fingertip of the hand 350. In contrast to this, none of the grid cells 310 and 311 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 341 and potentially the bounding box 340, the process can determine that the bounding box 341 includes even more of a pattern that corresponds to a fingertip than the bounding box 340. [0330] In view of this, the method can conclude that none of the bounding boxes 331 and 332 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0331] As both bounding boxes 340 and 341 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0332] While the smaller grid cell 340 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 341 may be regarded by the process to include a pattern that corresponds to a fingertip. [0333] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 341 and 340 to a pattern obtained from learning which indeed corresponds to a fingertip. [0334] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 341 is used as the starting point and its position and shapes modified or the smaller bounding box 340 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern….[0337] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0338] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 300 that contain the respective fingertip.--, in [0329]-[0338], and, --depicted in FIG. 18, for example, the bounding box 1841 centered around the point M in grid cell 1813 includes a portion of the fingertip of the hand 1850. In contrast to this, none of the grid cells 1810 and 1811 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 1841 and potentially the bounding box 1840, the process can determine that the bounding box 1841 includes even more of a pattern that corresponds to a fingertip than the bounding box 1840. [0565] In view of this, the method can conclude that none of the bounding boxes 1831 and 1832 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0566] As both bounding boxes 1840 and 1841 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0567] While the smaller grid cell 1840 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 1841 may be regarded by the process to include a pattern that corresponds to a fingertip. [0568] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 1841 and 1840 to a pattern obtained from learning which indeed corresponds to a fingertip. [0569] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 1841 is used as the starting point and its position and shapes modified or the smaller bounding box 1840 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern. [0570] This minimizing process can firstly comprise modifying the position of the bounding box (in the following, it will be assumed that the bounding box 1841 is used for the further calculations) by moving it a small amount into orthogonal directions first along the x-axis and then along the y-axis (or vice versa) as depicted in FIG. 18 around the center point M of the respective grid cell. The movement will be along the positive and the negative x-axis and y-axis and at each position, a comparison will be made to determine a difference function between the pattern obtained from the learning and the actual pattern identified in the image. This allows for calculating a two-dimensional function that represents the difference d(x,y) depending on the coordinates…..[0572] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0573] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 1800 that contain the respective fingertip or object carrying the biometric characteristic.--, in [0564]-[0573]; also see: -- [0071] It may be that the cropping operation comprises processing the image or the part of the image by a neural network, wherein processing the image or the part of the image by the neural network comprises processing the image by an encoder to obtain an encoded image and, after that, processing the encoded image by a decoder to obtain a decoded output image. [0072] Furthermore, the image or the part of the image provided to the neural network in the cropping operation for processing may comprise N×M pixels and the encoded image comprises n×m pixels, where n<N,m<M and the decoded output image comprises N×M pixels. [0073] Reducing the number of pixels when encoding the image results in a loss of information. When enlarging the image once again during the decoding, however, the most relevant information to distinguish the ROI from other portions of the image can be more easily discernable as not relevant information or very detailed information that is not necessary for identifying ROI is faded out with this procedure. [0074] In this regard, it can also be provided that distinguishing a portion of the image or the part of the image comprises distinguishing a portion of the decoded output image from another portion of the decoded output image. The distinguishing of the portions in the decoded image can be much easier compared to distinguishing the portion comprising the ROI from another portion of the original image. Thus, the processing power required for distinguishing a portion in the decoded output image from another portion in the decoded output image are reduced significantly compared to processing the original input image. [0075] More specifically, extracting the portion comprising the biometric characteristic can comprise identifying pixel in the decoded output image that are within the distinguished portion and, after that, identifying the pixels in the decoded output image that are in the distinguished portion with corresponding pixels in the original image or the part of the image and extracting, from the original image or the part of the image, the corresponding pixels, the extracted corresponding pixels constituting the portion of the image or the part of the image comprising the biometric characteristic.--, in [0071]-[0075]). Re Claim 7, GUPTA as modified by SARGENT and ZHANG further disclose wherein determining, with the center pixel as the starting point and based on the Markov chain, the cropping probability for each pixel in the training image comprises: setting, based on Markov chains along symmetric propagation directions starting from the center pixel, a cropping probability for each pixel along the symmetric propagation directions starting from the center pixel (see Gupta: e.g., -- [0436] The method of FIG. 12 begins with the provision of training data 1201. The training data may be constituted by a plurality of images of real objects as well as images of spoofs of real objects. For example, the images may comprise a number of images of real hands or fingers or the like and images of images (i.e. spoofs) of those objects. The images may be multiplied by using, from the same image, rotated, highlighted, darkened, enlarged or otherwise modified copies that are introduced as training data. In some embodiments, modifications involving image flips, image rotation and translation, shears, crops, multiplication to increase brightness and Gaussian blurs may be used to obtain a larger number of training images. Arbitrary combinations of the mentioned techniques may also be used. The values σ.sub.1 and σ.sub.2 provided according to item 1208 are the values indicating the “correct” output of the first node and second node of the last layer in the neural network that provide the probability of the image showing a spoof of an object or a real object. These values are provided for each image in the training data.--, in [0436]; and, -- specifically extracting a region of interest from the original image obtained and processed during the steps described above and before data comprising the biometric characteristic is sent to the third party computing device. In fact, with the embodiment now described, one embodiment of obtaining image data that only comprises the region of interest, like the fingertip carrying the fingerprint of interest, is described. [0448] This embodiment focuses on the extraction of a region of interest (ROI) from an image taken from an object of the user, where the image comprises a biometric characteristic that could be used to identify the user, but could also be used for any other purpose on the third party computing device. Such an object can be anything like a fingertip of one or more fingers of a hand of the user, the palm creases of a hand of the user or the face of the user or even the eye or the foot of the user. Each of these entities is known to carry biometric characteristics that can be used for identifying a user. [0449]… as a “cropping step” which separates a part of an image from another part of the image. This method is performed on the image that was already obtained during step 1 of the method according to FIG. 1. The step may occur after the identification step 2 and after the liveliness-detection and comparison steps 3 and 4. However, it can also occur in parallel to steps 3 and 4 or immediately after step 2. In some embodiments, it does not occur before step 2 since, in step 2 of FIG. 1, it is just determined whether there is an object carrying a biometric characteristic in the image at all. Processing the image in line with the embodiment to be described now would thus not be reasonable and potentially a waste of resources if it were performed before step 2 in FIG. 1.--, in [0447]-[0449], [0453]-[0457]; and, -- [0327] As explained above, it is assumed that the neural network is already perfectly learned for identifying a specific object, such as a fingertip. This involves that the neural network is able to identify a specific pattern of pixels that are most likely representing a fingertip. This might refer to specific patterns of color values or other characteristics like the brightness of those spots. It is, however, clear that the image 300 may arbitrarily show a fingertip which might not correspond in size and arrangement to a fingertip that was used for learning the neural network. [0328] With the help of the bounding boxes and the grid, however, it is possible for the neural network to identify the specific bounding box that will most likely comprise the fingertip. In order to identify this specific bounding box, the neural network (or an associated component that processes the image 300) compares the values of the pixels within each bounding box of each grid cell to a pattern of pixels that corresponds to a fingertip as was previously learned by the neural network. In this first stage, it is most unlikely that a perfect match will be found but there will be bounding boxes that are already more likely to contain at least a portion of a fingertip than other bounding boxes. [0329]…depicted in FIG. 6, for example, the bounding box 341 centered around the point M in grid cell 313 includes a portion of the fingertip of the hand 350. In contrast to this, none of the grid cells 310 and 311 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 341 and potentially the bounding box 340, the process can determine that the bounding box 341 includes even more of a pattern that corresponds to a fingertip than the bounding box 340. [0330] In view of this, the method can conclude that none of the bounding boxes 331 and 332 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0331] As both bounding boxes 340 and 341 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0332] While the smaller grid cell 340 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 341 may be regarded by the process to include a pattern that corresponds to a fingertip. [0333] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 341 and 340 to a pattern obtained from learning which indeed corresponds to a fingertip. [0334] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 341 is used as the starting point and its position and shapes modified or the smaller bounding box 340 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern….[0337] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0338] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 300 that contain the respective fingertip.--, in [0329]-[0338], and, --depicted in FIG. 18, for example, the bounding box 1841 centered around the point M in grid cell 1813 includes a portion of the fingertip of the hand 1850. In contrast to this, none of the grid cells 1810 and 1811 comprise bounding boxes that include a portion of a fingertip. When the method continues to evaluate the pixel values within the bounding box 1841 and potentially the bounding box 1840, the process can determine that the bounding box 1841 includes even more of a pattern that corresponds to a fingertip than the bounding box 1840. [0565] In view of this, the method can conclude that none of the bounding boxes 1831 and 1832 (and potentially other bounding boxes in other grid cells) includes a fingertip and can set their probability value in their corresponding B-vector to 0. [0566] As both bounding boxes 1840 and 1841 as centered around the point M comprise at least a portion of a fingertip, they may be considered to be likely to in fact comprise a fingertip and the probability value will be greater than 0 in a first step. [0567] While the smaller grid cell 1840 is almost completely filled with a pattern that could correspond to a fingertip, only the left border of the greater bounding box 1841 may be regarded by the process to include a pattern that corresponds to a fingertip. [0568] With this, the method may continue to calculate a loss function that determines the difference between the pattern identified within each of the bounding boxes 1841 and 1840 to a pattern obtained from learning which indeed corresponds to a fingertip. [0569] In the next step, the method will attempt to minimize this difference by modifying the size and the position of the respective bounding boxes. In this regard, it can be envisaged that the larger bounding box 1841 is used as the starting point and its position and shapes modified or the smaller bounding box 1840 is used as the starting point and its position and size are modified in order to minimize the differences to the learned pattern. [0570] This minimizing process can firstly comprise modifying the position of the bounding box (in the following, it will be assumed that the bounding box 1841 is used for the further calculations) by moving it a small amount into orthogonal directions first along the x-axis and then along the y-axis (or vice versa) as depicted in FIG. 18 around the center point M of the respective grid cell. The movement will be along the positive and the negative x-axis and y-axis and at each position, a comparison will be made to determine a difference function between the pattern obtained from the learning and the actual pattern identified in the image. This allows for calculating a two-dimensional function that represents the difference d(x,y) depending on the coordinates…..[0572] This can result in the bounding box being moved along the direction r to a new center point M′ where the function d(x,y) has a minimum. In a next step, the size of the respective bounding box at position M′ can be increased and reduced in order to determine whether with increasing or reducing the size in one or two directions (i.e. the height and/or the width) changes the value of a further difference function compared to the original pattern which can be denoted with e(h,b) depending on the height h and width b. This function is minimized such that for a specific bounding box having a position M′ and having a height h.sub.f and a width b.sub.f, the difference to the learned pattern is minimized. [0573] This bounding box will then be used as the final bounding box which has the greatest probability p of identifying those portions of the image 1800 that contain the respective fingertip or object carrying the biometric characteristic.--, in [0564]-[0573]; also see: -- [0071] It may be that the cropping operation comprises processing the image or the part of the image by a neural network, wherein processing the image or the part of the image by the neural network comprises processing the image by an encoder to obtain an encoded image and, after that, processing the encoded image by a decoder to obtain a decoded output image. [0072] Furthermore, the image or the part of the image provided to the neural network in the cropping operation for processing may comprise N×M pixels and the encoded image comprises n×m pixels, where n<N,m<M and the decoded output image comprises N×M pixels. [0073] Reducing the number of pixels when encoding the image results in a loss of information. When enlarging the image once again during the decoding, however, the most relevant information to distinguish the ROI from other portions of the image can be more easily discernable as not relevant information or very detailed information that is not necessary for identifying ROI is faded out with this procedure. [0074] In this regard, it can also be provided that distinguishing a portion of the image or the part of the image comprises distinguishing a portion of the decoded output image from another portion of the decoded output image. The distinguishing of the portions in the decoded image can be much easier compared to distinguishing the portion comprising the ROI from another portion of the original image. Thus, the processing power required for distinguishing a portion in the decoded output image from another portion in the decoded output image are reduced significantly compared to processing the original input image. [0075] More specifically, extracting the portion comprising the biometric characteristic can comprise identifying pixel in the decoded output image that are within the distinguished portion and, after that, identifying the pixels in the decoded output image that are in the distinguished portion with corresponding pixels in the original image or the part of the image and extracting, from the original image or the part of the image, the corresponding pixels, the extracted corresponding pixels constituting the portion of the image or the part of the image comprising the biometric characteristic.--, in [0071]-[0075]). Re Claim 14, GUPTA as modified by SARGENT however do not explicitly disclose the trained image processing network is used for performing gaze estimation on an image, and processing the cropped area to be processed by using the trained image processing network; ZHANG discloses the trained image processing network is used for performing gaze estimation on an image, and processing the cropped area to be processed by using the trained image processing network (see Zhang: e.g., -- An image processing system may receive image data from a camera of a user device and perform gaze prediction processing of the image data to predict one or more gaze patterns. The gaze prediction processing may include processing the image data using a neural network to detect faces and/or objects and generate an image feature map. The gaze prediction processing may include performing gaze direction prediction operations using the feature map and detected faces and/or objects to determine gaze direction probability data. The gaze prediction processing may include predicting a gaze pattern based on the gaze direction probability data and the image feature map. The gaze pattern may be short-term (e.g., atomic-level) or long-term (e.g., event-level).--, in abstract, and, --(19)… a user's gaze behavior may indicate attention towards a person or an object, or an attempt to attract a person's attention to the user or to an object. Gaze behavior information may help a system understand or otherwise determine what the user is looking at, and how the user feels, or what the user is referring to. A user interface system, such as a multi-modal user interface system that processes both speech and human gaze behavior may perform better when interacting with multiple users, who may also be interacting with each other. (20) Gaze behaviors may be represented by, or categorized into, gaze patterns. The user interface system may include one or more models trained to recognize various gaze patterns, and predict whether a person's gaze behavior as exhibited in image data corresponds to one of the learned gaze patterns. Gaze patterns may be categorized based on duration. “Atomic-level gaze patterns” may be brief in duration (e.g., lasting one or several image frames up to several seconds). Some atomic-level gaze patterns may include: Single: individual gaze behavior (e.g., directed toward the device or a person/object) Mutual: two people look into eyes of each other Avert: one person shifts away the gaze from another person's gaze Refer: one person introduces one target for another person using gaze Follow: one person accepts the gaze of another person and looking at the target referred by that person Share: two people are looking at the same target (21) “Event-level gaze patterns” may be longer in duration (e.g., longer than a second or several seconds), and may include one atomic-level gaze patterns followed by another. Some event-level gaze patterns may include: Non-communicative: Single gaze held for a period of time Mutual Gaze: Mutual gaze held for a period of time Gaze Aversion: starts with Mutual, follows by Avert, and finally Single Gaze Following: Follow and Share without Mutual Joint Attention: Starts with Mutual, then Refer, Follow and Share, and ends with Mutual again to confirm the Share event.--, in line 32, col. 2 through line 12, col. 3; and, -- The decoder portion of the network may take the collection of feature vectors and up-sample them using position information from the combined feature map 245 to produce a segmented output that represents a semantic segmentation of the input. In this case, the encoder-decoder network 250 can be trained to produce a gaze direction probability map 134. The gaze direction probability map 134 can include segments corresponding to bounding boxes around entities; for example, the first region 234a corresponding to the person and the second region 234b corresponding to the object. Values corresponding to those segments can represent a probability that the POI is gazing in the direction of that physical entity. In the example shown in FIGS. 1 and 2, the gaze direction probability map 134 may reflect a higher probability (e.g., closer to 1) that the first user 5a has a gaze directed toward the object 10, relative to a lower probability (e.g., closer to 0) that the first use 5a has a gaze directed toward the person (e.g., the second user 5b).--, in lines 16-40, col. 8; and, -- (47) The NN 131, the LSTM 380 and the LSTM 480 may be trained using annotated data sets. In some implementations, the models may be trained using the publicly available VACATION (Video gAze CommunicATION) dataset. The VACATION dataset includes 300 videos cropped from TV shows and movies from YouTube. The videos are stored in MPEG4 format with 640×360 spatial resolution. The length of individual videos ranges from 2.2 to 74.5 seconds, and average 13 seconds. The dataset includes a total of 96,993 frames and is 3,880 seconds long. The dataset includes annotations for, for example, human face and object bounding boxes, the attention of each person in the image (that is, to which person/object it directs its gaze) and both atomic-level and event-level gaze labels for each person in each frame. The NN 131 model may be trained using an Adam optimizer with a learning rate of 0.0001. Multi-class entropy may be chosen as the training loss. The network may learn over 10 epochs, taking 10 hours using a commercially available graphics card. The best model is chosen based on accuracy in processing a validation set. The LSTM 380 may be trained using information from 10 successive frames. During training of the second LSTM 480, more successive frames (e.g., 30, 60, or more) may be used to account for the relatively longer time period over which event-level gaze patterns may occur.--, in lines 1-25, col. 12); GUPTA (as modified by SARGENT) and ZHANG are combinable as they are in the same field of endeavor: training image processing neural network for classification and prediction including object recognition, segmentation, and bounding box identification. Therefore it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify GUPTA (as modified by SARGENT)’s method using ZHANG’s teachings by including wherein the trained image processing network is used for performing gaze estimation on an image, and processing the cropped area to be processed by using the trained image processing network to GUPTA (as modified by SARGENT)’s artificial intelligence model for object classification and detection in order to recognize various gaze patterns, and predict whether a person's gaze behavior as exhibited in image data corresponds to one of the learned gaze patterns (see ZHANG: e.g. in line 32, col. 2 through line 12, col. 3, and in lines 1-25, col. 12); GUPTA as modified by SARGENT and ZHANG further disclose to obtain the processing result for the image to be processed comprises: performing gaze estimation on the cropped area to be processed by using the trained image processing network, to obtain the processing result for the image to be processed (see Zhang: e.g., -- An image processing system may receive image data from a camera of a user device and perform gaze prediction processing of the image data to predict one or more gaze patterns. The gaze prediction processing may include processing the image data using a neural network to detect faces and/or objects and generate an image feature map. The gaze prediction processing may include performing gaze direction prediction operations using the feature map and detected faces and/or objects to determine gaze direction probability data. The gaze prediction processing may include predicting a gaze pattern based on the gaze direction probability data and the image feature map. The gaze pattern may be short-term (e.g., atomic-level) or long-term (e.g., event-level).--, in abstract, and, --(19)… a user's gaze behavior may indicate attention towards a person or an object, or an attempt to attract a person's attention to the user or to an object. Gaze behavior information may help a system understand or otherwise determine what the user is looking at, and how the user feels, or what the user is referring to. A user interface system, such as a multi-modal user interface system that processes both speech and human gaze behavior may perform better when interacting with multiple users, who may also be interacting with each other. (20) Gaze behaviors may be represented by, or categorized into, gaze patterns. The user interface system may include one or more models trained to recognize various gaze patterns, and predict whether a person's gaze behavior as exhibited in image data corresponds to one of the learned gaze patterns. Gaze patterns may be categorized based on duration. “Atomic-level gaze patterns” may be brief in duration (e.g., lasting one or several image frames up to several seconds). Some atomic-level gaze patterns may include: Single: individual gaze behavior (e.g., directed toward the device or a person/object) Mutual: two people look into eyes of each other Avert: one person shifts away the gaze from another person's gaze Refer: one person introduces one target for another person using gaze Follow: one person accepts the gaze of another person and looking at the target referred by that person Share: two people are looking at the same target (21) “Event-level gaze patterns” may be longer in duration (e.g., longer than a second or several seconds), and may include one atomic-level gaze patterns followed by another. Some event-level gaze patterns may include: Non-communicative: Single gaze held for a period of time Mutual Gaze: Mutual gaze held for a period of time Gaze Aversion: starts with Mutual, follows by Avert, and finally Single Gaze Following: Follow and Share without Mutual Joint Attention: Starts with Mutual, then Refer, Follow and Share, and ends with Mutual again to confirm the Share event.--, in line 32, col. 2 through line 12, col. 3; and, -- The decoder portion of the network may take the collection of feature vectors and up-sample them using position information from the combined feature map 245 to produce a segmented output that represents a semantic segmentation of the input. In this case, the encoder-decoder network 250 can be trained to produce a gaze direction probability map 134. The gaze direction probability map 134 can include segments corresponding to bounding boxes around entities; for example, the first region 234a corresponding to the person and the second region 234b corresponding to the object. Values corresponding to those segments can represent a probability that the POI is gazing in the direction of that physical entity. In the example shown in FIGS. 1 and 2, the gaze direction probability map 134 may reflect a higher probability (e.g., closer to 1) that the first user 5a has a gaze directed toward the object 10, relative to a lower probability (e.g., closer to 0) that the first use 5a has a gaze directed toward the person (e.g., the second user 5b).--, in lines 16-40, col. 8; and, -- (47) The NN 131, the LSTM 380 and the LSTM 480 may be trained using annotated data sets. In some implementations, the models may be trained using the publicly available VACATION (Video gAze CommunicATION) dataset. The VACATION dataset includes 300 videos cropped from TV shows and movies from YouTube. The videos are stored in MPEG4 format with 640×360 spatial resolution. The length of individual videos ranges from 2.2 to 74.5 seconds, and average 13 seconds. The dataset includes a total of 96,993 frames and is 3,880 seconds long. The dataset includes annotations for, for example, human face and object bounding boxes, the attention of each person in the image (that is, to which person/object it directs its gaze) and both atomic-level and event-level gaze labels for each person in each frame. The NN 131 model may be trained using an Adam optimizer with a learning rate of 0.0001. Multi-class entropy may be chosen as the training loss. The network may learn over 10 epochs, taking 10 hours using a commercially available graphics card. The best model is chosen based on accuracy in processing a validation set. The LSTM 380 may be trained using information from 10 successive frames. During training of the second LSTM 480, more successive frames (e.g., 30, 60, or more) may be used to account for the relatively longer time period over which event-level gaze patterns may occur.--, in lines 1-25, col. 12). . Re Claim 16, claim 16 the corresponding device claims to claim 2 respectively. Thus, claim 16 is rejected for the similar reasons as for claim 2. Furthermore, GUPTA as modified by SARGENT and ZHANG further disclose a device for training an image processing network, comprising a memory and a processor, wherein the memory is stored with a computer program executable on the processor, and the processor is configured to execute the computer program to the method (see GUPTA: e.g., --[0017] FIG. 7 shows the process of training the neural network in the identification operation--, and, --[0022] FIG. 12 shows the process of training the neural network used in the liveliness-detection operation--, and, --[0029] FIG. 19 presents a flowchart of the training of the neural network used in the cropping operation--, in [0017]-[0029]; and, --[0065] The neural network, according to some embodiments, is trained and/or otherwise adapted to be able to distinguish, through processing of the image, between those portions of the image that contain the region of interest and other portions of the image that do not contain the region of interest. This can be achieved in variety of ways and is thus not to be understood in a limiting way. That the neural network “distinguishes” that region comprising the ROI from another region is to be understood as the ability of the neural network to provide an output that distinguishes between the ROI and other regions of the image or makes it possible to distinguish between those regions. For example, the output could be an indication of pixels comprising the ROI but not other regions of the ROI. In any case, the outcome of the processing of the image by the neural network at least comprises that a first portion of the image comprising the region of interest is identified as different from another portion of the image. In this regard, it is noted that the specific size, shape of position of the region of interest is found out by the neural network during processing of the image and it is not preset.--, in [0065]). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to WEI WEN YANG whose telephone number is (571)270-5670. The examiner can normally be reached on 8:00 - 5:00 pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached on 571-272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WEI WEN YANG/Primary Examiner, Art Unit 2662
Read full office action

Prosecution Timeline

Dec 26, 2024
Application Filed
Sep 21, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743775
BIOMARKERS OF COLLAGEN FIBER ARCHITECTURE IN EPITHELIAL OVERIAN CANCER (EOC) PATIENTS
2y 11m to grant Granted Sep 22, 2026
Patent 12737879
Machine Learning for Detection of Diseases from External Anterior Eye Images
3y 9m to grant Granted Sep 15, 2026
Patent 12737880
SYSTEMS AND METHODS OF ANALYZING MICROBIOMES USING ARTIFICIAL INTELLIGENCE
3y 4m to grant Granted Sep 15, 2026
Patent 12738057
CUT-PASTE TRAINING AUGMENTATION FOR MACHINE LEARNING MODELS
2y 7m to grant Granted Sep 15, 2026
Patent 12737851
ENHANCED QUALITY BOREHOLE IMAGE GENERATION AND METHOD
2y 3m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
82%
Grant Probability
93%
With Interview (+11.5%)
2y 5m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 684 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month