Prosecution Insights
Last updated: October 02, 2026
Application No. 18/272,916

EXTRACTING FEATURES FROM SENSOR DATA

Non-Final OA §103
Filed
Jul 18, 2023
Priority
Jan 20, 2021 — GB 2100739.8 +1 more
Examiner
CAI, PHUONG HAU
Art Unit
2673
Tech Center
2600 — Communications
Assignee
Five AI Limited
OA Round
3 (Non-Final)
77%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
90 granted / 117 resolved
+14.9% vs TC avg
Strong +26% interview lift
Without
With
+26.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
26 currently pending
Career history
150
Total Applications
across all art units

Statute-Specific Performance

§101
22.3%
-17.7% vs TC avg
§103
42.6%
+2.6% vs TC avg
§102
23.3%
-16.7% vs TC avg
§112
11.4%
-28.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 117 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submissions, filed on July 24th, 2026, have been entered. Status of Claims Claims 1-17 and 20-22 are pending, claims 1, 9, 11, 13, 20 and 22 have been amended, claims 18-19 have been canceled. Claims 1-17 and 20-22 remains rejected. Response to Argument(s) 101 Rejection: The examiner finds the amendment to follow suggestion which was given to the Applicants’ representative Laura McJilton (registration number of 74,378) to clarify the limitation of performing matching between the numerical output value computed from the extracted features with at numerical transformation value parameterizing the transformation through the self-supervised regression loss function, so that learn values of the neural network of the encoder during training, hence, indicates an improvement in training of an encoder neural network through specific learning that include matching out parameter values driven by a self-supervised loss function, parameterizing a transformation for learning of the parameters of the neural network encoder. Rejection: In view of the Amendments to independent claims 1, 20 and 22 the previously applied prior art rejections are withdrawn. Applicants’ arguments are rendered moot in view of the new grounds of rejection set forth below. Particularly, in pages 11-14 of the Applicants’ remarks, the Applicants argue that the proposed prior arts of Xie and Park (previously used in the rejections of the independent claims), alone or in combination, does not teach or suggest the features of each of the independent claims 1, 20 and 22: “the first data representation is a transformed data representation of the second data representation; A self-supervised regression loss function configured to drive the at least one numerical output value to match at least one numerical transformation value parameterizing the transformation such that the values of one or more parameters of a neural network of the encoder are learned” As stated, the examiner find the feature of “the first data representation is a transformed data representation of the second data representation” to be a new amended feature that followed the suggestion during the telephonic interview, hence clarifies and narrows down the claims’ scope therefore, helps overcome the previously stated prior art 103 rejections. Regarding the limitation of “A self-supervised regression loss function configured to drive the at least one numerical output value to match at least one numerical transformation value parameterizing the transformation such that the values of one or more parameters of a neural network of the encoder are learned,” in support of the indicated argument that the prior arts do not teach or suggest this limitation, the Applicants centrally assert that Claim Objections Claims 1, 20 and 22 are objected to because of the following informalities: Claim 1, line 12, the reference “the at least one numerical transformation value parametrizing the transformation” is lack of antecedent basis support, since there is no prior instantiation of “an at least one numerical transformation value parametrizing a transformation” for the reference to have such antecedent reference, is suggested to be amended to “an at least one numerical transformation value parametrizing a transformation”. Appropriate correction is required to avoid 112(b) antecedent basis issue. Claim 20, line 17, the reference “the at least one numerical transformation value parametrizing the transformation” is lack of antecedent basis support, since there is no prior instantiation of “an at least one numerical transformation value parametrizing a transformation” for the reference to have such antecedent reference, is suggested to be amended to “an at least one numerical transformation value parametrizing a transformation”. Appropriate correction is required to avoid 112(b) antecedent basis issue. Claim 22, line 13, the reference “the at least one numerical transformation value parametrizing the transformation” is lack of antecedent basis support, since there is no prior instantiation of “an at least one numerical transformation value parametrizing a transformation” for the reference to have such antecedent reference, is suggested to be amended to “an at least one numerical transformation value parametrizing a transformation”. Appropriate correction is required to avoid 112(b) antecedent basis issue. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-8, 13-17 and 22 are rejected under 35 U.S.C. 103 as being unpatentable over Saining Xie et. al. (“PointContrast: Unsupervised Pre-Training for 3D Point Cloud Understanding, Dec. 2020, Part of the book series: Lecture Notes in Computer Science, Vol. 12348” hereinafter as “Xie”) in view of Minwoo Park et. al. (“US 2021/0166052 A1” hereinafter as “Park”) further in view of Dilip Krishnan et. al. (“US 2021/0326660 A1” hereinafter as “Krishnan”) and Jingyi Yu (“US 2020/0074658 A1” hereinafter as “Yu”). Regarding claim 1, Xie teaches a computer implemented method of training an encoder to extract features from sensor data, the method comprising (Abstract discloses “pre-training a network on a rich source set” and section 3.5 discloses “UNet architecture that has an encoder network…this architecture as a unified design for both the pre-training task…”): generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data (section 3.5 discloses “UNet architecture that has an encoder network…this architecture as a unified design for both the pre-training task…” wherein having training/pre-training of the network of encoder, Fig. 2 illustrates sensor to obtain x1 and x2 being a set of sensor data; moreover, Section 3.3 discloses “FCGF focuses on local descriptor learning” indicating a learning/training and “generate two views x1 and x2 that are aligned in the same world coordinates” wherein, aligned x1 and x2 are data representation of a set of sensor data used for learning/training); and training the encoder based on a supervised loss function applied to the training examples (section 3.1 discloses the encoder is trained based on supervised learning “pre-train an encoder network…we use full supervision” training using contrastive loss function; Section 3, 1st Par., discloses “we introduce our supervised pre-training solution…loss function (Sect. 3.4)” and Section 3.4, 1st Par., discloses “the first loss function, hardest-contrastive loss…”); wherein the encoder extracts respective features from the at least two data representations of each training example (Fig. 2 illustrates the encoder network extracts f1 and f2 features from the aligned x1 and x2; Section 3.4, 1st Par. discloses “matched pairs of points x1_i and x2_j from two views x1 and x2, and f1_i and f2_j are associated point features for the matched pair”), and at least one numerical output value is computed from the extracted features (Algorithm 1 of section 3.2 discloses “compute point features f1, f2…by f1 = NN(T1(x1)) and f2 = NN(T2(x2))” being the computation of such point features, indicating a numerical output value being associated with each of the point features being computed), wherein the supervised loss function is configured to drive the at least one numerical output value to match the at least one numerical value (Section 3.4, 1st 2 Pars., discloses “the first loss function…set of matched pairs of points x1_i and x2_j…the loss encourages a query q to be similar to its positive key k+ and dissimilar to, typically many, negative keys k-” having wherein the loss function to encourage matching between the feature points [at least one numerical output value and the at least one numerical as claimed]). However, Xie does not explicitly teach wherein the loss function being a regression loss function. Park teaches wherein the loss function being a regression loss function (FIG. 1 illustrates as the training can include regression loss and further disclosed in [0006]). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie of having a computer implemented method of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a supervised loss function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the supervised loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Park of having wherein the loss function being a regression loss function. Wherein having Xie’s method of having wherein the loss function being a regression loss function. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve training of an encoder network on producing accurate predictions based on using regression loss function. Since both Xie and Park both perform training of an encoder network on point cloud processing. Wherein Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Park’s system improves training of an encoder network on producing accurate predictions based on using regression loss function (see Park’s Par. [0006]). However, Xie in view of Park does not explicitly teach wherein the supervised loss function being a self-supervised loss function. Krishnan teaches wherein the supervised loss function being a self-supervised loss function (Par. [0036] discloses “contrastive learning loss for self-supervised representation learning. Next it is shown how this loss can be modified to be suitable for fully supervised learning, while simultaneously preserving properties important to the self-supervised approach”). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie in view of Park of having a computer implemented method of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a supervised loss regression function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the supervised regression loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Krishnan of having wherein the supervised loss function being a self-supervised loss function. Wherein having Xie’s method of having wherein the supervised loss function being a self-supervised loss function. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve training method using contrastive learning by enabling learning to occur simultaneously across multiple positive examples. Since both Xie and Krishnan both perform supervised contrastive learning for learning on classification tasks. Wherein Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Krishnan’s system improves training method using contrastive learning by enabling learning to occur simultaneously across multiple positive examples (see Krishnan’s Par. [0006]). However, Xie in view of Park and Krishnan does not explicitly teach wherein a first data representation is a transformed data representation of a second data representation, the at least one numerical value being the at least one numerical transformation value parameterizing the transformation. Yu teaches wherein a first data representation is a transformed data representation of a second data representation (Par. [0033] discloses “a first light field F, is captured by the LF camera at a first viewpoint k, and a second light field, F’, is captured by the LF camera at a second viewpoint k+1, and F’ is aligned to the world coordinates…within F, if the transformation between F and F’ is known” and Par. [0034] discloses “feature matching across the sub-aperture images to find matched features…between two different LF images” wherein having images of different viewpoints to match features through a transformation), the at least one numerical value being the at least one numerical transformation value parameterizing the transformation (Par. [0032] discloses “LF pose estimation…can be described in ray space…alpha and teta to parametrize the ray direction”; Par. [0033] discloses “a first light field F, is captured by the LF camera at a first viewpoint k, and a second light field, F’, is captured by the LF camera at a second viewpoint k+1, and F’ is aligned to the world coordinates…given a ray…within F, if the transformation between F and F’ is known” wherein having the matching between feature points being a transformation of another with ray direction parameterized by some vectors). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie in view of Park and Krishnan of having a computer implemented method of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a self-supervised loss regression function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the self-supervised regression loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Yu of having wherein a first data representation is a transformed data representation of a second data representation, the at least one numerical value being the at least one numerical transformation value parameterizing the transformation. Wherein Xie’s method of having wherein a first data representation is a transformed data representation of a second data representation, the at least one numerical value being the at least one numerical transformation value parameterizing the transformation. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve providing accurate results on 3D model reconstruction through utilizing ray transformation. Since both Xie and Yu both perform matching feature point across different viewpoints to world coordinates. Wherein, Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Yu’s system improves accurate results on 3D model reconstruction through utilizing ray transformation (see Yu’s Pars. [0003-0004]). Regarding claim 2, Xie in view of Park further in view of Krishnan and Yu teach the method of claim 1, Xie teaches wherein the respective features are respective local features contained in respective feature maps extracted from the at least two data representations (the respective features extracted from the at least two data representations, moreover, the features being local features according to Xie’s section 3, 1st Par. and shown in FIG. 1). Regarding claim 3, Xie in view of Park further in view of Krishnan and Yu teach the method of claim 2, Xie teaches wherein the transformation comprises a global transformation and the at least one numerical transformation value comprises a global transformation value (the transformation on features of global representations such as disclosed in section 3.1 of “directly training on object instances to obtain a global representation…”; Algorithm 1 shows “sample two transformations T1 and T2” wherein, transformation of feature points including global representation), wherein multiple numerical output values are computed from the extracted local features (the multiple numerical output values of algorithm 1 are being computed from features as shown in algorithm 1 being local features such as disclosed in section 3, 1st Par. of “briefly reviewing an inspirational local feature learning work Fully Convolutional Geometric Features”), and the loss function is configured to drive each of the multiple numerical output values to match the global transformation value (Section 3.4, 1st 2 Pars., discloses “the first loss function…set of matched pairs of points x1_i and x2_j…the loss encourages a query q to be similar to its positive key k+ and dissimilar to, typically many, negative keys k-” having wherein the loss function to encourage matching between the feature points [at least one numerical output value and the at least one numerical as claimed]; Algorithm 1 shows “sample two transformations T1 and T2” wherein, transformation of feature points including global representation). Regarding claim 4, Xie in view of Park further in view of Krishnan and Yu teach the method of claim 2, Xie teaches wherein the transformation comprises one or more local transformations and the at least one numerical transformation value comprises one or more local transformation values (the transformation which includes transformation on features of local features such as disclosed in section 3, 1st par., hence, it can be understood as the transformation being local transformation and include local transformation value such as the value computed in algorithm 1), wherein multiple local numerical output values are computed from the extracted local features (the multiple numerical output values of algorithms 1 are being computed from features as shown in algorithm 1 being local features such as disclosed in section 3, 1st Par.), and the loss function is configured to drive each of the local numerical output values to match a corresponding one of the local transformation values (Section 3.4, 1st 2 Pars., discloses “the first loss function…set of matched pairs of points x1_i and x2_j…the loss encourages a query q to be similar to its positive key k+ and dissimilar to, typically many, negative keys k-” having wherein the loss function to encourage matching between the feature points [at least one numerical output value and the at least one numerical as claimed]; the multiple numerical output values of algorithm 1 are being computed from features as shown in algorithm 1 being local features such as disclosed in section 3, 1st Par. of “briefly reviewing an inspirational local feature learning work Fully Convolutional Geometric Features”). Regarding claim 5, Xie in view of Park further in view of Krishnan and Yu teach the method of claim 4, Xie teaches wherein each local numerical output value is determined based on a mapping between a spatial location of a first of the data representations and a second spatial location of a second of the data representations (the local features of section 3, 1st par., as part of the computation of algorithm 1, being computed correspondence mapping between the points of the two data representations based on spatial location x1 and x2 location such as shown in algorithm 1). Regarding claim 6, Xie in view of Park further in view of Krishnan and Yu teach the method of claim 5, Xie teaches wherein the transformation is fully or partially geometric (“or” indicates a selection, therefore, only one of the options is the instant scope of the claim, the examiner selects “partially” for mapping which is a taught in Section 3, last 2 Pars., wherein the scanning for the processing of algorithm 1 including the transformation being partially scanned indicates partially geometric, by BRI) and the mapping is determined from the transformation (as shown in algorithm 1, of section 3.2, last Par., the computing of the correspondence mapping is determined according to the transformations T1 and T2). Regarding claim 7, Xie in view of Park further in view of Krishnan and Yu teach the method of claim 5, Xie teaches wherein each local numerical output value is computed by comparing a first vector or scalar and a second scalar or vector (Algorithm 1, the local numerical output value is computed by feature vectors of the correspondence mapping between the points which indicates comparing of the points for the mapping; here, the option “vector” is selected instead of the “scalar” for the mapping), wherein the first vector or scalar is defined by the first spatial location and the feature map of the first data representation (the feature vector of Algorithm 1, being defined by all the information including the spatial location as shown in Fig. 2, the feature map such as shown in algorithm 1), and the second vector or scalar is defined by the second spatial location and the feature map of the second data representation (since the feature vector of Algorithm 1, includes the vectors of both data representation, any of which is analogous to the 1st vector and any of the others would be the 2nd vector). Regarding claim 8, Xie in view of Park further in view of Yu teach the method of claim 7. However, Xie in view of Park further in view of Yu does not explicitly teach wherein the first and second vectors or scalars are computed from the feature maps using a trainable projection component that is trained simultaneously with the encoder. Krishnan teaches wherein the first and second vectors or scalars are computed from the feature maps using a trainable projection component (Par. [0007] discloses “a projection head neural network configured to process the embedding representation of the input image” Par. [0048] discloses “a projection network, which maps the normalized representation vector r into a projected representation suitable for computation of the contrastive loss”) that is trained simultaneously with the encoder (Par. [0060] discloses “consider the effects on the encoder…while simultaneously minimizing its denominator”; Par. [0036] discloses “contrastive learning loss for self-supervised representation learning…while simultaneously preserving properties important to the self-supervised approach”). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie in view of Park further in view of Yu of having a computer implemented method of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a supervised loss regression function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the supervised regression loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Krishnan of having wherein the first and second vectors or scalars are computed from the feature maps using a trainable projection component that is trained simultaneously with the encoder. Wherein having Xie’s method of having wherein the first and second vectors or scalars are computed from the feature maps using a trainable projection component that is trained simultaneously with the encoder. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve training method using contrastive learning by enabling learning to occur simultaneously across multiple positive examples. Since both Xie and Krishnan both perform supervised contrastive learning for learning on classification tasks. Wherein Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Krishnan’s system improves training method using contrastive learning by enabling learning to occur simultaneously across multiple positive examples (see Krishnan’s Par. [0006]). Regarding claim 13, Xie in view of Park further in view of Krishnan and Yu teach the method of claim 1, Xie teaches wherein the transformation comprises rescaling, translation, cropping and/or tearing as parameterized by the at least one numerical transformation value (“or” indicates a selection, the examiner selects “translation” for mapping, which is taught in Section 3.3, last Par., of “consider rigid transformation including rotation, translation…”). Regarding claim 14, Xie in view of Park further in view of Krishnan and Yu teach the method of claim 1, Xie teaches wherein the transformation comprises at least one non-geometric transformation (the transformation of Section 3.3, last Par., discloses the transformation comprises non-geometric transformation), such as addition of noise (the training can include natural noise information such as disclosed in section 4.6, 4th Par.), that is parameterized by the at least one numerical transformation value (the two data representations as shown in FIG. 2 of T1 and T2 relates to each other through a transformation such as, further shown in algorithm 1 of section 3.2, which is parameterized by transformation numerical value T1 and T2). Regarding claim 15, Xie in view of Park further in view of Krishnan and Yu teach the method of claim 4, Xie teaches wherein a 2D object detector (Section 3.5, 1st Par. discloses that the 2D ResNet block [2D object detector as claimed]) is applied to an image other than the at least two data representations (Section 3.5, 1st Par., discloses “2D ResBet basic block design…” and Section 1, 2nd Par., discloses “2D vision such as ResNets…point cloud network architecture” such as shown in FIG. 2 processing the two data representations) in order to determine the local transformations for one or more objects detected in the image (Section 3.1, last 2 Pars., discloses “the local geometric features…global representations but instead can capture dense/local features”), the image containing or associated with the sensor data (as shown in FIG. 2 of sensors capturing the sensor data). Regarding claim 16, Xie in view of Park further in view of Krishnan and Yu teach the method of claim 15, Xie teaches wherein the data representations encode views of the sensor data in a plane other than an image plane of the image (the data representations to encode views of the sensor data in plane such as shown in FIG. 2 to be other than an image plane since it’s in 3D point cloud). Regarding claim 17, Xie in view of Park further in view of Krishnan and Yu teach the method of a claim 1, Xie teaches wherein the data representations are image or voxel representations (“or” indicates a selection, the examiner selects “image” for mapping which is taught in Fig. 2 of Xie) and wherein the data representations are optionally image or voxel representations of 2D or 3D point clouds (“or” indicates a selection, the examiner selects 3D for mapping which is shown in FIG. 2 of the 3D point cloud of Xie). Regarding claim 22, Xie teaches a non-transitory medium embodying training computer-readable instructions program configured, when executed on one or more computer hardware processors (abstract discloses using of an encoder for computer processing hence, indicates the use of a computer to have computer components to perform computer component functions such as a processor to execute instructions of the invention stored in a memory, including a non-transitory medium such as a ROM or RAM), to train an encoder to extract features from sensor data by: (Abstract discloses “pre-training a network on a rich source set” and section 3.5 discloses “UNet architecture that has an encoder network…this architecture as a unified design for both the pre-training task…”): generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data (section 3.5 discloses “UNet architecture that has an encoder network…this architecture as a unified design for both the pre-training task…” wherein having training/pre-training of the network of encoder, Fig. 2 illustrates sensor to obtain x1 and x2 being a set of sensor data; moreover, Section 3.3 discloses “FCGF focuses on local descriptor learning” indicating a learning/training and “generate two views x1 and x2 that are aligned in the same world coordinates” wherein, aligned x1 and x2 are data representation of a set of sensor data used for learning/training); and training the encoder based on a supervised loss function applied to the training examples (section 3.1 discloses the encoder is trained based on supervised learning “pre-train an encoder network…we use full supervision” training using contrastive loss function; Section 3, 1st Par., discloses “we introduce our supervised pre-training solution…loss function (Sect. 3.4)” and Section 3.4, 1st Par., discloses “the first loss function, hardest-contrastive loss…”); wherein the encoder extracts respective features from the at least two data representations of each training example (Fig. 2 illustrates the encoder network extracts f1 and f2 features from the aligned x1 and x2; Section 3.4, 1st Par. discloses “matched pairs of points x1_i and x2_j from two views x1 and x2, and f1_i and f2_j are associated point features for the matched pair”), and at least one numerical output value is computed from the extracted features (Algorithm 1 of section 3.2 discloses “compute point features f1, f2…by f1 = NN(T1(x1)) and f2 = NN(T2(x2))” being the computation of such point features, indicating a numerical output value being associated with each of the point features being computed), wherein the supervised loss function is configured to drive the at least one numerical output value to match the at least one numerical transformation value parameterizing the transformation (Section 3.4, 1st 2 Pars., discloses “the first loss function…set of matched pairs of points x1_i and x2_j…the loss encourages a query q to be similar to its positive key k+ and dissimilar to, typically many, negative keys k-” having wherein the loss function to encourage matching between the feature points [at least one numerical output value and the at least one numerical as claimed]). However, Xie does not explicitly teach wherein the loss function being a regression loss function. Park teaches wherein the loss function being a regression loss function (as shown in FIG. 1 as the training can include regression loss and further disclosed in [0006]). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie of having a system of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a supervised loss function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the supervised loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Park of having wherein the loss function being a regression loss function. Wherein having Xie’s system of having wherein the loss function being a regression loss function. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve training of an encoder network on producing accurate predictions based on using regression loss function. Since both Xie and Park both perform training of an encoder network on point cloud processing. Wherein Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Park’s system improves training of an encoder network on producing accurate predictions based on using regression loss function (see Park’s Par. [0006]). However, Xie in view of Park does not explicitly teach wherein the supervised loss function being a self-supervised loss function. Krishnan teaches wherein the supervised loss function being a self-supervised loss function (Par. [0036] discloses “contrastive learning loss for self-supervised representation learning. Next it is shown how this loss can be modified to be suitable for fully supervised learning, while simultaneously preserving properties important to the self-supervised approach”). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie in view of Park of having a system of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a supervised loss regression function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the supervised regression loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Krishnan of having wherein the supervised loss function being a self-supervised loss function. Wherein having Xie’s system of having wherein the supervised loss function being a self-supervised loss function. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve training method using contrastive learning by enabling learning to occur simultaneously across multiple positive examples. Since both Xie and Krishnan both perform supervised contrastive learning for learning on classification tasks. Wherein Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Krishnan’s system improves training method using contrastive learning by enabling learning to occur simultaneously across multiple positive examples (see Krishnan’s Par. [0006]). However, Xie in view of Park and Krishnan does not explicitly teach wherein a first data representation is a transformed data representation of a second data representation, the at least one numerical value being the at least one numerical transformation value parameterizing the transformation. Yu teaches wherein a first data representation is a transformed data representation of a second data representation (Par. [0033] discloses “a first light field F, is captured by the LF camera at a first viewpoint k, and a second light field, F’, is captured by the LF camera at a second viewpoint k+1, and F’ is aligned to the world coordinates…within F, if the transformation between F and F’ is known” and Par. [0034] discloses “feature matching across the sub-aperture images to find matched features…between two different LF images” wherein having images of different viewpoints to match features through a transformation), the at least one numerical value being the at least one numerical transformation value parameterizing the transformation (Par. [0032] discloses “LF pose estimation…can be described in ray space…alpha and teta to parametrize the ray direction”; Par. [0033] discloses “a first light field F, is captured by the LF camera at a first viewpoint k, and a second light field, F’, is captured by the LF camera at a second viewpoint k+1, and F’ is aligned to the world coordinates…given a ray…within F, if the transformation between F and F’ is known” wherein having the matching between feature points being a transformation of another with ray direction parameterized by some vectors). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie in view of Park and Krishnan of having a system of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a self-supervised loss regression function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the self-supervised regression loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Yu of having wherein a first data representation is a transformed data representation of a second data representation, the at least one numerical value being the at least one numerical transformation value parameterizing the transformation. Wherein Xie’s system of having a first data representation is a transformed data representation of a second data representation, the at least one numerical value being the at least one numerical transformation value parameterizing the transformation. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve providing accurate results on 3D model reconstruction through utilizing ray transformation. Since both Xie and Yu both perform matching feature point across different viewpoints to world coordinates. Wherein, Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Yu’s system improves accurate results on 3D model reconstruction through utilizing ray transformation (see Yu’s Pars. [0003-0004]). Claims 9-10 are rejected under 35 U.S.C. 103 as being unpatentable over Saining Xie et. al. (“PointContrast: Unsupervised Pre-Training for 3D Point Cloud Understanding, Dec. 2020, Part of the book series: Lecture Notes in Computer Science, Vol. 12348” hereinafter as “Xie”) in view of Minwoo Park et. al. (“US 2021/0166052 A1” hereinafter as “Park”) further in view of Dilip Krishnan et. al. (“US 2021/0326660 A1” hereinafter as “Krishnan”) and Jingyi Yu (“US 2020/0074658 A1” hereinafter as “Yu”) and Yingzi Ma (“Self-Supervised Learning of 3D Point Clouds via Feature Transformation and Rotation Prediction, Nov. 2022, 2022 International Conference on Electrical, Computer, Communications and Mechatronics Engineering” hereinafter as “Ma”). Regarding claim 9, Xie in view of Park further in view of Krishnan and Yu teach the method of claim 7, Xie teaches wherein the transformation comprises global rotation and the at least one numerical transformation value (the transformation include rotation including global representation as disclosed in Section 3.3, 2nd Par., hence, analogous to global rotation as claimed) comprises a global rotation angle (including the rotation as discussed previously of Section 3.3, 2nd Par., a rotation value/data indicates a rotation angle, such as shown in FIG. 2); and the loss function is configured to drive each of the local numerical output values to match the global rotation angle (as shown in algorithm 1, the contrastive loss is to encourage the numerical output value of f1 and f2 to match through a transformation to match the transformed point cloud such as disclosed in Section 3.3, 2nd Par., include matching value with a global rotation angle, by BRI). However, Xie in view of Park further in view of Krishnan and Yu does not explicitly teach wherein at least one local numerical output value is computed as an angular separation between the first vector and the second vector. Ma teaches wherein at least one local numerical output value is computed as an angular separation between the first vector and the second vector (section III.B discloses the transformation being transformation rotation including an angle as separation between two vectors, hence, can be combined with Xie to teach the claimed limitation, by BRI). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie in view of Park further in view of Krishnan and Yu of having a computer implemented method of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a self-supervised loss regression function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the self-supervised regression loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Ma of having wherein at least one local numerical output value is computed as an angular separation between the first vector and the second vector. Wherein, Xie’s method of having at least one local numerical output value is computed as an angular separation between the first vector and the second vector. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve supervised settings on point cloud analysis task learning to be more demonstrative on better performance and effectiveness. Since both Xie and Ma both perform matching feature point across different viewpoints to world coordinates. Wherein Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Ma’s system improves supervised settings on point cloud analysis task learning to be more demonstrative on better performance and effectiveness (see Ma’s Abstract). Regarding claim 10, Xie in view of Park further in view of Krishnan and Yu teach the method of claim 7, Xie teaches wherein the transformation comprises local rotation and the at least one at least one numerical transformation value (the one numerical transformation and global transformation, moreover, the transformation include rotation as disclosed in Section 3.3, 2nd Par., wherein, is analogous to local rotation) comprises a local rotation angle (including the rotation as discussed previously of Section 3.3, 2nd Par., a rotation value/data indicates a rotation angle, such as shown in FIG. 2); and the loss function encourages each of the local numerical output values to match the local rotation angle (as shown in algorithm 1, the contrastive loss is to encourage the numerical output value of f1 and f2 to match through a transformation to match the transformed point cloud such as disclosed in Section 3.3, last Par.; by BRI, covers the scope of the claim, here the matching would include the information of the transformation rotation hence would include matching value with a local rotation angle; Section 3.4, 1st 2 Pars., discloses “the first loss function…set of matched pairs of points x1_i and x2_j…the loss encourages a query q to be similar to its positive key k+ and dissimilar to, typically many, negative keys k-” having wherein the loss function to encourage matching between the feature points [at least one numerical output value and the at least one numerical as claimed]). However, Xie in view of Park further in view of Krishnan and Yu does not explicitly teach wherein at least one local numerical output value is computed as an angular separation between the first vector and the second vector. Ma teaches wherein at least one local numerical output value is computed as an angular separation between the first vector and the second vector (section III.B discloses the transformation being transformation rotation including an angle as separation between two vectors). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie in view of Park further in view of Krishnan and Yu of having a computer implemented method of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a self-supervised loss regression function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the self-supervised regression loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Ma of having wherein at least one local numerical output value is computed as an angular separation between the first vector and the second vector. Wherein, Xie’s method of having at least one local numerical output value is computed as an angular separation between the first vector and the second vector. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve supervised settings on point cloud analysis task learning to be more demonstrative on better performance and effectiveness. Since both Xie and Ma both perform matching feature point across different viewpoints to world coordinates. Wherein Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Ma’s system improves supervised settings on point cloud analysis task learning to be more demonstrative on better performance and effectiveness (see Ma’s Abstract). Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Saining Xie et. al. (“PointContrast: Unsupervised Pre-Training for 3D Point Cloud Understanding, Dec. 2020, Part of the book series: Lecture Notes in Computer Science, Vol. 12348” hereinafter as “Xie”) in view of Minwoo Park et. al. (“US 2021/0166052 A1” hereinafter as “Park”) further in view of Dilip Krishnan et. al. (“US 2021/0326660 A1” hereinafter as “Krishnan”) and Jingyi Yu (“US 2020/0074658 A1” hereinafter as “Yu”) and Qiangeng Xu et. al. (“Grid-GCN for Fast and Scalable Point Cloud Learning, 2020, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5661-5670” hereinafter as “Xu”). Regarding claim 11, Xie in view of Park further in view of Krishnan and Yu teach the method of claim 7, wherein the mapping (the mapping as discussed above in claim 7). However, Xie in view of Park further in view of Krishnan and Yu does not explicitly teach the mapping is from a grid cell of the first data representation to a grid cell of the second representation, wherein the first and second spatial locations are grid cell locations. Xu teaches the mapping is from a grid cell of the first data representation to a grid cell of the second representation (FIG. 1 illustrates the mapping includes gridConv and the data representation in grid cells), wherein the first and second spatial locations are grid cell locations (Section 5.2, 1st Par., discloses “sample 8192 points during training and 3 spatial channels for each point”; Section 1, 1st Par., discloses “Volumetric models are a family of models that transfer point cloud to spatially quantized voxel grids”). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie in view of Park further in view of Krishnan and Yu of having a computer implemented method of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a self-supervised loss regression function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the self-supervised regression loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Xu of having wherein the mapping is from a grid cell of the first data representation to a grid cell of the second representation, wherein the first and second spatial locations are grid cell locations. Wherein, Xie’s method of having the mapping is from a grid cell of the first data representation to a grid cell of the second representation, wherein the first and second spatial locations are grid cell locations. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve point cloud learning for more accurate point cloud classification task by considering quantized information of grid points and grid space. Since both Xie and Xu both perform point cloud learning for classification tasks. Wherein Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Xu’s system improves point cloud learning for more accurate point cloud classification task by considering quantized information of grid points and grid space (see Xu’s Abstract). Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Saining Xie et. al. (“PointContrast: Unsupervised Pre-Training for 3D Point Cloud Understanding, Dec. 2020, Part of the book series: Lecture Notes in Computer Science, Vol. 12348” hereinafter as “Xie”) in view of Minwoo Park et. al. (“US 2021/0166052 A1” hereinafter as “Park”) further in view of Dilip Krishnan et. al. (“US 2021/0326660 A1” hereinafter as “Krishnan”) and Jingyi Yu (“US 2020/0074658 A1” hereinafter as “Yu”) and Qiangeng Xu et. al. (“Grid-GCN for Fast and Scalable Point Cloud Learning, 2020, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5661-5670” hereinafter as “Xu”) and Michael Goersele et. al. (“Ambient Point Clouds for View Interpolation, July 2010, SIGGRAPH’ 10: ACM SIGGRAPH 2010 papers, article number 95, pp. 1-6” hereinafter as “Goersele”). Regarding claim 12, Xie in view of Park further in view of Krishnan and Yu and Xu teach the method of claim 7. However, Xie in view of Park further in view of Krishnan and Yu does not explicitly teach the mapping is from a grid cell of the first data representation to a region of the second representation spanning multiple grid cells thereof, the second vector or scalar determined via interpolation of vectors of scalars of the multiple grid cells. Xu teaches the mapping is from a grid cell of the first data representation to a region of the second representation spanning multiple grid cells thereof (FIG. 1 illustrates the mapping includes gridConv and the data representation in grid cells; moreover, since a region of a data representation shown in FIG. 1 of Xu such as a box of FIG. 1 would capture grid cells in all directions x, y, x hence can be understood to be analogous to multiple grid cells as claimed). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie in view of Park further in view of Krishnan and Yu of having a computer implemented method of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a self-supervised loss regression function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the self-supervised regression loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Xu of having the mapping is from a grid cell of the first data representation to a region of the second representation spanning multiple grid cells thereof. Wherein, Xie’s method of having the mapping is from a grid cell of the first data representation to a region of the second representation spanning multiple grid cells thereof. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve point cloud learning for more accurate point cloud classification task by considering quantized information of grid points and grid space. Since both Xie and Xu both perform point cloud learning for classification tasks. Wherein Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Xu’s system improves point cloud learning for more accurate point cloud classification task by considering quantized information of grid points and grid space (see Xu’s Abstract). However, Xie in view of Park further in view of Krishnan and Yu and Xu does not explicitly teach the second vector or scalar determined via interpolation of vectors of scalars of the multiple grid cells. Goersele teaches the second vector or scalar determined via interpolation of vectors of scalars of the multiple grid cells (the two vectors are being mapped using interpolation of such as shown in FIG. 3; Section 5, 1st Par., discloses “varying progress parameter p(t) instead of linear time t as the interpolation parameter…”). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie in view of Park further in view of Krishnan and Yu and Xu of having a computer implemented method of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a self-supervised loss regression function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the self-supervised regression loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Goersele of having the second vector or scalar determined via interpolation of vectors of scalars of the multiple grid cells. Wherein, Xie’s method of having the second vector or scalar determined via interpolation of vectors of scalars of the multiple grid cells. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve 3D point cloud rendering more effective by interpolation approach. Since both Xie and Goersele both perform point cloud learning for 3D rendering tasks. Wherein Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Goersele’s system improves 3D point cloud rendering more effective by interpolation approach (see Goersele’s Abstract). Claims 20-21 are rejected under 35 U.S.C. 103 as being unpatentable over Saining Xie et. al. (“PointContrast: Unsupervised Pre-Training for 3D Point Cloud Understanding, Dec. 2020, Part of the book series: Lecture Notes in Computer Science, Vol. 12348” hereinafter as “Xie”) in view of Minwoo Park et. al. (“US 2021/0166052 A1” hereinafter as “Park”) further in view of Dilip Krishnan et. al. (“US 2021/0326660 A1” hereinafter as “Krishnan”) and Jingyi Yu (“US 2020/0074658 A1” hereinafter as “Yu”) and Haoming Lu et. al. (“Deep Learning for 3D Point Cloud Understanding: A Survey, Sept. 2020, Computer Vision and Pattern Recognition, Machine Learning” hereinafter as “Lu”). Regarding claim 20, Xie teaches a computer system comprising: at least one memory configured to store computer-readable instructions; at least one hardware processor coupled to the at least one memory and configured to execute the computer-readable instructions, which upon execution cause the at least one hardware processor to extract features from sensor data, by: (abstract discloses using of an encoder for computer processing hence, indicates the use of a computer to have computer components to perform computer component functions such as a processor to execute instructions of the invention stored in a memory; abstract and section 3.5 disclose straining of an encoder for processing data from sensor): generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data (section 3.5 discloses “UNet architecture that has an encoder network…this architecture as a unified design for both the pre-training task…” wherein having training/pre-training of the network of encoder, Fig. 2 illustrates sensor to obtain x1 and x2 being a set of sensor data; moreover, Section 3.3 discloses “FCGF focuses on local descriptor learning” indicating a learning/training and “generate two views x1 and x2 that are aligned in the same world coordinates” wherein, aligned x1 and x2 are data representation of a set of sensor data used for learning/training); and training the encoder based on a supervised loss function applied to the training examples (section 3.1 discloses the encoder is trained based on supervised learning “pre-train an encoder network…we use full supervision” training using contrastive loss function; Section 3, 1st Par., discloses “we introduce our supervised pre-training solution…loss function (Sect. 3.4)” and Section 3.4, 1st Par., discloses “the first loss function, hardest-contrastive loss…”); wherein the encoder extracts respective features from the at least two data representations of each training example (Fig. 2 illustrates the encoder network extracts f1 and f2 features from the aligned x1 and x2; Section 3.4, 1st Par. discloses “matched pairs of points x1_i and x2_j from two views x1 and x2, and f1_i and f2_j are associated point features for the matched pair”), and at least one numerical output value is computed from the extracted features (Algorithm 1 of section 3.2 discloses “compute point features f1, f2…by f1 = NN(T1(x1)) and f2 = NN(T2(x2))” being the computation of such point features, indicating a numerical output value being associated with each of the point features being computed), wherein the supervised loss function is configured to drive the at least one numerical output value to match the at least one numerical transformation value parameterizing the transformation (Section 3.4, 1st 2 Pars., discloses “the first loss function…set of matched pairs of points x1_i and x2_j…the loss encourages a query q to be similar to its positive key k+ and dissimilar to, typically many, negative keys k-” having wherein the loss function to encourage matching between the feature points [at least one numerical output value and the at least one numerical as claimed]); wherein the encoder is configured to receive an input sensor data representation and extract features therefrom (as shown in FIG. 2 of the encoder network taking in the data representations from sensors and extract features from them). However, Xie does not explicitly teach wherein the loss function being a regression loss function. Park teaches wherein the loss function being a regression loss function (as shown in FIG. 1 as the training can include regression loss and further disclosed in [0006]). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie of having a computer implemented method of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a supervised loss function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the supervised loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Park of having wherein the loss function being a regression loss function. Wherein having Xie’s method of having wherein the loss function being a regression loss function. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve training of an encoder network on producing accurate predictions based on using regression loss function. Since both Xie and Park both perform training of an encoder network on point cloud processing. Wherein Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Park’s system improves training of an encoder network on producing accurate predictions based on using regression loss function (see Park’s Par. [0006]). However, Xie in view of Park does not explicitly teach wherein the supervised loss function being a self-supervised loss function. Krishnan teaches wherein the supervised loss function being a self-supervised loss function (Par. [0036] discloses “contrastive learning loss for self-supervised representation learning. Next it is shown how this loss can be modified to be suitable for fully supervised learning, while simultaneously preserving properties important to the self-supervised approach”). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie in view of Park of having a computer implemented method of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a supervised loss regression function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the supervised regression loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Krishnan of having wherein the supervised loss function being a self-supervised loss function. Wherein having Xie’s method of having wherein the supervised loss function being a self-supervised loss function. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve training method using contrastive learning by enabling learning to occur simultaneously across multiple positive examples. Since both Xie and Krishnan both perform supervised contrastive learning for learning on classification tasks. Wherein Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Krishnan’s system improves training method using contrastive learning by enabling learning to occur simultaneously across multiple positive examples (see Krishnan’s Par. [0006]). However, Xie in view of Park and Krishnan does not explicitly teach wherein a first data representation is a transformed data representation of a second data representation, the at least one numerical value being the at least one numerical transformation value parameterizing the transformation. Yu teaches wherein a first data representation is a transformed data representation of a second data representation (Par. [0033] discloses “a first light field F, is captured by the LF camera at a first viewpoint k, and a second light field, F’, is captured by the LF camera at a second viewpoint k+1, and F’ is aligned to the world coordinates…within F, if the transformation between F and F’ is known” and Par. [0034] discloses “feature matching across the sub-aperture images to find matched features…between two different LF images” wherein having images of different viewpoints to match features through a transformation), the at least one numerical value being the at least one numerical transformation value parameterizing the transformation (Par. [0032] discloses “LF pose estimation…can be described in ray space…alpha and teta to parametrize the ray direction”; Par. [0033] discloses “a first light field F, is captured by the LF camera at a first viewpoint k, and a second light field, F’, is captured by the LF camera at a second viewpoint k+1, and F’ is aligned to the world coordinates…given a ray…within F, if the transformation between F and F’ is known” wherein having the matching between feature points being a transformation of another with ray direction parameterized by some vectors). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie in view of Park and Krishnan of having a computer implemented method of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a self-supervised loss regression function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the self-supervised regression loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Yu of having wherein a first data representation is a transformed data representation of a second data representation, the at least one numerical value being the at least one numerical transformation value parameterizing the transformation. Wherein Xie’s method of having wherein a first data representation is a transformed data representation of a second data representation, the at least one numerical value being the at least one numerical transformation value parameterizing the transformation. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve providing accurate results on 3D model reconstruction through utilizing ray transformation. Since both Xie and Yu both perform matching feature point across different viewpoints to world coordinates. Wherein Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Yu’s system improves accurate results on 3D model reconstruction through utilizing ray transformation (see Yu’s Pars. [0003-0004]). However, Xie in view of Park further in view of Krishnan and Yu does not explicitly teach a perception component; and the perception component is configured to use the extracted features to interpret the input sensor data representation. Lu teaches a perception component (the features of the point cloud also include perception features such as disclosed in section 5.2.2, 3rd to the last par.); and the perception component is configured to use the extracted features to interpret the input sensor data representation (the perception features are being used to generate conceptual models to enhance feature [section 5.2.2, 3rd to the last par.] to interpret the input sensor data). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie in view of Park and Krishnan and Yu of having a computer implemented method of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a self-supervised loss regression function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the self-supervised regression loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Lu of having a perception component; and the perception component is configured to use the extracted features to interpret the input sensor data representation. Wherein Xie’s method of having a perception component; and the perception component is configured to use the extracted features to interpret the input sensor data representation. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve providing accurate results on 3D model reconstruction through utilizing ray transformation. Since both Xie and Lu both perform deep learning for 3D point cloud. Wherein Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Lu’s system improves learning about 3D point cloud more effectively and enhance understanding of point cloud features (see Lu’s abstract and section 5.2.2, 3rd to the last Par.). Regarding claim 21, Xie in view of Park and Krishnan and Yu and Lu teach the computer system of claim 20. However, Xie in view of Park and Krishnan and Yu does not explicitly teach wherein the perception component is configured to perform a task on the extracted features. Lu teaches wherein the perception component is configured to perform a task on the extracted features (the features of the point cloud also include perception features such as disclosed in section 5.2.2, 3rd to the last par.; the perception features are being used to generate conceptual models to enhance feature [section 5.2.2, 3rd to the last par.] to interpret the input sensor data). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie in view of Park and Krishnan and Yu of having a computer implemented method of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a self-supervised loss regression function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the self-supervised regression loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Lu of having wherein the perception component is configured to perform a task on the extracted features. Wherein, Xie’s method of having wherein the perception component is configured to perform a task on the extracted features. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve providing accurate results on 3D model reconstruction through utilizing ray transformation. Since both Xie and Lu both perform deep learning for 3D point cloud. Wherein Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Lu’s system improves learning about 3D point cloud more effectively and enhance understanding of point cloud features (see Lu’s abstract and section 5.2.2, 3rd to the last Par.). However, Xie in view of Krishnan and Yu and Lu does not explicitly teach the task being a regression task. Park teaches the task being a regression task (as shown in FIG. 1 as the training can include regression loss and further disclosed in [0006]). Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Xie in view of Krishnan and Yu and Lu of having a computer implemented method of training an encoder to extract features from sensor data, the method comprising: generating a plurality of training examples, each training example comprising at least two data representations of a set of sensor data; and training the encoder based on a supervised loss function applied to the training examples; wherein the encoder extracts respective features from the at least two data representations of each training example, and at least one numerical output value is computed from the extracted features, wherein the supervised loss function is configured to drive the at least one numerical output value to match the at least one numerical, with the teachings of Park of having wherein the loss function being a regression loss function. Wherein having Xie’s method of having wherein the loss function being a regression loss function. The motivation behind the modification would have been to use contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved, and further to improve training of an encoder network on producing accurate predictions based on using regression loss function. Since both Xie and Park both perform training of an encoder network on point cloud processing. Wherein Xie’s system improves training of an encoder network by using contrastive loss for pre-training on unified triplet architecture of neural network so that improvement on machine learning tasks such as segmentation and detection are achieved (see Xie’s Abstract), and Park’s system improves training of an encoder network on producing accurate predictions based on using regression loss function (see Park’s Par. [0006]). Pertinent Prior Art(s) The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Stoytchev, Alexander et. al., “US 2023/0385372 A1”, discloses methods for encoding, decoding, and matching patterns in collections of signals. These methods use weighting functions to scale the signals. This scaling enables the use of signals of arbitrary duration, wherein the signals may include discrete sequences and spike trains. In the most general case, the signals can be represented using functionals, which extends the expressive power of the methods. Further disclosed herein are embodiments of a system that performs these methods. Bonnery, Christophe et. al., “US 8175375 B2”, discloses a method of compression of videotelephony images characterized by: creating a learning base containing images; centering the learning base about zero; determining component images by principal component analysis; and keeping a number of significant principal components. He, Pengcheng et. al., “US 12061876 B2”, discloses facilitating the building and use of natural language understanding models. The systems and methods identify a plurality of tokens and use them to generate one or more pre-trained natural language models using a transformer. The transformer disentangles the content embedding and positional embedding in the computation of its attention matrix. Systems and methods are also provided to facilitate self-training of the pre-trained natural language model by utilizing multi-step decoding to better reconstruct masked tokens and improve pre-training convergence. Wang, Yan et. al., “US 11669558 B2”, discloses technique for generating a dense embedding vector that provides a distributed representation of input text. The technique includes: generating an input term-frequency (TF) vector of dimension g that includes frequency information relating to frequency of occurrence of terms in an instance of input text; using a TF-modifying component to modify the term-specific frequency information in the input TF vector by respective machine-trained weighting factors, to produce an intermediate vector of dimension g; using a projection component to project the intermediate vector of dimension g into an embedding vector of dimension k, where k is less than g. Both the TF-modifying component and the projection component use respective machine-trained neural networks. An application performs any of a retrieval-based function, a recognition-based function, a recommendation-based function, a classification-based function, etc. based on the embedding vector. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHUONG HAU CAI whose telephone number is (571)272-9424. The examiner can normally be reached M-F 8:30 am - 5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PHUONG HAU CAI/ Examiner, Art Unit 2673 /CHINEYERE WILLS-BURNS/Supervisory Patent Examiner, Art Unit 2673
Read full office action

Prosecution Timeline

Show 1 earlier event
Oct 23, 2025
Non-Final Rejection mailed — §103
Jan 23, 2026
Response Filed
Apr 24, 2026
Final Rejection mailed — §103
Jun 29, 2026
Applicant Interview (Telephonic)
Jul 01, 2026
Examiner Interview Summary
Jul 24, 2026
Request for Continued Examination
Jul 28, 2026
Response after Non-Final Action
Sep 23, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12700217
SYSTEM AND METHOD FOR PROCESSING TRAINING DATASET ASSOCIATED WITH SYNTHETIC IMAGE
3y 8m to grant Granted Aug 04, 2026
Patent 12688683
Method and System for Optimization of a Human-Machine Team for Geographic Region Digitization
2y 2m to grant Granted Jul 21, 2026
Patent 12682605
SYSTEMS, METHODS, AND APPARATUS FOR IMAGE CLASSIFICATION WITH DOMAIN INVARIANT REGULARIZATION
3y 10m to grant Granted Jul 14, 2026
Patent 12639955
AUTOMATED VEHICLE IDENTIFICATION BASED ON CAR-FOLLOWING DATA WITH MACHINE LEARNING
3y 9m to grant Granted May 26, 2026
Patent 12632931
INSPECTION SYSTEM, IMAGE PROCESSING METHOD, AND DEFECT INSPECTION DEVICE
3y 7m to grant Granted May 19, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
77%
Grant Probability
99%
With Interview (+26.4%)
2y 11m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 117 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month