Prosecution Insights
Last updated: August 15, 2026
Application No. 18/293,246

PERCEPTION OF 3D OBJECTS IN SENSOR DATA

Final Rejection §102§103
Filed
Jan 29, 2024
Priority
Jul 29, 2021 — GB 2110950.9 +1 more
Examiner
BARHAM, RYAN ALLEN
Art Unit
2613
Tech Center
2600 — Communications
Assignee
Five AI Limited
OA Round
3 (Final)
56%
Grant Probability
Moderate
4-5
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 56% of resolved cases
56%
Career Allowance Rate
9 granted / 16 resolved
-5.7% vs TC avg
Strong +54% interview lift
Without
With
+53.8%
Interview Lift
resolved cases with interview
Typical timeline
2y 4m
Avg Prosecution
23 currently pending
Career history
39
Total Applications
across all art units

Statute-Specific Performance

§101
2.4%
-37.6% vs TC avg
§103
49.6%
+9.6% vs TC avg
§102
44.9%
+4.9% vs TC avg
§112
2.4%
-37.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 16 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1, 3-4, 6-9, 11-12, 14, 16-17, and 19-23 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Li (US 11354847 B2). Regarding claim 1, Li teaches a computer-implemented method of locating and modelling a 3D object captured in multiple time-series of sensor data of multiple sensor modalities (col. 1, lines 43-49: “A 3D object reconstruction neural network system learns to predict a 3D representation of an object from a video that includes the object. An object in a video maintains temporal consistency, having a consistent shape and consistent texture across multiple frames. The temporal consistency of the object is exploited to reconstruct a dynamic 3D representation of the object from an unlabeled video.”), the method comprising: optimizing a cost function applied to the multiple time-series of sensor data (col. 6, lines 20-26: “The shape model is represented by an identity or base shape V.sub.base and an offset or deformation term ΔV, in which the identity shape V.sub.base intuitively corresponds to the “identity” of the instance, e.g., a duck, or a flying bird, etc. During online adaptation, the neural network model 150 is trained to predict consistent V.sub.base to preserve the identity shape over time, via a swapping loss function.”), wherein the cost function aggregates over time and the multiple sensor modalities (col. 10, lines 21-25: “In an embodiment, the part maps are aggregated to produce a video level part UV map 235. By aggregating the part UV maps, i.e., averaging, noise is minimized in each individual part UV map.”), and is defined over a set of variables, the set of variables comprising: one or more shape parameters of a 3D object model (col. 8, lines 43-47: “The loss function 215 may adjust parameters of the shape decoder 115, motion decoder 120, texture decoder 125, and/or camera pose unit 225 to reduce differences between the rendered images and video frames.”), a time sequence of poses of the 3D object model, each pose comprising a 3D object location and 3D object orientation (col. 8, lines 27-33: “The camera pose unit 225 receives the features and predicts a camera pose (position and orientation), θ of the identity shape based on the features. The shape decoder 115, motion decoder 120, texture decoder 125, and camera pose unit 225 may be configured to jointly predict the identity shape, offsets, texture image, and camera pose.”), and one or more motion parameters of a motion model for the 3D object (col. 8, lines 43-47, as above); wherein the cost function penalizes inconsistency between: the multiple time-series of sensor data and the set of variables (col. 10, lines 40-46: “Discrepancies between the parts rendered back to the 2D space and the part propagations, for each frame, are penalized. In an embodiment, the loss function 215 may be configured to update parameters of the neural network model 150 based on the discrepancies. In an embodiment, the parameters are updated to encourage consistency between the rendered images and the propagated part maps.”), and the time sequence of poses and the motion model (col. 10, lines 40-46, as above), wherein the object belongs to a known object class, and the 3D object model or the cost function encodes expected 3D shape information associated with the known object class, whereby the 3D object is located at multiple time instants and modelled, and motion of the 3D object is modelled, by tuning each pose, the shape parameters and the one or more motion parameters, with the objective of optimizing the cost function (col. 2, line 66 – col. 3, line 3: “Prior to inferencing, the neural network system is trained to jointly predict the shape, texture, and camera pose of an image for category-specific 3D reconstruction using a collection of single-view images of the same category.”). Regarding claim 3, Li teaches the method of claim 1, wherein at least one of the multiple time-series of sensor data comprises a piece of sensor data which is not aligned in time with any pose of the time sequence of poses, the method comprising: using the motion model to compute, from the time sequence of poses, an interpolated pose that coincides in time with the piece of sensor data (col. 5, lines 45-49: “In an embodiment, the identity shape is computed as a sum of component shapes included in the set of learned shape bases and each component shape is corresponding scaled by a coefficient generated by the neural network model.”), wherein the cost function penalizes inconsistency between the piece of sensor data and the interpolated pose (col. 10, lines 40-46, as above in claim 1 rejection). Regarding claim 4, Li teaches the method of claim 3, wherein the at least one time-series of sensor data comprises a time-series of images, and the piece of sensor data is an image (col. 3, lines 42-44: “The encoder 105 extracts features 110 from each frame (e.g., image) in the video.”). Regarding claim 6, Li teaches the method of claim 1, wherein: the variables additionally comprise one or more object dimensions for scaling the 3D object model, the shape parameters being independent of the object dimensions (col. 5, lines 45-53: “In an embodiment, the identity shape is computed as a sum of component shapes included in the set of learned shape bases and each component shape is corresponding scaled by a coefficient generated by the neural network model. In an embodiment, shape offsets (e.g., non-rigid motion deformations) are computed and applied to the vertices of the identity shape to predict the 3D shape representation of the object. In an embodiment, the 3D shape representation is a mesh of vertices that define faces.”); or the shape parameters of the 3D object model encode both 3D object shape and object dimensions (col. 4, lines 25-27: “The offsets 118 encode the object's asymmetric non-rigid motion, defining deformations for each vertex in the identity shape 116.”). Regarding claim 7, Li teaches the method of claim 1, wherein the cost function additionally penalizes each pose to an extent the pose violates an environmental constraint (col. 12, line 65 – col. 13, line 8: “The ARAP constraint is an additional loss objective that may be used for self-supervised training of the 3D object construction system 100. The identity shape is smooth by construction. However, application of the offsets predicted by the motion decoder 120 may produce discontinuities in the 3D mesh representation. The ARAP constraint is used to ensure that edge lengths of the 3D mesh are maintained even when the 3D mesh is rotated. ARAP is a self-supervised regularization that maintains rigidity and can be used on individual input images and video sequences.”). Regarding claim 8, Li teaches the method of claim 7, wherein the environmental constraint is defined relative to a known 3D road surface (col. 26, lines 47-52: “Furthermore, images generated applying one or more of the techniques disclosed herein may be used to train, test, or certify DNNs used to recognize objects and environments in the real world. Such images may include scenes of roadways, factories, buildings, urban settings, rural settings, humans, animals, and any other physical object or real-world setting.”). Regarding claim 9, Li teaches the method of claim 8, wherein each pose is used to locate the 3D object model relative to the road surface, and the environmental constraint penalizes each pose to the extent the 3D object model does not lie on the known 3D road surface (col. 29, lines 32-41: “The rasterization stage 660 may be configured to utilize the vertices of the geometric primitives to setup a set of plane equations from which various attributes can be interpolated. The rasterization stage 660 may also compute a coverage mask for a plurality of pixels that indicates whether one or more sample locations for the pixel intercept the geometric primitive. In an embodiment, z-testing may also be performed to determine if the geometric primitive is occluded by other geometric primitives that have already been rasterized.”). Regarding claim 11, Li teaches the method of claim 1, wherein at least one of the sensor modalities is such that the poses and the shape parameters are not uniquely derivable from that sensor modality alone (col. 4, lines 51-54: “The UV texture space provides a parameterization that is invariant to object deformation. Therefore, over time, the image texture for the images in the video should be constant or invariant to shape deformation.”). Regarding claim 12, Li teaches the method of claim 1, wherein: one of the multiple time-series of sensor data is a time-series of radar data encoding measured Doppler velocities, wherein the time sequence of poses and the 3D object model are used to compute expected Doppler velocities, and the cost function penalizes discrepancy between the measured Doppler velocities and the expected Doppler velocities; or one of the multiple time-series of sensor data is a time-series of images, and the cost function penalizes an aggregate reprojection error between (i) the images and (ii) the time sequence of poses and the 3D object model (col. 10, lines 50-56: “In an embodiment, instead of minimizing the discrepancy between the rendered part map (e.g., rendered image, such as the rendered image 238) and the propagated part map of a frame, it may be more robust to penalize the geometric distance between the projections of vertices assigned to each part with 2D points sampled from the corresponding part.” NOTE: use of the term “frame” implies a time-series of images, i.e. a video.); or one of the multiple time-series of sensor data is a time-series of lidar data, wherein the cost function is based on a point-to-surface distance between lidar points and a 3D surface defined by the parameters of the 3D object model, wherein the point-to-surface distance is aggregated across all points of the lidar data. NOTE: the repeated use of the word “or” in the above claim implies that only one of the aforementioned time-series of sensor data is necessary to encompass the entirety of the claim. Regarding claim 14, Li teaches the method of claim 12, wherein a semantic keypoint detector is applied to each image (col. 11, lines 37-42: “When the 3D object construction system 100 is trained using weak supervision, 2D keypoints may be provided as ground truth annotations that semantically associate different instances of the object. For example, an annotated frame 240 includes multiple keypoints, such as a tail keypoint 248 at the tip of the bird's tail.”), and the reprojection error is defined on semantic keypoints of the object (col. 11, lines 43-46: “When the 2D keypoints are projected onto the 3D representation (e.g., mesh surface), the same semantic keypoint for different object instances should be matched to the same face on the mesh surface.”). Regarding claim 16, Li teaches the method of claim 12, wherein the 3D object model is encoded as a distance field (col. 10, lines 50-56: “In an embodiment, instead of minimizing the discrepancy between the rendered part map (e.g., rendered image, such as the rendered image 238) and the propagated part map of a frame, it may be more robust to penalize the geometric distance between the projections of vertices assigned to each part with 2D points sampled from the corresponding part.”). Regarding claim 17, Li teaches the method of claim 1, wherein: the expected 3D shape information is encoded in the 3D object model (col. 3, lines 40-42: “The 3D object construction system 100 includes a neural network model comprising at least an encoder 105, shape decoder 115, and motion decoder 120.”), the 3D object model learned from a set of training data comprising example objects of the known object class (col. 2, line 66 – col. 3, line 3: “Prior to inferencing, the neural network system is trained to jointly predict the shape, texture, and camera pose of an image for category-specific 3D reconstruction using a collection of single-view images of the same category.”); or the expected 3D shape information is encoded in a regularization term of the cost function, which penalizes discrepancy between the 3D object model and a 3D shape prior for the known object class (col. 4, lines 55-63: “By enforcing the predicted values for the texture images 106 to be consistent in the UV texture space across different frames of a video, the neural network model can be regularized to generate coherent reconstructions over time during inferencing. The temporal invariance may be used as self-supervised signals to tune the neural network model. In an embodiment, during inferencing, the 3D object construction system 100 is adapted using self-supervised regularization based on shape invariance and texture invariance.”). Regarding claim 19, Li teaches the method of claim 1, comprising: using an object classifier to determine the known class of the object from multiple available object classes, the multiple object classes associated with respective expected 3D shape information (col. 2, line 66 – col. 3, line 7: “Prior to inferencing, the neural network system is trained to jointly predict the shape, texture, and camera pose of an image for category-specific 3D reconstruction using a collection of single-view images of the same category. A first example category may include, but is not limited to, birds (including ducks). A second example category may be horses. In general, a category includes animals that have a similar structure, such as animals within a single species.”). Regarding claim 20, Li teaches the method of claim 1, wherein the one or more shape parameters are applied to each pose of the time sequence of poses for modelling a rigid object (col. 13, lines 4-8: “The ARAP constraint is used to ensure that edge lengths of the 3D mesh are maintained even when the 3D mesh is rotated. ARAP is a self-supervised regularization that maintains rigidity and can be used on individual input images and video sequences.”). Regarding claim 21, Li teaches the method of claim 1, wherein the 3D object model is a deformable model, with at least one of the shape parameters varied across frames (col. 13, lines 30-33: “Non-rigid motion deformations of the 3D shape representations are predicted for the frames and applied to identity shapes predicted for the frames to produce the 3D shape representations of the object.”). Claim 22 is substantially similar to claim 1, and differs only in that it teaches a system rather than a method. As such, it is rejected on a similar basis to claim 1. Claim 23 is substantially similar to claim 1, and differs only in that it teaches a medium rather than a method. As such, it is rejected on a similar basis to claim 1. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 5 and 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li (US 11354847 B2) as applied to claims 1 and 3 above, and further in view of Keilaf (US 10698114 B2). Regarding claim 5, Li teaches the method of claim 3, but fails to teach wherein the at least one time-series of sensor data comprises a time-series of lidar or radar data, the piece of sensor data is an individual lidar or radar return, and the interpolated pose coincides with a return time of the lidar or radar return. Keilaf teaches wherein the at least one time-series of sensor data comprises a time-series of lidar or radar data, the piece of sensor data is an individual lidar or radar return, and the interpolated pose coincides with a return time of the lidar or radar return (col. 16, lines 38-46: “In some embodiments, the light deflector may be moved such that during a scanning cycle of the LIDAR FOV the light deflector is located at a plurality of different instantaneous positions. In other words, during the period of time in which a scanning cycle occurs, the deflector may be moved through a series of different instantaneous positions/orientations, and the deflector may reach each different instantaneous position/orientation at a different time during the scanning cycle.”). It would have been obvious to one familiar in the art prior to the effective filing date of the claimed invention to utilize lidar or radar in the 3D modeling method of Li, as both technologies are in the same field of endeavor of 3D object identification, classification, and modeling. Making use of lidar and radar could obviously enable the invention of Li to be utilized in autonomous vehicles, a technology which Li makes mention of but does not explore in-depth (col. 26, lines 56-58: “Furthermore, such images may be used to train, test, or certify DNNs that are employed in autonomous vehicles to navigate and move the vehicles through the real world.”). Regarding claim 10, Li teaches the method of claim 1, wherein the multiple sensor modalities comprise an image modality (col. 25, line 66 – col. 26, line 6: “A deep neural network (DNN) model includes multiple layers of many connected nodes (e.g., perceptrons, Boltzmann machines, radial basis functions, convolutional layers, etc.) that can be trained with enormous amounts of input data to quickly solve complex problems with high accuracy. In one example, a first layer of the DNN model breaks down an input image of an automobile into various sections and looks for basic patterns such as lines and angles.”), but fails to teach a lidar modality or a radar modality. Keilaf teaches wherein the multiple sensor modalities comprise two or more of: an image modality, a lidar modality, and a radar modality (col. 80, lines 34-45: “In other embodiments, the at least one processor 118 may receive an identification of a region of interest or at least one indicator of a region of interest from one or more sources peripheral to LIDAR system 100. For example, as shown in FIG. 22, such an identification or indicator may be received from vehicle navigation system 2201, radar 2203, camera 2205, GPS 2207, or another LIDAR 2209. Such indicators or identifiers may be associated with mapped objects or features, directional headings, etc. from the vehicle navigation system or one or more objects, clusters of objects, etc. detected by radar 2203 or LIDAR 2209 or camera 2205, etc.”). It would have been obvious to one familiar in the art prior to the effective filing date of the claimed invention to utilize lidar or radar in the 3D modeling method of Li, as both technologies are in the same field of endeavor of 3D object identification, classification, and modeling. Making use of lidar and radar could obviously enable the invention of Li to be utilized in autonomous vehicles, a technology which Li makes mention of but does not explore in-depth (col. 26, lines 56-58: “Furthermore, such images may be used to train, test, or certify DNNs that are employed in autonomous vehicles to navigate and move the vehicles through the real world.”). Response to Arguments Applicant’s arguments, see Remarks, filed 05/19/2026, with respect to the rejection(s) of claim(s) 1, 22, and 23 over Keilaf (US 10698114 B2) and Vineet (US 12283119 B2) have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Li (US 11354847 B2). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to RYAN A BARHAM whose telephone number is (571)272-4338. The examiner can normally be reached Mon-Fri, 8:30am-5pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu, can be reached at (571) 272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /RYAN ALLEN BARHAM/Examiner, Art Unit 2613 /XIAO M WU/Supervisory Patent Examiner, Art Unit 2613
Read full office action

Prosecution Timeline

Jan 29, 2024
Application Filed
Oct 01, 2025
Non-Final Rejection mailed — §102, §103
Jan 02, 2026
Response Filed
Feb 19, 2026
Non-Final Rejection mailed — §102, §103
May 19, 2026
Response Filed
Jun 16, 2026
Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12657785
SIMULATING SHUTTER ROLLING EFFECT
2y 4m to grant Granted Jun 16, 2026
Patent 12646187
METHOD AND DEVICE FOR ALIGNING LASER POINT CLOUD AND IMAGE BASED ON DEEP LEARNING
2y 5m to grant Granted Jun 02, 2026
Patent 12639935
Visual Analytics Framework for Explainable Data Slicing-Based Model Validation
2y 5m to grant Granted May 26, 2026
Patent 12633031
STOCHASTIC TEXTURE FILTERING
2y 4m to grant Granted May 19, 2026
Patent 12564345
MEDICAL APPARATUS, AND IMAGE GENERATION METHOD FOR VISUALIZING TEMPORAL TRENDS OF BIOMAGNETIC DATA ON AN ORGAN MODEL
2y 10m to grant Granted Mar 03, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

4-5
Expected OA Rounds
56%
Grant Probability
99%
With Interview (+53.8%)
2y 4m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 16 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month