Prosecution Insights
Last updated: August 16, 2026
Application No. 19/044,430

Plane Estimation via 3D-Consistent Embeddings

Non-Final OA §103§112
Filed
Feb 03, 2025
Priority
Feb 02, 2024 — provisional 63/549,154
Examiner
AHMAD, NAUMAN UDDIN
Art Unit
Tech Center
Assignee
Niantic, Inc.
OA Round
1 (Non-Final)
77%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
34 granted / 44 resolved
+17.3% vs TC avg
Strong +23% interview lift
Without
With
+23.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
32 currently pending
Career history
74
Total Applications
across all art units

Statute-Specific Performance

§101
4.6%
-35.4% vs TC avg
§103
72.7%
+32.7% vs TC avg
§102
4.2%
-35.8% vs TC avg
§112
15.6%
-24.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 44 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claims 1-2, 6-7, 11-12 and 16-17 objected to because of the following informalities: typos as follows. Claim 1 and 11, “the distance” should be “a distance”. Claim 2 and 12 “labels to respective the” should read “labels respective to the”. Claims 6 and 16 “the 3D scene; determining that the first area extends across at least two of the plane instances;” should read “the [[3D]] scene; determining that the first area extends across at least two of the 3D plane instances;”. Claims 6 and 16 “the area” should read “the first area”. Claims 7 and 17 “second MLP the second room” should read “second MLP on the second room”. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 4-6, 9, 14-16 and 19 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 4 and 14 recites the limitation "generating 2D embeddings" in line 5. There is insufficient antecedent basis for this limitation in the claim. This is because it is unclear if this is the same as the 2D embedding from “2D embedding generated based on training images” in line 2. Claims 5 and 15 recites the limitation "into 3D plane instances…clustering of 3D plane instances…clustering of 3D plane instances" in lines 2-3 and 5. There is insufficient antecedent basis for this limitation in the claim. This is because it’s unclear if this is referring to same 3D plane instances as parent claims 1 and 11 or newer/different instances of 3D plane instances. Claims 6 and 16 recites the limitation "type of AR object" in lines 4-5 and 5-6 (respectively of each claim). There is insufficient antecedent basis for this limitation in the claim. This is because it’s unclear if this is same as “AR object” in lines 2 and 3 (respectively of each claim) as aforementioned. Claims 9 and 19 recites the limitation "a scene" in lines 1-2. There is insufficient antecedent basis for this limitation in the claim. This is because it’s unclear if this is the same scene as aforementioned in parent claims. Note. Most likely these claims depend on some dependent claim or are missing elements. In order to fix this issue, dependency should be reviewed and any first instance of an element should be made clear that it’s a first instance and should be referred to as “a” or “an” instead of “the”, and if multiple instances exist, further instances should be further distinguished for example by saying “first”, “second”, and/or “third” etc. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-3 and 11-13 is/are rejected under 35 U.S.C. 103 as being unpatentable over GHAZVINIAN ZANJANI (U.S. Patent Application Publication No. 2024/0386650), hereinafter referenced as Ghaz, in view of Xie et al. (U.S. Patent Application Publication No. 2021/0256680), hereinafter referenced as Xie. Regarding claim 1, Ghaz teaches a computer-implemented method comprising: determining a camera pose for each image of a sequence of images of a scene, (paragraph 66 teaches “plurality of images depicting multiple views of a scene by a camera…view frustrum of each image can be determined based on camera pose information”); camera pose information for each image shows camera pose determined for each image in the sequence of images of the scene; the scene comprising planar surfaces in the scene (paragraph 43 teaches “a plurality of planar meshes can be generated based on image data of a scene” and fig. 3 shows planar surfaces in scene as well); planar meshes indicates the scene comprises planar surfaces; generating a three-dimensional (3D) mesh using the sequence of images and respective camera poses, (paragraph 64 teaches “Based on the 2D input images 312-318 and the extrinsic and/or intrinsic information of monocular moving camera….generate 3D meshes corresponding to planar surfaces within the scene”); based on images and information of camera shows the generation of 3D mesh here would be done using such; the 3D mesh comprising vertices corresponding to respective locations in the scene (paragraph 64 teaches “label each vertex of a plurality of vertices included in the 3D mesh”); vertices included in 3D mesh indicate vertices which correspond to respective locations since the mesh is a model of physical environment which has locations (therefore vertex at specific location in mesh would correspond to that respective location in environment/scene); mapping vertices of the 3D mesh to a 3D embedding space using a machine learning model, (paragraph 113 teaches “rendering engine 470 can determine semantic information {s.sub.i}… a respective embedding s, can be determined for each of the i vertices” and paragraph 143 teaches “the MLP machine learning network can be included in and/or implemented by the rendering engine 470”); since rendering engine implements/includes neural network, this is done by using a machine learning model and respective embedding determined for each of the vertices shows mapping vertices (of aforementioned 3D mesh) to a 3D embedding space; clustering the vertices of the 3D mesh into 3D plane instances based on respective 3D embeddings of the 3D embedding space, wherein the 3D plane instances correspond to the respective planar surfaces in the scene (paragraph 83 teaches “Pixels lying on the same planar surface (e.g., same plane) can be associated with same or highly similar values of the planar depth d.sub.pl. For example, the pixels associated with the planar surface of the table seen in D-map 630a may have same or highly similar planar depth values d.sub.pl, the pixels associated with the planar surface of the floor (or wall) seen in D-map 630a may have same or similar planar depth values”); this shows pixels (including vertices of 3D mesh) clustered (using same or highly similar depth values) into plane instances such as floor or wall, each of which plane instance is a respective (thus corresponds to) planar surface in the scene (since floor would be its own planar surface versus the wall); determining a location of a virtual element in an augmented reality (AR) environment based on the 3D plane instances (paragraph 37 teaches “3D mesh generated using 3D planar reconstruction can be used to determine accurate 3D depth information to more accurately perform object placement for one or more virtual objects that are placed and anchored onto planar surfaces, …3D meshes of an environment can be used to identify one or more planar surfaces that can be used (e.g., by an XR or AR navigation application)”); determine depth information to accurately place object shows determining location of virtual element/object (in AR) and this is done using planar instances thus based on such; and causing placement of the virtual element at the location for display at a client device (paragraph 37 teaches “one or more virtual objects that are placed and anchored onto planar surfaces, and to improve the realism of the XR scene” and paragraph 309 teaches “further comprising a display configured to output the reconstructed planar mesh based on rendering”); this shows virtual object/element placed at the aforementioned accurate location determined and displaying based on the rendering shows this would be for display at client device(which has the display). However, Ghaz fails to teach wherein a given pair of vertices on the same planar surface in the scene correspond to a pair of 3D embeddings of the 3D embedding space, the distance between the pair of 3D embeddings not exceeding a threshold distance; However, Xie teaches wherein a given pair of vertices on the same planar surface in the scene correspond to a pair of 3D embeddings of the 3D embedding space, the distance between the pair of 3D embeddings not exceeding a threshold distance (Xie, paragraph 145 teaches “If a distance between two embedding vectors corresponding to two corner points is less than the threshold, it indicates that the top-left corner point and the bottom-right corner point belong to a same object”); the two embedding vectors shows a pair of 3D embeddings of the embedding space (when viewed in combination) and since they correspond to two/pair of corner/vertex points of same object, they would be on the same planar surface, lastly, since distance between two embedding vectors is less than threshold, the pair of 3D embeddings are not exceeding threshold distance. Xie is considered to be analogous art because it is reasonably pertinent to the problem faced by the inventor of embeddings corresponding to vertices in pairs. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Ghaz's invention with the correspondence and distance threshold techniques of Xie to ensure target detection accuracy can be improved (Xie, paragraph 30). This would led to more realistic renderings. Regarding claim 2, the combination of Ghaz and Xie teaches wherein the vertices are a first subset of vertices of the 3D mesh and wherein a second subset of vertices of the 3D mesh are not clustered, further comprising: (Ghaz, paragraph 83 teaches “Pixels lying on the same planar surface (e.g., same plane) can be associated with same or highly similar values of the planar depth d.sub.pl. For example, the pixels associated with the planar surface of the table seen in D-map 630a may have same or highly similar planar depth values d.sub.pl, the pixels associated with the planar surface of the floor (or wall) seen in D-map 630a may have same or similar planar depth values”); this shows pixels (including vertices of 3D mesh) clustered (using same or highly similar depth values) into plane instances such as floor or wall which can be considered first subset of vertices, however, there is no mention of vertices of ceiling plane instance thus second subset of vertices (of ceiling) would not be clustered using the depth values; assigning a set of labels to respective the 3D plane instances (Ghaz, paragraph 128 teaches “assign a unique label to the planar surfaces in the deformed 3D mesh”); this shows set of labels assigned to respective planar surfaces thus respective 3D plane instances (since each of the planar surface corresponds to planar instance as aforementioned and explained in claim 1 of this action above); and grouping the second subset of vertices with respective labeled 3D plane instances (Ghaz, paragraph 114 teaches “associate a unique label to all vertices that belong to a representative planar surface in the 3D mesh of the scene (e.g., in the 3D mesh that includes the plurality of 3D planar meshes).”); label vertices of planar surface shows second subset of vertices of 3D mesh labeled (and grouped using such) as well as with respective labeled 3D plane instances (since second subset of vertices on same/representative plane instance would be associated with plane instance such as ceiling for example). Regarding claim 3, the combination of Ghaz and Xie teaches further comprising: identifying a plane instance of the 3D plane instances comprising fewer than a threshold vertex count (Xie, paragraph 10 teaches “In this implementation, a key point feature in the input image is extracted to determine whether the calibration area in the target frame includes the target feature point”); this shows determining/identifying a frame (also plane instance of the 3D plane instances since frame would have plane instance thereof when viewed in combination), wherein there’s a fewer than threshold (of one) vertex/feature (vertex would correspond to feature when viewed in combination) point count; and removing the identified plane instance from the 3D plane instances (Xie, paragraph 10 teaches “and further remove an erroneous target frame”); removing a frame would remove identified plane instance of the specific frame from the 3D plane instances (when viewed in combination). The same motivations used in claim 1 apply here in claim 3. Regarding claim 11, the non-transitory computer readable medium claim 11 recites similar limitations as method claim 1, and thus is rejected under similar rationale. In addition Ghaz paragraph 11 teaches “aspects include a non-transitory processor-readable storage medium having stored thereon processor-executable instructions configured to cause a processor of a device to perform operations of any of the methods summarized above”. Regarding claim 12, the non-transitory computer readable medium claim 12 recites similar limitations as method claim 2, and thus is rejected under similar rationale. Regarding claim 13, the non-transitory computer readable medium claim 13 recites similar limitations as method claim 3, and thus is rejected under similar rationale. Claim(s) 5 and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Ghaz and Xie as applied to claim 1 and 11 above, and further in view of Isack et al. (Energy-Based Geometric Multi-model Fitting), hereinafter referenced as Isack. Regarding claim 5, the combination of Ghaz and Xie teaches wherein clustering the vertices of the 3D mesh into 3D plane instances comprises: in response to receiving a request for a real-time clustering of 3D plane instances, applying a mean-shift algorithm to the vertices of the 3D mesh (Ghaz, paragraph 85 teaches “segmentation information 440 can be generated based on performing clustering over the feature vector including D-map 430, normal map 427, and positional encoding 432. In one illustrative example, the planar segmentation information 440 can be generated based on performing clustering using a mean-shift clustering algorithm (e.g., based on performing mean-shift clustering)”); since the clustering here is based on performing mean-shift algorithm, that means that in response to receiving request for real-time clustering of 3D plane instances, the mean-shift algorithm would be applied to the vertices of the 3D mesh; However, the combination of Ghaz and Xie fails to teach and in response to receiving a request for an asynchronous clustering of 3D plane instances, applying a random sample consensus (RANSAC) algorithm to the vertices of the 3D mesh by: randomly sampling proposed 3D plane instances, determining an inlier count for each proposed 3D plane instance sampled, and selecting a proposed 3D plane instance having a threshold number of inlier vertices. However, Isack teaches and in response to receiving a request for an asynchronous clustering of 3D plane instances, applying a random sample consensus (RANSAC) algorithm to the vertices of the 3D mesh by: randomly sampling proposed 3D plane instances, (Isack, abstract teaches “data points should be clustered based on geometric proximity to models…regularity of inlier clusters…Our proposed approach (PEARL) combines model sampling from data points as in RANSAC with iterative re-estimation of inliers and models’ parameters based on a global regularization functional”); this shows RANSAC applied and randomly sampling proposed 3D plane instances (since those plane instances would be data points when viewed in combination), also, one of ordinary skill in the art would understand that this would be done in response to request for an asynchronous clustering of 3D plane instances for added efficiency and for the multiple mentions of clustering; determining an inlier count for each proposed 3D plane instance sampled, (Isack, page 125, last column teaches “RANSAC’s energy E(L) counts inliers for L using 0–1 measure”); counts inlier means a inlier count would be determined and this is for each proposed 3D plane instance sampled because it is to minimize errors thus would need to account for all values (each 3D plane instance sampled); and selecting a proposed 3D plane instance having a threshold number of inlier vertices (Isack, page 125, left column, second paragraph teaches “The main goal of RANSAC is to find parameters L of the model with the largest number of inliers within some threshold T. This can be represented as minimization of energy”); this shows selection of parameters that have a threshold number of inliers (when viewed in combination that number of inliers would be for vertices and the selection of parameters would show selecting proposed 3D plane instance). Isack is considered to be analogous art because it is reasonably pertinent to the problem faced by the inventor of inliers and RANSAC algorithm. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Ghaz and Xie with the RANSAC techniques of Isack to significantly improve probability of an accurate model reconstruction from rough initial guesses (Isack, page 129, left column, second paragraph). This would increase/improve realism and user experience. Regarding claim 15, the non-transitory computer readable medium claim 15 recites similar limitations as method claim 5, and thus is rejected under similar rationale. Claim(s) 6, 8-9, 16 and 18-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Ghaz and Xie as applied to claim 1 and 11 above, and further in view of Rau et al. (U.S. Patent Application Publication No. 2022/0051048), hereinafter referenced as Rau. Regarding claim 6, the combination of Ghaz and Xie teaches further comprising: receiving a request to place an AR object on a first area of the 3D scene (Ghaz, paragraph 37 teaches “where the XR game is designed to take place on a virtual table with objects appearing to interact with the real-world surface”); this shows a request must be received to place virtual/AR object on first area/table of scene; determining that the first area extends across at least two of the plane instances (Ghaz, paragraph 91 teaches “edges (e.g., faces F) across distinct planar segments can be suppressed to ensure that the obtained mesh is planar and does not include edges across two planes with different planar depth”); ensuring to not include edges across two planes shows a determination that first area extends across at least two of the plane instances; However, the combination of Ghaz and Xie fails to teach identifying a type of AR object; and in response to determining that the type of AR object cannot be placed on the area, recommending a second area that does not extend beyond one of the at least two plane instances. However, Rau teaches identifying a type of AR object (paragraph 28 teaches “(4) data associated with virtual elements in the virtual world (e.g., positions of virtual elements, types of virtual elements”); this shows type of AR/virtual element/object identified; and in response to determining that the type of AR object cannot be placed on the area, recommending a second area that does not extend beyond one of the at least two plane instances (Rau, paragraph 37 teaches “may generate virtual content or adjust virtual content according to other information received from other components of the client device 110. For example, the gaming module 210 may adjust a virtual object to be displayed on the user interface according to a depth map of the scene captured in the image data (e.g., as generated by a depth estimation model).”); one of ordinary skill in the art would understand that the adjustment must be done once an object (thus type of AR object) can’t be placed on the area and that adjustment follows a recommended second area which would not extend beyond one of the at least two plane instances so that the object could be fully visible (hence displayed on user interface). Rau is considered to be analogous art because it is reasonably pertinent to the problem faced by the inventor of parallel reality game with type of AR object being placed in certain areas. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Ghaz and Xie with the game and AR object placement techniques of Rau to ensure such a method is more efficient and can be significantly cheaper, especially for processing a large number of images (Rau, paragraph 7). This would be partly due to the adjustment of AR object techniques instead of having to re-render the whole object in a different location which would use more processing resources. Regarding claim 8, the combination of Ghaz, Xie and Rau teaches wherein the augmented reality environment is a parallel reality game (Rau, paragraph 22 teaches “Various embodiments are described in the context of a parallel reality game that includes augmented reality content in a virtual world geography that parallels at least a portion of the real-world geography”); this shows the AR environment is a parallel reality game. The same motivations used in claim 6 apply here in claim 8. Regarding claim 9, the combination of Ghaz, Xie and Rau teaches wherein the sequence of images of a scene comprises images of a room taken at two or more angles about a fixed spot in the room (Rau, paragraph 51 teaches “Different images in the image store 410 may be captured at…, the same position but different orientations,”); same position but different orientation indicates the sequence of images of scene would have images of room taken at two or more angles (different orientations thus different angles of rotation) about a fixed spot/position in the room. The same motivations used in claim 6 apply here in claim 9. Regarding claim 16, the non-transitory computer readable medium claim 16 recites similar limitations as method claim 6, and thus is rejected under similar rationale. Regarding claim 18, the non-transitory computer readable medium claim 18 recites similar limitations as method claim 8, and thus is rejected under similar rationale. Regarding claim 19, the non-transitory computer readable medium claim 19 recites similar limitations as method claim 9, and thus is rejected under similar rationale. Claim(s) 7 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Ghaz and Xie as applied to claim 1 and 11 above, and further in view of Hu et al. (U.S. Patent Application Publication No. 2022/0383572), hereinafter referenced as Hu. Regarding claim 7, the combination of Ghaz and Xie teaches wherein the scene is a first room (Ghaz, paragraph 61 teaches “the scene 310 depicted in FIG. 3 is an indoor scene corresponding to an interior room of a house or building”); and wherein the machine learning model is a first multilayer perceptron (MLP), further comprising: (Ghaz, paragraph 143 teaches “the MLP machine learning network can be included in and/or implemented by the rendering engine 470); this shows aforementioned machine learning model as MLP; in response to determining that a user has entered a second room: prompting the user to capture images of the second room (Ghaz, paragraph 61 teaches “a second image 314 can be captured from a second location of the monocular moving camera 311); one of ordinary skill in the art would understand that moving camera inside house indicates user entering second room and second image captured from second location (thus second room) indicates prompting user to capture images of second room. However, the combination of Ghaz and Xie fails to teach training a second MLP the second room. However, Hu teaches training a second MLP the second room (Hu, claim 17 teaches “each of the multiple correlated feature vectors of the multiple rooms into the second MLP to obtain an optimized bounding box corresponding to each room”); this shows second MLP trained on the second room. Hu is considered to be analogous art because it is reasonably pertinent to the problem faced by the inventor of multiple MLP and multiple rooms. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Ghaz and Xie with the second MLP techniques of Hu to improve the quality of the floor plan image (Hu, paragraph 6). This would be done due to the multiple MLP leading to a better quality due to specialized neural networks. Regarding claim 17, the non-transitory computer readable medium claim 17 recites similar limitations as method claim 7, and thus is rejected under similar rationale. Claim(s) 10 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Ghaz and Xie as applied to claim 1 and 11 above, and further in view of CHERNOV et al. (U.S. Patent Application Publication No. 2017/0046868), hereinafter referenced as CHERNOV. Regarding claim 10, the combination of Ghaz and Xie teaches wherein generating the 3D mesh using the sequence of images and the respective camera poses comprises: estimating depth maps based on the sequence of images and the respective camera poses (Ghaz, paragraph 72 teaches “a depth estimation machine learning network 422 can be used to generate a predicted depth map 425 based on the 2D input image 410”); this shows estimation of depth maps which is based on sequences of images and respective camera poses since based on input image; However, the combination of Ghaz and Xie fails to explicitly teach fusing the depth maps into a truncated signed distance function (TSDF); and extracting the 3D mesh based on the TSDF. However, CHERNOV teaches fusing the depth maps into a truncated signed distance function (TSDF) (CHERNOV, claim 9 teaches “by fusing the estimated depth maps by using 3D voxel truncated signed distance function (TSDF)”); this shows fusion of depth maps into TSDF; and extracting the 3D mesh based on the TSDF (CHERNOV, paragraph 124 teaches “During operation 1204 for isosurface extraction, a 3D mesh 1206 is reconstructed” and claim 9 teaches “the generating of the surface mesh further comprises: generating the surface mesh…using 3D voxel truncated signed distance function (TSDF).”); this shows extraction of 3D mesh would be based on the TSDF. CHERNOV is considered to be analogous art because it is reasonably pertinent to the problem faced by the inventor of fusion of depth maps and using TSDF. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Ghaz and Xie with the TSDF techniques of CHERNOV to serve as an iterative solution of the minimization of an energy function (CHERNOV, paragraph 120). This would lead to better alignment (increased accuracy) and reduction of error. Regarding claim 20, the non-transitory computer readable medium claim 20 recites similar limitations as method claim 10, and thus is rejected under similar rationale. Allowable Subject Matter Claims 4 and 14 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: Regarding claim 4, the closest prior art of (or combination of) Ghaz and Xie teaches generating 3D embeddings for vertices of the 3D training mesh using the machine learning model (Ghaz, paragraph 113 teaches “rendering engine 470 can determine semantic information {s.sub.i}… a respective embedding s, can be determined for each of the i vertices” and paragraph 143 teaches “the MLP machine learning network can be included in and/or implemented by the rendering engine 470”); this shows embeddings for vertices (3D since for 3D vertices) being generated using MLP/machine learning, and it would be for the training mesh when the MLP is being trained; determining a first distance between a pair of 3D embeddings corresponding to a pair of vertices in the 3D training mesh, the pair of vertices corresponding to a pair of pixels in a training image of the scene (Xie, paragraph 145 teaches “If a distance between two embedding vectors corresponding to two corner points is less than the threshold, it indicates that the top-left corner point and the bottom-right corner point belong to a same object”); the two embedding vectors shows a pair of 3D embeddings of the embedding space (when viewed in combination) and since they correspond to two/pair of corner/vertex points of same object (from pair of pixels in training image of the scene), they would be on the same planar surface such as a 3D training mesh, also, since distance between two embedding vectors is mentioned, the first distance between such must be determined; Xue et al. (U.S. Patent Application Publication No. 2023/0281207), hereinafter referenced as Xue teaches generating 2D embeddings for pixels of the training images of the scene (Xue, paragraph 42 teaches “for example, a two-dimensional embedding model was used in the prior art); this shows 2D embeddings generated for pixels which would be of training images of the scene when viewed in combination; determining a second distance between a pair of 2D embeddings of respective pixels corresponding to the pair of vertices in the 3D training mesh (Xue, paragraph 42 teaches “all that would matter would be the distance between two embedded points in the two-dimensional space, as the distance would indicate similarity of the underlying data that was embedded”); this shows distance between pair of 2D embeddings determined (which is second distance when viewed in combination and would correspond to pair of vertices in the 3D training mesh since indicates underlying data that was embedded); However, the combination of Ghaz and Xie fails to teach further comprising training the machine learning model using 2D embeddings generated based on training images of the scene by: and updating the machine learning model based on a comparison of the first distance and the second distance. Furthermore, no prior art of record either alone or in combination teaches further comprising training the machine learning model using 2D embeddings generated based on training images of the scene by: and updating the machine learning model based on a comparison of the first distance and the second distance when read in light of the rest of the limitations in claim 4 and the claims to which claim 4 depends and thus claim 4 contains allowable subject matter. Regarding claim 14, the prior art of record either alone or in combination fails to teach wherein the operations further comprise training the machine learning model using 2D embeddings generated based on training images of the scene by: generating 2D embeddings for pixels of the training images of the scene; determining a second distance between a pair of 2D embeddings of respective pixels corresponding to the pair of vertices in the 3D training mesh; and updating the machine learning model based on a comparison of the first distance and the second distance when read in light of the rest of the limitations in claim 14 and the claims to which claim 14 depends and thus claim 14 contains allowable subject matter. The same reasoning for the indication of allowable subject matter in claim 4 applies here to claim 14. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Kim et al. (U.S. Patent Application Publication No. 2023/0206567) abstract teaches “an initial mesh with vertices that represents a physical environment and a depth map that indicates a geometry of real objects within the physical environment. The AR system is configured to represent the real objects in the physical environment by displacing the vertices of the mesh”; this shows to place virtual element in AR by displacing vertices of a mesh that represents a physical environment. Any inquiry concerning this communication or earlier communications from the examiner should be directed to NAUMAN U AHMAD whose telephone number is (703)756-5306. The examiner can normally be reached Monday - Friday 9:00am - 5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at (571) 272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /N.U.A./Examiner, Art Unit 2611 /KEE M TUNG/Supervisory Patent Examiner, Art Unit 2611
Read full office action

Prosecution Timeline

Feb 03, 2025
Application Filed
Jul 24, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705803
PSEUDO VASCULAR PATTERN GENERATION DEVICE AND METHOD OF GENERATING PSEUDO VASCULAR PATTERN
2y 3m to grant Granted Aug 11, 2026
Patent 12700184
SYSTEM AND METHODS FOR REFINING ROOM SEGMENTS TO IMPROVE AESTHETIC QUALITY FOR END-USER APPLICATIONS
2y 2m to grant Granted Aug 04, 2026
Patent 12682515
METHOD AND DEVICE WITH IMAGE GENERATION BASED ON NEURAL SCENE REPRESENTATION
2y 10m to grant Granted Jul 14, 2026
Patent 12670637
System and Method for Creating a Design Tool Using a Clockwise Fill Rule
2y 5m to grant Granted Jun 30, 2026
Patent 12664616
IMAGE PROCESSING METHOD, MODEL TRAINING METHOD, APPARATUS, MEDIUM AND DEVICE
3y 1m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
77%
Grant Probability
99%
With Interview (+23.3%)
2y 6m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 44 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month