Prosecution Insights
Last updated: August 06, 2026
Application No. 19/052,473

NEURAL RADIANCE FIELDS FOR ORTHOGRAPHIC IMAGERY

Non-Final OA §103§DP
Filed
Feb 13, 2025
Priority
Feb 16, 2024 — provisional 63/554,620 +1 more
Examiner
TSWEI, YU-JANG
Art Unit
Tech Center
Assignee
1000786269 Ontario Inc.
OA Round
1 (Non-Final)
84%
Grant Probability
Favorable
1-2
OA Rounds
9m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
386 granted / 458 resolved
+24.3% vs TC avg
Strong +17% interview lift
Without
With
+17.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 2m
Avg Prosecution
42 currently pending
Career history
502
Total Applications
across all art units

Statute-Specific Performance

§101
6.2%
-33.8% vs TC avg
§103
71.7%
+31.7% vs TC avg
§102
6.4%
-33.6% vs TC avg
§112
7.6%
-32.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 458 resolved cases

Office Action

§103 §DP
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 1-14 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-5 of U.S. Patent No. 12,462,334 B2 (issued from U.S. Application No. 19/198,205). Although the claims at issue are not identical, they are not patentably distinct from each other because they both claim the same subject matters and limitations as explained below. Claims 1, 2, 3, 4, and 5 are determined to be obvious in light of claim 1 of U.S. Patent No. 12,462,334 B2 based on reasons below for having similar limitations. Instant application (19/052,473) claims 1, 2, 3, 4, and 5 U.S. Patent No. 12,462,334 B2 claim 1 (Claim 1) A method comprising: accessing one or more source images depicting a scene from arbitrary points of view; 1. A method comprising: accessing a plurality of source images depicting a scene from multiple points of view; (Claim 1) encoding the one or more source images into one or more corresponding feature maps; (Claim 2) The method of claim 1, wherein encoding the one or more source images involves: for each source image, encoding the source image into a series of multiscale feature maps. encoding at least some of the plurality of source images each into a corresponding series of multiscale feature maps; (Claim 1) receiving an indication of an orthographic view of the scene; and receiving a definition of an orthographic projection of the scene; and (Claim 1) generating an orthographic image of the scene from the orthographic view, wherein the generating is based at least in part on decoding the indication of the orthographic view and is further based at least in part on at least some of the features of the feature maps into which the source images were encoded. generating an orthographic image of the scene corresponding to the orthographic projection, wherein the generating involves: … [the global attention / convolution+upsampling / depth-map / back-projection / local attention pipeline recited below]. (Claim 3) The method of claim 1, wherein decoding the indication of the orthographic view into the orthographic image involves: applying global attention to a set of higher-level features of the scene extracted from the source images; and applying global attention, based on the definition of the orthographic projection, across a set of higher-level features of the series of multiscale feature maps of at least some of the encoded source images, to produce a first set of decoded features; applying a convolutional and upsampling layer to the first set of decoded features to produce a second set of decoded features; (Claim 4) The method of claim 3, wherein applying local attention to the lower-level features of the scene involves: generating a depth map for the scene corresponding to the orthographic view; and generating a depth map for the scene corresponding to the orthographic projection of the scene based on the second set of decoded features; (Claim 4) determining, with reference to the depth map, a limited set of features to be included in a local attention calculation. (Claim 5) The method of claim 4, wherein determining the limited set of features to be included in a local attention calculation involves: for each point on the depth map corresponding to a pixel to be rendered in the orthographic image, back-projecting the point through the orthographic view to determine one or more features to be used to decode the pixel to be rendered. for each point on the depth map corresponding to a pixel that should be rendered in the orthographic image, back-projecting the point through the orthographic projection to determine a set of lower-level features of the series of multiscale feature maps to be used to decode the pixel; and (Claim 3) applying local attention to a set of lower-level features of the scene extracted from the source images. applying local attention across the set of lower-level features to produce a third set of decoded features from which pixel information can be determined. Although the claims at issue are not identical, they are not patentably distinct from each other. Claim 1 of U.S. Patent No. 12,462,334 B2 is a single, narrower species claim that recites every limitation of instant claims 1, 2, 3, 4, and 5 collectively. For example, claim 1 of the patent discloses “accessing a plurality of source images depicting a scene from multiple points of view” — which reads on instant claim 1’s “accessing one or more source images depicting a scene from arbitrary points of view”. Claim 1 of the patent also discloses “encoding at least some of the plurality of source images each into a corresponding series of multiscale feature maps” — which reads on instant claim 1’s “encoding the one or more source images into one or more corresponding feature maps” and also on instant claim 2’s “for each source image, encoding the source image into a series of multiscale feature maps”. While instant claim 1 broadly recites “receiving an indication of an orthographic view of the scene” and “generating an orthographic image … based at least in part on decoding the indication of the orthographic view”, claim 1 of the patent expressly recites “receiving a definition of an orthographic projection of the scene” and “generating an orthographic image of the scene corresponding to the orthographic projection,” wherein the generating involves applying global attention based on that definition across encoded features and producing a decoded result — disclosing the same input definition/indication and the same decoded-features-based generation. Instant claim 3 further requires “applying global attention to a set of higher-level features” and “applying local attention to a set of lower-level features”; both are expressly recited in claim 1 of the patent (“applying global attention … across a set of higher-level features” and “applying local attention across the set of lower-level features”). Instant claim 4 adds “generating a depth map for the scene corresponding to the orthographic view” and “determining, with reference to the depth map, a limited set of features to be included in a local attention calculation” — both expressly recited in claim 1 of the patent (“generating a depth map for the scene corresponding to the orthographic projection” and “back-projecting the point through the orthographic projection to determine a set of lower-level features … to be used to decode the pixel”). Instant claim 5 requires “for each point on the depth map corresponding to a pixel to be rendered in the orthographic image, back-projecting the point through the orthographic view to determine one or more features to be used to decode the pixel to be rendered” — which is recited in claim 1 of the patent verbatim in substance (“for each point on the depth map corresponding to a pixel that should be rendered in the orthographic image, back-projecting the point through the orthographic projection to determine a set of lower-level features … to be used to decode the pixel”). Because claim 1 of the patent recites every limitation of instant claims 1, 2, 3, 4, and 5 (any practice of patent claim 1 necessarily practices each of instant claims 1–5), instant claims 1, 2, 3, 4, and 5 are not patentably distinct from claim 1 of the patent. Claim 6 is determined to be obvious in light of claim 2 of U.S. Patent No. 12,462,334 B2 based on reasons below for having similar limitations. Instant application (19/052,473) claim 6 U.S. Patent No. 12,462,334 B2 claim 2 (Claim 6) The method of claim 1, wherein at least one of the one or more source images provides at least some coverage of the scene from a substantially overhead point of view. 2. The method of claim 1, wherein at least one of the source images provides at least some coverage of the scene from a substantially overhead point of view. Although the claims at issue are not identical, they are not patentably distinct from each other. Claim 2 of the patent recites the same limitation as instant claim 6 essentially verbatim — “at least one of the source images provides at least some coverage of the scene from a substantially overhead point of view.” The only difference is the inconsequential phrase “of the one or more” in the instant claim, which does not change the scope under the broadest reasonable interpretation. Hence instant claim 6 is not patentably distinct from claim 2 of the patent. Claim 7 is determined to be obvious in light of claim 3 of U.S. Patent No. 12,462,334 B2 based on reasons below for having similar limitations. Instant application (19/052,473) claim 7 U.S. Patent No. 12,462,334 B2 claim 3 (Claim 7) The method of claim 1, wherein the indication of the orthographic view of the scene comprises an embedded representation of a set of camera parameters that defines the orthographic view. 3. The method of claim 1, wherein the definition of the orthographic projection of the scene comprises an embedded representation of a set of camera parameters corresponding to the orthographic projection. Although the claims at issue are not identical, they are not patentably distinct from each other. Both claims recite that the input characterization of the target view (“indication of the orthographic view” in the instant claim and “definition of the orthographic projection” in the patent claim) “comprises an embedded representation of a set of camera parameters” that defines/corresponds to the orthographic view/projection. The terms “defines” and “corresponding to” are interchangeable in this context, and “orthographic view” and “orthographic projection” refer to the same target. Hence instant claim 7 is not patentably distinct from claim 3 of the patent. Claim 8 is determined to be obvious in light of claim 4 of U.S. Patent No. 12,462,334 B2 based on reasons below for having similar limitations. Instant application (19/052,473) claim 8 U.S. Patent No. 12,462,334 B2 claim 4 (Claim 8) The method of claim 1, further comprising combining the orthographic image with other orthographic images to generate an orthomosaic. 4. The method of claim 1, further comprising combining the orthographic image with other orthographic images to generate an orthomosaic. Claim 9 is determined to be obvious in light of claim 5 of U.S. Patent No. 12,462,334 B2 based on reasons below for having similar limitations. Instant application (19/052,473) claim 9 U.S. Patent No. 12,462,334 B2 claim 5 (Claim 9) The method of claim 8, wherein combining the orthographic image with other orthographic images to generate the orthomosaic involves: decoding a first orthographic image patch corresponding to a first orthographic view; and 5. The method of claim 4, wherein combining the orthographic image with other orthographic images to generate the orthomosaic involves: decoding a first orthographic image patch corresponding to a first orthographic view; and (Claim 9) decoding a second orthographic image patch corresponding to a second orthographic view, wherein the second image patch at least partly overlaps the first orthographic image patch, and wherein decoding the second orthographic image patch is based at least in part on feature information decoded for the first orthographic image patch, resulting in a set of blended features in at least an area in which the second orthographic image patch overlaps the first orthographic image patch. decoding a second orthographic image patch corresponding to a second orthographic view, wherein the second orthographic image patch at least partly overlaps the first orthographic image patch, and wherein decoding the second orthographic image patch is based at least in part on feature information decoded for the first orthographic image patch, resulting in a set of blended features in at least an area in which the second orthographic image patch overlaps the first orthographic image patch. The two claims are substantively identical, both reciting (i) decoding a first orthographic image patch corresponding to a first orthographic view, (ii) decoding a second orthographic image patch corresponding to a second orthographic view that at least partly overlaps the first patch, and (iii) decoding of the second patch based on feature information decoded for the first patch, resulting in blended features in the overlap area. The only difference is “second image patch” vs. “second orthographic image patch” in one location, which does not alter scope because the antecedent “second orthographic image patch” unambiguously refers to the same patch. Hence instant claim 9 is not patentably distinct from claim 5 of the patent. Claims 10 and 11 are determined to be obvious in light of claims 1, 4, and 5 of U.S. Patent No. 12,462,334 B2 based on reasons below for having similar limitations. Instant application (19/052,473) claims 10 and 11 U.S. Patent No. 12,462,334 B2 claims 1, 4, and 5 (Claim 10) A method comprising: accessing a set of source images covering an area of interest; (Patent claim 1) A method comprising: accessing a plurality of source images depicting a scene from multiple points of view; (Claim 10) for each source image, encoding the source image; (Claim 11) The method of claim 10, wherein: encoding each source image comprises encoding each source image into a series of multiscale feature maps; (Patent claim 1) encoding at least some of the plurality of source images each into a corresponding series of multiscale feature maps; (Claim 10) receiving a set of orthographic views, wherein the set of orthographic views include at least a first orthographic view and a second orthographic view, wherein the first and second orthographic views at least partly overlap one another; and (Patent claim 1) receiving a definition of an orthographic projection of the scene; and (Patent claim 4) The method of claim 1, further comprising combining the orthographic image with other orthographic images to generate an orthomosaic. (Patent claim 5) decoding a first orthographic image patch corresponding to a first orthographic view; and decoding a second orthographic image patch corresponding to a second orthographic view, wherein the second orthographic image patch at least partly overlaps the first orthographic image patch … (Claim 10) for each orthographic view, generating an orthographic image patch corresponding to the orthographic view through a method for depth-guided novel view synthesis; (Patent claim 1) generating an orthographic image of the scene corresponding to the orthographic projection, wherein the generating involves: applying global attention …; applying a convolutional and upsampling layer …; generating a depth map for the scene corresponding to the orthographic projection …; for each point on the depth map … back-projecting the point through the orthographic projection …; and applying local attention … (Claim 10) wherein generating the orthographic image patch for at least the second orthographic view involves incorporating at least some underlying feature information in an overlapping area between the image patch corresponding to the first orthographic view and the image patch corresponding to the second orthographic view. (Claim 11) and incorporating at least some underlying feature information generated for one or more previous image patches comprises incorporating feature information from multiple of such feature maps. (Patent claim 5) … decoding the second orthographic image patch is based at least in part on feature information decoded for the first orthographic image patch, resulting in a set of blended features in at least an area in which the second orthographic image patch overlaps the first orthographic image patch. (Patent claim 1) … global attention … across a set of higher-level features of the series of multiscale feature maps …; … back-projecting the point through the orthographic projection to determine a set of lower-level features of the series of multiscale feature maps … (drawing on multiple multiscale feature maps). Although the claims at issue are not identical, they are not patentably distinct from each other. Patent claim 1 discloses “accessing a plurality of source images depicting a scene from multiple points of view,” which reads on instant claim 10’s “accessing a set of source images covering an area of interest” — “multiple points of view of a scene” corresponds to “covering an area of interest.” Patent claim 1’s “encoding at least some of the plurality of source images each into a corresponding series of multiscale feature maps” reads on both instant claim 10’s “for each source image, encoding the source image” and instant claim 11’s further requirement that “encoding each source image comprises encoding each source image into a series of multiscale feature maps.” Instant claim 10 further recites “receiving a set of orthographic views … at least partly overlap one another” and “for each orthographic view, generating an orthographic image patch corresponding to the orthographic view through a method for depth-guided novel view synthesis.” Patent claim 1 expressly recites the full depth-guided novel view synthesis pipeline (global attention → conv/upsampling → depth map → back-projection → local attention) that produces an orthographic image for a single orthographic projection; patent claim 4 (depending from patent claim 1) requires “combining the orthographic image with other orthographic images to generate an orthomosaic,” thereby producing image patches for multiple orthographic views; and patent claim 5 (depending from patent claim 4) expressly recites decoding a first orthographic image patch and a second orthographic image patch in which the second “at least partly overlaps the first orthographic image patch.” Together, patent claims 1, 4, and 5 disclose generating multiple, mutually-overlapping orthographic image patches through a depth-guided novel-view-synthesis method, as required by instant claim 10. Finally, instant claim 10 requires “incorporating at least some underlying feature information in an overlapping area between” the first and second orthographic image patches, and instant claim 11 further requires that such incorporation use “feature information from multiple of such [multiscale] feature maps.” Patent claim 5 expressly recites that “decoding the second orthographic image patch is based at least in part on feature information decoded for the first orthographic image patch, resulting in a set of blended features in at least an area in which the second orthographic image patch overlaps the first orthographic image patch,” which discloses the incorporation of underlying feature information in the overlapping area. And because patent claim 1 expressly draws on both “higher-level features of the series of multiscale feature maps” and “lower-level features of the series of multiscale feature maps” during the decoding, the feature information incorporated across overlapping patches is drawn from multiple of the multiscale feature maps, satisfying instant claim 11. Therefore, claims 1, 4, and 5 of the patent collectively disclose all limitations of instant claims 10 and 11, and instant claims 10 and 11 are not patentably distinct from claims 1, 4, and 5 of the patent. Claims 12, 13, and 14 are determined to be obvious in light of claim 1 of U.S. Patent No. 12,462,334 B2 based on reasons below for having similar limitations. Instant application (19/052,473) claims 12, 13, and 14 U.S. Patent No. 12,462,334 B2 claim 1 (Claim 12) A method comprising: accessing source images covering an area of interest; 1. A method comprising: accessing a plurality of source images depicting a scene from multiple points of view; (Claim 12) encoding the source images into a set of encoded features; encoding at least some of the plurality of source images each into a corresponding series of multiscale feature maps; (Claim 12) defining an orthographic view; and receiving a definition of an orthographic projection of the scene; and (Claim 12) decoding the orthographic view into an orthographic image depicting at least a portion of the area of interest based at least in part on the encoded features. generating an orthographic image of the scene corresponding to the orthographic projection, wherein the generating involves: applying global attention, based on the definition of the orthographic projection, across a set of higher-level features of the series of multiscale feature maps …, to produce a first set of decoded features; applying a convolutional and upsampling layer … to produce a second set of decoded features; (Claim 13) The method of claim 12, wherein decoding the orthographic view into the orthographic image involves predicting a depth map for at least part of the area of interest and decoding the orthographic image with reference to the encoded features that can be projected to from points on the depth map visible through the orthographic view. generating a depth map for the scene corresponding to the orthographic projection of the scene based on the second set of decoded features; for each point on the depth map corresponding to a pixel that should be rendered in the orthographic image, back-projecting the point through the orthographic projection to determine a set of lower-level features of the series of multiscale feature maps to be used to decode the pixel; (Claim 14) The method of claim 13, wherein the depth map is leveraged in a local attention mechanism that is applied as part of a process for decoding the orthographic view into the orthographic image. and applying local attention across the set of lower-level features to produce a third set of decoded features from which pixel information can be determined. Although the claims at issue are not identical, they are not patentably distinct from each other. Patent claim 1 discloses “accessing a plurality of source images depicting a scene from multiple points of view,” which reads on instant claim 12’s “accessing source images covering an area of interest” — source images “depicting a scene from multiple points of view” cover an area of interest. Patent claim 1’s “encoding at least some of the plurality of source images each into a corresponding series of multiscale feature maps” reads on instant claim 12’s broader “encoding the source images into a set of encoded features” — the recited “multiscale feature maps” are a species of “encoded features.” Patent claim 1’s “receiving a definition of an orthographic projection of the scene” reads on instant claim 12’s “defining an orthographic view” — the patent claim’s receipt of a “definition” inherently involves an orthographic view being defined, and “orthographic projection” and “orthographic view” refer to the same target view. Patent claim 1’s “generating an orthographic image … corresponding to the orthographic projection” using the global attention → conv/upsampling pipeline reads on instant claim 12’s “decoding the orthographic view into an orthographic image … based at least in part on the encoded features.” Instant claim 13 further requires “predicting a depth map for at least part of the area of interest” and “decoding the orthographic image with reference to the encoded features that can be projected to from points on the depth map visible through the orthographic view.” Patent claim 1 expressly recites “generating a depth map for the scene corresponding to the orthographic projection” and, “for each point on the depth map …, back-projecting the point through the orthographic projection to determine a set of lower-level features … to be used to decode the pixel” — “back-projecting” the point through the orthographic projection is the same geometric operation as identifying “encoded features that can be projected to from points on the depth map visible through the orthographic view.” Instant claim 14 further requires that “the depth map is leveraged in a local attention mechanism that is applied as part of a process for decoding the orthographic view into the orthographic image.” Patent claim 1 expressly satisfies this: the depth map is used to determine, by back-projection, the set of lower-level features, and “applying local attention across the set of lower-level features” is then performed as part of decoding the orthographic image. Therefore, patent claim 1 discloses every limitation of instant claims 12, 13, and 14, and instant claims 12, 13, and 14 are not patentably distinct from claim 1 of the patent. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-2, 6-8, 10-14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sajjadi et al. (US 20240096001 A1, hereinafter Sajjadi) in view of Pylvaenaeinen et al. ( US 20180276875 A1, hereinafter Pylvaenaeinen). Regarding Claim 1, Sajjadi teaches a method comprising: (Sajjadi, "computer-implemented method to generate novel views of a scene more efficiently", [0009]; Sajjadi expressly teaches a method.); accessing one or more source images depicting a scene from arbitrary points of view (Sajjadi, FIG. 1; Paragraph [0061]; "a computing system can obtain one or more input images 12 that depict a scene. In some implementations, the one or more input images 12 can be a plurality of input images respectively captured at a plurality of different poses <read on arbitrary points of view > relative to the scene", [0062]; "the one or more input images 12 are unposed images that have an unspecified pose relative to the scene"); encoding the one or more source images into one or more corresponding feature maps (Sajjadi, FIG. 2; Paragraph [0063], "processing, by the computing system, the one or more input images 12 with a convolutional neural network 16 to respectively generate the one or more image embeddings 14"; [0073]; "given RGB inputs that are optionally posed, a CNN extracts patch features <read on feature maps> , onto which learned embeddings for 2D position and camera ID are added"; [0065], "process the one or more image embeddings 14 with a machine-learned encoder model 18 to generate a scene embedding 20"); receiving an indication of an [[ orthographic ]] view of the scene; and (Sajjadi, Paragraph [0066], "The computing system can obtain ray data 22 descriptive of one or more ray castings for a predicted image of the scene. As examples, the ray data 22 can include one or more sets of 5- or 6-D ray information that respectively correspond to one or more pixels in the predicted image"; generating an [[ orthographic ]] image of the scene from the [[ orthographic ]] view, wherein the generating is based at least in part on decoding the indication of the [[ orthographic ]] view and is further based at least in part on at least some of the features of the feature maps into which the source images were encoded (Sajjadi, FIG. 1; Paragraph [0067], "process the scene embedding 20 and the ray data 22 with a machine-learned decoder model 24 to generate synthesized image data for the one or more ray castings for the predicted image 26 of the scene"; [0073], "The decoder attends into the Scene Representation using a given 6D ray pose, leading to the final RGB output"). But Sajjadi does not explicitly disclose [[ receiving an indication of an ]] orthographic [[ view of the scene ]]; [[ generating an ]] orthographic [[ image of the scene from the ]] orthographic [[ view, wherein the generating is based at least in part on decoding the indication of the ]] orthographic [[ view and is further based at least in part on at least some of the features of the feature maps into which the source images were encoded ]] However, Pylvaenaeinen teaches receiving an indication of an orthographic view of the scene; and (Pylvaenaeinen, FIG. 4; Paragraph [0030], "preprocessed into selected (specific) orthographic views such as top-down and oblique"; [0053] "Orthographic views do not have real perspective, so parallel lines remain parallel. Therefore, the orthographic view along any compass direction produces a rectangular grid) generating an orthographic image of the scene from the orthographic view, wherein the generating is based at least in part on decoding the indication of the orthographic view and is further based at least in part on at least some of the features of the feature maps into which the source images were encoded. (Pylvaenaeinen, Abstract, "rendered to the desired orthographic projection", [0030]; "preprocessed into selected (specific) orthographic views such as top-down and oblique"; Pylvaenaeinen and Sajjadi are analogous since both of them are dealing with operate within the same neural scene rendering / novel view synthesis framework. Sajjadi teaches receiving an indication of a target view (ray data / camera ray) but does not specify it is an orthographic view and decoder-based generation of an image from the ray indication using features of the scene embedding (i.e., the feature maps). Pylvaenaeinen teaches defining orthographic target views (top-down, oblique) and the rendered target view is an orthographic projection. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate orthographic view taught by Pylvaenaeinen into modified invention of Sajjadisuch that system will be able to enable rendering of aerial/orthographic projections for mapping and visualization. The motivation is to generating an orthographic image by decoding the orthographic-view indication against the encoded feature maps as a view-type definition without requiring any fundamental redesign, making the combination straightforward and predictable. Regarding Claim 10, Sajjadi teaches a method comprising: (Sajjadi, [0009]; Sajjadi expressly teaches a method.) accessing a set of source images covering an area of interest (Sajjadi, Paragraph [0062], "a computing system can obtain one or more input images 12 that depict a scene"; "the one or more input images 12 can be a plurality of input images respectively captured at a plurality of different poses relative to the scene"); for each source image, encoding the source image (Sajjadi, FIG. 2; Paragraph [0073], "a CNN extracts patch features"; [0065], "process the one or more image embeddings 14 with a machine-learned encoder model 18 to generate a scene embedding 20"; Sajjadi teaches encoding each source image via the CNN and encoder transformer.) receiving a set of [[ orthographic ]] views, wherein the set of orthographic views include at least a first orthographic view and a second [[ orthographic ]] view, wherein the first and second [[ orthographic ]] views at least partly overlap one another; and (Sajjadi, Paragraph [0066], "The computing system can obtain ray data 22 descriptive of one or more ray castings for a predicted image of the scene) for each [[ orthographic ]] view, generating an [[ orthographic ]] image patch corresponding to the [[ orthographic ]] view through a method for depth-guided novel view synthesis (Sajjadi, Paragraph [0067], "process the scene embedding 20 and the ray data 22 with a machine-learned decoder model 24 to generate synthesized image data"; (teaches novel-view synthesis from encoded features); [[ wherein generating the orthographic image patch for at least the second orthographic view involves incorporating at least some underlying feature information in an overlapping area between the image patch corresponding to the first orthographic view and the image patch corresponding to the second orthographic view ]] But Sajjadi does not explicitly disclose orthographic image and orthographic image vies; wherein generating the orthographic image patch for at least the second orthographic view involves incorporating at least some underlying feature information in an overlapping area between the image patch corresponding to the first orthographic view and the image patch corresponding to the second orthographic view. However, Pylvaenaeinen teaches accessing a set of source images covering an area of interest (Pylvaenaeinen, Paragraph [0026], "The input data can comprise hundreds of miles of street-side (also "street-level") data comprising a full 360-degree panorama data captured at regular distance intervals"); receiving a set of orthographic views, wherein the set of orthographic views include at least a first orthographic view and a second orthographic view, wherein the first and second orthographic views at least partly overlap one another; and (Pylvaenaeinen, FIG. 5; Paragraph [0030], "preprocessed into selected (specific) orthographic views such as top-down and oblique"; [0009], "Tiles in the same geographical area, yet rendered independently from different data"; [0040], "the same tile (in terms of area on the map, i.e., the same quad key) can be produced by multiple nodes",;); for each orthographic view, generating an orthographic image patch corresponding to the orthographic view through a method for depth-guided novel view synthesis (Pylvaenaeinen, Paragraph [0008], "A depth map can be generated by rendering the 3D surface onto a z-buffer. The z-buffer is used to determine if pixels originating from some other triangle (polygon) are to be processed to overwrite a previous rendering"; [0056], "both a color image as well as a depth map are maintained for every map pixel"; Abstract, "rendered to the desired orthographic projection";) Pylvaenaeinen and Sajjadi are analogous since both of them are dealing with operate within the same neural scene rendering / novel view synthesis framework. Sajjadi teaches receiving target view indications for multiple predicted images and accessing a set of source images also provide novel view synthesis but is geometry-free. Pylvaenaeinen reinforces that these images cover a geographical area of interest and provide a multiple orthographic views/tiles, including overlapping coverage of the same area and further provide depth-map–guided orthographic rendering via z-buffer. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate square tiled orientations depth-guided rendering and orthographic views taught by Pylvaenaeinen into modified invention of Sajjadi such that system will be able to resolve occlusions and select correct surface points to provide best viewing result. Regarding Claim 2, the combination of Sajjadi and Pylvaenaeinen teaches the invention in Claim 1. The combination further teaches wherein encoding the one or more source images involves: for each source image, encoding the source image into a series of multiscale feature maps (Sajjadi, FIG. 2, Paragraph [0073], "a CNN extracts patch features, onto which learned embeddings for 2D position and camera ID are added"; [0063], "processing, by the computing system, the one or more input images 12 with a convolutional neural network 16 to respectively generate the one or more image embeddings 14"; it is noted per-image CNN encoding into patch feature maps. By using CNN architectures producing a hierarchy of multiscale feature maps are a well-known design choice). Regarding Claim 6, the combination of Sajjadi and Pylvaenaeinen teaches the invention in Claim 1. The combination further teaches wherein at least one of the one or more source images provides at least some coverage of the scene [[ from a substantially overhead point of view ]] (Sajjadi, Paragraph [0062], "the one or more input images 12 can be a plurality of input images respectively captured at a plurality of different poses relative to the scene"; But Sajjadi does not explicitly disclose the scene from a substantially overhead point of view. However, Pylvaenaeinen teaches wherein at least one of the one or more source images provides at least some coverage of the scene from a substantially overhead point of view (Pylvaenaeinen, Paragraph [0022], "An aerial view is better suited for quickly exploring an area"; [0030], "This also occurs for aerial imaging, which has perspective, and which is converted into an orthographic projection, tiled"); Pylvaenaeinen and Sajjadi are analogous since both of them are dealing with operate within the same neural scene rendering / novel view synthesis framework. Sajjadi teaches receiving target view indications for multiple predicted images and accessing a set of source images with arbitrary input poses. Pylvaenaeinen teaches contemplates aerial (overhead) source imagery being incorporated into the mapping pipeline. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate teaching by Pylvaenaeinen into modified invention of Sajjadi such that system will be able to enrich the scene representation with overhead coverage suited to Sajjadi's stated mapping/visualization use case. Regarding Claim 7, the combination of Sajjadi and Pylvaenaeinen teaches the invention in Claim 1. The combination further teaches wherein the indication of the [[ orthographic ]] view of the scene comprises an embedded representation of a set of camera parameters that defines the [[ orthographic ]] view (Sajjadi, Paragraph [0066], "the ray data 22 can include one or more sets of 5- or 6-D ray information that respectively correspond to one or more pixels in the predicted image"; [0069],"generating query data elements from the ray data"; [0073], "The decoder attends into the Scene Representation using a given 6D ray pose"; it is noted that the indication of the target view is supplied as 5-/6-D ray (camera-parameter) data that is processed as a query embedding by the decoder — i.e., an embedded representation of camera parameters defining the target view.). But Sajjadi does not explicitly disclose orthographic view of the scene. However, Pylvaenaeinen teaches an orthographic view of the scene (Pylvaenaeinen, FIG. 4; Paragraph [0030], "preprocessed into selected (specific) orthographic views such as top-down and oblique"; , [0053] "Orthographic views do not have real perspective, so parallel lines remain parallel. Therefore, the orthographic view along any compass direction produces a rectangular grid). As explained in rejection of claim 1, the obviousness for combining of orthographic view of Pylvaenaeinen into Sajjadi is provided above. Regarding Claim 8, the combination of Sajjadi and Pylvaenaeinen teaches the invention in Claim 1. The combination further teaches further comprising combining the orthographic image with other orthographic images to generate an orthomosaic (Pylvaenaeinen, FIG. 5; Paragraph [0030], "preprocessed into selected (specific) orthographic views such as top-down and oblique"; [0057], "Multiple partial renderings are then merged into a single (partial) rendering"; [0029], "The pyramid map is created by taking all the scanned data points... and then projecting these into a global bird's-eye view"). As explained in rejection of claim 1, the obviousness for combining of orthographic view of Pylvaenaeinen into Sajjadi is provided above. Regarding Claim 11, the combination of Sajjadi and Pylvaenaeinen teaches the invention in Claim 10. The combination further teaches wherein: encoding each source image comprises encoding each source image into a series of multiscale feature maps; and incorporating at least some underlying feature information generated for one or more previous image patches comprises incorporating feature information from multiple of such feature maps (Sajjadi, Paragraph [0073], "a CNN extracts patch features"; [0065], "process the one or more image embeddings 14 with a machine-learned encoder model 18 to generate a scene embedding 20"; it is noted multiscale aspect is a well-known CNN technique); Sajjadi does not explicitly disclose but Pylvaenaeinen teaches and incorporating at least some underlying feature information generated for one or more previous image patches comprises incorporating feature information from multiple of such feature maps (Pylvaenaeinen, Paragraph [0045], "merge multiple renderings of image tiles of a same area, as part of the distributed processing, using the associated depth maps"; [0057], "Multiple partial renderings are then merged into a single (partial) rendering by combining the color and height raster"; it is noted cross-patch incorporation from multiple feature levels is a predictable extension from merge across multiple LoD layers). Pylvaenaeinen and Sajjadi are analogous since both of them are dealing with operate within the same neural scene rendering / novel view synthesis framework. Sajjadi teaches CNN-based per-image encoding. Pylvaenaeinen teaches prior-patch information in overlapping regions. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate teaching by Pylvaenaeinen into modified invention of Sajjadi such that system will be able to capturing both fine and coarse cross-patch coherence in overlap regions Regarding Claim 12, Sajjadi teaches a method comprising: (Sajjadi, "computer-implemented method to generate novel views of a scene more efficiently", [0009]; Sajjadi expressly teaches a method.); accessing source images covering an area of interest (Sajjadi, FIG. 1; Paragraph [0061]; "a computing system can obtain one or more input images 12 that depict a scene. In some implementations, the one or more input images 12 can be a plurality of input images respectively captured at a plurality of different poses <read on arbitrary points of view > relative to the scene", [0062]; "the one or more input images 12 are unposed images that have an unspecified pose relative to the scene"); encoding the source images into a set of encoded features (Sajjadi, FIG. 2; Paragraph [0063], "processing, by the computing system, the one or more input images 12 with a convolutional neural network 16 to respectively generate the one or more image embeddings 14"; [0073]; "given RGB inputs that are optionally posed, a CNN extracts patch features <read on feature maps> , onto which learned embeddings for 2D position and camera ID are added"; [0065], "process the one or more image embeddings 14 with a machine-learned encoder model 18 to generate a scene embedding 20"); defining an [[ orthographic ]] view (Sajjadi, Paragraph [0009], “One example aspect is directed to a computer-implemented method to generate novel views of a scene”); and decoding the [[ orthographic ]] view into an [[ orthographic ]] image depicting at least a portion of the area of interest based at least in part on the encoded features (Sajjadi, FIG. 1; Paragraph [0067], "process the scene embedding 20 and the ray data 22 with a machine-learned decoder model 24 to generate synthesized image data for the one or more ray castings for the predicted image 26 of the scene"; [0073], "The decoder attends into the Scene Representation using a given 6D ray pose, leading to the final RGB output"). But Sajjadi does not explicitly disclose defining orthographic view into an orthographic image. However, Pylvaenaeinen teaches defining an orthographic view and (Pylvaenaeinen, FIG. 4; Paragraph [0030], "preprocessed into selected (specific) orthographic views such as top-down and oblique"; , [0053] "Orthographic views do not have real perspective, so parallel lines remain parallel. Therefore, the orthographic view along any compass direction produces a rectangular grid); and decoding the orthographic view into an orthographic image depicting at least a portion of the area of interest based at least in part on the encoded features (Pylvaenaeinen, Abstract, "rendered to the desired orthographic projection", [0030]; "preprocessed into selected (specific) orthographic views such as top-down and oblique"). Pylvaenaeinen and Sajjadi are analogous since both of them are dealing with operate within the same neural scene rendering / novel view synthesis framework. Sajjadi teaches receiving an indication of a target view (ray data / camera ray) but does not specify it is an orthographic view and decoder-based generation of an image from the ray indication using features of the scene embedding (i.e., the feature maps). Pylvaenaeinen teaches defining orthographic target views (top-down, oblique) and the rendered target view is an orthographic projection. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate orthographic view taught by Pylvaenaeinen into modified invention of Sajjadisuch that system will be able to enable rendering of aerial/orthographic projections for mapping and visualization. The motivation is to generating an orthographic image by decoding the orthographic-view indication against the encoded feature maps as a view-type definition without requiring any fundamental redesign, making the combination straightforward and predictable. Regarding Claim 13, the combination of Sajjadi and Pylvaenaeinen teaches the invention in Claim 12. The combination further teaches wherein decoding the [[ orthographic ]] view into the [[ orthographic ]] image involves predicting a depth map for at least part of the area of interest and decoding the [[orthographic]] image with reference to the encoded features that can be projected to from points on the depth map visible through the [[orthographic]] view (Sajjadi, [0054]; "Moreover, while example implementations are discussed with respect to generation of synthetic RGB imagery, the learned latent scene representation also encodes sufficient information for performing semantic segmentation and other image or scene analysis tasks, even though it was trained only for novel view synthesis."; [0094]; "As another example, the image processing task can be depth estimation, where the image processing output defines, for each pixel in the one or more images, a respective depth value."; [0009]; "processing, by the computing system, the scene embedding and the ray data with a machine-learned decoder model to generate synthesized image data for the one or more ray castings for the predicted image of the scene;"; [0015]; "generating key and value data elements from the scene embedding; generating query data elements from the ray data; and processing, by the computing system, the key, value, and query data elements with the machine-learned decoder model 24 to generate the synthesized image data for the one or more ray castings." [0069], “the machine-learned decoder model 24 can include a self-attention model; and processing, by the computing system, the scene embedding 20 and the ray data 22 with the machine-learned decoder model”; [0069]; Sajjadi teaches decoding a view (via ray data) using encoded features (scene embedding/key/value). Sajjadi does not explicitly disclose the wherein decoding the orthographic view into the orthographic image. However, Pylvaenaeinen teaches involves predicting a depth map for at least part of the area of interest and decoding the orthographic image with reference to the encoded features that can be projected to from points on the depth map visible through the orthographic view (Pylvaenaeinen, [0008]; " The panorama data is projected onto this 3D triangular surface, and then rendered to the desired orthographic projection. A depth map can be generated by rendering the 3D surface onto a z-buffer. The z-buffer is used to determine if pixels originating from some other triangle (polygon) are to be processed to overwrite a previous rendering"; [0055]; "For the final rendering to appear correct, pixels closer to the virtual camera... occlude pixels further away... To recover the highest point, every point is rendered and the elevation of the source point (currently being rendered) is tracked for each map pixel. In other words, both a color image as well as a depth map are maintained for every map pixel (x,y)."; [0060]; "For continuous scanning systems that do not record color, to assign a color to a scanned point, the point is projected back to one of the panorama images that captured the scene."); As explained in rejection of claim 12, the obviousness for combining of orthographic view of Pylvaenaeinen into Sajjadi is provided above. Regarding Claim 14, the combination of Sajjadi and Pylvaenaeinen teaches the invention in Claim 13. The combination further teaches wherein the depth map is leveraged in a local attention mechanism that is applied as part of a process for decoding the [[ orthographic ]] view into the [[ orthographic ]] image (Sajjadi, [0069]; "In some implementations, the machine-learned decoder model 24 can include a self-attention model; and processing, by the computing system, the scene embedding 20 and the ray data 22 with the machine-learned decoder model 24 can includes: generating key and value data elements from the scene embedding 20; generating query data elements from the ray data 22; and processing, by the computing system, the key, value, and query data elements with the machine-learned decoder model 24 to generate the synthesized image data for the one or more ray castings."; [0106]; " . For example, each ray can be used to determine a pixel the slot that had the highest weight in the Mixing Block (and each slot gets a different color for the visualization Block (and each slot gets a different color for the visualization "). But Sajjadi does not explicitly disclose orthographic view into the orthographic image. However, Pylvaenaeinen teaches wherein the depth map is leveraged in a local attention mechanism that is applied as part of a process for decoding the orthographic view into the orthographic image (Pylvaenaeinen, [0008]; "A depth map can be generated by rendering the 3D surface onto a z-buffer. The z-buffer is used to determine if pixels originating from some other triangle (polygon) are to be processed to overwrite a previous rendering.", [0008]; Pylvaenaeinen, [0055]; "For the final rendering to appear correct, pixels closer to the virtual camera... occlude pixels further away... To recover the highest point, every point is rendered and the elevation of the source point (currently being rendered) is tracked for each map pixel (x,y)."). As explained in rejection of claim 13, the obviousness for combining of orthographic view of Pylvaenaeinen into Sajjadi is provided above. Claim(s) 3 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sajjadi et al. (US 20240096001 A1, hereinafter Sajjadi) in view of Pylvaenaeinen et al. ( US 20180276875 A1, hereinafter Pylvaenaeinen) as applied to Claim 1 above and further in view of Wang et al. (US 20250173958 A1, hereinafter Wang). Regarding Claim 3, the combination of Sajjadi and Pylvaenaeinen teaches the invention in Claim 1. The combination further teaches decoding the indication of the orthographic view into the orthographic image involves (Sajjadi, Paragraph [0067], "process the scene embedding 20 and the ray data 22 with a machine-learned decoder model 24 to generate synthesized image data for the one or more ray castings for the predicted image 26 of the scene"; Sajjadi, Paragraph [0073], "The decoder attends into the Scene Representation using a given 6D ray pose, leading to the final RGB output" <read on decoding the indication of the view via attention>). Sajjadi does not explicitly disclose but Pylvaenaeinen teaches orthographic view into the orthographic image (2, FIG. 4; Paragraph [0030], "preprocessed into selected (specific) orthographic views such as top-down and oblique"; , [0053] "Orthographic views do not have real perspective, so parallel lines remain parallel. Therefore, the orthographic view along any compass direction produces a rectangular grid). As explained in rejection of claim 1, the obviousness for combining of orthographic view of Pylvaenaeinen into Sajjadi is provided above. But the combination does not explicitly disclose applying global attention to a set of higher-level features of the scene extracted from the source images; and applying local attention to a set of lower-level features of the scene extracted from the source images. However, Wang teaches applying global attention to a set of higher-level features of the scene extracted from the source images (Wang, Paragraph [0032], "The global controller 202 b may be configured to enable the first sub-model 102 to absorb global structural information <read on higher-level features> from the input image 101"; Wang, Paragraph [0032], "A control of the global controller may be implemented by inputting the vector from the global controller 202 b into the cross-attention layers <read on applying global attention> of the diffusion block 204"; Wang, Paragraph [0036], "The global controller 202 b may integrate global embedding 504 into the first sub-model 102. The global embedding 504 may comprise a 1024-dimension vector with a token length of 4 (denoted as fg)"); and applying local attention to a set of lower-level features of the scene extracted from the source images (Wang, Paragraph [0031], "The local controller 202 a may be configured to enable the first sub-model 102 to capture detailed structural information <read on lower-level features> from the input image 101"; Wang, Paragraph [0031], "A control of the local controller 202 a may be implemented by inputting the balanced local features into cross-attention layers <read on applying local attention> of the diffusion block 204"; Wang, Paragraph [0037], "The hidden features 502 from the CLIP encoder may be resampled by a resampling component 508 of the local controller 202 a. The hidden features may be resampled before global pooling. The local controller 202 a may generate balanced local features"). Wang and Sajjadi are analogous since both deal with neural-network-based view/image synthesis that decodes target image content from features extracted by encoders of input images using attention. Sajjadi provides a way of decoding scene features using a self-attention decoder against a single attention pathway. Wang provides a way of hierarchically conditioning the decoder using two distinct cross-attention pathways which is a global controller that injects coarse/global structural information (higher-level features) and a local controller that injects detailed/local structural information (lower-level features) into the diffusion decoder's cross-attention layers. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate the global controller and local controller cross-attention scheme taught by Wang into the modified invention of Sajjadi such that the decoder applies global attention to higher-level features and local attention to lower-level features when decoding the orthographic view indication. Claim(s) 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sajjadi et al. (US 20240096001 A1, hereinafter Sajjadi) in view of Pylvaenaeinen et al. ( US 20180276875 A1, hereinafter Pylvaenaeinen), further in view of Wang et al. (US 20250173958 A1, hereinafter Wang) as applied to Claim 3 above, further in view of Shrivastava (US 11107228 B1). Regarding Claim 4, the combination of Sajjadi, Pylvaenaeinen and Wang teaches the invention in Claim 3. The combination further teaches orthographic view (Pylvaenaeinen, FIG. 4; Paragraph [0030], "preprocessed into selected (specific) orthographic views such as top-down and oblique"; , [0053] "Orthographic views do not have real perspective, so parallel lines remain parallel. Therefore, the orthographic view along any compass direction produces a rectangular grid). As explained in rejection of claim 1, the obviousness for combining of orthographic view of Pylvaenaeinen into Sajjadi is provided above. The combination does not explicitly disclose but Shrivastava teaches generating a depth map for the scene corresponding to the orthographic view (Shrivastava, Column 2, Line 27-30, "generating, via the deep neural network, a depth map corresponding to the image having the first perspective"; Line 33-35, "generating a depth map corresponding to the image having the second perspective <read on a depth map corresponding to the orthographic view>"; Shrivastava, 10, Line 62-65, "The depth map 418, the reshaped rotation matrix, and the translation matrix T are provided to a transformed depth map decoder 606, and the transformed depth map decoder 606 generates the depth map 608 corresponding to an image having the second perspective"); and determining, with reference to the depth map, a limited set of features to be included in a local attention calculation (Shrivastava, Column 8, Line 31-32, "The computer 120 uses the generated depth map to generate a point cloud"; Line 62-65, "Using the generated 3D coordinates for the point cloud (x,y,z), the computer 120 can use a camera pose transformation matrix to transform the point cloud (x,y,z) coordinates corresponding to the first perspective to point cloud (x′, y′, z′) coordinates corresponding to a second perspective"; Column 9, Line 14-20, "Using the generated point cloud (x′, y′, z′) coordinates, each point can then be projected onto a new depth map corresponding to the image having the second perspective" <read on the depth map narrows/limits which features (point-cloud points actually projected to the target view) are considered in subsequent decoding>). Shrivastava and Sajjadi are analogous since both deal with neural-network-based novel-view image synthesis that uses scene geometry to drive generation of an image from a different viewpoint. Sajjadi provide a way of decoding a target view from encoded source-image features using attention. Shrivastava provides a way of generating, via a deep neural network, a depth map corresponding to the target (second) perspective and using that depth map (via a derived point cloud) to determine which geometric points/features actually correspond to the target view. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate the target-view depth map generation and depth-driven feature/point selection taught by Shrivastava into the modified invention of Sajjadi such that a depth map for the orthographic view is generated and used to confine the local attention calculation to a limited set of features identified by the depth map. The motivation is to reduce computational burden and improve geometric accuracy the generator learns the relationship between the depth map, the 2D projected RGB image, and the camera intrinsic matrix, and to support per-pixel reasoning needed for novel-view rendering. Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sajjadi et al. (US 20240096001 A1, hereinafter Sajjadi) in view of Pylvaenaeinen et al. ( US 20180276875 A1, hereinafter Pylvaenaeinen), further in view of Wang et al. (US 20250173958 A1, hereinafter Wang) and Shrivastava (US 11107228 B1) as applied to Claim 4 above and further in view of Beretel et al. (US 20220139036 A1, hereinafter Beretel). Regarding Claim 5, the combination of Sajjadi, Pylvaenaeinen, Wang and Shrivastava teaches the invention in Claim 4 of determining the limited set of features to be included in a local attention calculation. The combination does not explicitly disclose but Beretel teaches for each point on the depth map corresponding to a pixel to be rendered in the orthographic image, back-projecting the point through the orthographic view to determine one or more features to be used to decode the pixel to be rendered (Beretel, Paragraph [0101], "for a given posed input view, the system may create a dense triangulated mesh assigning a vertex to the center of each pixel and back-project its depth to world space" <read on for each point on the depth map corresponding to a pixel to be rendered…back-projecting the point through the (orthographic) view>; Beretel, Paragraph [0097], "Depth maps may be computed from MPIs using a weighted summation of the depth of each layer, weighted by the MPI’s per-pixel a value"; Beretel, Paragraph [0100], "the system may instead use the depth maps extracted from the MPIs to create a per-view geometry and render individual contributions from the neighboring input views" <read on the back-projected geometry is what selects which neighboring-view features contribute to (i.e., decode) the pixel>; Beretel, Paragraph [0011], "the one or more position guides may include a respective plurality of position values for a designated one of the input images. Each of the plurality of position values may identify a position in the three-dimensional coordinate space associated with a respective pixel position in the designated input image" <read on features (position guides) that are used to decode a pixel are selected per-pixel via back-projected 3D positions>). Beretel and Sajjadi are analogous since both deal with novel-view image synthesis driven by per-pixel depth/geometry that maps an output pixel to its supporting source-image content. Sajjadi provide a way of using a target-view depth map to constrain feature selection for a target-view decoder. Beretel provides a way of, per pixel of the target view, back-projecting the per-pixel depth into world space (creating a vertex at the center of each pixel) and using that per-pixel back-projected geometry to retrieve/blend contributions from neighboring source views. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate the per-pixel depth back-projection of Beretel into the modified invention of Sajjadi such that, for each point on the depth map corresponding to a pixel to be rendered in the orthographic image, the point is back-projected through the orthographic view to determine the source-image features used to decode that pixel. Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sajjadi et al. (US 20240096001 A1, hereinafter Sajjadi) in view of Pylvaenaeinen et al. ( US 20180276875 A1, hereinafter Pylvaenaeinen) as applied to Claim 8 above and further in view of He et al. (US 20210312698 A1, hereinafter He). Regarding Claim 9, the combination of Sajjadi and Pylvaenaeinen teaches the invention in Claim 8. The combination further teaches combining the orthographic image with other orthographic images to generate the orthomosaic involves: decoding a first orthographic image patch corresponding to a first orthographic view; and decoding a second orthographic image patch corresponding to a second orthographic view (Sajjadi, Paragraph [0067], "process the scene embedding 20 and the ray data 22 with a machine-learned decoder model 24 to generate synthesized image data for the one or more ray castings for the predicted image 26 of the scene"; Sajjadi, Paragraph [0066], "The computing system can obtain ray data 22 descriptive of one or more ray castings for a predicted image of the scene"; Sajjadi teaches decoding multiple predicted images from rays/views) — combined with Pylvaenaeinen, which teaches the orthographic nature of those views (Pylvaenaeinen, Paragraph [0030], "preprocessed into selected (specific) orthographic views such as top-down and oblique"; Pylvaenaeinen, Paragraph [0040], "the same tile (in terms of area on the map, i.e., the same quad key) can be produced by multiple nodes" <read on multiple overlapping orthographic patches>). But the combination does not explicitly disclose wherein the second image patch at least partly overlaps the first orthographic image patch, and wherein decoding the second orthographic image patch is based at least in part on feature information decoded for the first orthographic image patch, resulting in a set of blended features in at least an area in which the second orthographic image patch overlaps the first orthographic image patch. However, He teaches wherein the second image patch at least partly overlaps the first orthographic image patch (He, Paragraph [0053], "the term “overlap” can refer to border portions of multiple images that include similar visual features. For instance, an overlap can include a border portion of a first image patch that is similar to a boarder portion of a second image patch"; He, Paragraph [0078], "the novel-view synthesis system 106 can subdivide each source image (Si) into image patches {Pi n}n=1 N via a sliding window with overlaps"; He, Paragraph [0112], "the novel-view synthesis system 106 can divide the lower-dimension frustum feature (e.g., {hn}) into frustum feature patches (e.g., {hn}n=1 N) by utilizing a sliding window approach (with overlaps) along the width and height of the lower-dimension frustum feature"); and wherein decoding the second orthographic image patch is based at least in part on feature information decoded for the first orthographic image patch, resulting in a set of blended features in at least an area in which the second orthographic image patch overlaps the first orthographic image patch (He, Paragraph [0113], "the novel-view synthesis system 106 utilizes a 2D U-Net 724 on the frustum feature patches {hn} to render image patches {{circumflex over (P)}n}. Indeed, the novel-view synthesis system 106 can blend the rendered image patches {{circumflex over (P)}n} to render a 2D view 728 of the object from the target viewpoint" <read on blended features in overlapping area>; He, Paragraph [0113], "the novel-view synthesis system 106 can blend (or composite) all N rendered patches {{circumflex over (P)}n}n=1 N into a target image raster. Furthermore, the novel-view synthesis system 106 can crop overlapped regions of the composited patches {{circumflex over (P)}n}n=1 N to reduce seam artifacts"; He, Paragraph [0052], "the term 'neural renderer' refers to a machine learning based renderer that decodes feature representations (e.g., frustum features) into images" <read on decoding patches based on feature representations that are blended in the overlap region>). He and Sajjadi are analogous since both deal with neural-network-based novel-view image synthesis that decodes image content patch-by-patch from feature representations and combines those patches into a larger output image. Sajjadi provide a way of decoding predicted images from encoded scene features for ray-cast views (including orthographic views per Pylvaenaeinen). He provides a way of (i) subdividing source images and frustum-features into patches via a sliding window with overlaps so that adjacent patches share border regions, and (ii) blending/compositing the neurally-rendered overlapping patches into a single target raster so that the overlapping regions are blended rather than seam-discontinuous. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate the overlapping-patch sampling and feature-level blending of overlapping patches taught by He into the modified invention of Sajjadi such that, when generating an orthomosaic from multiple orthographic patches, decoding the second orthographic image patch incorporates feature information decoded for the first orthographic image patch in the overlap region, producing a set of blended features there. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 20220058859 A1 UV MAPPING ON 3D OBJECTS WITH THE USE OF ARTIFICIAL INTELLIGENCE US 20210166351 A1 SYSTEMS AND METHODS FOR DETECTING AND CORRECTING ORIENTATION OF A MEDICAL IMAGE US 20210158561 A1 IMAGE VOLUME FOR OBJECT POSE ESTIMATION US 20210150671 A1 SYSTEM, METHOD AND COMPUTER-ACCESSIBLE MEDIUM FOR THE REDUCTION OF THE DOSAGE OF GD-BASED CONTRAST AGENT IN MAGNETIC RESONANCE IMAGING Any inquiry concerning this communication or earlier communications from the examiner should be directed to YUJANG TSWEI whose telephone number is (571)272-6669. The examiner can normally be reached 8:30am-5:30pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached on (571) 272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YuJang Tswei/Primary Examiner, Art Unit 2614
Read full office action

Prosecution Timeline

Feb 13, 2025
Application Filed
Jul 15, 2026
Non-Final Rejection mailed — §103, §DP (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12675993
AUGMENTED, VIRTUAL AND MIXED-REALITY CONTENT SELECTION & DISPLAY FOR BANK NOTE
4y 4m to grant Granted Jul 07, 2026
Patent 12670628
COMPOSITIONAL IMAGE GENERATION AND MANIPULATION
2y 9m to grant Granted Jun 30, 2026
Patent 12657909
AUGMENTED, VIRTUAL AND MIXED-REALITY CONTENT SELECTION & DISPLAY FOR BILLBOARDS
4y 3m to grant Granted Jun 16, 2026
Patent 12629233
ALIGNER FINISHING LINE TRIMMING AND ALIGNERS HAVING TRIMMED FINISHING LINES
2y 2m to grant Granted May 19, 2026
Patent 12579805
AUGMENTED, VIRTUAL AND MIXED-REALITY CONTENT SELECTION & DISPLAY FOR TRAVEL
4y 0m to grant Granted Mar 17, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
84%
Grant Probability
99%
With Interview (+17.2%)
2y 2m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 458 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month