DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 04/03/2026 has been entered.
Response to Amendment
1 This action is in response to the amendment filed on 04/03/2026. Claims 1, 7-8, and 14-15 have been amended. Claims 1-3, 5-10, and 12-15 remain rejected.
Response to Arguments
2 Applicant’s arguments with respect to claims 1, 8, and 15 filed on 04/03/2026, with respect to the rejection under 35 U.S.C. § 103 regarding that the prior art does not teach the following but not limited to “wherein the attention maps indicate defect locations in the 2D image frames; …wherein the 3D attention model displays infrastructure damage conditions around a current inspection location, enabling inspectors to evaluate defect severity based on spatial context.”. This argument has been considered, but are moot due to new grounds of rejection, alongside with modifications surrounding previously used prior art.
3 Regarding claims 2-3, 5-7, and 9-14, they directly/indirectly depend on independent claims 1 and 8 respectively. Applicant does not argue anything other than independent claims 1, 8, and 15. The limitations in those claims, in conjunction with combination, was mostly previously established as explained, with some modifications surrounding some amendments from the dependent claims.
Claim Rejections - 35 USC § 103
4 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
5 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
6 Claim(s) 1-3, 5-6, 8-10, 12-13, and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over N. Li, F. Chang and C. Liu, "Spatial-Temporal Cascade Autoencoder for Video Anomaly Detection in Crowded Scenes," in IEEE Transactions on Multimedia, vol. 23, pp. 203-215, 2021, doi: 10.1109/TMM.2020.2984093 (hereinafter Li) in view of Lodato et al. (US 20180350134 A1) and Garg et al. (US 20230088335 A1).
7 Regarding claim 1, Li teaches a method for generating a 3D attention model from use of a trained classifier configured to generate an attention map from 2D image frames and a 3D reconstruction process configured to generate a 3D reconstructed representation from the 2D image frames ([Section III] reciting “The proposed method is based on a cascade classifier, called ST-CaAE, which is established based on two spatial-temporal autoencoders…”; [Section III. A] reciting “Multiple state-of-the-art works [16], [39], [42], [51] have split each video frame into many fixed-size non-overlapping image patches or video cuboids for local anomalies detection. Inspired by these methods, we collect raw 3D video cuboids from video frames utilizing a sliding window with size w×h×t, where w and h are the width and height of the sliding window, respectively, and t is the temporal depth (the number of patches at the same spatial location in continuous frames that are stacked together to construct a raw video cuboid) … In order to obtain the gradient cuboids, we first employ the method in [16] to calculate the magnitude of the 3D gradient at each pixel within each video frame to construct its 3D gradient map.”; [Section III. B] reciting “Four 3D deconvolutional layers (corresponding to each 3D convolutional layer) are stacked to form the reconstruction part, which rebuilds the input cuboid from the latent vector z.”), the method comprising:
for an input of the 2D image frames: creating, through the 3D reconstruction process, the 3D reconstructed representation using the 2D image frames after data collection of an inspection process ([Section I] reciting “We adopt a popular two-stream network that employs 3D gradient and optical flow maps for appearance and motion anomaly detection, respectively.”; [Section III. A] reciting “In order to obtain the gradient cuboids, we first employ the method in [16] to calculate the magnitude of the 3D gradient at each pixel within each video frame to construct its 3D gradient map. The 3D gradient map of each frame has three channels. The first and second channel include the values in the horizontal and vertical dimensions of the video, respectively; these describe the external pose or shape of an object. The third channel records values from the temporal dimension of the video; this channel characterizes the appearance changing of objects with time. Then, we apply the sliding window mentioned above to the 3D gradient maps to produce gradient cuboids.”), wherein the 3D reconstruction process comprises:
wherein the 3D reconstruction process comprises:
the 3D reconstructed representation associated; with the mapping to the 2D image frames ([Section III. A] reciting “In order to obtain the gradient cuboids, we first employ the method in [16] to calculate the magnitude of the 3D gradient at each pixel within each video frame to construct its 3D gradient map. The 3D gradient map of each frame has three channels. The first and second channel include the values in the horizontal and vertical dimensions of the video, respectively; these describe the external pose or shape of an object. The third channel records values from the temporal dimension of the video; this channel characterizes the appearance changing of objects with time. Then, we apply the sliding window mentioned above to the 3D gradient maps to produce gradient cuboids.”);
executing the trained classifier on the 2D image frames to generate attention maps of the 2D image frames, wherein the attention maps indicate defect locations in the 2D image frames ([Section I] reciting “We adopt a popular two-stream network that employs 3D gradient and optical flow maps for appearance and motion anomaly detection, respectively.”; [Section III] reciting “The proposed method is based on a cascade classifier, called ST-CaAE, which is established based on two spatial-temporal autoencoders…”; [Section III A.] reciting “Based on the fusion of the appearance and motion anomaly scores, the unnecessary normal cuboids are removed, and the suspicious anomalous cuboids are identified. In stage 2, similarly, the related gradient and optical flow cuboids of suspicious anomalous cuboids are extracted and input to the ST-CAE in the two streams to calculate the appearance and anomaly scores of each local patch in the suspicious anomalous cuboids using the reconstruction error based strategy…The extracted gradient and optical flow cuboids are input to the two-stream ST-CaAE, which automatically extracts high-level spatial-temporal features from the video sequences.”);
8 Li does not explicitly teach reconstructing a 3D image model from the 2D image frames; determining a mapping between the 2D image frames and the 3D image model; and providing the 3D image model and the mapping as the 3D reconstructed representation, … projecting the attention maps of the 2D image frames to the 3D reconstructed representation based on the mapping to the 2D image frames and use the attention maps to sample colors of the 3D image model instead of the 2D image frames; and storing the 3D attention model comprising the associated 3D attention maps and the 3D reconstructed representation, wherein the 3D attention model displays infrastructure damage conditions around a current inspection location, enabling inspectors to evaluate defect severity based on spatial context.
9 Lodato teaches reconstructing a 3D image model from the 2D image frames ([0109] reciting “In certain examples, system 100 may be incorporated in a media player device that may receive 2D color data, depth data, and metadata for a captured 3D scene and use the received data to render virtual reality content, as described herein, for presentation to a user of the media player device such that the user of the media player device may experience a virtual reconstruction of the 3D scene from a dynamically selected viewpoint within the virtual reconstruction of the 3D scene.”);
determining a mapping between the 2D image frames and the 3D image model; and providing the 3D image model and the mapping as the 3D reconstructed representation ([0084] reciting “The generated image view may be represented in any suitable way, including as data (e.g., data representative of fragments) mapped (e.g., rasterized) from 3D coordinates in virtual 3D space 504 to a set of 2D image coordinates in an image plane such that the data may be used to generate and output display screen data (e.g., pixel data) representative of the image view for display on a 2D display screen.”; [0091] reciting “Because rendering facility 102 may accumulate partial 3D meshes 1202 that have overlapping sections associated with common 3D coordinates in virtual 3D space 504, rendering facility 102 may select multiple primitives or samples, from multiple partial 3D meshes 1202, for a common 3D coordinate in virtual 3D space 504 to be mapped to a common 2D coordinate of the image view 1302 as represented in the frame buffer.”)…
projecting the attention maps of the 2D image frames to the 3D reconstructed representation based on the mapping to the 2D image frames and use the attention maps to sample colors of the 3D image model instead of the 2D image frames ([0023] reciting “The metadata may include information associated with the 3D scene, such as information about the plurality of capture devices, that is usable by the rendering system to project the 2D color data and depth data into a virtual 3D space to produce virtual representations of the 3D scene in the virtual 3D space such that the projected data may be used by the rendering system to render a view of the virtual 3D space”; [0025] reciting “In certain examples, the rendering system may determine an accumulation region and accumulate and blend only primitives or fragments that are located within the accumulation region.”; [0085] reciting “Rendering facility 102 may generate an image view of the virtual 3D space from an arbitrary viewpoint within the virtual 3D space by accumulating partial 3D meshes projected into the virtual 3D space and blending color samples for the partial 3D meshes to form the image view of the virtual 3D space.”; [0091] reciting “Because rendering facility 102 may accumulate partial 3D meshes 1202 that have overlapping sections associated with common 3D coordinates in virtual 3D space 504, rendering facility 102 may select multiple primitives or samples, from multiple partial 3D meshes 1202, for a common 3D coordinate in virtual 3D space 504 to be mapped to a common 2D coordinate of the image view 1302 as represented in the frame buffer.”); and storing the 3D attention model comprising the associated 3D attention maps and the 3D reconstructed representation ([0046] reciting “Storage facility 104 may further include any other data as may be used by rendering facility 102 to form an image view of a virtual representation of a 3D scene in a virtual 3D space from an arbitrary viewpoint within the virtual 3D space as may serve a particular implementation.”).
10 It would have been obvious to one with ordinary skill before the effective filing date of the claimed invention, to have modified the method (taught by Li) to incorporate the teachings of Lodato to provide a method that can use the 2d image frames that are provided from Li to create a 3d model from various mappings , as well as to provide a method that can project various 2d to 3d related content that function similarly to attention maps that can get specific partials, to use them for sampling colors for 3d related material, and to store them using the actual attention maps that were taught by Li. Doing so would generate an image view of the virtual 3D space from an arbitrary viewpoint within the virtual 3D space by accumulating partial 3D meshes projected into the virtual 3D space and blending color samples for the partial 3D meshes to form the image view of the virtual 3D space as stated by Lodato ([0085] recited).
11 Li in view of Lodato does not explicitly teach wherein the 3D attention model displays infrastructure damage conditions around a current inspection location, enabling inspectors to evaluate defect severity based on spatial context.
12 Garg teaches wherein the 3D attention model displays infrastructure damage conditions around a current inspection location, enabling inspectors to evaluate defect severity based on spatial context ([0014] reciting “In one embodiment, a road and infrastructure analysis tool can measure 3D sizes and localize positions on a map to detect road attributes. The road and infrastructure analysis tool can detect sidewalks, road lanes, crosswalks, street signs, street lights, trees, bridges, etc. The road and infrastructure analysis tool can detect and identify road conditions that may need maintenance or otherwise require attention, for example, potholes, cracks, ruts, faded road markings, etc. The road and infrastructure analysis tool can also detect infrastructure deterioration, for example, corrosion, rot, crumbing cement, etc. using trained neural networks.”).
13 It would have been obvious to one with ordinary skill before the effective filing date of the claimed invention, to have modified the method (taught by Li in view of Lodato) to incorporate the teachings of Garg to provide a method that can display types of infrastructure damage from a location to evaluate certain types of defects from the maps that can be provided by Li in view of Lodato. Doing so would detect road hazards in a real road setting as stated by Garg ([Abstract] recited).
14 Regarding claim 2, Li in view of Lodato and Garg teaches the method of claim 1 (see claim 1 rejection above), wherein the trained classifier is trained against labeled 2D image frames classified as normal or anomalous and configured to output a classification for an input 2D image frame as normal or anomalous and the attention map indicating defects in the 2D image frames labeled as anomalous by the trained classifier (Li; [Section III. A] reciting “Based on the fusion of the appearance and motion anomaly scores, the unnecessary normal cuboids are removed, and the suspicious anomalous cuboids are identified. In stage 2, similarly, the related gradient and optical flow cuboids of suspicious anomalous cuboids are extracted and input to the ST-CAE in the two streams to calculate the appearance and anomaly scores of each local patch in the suspicious anomalous cuboids using the reconstruction error based strategy.”; [Section I] reciting “We adopt a popular two-stream network that employs 3D gradient and optical flow maps for appearance and motion anomaly detection, respectively.”).
15 Regarding claim 3, Li in view of Lodato and Garg teaches the method of claim 1 (see claim 1 rejection above), wherein the executing the trained classifier comprises: for the input of a 2D image frame from the 2D image frames: generating the attention map for the 2D image frame and an anomaly score (Li; [Section III]; reciting “Then, the gradient and optical flow cuboids are input to the ST-AAE in the two streams to obtain the appearance and motion anomaly scores using Gaussian distribution. Based on the fusion of the appearance and motion anomaly scores, the unnecessary normal cuboids are removed, and the suspicious anomalous cuboids are identified.”); and
weighing the generated attention map with the anomaly score (Li; [Section III. B; 2) “Training Strategy of ST-AAE”] reciting “In order to detect anomalous cuboids in the testing stage, our proposed deep model ST-AAE first attempts to train the distribution of generated codes z of normal cuboids to approach the prior p(z) … To ensure that the encoder is able to confuse the discriminator, Steps 5 and 7 update the weights of the encoder by maximizing the probability that the latent space vector z (generated by the encoder) is sampled from the prior distribution p(z).”).
16 Regarding claim 5, Li in view of Lodato and Garg teaches the method of claim 1 (see claim 1 rejection above),
17 Lodato from claim 1 can further teach the limitations, specifically further comprising providing a user interface configured to display the 2D image frames ([0084] reciting “The generated image view may be represented in any suitable way, including as data (e.g., data representative of fragments) mapped (e.g., rasterized) from 3D coordinates in virtual 3D space 504 to a set of 2D image coordinates in an image plane such that the data may be used to generate and output display screen data (e.g., pixel data) representative of the image view for display on a 2D display screen.”; [0163] reciting “In certain embodiments, I/O module 2208 is configured to provide graphical data to a display for presentation to a user. The graphical data may be representative of one or more graphical user interfaces and/or any other graphical content as may serve a particular implementation. I/O module 2208 may be omitted from certain implementations.”);
the attention maps of the 2D image frames, the 3D reconstructed representation and the associated 3D attention maps ([0101] reciting “Rendering facility 102 may output (e.g., write) the determined weighted average of the fragments to output buffer 1404 (e.g., in RGBA color model format as shown or in any other suitable format), which may be used to provide a display of the 2D image view on a display screen for viewing by a user associated with the display screen.”; [0130] reciting “To facilitate user 1608 in experiencing an image view of a virtual 3D space, media player device 1602 may include or be associated with at least one display screen (e.g., a head-mounted display screen built into a head-mounted virtual reality device or a display screen of a mobile device mounted to the head of the user with an apparatus such as a cardboard apparatus)”; [0084] reciting “Because the partial 3D meshes projected into the virtual 3D space produce partial virtual representations of a captured 3D scene (i.e., partial virtual reconstructions of one or more objects in the 3D scene)… The generated image view may be represented in any suitable way, including as data (e.g., data representative of fragments) mapped (e.g., rasterized) from 3D coordinates in virtual 3D space 504 to a set of 2D image coordinates in an image plane such that the data may be used to generate and output display screen data (e.g., pixel data) representative of the image view for display on a 2D display screen.”; [0153] reciting “…to display a perspective image view of the virtual 3D space that has been formed by the virtual reality content rendering system generating, accumulating, and blending partial 3D meshes as described herein.”).
18 It would have been obvious to one with ordinary skill before the effective filing date of the claimed invention, to have modified the method (taught by Li in view of Lodato and Garg) to incorporate additional teachings of Lodato to provide a type of user interface that can be used to display the frames, as well as the maps, the 3d reconstructions, and the 3d maps all provided by the teachings of Li in view of Lodato and Garg. Doing so would generate an image view of the virtual 3D space from an arbitrary viewpoint within the virtual 3D space by accumulating partial 3D meshes projected into the virtual 3D space and blending color samples for the partial 3D meshes to form the image view of the virtual 3D space as stated by Lodato ([0085] recited).
19 Regarding claim 6, Li in view of Lodato and Garg teaches the method of claim 1 (see claim 1 rejection above), wherein the 2D image frames are extracted from a recording of an infrastructure inspection video comprising infrastructure undergoing the inspection process (Li; [Section III. A] reciting “Multiple state-of-the-art works [16], [39], [42], [51] have split each video frame into many fixed-size non-overlapping image patches or video cuboids for local anomalies detection. Inspired by these methods, we collect raw 3D video cuboids from video frames utilizing a sliding window with size w×h×t, where w and h are the width and height of the sliding window, respectively, and t is the temporal depth (the number of patches at the same spatial location in continuous frames that are stacked together to construct a raw video cuboid).”; [Section IV. C. “The UCSD Dataset”] reciting “The UCSD dataset was captured from a campus using a static camera.”; See Figure 7 below).
PNG
media_image1.png
426
588
media_image1.png
Greyscale
20 Claim(s) 8 and 15 has similar limitations as of claim 1, therefore it is rejected under the same rationale as claim 1.
21 Claim 9 has similar limitations as of claim 2, therefore it is rejected under the same rationale as claim 2.
22 Claim 10 has similar limitations as of claim 3, therefore it is rejected under the same rationale as claim 3.
23 Claim 12 has similar limitations as of claim 5, therefore it is rejected under the same rationale as claim 5.
24 Claim 13 has similar limitations as of claim 6, therefore it is rejected under the same rationale as claim 6.
25 Claim(s) 7 and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over N. Li, F. Chang and C. Liu, "Spatial-Temporal Cascade Autoencoder for Video Anomaly Detection in Crowded Scenes," in IEEE Transactions on Multimedia, vol. 23, pp. 203-215, 2021, doi: 10.1109/TMM.2020.2984093 (hereinafter Li) in view of Lodato et al. (US 20180350134 A1) and Garg et al. (US 20230088335 A1) as of claim 1, further in view of Sema et al. (US 20240043053 A1).
26 Regarding claim 7, Li in view of Lodato and Garg teach the method of claim 1 (see claim 1 rejection above), but does not explicitly teach wherein the 3D reconstruction process further utilizes 3D sensors to generate the 3D reconstructed representation, and wherein the 3D sensors provide depth information for each 2D image frame, and wherein the depth information is combined with the 2D image frame to map pixels in the 2D image frame to points in 3D coordinates for accurate analysis of damages in infrastructures.
27 Sema teaches wherein the 3D reconstruction process further utilizes 3D sensors to generate the 3D reconstructed representation, and wherein the 3D sensors provide depth information for each 2D image frame, and wherein the depth information is combined with the 2D image frame to map pixels in the 2D image frame to points in 3D coordinates for accurate analysis of damages in infrastructures ([Abstract] reciting “In a method for detecting obstacles for a rail vehicle, 3D sensor data is detected from a surrounding region, 3D image data is generated from the 3D sensor data, and 2D image data is generated on the basis of the 3D image data. A 2D anomaly mask is ascertained or generated by comparing the 2D image data with reference image data which is free of a collision obstacle. In the process, image regions are identified as mask regions in the 2D image data which differ from the corresponding image regions in the reference image data. By fusing the 2D anomaly mask with the 3D image data, a 3D anomaly mask is generated in the 3D image data. Finally, the 3D image data which is part of the 3D anomaly mask is interpreted as a possible collision obstacle.”; [0065] reciting “Furthermore, the obstacle detection facility 50 also comprises a projection unit 53, which generates 2D image data 2D-BD on the basis of the 3D pixel data 3D-BD by way of a projection. Part of the obstacle detection facility 50 is also a mask-generating unit 54, which is configured to generate a 2D anomaly mask 2D-AM by comparing the 2D image data 2D-BD with reference image data R-BD, free of an obstacle, generated by an AI-based reconstruction.”).
28 It would have been obvious to one with ordinary skill before the effective filing date of the claimed invention, to have modified the method (taught by Li in view of Lodato and Garg) to incorporate the teachings of Sema to provide a type of 3d sensor that can contain depth information for various frames for analysis of certain damages, utilizing the frame and map data provided by Li in view of Lodato and Garg. Doing so would allow obstacle detection is fast and flexible as stated by Sema ([0031] recited).
29 Claim 14 has similar limitations as of claim 7, therefore it is rejected under the same rationale as claim 7.
Conclusion
30 The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Burton et al. (US 20170280130 A1) teaches reconstructing a 3D image model from the 2D frames, as well as can determine mapping based on the models.
Bell et al. (US 20160055268 A1) teaches 3D sensors for the 3D reconstructions, as well as to provide depth information and pixels.
31 Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOHNNY TRAN LE whose telephone number is (571)272-5680. The examiner can normally be reached Mon-Thu: 7:30am-5pm; First Fridays Off; Second Fridays: 7:30am-4pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached at (571) 272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JOHNNY T LE/Examiner, Art Unit 2614
/KENT W CHANG/Supervisory Patent Examiner, Art Unit 2614