Prosecution Insights
Last updated: October 01, 2026
Application No. 18/872,849

METHOD, DEVICE AND SYSTEM FOR DETECTING DYNAMIC OCCLUSION

Non-Final OA §103§112
Filed
Dec 09, 2024
Priority
Jul 01, 2022 — SG 10202250352P +1 more
Examiner
ALLISON, ANDRAE S
Art Unit
2673
Tech Center
2600 — Communications
Assignee
Grabtaxi Holdings Pte. Ltd.
OA Round
1 (Non-Final)
84%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
69%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
808 granted / 961 resolved
+22.1% vs TC avg
Minimal -16% lift
Without
With
+-15.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
15 currently pending
Career history
984
Total Applications
across all art units

Statute-Specific Performance

§101
11.5%
-28.5% vs TC avg
§103
49.2%
+9.2% vs TC avg
§102
18.2%
-21.8% vs TC avg
§112
13.8%
-26.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 961 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 08/21/2013 have been entered and considered. Initialed copies of the PTO-1449 by the Examiner are attached. Claim Objections Claim 14 is objected to because of the following informalities: Claim 14 is missing the means to carry out the function. The function of “determine whether each voxel in the voxel grid is in a dynamically occluded state” is recited in lines 15-16. In order to avoid clarity issues and prevent a rejection under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, the Examiner suggests Applicant amend the claim to include an “image aggregator” as shown in Fig 2 and described in [p][0053][0057] of the filed specification. Appropriate correction is required. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that use the word “means” or “step” but are nonetheless not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph because the claim limitation(s) recite(s) sufficient structure, materials, or acts to entirely perform the recited function. If applicant intends to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to remove the structure, materials, or acts that performs the claimed function; or (2) present a sufficient showing that the claim limitation(s) does/do not recite sufficient structure, materials, or acts to perform the claimed function. Claims 1 and 14 recite limitations that use words like “means” (or “step”) or similar terms with functional language and do invoke 35 U.S.C. 112(f): Claim 1; recites the limitation, “step of receiving …..” [Line 3]. Claim 1; recites the limitation, “step of determining…..” [Line 5]. Claims 1; recites the limitation, “step of generating a corresponding depth…...” [Line 7]. Claim 1; recites the limitation, “step of generating corresponding semantic ……,” [Line 8]. Claim 1; recites the limitation, “step of grouping ……,” [Line 9]. Claim 1; recites the limitation, “step of generating a voxel ……,” [Line 11]. Claim 1; recites the limitation, “step of determining ……,” [Line 12]. Claim 14; recites the limitation, “input module, configured to…..” [Line 3] Claim 14; recites the limitation, “a device pose module, configured to…..” [Line 5]. Claim 14; recites the limitation, “a depth map generation module, configured to…..” [Line 8]. Claim 14; recites the limitation, “a segmentation module, configured to…...” [Line 9]. Claim 14; recites the limitation, “an image aggregator module, configured to……,” [Line 11]. Claim 14; recites the limitation, “a voxel grid state estimator, configured to……,” [Line 14]. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. After a careful analysis, as disclosed above, and a careful review of the specification the following limitations in claims 1 and 14: (i) “an input module” (Fig. 5, #510. Paragraph [0058, 0076 and 0084-0085]- a first and second receiving module is describe as the user input unit 707 is configured to receive a first gesture input performed on first message content, where the first gesture input is a gesture. The electronic device 500 may further include a second receiving module. The user input unit 707 includes a touch panel 7071 and another input device 7072. The touch panel 7071 is also referred to as a touchscreen, and may collect a touch operation performed by a user on or near the touch panel 7071 (such as an operation performed by a user on the touch panel 7071 or near the touch panel 7071 by using any proper object or accessory, such as a finger or a stylus). The touch panel 7071 may include two parts: a touch detection apparatus and a touch controller. The second receiving module is configured to receive a second gesture input performed on first message content on a communication interface. The first and second receiving module is illustrated in Fig. 5, as a black box (i) “a device pose module” (Fig. 5, #510. Paragraph [0058, 0076 and 0084-0085]- a first and second receiving module is describe as the user input unit 707 is configured to receive a first gesture input performed on first message content, where the first gesture input is a gesture. The electronic device 500 may further include a second receiving module. The user input unit 707 includes a touch panel 7071 and another input device 7072. The touch panel 7071 is also referred to as a touchscreen, and may collect a touch operation performed by a user on or near the touch panel 7071 (such as an operation performed by a user on the touch panel 7071 or near the touch panel 7071 by using any proper object or accessory, such as a finger or a stylus). The touch panel 7071 may include two parts: a touch detection apparatus and a touch controller. The second receiving module is configured to receive a second gesture input performed on first message content on a communication interface. The first and second receiving module is illustrated in Fig. 5, as a black box #510 thus have sufficient structure or material wherein is a touch panel.). (ii) “a depth map generation module” (Fig. 5, #520. Paragraph [0060 and 0088]- a processing module is described as the processing module 520 may include at least one of the following: a first processing unit, a second processing unit, a third processing unit, and a fourth processing unit. The processor 710 is a control center of the electronic device, connects all parts of the entire electronic device by using various interfaces and lines, and performs various functions of the electronic device and data processing by running or executing a software program and/or a module that are/is stored in the memory 709 and by invoking data stored in the memory 709, to overall monitor the electronic device. The processor 710 may include one or more processing units. In some embodiments, an application processor and a modem processor may be integrated into the processor 710. The processing module is illustrated in Fig. 5, as a black box #520 thus have sufficient structure or material wherein is a processor.). (iii) “a segmentation module” (Fig. 5, #530. Paragraph [0062, 0078 and 0082]- a sending module is described as the sending module 530 may include a first determining unit, a second determining unit, and a sending unit. The radio frequency unit 701 may be configured to receive and send information or a signal in a call process. Specifically, after receiving downlink data from a base station, the radio frequency unit 701 sends the downlink data to the processor 710 for processing. In addition, the radio frequency unit 701 sends uplink data to the base station. Usually, the radio frequency unit 701 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer, and the like. In addition, the radio frequency unit 701 may communicate with a network and another device through a wireless communication system. The sending module is illustrated in Fig. 5, as a black box #530 thus have sufficient structure or material wherein is an antenna, transceiver and amplifier.). (iv) “an image aggregator module” (Fig. 5. Paragraph [0058 and 0059]- a enlarging module is described as the electronic device 500 may further include a second receiving module and an enlarging module. The second receiving module is configured to receive a second gesture input performed on first message content on a communication interface; and the enlarging module is configured to display the first message content through enlarging in response to the second gesture input. The enlarging module is not illustrated thus have no sufficient structure or material.). (v) “a voxel grid state estimator” (Fig. 5. Paragraph [0062 and 0088]- a first and second determining unit is described as the sending module 530 may include a first determining unit, a second determining unit, and a sending unit. The processor 710 is a control center of the electronic device, connects all parts of the entire electronic device by using various interfaces and lines, and performs various functions of the electronic device and data processing by running or executing a software program and/or a module that are/is stored in the memory 709 and by invoking data stored in the memory 709, to overall monitor the electronic device. The processor 710 may include one or more processing units. In some embodiments, an application processor and a modem processor may be integrated into the processor 710. The first and second determining unit is illustrated in Fig. 5, thus have sufficient structure or material wherein is a processor.). (vi) “sending unit” (Fig. 5. Paragraph [0062-0063 and 0088]- a sending unit is described as the sending module 530 may include a first determining unit, a second determining unit, and a sending unit. The radio frequency unit 701 may be configured to receive and send information or a signal in a call process. Specifically, after receiving downlink data from a base station, the radio frequency unit 701 sends the downlink data to the processor 710 for processing. In addition, the radio frequency unit 701 sends uplink data to the base station. Usually, the radio frequency unit 701 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer, and the like. In addition, the radio frequency unit 701 may communicate with a network and another device through a wireless communication system. The sending unit is illustrated in Fig. 5, thus have sufficient structure or material wherein is an antenna, transceiver and amplifier.). If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-5, 10-11 and 13-19 are rejected under 35 U.S.C. 103 as being unpatentable over Shepel et al (NPL titled: Occupancy Grid Generation With Dynamic Obstacle Segmentation in Stereo Images) in view of Tong et al (English translation of CN 112435262 A) Regarding independent claim 1, Shepel teaches a method for detecting dynamic occlusion on one or more images associated with a location of interest (creating a high-quality method for generating an occupancy map using stereo camera data and odometry from an inertial navigation system – see section III, [p][001]) comprising the steps of: receiving a plurality of image data files associated with a location of interest ([r]aw images from cameras are rectified and transferred to the stereo matching algorithm – see section IV, [p][001]), each image data file associated with at least a part of the location of interest (note that since the input are stereo images, the images captured would capture different field of view - see section IV, [p][001] ); for each image determining the position and orientation of an image capturing device relative to the location of interest and generating device pose information ([a]lso, to build an accumulated grid, the proposed approach requires camera position and its orientation using the inertial navigation system (INS) - see section IV, [p][001] and Fig 1.); generating a corresponding depth map ([d]uring the experiments, the rSGM [44] algorithm was used since this algorithm generates a dense depth map and dense map generation - see section IV, [p][001] and Fig 1); generating a corresponding semantic segmentation (image segmentation – see Fig 1); grouping device pose information, depth map and semantic segmentation to form an image group ([g]eneral scheme of the proposed approach. A point cloud with semantic labeling together with data from inertial navigation system (INS) is fed as input to the occupancy grid generation module - see Fig 1); generating a voxel grid associated with the image group ([f]irst, an one-shot (generated from a single frame) occupancy map is built and One-Shot Occupancy Grid: Each cell of the one-shot occupancy map MO can be in one of three states: cF- free, cO- occupied, cU- unknown – see section IV, subsection B, [p][001] ); and determining whether each voxel in the voxel grid is in a dynamically occluded state (First, an one-shot (generated from a single frame) occupancy map is built. Second, multiple one-shot grids are used to build an accumulated occupancy map and we report precision (P) and recall (R) metrics calculated for occupancy grid cells marked as obstacles for different cell sizes. The results of comparing maps for raw point clouds and for clouds with semantics are shown in Table VI – see section IV, subsection B, [p][001] and section V, subsection B, [p][001]). Shepel does not explicitly teach grouping the image based on coordinates of the location of interest. Tong explicitly teaches grouping the image based on coordinates of the location of interest (judging and inserting a key frame, and performing point cloud processing through a local graph building process to obtain a sparse point cloud map and s1.3: and transferring the coordinates in the camera coordinate system to the pixel coordinate system –see page 4, [p][002][013]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Shepel of a device for detecting dynamic occlusion on one or more images associated with a location of interest comprising an input module configured to receive a plurality of image data files associated with a location of interest; with the teachings of Tong grouping the image based on coordinates of the location of interest. Wherein having Shepel grouping the image based on coordinates of the location of interest. The motivation behind the modification would have been for detecting a dynamic environment information by inserting a key frame and generating an occupancy map of the surrounding space from noisy point clouds obtained from one or several stereo cameras since both Shepel and Tong relates to dynamic environment information detection, wherein Shepel generates an occupancy map of the surrounding space from noisy point clouds obtained from one or several stereo cameras while Tong detects a dynamic environment information by inserting a key frame (Please see Shepel et al (NPL titled: Occupancy Grid Generation With Dynamic Obstacle Segmentation in Stereo Images), see Abstract in view of Tong et al (English translation of CN 112435262 A), see page 2, [p][001]). Regarding claim 2, Shepel teaches the method of claim 1, PNG media_image1.png 87 12 media_image1.png Greyscale wherein the step of determining whether each voxel in the voxel grid is in a dynamically occluded state includes selecting a state from a set of states comprising the following states: unseen, dynamically occluded, void, and occupied (see section 5, B.1). Regarding claim 3, Shepel in view of Tong teach the method of claim 1, further comprising the step of generating a voxel grid state array comprising the states of the each of the voxel in the voxel grid (see section 5, B.1). Regarding claim 4, Shepel in view of Tong teach the method of claim 3, wherein the voxel grid state array is a one- dimensional array (the points in the cell are soted by z-coordinates in an increasing order - see section 5, B.1). Regarding claim 5, Shepel in view of Tong teach the method of claim 3, wherein at an initialization of the voxel grid state array, the state of every voxel is set to an unseen state (unknown - see section 5, B.1). Regarding claim 10, Shepel in view of Tong teach the method of any one of the preceding claims1, wherein the step of grouping comprises matching the location of interest with at least one feature on a reference map (depending on the cell type, the probability value is associated with each cell with corresponding pcO=0.95 , pcU=0.5 or pcF=0.4 , related to each other by inequalities pcO>pcU>pcFsee section IV, subsection B, [p][003]). Regarding claim 11 Shepel in view of Tong teach the method of claim 10, wherein the step of generating a voxel grid comprises determining a length, a width and a height of the voxel grid based on the at least one feature on the reference map (see section II, [p][002]). Regarding claim 13, Shepel in view of Tong teach the method of claim 1, Shepel explicitly teaches wherein the step of generating a corresponding semantic segmentation of the image comprises using a trained convolutional neural network model to generate semantic labels associated with one or more features on the image ([t]his article uses a supervised approach to train a neural network model; therefore, the final quality of the segmentation algorithm significantly depends on the choice of the loss function – see section IV, subsection A, step 3, [p][001]). Regarding independent claim 14, Shepel teaches a device for detecting dynamic occlusion on one or more images associated with a location of interest (device for creating a high-quality for generating an occupancy map using stereo camera data and odometry from an inertial navigation system – see section III, [p][001]) comprising: an input module (stereo camera – see Fig 1) configured to receive a plurality of image data files associated with a location of interest ([r]aw images from cameras are rectified and transferred to the stereo matching algorithm – see section IV, [p][001]), each image data file associated with at least a part of the location of interest (note that since the input are stereo images, the images captured would capture different field of view - see section IV, [p][001]); a device pose module (INS – see Fig 1) configured to determine the position and orientation of an image capturing device relative to the location of interest and generating device pose information ([a]lso, to build an accumulated grid, the proposed approach requires camera position and its orientation using the inertial navigation system (INS) - see section IV, [p][001] and Fig 1.); a depth map generation module (stereo camera hardware.- see section V, subsection B, [p][006] and Fig 1) configured to generate a corresponding depth map ([d]uring the experiments, the rSGM [44] algorithm was used since this algorithm generates a dense depth map and dense map generation - see section IV, [p][001] and Fig 1); a segmentation module (segmentation module – see section III, [p][001]) configured to generating a corresponding semantic segmentation (image segmentation – see Fig 1); an image aggregator module (occupancy grid generation – see Fig 1) configured to group device pose information, depth map and semantic segmentation to form an image group ([g]eneral scheme of the proposed approach. A point cloud with semantic labeling together with data from inertial navigation system (INS) is fed as input to the occupancy grid generation module - see Fig 1); a voxel grid state estimator (one shot occupancy grid – see Fig 1) configured to generate a voxel grid associated with the image group ([f]irst, an one-shot (generated from a single frame) occupancy map is built and One-Shot Occupancy Grid: Each cell of the one-shot occupancy map MO can be in one of three states: cF- free, cO- occupied, cU- unknown – see section IV, subsection B, [p][001] ); and determining whether each voxel in the voxel grid is in a dynamically occluded state (First, an one-shot (generated from a single frame) occupancy map is built. Second, multiple one-shot grids are used to build an accumulated occupancy map and we report precision (P) and recall (R) metrics calculated for occupancy grid cells marked as obstacles for different cell sizes. The results of comparing maps for raw point clouds and for clouds with semantics are shown in Table VI – see section IV, subsection B, [p][001] and section V, subsection B, [p][001]). Shepel does not explicitly teach grouping the image based on coordinates of the location of interest. Tong explicitly teaches grouping the image based on coordinates of the location of interest (judging and inserting a key frame, and performing point cloud processing through a local graph building process to obtain a sparse point cloud map and s1.3: and transferring the coordinates in the camera coordinate system to the pixel coordinate system –see page 4, [p][002][013]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Shepel of a device for detecting dynamic occlusion on one or more images associated with a location of interest comprising an input module configured to receive a plurality of image data files associated with a location of interest; with the teachings of Tong grouping the image based on coordinates of the location of interest. Wherein having Shepel grouping the image based on coordinates of the location of interest. The motivation behind the modification would have been for detecting a dynamic environment information by inserting a key frame and generating an occupancy map of the surrounding space from noisy point clouds obtained from one or several stereo cameras since both Shepel and Tong relates to dynamic environment information detection, wherein Shepel generates an occupancy map of the surrounding space from noisy point clouds obtained from one or several stereo cameras while Tong detects a dynamic environment information by inserting a key frame (Please see Shepel et al (NPL titled: Occupancy Grid Generation With Dynamic Obstacle Segmentation in Stereo Images), see Abstract in view of Tong et al (English translation of CN 112435262 A), see page 2, [p][001]). Regarding claim 15, Shepel in view of Tong teach the device of claim 14, wherein determination of whether each voxel in the voxel grid is in a dynamically occluded state includes selecting a state from a set of states comprising the following: unseen state, dynamically occluded state, void state, and occupied state (see section 5, B.1). Regarding claim 16, Shepel in view of Tong teach the device of claim14, wherein the voxel grid state estimator is further configured to generate a voxel grid state array comprising the states of the each of the voxel in the voxel grid (see section 5, B.1). Regarding claim 17, Shepel in view of Tong teach the device of claim 16, wherein the voxel grid state array is a one-dimensional array (the points in the cell are soted by z-coordinates in an increasing order - see section 5, B.1). Regarding claim 18, Shepel in view of Tong teach the system for updating a voxel grid state array comprising the device of claim 16, further comprising an updater to check if the voxel is previously detected to be in a dynamic occluded state and subsequently in a void state (unknown - see section 5, B.1) or occupied state. Regarding claim 19, Shepel in view of Tong teach the system of claim 18, wherein if the voxel is detected to be in a void state or occupied state, the system is configured to update the voxel grid state array associated with the change of state(s) (update the state of each cell - section IV, subsection B, step 2). Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Shepel et al (NPL titled: Occupancy Grid Generation With Dynamic Obstacle Segmentation in Stereo Images) in view of Tong et al (English translation of CN 112435262 A) as applied to claim 1 further in view of Leung et al (NPL titled: Embedded Voxel Colouring with Adaptive Threshold Selection Using Globally Minimal Surfaces). Regarding claim 6, Shepel in view of Tong teach the method of claim 5, Shepel in view of Tong does not explicitly teach further comprising the step of reprojecting the voxel onto a two-dimensional image plane based on the device pose information, and obtaining an associated two-dimensional pixel. However, Leung explicitly teaches further comprising the step of reprojecting the voxel onto a two-dimensional image plane based on the device pose information, and obtaining an associated two-dimensional pixel ([t]his record stores: (i) the list of pixels to which that voxel projects, (ii) the distance of that voxel to each camera and (iii) intermediate values of the metric computations. The pixel-list is maintained to avoid inefficiently reprojecting the voxel onto each camera. – see section 5, [p][004]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Shepel as modified by Tong of a device for detecting dynamic occlusion on one or more images associated with a location of interest comprising an input module configured to receive a plurality of image data files associated with a location of interest; with the teachings of Leung further comprising the step of reprojecting the voxel onto a two-dimensional image plane based on the device pose information, and obtaining an associated two-dimensional pixel. Wherein having Shepel further comprising the step of reprojecting the voxel onto a two-dimensional image plane based on the device pose information, and obtaining an associated two-dimensional pixel. The motivation behind the modification would have been for performs volumetric reconstruction which considers occlusion modelling in detail and generating an occupancy map of the surrounding space from noisy point clouds obtained from one or several stereo cameras since both Shepel and Leung relates to dynamic environment information detection, wherein Shepel generates an occupancy map of the surrounding space from noisy point clouds obtained from one or several stereo cameras while Leung performs volumetric reconstruction which considers occlusion modelling in detail (Please see Shepel et al (NPL titled: Occupancy Grid Generation With Dynamic Obstacle Segmentation in Stereo Images), see Abstract in view of Leung et al (NPL titled: Embedded Voxel Colouring with Adaptive Threshold Selection Using Globally Minimal Surfaces), see section 7, [p][001] ). Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Shepel et al (NPL titled: Occupancy Grid Generation With Dynamic Obstacle Segmentation in Stereo Images) in view of Tong et al (English translation of CN 112435262 A) as applied to claim 1 further in view of Leung et al (NPL titled: Embedded Voxel Colouring with Adaptive Threshold Selection Using Globally Minimal Surfaces) further in view of Alface (Pub No.: US20170032569 ) Regarding claim 7, Shepel in view of Tong and Leung teach the method of claim 6, Shepel in view of Tong and Leung does not teach further comprising a step of determining if the pixel is out of an image border specified by the image resolution, and assigning the dynamic occluded state to the voxel if the associated pixel is determined to be within the image border. However, Alface explicitly teaches further comprising a step of determining if the pixel is out of an image border specified by the image resolution (generating a digital environment in the form of a 3D model representing the outer borders - see [p][0013]), and assigning the dynamic occluded state to the voxel if the associated pixel is determined to be within the image border (when observing the object via one camera, there is a visible side and an occluded side at the object – see [p][0021]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Shepel as modified by Tong and Leung of a device for detecting dynamic occlusion on one or more images associated with a location of interest comprising an input module configured to receive a plurality of image data files associated with a location of interest; with the teachings of Alface further comprising a step of determining if the pixel is out of an image border specified by the image resolution, and assigning the dynamic occluded state to the voxel if the associated pixel is determined to be within the image border. Wherein having Shepel further comprising a step of determining if the pixel is out of an image border specified by the image resolution, and assigning the dynamic occluded state to the voxel if the associated pixel is determined to be within the image border. The motivation behind the modification would have been for keeping information in the 3D model that falls out of the view of a second image and generating an occupancy map of the surrounding space from noisy point clouds obtained from one or several stereo cameras since both Shepel and Alface relates to dynamic environment information detection, wherein Shepel generates an occupancy map of the surrounding space from noisy point clouds obtained from one or several stereo cameras while Alface allows to keep information in the 3D model that falls out of the view of a second image. (Please see Shepel et al (NPL titled: Occupancy Grid Generation With Dynamic Obstacle Segmentation in Stereo Images), see Abstract in view of Alface (Pub No.: US20170032569 ), [p][0012] ). Claims 12 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Shepel et al (NPL titled: Occupancy Grid Generation With Dynamic Obstacle Segmentation in Stereo Images) in view of Tong et al (English translation of CN 112435262 A) as applied to claim 1 further in view of Garg et al (Pub No.: 20220343525). Regarding claim 12, Shepel in view of Tong teach the method of claims 1, Shepel in view of Tong does not explicitly teach wherein the step of generating the corresponding depth map of the image comprises using a trained deep learning model or a structure-from-motion (SfM) algorithm to estimate the depth map using the image as the only input. However, Garg explicitly teaches wherein the step of generating the corresponding depth map of the image comprises using a trained deep learning model (a trained neural network that can use the images and depth estimates to produce a joint depth map for the scene – see [p][0107]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Shepel as modified by Tong of a device for detecting dynamic occlusion on one or more images associated with a location of interest comprising an input module configured to receive a plurality of image data files associated with a location of interest; with the teachings of Garg wherein the step of generating the corresponding depth map of the image comprises using a trained deep learning model. Wherein having Shepel wherein the step of generating the corresponding depth map of the image comprises using a trained deep learning model. The motivation behind the modification would have been for modifying an image representing the scene based on the joint depth map and generating an occupancy map of the surrounding space from noisy point clouds obtained from one or several stereo cameras since both Shepel and Garg relates to dynamic environment information detection, wherein Shepel generates an occupancy map of the surrounding space from noisy point clouds obtained from one or several stereo cameras while Garg modifies an image representing the scene based on the joint depth map (Please see Shepel et al (NPL titled: Occupancy Grid Generation With Dynamic Obstacle Segmentation in Stereo Images), see Abstract in view of Garg et al (Pub No.: 20220343525)., see Abstract). Regarding claim 20, Shepel in view of Tong teaches a Shepel teaches a comprising instructions (new algorithm for generating an occupancy map – see Abstract), which, when executed by one or more processors (CPU – see section IV, [p][001]), cause the execution of the method according claim 1. Shepel in of Tong does not explicitly teach non-transitory computer-readable storage medium. However, Garg explicitly teaches non-transitory computer-readable storage medium (104 – see Fig 1) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Shepel as modified by Tong of a device for detecting dynamic occlusion on one or more images associated with a location of interest comprising an input module configured to receive a plurality of image data files associated with a location of interest; with the teachings of Garg non-transitory computer-readable storage medium. Wherein having Shepel non-transitory computer-readable storage medium. The motivation behind the modification would have been for modifying an image representing the scene based on the joint depth map and generating an occupancy map of the surrounding space from noisy point clouds obtained from one or several stereo cameras since both Shepel and Garg relates to dynamic environment information detection, wherein Shepel generates an occupancy map of the surrounding space from noisy point clouds obtained from one or several stereo cameras while Garg modifies an image representing the scene based on the joint depth map (Please see Shepel et al (NPL titled: Occupancy Grid Generation With Dynamic Obstacle Segmentation in Stereo Images), see Abstract in view of Garg et al (Pub No.: 20220343525), see Abstract). Allowable Subject Matter Claims 8-9 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Li et al (Pub No.: 20220180548) discloses determining an image feature corresponding to a point cloud of an input image (110). Semantic segmentation information, instance mask information, and keypoint information of an object are determined (120) based on the image feature. A pose of the object is estimated based on a combination of the information (130). A point cloud feature is extracted based on depth image. A three-dimensional grid is generated based on point cloud information corresponding to the input image, and the instance information is determined. A target predicted value of a key point is determined corresponding to each keypoint and point cloud. FLEISHMAN et al (20190043203) discloses a method wherein current frame is semantically segmented to form a segmented or label frame, and this is repeated for individual current frames in a video sequence. Each semantically segmented frame, depending on a camera (or sensor) pose used to form the current frame, are then used to update the semantic model. This is typically performed without factoring the sequence or history of semantic updating that occurred previously during a video sequence while performing the semantic segmentation of the current frame to form the segmented frame. This results in a significantly less accurate analysis resulting in errors and inaccuracies in semantic assignments to vertices or voxels in the semantic model. Sokolova et al (US Patent No.: 12675950) discloses a method for 3D scene reconstruction and visualization, may include, using at least one processor: obtaining a trained base neural network by training the base neural network for obtaining distance information for voxels of a real scene; operating the trained base neural network for obtaining the distance information of an input sequence of frames of the real scene; inputting the distance information to an algorithm that outputs a 3D reconstruction of the real scene; obtaining a 3D visualization of the real scene by rendering the 3D reconstruction of the real scene; and instructing at least one display to display the 3D visualization of the real scene. Blechschmidt et al (US Patent No.: 12175162) discloses systems and methods that determine relationships between objects based on an original semantic mesh of vertices and faces that represent the 3D geometry of a physical environment. Such an original semantic mesh may be generated and used to provide input to a machine learning model that estimates relationships between the objects in the physical environment. For example, the machine learning model may output a graph of nodes and edges indicating that a vase is on top of a table or that a particular instance of a vase, V1, is on top of a particular instance of a table, T1. Inquiries Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANDRAE S ALLISON whose telephone number is (571)270-1052. The examiner can normally be reached on Monday-Friday 9am-5pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns, can be reached on (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ANDRAE S ALLISON/Primary Examiner, Art Unit 2673 July 10, 2026
Read full office action

Prosecution Timeline

Dec 09, 2024
Application Filed
Jul 15, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725399
Systems and Methods for Detecting a Travelling Object Vortex
3y 9m to grant Granted Sep 01, 2026
Patent 12725402
METHOD TO CLASSIFY QUENCH PATTERNS OF HEAT-TREATED COATED MINERAL GLASSES AND PREDICTS THE OPTICAL VISIBILITY THEREOF
2y 9m to grant Granted Sep 01, 2026
Patent 12718323
MASSIVELY PARALLEL AMPLITUDE-ONLY OPTICAL PROCESSING SYSTEM AND METHODS FOR MACHINE LEARNING
3y 6m to grant Granted Aug 25, 2026
Patent 12711786
ASSOCIATING TWO DIMENSIONAL LABEL DATA WITH THREE-DIMENSIONAL POINT CLOUD DATA
2y 2m to grant Granted Aug 18, 2026
Patent 12700249
SYSTEMS AND METHODS FOR TRAINING MACHINE LEARNING ON NOISY AND INACCURATE IMMUNOSTAINS
3y 1m to grant Granted Aug 04, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
84%
Grant Probability
69%
With Interview (-15.5%)
2y 9m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 961 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month