Prosecution Insights
Last updated: October 01, 2026
Application No. 19/014,944

SYSTEMS AND METHODS FOR COHERENT MONITORING

Non-Final OA §102§103§112
Filed
Jan 09, 2025
Priority
Jan 31, 2019 — provisional 62/799,292 +4 more
Examiner
ALLEN, KYLA GUAN-PING TI
Art Unit
Tech Center
Assignee
Palantir Technologies Inc.
OA Round
1 (Non-Final)
89%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 89% — above average
89%
Career Allowance Rate
65 granted / 73 resolved
+29.0% vs TC avg
Strong +17% interview lift
Without
With
+16.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
26 currently pending
Career history
89
Total Applications
across all art units

Statute-Specific Performance

§101
10.1%
-29.9% vs TC avg
§103
54.5%
+14.5% vs TC avg
§102
14.5%
-25.5% vs TC avg
§112
19.4%
-20.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 73 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are pending regarding this application. Information Disclosure Statement The information disclosure statement (IDS) submitted on 01/09/2025 is considered and attached. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 8 and 18 recite the limitation "wherein the detecting one or more events associated with the tracked object comprises" in the preamble. There is insufficient antecedent basis for this limitation in the claims. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1, 2, 5, 11, 12, and 15 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Chang et al. (U.S. Publication No. 2019/0205651 A1), hereinafter Chang. Regarding claim 1, Chang teaches a system for intelligently monitoring an environment (Chang, see FIG. 9 and para. [0017]), comprising: one or more processors (Chang teaches “the processing device 5102 may include one or more processors and memory that stores computer-executable instructions that are executed by the one or more processors” in para. [0308]); and a memory storing instructions that, when executed by the one or more processors (Chang teaches “the processing device 5102 may include one or more processors and memory that stores computer-executable instructions that are executed by the one or more processors. The processing device 5102 may execute a video player application 5200” in para. [0308]), cause the system to: obtain content representing an environment, the content comprising a plurality of video frames (Chang teaches “at a step 4412, a broadcast video feed may be ingested, which may consist of an un-calibrated and un-synchronized video feed” in para. [0267]. See further para. [0264]-[0267] which describes the potential environments for which the content may represent); identify, based on the content, one or more discrete objects or events observed within the environment (Chang teaches “at a step 4404 objects may be detected within the broadcast video feed 4404, such as using machine-based object-recognition technologies. Objects may include players (including based on recognizing player numbers), body parts of players (e.g., heads of players, torsos of players, etc.) equipment (such as the ball in a basketball game), and many others” in para. [0267]); generate a media representation that augments the one or more discrete objects or events (Chang teaches “the user history may indicate that the user watches media segments of a particular type of event (e.g., dunks), but skips over other types of events (e.g., blocked shots). In this example, the dynamic video module 5020 may infer that the user prefers to consume media segments of dunks over media segments of blocked shots. In operation, the dynamic video module 5020 can utilize the indicated, predicted, and/or inferred user preferences to determine which media segments to include in the dynamic video and/or the duration of the media segments (e.g., should the media segment be shorter or longer” in para. [0291]. See also para. [0281] wherein the above process occurs in the context of additionally providing a graphical representation in conjunction with the dynamic synchronized media content. Here, the generated dynamic video that augments a dunk event is interpreted as equivalent to the claimed generated media representation that augments an object or event); and augment a graphical representation of the environment with the generated media representation (Chang teaches “the SDK 4804 may incorporate the additional content feeds into the synchronized media content 4810, by augmenting the dunk in the live or VOD feed with the graphical augmentation and the statistics” in para. [0281]). Regarding claim 2, Chang teaches the system of claim 1, wherein the media representation comprises different frames corresponding to different perspectives of the one or more discrete objects or events captured by sensors at different orientations and positions (Chang teaches “Synchronization among the output streams may enable combining and/or switching 4528 seamlessly among alternative video feeds (e.g., different angles, encoding, augmentations or the like) and data feeds of a live streamed event” in para. [0272]). Regarding claim 5, Chang teaches the system of claim 1, wherein instructions that, when executed by the one or more processors, further causes the system to: present the graphical representation of the environment (Chang teaches “the SDK 4804 may incorporate the additional content feeds into the synchronized media content 4810, by augmenting the dunk in the live or VOD feed with the graphical augmentation and the statistics” in para. [0281]); and simultaneously present a playback of the one or more discrete objects or events in a separate pane (Chang teaches “the step of creating the augmented video content may be done by a server and then transmitted to a user's client device for playback” in para. [0537]. Chang additionally teaches “available pieces of information and augmentation elements may be selected individually or in combination. In embodiments, combinations of audio, video, information, augmentation, replays, and the like may constitute channels for end-users to choose from. The smart pipe may contain sufficient indexed and aligned content to create derived content and interactive apps tied to live and recorded games” in para. [0278]. Chang further teaches simultaneously playing dynamic video (playback) with the graphical representation as shown in para. [0281]). Regarding claim 11, Chang teaches a method being implemented by a computing system including one or more physical processors and storage media storing machine-readable instructions (Chang teaches “the processing device 5102 may include one or more processors and memory that stores computer-executable instructions that are executed by the one or more processors” in para. [0308]), the method comprising: obtaining content representing an environment, the content comprising a plurality of video frames (Chang teaches “at a step 4412, a broadcast video feed may be ingested, which may consist of an un-calibrated and un-synchronized video feed” in para. [0267]. See further para. [0264]-[0267] which describes the potential environments for which the content may represent); identifying, based on the content, one or more discrete objects or events observed within the environment (Chang teaches “at a step 4404 objects may be detected within the broadcast video feed 4404, such as using machine-based object-recognition technologies. Objects may include players (including based on recognizing player numbers), body parts of players (e.g., heads of players, torsos of players, etc.) equipment (such as the ball in a basketball game), and many others” in para. [0267]); generating a media representation that augments the one or more discrete objects or events (Chang teaches “the user history may indicate that the user watches media segments of a particular type of event (e.g., dunks), but skips over other types of events (e.g., blocked shots). In this example, the dynamic video module 5020 may infer that the user prefers to consume media segments of dunks over media segments of blocked shots. In operation, the dynamic video module 5020 can utilize the indicated, predicted, and/or inferred user preferences to determine which media segments to include in the dynamic video and/or the duration of the media segments (e.g., should the media segment be shorter or longer” in para. [0291]. See also para. [0281] wherein the above process occurs in the context of additionally providing a graphical representation in conjunction with the dynamic synchronized media content. Here, the generated dynamic video that augments a dunk event is interpreted as equivalent to the claimed generated media representation that augments an object or event); and augmenting a graphical representation of the environment with the generated media representation (Chang teaches “the SDK 4804 may incorporate the additional content feeds into the synchronized media content 4810, by augmenting the dunk in the live or VOD feed with the graphical augmentation and the statistics” in para. [0281]). Regarding claim 12, Chang teaches the method of claim 11, wherein the media representation comprises different frames corresponding to different perspectives of the one or more discrete objects or events captured by sensors at different orientations and positions (Chang teaches “Synchronization among the output streams may enable combining and/or switching 4528 seamlessly among alternative video feeds (e.g., different angles, encoding, augmentations or the like) and data feeds of a live streamed event” in para. [0272]). Regarding claim 15, Chang teaches the method of claim 11, further comprising: presenting the graphical representation of the environment (Chang teaches “the SDK 4804 may incorporate the additional content feeds into the synchronized media content 4810, by augmenting the dunk in the live or VOD feed with the graphical augmentation and the statistics” in para. [0281]); and simultaneously presenting a playback of the one or more discrete objects or events in a separate pane (Chang teaches “the step of creating the augmented video content may be done by a server and then transmitted to a user's client device for playback” in para. [0537]. Chang additionally teaches “available pieces of information and augmentation elements may be selected individually or in combination. In embodiments, combinations of audio, video, information, augmentation, replays, and the like may constitute channels for end-users to choose from. The smart pipe may contain sufficient indexed and aligned content to create derived content and interactive apps tied to live and recorded games” in para. [0278]. Chang further teaches simultaneously playing dynamic video (playback) with the graphical representation as shown in para. [0281]). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 3, 4, 7, 8, 13, 14, 17, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Chang et al. (U.S. Publication No. 2019/0205651 A1), hereinafter Chang in view of Sai (U.S. Publication No. 2019/0057262 A1). Regarding claim 3, Chang teaches the system of claim 1, wherein the generating of the media representation comprises automatically tagging the one or more discrete objects or events based on a confidence level of the one or more discrete objects or events matching one or more respective (Chang teaches “machine learning algorithms are designed to output a measure of confidence. […] If an example is labeled by the machine and has confidence above the threshold, the event goes into the canonical event datastore 210 and nothing further is done” in para. [0177]). Chang fails to teach matching one or more respective templates that define respective characteristics of the one or more discrete objects or events. However, Sai teaches matching one or more respective templates that define respective characteristics of the one or more discrete objects or events (Sai teaches “modeled representations relate to characteristics of objects with parameters filtered to a light range or light ranges which improve detection of shapes” wherein “by way of example, a modeled representation of pedestrian may include a shape profile for a pedestrian walking towards, away, transverse in either direction (e.g., left to right, and right to left), wherein the model representation may be matched to image data in one or more sizes” in para. [0041] and “the second feature extraction of block 210 includes identifying one or more objects by detecting artifacts in the image data associated with thermal characteristics in the image data” as shown in para. [0042] and FIG. 2. Here, the modeled representation is interpreted as equivalent to the claimed templates that define respective characteristics). Chang and Sai are both considered to be analogous to the claimed invention because they are in the same field of identifying objects through video analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Chang to incorporate the teachings of Sai and include “matching one or more respective templates that define respective characteristics of the one or more discrete objects or events”. The motivation for doing so would have been that “modeled representations relate to characteristics of objects with parameters filtered to a light range or light ranges which improve detection of shapes”, as suggested by Sai in para. [0041]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Chang with Sai to obtain the invention specified in claim 3. Regarding claim 4, Chang and Sai teach the system of claim 3, wherein the instructions that, when executed by the one or more processors, further cause the system to: obtain a missed detection of the one or more discrete objects or events (Sai teaches identifying incorrectly identified objects in para. [0058]. Here, a correct detection of the object is “missed”, which is interpreted as equivalent to the claimed “missed detection of the one or more discrete objects); and based on the missed detection, adjust a criteria of detecting the one or more discrete objects or events based on features of the one or more discrete objects or events (Sai teaches that “module 540 receives joint features 535 and modifies the parameters employed by RGB feature extractor, such as thermal modeled parameters. Correction module 550 is configured to provide feature extraction module 515 with updates to modify the parameters of RGB feature extractor 520 when for areas of objects identified by RGB feature extractor 520 which are incorrectly identified” in para. [0058]. Here, the parameters are used in the modeled representations (templates) as shown in para. [0041]). Similar motivation as applied to claim 3 can be applied here to claim 4. Regarding claim 7, Chang and Sai teach the system of claim 4, wherein the adjusting of the criteria comprises adjusting the one or more respective templates (Sai teaches that “module 540 receives joint features 535 and modifies the parameters employed by RGB feature extractor, such as thermal modeled parameters. Correction module 550 is configured to provide feature extraction module 515 with updates to modify the parameters of RGB feature extractor 520 when for areas of objects identified by RGB feature extractor 520 which are incorrectly identified” in para. [0058]. Here, the parameters are used in the modeled representations (templates) as shown in para. [0041]). Similar motivation as applied to claim 3 can be applied here to claim 7. Regarding claim 8, Chang teaches the system of claim 1, wherein the detecting one or more events associated with the tracked object comprises: detecting changes in the environment; identifying candidate events based on the detected changes (Chang teaches “the tracking camera video feed that was processed to detect and track objects may be further processed at a step 4410 by using spatiotemporal pattern recognition (such as machine-based spatiotemporal pattern recognition as described throughout this disclosure) to identify one or more events, which may be a wide range of events as described throughout this disclosure, such as events that correspond to patterns in a game or sport” in para. [0251]); and comparing the candidate events (Chang teaches that, “the tasks of object detection and recognition may be performed on the basis of knowledge of known calibration parameters of the cameras in the tracking system and known properties of the objects being detected such as their size, orientation, or positions etc”, wherein “perspectives and distortions introduced by the cameras can be undone by applying a transformation such that the objects being detected may have a consistent scale and orientation in transformed images” and “the transformed images may be used as inputs to detection and recognition algorithms by image processing devices” as shown in para. [0434]). Chang fails to teach specifically comparing the candidate events with templates and accounting for differences specifically between the object and corresponding objects in the template. However, Sai teaches comparing the candidate (Sai teaches comparing model representations of objects with images to detect objects, “wherein the model representation may be matched to image data in one or more sizes” in para. [0041]. Here, the model representations are interpreted as equivalent to the claimed template. Additionally, since the model representation may be matched to the image data in one or more sizes, it is inherent that the model representations account for scaling). As such, Chang’s teaching of accounting for scaling and orientation distortions caused by the camera to perform more accurate object/event detection can be combined with Sai’s teaching of comparing objects with templates while accounting for a scaling difference between the object and corresponding objects in the templates to teach the above limitation of claim 8. Chang and Sai are both considered to be analogous to the claimed invention because they are in the same field of identifying objects through video analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Chang to incorporate the teachings of Sai and include “comparing the candidate events with templates and accounting for differences specifically between the object and corresponding objects in the template”. The motivation for doing so would have been that “modeled representations relate to characteristics of objects with parameters filtered to a light range or light ranges which improve detection of shapes”, as suggested by Sai in para. [0041]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Chang with Sai to obtain the invention specified in claim 8. Regarding claim 13, Chang teaches the method of claim 11, wherein the generating of the media representation comprises automatically tagging the one or more discrete objects or events based on a confidence level of the one or more discrete objects or events matching one or more respective (Chang teaches “machine learning algorithms are designed to output a measure of confidence. […] If an example is labeled by the machine and has confidence above the threshold, the event goes into the canonical event datastore 210 and nothing further is done” in para. [0177]). Chang fails to teach matching one or more respective templates that define respective characteristics of the one or more discrete objects or events. However, Sai teaches matching one or more respective templates that define respective characteristics of the one or more discrete objects or events (Sai teaches “modeled representations relate to characteristics of objects with parameters filtered to a light range or light ranges which improve detection of shapes” wherein “by way of example, a modeled representation of pedestrian may include a shape profile for a pedestrian walking towards, away, transverse in either direction (e.g., left to right, and right to left), wherein the model representation may be matched to image data in one or more sizes” in para. [0041] and “the second feature extraction of block 210 includes identifying one or more objects by detecting artifacts in the image data associated with thermal characteristics in the image data” as shown in para. [0042] and FIG. 2. Here, the modeled representation is interpreted as equivalent to the claimed templates that define respective characteristics). Chang and Sai are both considered to be analogous to the claimed invention because they are in the same field of identifying objects through video analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Chang to incorporate the teachings of Sai and include “matching one or more respective templates that define respective characteristics of the one or more discrete objects or events”. The motivation for doing so would have been that “modeled representations relate to characteristics of objects with parameters filtered to a light range or light ranges which improve detection of shapes”, as suggested by Sai in para. [0041]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Chang with Sai to obtain the invention specified in claim 13. Regarding claim 14, Chang and Sai teach the method of claim 13, further comprising: obtaining a missed detection of the one or more discrete objects or events (Sai teaches identifying incorrectly identified objects in para. [0058]. Here, a correct detection of the object is “missed”, which is interpreted as equivalent to the claimed “missed detection of the one or more discrete objects); and based on the missed detection, adjusting a criteria of detecting the one or more discrete objects or events based on features of the one or more discrete objects or events (Sai teaches that “module 540 receives joint features 535 and modifies the parameters employed by RGB feature extractor, such as thermal modeled parameters. Correction module 550 is configured to provide feature extraction module 515 with updates to modify the parameters of RGB feature extractor 520 when for areas of objects identified by RGB feature extractor 520 which are incorrectly identified” in para. [0058]. Here, the parameters are used in the modeled representations (templates) as shown in para. [0041]). Similar motivation as applied to claim 13 can be applied here to claim 14. Regarding claim 17, Chang and Sai teach the method of claim 14, wherein the adjusting of the criteria comprises adjusting the one or more respective templates (Sai teaches that “module 540 receives joint features 535 and modifies the parameters employed by RGB feature extractor, such as thermal modeled parameters. Correction module 550 is configured to provide feature extraction module 515 with updates to modify the parameters of RGB feature extractor 520 when for areas of objects identified by RGB feature extractor 520 which are incorrectly identified” in para. [0058]. Here, the parameters are used in the modeled representations (templates) as shown in para. [0041]). Similar motivation as applied to claim 13 can be applied here to claim 17. Regarding claim 18, Chang teaches the method of claim 11, wherein the detecting one or more events associated with the tracked object comprises: detecting changes in the environment; identifying candidate events based on the detected changes (Chang teaches “the tracking camera video feed that was processed to detect and track objects may be further processed at a step 4410 by using spatiotemporal pattern recognition (such as machine-based spatiotemporal pattern recognition as described throughout this disclosure) to identify one or more events, which may be a wide range of events as described throughout this disclosure, such as events that correspond to patterns in a game or sport” in para. [0251]); and comparing the candidate events (Chang teaches that, “the tasks of object detection and recognition may be performed on the basis of knowledge of known calibration parameters of the cameras in the tracking system and known properties of the objects being detected such as their size, orientation, or positions etc”, wherein “perspectives and distortions introduced by the cameras can be undone by applying a transformation such that the objects being detected may have a consistent scale and orientation in transformed images” and “the transformed images may be used as inputs to detection and recognition algorithms by image processing devices” as shown in para. [0434]). Chang fails to teach specifically comparing the candidate events with templates and accounting for differences specifically between the object and corresponding objects in the template. However, Sai teaches comparing the candidate (Sai teaches comparing model representations of objects with images to detect objects, “wherein the model representation may be matched to image data in one or more sizes” in para. [0041]. Here, the model representations are interpreted as equivalent to the claimed template. Additionally, since the model representation may be matched to the image data in one or more sizes, it is inherent that the model representations account for scaling). As such, Chang’s teaching of accounting for scaling and orientation distortions caused by the camera to perform more accurate object/event detection can be combined with Sai’s teaching of comparing objects with templates while accounting for a scaling difference between the object and corresponding objects in the templates to teach the above limitation of claim 8. Chang and Sai are both considered to be analogous to the claimed invention because they are in the same field of identifying objects through video analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Chang to incorporate the teachings of Sai and include “comparing the candidate events with templates and accounting for differences specifically between the object and corresponding objects in the template”. The motivation for doing so would have been that “modeled representations relate to characteristics of objects with parameters filtered to a light range or light ranges which improve detection of shapes”, as suggested by Sai in para. [0041]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Chang with Sai to obtain the invention specified in claim 18. Claims 6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Chang et al. (U.S. Publication No. 2019/0205651 A1), hereinafter Chang in view of Vendas Da Costa et al. (U.S. Publication No. 2021/0366097 A1), hereinafter Vendas Da Costa. Regarding claim 6, Chang teaches the system of claim 1. Chang fails to teach wherein the memory stored instructions that, when executed by the one or more processors, further causes the system to: identify an operational status of the one or more discrete objects or events; and overlay an indication of the operational status of the one or more discrete objects or events. However, Vendas Da Costa teaches wherein the memory stored instructions that, when executed by the one or more processors, further causes the system to: identify an operational status of the one or more discrete objects or events; and overlay an indication of the operational status of the one or more discrete objects or events (Vendas Da Costa teaches that “an important feature of the planning module is the input and presentation of bathymetry information 32 through 3D visualization. As seen on the Navigation Interface, waypoints 33 and checkpoints 34 are superimposed onto the video feed. These elements may be identified, for example, by number, and/or by distance from a reference point. In other words, in addition to superimposing the technical specifications and status information 30 for the ROV 1 or other relevant structures, the Navigation Interface also provides GPS-determined positions for navigation and pilot information” as shown in para. [0105]. See Fig. 6A and para. [0103]-[0108]). Chang and Vendas Da Costa are both considered to be analogous to the claimed invention because they are in the same field of analyzing an environment and using a user interface to display relevant details regarding the objects/events in a video. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Chang to incorporate the teachings of Vendas Da Costa and include “wherein the memory stored instructions that, when executed by the one or more processors, further causes the system to: identify an operational status of the one or more discrete objects or events; and overlay an indication of the operational status of the one or more discrete objects or events”. The motivation for doing so would have been to “provide immersive visualization of ROV's operation” and “enable[] fast analysis and a comprehensive understanding of operations”, as suggested by Vendas Da Costa in para. [0103] and para. [0112], respectively. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Chang with Vendas Da Costa to obtain the invention specified in claim 6. Regarding claim 16, Chang teaches the method of claim 11. Chang fails to teach identifying an operational status of the one or more discrete objects or events; and overlaying an indication of the operational status of the one or more discrete objects or events. However, Vendas Da Costa teaches identifying an operational status of the one or more discrete objects or events; and overlaying an indication of the operational status of the one or more discrete objects or events (Vendas Da Costa teaches that “an important feature of the planning module is the input and presentation of bathymetry information 32 through 3D visualization. As seen on the Navigation Interface, waypoints 33 and checkpoints 34 are superimposed onto the video feed. These elements may be identified, for example, by number, and/or by distance from a reference point. In other words, in addition to superimposing the technical specifications and status information 30 for the ROV 1 or other relevant structures, the Navigation Interface also provides GPS-determined positions for navigation and pilot information” as shown in para. [0105]. See Fig. 6A and para. [0103]-[0108]). Chang and Vendas Da Costa are both considered to be analogous to the claimed invention because they are in the same field of analyzing an environment and using a user interface to display relevant details regarding the objects/events in a video. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Chang to incorporate the teachings of Vendas Da Costa and include “identifying an operational status of the one or more discrete objects or events; and overlaying an indication of the operational status of the one or more discrete objects or events”. The motivation for doing so would have been to “provide immersive visualization of ROV's operation” and “enable[] fast analysis and a comprehensive understanding of operations”, as suggested by Vendas Da Costa in para. [0103] and para. [0112], respectively. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Chang with Vendas Da Costa to obtain the invention specified in claim 16. Claims 9 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Chang et al. (U.S. Publication No. 2019/0205651 A1), hereinafter Chang in view of Bratton et al. (U.S. Publication No. 2013/0076908 A1), hereinafter Bratton. Regarding claim 9, Chang teaches the system of claim 1. Chang fails to teach wherein the instructions further cause the system to: determine a view field of a sensor capturing the content; and display an indication of the determined view field. However, Bratton teaches wherein the instructions further cause the system to: determine a view field of a sensor capturing the content; and display an indication of the determined view field (Bratton teaches that “the orientation of the camera is indicated in text 82 overlaid on the video display. In the illustrated example, this text 82 is generated by the camera. Here, the camera is directed at azimuth +15.68 and at elevation −2.75. These direction indicators may be in degrees or some other increment. The display also includes camera generated text 84 that indicates which camera is the source of the video data, here it is camera 1, as well as a camera generated P/T function 86, which is indicated here as 00” in para. [0059]). Chang and Bratton are both considered to be analogous to the claimed invention because they are in the same field of analyzing an environment and using a user interface to display relevant details regarding the objects/events in a video. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Chang to incorporate the teachings of Bratton and include “wherein the instructions further cause the system to: determine a view field of a sensor capturing the content; and display an indication of the determined view field”. The motivation for doing so would have been “to receive and display multiple video signals for simultaneous and/or serial display on the portable device. Control of the video signals is preferably by manipulation of the touch screen of the portable device” and “enable[] persons who may be away from a docking station to add the capability to view a set of video signals”, as suggested by Bratton in para. [0007] and para. [0008], respectively. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Chang with Bratton to obtain the invention specified in claim 9. Regarding claim 19, Chang teaches the method of claim 11. Chang fails to teach determining a view field of a sensor capturing the content; and displaying an indication of the determined view field. However, Bratton teaches determining a view field of a sensor capturing the content; and display an indication of the determined view field (Bratton teaches that “the orientation of the camera is indicated in text 82 overlaid on the video display. In the illustrated example, this text 82 is generated by the camera. Here, the camera is directed at azimuth +15.68 and at elevation −2.75. These direction indicators may be in degrees or some other increment. The display also includes camera generated text 84 that indicates which camera is the source of the video data, here it is camera 1, as well as a camera generated P/T function 86, which is indicated here as 00” in para. [0059]). Chang and Bratton are both considered to be analogous to the claimed invention because they are in the same field of analyzing an environment and using a user interface to display relevant details regarding the objects/events in a video. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Chang to incorporate the teachings of Bratton and include “determining a view field of a sensor capturing the content; and displaying an indication of the determined view field”. The motivation for doing so would have been “to receive and display multiple video signals for simultaneous and/or serial display on the portable device. Control of the video signals is preferably by manipulation of the touch screen of the portable device” and “enable[] persons who may be away from a docking station to add the capability to view a set of video signals”, as suggested by Bratton in para. [0007] and para. [0008], respectively. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Chang with Bratton to obtain the invention specified in claim 19. Claims 10 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Chang et al. (U.S. Publication No. 2019/0205651 A1), hereinafter Chang in view of Inoue (JP 2001-45409 A, see attached English translation for citations). Regarding claim 10, Chang teaches the system of claim 1. While Chang further teaches augmenting of the graphical representation of the environment with the generated media representation (See claim 1) and displaying particular frames corresponding to flagged objects/events (See para. [0271]-[0277], wherein only events with a high interest level are displayed to the user), Chang fails to teach wherein the augmenting of the graphical representation of the environment with the generated media representation comprises presenting a snapshot of a particular frame corresponding to the one or more detected events or objects, wherein the snapshot corresponds to a flagged event. However, Inoue teaches wherein the augmenting of the graphical representation of the environment with the generated media representation comprises presenting a snapshot of a particular frame corresponding to the one or more detected events or objects, wherein the snapshot corresponds to a flagged event (Inoue teaches “the recording control unit 19 controls the recording and reproduction processing unit 19 to add a still image flag to a corresponding still image in the moving image to be shot, that is, a still image shot at the moment when the shooting of the still image is instructed. The still image flag is identification information indicating that the still image to which the still image flag is added is a still image to be displayed independently” in para. [0035], wherein “the still image to which the information has been added is displayed as a child image of the picture-in-picture image, and in a scene considered to be particularly important at the time of shooting” as shown in para. [0010]). Here, Inoue’s teaching of flagging a frame of a moving image and displaying it to the user as a overlay on the moving image can be combined with Chang’s teaching of flagging particular frames corresponding to the one or more detected events or objects to teach the above limitation of claim 10. Chang and Inoue are both considered to be analogous to the claimed invention because they are in the same field of analyzing an environment and using a user interface to display relevant details regarding the content in a video. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Chang to incorporate the teachings of Inoue and include “wherein the augmenting of the graphical representation of the environment with the generated media representation comprises presenting a snapshot of a particular frame corresponding to the one or more detected events or objects, wherein the snapshot corresponds to a flagged event”. The motivation for doing so would have been that, “in order to solve [the time and effort it takes to select an arbitrary image in the moving image], the present invention adds identification information to a desired still image in a moving image when capturing a moving image, and adds the identification information when reproducing the moving image. The main point is that a still image to which is added is displayed using a part of the display screen”, as suggested by Inoue in para. [0006] and para. [0008]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Chang with Inoue to obtain the invention specified in claim 10. Regarding claim 20, Chang teaches the method of claim 11. While Chang further teaches augmenting of the graphical representation of the environment with the generated media representation (See claim 1) and displaying particular frames corresponding to flagged objects/events (See para. [0271]-[0277], wherein only events with a high interest level are displayed to the user), Chang fails to teach wherein the augmenting of the graphical representation of the environment with the generated media representation comprises presenting a snapshot of a particular frame corresponding to the one or more detected events or objects, wherein the snapshot corresponds to a flagged event. However, Inoue teaches wherein the augmenting of the graphical representation of the environment with the generated media representation comprises presenting a snapshot of a particular frame corresponding to the one or more detected events or objects, wherein the snapshot corresponds to a flagged event (Inoue teaches “the recording control unit 19 controls the recording and reproduction processing unit 19 to add a still image flag to a corresponding still image in the moving image to be shot, that is, a still image shot at the moment when the shooting of the still image is instructed. The still image flag is identification information indicating that the still image to which the still image flag is added is a still image to be displayed independently” in para. [0035], wherein “the still image to which the information has been added is displayed as a child image of the picture-in-picture image, and in a scene considered to be particularly important at the time of shooting” as shown in para. [0010]). Here, Inoue’s teaching of flagging a frame of a moving image and displaying it to the user as a overlay on the moving image can be combined with Chang’s teaching of flagging particular frames corresponding to the one or more detected events or objects to teach the above limitation of claim 20. Chang and Inoue are both considered to be analogous to the claimed invention because they are in the same field of analyzing an environment and using a user interface to display relevant details regarding the content in a video. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Chang to incorporate the teachings of Inoue and include “wherein the augmenting of the graphical representation of the environment with the generated media representation comprises presenting a snapshot of a particular frame corresponding to the one or more detected events or objects, wherein the snapshot corresponds to a flagged event”. The motivation for doing so would have been that, “in order to solve [the time and effort it takes to select an arbitrary image in the moving image], the present invention adds identification information to a desired still image in a moving image when capturing a moving image, and adds the identification information when reproducing the moving image. The main point is that a still image to which is added is displayed using a part of the display screen”, as suggested by Inoue in para. [0006] and para. [0008]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Chang with Inoue to obtain the invention specified in claim 20. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Girgensohn et al. (U.S. Publication No. 2009/0199099 A1) teaches augmenting a video with a snapshot corresponding to a flagged event. Dal Mutto et al. (US 20130343605 A1) teaches systems and methods for tracking human hands using parts based template matching. Ramaswamy et al (U.S. Patent No. 9224060 B1) teaches using different templates for object tracking based on a threshold confidence value. Any inquiry concerning this communication or earlier communications from the examiner should be directed to KYLA G ALLEN whose telephone number is (703)756-5315. The examiner can normally be reached M-F 7:30am - 4:30pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, John Villecco can be reached on (571) 272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Kyla Guan-Ping Tiao Allen/ Examiner, Art Unit 2661 /AARON W CARTER/Primary Examiner, Art Unit 2661
Read full office action

Prosecution Timeline

Jan 09, 2025
Application Filed
Sep 14, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743768
ADAPTIVE AND ROBUST DETECTION OF LARGE DEFECTS AND IMAGE MISALIGNMENT
2y 7m to grant Granted Sep 22, 2026
Patent 12738022
METHOD, APPARATUS, READABLE MEDIUM AND ELECTRONIC DEVICE OF KEY-VALUE MATCHING
2y 3m to grant Granted Sep 15, 2026
Patent 12711732
IMAGE PROCESSING DEVICE AND IMAGE PROCESSING METHOD
2y 8m to grant Granted Aug 18, 2026
Patent 12705787
DIFFERENTIABLE MAPS FOR LANDMARK LOCALIZATION
3y 1m to grant Granted Aug 11, 2026
Patent 12705890
MEMORY-BASED VIDEO OBJECT SEGMENTATION
3y 0m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
89%
Grant Probability
99%
With Interview (+16.7%)
2y 10m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 73 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month