Prosecution Insights
Last updated: October 02, 2026
Application No. 18/897,091

METHOD, APPARATUS AND STORAGE MEDIUM FOR VEHICLE FINDING

Final Rejection §103
Filed
Sep 26, 2024
Priority
Sep 28, 2023 — CN 202311281016.X +1 more
Examiner
ZEWEDE, ASTEWAYE GETTU
Art Unit
2481
Tech Center
2400 — Computer Networks
Assignee
Volvo Group
OA Round
4 (Final)
81%
Grant Probability
Favorable
5-6
OA Rounds
4m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 81% — above average
81%
Career Allowance Rate
47 granted / 58 resolved
+23.0% vs TC avg
Strong +38% interview lift
Without
With
+37.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 4m
Avg Prosecution
11 currently pending
Career history
77
Total Applications
across all art units

Statute-Specific Performance

§101
2.2%
-37.8% vs TC avg
§103
69.7%
+29.7% vs TC avg
§102
11.0%
-29.0% vs TC avg
§112
6.6%
-33.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 58 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of Claims 2. This Office Action is in response to the amendment filed on 07/23/2026. Claims 1,4,7,14, and 16 have been amended, claims 9,12, and 19-20 are cancelled. Thus, claims 1-8, 10-11, and 13-18 are pending for examination. Information Disclosure Statement 3. The information disclosure statement (IDS) submitted on 08/20/2025 is in complies with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Priority Receipt is acknowledged of certified copies of papers submitted under 35 U.S.C. 119(a)- (d), which papers have been placed of record in the file. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness 6. Claims 1-3, 6-7, 14-15, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over LI, Xu-hui (CN-116434589-A) (translation provided and citation given from the translated document) hereinafter “Li” in view of Sun et al. (US-12154431-B2) hereinafter “Sun” in view of Dolgov et al (US-20190019349-A1) hereinafter “Dolgov” further in view of Yamamoto et al (US 20190019349 A1) hereinafter “Yamamoto”. Regarding Claim 1 Li-Sun-Dolgov-Yamamoto Li discloses (Currently Amended) 1. A method for vehicle finding (Li, [0005] “an intelligent vehicle finding method”), comprising: reading video data . . . recorded by a vehicle camera in response to a vehicle finding request from a user; (Li, [0031] “Receive the car search request, read the picture data of the external device camera in real time, and obtain real-time picture information;”) extracting feature point data for a plurality of video frames of…;(Li, [0036] “multiple video frames in the video” Claim 5 “Determine whether there is a specific sign in the real-time screen information; and the specific sign is the signboard on the road or the garage location number; … extract the coordinates in the identification information; The coordinates in the identification information are determined as the matching position of the current position of the external device camera,,,” i.e., real-time screen information refer to the live camera feed (video data)) and matching the feature point data for a plurality of video frames ,(Li, [0036] “multiple video frames in the video”) with a pre-acquired map of a parking environment to obtain coordinates of vehicle in a vector map. (Li, [0021] “…, the current position of the external device camera is matched on the preset three-dimensional map, and the real-time position of the current position on the three-dimensional map is obtained, …, and the three--dimensional map includes the parking location;”) Li does not explicitly disclose …from a predetermined period of time or a predetermined distance before a vehicle stopped… the video data on a frame-by-frame basis, wherein the frame-by-frame basis is in reverse order from a later video frame when the vehicle stopped to an earlier video frame when the predetermined period of time or the predetermined distance started before the vehicle stopped or is dynamically selected based on a vehicle finding scenario; However, in the same field of endeavor Sun discloses more explicitly the following: …from a predetermined period of time or a predetermined distance before a vehicle stopped…(Sun, Col, 30, lines 8-10 “a video recording the period of time between entering the parking lot and the car stopping at the parking space can be provided ,…”) It would have been obvious to one of ordinary skill in the art to modify Li in view of Sun to utilize video data captured during a predetermined period of time before the vehicle stops, in order to improve vehicle searching efficiency and provide a faster and more accurate determination of the vehicle’s position,, as taught by Sun (Col.3, lines 15-16) Li-Sun does not explicitly disclose the video data on a frame-by-frame basis, wherein the frame-by-frame basis is in reverse order from a later video frame when the vehicle stopped to an earlier video frame when the predetermined period of time or the predetermined distance started before the vehicle stopped or is dynamically selected based on a vehicle finding scenario; However, in the same field of endeavor Dolgov discloses more explicitly the following: the video data….from a later video frame when the vehicle stopped to an earlier video frame when the predetermined period of time or the predetermined distance started before the vehicle or is dynamically selected based on a vehicle finding scenario; (Dolgov, [0136] “…at block 602, the computing system operates by determining that the autonomous vehicle has stopped based on data received from the autonomous vehicle.” [0139] “The computing system may provide images or video (i.e. a plurality of images)…..The video or images provided to the operator may correspond to a predetermined period of time before the review criteria was met. The predetermined period of time may be based on a context in which the one or more review criteria was or were met.” One of ordinary skill in the art would have been motivated to incorporate the teaching of Dolgov with Li-Sun to provide information to a remote assistance operator when remote-assistance triggering criteria are met, thereby enabling the remote-assistance operator to review the data and provide input to the autonomous vehicle (Dolgov, Abs.) Li-Sun-Dolgov does not expressly disclose on a frame-by-frame basis, wherein the frame-by-frame basis is in reverse order… However, in the same field of endeavor Yamamoto discloses more explicitly the following: on a frame-by-frame basis, wherein the frame-by-frame basis is in reverse order... (Yamamoto [0123] “The plurality of frames is frames captured in a reverse direction (past direction) in time series, and the recognition target tracking unit 124 executes processing of tracking the recognition target in the reverse direction in time series.”) One of ordinary skill in the art would have been motivated to incorporate the teaching of Yamamoto with Li-Sun-Dolgov “to improve the performance of the recognizer so that erroneous recognition is not repeated.” (Yamamoto, [0004} Note: The motivation that was utilized in the rejection of claim 1 applies equally as well to claims 2-3, 6,14, 15, and 18. Regarding Claim 2 Li-Sun-Dolgov-Yamamoto Li-Sun-Dolgov-Yamamoto discloses (Original): The method according to claim 1, wherein the matching is performed in a cloud, and the method further comprises sending the coordinates of the vehicle in the vector map from the cloud to the user. (Sun, Col, 28, lines 45-48, “The cloud analyzes the received video frames, obtains useful information (key guidance content), and generates a video (updated to the multimedia guidance content of the key route video) to be delivered to the user.” i.e., performing the matching remotely) Regarding Claim 3 Li-Sun-Dolgov-Yamamoto Li-Sun-Dolgov-Yamamoto discloses 3. (Original): The method according to claim 2, wherein the pre-acquired map of the parking environment is a fusion map that is pre-stored in the cloud including a point cloud base map and a vector map. (Li, Claim 3” The three dimensional map and the preset local map are analyzed to obtain the coincident coordinate information; The coincident coordinate information and the geographical coordinates of the vehicle search origin are compared and analyzed, and the judgment results are obtained.”) Regarding Claim 6 Li-Sun-Dolgov-Yamamoto Li-Sun-Dolgov-Yamamoto discloses (Original): The method according to claim 1, wherein the video data is panoramic video data around the vehicle obtained by a plurality of cameras with different shooting orientations in the vehicle camera. (Sun, Col, 5, lines 36-40 “The in-vehicle acquisition …includes cameras such as a camera of a driving recorder mounted on the vehicle, a surround view camera, a 360-degree panoramic camera around the vehicle, and the like.”) Regarding Claim 7 Li-Sun-Dolgov-Yamamoto Li discloses 7. (Currently Amended) A method for vehicle finding (Li, [0005] “an intelligent vehicle finding method”), comprising: reading video data . . . recorded by a vehicle camera in response to a vehicle finding request from a user side; (Li, [0031] “Receive the car search request, read the picture data of the external device camera in real time, and obtain real-time picture information;) . . . determining priorities for respective recognition results according to preset weights; (Li, [0058] “if the weight coefficient is greater than or equal to the preset weight coefficient, determine the existence of specific marks in the real-time picture information; If the weight factor is less than the preset weight coefficient, it is determined that there is no specific identification in the real-time picture information.”) and Li does not explicitly disclose … a predetermined period of time or a predetermined distance before a vehicle stopped… extracting candidate target objects in respective video frames based on the video data, wherein the video frames are taken in reverse order from a later video frame when the vehicle stopped to an earlier video frame when the predetermined period of time or the predetermined distance started before the vehicle stopped or the order is dynamically selected based on a vehicle finding scenario; performing character and/or symbol recognition on the respective candidate target objects to obtain recognition results; determining a position of the vehicle based on the recognition result with the highest priority. However, in the same field of endeavor Sun discloses more explicitly the following: … a predetermined period of time or a predetermined distance before a vehicle stopped … (Sun, Col, 30, lines 8-10 “a video recording the period of time between entering the parking lot and the car stopping at the parking space can be provided ,…”) ,…”) extracting candidate target objects in respective video frames based on the video data, …;(Sun, Col, 25, lines 36-50 “selecting a plurality of video frames …extract the key information… device may perform a series of processing processes such as extracting key frames, cropping, extracting important indication information,…”) performing character and/or symbol recognition on the respective candidate target objects to obtain recognition results; (Sun, Col, 25, lines 44-50“…extracts the key guidance information from each video frame of the multimedia guidance content through text recognition, and then determines the video frame including the key guidance information as the key frame,…”) determining a position of the vehicle based on the recognition result with the highest priority.(Sun, Col, 25, lines 29-42 “…, when the video frame includes an image having a number of parking space, the electronic device may recognize the number in the video frame, which is the key guidance information, and the video frame is the key frame…..Once the electronic device has recognized the key guidance information, the video frame including the key guidance information is extracted from the multimedia guidance content to obtain the key frame. In some embodiments, the electronic device may further determine a corresponding position of the key frame on the progress bar, and add a progress prompt identifier on the position.”) Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the application to modify the teachings of Li by incorporating the feature disclosed in Sun of extracting candidate target objects in respective video frames based on the video and performing character and/or symbol recognition of those objects. Sun further teaches determining a position of a user's vehicle based on the recognition result with the highest priority. Incorporating these features into Li would have enhanced the Li-Sun system as outlined above. One ordinary skill in the art would have been motivated to incorporate Sun’s future into Li in order to “reducing a time for the target object to obtain the key guidance information, and continuing to improve the efficiency of finding the vehicle.” (Sun, Col, 17, lines 3-5) Li-Sun does not explicitly disclose wherein the video frames are taken in reverse order from a later video frame when the vehicle stopped to an earlier video frame when the predetermined period of time or the predetermined distance started before the vehicle stopped or the order is dynamically selected based on a vehicle finding scenario; However, in the same field Dolgov discloses more explicitly the following: wherein . . . from a later video frame when the vehicle stopped to an earlier video frame when the predetermined period of time or the predetermined distance started before the vehicle stopped or the order is dynamically selected based on a vehicle finding scenario; However, in the same field of endeavor Dolgov discloses more explicitly the following: wherein . . . from a later video frame when the vehicle stopped to an earlier video frame when the predetermined period of time or the predetermined distance started before the vehicle stopped or the order is dynamically selected based on a vehicle finding scenario; (Dolgov, [0136] “…at block 602, the computing system operates by determining that the autonomous vehicle has stopped based on data received from the autonomous vehicle.” [0139] “The computing system may provide images or video (i.e. a plurality of images)…..The video or images provided to the operator may correspond to a predetermined period of time before the review criteria was met. The predetermined period of time may be based on a context in which the one or more review criteria was or were met.” One of ordinary skill in the art would have been motivated to incorporate the teaching of Dolgov with Li-Sun to provide information to a remote assistance operator when remote-assistance triggering criteria are met, thereby enabling the remote-assistance operator to review the data and provide input to the autonomous vehicle (Dolgov, Abs.) Li-Sun-Dolgov does not expressly disclose the video frames are taken in reverse order … However, in the same field of endeavor Yamamoto discloses more explicitly the following: the video frames are taken in reverse order … (Yamamoto [0123] “The plurality of frames is frames captured in a reverse direction (past direction) in time series, and the recognition target tracking unit 124 executes processing of tracking the recognition target in the reverse direction in time series.”) One of ordinary skill in the art would have been motivated to incorporate the teaching of Yamamoto with Li-Sun-Dolgov “to improve the performance of the recognizer so that erroneous recognition is not repeated.” (Yamamoto, [0004]) Regarding Claim 14 Li-Sun-Dolgov-Yamamoto Li discloses 14. (Currently Amended) An apparatus for vehicle finding (Li, [0001] “...an intelligent vehicle finding method, apparatus…) comprising: a camera, recording video data of the vehicle's surroundings . . .; (Li, [0032] “…external device camera, shoots the surrounding pictures through the external device camera…”) a memory (Fig. 4, memory 520,) having stored computer instructions thereon; (Li, [0093] “…the computer device includes a memory and a processor, the memory stores computer-readable instructions…”) and a processor (Fig. 4, “the processor 510”), wherein the instructions, when executed by the processor, cause the processor to perform: (Li, [0020] “…the computer readable storage medium is stored with instructions, when it is running on the computer, so that the computer performs...”) The remaining limitations of independent claim 14 recite features that are substantially similar to those set forth in independent claim 1. Accordingly, the reasoning and analysis provided with respect to claim 1 apply equally to claim 14. Regarding Claim 15 Li-Sun-Dolgov-Yamamoto Li-Sun-Dolgov-Yamamoto discloses 15 (Original): The apparatus according to claim 14, wherein the matching is performed in a cloud, and the instructions further cause the processor to perform sending the coordinates of the vehicle in the vector map from the cloud to the user, (Sun, Col, 28, lines 45-48, “The cloud analyzes the received video frames, obtains useful information (key guidance content), and generates a video (updated to the multimedia guidance content of the key route video) to be delivered to the user.” i.e., performing the matching remotely) and wherein the pre-acquired map of the parking environment is a fusion map that is pre-stored in the cloud including a point cloud base map and a vector map. (Li, claim 3” The three dimensional map and the preset local map are analyzed to obtain the coincident coordinate information; The coincident coordinate information and the geographical coordinates of the vehicle search origin are compared and analyzed, and the judgment results are obtained.”) Regarding Claim 18 Li-Sun-Dolgov-Yamamoto Li-Sun-Dolgov-Yamamoto discloses 18: (Original) The apparatus according to claim 14, wherein the video data is panoramic video data around the vehicle obtained by a plurality of cameras with different shooting orientations in the vehicle camera. (Sun, Col, 5, lines 36-40 “The in-vehicle acquisition …includes cameras such as a camera of a driving recorder mounted on the vehicle, a surround view camera, a 360-degree panoramic camera around the vehicle, and the like.”) Claim Rejections - 35 USC § 103 Claims 4 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Li-Sun-Dolgov-Yamamoto in view of Rublee et al. (US 20240199152 A1) hereinafter “Rublee”. Regarding Claim 4 Li-Sun-Dolgov-Yamamoto-Rublee Li-Sun-Dolgov-Yamamoto discloses 4(Currently Amended): The method according to claim 1, . . . performing screening on the feature point data to exclude feature point data corresponding to video frames with confidence lower than a confidence threshold; (Sun, Col, 19, lines 34-51”the electronic device may also remove redundant information in the parking route video or the parking route image set, for example, selecting a plurality of video frames or images that are more similar than a certain threshold, removing the fuzzier video frames or images, thereby determining the parking route video or the parking route image set on which redundant information removal is performed as the multimedia guidance content. Certainly, the electronic device may also recognize and extract the key information in the parking route video or the parking route image set, and then add an enhancement effect to the key information. The enhanced parking route video or the parking route image set may be used as the multimedia guidance content. That is, the electronic device may perform a series of processing processes such as extracting key frames, cropping, extracting important indication information, and the like from the acquired parking route video or parking route image set,”) and sending the screened feature point data to the cloud for the matching. (Sun, Col, 28, lines 45-48, “The cloud analyzes the received video frames, obtains useful information (key guidance content), and generates a video (updated to the multimedia guidance content of the key route video) to be delivered to the user.” i.e., performing the matching remotely) Li-Sun-Dolgov-Yamamoto do not explicitly disclose wherein the extracting feature point data based on the video data comprises: performing simultaneous localization and mapping (SLAM) modeling on the video data that is in the predetermined period of time or the predetermined distance before the vehicle stopped to extract feature point data; However, in the same field of endeavor Rublee discloses more explicitly the following: wherein the extracting feature point data based on the video data comprises: performing simultaneous localization and mapping (SLAM) modeling on the video data that is in the predetermined period of time or the predetermined distance before the vehicle stopped to extract feature point data; (Rublee, [0091] “where the vehicle moves forward (or backward) a set distance (e.g., 12 inches) and stops.” [0094] “In non-limiting embodiments, a simultaneous localization and mapping (SLAM) algorithm may be utilized to monitor the location of the modular vehicle while forming or adjusting a mapping of an environment in which the vehicle is moving. This can allow to follow a pre-set route automatically and accurately with camera data … The camera data may be generated by a video camera installed on the front side of the modular vehicle and connected to the intra-vehicle network.”) Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the application to modify the teachings of Li-Sun-Dolgov-Yamamoto by incorporating the feature disclosed in Rublee of “performing simultaneous localization and mapping (SLAM) modeling on video data that is in a period of time before the vehicle stops to extract feature point data” in order to enhance the system of Li-Sun-Dolgov-Yamamoto. One of ordinary in the art would have been motivated to incorporate Rublee’s feature into Li-Sun-Dolgov-Yamamoto “to improve the efficiency of vehicle search.”(Li,[0004]). Regarding Claim 16 Li-Sun-Dolgov-Yamamoto-Rublee Li-Sun-Dolgov-Yamamoto discloses 16 (Currently Amended): The apparatus according to claim 14, . . . performing screening on the feature point data to exclude feature point data corresponding to video frames with confidence lower than a confidence threshold; (Sun, Col, 19, lines 34-51”the electronic device may also remove redundant information in the parking route video or the parking route image set, for example, selecting a plurality of video frames or images that are more similar than a certain threshold, removing the fuzzier video frames or images, thereby determining the parking route video or the parking route image set on which redundant information removal is performed as the multimedia guidance content. Certainly, the electronic device may also recognize and extract the key information in the parking route video or the parking route image set, and then add an enhancement effect to the key information. The enhanced parking route video or the parking route image set may be used as the multimedia guidance content. That is, the electronic device may perform a series of processing processes such as extracting key frames, cropping, extracting important indication information, and the like from the acquired parking route video or parking route image set,”) and sending the screened feature point data to the cloud for the matching. (Sun, Col, 28, lines 45-48, “The cloud analyzes the received video frames, obtains useful information (key guidance content), and generates a video (updated to the multimedia guidance content of the key route video) to be delivered to the user.” i.e., performing the matching remotely) Li-Sun-Dolgov-Yamamoto do not explicitly disclose wherein the extracting feature point data based on the video data comprises: performing simultaneous localization and mapping (SLAM) modeling on the video data that is in the predetermined period of time or the predetermined distance before the vehicle stopped to extract feature point data; However, in the same field of endeavor Rublee discloses more explicitly the following: wherein the extracting feature point data based on the video data comprises: performing simultaneous localization and mapping (SLAM) modeling on the video data that is in the predetermined period of time or the predetermined distance before the vehicle stopped to extract feature point data; (Rublee, [0091] “where the vehicle moves forward (or backward) a set distance (e.g., 12 inches) and stops.” [0094] “In non-limiting embodiments, a simultaneous localization and mapping (SLAM) algorithm may be utilized to monitor the location of the modular vehicle while forming or adjusting a mapping of an environment in which the vehicle is moving. This can allow to follow a pre-set route automatically and accurately with camera data … The camera data may be generated by a video camera installed on the front side of the modular vehicle and connected to the intra-vehicle network.”) Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the application to modify the teachings of Li-Sun-Dolgov-Yamamoto by incorporating the feature disclosed in Rublee of “performing simultaneous localization and mapping (SLAM) modeling on video data that is in a period of time before the vehicle stops to extract feature point data” in order to enhance the system of Li-Sun-Dolgov-Yamamoto. One of ordinary in the art would have been motivated to incorporate Rublee’s feature into Li-Sun-Dolgov-Yamamoto “to improve the efficiency of vehicle search.” (Li, [0004]). Claim Rejections - 35 USC § 103 Claim 5 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Li-Sun-Dolgov-Yamamoto-Rublee further in view of ABE; SHINICHIRO (US-20230206491-A1) hereinafter “Abe”. Regarding Claim 5 Li-Sun-Dolgov-Yamamoto-Rublee-Abe Li-Sun-Dolgov-Yamamoto-Rublee discloses 5 (Previously Presented): The method according to claim 4, Li-Sun-Dolgov-Yamamoto-Rublee does not explicitly disclose wherein: in a case where a number of feature points that have been extracted in a current video frame is determined to be more than a number threshold, the extracting of feature points of the current video frame is stopped, the extracted feature points are sent to the cloud, and the extracting of feature points of the next video frame is started; and in a case where a number of all feature points extracted in the current video frame is determined not to be more than the number threshold, all the feature points extracted in the current video frame are excluded, and the extracting of feature point of the next video frame is started. However, in the same field of endeavor Abe discloses more explicitly the following: wherein: in a case where a number of feature points that have been extracted in a current video frame is determined to be more than a number threshold, all the extracting of feature points of the current video frame is stopped, the extracted feature points are sent to the cloud, and the extracting of feature point of the next video frame is started. (Abe, [0332] “The prescribed condition is, for example, a condition such as whether the number of extracted feature points is equal to or larger than a prescribed threshold, and whether an identification level of the extracted feature point is equal to or larger than a prescribed threshold level. The determination processing in step S104 is executed as processing of determining whether or not the feature point extracted in step S103 is a feature point at a level at which a marker region can be reliably extracted from various camera-captured images. [0333] In a case where it is determined in step S104 that the feature point data of the marker region extracted in step S103 satisfies the previously prescribed condition, the process proceeds to step S105.” [ i.e., when the extraction points satisfy the threshold, the process proceeds to the next step S105—which is equivalent to stopping extraction of the current frame and using the results (such as transmitting/storing, which in practical system may include sending to the cloud.]), and in a case where a number of all feature points extracted in the current video frame is determined not to be more than the number threshold, all the feature points extracted in the current video frame are excluded, and the extracting of feature point of the next video frame is started. (Abe, [0332] The prescribed condition is, for example, a condition such as whether the number of extracted feature points is equal to or larger than a prescribed threshold, and whether an identification level of the extracted feature point is equal to or larger than a prescribed threshold level. The determination processing in step S104 is executed as processing of determining whether or not the feature point extracted in step S103 is a feature point at a level at which a marker region can be reliably extracted from various camera-captured images [0334] On the other hand, in a case where it is determined that the feature point data of the marker region extracted in step S103 does not satisfy the previously prescribed condition, the process returns to step S102. [ i.e., when the extraction points do not satisfy the threshold, the process returns to a new frame S102—which is equivalent to excluding the current frame’s points and starting the next frame.]) Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the application to modify the teachings of Li-Sun-Dolgov-Yamamoto-Rublee with Abe in order to incorporate a conditional feature point extraction process. Specifically, the system would stop extracting feature points once a threshold is reached, send the extraction points to the cloud, and then proceed with extraction on the next video frame, exclude feature points if below the threshold as suggested by Abe. The reasoning is that this approach enables “a program capable of tracking a flight path and a destination of a mobile device such as a drone, for example, and a target to be followed with high accuracy.” (Abe, [0001]) Note: The motivation that was utilized in the rejection of claim 5 applies equally as well to claim 17 Regarding Claim 17 Li-Sun-Dolgov-Yamamoto-Rublee-Abe Li-Sun-Dolgov-Yamamoto-Rublee disclose (Previously Presented): The apparatus according to claim 14, Li-Sun-Dolgov-Yamamoto-Rublee does not explicitly disclose wherein: in a case where a number of feature points that have been extracted in a current video frame is determined to be more than a number threshold, the extracting of feature points of the current video frame is stopped, the extracted feature points are sent to the cloud, and the extracting of feature points of the next video frame is started; and in a case where a number of all feature points extracted in the current video frame is determined not to be more than the number threshold, all the feature points extracted in the current video frame are excluded, and the extracting of feature point of the next video frame is started. However, in the same field of endeavor Abe discloses more explicitly the following: wherein: in a case where a number of feature points that have been extracted in a current video frame is determined to be more than a number threshold, the extracting of feature points of the current video frame is stopped, the extracted feature points are sent to the cloud, and the extracting of feature point of the next video frame is started. (Abe, [0332] “The prescribed condition is, for example, a condition such as whether the number of extracted feature points is equal to or larger than a prescribed threshold, and whether an identification level of the extracted feature point is equal to or larger than a prescribed threshold level. The determination processing in step S104 is executed as processing of determining whether or not the feature point extracted in step S103 is a feature point at a level at which a marker region can be reliably extracted from various camera-captured images. [0333] In a case where it is determined in step S104 that the feature point data of the marker region extracted in step S103 satisfies the previously prescribed condition, the process proceeds to step S105.” [ i.e., when the extraction points satisfy the threshold, the process proceeds to the next step S105—which is equivalent to stopping extraction of the current frame and using the results (such as transmitting/storing, which in practical system may include sending to the cloud.]), and in a case where a number of all feature points extracted in the current video frame is determined not to be more than the number threshold, all the feature points extracted in the current video frame are excluded, and the extracting of feature point of the next video frame is started. (Abe, [0332] The prescribed condition is, for example, a condition such as whether the number of extracted feature points is equal to or larger than a prescribed threshold, and whether an identification level of the extracted feature point is equal to or larger than a prescribed threshold level. The determination processing in step S104 is executed as processing of determining whether or not the feature point extracted in step S103 is a feature point at a level at which a marker region can be reliably extracted from various camera-captured images [0334] On the other hand, in a case where it is determined that the feature point data of the marker region extracted in step S103 does not satisfy the previously prescribed condition, the process returns to step S102. [ i.e., when the extraction points do not satisfy the threshold, the process returns to a new frame S102—which is equivalent to excluding the current frame’s points and starting the next frame.]) Claim Rejections - 35 USC § 103 Claims 8 is rejected under 35 U.S.C. 103 as being unpatentable over Li-Sun-Dolgov-Yamamoto in view of Mao et al. (US-20210365707-A1) hereinafter “Mao”. Regarding Claim 8 Li-Sun-Dolgov-Yamamoto-Mao Li-Sun-Dolgov-Yamamoto discloses 8. (Original): The method according to claim 7, Li-Sun-Dolgov-Yamamoto do not explicitly disclose wherein the extracting candidate target objects in respective video frames based on the video data comprises: performing target detection on the respective video frames to generate one or more prediction boxes containing candidate target objects; determining confidences for the one or more prediction boxes; determining a video frame in which the highest confidence for the prediction boxes is more than a threshold as a valid video frame; and extracting the candidate target objects in the prediction boxes in respective valid video frames. However, in the same field of endeavor Mao discloses more explicitly the following: wherein the extracting candidate target objects in respective video frames based on the video data (Mao, [0246] “At block 1344, the process 1340 performs visual processing to process the video data … to detect one or more candidate target objects.” comprises: performing target detection on the respective video frames to generate one or more prediction boxes containing candidate target objects; (Mao, [0194] “the process 930 can include generating a bounding box for the target object or ROI in the initial frame in the sequence of frames (e.g., video). For instance, an ROI can be determined for the object, and the bounding box can be generated to represent the ROI.”) determining confidences for the one or more prediction boxes; (Mao, [0336] “A confidence score is provided that indicates how certain it is that the predicted bounding box actually encloses an object.….”) determining a video frame in which the highest confidence for the prediction boxes is more than a threshold as a valid video frame; (Mao. [0337] “the confidence score for a bounding box and the class prediction are combined into a final score that indicates the probability that that bounding box contains a specific type of object…Many of the bounding boxes will have very low scores, in which case only the boxes with a final score above a threshold (e.g., above a 30% probability, 40% probability, 50% probability, or other suitable threshold) are kept. FIG. 46C shows an image with the final predicted bounding boxes and classes, including a dog, a bicycle, and a car. As shown, from the 4645 total bounding boxes that were generated, only the three bounding boxes shown in FIG. 46C were kept because they had the best final scores”) and extracting the candidate target objects in the prediction boxes in respective valid video frames. (Mao, [0337] “The confidence score for a bounding box and the class prediction are combined into a final score that indicates the probability that that bounding box contains a specific type of object.”) Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the application to modify the teachings of Li-Sun-Dolgov-Yamamoto with Mao to create the system of Li-Sun-Dolgov-Yamamoto as outlined above in order to perform target detection on the respective video frames to generate one or more prediction boxes containing candidate target objects; determine confidences for the one or more prediction boxes; determine a video frame in which the highest confidence for the prediction boxes is more than a threshold as a valid video frame; and extract the candidate target objects in the prediction boxes in respective valid video frames. as suggested by Mao. One ordinary in the art would have been motivated to incorporate Mao’s feature with Li-Sun-Dolgov-Yamamoto to “detect and track the object in one or more frames of the sequence of frames.” (Mao, [0225]) Regarding Claim 9 (Canceled) Claim Rejections - 35 USC § 103 Claims 10-11, and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Li-Sun-Dolgov-Yamamoto in view of Haider et al. (US 10990826 B1) hereinafter “Haider”. Regarding Claim 10 Li-Sun-Dolgov-Yamamoto-Haider Li-Sun-Dolgov-Yamamoto discloses 10 (Previously Presented): The method according to claim 7, Li-Sun-Dolgov-Yamamoto do not explicitly disclose wherein the extracting the candidate target objects in the prediction boxes in respective valid video frames comprises: determining a prediction box with the highest confidence in the valid video frame as a valid prediction box; and selecting a subset of valid prediction boxes and extract candidate target objects therefrom. However, in the same field of endeavor Haider discloses more explicitly the following: wherein the extracting the candidate target objects in the prediction boxes in respective valid video frames comprises: determining a prediction box with the highest confidence in the valid video frame as a valid prediction box; (Haider, Col, 7 lines 63-67 “NMS 208 is capable of selecting one of the object detections in the group based on the highest confidence score. The object detections of each group that are not selected for the frame being processed are suppressed (e.g., discarded).;”) selecting a subset of valid prediction boxes and extract candidate target objects therefrom. (Haider, Col, 8, lines 11-15, “Pruner 210 is capable of processing the remaining detections in a given frame…Pruner 210 is capable of removing those object detections that have been suppressed by NMS 208 from the data set for the current frame being processed….) Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the application to modify the teachings of Li-Sun-Dolgov-Yamamoto with Haider to create the system of Li-Sun-Dolgov-Yamamoto as outlined above. In particular, Haider suggests that “extracting the candidate target objects in the prediction boxes in respective valid video frames comprises: determining a prediction box with the highest confidence in the valid video frame as a valid prediction box and selecting a subset of valid prediction boxes and extract candidate target objects therefrom.” This modification improves “accuracy and real-time performance of object detection.” (Haider, Col, 2, lines 63-64.) Note: The motivation that was utilized in the rejection of claim 10 applies equally as well to claims 11 and 13. Regarding Claim 11 Li-Sun-Dolgov-Yamamoto-Haider Li-Sun-Dolgov-Yamamoto-Mao-Haider disclose 11 (Original): The method according to claim 10, wherein the preset weights include at least one or more of: weights allocated sequentially based on timestamp of video frame corresponding to recognition result, weights allocated based on confidence for valid prediction box associated with recognition result, and weights allocated based on repetition frequency of recognition result among all recognition results. (Haider, Col, 12, lines, 58-67“the system is capable of propagating object detections in the selected frames to adjacent frames bidirectionally …object detections are propagated from each selected frame to a location in the adjacent, or neighboring, frames based upon the vector flow data for each of the object detections.”) Regarding Claim 12 (Canceled) Regarding Claim 13 Li-Sun-Dolgov-Yamamoto-Haider Li-Sun-Dolgov-Yamamoto-Haider discloses 13 (Original): The method according to claim 10, further comprising processing, the determined position of a user's vehicle into a standardized format and transmitting to the user side. (Li, [0043] “The external device…, obtains the location information of the external device, the location information includes coordinates, time and object, and sends the location information of the external device to the external device, The external device receives the location information of the external device, and determines the coordinates in the location information as the geographical coordinates of the vehicle seeker, and cooperates with the base station information…satellite positioning of the external device to …,”) Regarding Claim 19-20 (Canceled) Pertinent Prior Art The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. OTSUKI et al. (US-20240280371-A1) Liu et al. (US-20210229292-A1) Kim et al. (US-20260077760-A1) Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ASTEWAYE GETTU ZEWEDE whose telephone number is (703)756-1441. The examiner can normally be reached Mo-Fr 8:30 am to 5:30 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, William Vaughn can be reached at (571)272-3922. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ASTEWAYE GETTU ZEWEDE/Examiner, Art Unit 2481 /WILLIAM C VAUGHN JR/Supervisory Patent Examiner, Art Unit 2481
Read full office action

Prosecution Timeline

Show 3 earlier events
Dec 01, 2025
Response Filed
Feb 13, 2026
Final Rejection mailed — §103
Mar 09, 2026
Response after Non-Final Action
Mar 31, 2026
Request for Continued Examination
Apr 08, 2026
Response after Non-Final Action
Apr 28, 2026
Non-Final Rejection mailed — §103
Jul 23, 2026
Response Filed
Sep 01, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12739411
NETWORK DEVICE AND ERROR HANDLING
4y 8m to grant Granted Sep 15, 2026
Patent 12732700
MULTIPLE POSITION ROLLING SHUTTER IMAGING DEVICE
4y 2m to grant Granted Sep 08, 2026
Patent 12720011
DUPLICATE FRAME DETECTION IN MULTI-CAMERA VIEWS FOR AUTONOMOUS SYSTEMS AND APPLICATIONS
3y 5m to grant Granted Aug 25, 2026
Patent 12718590
INFORMATION PROCESSING DEVICE
2y 5m to grant Granted Aug 25, 2026
Patent 12701269
IMAGE/VIDEO ENCODING/DECODING METHOD AND DEVICE
1y 5m to grant Granted Aug 04, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
81%
Grant Probability
99%
With Interview (+37.9%)
2y 4m (~4m remaining)
Median Time to Grant
High
PTA Risk
Based on 58 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month