Prosecution Insights
Last updated: August 17, 2026
Application No. 18/934,231

SYSTEM AND METHOD FOR DETECTING OBSTACLES AND ALERTING USERS USING ARTIFICIAL INTELLIGENCE IN MACHINE VISION CAMERAS

Non-Final OA §101§103§112
Filed
Oct 31, 2024
Examiner
BAYNES, SAMUEL DAVID
Art Unit
2665
Tech Center
2600 — Communications
Assignee
Zebra Technologies Corporation
OA Round
1 (Non-Final)
86%
Grant Probability
Favorable
1-2
OA Rounds
8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 86% — above average
86%
Career Allowance Rate
6 granted / 7 resolved
+23.7% vs TC avg
Strong +25% interview lift
Without
With
+25.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
12 currently pending
Career history
20
Total Applications
across all art units

Statute-Specific Performance

§101
11.3%
-28.7% vs TC avg
§103
53.5%
+13.5% vs TC avg
§102
9.9%
-30.1% vs TC avg
§112
22.5%
-17.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 7 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement No information disclosure statement(s) (IDS) has been received, nor considered, by the Examiner at the time of the writing of the instant non-final office action. Specification The abstract of the disclosure is objected to because it contains legal phraseology, i.e. “said data,” found on line 3. A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b). Applicant is reminded of the proper language and format for an abstract of the disclosure. The abstract should be in narrative form and generally limited to a single paragraph on a separate sheet within the range of 50 to 150 words in length. The abstract should describe the disclosure sufficiently to assist readers in deciding whether there is a need for consulting the full patent text for details. The language should be clear and concise and should not repeat information given in the title. It should avoid using phrases which can be implied, such as, “The disclosure concerns,” “The disclosure defined by this invention,” “The disclosure describes,” etc. In addition, the form and legal phraseology often used in patent claims, such as “means” and “said,” should be avoided. Claim Objections Claims 5-7, and 14 are objected to because of the following informalities: Claim 5 line 1, “where the obstacle” should read “wherein the obstacle”. Claims 5-7 are also objected to because inconsistent terminology is used to describe the information provided by the user. Claims 5 and 14 recite “input of user data” and later “user data input,” while claims 6 and 7 refer to “the user input data.” Regarding claims 5 and 14, it is unclear whether “user data input” refers to the previously recited “user data” or to the act of inputting the user data. Claims 6 and 7 further recite “the user input data,” although the term is not expressly introduced in claim 5. Although the intended meaning is reasonably clear in view of the specification, applicant is required to amend the claim to use consistent terminology for clarity and provide clear antecedent basis. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 3 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 3, lines 4-6, recites the limitation “providing the alert data on the environmental data to the trained generative Al model and generating the textual description of the detected obstacle to include a description derived from the environmental data.” Specifically, line 4 of claim 3 recites “providing the alert data on the environmental data to the trained generative AI model.” Although “the alert data” and “the environmental data” each have antecedent basis, claims 1 and 3 do not establish the recited relationship between the alert data and the environmental data. In particular, it is unclear whether “on” requires the alert data to be based on, associated with, included in, overlaid upon, or otherwise combined with the environmental data. The claim does not recite an operation that creates or establishes such a relationship before the combined phrase is used. Consequentially, it is further unclear how “the alert data on the environmental data” is provided to the trained generative AI model for subsequential generation of textual descriptions derived from the environmental data because it is unclear what the structure and/or meaning of “the alert data on the environment data” constitutes. Also, it is unclear how the trained generative AI model can both be (1) “provided the alert data on the environmental data” for processing, while also, the trained generative AI model’s results are (2) “derived from the environmental data”. As currently written, it’s unclear if the trained generative AI’s model is using “the alert data,” “the alert data on the environmental data” (which is further surrounded by unclarity, detailed above), “the environmental data,” or some type of combination of alert data and environmental data to generate the textual description of the detected obstacle. Accordingly, claim 3 is indefinite because the metes and bounds of the claimed step cannot be determined with reasonable certainty and rejected under 35 U.S.C. 112(b). Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim(s) 1-3, 9-12, and 17-19 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The limitations, under their broadest reasonable interpretation, cover mental process (concept performed in a human mind, including as observation, evaluation, judgment, opinion, organizing human activity and mathematical concepts and calculations). The claim(s) recite(s) steps for detecting an obstacle using an imaging device (i.e. observation, evaluation, judgment, opinion). This judicial exception is not integrated into a practical application because the steps do not add meaningful limitations to be considered specifically applied to a particular technological problem to be solved. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the steps of the claimed invention can be done mentally and no additional features in the claims would preclude them from being performed as such except for the generic computer elements at high level of generality (e.g. processor, memory, operating system, etc.). According to the USPTO guidelines, a claim is directed to non-statutory subject matter if: STEP 1: the claim does not fall within one of the four statutory categories of invention (process, machine, manufacture or composition of matter), or STEP 2: the claim recites a judicial exception, e.g. an abstract idea, without reciting additional elements that amount to significantly more than the judicial exception as determined using the following analysis: STEP 2A (PRONG 1): Does the claim recite an abstract idea, law of nature, or natural phenomenon? STEP 2A (PRONG 2): Does the claim recite additional elements that integrate the judicial exception into a practical application? STEP 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? Using the two-step inquiry, it is clear that independent claims 1 and 12 are directed to an abstract idea as shown below: ► STEP 1: Do the claims fall within one of the statutory categories? YES. Claims 1-3 and 9-11 are directed to a method and claims 12 and 17-19 are directed to an imaging device. ► STEP 2A (PRONG 1): Is the claim directed to a law of nature, a natural phenomenon or an abstract idea? YES, the claims are directed toward a mental process and/or mathematical concepts (i.e. abstract idea). With regard to STEP 2A (PRONG 1), the guidelines provide three groupings of subject matter that are considered abstract ideas: Mathematical concepts - mathematical relationships, mathematical formulas or equations, mathematical calculations; Certain methods of organizing human activity - fundamental economic principles or practices (including hedging, insurance, mitigating risk); commercial or legal interactions (including agreements in the form of contracts; legal obligations; advertising, marketing or sales activities or behaviors; business relations); managing personal behavior or relationships or interactions between people (including social activities, teaching, and following rules or instructions); and Mental processes - concepts that are practicably performed in the human mind (including an observation, evaluation, judgment, opinion). Independent claim(s) 1 and 12 comprise a mental process that can be practicably performed in the human mind (or generic computers or components configured to perform the method) and, therefore, an abstract idea. Regarding independent claim(s) 1 and 12, under step 2A prong 1, claims 1 recite the mental steps: “analyzing, using the obstacle detection model, the image data to identify an obstacle in a field of view (FOV) of the imaging device” and “responsive to identifying the obstacle in the FOV, generating alert data at the imaging device, providing the alert data to a trained generative artificial intelligence (Al) model, at the imaging device, the generative Al model trained to generate a textual description of the alert data.” One of ordinary skill in the art could reasonably analyze image data and identify an obstacle in a field of view of an imaging device (i.e. observe and evaluate), generate data relating to an alert based on the identify (i.e. further evaluate), and generate a textual description of the alert data either mentally or using a pen and paper. The mere nominal recitation that the various steps are being executed by a device/in a device (e.g. imaging device) and the use of a trained generative artificial intelligence (Al) model does not take the limitations out of the mental process grouping. Thus, the claims recite a mental process. These limitations, as drafted, is a simple process that, under their broadest reasonable interpretation, covers performance of the limitations in the mind or by a human. The Examiner notes that under MPEP 2106.04(a)(2)(III), the courts consider a mental process (thinking) that "can be performed in the human mind, or by a human using a pen and paper" to be an abstract idea. CyberSource Corp. v. Retail Decisions, Inc., 654 F.3d 1366, 1372, 99 USPQ2d 1690, 1695 (Fed. Cir. 2011). As the Federal Circuit explained, "methods which can be performed mentally, or which are the equivalent of human mental work, are unpatentable abstract ideas the 'basic tools of scientific and technological work' that are open to all."' 654 F.3d at 1371, 99 USPQ2d at 1694 (citing Gottschalk v. Benson, 409 U.S. 63, 175 USPQ 673 (1972)). See also Mayo Collaborative Servs. v. Prometheus Labs. Inc., 566 U.S. 66, 71, 101 USPQ2d 1961, 1965 ('"[M]ental processes[] and abstract intellectual concepts are not patentable, as they are the basic tools of scientific and technological work'" (quoting Benson, 409 U.S. at 67, 175 USPQ at 675)); Parker v. Flook, 437 U.S. 584,589, 198 USPQ 193, 197 (1978) (same). STEP 2A (PRONG 2): Does the claim recite additional elements that integrate the judicial exception into a practical application? NO, the claims do not recite additional elements that integrate the judicial exception into a practical application. With regard to STEP 2A (prong 2), whether the claim recites additional elements that integrate the judicial exception into a practical application, the guidelines provide the following exemplary considerations that are indicative that an additional element (or combination of elements) may have integrated the judicial exception into a practical application: an additional element reflects an improvement in the functioning of a computer, or an improvement to other technology or technical field; an additional element that applies or uses a judicial exception to affect a particular treatment or prophylaxis for a disease or medical condition; an additional element implements a judicial exception with, or uses a judicial exception in conjunction with, a particular machine or manufacture that is integral to the claim; an additional element effects a transformation or reduction of a particular article to a different state or thing; and an additional element applies or uses the judicial exception in some other meaningful way beyond generally linking the use of the judicial exception to a particular technological environment, such that the claim as a whole is more than a drafting effort designed to monopolize the exception. While the guidelines further state that the exemplary considerations are not an exhaustive list and that there may be other examples of integrating the exception into a practical application, the guidelines also list examples in which a judicial exception has not been integrated into a practical application: an additional element merely recites the words "apply it" (or an equivalent) with the judicial exception, or merely includes instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea; an additional element adds insignificant extra-solution activity to the judicial exception; an additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use. Independent claim(s) 1 and 12 disclose trained generative artificial intelligence (Al) model, processors, imaging sensors, and memories storing computer instructions, which are generic computer components and/or insignificant pre/post-solution extra activity that do not add a meaningful limitation to the abstract idea because they amount to simply implementing the abstract idea in the apparatus claim. See MPEP 2106.05(g). These components are recited at a broad level such that it amounts to “apply it” of generic computing processing components. Additionally, the claims recite data gathering and communication steps of “capturing, at the imaging device, image data; executing, at the imaging device, one or more machine vision jobs on the image data; providing the image data to an obstacle detection model executing at the imaging device, the obstacle detection model being a trained machine learning model” (claim 1) and “one or more processors; one or more imaging sensors to capture image data over one or more fields of view (FOVs) of the imaging device; and one or more memories including computer-executable instructions stored thereon that, when executed by the one or more processors, cause the imaging device to: capture, via the imaging sensors, image data; execute one or more machine vision jobs on the image data” (claim 12) and adds insignificant extra-solution activity generally linking technological environments, reciting “communicating the textual description from the imaging device for receipt at a user computing device communicatively coupled to the imaging device”. Thus, claims 1 and 12 do not recite any of the exemplary considerations that are indicative of an abstract idea having been integrated into a practical application. These limitations are recited at a high level of generality (i.e. as a general action or change being taken based on the results of the acquiring step) and/or amounts to mere post solution actions, which is a form of insignificant extra-solution activity. Further, the claims are claimed generically and are operating in their ordinary capacity such that they do not use the judicial exception in a manner that imposes a meaningful limit on the judicial exception. Accordingly, even in combination, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. STEP 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? NO, the claims do not recite additional elements that amount to significantly more than the judicial exception. With regard to STEP 2B, whether the claims recite additional elements that provide significantly more than the recited judicial exception, the guidelines specify that the pre-guideline procedure is still in effect. Specifically, that examiners should continue to consider whether an additional element or combination of elements: adds a specific limitation or combination of limitations that are not well-understood, routine, conventional activity in the field, which is indicative that an inventive concept may be present; or simply appends well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception, which is indicative that an inventive concept may not be present. Independent claim(s) 1 and 12 do not recite any additional elements that are not well understood, routine or conventional. The use of generic computer elements is a routine, well understood and conventional process that is performed by computers. Regarding claim 2, claim adds receiving textual description of the alert data at the user computing device, this is utilizing generic computer processing and fails to remedy the abstract idea of claim 1. Regarding claim 3, claim adds obtaining environmental data and providing alert data found in the environmental data to the trained generative AI model to generate textual description, this is further generic computer processing to perform a mental process (i.e. evaluation and judgement) and fails to remedy the abstract idea of claim 1. Regarding claims 9 and 17, claims add an obstacle detection model (i.e. broad level AI-based model that amounts to “apply it” of a generic computing processing component), and capturing and providing stream image data to the obstacle detection model to analyze and identify the obstacle, this is generic processing to perform a mental process (e.g. analyze/evaluate observed data) and fails to remedy the abstract idea of claim 1. Regarding claims 10 and 18, claim adds grouping image data based on time and providing the groupings to the obstacle detection model, this is generic computer processing to perform a mental process (e.g. observation and evaluation) and fails to remedy the abstract idea of claim 1. Regarding claims 11 and 19, claim adds tracking the obstacle over time and predicting a collision of the obstacle with an object, generating obstacle collision alert data, providing the alert data to the trained generative AI model, and generating a textual description of the alert data, this amounts to generic processing to perform a mental process (e.g. observation, evaluation, and communication that could be performed with a pen and paper) and fails to remedy the abstract idea of claim 1. Thus, since method claims 1-3 and 9-11 and imaging device claims 12 and 17-19 are: (a) directed toward an abstract idea, (b) do not recite additional elements that integrate the judicial exception into a practical application, and (c) do not recite additional elements that amount to significantly more than the judicial exception, it is clear that Claims 1-3 and 9-11 and claims 12 and 17-19 are not eligible subject matter under 35 U.S.C 101. Claim(s) 4-8 and 13-16 are not directed to an abstract idea since the claims recite additional elements that integrate the judicial exception into a practical application and add significantly more than the judicial exception (e.g. improving and training an “obstacle detection model” to be more than a broad generic component that is simply being applied to perform mental processes). Therefore, 4-8 and 13-16 are not rejected under 35 U.S.C. 101. Amending claim 1 to include at least claim 4 and/or claims 4-8 may overcome 101 rejections for claim 1 and its respective dependent claims’ 101 rejections. Amending claim 12 to include claim 13 and/or claims 13-16 may overcome 101 rejections for claim 12 and its respective dependent claims’ 101 rejections. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 1-6, 8-10, and 12-18 are provisionally rejected on the grounds of nonstatutory double patenting as being unpatentable over the corresponding claims of U.S. Patent Application No. 18/934,236 (see chart below) in view of Zhang et al. (US 20230363609 A1). Although the claims at issue are not identical, they are not patentably distinct from each other because the claims at issue are broader in scope and/or are encompassed in the claims of the reference application (US 18/934,236) and/or encompassed in the teachings of Zhang. Specifically, the reference application teaches substantially the same method and imaging device architecture, image acquisition, machine learning analysis, alert generation, generative AI processing, user communication, user feedback/prompting, model updating, image stream processing and time-based segmentation, but are directed to detecting a physical shift of the imaging device, rather than the detection of an obstacle. Zhang teaches applying a trained obstacle detection model to analyze image data and determine whether an obstacle is present within the field of view (e.g. obstacle/no-obstacle classification) and alert data operations (e.g. generating notifications/sending out alerts), see Zhang Abstract, ¶ [0022], ¶¶ [0112]-[0115], and ¶¶ [0277]-[0279]). It would have been obvious to substitute Zhang’s known obstacle detection model for the physical shift detection model of the reference application because both analyze captured image data to detect an image condition using substantially the same processing workflow. Claims 11 and 19 are provisionally rejected on the grounds of nonstatutory double patenting as being unpatentable over the corresponding claims of U.S. Patent Application No. 18/934,236 (see chart below) in view of Calmer et al. (US 11643102 B1). Although the claims at issue are not identical, they are not patentably distinct from each other because the claims at issue are broader in scope and/or are encompassed in the claims of the reference application (US 18/934,236) and/or encompassed in the teachings of Calmer. The corresponding claims of the reference application recite substantially the same image processing, alert generation, and generative AI operations as the instant claims, but do not expressly recite tracking an obstacle over time, predicting a collision, and generating obstacle collision alert data. However, Calmer teaches tracking detected objects over time, predicting a likely collision, and generating corresponding collision event data (see Calmer column 10 line 51 through column 11 line 50, column 12 line 17 through column 13 line 39, column 15 lines 5-13, and column 15 line 51 through column 16 line 16, as further detailed in the claim 11’s 103 rejection found below). It would have been obvious to incorporate Calmer’s known collision prediction techniques into the otherwise substantially identical workflow of the reference claims. Please refer to the chart below depicting the instant application’s claims mapped to the matching claims from reference US Application No. 18/934,236. The mapping depicts the reference application’s claims that correspond to the instant application, in addition to the secondary reference (i.e. Zhang or Calmer), applied to reject the corresponding instant claim under nonstatutory double patenting. The teachings of Zhang and Calmer, discussed above, are applicable to respective mapped claims below. For sake of brevity, independent claims are fully recited and the dependent claims are mapped using their respective reference claim numbers and teachings of Zhang or Calmer. Please refer to their respective disclosures for further clarity. Instant claims US Application No. 18/934,236 claims Claim 1: A method of detecting an obstacle using an imaging device, the method comprising: capturing, at the imaging device, image data; executing, at the imaging device, one or more machine vision jobs on the image data; providing the image data to an obstacle detection model executing at the imaging device, the obstacle detection model being a trained machine learning model; analyzing, using the obstacle detection model, the image data to identify an obstacle in a field of view (FOV) of the imaging device; responsive to identifying the obstacle in the FOV, generating alert data at the imaging device, and providing the alert data to a trained generative artificial intelligence (Al) model, at the imaging device, the generative Al model trained to generate a textual description of the alert data; and communicating the textual description from the imaging device for receipt at a user computing device communicatively coupled to the imaging device. Claim 1 (18/934,236): A method of detecting a executing, at the imaging device, one or more machine vision jobs on the image data; providing the image data to a analyzing, using the responsive to identifying model trained to generate a textual description of the and communicating the textual description from the imaging device for receipt at a user computing device communicatively coupled to the imaging device.in view of Zhang’s obstacle detection model teachings discussed and applied above. 2 2 in view of Zhang teachings applied above. 3 3 in view of Zhang teachings applied above. 4 4 in view of Zhang teachings applied above (i.e. obstacle/no obstacle classification). 5 5 in view of Zhang teachings applied above. 6 6 in view of Zhang teachings applied above. 8 7 in view of Zhang teachings applied above. 9 9 in view of Zhang teachings applied above. 10 10 in view of Zhang teachings applied above. 11 1 + 9 + 10 in view of Calmer teachings applied above Claim 12: An imaging device comprising: one or more processors; one or more imaging sensors to capture image data over one or more fields of view (FOVs) of the imaging device; and one or more memories including computer-executable instructions stored thereon that, when executed by the one or more processors, cause the imaging device to: capture, via the imaging sensors, image data; execute one or more machine vision jobs on the image data; provide the image data to an obstacle detection model at the imaging device, the obstacle detection model being a trained machine learning model; analyze, using the obstacle detection model, the image data to identify an obstacle in a field of view (FOV) of the imaging device; responsive to identifying the obstacle in the FOV, generate alert data and provide the alert data to a trained generative artificial intelligence (Al) model, at the imaging device, trained to generate a textual description of the alert data; and communicate the textual description from the imaging device for receipt at a user computing device communicatively coupled to the imaging device. Claim 11: An imaging device comprising: one or more processors; one or more imaging sensors to capture image data over one or more fields of view (FOVs) of the imaging device; and one or more memories including computer-executable instructions stored thereon that, when executed by the one or more processors, cause the imaging device to: capture, via the imaging sensors, image data; execute one or more machine vision jobs on the image data; provide the image data to a analyze, using the responsive to identifying the and communicate the textual description from the imaging device for receipt at a user computing device communicatively coupled to the imaging device. in view of Zhang’s obstacle detection model teachings discussed and applied above. 13 12 in view of Zhang teachings applied above. 14 13 in view of Zhang teachings applied above (i.e. obstacle/no obstacle classification). 15 14 in view of Zhang teachings applied above. 16 15 in view of Zhang teachings applied above. 17 17 in view of Zhang teachings applied above. 18 18 in view of Zhang teachings applied above. 19 11 + 17 + 18 in view of Calmer teachings applied above. This is a provisional nonstatutory double patenting rejection because the patentably indistinct claims have not in fact been patented. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-3, 9-12, and 17-19 are rejected under 35 U.S.C. 103 as being unpatentable over Calmer et al (US 11643102 B1; hereafter “Calmer”) in view of Pertsel et al. (US 12654713 B1; hereafter “Pertsel”). Regarding claim 1 (method) and 12 (imaging device), Calmer teaches: A method of detecting an obstacle using an imaging device (column 1 lines 7-9 “present disclosure relate to devices, systems, and methods that provide real-time safety event detection”; Abstract “A vehicle dash cam may be configured to execute one or more neural networks (and/or other artificial intelligence), such as based on input from one or more of the cameras and/or other sensors associated with the dash cam, to intelligently detect safety events in real-time.”; column 8 line 16 “the vehicle device 114 comprises a dash cam”; see FIGs. 1B-1D), [claim 12] An imaging device comprising: one or more processors; one or more imaging sensors to capture image data over one or more fields of view (FOVs) of the imaging device; and one or more memories including computer-executable instructions stored thereon that, when executed by the one or more processors, cause the imaging device to: capture, via the imaging sensors, image data; [perform subsequent limitations] (Refer back to column 1 lines 7-9 “present disclosure relate to devices,” Abstract, column 8 lines 10-19, and FIGs. 1B-1D. Calmer further teaches the devices are “configured to identify features within sensor data, such as in images from one or more of the outward-facing or inward-facing cameras” (column 6 lines 40-47) and hardware configurations (e.g. memories including stored computer-executable instructions that cause the system to execute the processing configurations), including “a computer readable storage medium (or mediums) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure” (column 19 lines 18-22; see more hardware configurations throughout column 19 line 16 through column 21 line 28) the method comprising: capturing, at the imaging device, image data (Column 8 lines 21-22 “The sensors 112 may include, for example, one or more inward-facing camera and one or more outward-facing camera.”; column 3 lines 28-30 “The dash cam… provides real-time alerts based on processing of video data from one or more cameras of the dash cam.”); executing, at the imaging device, one or more machine vision jobs on the image data (Column 3 lines 30-31 “The safety event detection is performed local to the dash cam…”; column 6 lines 56-61 “the vehicle device can process video data locally to identify various associated features, such as detection of an object… characteristics of the object…, location of the object within the image files of the video”; column 10 lines 33-35 “One or more computer processors… that enable concurrent neural networks for real-time artificial intelligence”); providing the image data to an obstacle detection model executing at the imaging device, the obstacle detection model being a trained machine learning model (Column 11 lines 3-8 “the vehicle device 114 advantageously executes one or more event models (e.g., neural networks) on sensor data, such as video data, to detect safety events, such as a tailgating, forward collision risk, and/or distracted driver event.”; column 7 lines 1-5 “the feature detection module and/or event models…can include a machine learning component”; column 7 lines 29-31 “certain embodiments herein may use…convolutional neural networks, deep networks”; column 15 lines 2-4 “The neural network may be trained at the event analysis system and periodically provided to the vehicle device for improved detection”); analyzing, using the obstacle detection model, the image data to identify an obstacle in a field of view (FOV) of the imaging device (column 6 lines 56-58 “the vehicle device can process video data locally to identify various associated features, such as detection of an object (e.g., a person or a vehicle)”; column 5 lines 46-51 “metadata regarding a triggered event may include a location of an object that triggered the event, such as a vehicle in which a forward collision warning (“FCW”) or tailgating safety event has triggered”; column 16 lines 2-4 “Metadata may further include information about other vehicles or objects within the field of view of the cameras”; column 12 lines 61-65 “metadata shows bounding boxes 420, 422 indicating objects identified by the feature detection neural network” Under the broadest reasonable interpretation, a detected vehicle or other detected object within the camera field of view that triggers a forward-collision or tailgating safety event (as taught by Calmer above) reasonably teaches the claimed “obstacle”. The instant application’s specification explains that “Obstacles can be any item in the environment” that interferes with imaging, that “Obstacles may be fixed in position or may be moving” and may include “foreign objects,” thereby demonstrating that the claimed obstacle is not limited to a particular conveyor belt implementation or type of object (see ¶¶ [0072]-[0073] of the instant application’s specification). Therefore, the methods for detection, detection neural network, and other logical and/or analytical models taught by Calmer with respect to foreign objects within the camera field of view (i.e. objects within the camera field of view that trigger a forward-collision or tailgating safety event) reasonably correspond to the claimed logical and/or analytical operations with respect to obstacles (e.g. methods for detecting obstacles, object detection model, etc.).); responsive to identifying the obstacle in the FOV, generating alert data at the imaging device, and providing the alert data to a (Abstract “Detection of a safety event may trigger an in-cab alert”; column 15 lines 5-9 “at block 212, if a safety event has been triggered… an in-vehicle alert 214 is provided within the vehicle and event data associated with the event is identified and transmitted to the event analysis system”; column 15 line 51 through column 16 line 12 “the event data 219 that is transmitted to the event analysis system…may include metadata and only a limited (or no) other asset data…metadata…may include location of the object that triggered the event…severity of the event…[and] confidence level” Under the broadest reasonable interpretation, Calmer’s teaching of event data identified in response to the triggered alert constitutes the claimed “alert data”.); and communicating (column 15 lines 14-24 “alerts may also be transmitted to one or more devices external to the vehicle… The alert may be delivered via SMS, text message, or application-specific alert, or may be delivered via the safety dashboard, email, and/or via any other communication medium.”). Calmer fails to explicitly disclose: providing the alert data to a trained generative artificial intelligence (Al) model, at the imaging device, the generative Al model trained to generate a textual description of the alert data, and communicating the textual description from the imaging device for receipt at a user computing device. However, Calmer expressly teaches local execution of trained neural networks on the vehicle device (see Calmer column 3 lines 24-34; column 9 line 24 through column 10 line 35), thereby providing a suitable platform for incorporating additional artificial intelligence models. In a related art, Pertsel teaches: providing the alert data to a trained generative artificial intelligence (Al) model (see Pertsel column 35 lines 34-45 “Generally, the AI module 270 may comprise hardware configured to implement DAGs and/or LLMs…. The video-to-text AI module 272 may be configured to receive the video data from the signal VDATA or the video data from the signal EFRM. For example, the video-to-text AI module 272 may be configured to operate on … the subset of the video frames likely to comprise an event of interest”); the generative Al model trained to generate a textual description of the alert data (column 35 lines 4-28 “The video-to-text analysis may generate a plain text (e.g., natural language that may be human readable) description of the content of the video frames.”; column 35 lines 52-67 “The video-to-text AI module 272 may be configured to perform an analysis of the video data and generate the smart metadata… The smart metadata may comprise a full text description of the video frames. The smart metadata may comprise a plain language description of the objects in the video frames, the context of the video frames, the colors in the video frames, the arrangement of the visual elements in the video frames, the behavior of objects in the video frames, the location of objects in the video frames, the types of objects in the video frames, etc.” Pertsel further teaches that “the video-to-text analysis may be a transformer network,” may use “large language models (e.g., BLIP-2),” may perform “vision-to-language generative learning,” may be based on “a Flamingo80B model” and may provide “zero-shot image-to-text generation that may follow natural language instructions” (column 36 lines 31-60).); and communicating the textual description from the imaging device for receipt at a user computing device coupled to the imaging device (Pertsel expressly teaches generating the textual description at an edge imaging device, incorporating that description into a notification signal, communicating the notification through the edge device’s communication interface, and displaying or outputting the natural text description on a separate, communicatively connected, user device (column 30 lines 64-67 “the AI models may be implemented on the edge devices 100a-100n (e.g., without offloading computation to a cloud computing service); column 31 lines 15-18 “The user device 252 may be a device separate from the edge devices 100a-100n… [and] may be configured to connect to the edge devices 100a-100n (e.g., via the communication interface 254)”; column 32 lines 13-23 “The communication interface 254 may be configured to generate a signal (e.g., NOTIFY). In one example, the signal NOTIFY may be generated in response to the signal SCORE. The signal NOTIFY may be presented to the user device 252.”; column 32 lines 24-47 “The signal SMETA may comprise text description metadata (e.g., smart metadata)…The signal NOTIFY may comprise a notification, a type of alert, audio, text and/or video.”; column 39 lines 26-34 “Providing the signal SCORE with the smart metadata may enable the signal NOTIFY to comprise a human readable description of the event detected….the signal NOTIFY may comprise a plain language description of the event…”; column 46 lines 47-63 “the signal NOTIFY may comprise the video frames only” or “the signal NOTIFY may comprise the smart metadata describing the video frames only”; column 46 line 64 through column 47 line 5 “The user interface 302 may display and/or output the video results and/or the natural text answer/description from the signal NOTIFY on the user device 252”; See FIG. 4, Capture Device (i.e. imaging device) pipeline sends video to Video-To-Text AI Module 272, then sends score to Communication Interface 254 and subsequently sends NOTIFY (i.e. notifies) to User Device 252, wherein “the user device 252 may be a smartphone…a desktop computer, a laptop computer, a tablet computing device, a smartwatch, a security terminal, etc.” (column 31 lines 29-35).). Thus, Pertsel teaches communicating the textual description from the imaging device for receipt at a user computing device, as claimed.). Calmer and Pertsel are analogous to the claimed invention because each concerns a camera-based edge device that analyzes captured image data using AI to detect objects or events and communicates the resulting alert information, with Pertsel further teaching the generation and communication of a natural language description of the detected event. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the vehicle-device of Calmer to incorporate the video-to-text generative AI model taught by Pertsel so that the event data generated in response to a detected obstacle is automatically converted into a natural-language description before being communicated to a user. Calmer already performs local neural network processing to provide alerts, while Pertsel teaches that transformer and LLM based vision language models generate plain language description of detected objects, behaviors, locations, and context from event-related image data. Incorporating the known vision language model taught by Pertsel into Calmer’s existing local artificial intelligence architecture would have predictably improved the usefulness of the generated alerts by providing more informative, readable notifications without changing the underlying obstacle detection operation. Independent claim 12 recites an imaging device configured to perform operations substantially corresponding to the method of claim 1. Accordingly, Calmer and Pertsel teach or suggest the limitations of claim 12 for substantially the same reasons discussed above with respect to claim 1. Further, with respect to the imaging device claim 12 (and imaging device claims 13-19 discussed below), it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to store computer-executable instructions on one or more memories in an imaging device that, when executed by one or more processors, cause the imaging device to perform the combined obstacle detection method taught by Calmer and Pertsel (and further by Calmer, Pertsel, Zhang, and Gonzalez throughout claims 13-19), as taught by Pertsel, because Pertsel teaches implementing analogous processing operations through stored computer executable instructions. Doing so would have presented predictable use of Pertsel’s known implementation technique to perform the combine functionality of detecting an obstacle using an imaging device. [ EXAMINER NOTES, the imaging device claims 13-19 discussed below are shown to be taught by Calmer, Pertsel, Zhang, and Gonzales with respect to corresponding claims 2-11. For sake of brevity, examiner has shown the teachings of Calmer, Pertsel, Zhang, and Gonzales by using the claim language of method claims 2-11 that correspond to operations performed by the imaging device found in claims 13-19. While examiner has not included all of the claim language of claims 13-19 relating to the imaging device and specifically the computer-executable instructions for performing the obstacle detection methods, it would have been obvious to a person of ordinary skill in the art, prior the effective filing date of the claimed invention, to modify the teachings of Calmer, Pertsel, Zhang, and Gonzales found throughout claims 2-11 to be implemented using the imaging device architecture and computer-executable instruction-based technique taught by Pertsel (as seen in claims 1 and 12s’ 103 rejection above) because doing so would have presented predictable use of Pertsel’s known implementation technique to perform the combine functionality of Calmer, Pertsel, Zhang, and Gonzales’ teachings, discussed below with respect to claims 2-11. ] Regarding claim 2, Calmer and Pertsel teach the method of claim 1, including user computing device and textual description of the alert data. Perstel further teaches: receiving, at the user computing device, the textual description of the alert data (Under the broadest reasonable interpretation, the notification (NOTIFY) containing the human readable event description, taught by Pertsel, corresponds to the claimed textual description of the alert data (see column 39 lines 26-28 “the signal NOTIFY to comprise a human readable description of the event detected.”; column 46 lines 54-56 “the signal NOTIFY may comprise the smart metadata describing the video frames only”; column 32 lines 22-23 “The signal NOTIFY may be presented to the user device 252.”; column 46 lines 65 through column 47 line 1 “The user interface 302 may display and/or output…the natural text answer/description from the signal NOTIFY on the user device 252”). Regarding claim 3, as best understood in light of the 35 U.S.C. 112(b) issue addressed above, Calmer and Pertsel teach the method of claim 1. Calmer further teaches: obtaining environmental data containing information on an environment in which the imaging device is located for capturing image data (Calmer teaches obtaining metadata including other vehicles or objects found in a scene, a location of an object that triggered an event, and event location (i.e. alert data, environmental data, and alert data corresponding to environmental data) within the field of view of cameras (Column 5 lines 3-4 “ Sensor Data: any data obtained by the vehicle device, such as asset data and metadata.”; column 5 lines 40-61 “metadata regarding a triggered event may include a location of an object that triggered the event…Metadata may include information about other vehicles within the scene in the case of tailgating or FCW event...[and] event location”; column 16 lines 2-4 “Metadata may further include information about other vehicles or objects within the field of view of the cameras”); and providing the alert data on the environmental data (Calmer teaches environmental data that is correlated to the alert data (i.e. when an event is triggered) (column 15 line 51 through column 16 line 12 “the event data 219 that is transmitted to the event analysis system…may include metadata and only a limited (or no) other asset data…metadata…may include location of the object that triggered the event…severity of the event…[and] confidence level”).) Calmer fails to explicitly disclose: providing the alert data on the environmental data to the trained generative Al model and generating the textual description of the detected obstacle to include a description derived from the environmental data. However, as shown with respect to claim 1’s 103 rejection, Calmer expressly teaches local execution of trained neural networks on the vehicle device (see Calmer column 3 lines 24-34; column 9 line 24 through column 10 line 35), thereby providing a suitable platform for incorporating additional artificial intelligence models and Pertsel teaches providing alert data to a trained generative AI model, the generative AI model trained to generate a textual description of the alert data (For sake of brevity, refer back to teachings found in 103 rejection for claim 1 and Pertsel column 35 line 4 through column 36 line 60). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the vehicle-detection system of Calmer, previously modified by Pertsel to generate textual descriptions of alert data, to incorporate Calmer’s further teachings utilizing environmental data corresponding to alert data in order to generate a textual description of the environment where the obstacle was detected. Doing so would enhance subsequent training of the object detection model taught by Calmer and Pertsel (taught in claim 1) by accounting for where obstacles have occurred in the past, and provide more meaningful information in a human-readable format to the user that may be used in obstacle analysis. Regarding claim 9, Calmer and Pertsel teach the method of claim 1. Calmer further teaches: capturing, at the imaging device, a stream of image data (see Calmer column 3 lines 24-34 “The dash cam is installable into existing vehicles and provides real-time alerts based on processing of video data from one or more cameras of the dash cam”; column 10 lines 51-59 “sensor data output from the multiple sensors …may be recorded… video data and metadata from one or more sensors may be stored”); and providing the stream of image data to the obstacle detection model and analyzing, using the obstacle detection model, the stream of image data to identify the obstacle (column 11 lines 3-8 “the vehicle device 114 advantageously executes one or more event models (e.g., neural networks) on sensor data, such as video data, to detect safety events, such as a tailgating, forward collision risk,”; column 6 lines 56-61 “the vehicle device can process video data locally to identify various associated features, such as detection of an object (e.g., a person or a vehicle)”; column 11 lines 35-37 “ detection models executed at the vehicle device are performed on downsampled images from the video feed.” Under the broadest reasonable interpretation, Calmer’s continuously acquired camera video feed constitutes the claimed stream of image data, and a detected vehicle or other object producing a forward collision or tailgating condition reasonably corresponds to the claimed obstacle.). Regarding claim 10, Calmer and Pertsel teach the method of claim 9, including providing the stream of image data to the obstacle detection model. Calmer further teaches: wherein providing the stream of image data to the obstacle detection model further comprises: performing a group segmentation on the stream of image data to generate a plurality of time- based groupings of the image data (column 13 line 67 through column 14 line 2 “the distracted driver safety event may be configured to only trigger after the head pose angle and confidence level exceed the threshold levels for a minimum time period, such as 5 seconds.”; A tailgating event may be required to occur for a configurable duration (column 12 lines 17-40); The system may utilize “a 20 second video segment” (column 12 line 58 through column 13 line 3); Additional video corresponding to “a particular time period” may be requested (see column 18 lines 5-35); Under the broadest reasonable interpretation, Calmer’s teachings of video segments and defined event-analysis time periods reasonably correspond to the claimed plurality of time-based groupings of image data.); and providing the plurality of time-based groupings of the image data to the obstacle detection model (column 11 lines 3-9 “the vehicle device 114 advantageously executes one or more event models (e.g., neural networks) on sensor data, such as video data”; column 11 lines 35-37 “the event detection models executed at the vehicle device are performed on downsampled images from the video feed”; As discussed above with respect to claim 9, Calmer teaches providing the captured image data to the obstacle detection model for analysis. Accordingly, the discussed time-based groupings of the captured image data are likewise provided to the obstacle detection model.). Regarding claim 11, Calmer and Pertsel teach the method of claim 10. Calmer further teaches: tracking, at the imaging device, the obstacle over time and predicting a collision of the obstacle with an object (Calmer teaches a tailgating condition of 30 seconds, wherein “a distance to a car in front of the vehicle, a ‘leading vehicle,’ is determined, such as by analysis of video data, to be less than a threshold distance or travel time)” (column 12 lines 16-40) and further teaches forward collision determination considers “whether a collision course with a leading vehicle is likely” and a neural network executed at the vehicle device determines whether “a forward collision is likely” based on “analysis of video and/or other sensor data” and whether “a time to collision threshold has been passed” (column 13 lines 13-39).); generating an obstacle collision alert data (column 15 lines 5-9 “if a safety event has been triggered the method continues to block 214 where an in-vehicle alert 214 is provided within the vehicle and event data associated with the event is identified and transmitted”; The metadata transmitted may include “object that triggered the event, such as the lead vehicle in the case of a forward collision warning or tailgating” (column 15 line 51 through column 16 line 16).), Pertsel further teaches the discrepancies of the claim not taught by Calmer, including: providing the obstacle collision alert data to the trained generative Al model (see Pertsel column 35 lines 39-51 “The video-to-text AI module 272 may be configured to receive the video data…The video-to-text AI module 272 may be configured to generate the signal SMETA”; column 35 lines 29-32 “The circuit 272 may implement a video-to-text AI module (or model).”); and generating the textual description to include a description derived from the obstacle collision alert data (column 39 lines 26-28 “Providing the signal SCORE with the smart metadata may enable the signal NOTIFY to comprise a human readable description of the event detected.”; column 43 lines 45-48 “The signal ANS may comprise the natural text description.”). Under the broadest reasonable interpretation of the claim, Calmer’s detected collision event data corresponds to the claimed obstacle collision alert data, and Pertsel’s video-to-text AI generates a natural language description derived from that detected event. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the object detection system taught by Calmer and Pertsel to provide Calmer’s detected obstacle event data to Pertsel’s trained generative AI model so that detected collision events are communicated as natural language descriptions, thereby predictably improving users’ ability to understand obstacle collision alerts. Regarding claim 17, Calmer and Pertsel teach the imaging device of claim 12. Claim 17 recites an imaging device configured to perform the image stream operations substantially corresponding to the method of claim 9. As discussed above with respect to claim 12, Pertsel teaches the recited imaging device architecture (e.g. processor(s), memory storing instructions to execute operations, etc.). Accordingly, Calmer and Pertsel teach the limitations of claim 17 for the same reasons discussed above with respect to claim 9. Regarding claim 18, Calmer and Pertsel teach the imaging device of claim 17. Claim 18 recites an imaging device configured to perform the group segmentation operations substantially corresponding to the method of claim 10. As discussed above with respect to claim 12, Pertsel teaches the recited imaging device architecture (e.g. processor(s), memory storing instructions to execute operations, etc.). Accordingly, Calmer and Pertsel teach the limitations of claim 18 for the same reasons discussed above with respect to claim 10. Regarding claim 19, Calmer and Pertsel teach the imaging device of claim 18. Claim 19 recites an imaging device configured to perform the obstacle tracking, collision prediction, and textual description operations substantially corresponding to the method of claim 11. As discussed above with respect to claim 12, Pertsel teaches the recited imaging device architecture (e.g. processor(s), memory storing instructions to execute operations, etc.). Accordingly, Calmer and Pertsel teach the limitations of claim 19 for the same reasons discussed above with respect to claim 11. Claim(s) 4 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Calmer et al (US 11643102 B1; hereafter “Calmer”) in view of Pertsel et al. (US 12654713 B1; hereafter “Pertsel”) as applied to claims 1 and 12 above, and in further view of Zhang et al. (US 20230363609 A1; hereafter “Zhang”). Regarding claims 4 and 13 Calmer and Pertsel teach the method of claim 1 and the imaging device of claim 12. Calmer and Pertsel fail to explicitly disclose: wherein the obstacle detection model is trained to classify image data as indicating no presence of obstacles in the FOV or presence of an obstacle in the FOV. In a related art, Zhang teaches: wherein the obstacle detection model is trained to classify image data as indicating no presence of obstacles in the FOV or presence of an obstacle in the FOV (Zhang teaches “using a deep learning trained classifier to detect obstacles and pathways in an environment” based on “image information captured by at least one visual spectrum-capable camera…using an ensemble of trained neural network classifiers” to determine the identity for objects corresponding to features extracted from images (see Zhang Abstract and ¶ [0022]). Zhang further teaches a “Barrier range discretizes the field of view of the camera on the robot into angular bins,” where each bin contains barrier information, increasing an occupied probability where a barrier-range item “has obstacle,” and decreasing the occupied probability where the barrier-range item “has no obstacle” (¶¶ [0277]-[0279]; FIGs. 31A-31C). Thus, Zhang teaches applied a trained image-based obstacle classifier to distinguish portions of the camera field of view containing an obstacle from portions having no obstacle, thereby classifying the analyzed image information according to obstacle presence or absence.). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to train the obstacle detection model of Calmer, previously modified by Pertsel, using the obstacle classification techniques of Zhang because all three references utilize trained machine learning models to analyze camera image data for obstacle or object detection, and incorporating the expressly taught obstacle/non-obstacle classification would have predictably enabled the obstacle detection model taught by Calmer and Pertsel to expressly distinguish image data indicating the presence of an obstacle, thereby improving the accuracy and robustness of determining whether an obstacle is present. Claim(s) 5-8 and 14-16 are rejected under 35 U.S.C. 103 as being unpatentable over Calmer et al (US 11643102 B1; hereafter “Calmer”) in view of Pertsel et al. (US 12654713 B1; hereafter “Pertsel”), in further view of Zhang et al. (US 20230363609 A1; hereafter “Zhang”), and in further view of Gonzalez et al. (US 20190228262 A1; hereafter “Gonzalez”). Regarding claims 5 Calmer, in view of Pertsel and Zhang, teach the method of claim 4, including training the obstacle detection model to classify image data. As shown above, the combination of Calmer, Pertsel and Zhang teach a trained obstacle detection model and natural language communication between an imaging device and a user computing device. Pertsel further teaches: communication between an edge device and a separate user device, including presenting prompts and receiving natural-language user input (Pertsel column 43 lines 6-16 “The signal QUERY may comprise a plain text and/or natural language description”; column 44 lines 33-39 “The LLM AI module 310 may be configured to receive the question/request from the signal QUERY.”; column 46 lines 50-52 “The communication interface 254 may generate the signal NOTIFY”; column 46 line 64 through column 47 line 1 “The user interface 302 may display… the natural text answer/description from the signal NOTIFY”). Calmer further teaches: using user-verified detections as training data to improve the detection model, disclosing the accurate detections are used as “positive training data for updating a neural network” and inaccurate detections are used as “negative training data for updating the neural network” (column 14 line 53 through column 15 line 4). The combination of Calmer, Pertsel, and Zhang fail to explicitly disclose: where the obstacle detection model is further trained to classify image data as equivocal, indicating no determination of a presence of an obstacle in the FOV, the method further comprising: in response to the obstacle detection model classifying image data as equivocal, the generative Al model generating a prompt request; communicating the prompt request from the imaging device to the user computing device for generating a prompt at the user computing device for input of user data comprising natural language descriptions at the prompt; and receiving the user data input to the prompt, from the user computing device to the imaging device for updating the trained machine learning model. In a related art, Gonzalez teaches: (Gonzales ¶ [0037] “a trigger includes instances in which an image classifier generates a confidence score for a classified object that falls below a given threshold.”; ¶ [0044] “the vehicle control system 103 may determine that a confidence score associated with a classified object is below a specified threshold”; Under the broadest reasonable interpretation, a below-threshold confidence score corresponds to an equivocal classification because the classifier has not made a sufficiently reliable determination regarding the detected object.), the method further comprising: in response to (Gonzalez teaches generating clarification prompt responsive to the uncertain classification (i.e. image data classified as equivocal) (¶ [0044] “the vehicle control system 103 detects a trigger to initiate an annotation prompt”; ¶ [0037] “Once a trigger is detected, the initiation component 232 may generate an annotation prompt…. The prompt may request that the user identify… the correct label”). (¶ [0051] “present, via a user interface, the annotation prompt; receive, via the user interface, user input indicative of a response to the annotation prompt”); (¶ [0051] “receive, via the user interface, user input indicative of a response to the annotation prompt by a user”; ¶ [0039] “update component 236 is configured to update…model data 208 based on responses provided by the user”). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the user assisted annotation techniques of Gonzales into the obstacle detection and generative AI system of Calmer, as modified by Pertsel, and Zhang, because Gonzalez teaches resolving uncertain image classifications through user annotation, while the prior combination of Calmer, Pertsel, and Zhang already teaches obstacle detection, generative AI-generated natural language communication, and using verified user feedback as training data to improve a neural network. Doing so would have predictably enabled uncertain (equivocal) obstacle classifications to be confirmed through user feedback, thereby improving the accuracy of the trained obstacle detection model through additional verified training data. Regarding claim 6, Calmer, in view of Pertsel, Zhang, and Gonzalez, teach the method of claim 5, including receiving natural language user data responsive to a prompt generated for an equivocal obstacle classification. Pertsel further teaches: providing the user input data received from the user computing device to the generative Al model of the imaging device (Pertsel column 43 lines 6-16 “The signal QUERY may comprise a plain text and/or natural language description”; column 44 lines 33-34 “The LLM AI module 310 may be configured to receive the question/request from the signal QUERY.”; column 47 lines 7-30 “the query module 304 (or the LLM AI module 310) may be one distinct AI model (e.g., a LLM) implemented to understand the criteria for the triggers and/or an answer to a query presented by the user. In some embodiments, all four of the AI models may be implemented locally by the processor 102. [at the edge device]”); generating, (column 44 lines 40-65 “the LLM AI module 310 may translate the natural language input into computer readable information”; column 31 lines 62-64 “The user device 252 may enable the end user to tag video captured for providing training data to the various AI models”; Thus Pertsel teaches using the locally executed LLM to convert natural language user input into machine readable information and using user-provided labels or tags as AI model training data); Calmer further teaches: providing the updated training data to the trained machine learning model for updating the trained machine learning model (column 14 lines 62-66 “an indication of an accurate detection of a distracted driver event may cause the event data to be used as positive training data for updating a neural network configured to detect distracted driver events, while an indication of an inaccurate detection of a distracted driver event may cause the event data to be used as negative training data for updating the neural network. The neural network may be trained at the event analysis system and periodically provided to the vehicle device for improved detection”). Gonzales also teaches: updating model data based on the prompted user response (¶ [0051] “receive, via the user interface, user input indicative of a response to the annotation prompt”; [0039] “The update component 236 is configured to update the…model data 208 based on responses provided by the user”). The combination of Calmer, Palmer, Zhang, and Gonzalez fail to explicitly disclose that the generative AI model itself generates updated training data from the received user data for updating the trained obstacle detection model. However, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention modify the object detection system previously taught by Calmer, Palmer, Zhang, and Gonzalez so that Pertsel’s locally executed generative AI model converts the user provided natural language responses obtained through Gonzalez’s annotation process into updated training data for Calmer’s trained obstacle detection model. Doing so merely applies known functions of the references according to their established purposes. Specifically, the modification would represent a predictable use of converting natural language user feedback into machine readable information and using verified user feedback to update the machine learning model, thereby improving the performance and accuracy of the trained obstacle detection model by continuously refining it based on user verified information. Regarding claim 7, Calmer, in view of Pertsel, Zhang, and Gonzalez, teach the method of claim 5. In claims 1, 4, and 5 (claims from which claim 7 depends upon), the combination of prior art references taught classifying the image data as indicating no presence of obstacles in the FOV, presence of an obstacle in the FOV, or equivocal, and further disclosed (with respect to claim 5) prompting a user for input when an obstacle classification is equivocal, receiving the responsive user data, and updating the trained obstacle detection model using the user data. Pertsel further teaches: receiving user provided labels as training input (column 43 lines 6-16 “The signal QUERY may comprise a plain text and/or natural language description”; The user may “tag video captured for providing training data to the various AI models” (column 31 lines 62-34).). Gonzalez further teaches: requesting that the user identify “the correct label” from a list of labels when the classifier confidence falls below a threshold (see Gonzalez ¶ [0037]-[0045]). While Calmer, in view of Pertsel, Zhang, and Gonzalez fail to explicitly disclose: wherein the user input data comprises a label indicating no presence of obstacles in the FOV, presence of an obstacle is the FOV, or equivocal, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to configure the user interface of the combined system of Calmer, Pertsel, Zhang, and Gonzalez to receive user input in the form of labels corresponding to the previously established obstacle classifications (i.e., no obstacle present, obstacle present, or equivocal) because Gonzalez teaches presenting a user with selectable labels to resolve uncertain classifications, Pertsel teaches receiving user-provided labels for AI training at the edge device, and Zhang taught the underlying obstacle present and obstacle absent classifications (see Pertsel and Gonzalez teachings’, found above, and Zhang’s teachings with respect to claims 4 and 13). Combining these known techniques would predictably provide standardized user feedback to correspond to the obstacle classifications already generated by the obstacle detection model, thereby increasing the reliability and accuracy of the model by further incorporating user verified feedback into the update process. Regarding claim 8 Calmer, Pertsel, and Zhang teach the method of claim 4. Calmer further teaches: (see FIG. 2; column 15 lines 5-9 “event data associated with the event is identified and transmitted to the event analysis system (block 216)”; column 17 lines 41-44 “ the event analysis system receives the event data 219…and stores the event data for further analysis at block 210”; column 18 lines 5-20 “ At block 230, additional event data is requested from the vehicle device”); (column 17 lines 41 through column 18 line 4 “high-fidelity event detection models, such as higher precision neural networks than are executed on the vehicle device, may be executed to determine whether the triggered event was accurately detected…the event models applied in the event analysis system may take as inputs additional sensor data, such as full video data”); providing the image data and the re-analyzed determination to an instantiation of the obstacle detection model executed (Calmer teaches the event analysis stem executes higher-fidelity event detection models using the received event data to evaluate the detected event (i.e. detected obstacle), see column 17 line 41 through column 18 line 20.); and updating training of the instantiation of the obstacle detection model to generate an updated obstacle detection model (Column 18 line 55 through column 19 line 3 “event data may become part of a training data set for updating/improving a neural network configured to detect the safety event… the models can be regenerated on a periodic basis”); and communicating the updated obstacle detection model (column 15 lines 2-4 “The neural network may be trained at the event analysis system and periodically provided to the vehicle device for improved detection of safety events.”). Calmer, Pertsel, and Zhang fail to explicitly disclose: the obstacle detection model being further trained to classify image data as equivocal, indicating no determination of a presence of an obstacle in the FOV, and triggering the communication and re-analysis workflow specifically in response to the equivocal classification. In a related art, Gonzales teaches: classify image data as equivocal, indicating no determination of an obstacle (Gonzales ¶ [0037] “a trigger includes instances in which an image classifier generates a confidence score for a classified object that falls below a given threshold.”; ¶ [0044] “the vehicle control system 103 may determine that a confidence score associated with a classified object is below a specified threshold”; Under the broadest reasonable interpretation, a below-threshold confidence score corresponds to an equivocal classification because the classifier has not made a sufficiently reliable determination regarding the detected object.); in response to (Gonzalez teaches generating clarification prompt responsive to the uncertain classification (i.e. image data classified as equivocal) (¶ [0044] “the vehicle control system 103 detects a trigger to initiate an annotation prompt”; ¶ [0037] “Once a trigger is detected, the initiation component 232 may generate an annotation prompt…. The prompt may request that the user identify… the correct label”) It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the user-assisted annotation techniques of Gonzalez into the obstacle detection system of Calmer, Pertsel, and Zhang using the communicatively connected user computing device previously established by Pertsel, so that uncertain obstacle classifications are resolved through user feedback and used to update the obstacle detection model, thereby predictably improving obstacle detection accuracy. Regarding claim 14, Calmer, in view of Pertsel and Zhang, teach the imaging device of claim 13. Claim 14 recites an imaging device configured to perform operations substantially corresponding to the method of claim 5. As discussed above with respect to claim 12, Pertsel teaches the recited imaging device architecture (e.g. processor(s), memory storing instructions to execute operations, etc.). Accordingly, Calmer, Pertsel, Zhang, and Gonzalez teach the limitations of claim 14 for the same reasons discussed above with respect to claim 5. Regarding claim 15, Calmer, in view of Pertsel, Zhang, and Gonzalez, teach the imaging device of claim 14. Claim 15 recites an imaging device configured to perform operations substantially corresponding to the method of claim 6. As discussed above with respect to claim 12, Pertsel teaches the recited imaging device architecture (e.g. processor(s), memory storing instructions to execute operations, etc.). Accordingly, Calmer, Pertsel, Zhange, and Gonzalez teach the limitations of claim 15 for the same reasons discussed above with respect to claim 6. Regarding claim 16, Calmer, in view of Pertsel and Zhang, teach the imaging device of claim 13. Claim 16 recites an imaging device configured to perform operations substantially corresponding to the method of claim 8. As discussed above with respect to claim 12, Pertsel teaches the recited imaging device architecture (e.g. processor(s), memory storing instructions to execute operations, etc.). Accordingly, Calmer, Pertsel, Zhang, and Gonzalez teach the limitations of claim 16 for the same reasons discussed above with respect to claim 8. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to SAMUEL DAVID BAYNES whose telephone number is (571)272-0607. The examiner can normally be reached Monday - Friday 8:00 am - 5:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Stephen R Koziol can be reached at (408)918-7630. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SDB/ Samuel D. Baynes Examiner, Art Unit 2665 /Stephen R Koziol/Supervisory Patent Examiner, Art Unit 2665
Read full office action

Prosecution Timeline

Oct 31, 2024
Application Filed
Jul 24, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12685438
METHOD AND APPARATUS FOR DETECTING PENETRATION DEPTH OF RIBOFLAVIN IN CORNEA
2y 2m to grant Granted Jul 21, 2026
Patent 12688557
EXTENDED U-NET FOR MULTI-INFORMATION EXTRACTION AND APPLICATION METHOD THEREOF IN LOW-DOSE X-RAY IMAGING
1y 11m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 2 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
86%
Grant Probability
99%
With Interview (+25.0%)
2y 5m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 7 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month