Prosecution Insights
Last updated: August 17, 2026
Application No. 18/935,234

JOINT ASSET AND DEFECT DETECTION MACHINE LEARNING MODEL

Non-Final OA §101§103
Filed
Nov 01, 2024
Priority
Nov 06, 2023 — provisional 63/596,355
Examiner
PHAM, NHUT HUY
Art Unit
Tech Center
Assignee
X Development LLC
OA Round
1 (Non-Final)
81%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 81% — above average
81%
Career Allowance Rate
58 granted / 72 resolved
+20.6% vs TC avg
Strong +25% interview lift
Without
With
+24.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
24 currently pending
Career history
91
Total Applications
across all art units

Statute-Specific Performance

§101
9.4%
-30.6% vs TC avg
§103
59.4%
+19.4% vs TC avg
§102
13.8%
-26.2% vs TC avg
§112
15.0%
-25.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 72 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION The United States Patent & Trademark Office appreciates the application that is submitted by the inventor/assignee. The United States Patent & Trademark Office reviewed the following application and has made the following comments below. Information Disclosure Statement The information disclosure statement (IDS) submitted on 02/11/2026 is considered and attached. Priority This application claims benefit of provisional benefit of application 63/596,355 on 11/06/2023. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-5, 7, 15, 17-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. When reviewing independent claim 1, and based upon consideration of all of the relevant factors with respect to the claim as a whole, claim(s) 1 are held to claim an abstract idea without reciting elements that amount to significantly more than the abstract idea and is/are therefore rejected as ineligible subject matter under 35 U.S.C. 101. The Examiner will analyze Claim 1, and similar rationale applies to independent claim/s 15. The rationale, under MPEP § 2106, for this finding is explained below: The claimed invention (1) must be directed to one of the four statutory categories, and (2) must not be wholly directed to subject matter encompassing a judicially recognized exception, as defined below. The following two step analysis is used to evaluate these criteria. Step 1: Is the claim directed to one of the four patent-eligible subject matter categories: process, machine, manufacture, or composition of matter? When examining the claim under 35 U.S.C. 101, the Examiner interprets that the claims is related to a process since the claim is directed to a method. Step 2a, Prong 1: Does the claim wholly embrace a judicially recognized exception, which includes laws of nature, physical phenomena, and abstract ideas, or is it a particular practical application of a judicial exception? YES, the claims are directed toward a mental process (i.e., abstract idea). With regard to STEP 2A (PRONG 1), the guidelines provide three groupings of subject matter that are considered abstract ideas: Mathematical concepts – mathematical relationships, mathematical formulas or equations, mathematical calculations; Certain methods of organizing human activity – fundamental economic principles or practices (including hedging, insurance, mitigating risk); commercial or legal interactions (including agreements in the form of contracts; legal obligations; advertising, marketing or sales activities or behaviors; business relations); managing personal behavior or relationships or interactions between people (including social activities, teaching, and following rules or instructions); and Mental processes – concepts that are practicably performed in the human mind (including an observation, evaluation, judgment, opinion). The method in claim 1 comprise a mental process that can be practicably performed in the human mind therefore, an abstract idea. Claim 1 recites: Generating embeddings for one or more classification labels of the one or more objects in the input images, each embedding corresponding to a classification label and comprising a mapping between the classification label and a subset of feature vectors (a human can review an image, create an image embedding/patch corresponding to an object label by drawing an outline for an image region associated with an object and annotating the type of the object onto the outline, using a pen and paper, as a mental process as an abstract idea); determining a likelihood of an object from the one or more objects in the input image containing a type of defect (a human can review an image region that contains an object and annotation about object type, determine whether the object is defective and give an estimation regarding his confidence of the determination, using a pen and paper, as a mental process as an abstract idea); These limitations, as drafted, is a simple process that, under their broadest reasonable interpretation, covers performance of the limitations in the mind or by a human. The Examiner notes that under MPEP 2106.04(a)(2)(III), the courts consider a mental process (thinking) that “can be performed in the human mind, or by a human using a pen and paper" to be an abstract idea. CyberSource Corp. v. Retail Decisions, Inc., 654 F.3d 1366, 1372, 99 USPQ2d 1690, 1695 (Fed. Cir. 2011). As the Federal Circuit explained, "methods which can be performed mentally, or which are the equivalent of human mental work, are unpatentable abstract ideas the ‘basic tools of scientific and technological work’ that are open to all.’" 654 F.3d at 1371, 99 USPQ2d at 1694 (citing Gottschalk v. Benson, 409 U.S. 63, 175 USPQ 673 (1972)). See also Mayo Collaborative Servs. v. Prometheus Labs. Inc., 566 U.S. 66, 71, 101 USPQ2d 1961, 1965 ("‘[M]ental processes[] and abstract intellectual concepts are not patentable, as they are the basic tools of scientific and technological work’" (quoting Benson, 409 U.S. at 67, 175 USPQ at 675)); Parker v. Flook, 437 U.S. 584, 589, 198 USPQ 193, 197 (1978) (same). The courts do not distinguish between mental processes that are performed entirely in the human mind and mental processes that require a human to use a physical aid (e.g., pen and paper or a slide rule) to perform the claim limitation. See, e.g., Benson, 409 U.S. at 67, 65, 175 USPQ at 674-75, 674 (noting that the claimed "conversion of [binary-coded decimal] numerals to pure binary numerals can be done mentally," i.e., "as a person would do it by head and hand."); Synopsys, Inc. v. Mentor Graphics Corp., 839 F.3d 1138, 1139, 120 USPQ2d 1473, 1474 (Fed. Cir. 2016) (holding that claims to a mental process of "translating a functional description of a logic circuit into a hardware component description of the logic circuit" are directed to an abstract idea, because the claims "read on an individual performing the claimed steps mentally or with pencil and paper"). Nor do the courts distinguish between claims that recite mental processes performed by humans and claims that recite mental processes performed on a computer. As the Federal Circuit has explained, "[c]ourts have examined claims that required the use of a computer and still found that the underlying, patent-ineligible invention could be performed via pen and paper or in a person’s mind." Versata Dev. Group v. SAP Am., Inc., 793 F.3d 1306, 1335, 115 USPQ2d 1681, 1702 (Fed. Cir. 2015). See also Intellectual Ventures I LLC v. Symantec Corp., 838 F.3d 1307, 1318, 120 USPQ2d 1353, 1360 (Fed. Cir. 2016) (‘‘[W]ith the exception of generic computer-implemented steps, there is nothing in the claims themselves that foreclose them from being performed by a human, mentally or with pen and paper.’’); Mortgage Grader, Inc. v. First Choice Loan Servs. Inc., 811 F.3d 1314, 1324, 117 USPQ2d 1693, 1699 (Fed. Cir. 2016) (holding that computer-implemented method for "anonymous loan shopping" was an abstract idea because it could be "performed by humans without a computer"). Because both product and process claims may recite a "mental process", the phrase "mental processes" should be understood as referring to the type of abstract idea, and not to the statutory category of the claim. The courts have identified numerous product claims as reciting mental process-type abstract ideas, for instance the product claims to computer systems and computer-readable media in Versata Dev. Group. v. SAP Am., Inc., 793 F.3d 1306, 115 USPQ2d 1681 (Fed. Cir. 2015). As such, a person could identify an object in an image, and create a rectangle to capture a region of interest (bounding box of object) associated with an object; determine, based on identified region, whether the object in the region is damaged, and give an estimation regarding his confidence of the determination. The mere nominal recitation that the various steps are being executed by a device/in a device (e.g. processing unit) does not take the limitations out of the mental process grouping. Thus, the claims recite a mental process. If a claim limitation, under its broadest reasonable interpretation, covers performance of a mental step which could be performed with a simple tool such as a pen and paper, then it falls within the “mental steps” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. Step 2a, Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? NO, the claims do not recite additional elements that integrate the judicial exception into a practical application. With regard to STEP 2A (prong 2), whether the claim recites additional elements that integrate the judicial exception into a practical application, the guidelines provide the following exemplary considerations that are indicative that an additional element (or combination of elements) may have integrated the judicial exception into a practical application: an additional element reflects an improvement in the functioning of a computer, or an improvement to other technology or technical field; an additional element that applies or uses a judicial exception to affect a particular treatment or prophylaxis for a disease or medical condition; an additional element implements a judicial exception with, or uses a judicial exception in conjunction with, a particular machine or manufacture that is integral to the claim; an additional element effects a transformation or reduction of a particular article to a different state or thing; and an additional element applies or uses the judicial exception in some other meaningful way beyond generally linking the use of the judicial exception to a particular technological environment, such that the claim as a whole is more than a drafting effort designed to monopolize the exception. While the guidelines further state that the exemplary considerations are not an exhaustive list and that there may be other examples of integrating the exception into a practical application, the guidelines also list examples in which a judicial exception has not been integrated into a practical application: an additional element merely recites the words “apply it” (or an equivalent) with the judicial exception, or merely includes instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea; an additional element adds insignificant extra-solution activity to the judicial exception; and an additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use. Claims 1-5, 7, 15, 17-20 do not recite any of the exemplary considerations that are indicative of an abstract idea having been integrated into a practical application. The steps “generating, by one or more deep neural networks…”, “determining, by a plurality of defect classifier…” amounts to merely using a computer as a tool to perform the claimed mental process. Implementing an abstract idea on a computer does not integrate a judicial exception into a practical application (See MPEP 2106.05(f)). The steps “receiving input data comprising an input image …”, “generating an output image comprising …” merely constitutes activity involving data gathering and data outputting. Such extra-solution activity does not integrate the abstract idea into a practical application. Please see MPEP §2106.05(g). Claim 15 recites “one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations” amounts to merely using a computer as a tool to perform the claimed mental process. Implementing an abstract idea on a computer does not integrate a judicial exception into a practical application (See MPEP 2106.05(f)). These limitations are recited at a high level of generality (i.e. as a general action or change being taken based on the results of the acquiring step) and amounts to mere post solution actions, which is a form of insignificant extra-solution activity. Further, the claims are claimed generically and are operating in their ordinary capacity such that they do not use the judicial exception in a manner that imposes a meaningful limit on the judicial exception. Accordingly, even in combination, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. Step 2b: If a judicial exception into a practical application is not recited in the claim, the Examiner must interpret if the claim recites additional elements that amount to significantly more than the judicial exception. With regard to STEP 2B, whether the claims recite additional elements that provide significantly more than the recited judicial exception, the guidelines specify that the pre-guideline procedure is still in effect. Specifically, that examiners should continue to consider whether an additional element or combination of elements: adds a specific limitation or combination of limitations that are not well-understood, routine, conventional activity in the field, which is indicative that an inventive concept may be present; or simply appends well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception, which is indicative that an inventive concept may not be present. With regard to (2b) the Guidance provided the following examples of limitations that may be enough to qualify as “significantly more" when recited in a claim with a judicial exception: Improvement to another technology or technical field Improvement to functioning of computer itself and/or applying the judicial exception with, or by use of, a particular machine Effecting a transformation or reduction of a particular article to a different state or thing. Adding a specific limitation other that what is well understood, routine and conventional in the field, or adding unconventional steps that confine the claim to a particular useful application Meaningful limitation beyond generally linking the use of an abstract idea to a particular technological environment. The Guidance further set forth limitations that were found not to be enough to qualify as “significantly more” when recited in a claim with a judicial exception include: Adding words to “apply it” (or an equivalent) with the judicial exception or mere instructions to implement abstract ideas on a computer Simply appending well-understood, routine and conventional activities previously known to the industry specified at a high level of generality to the judicial exception, e.g. a claim to an abstract idea requiring no more than a generic Computer to perform generic computer functions that are well -understood, routine and conventional activities previously known to the industry. Adding insignificant extra-solution activity to the judicial exception, e.g. mere data gathering in conjunction with a law of nature or abstract idea Generally linking the use of the judicial exception to a particular technological environment or field of use. Claims 1-5, 7, 15, 17-20 do not recite any additional elements that are not well-understood, routine or conventional. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The above identified additional computer components, using instructions to apply the judicial exception, are merely generic computer components that are well-known, routine, and conventional as is evidenced by Bancorp Services v. Sun Life (Fed. Cir. 2012) and Alice Corp. v. CLS Bank (2014). Thus, since claims 1 and 15 are: (a) directed toward an abstract idea, (b) do not recite additional elements that integrate the judicial exception into a practical application, and (c) do not recite additional elements that amount to significantly more than the judicial exception, claims 1 and 15 are not eligible subject matter under 35 U.S.C 101. Similar analysis is made for the dependent claims 2-5, 7 and 17-20 and the dependent claims are similarly identified as: being directed towards an abstract idea, not reciting additional elements that integrate the judicial exception into a practical application, and not reciting additional elements that amount to significantly more than the judicial exception. Regarding Claim 6, the claim depends on claim 1. Thus, claim 6 recites “Mental Processes”. Claim 6 further recites additional elements: “generating, using the output images, a model representation of an electric grid comprising the one or more utility assets from the input images” The combination of the additional elements integrates the “Mental Processes” abstract idea into a practical application. Specifically, as discussed in the paragraph [0056] and [0090] of the originally filed specification of the subject application, utilizing the object detection and defect classification result to obtain output image with annotation, and use the obtained output image to generate model representation or an interactive image map of the electric grid constitutes an improvement to the technical field of object detection. As such, the additional elements of claim 6 [in combination with all the limitations of claim 1] integrate the “Mental Processes” into a practical application. Therefore, claim 6 recites eligible subject matter. Regarding Claim 16, the claim depends on claim 15. Thus, claim 16 recites “Mental Processes”. Claim 16 further recites additional elements: “obtaining a plurality of training examples …”, “generating a plurality of groupings …”, “applying … an activation function …”, “generating … a first additional class …”, “generating a first additional grouping …”, “sampling … the feature data …”, “generating … a predicted annotation”, “updating one or more weights of at least one defect classifier …” The combination of the additional elements integrates the “Mental Processes” abstract idea into a practical application. Specifically, as discussed in the paragraph [0058] and [0062] of the originally filed specification of the subject application, utilizing dynamically grouping training examples, sampling features, and updating classifier weights to handle class imbalances constitutes an improvement to the technical field of training neural networks for object detection and classification. As such, the additional elements of claim 16 [in combination with all the limitations of claim 15] integrate the “Mental Processes” into a practical application. Therefore, claim 16 recites eligible subject matter. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-5, 7, 15, 17-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Nguyen et al. (Nguyen, Van Nhan and Jenssen, Robert, and Davide Roverso. "Intelligent monitoring and inspection of power line components powered by UAVs and deep learning." IEEE, published 2019, hereinafter Nguyen) in view of Karan (US-20210295082-A1, hereinafter Karan). CLAIM 1 In regards to Claim 1, Nguyen teaches a method for joint asset and defect detection (Nguyen, Abstract: “vision-based power line inspection: (i) the lack of training data; (ii) class imbalance; and (iii) the detection of small components and faults.”), the method comprising: receiving input data comprising an input image of a utility asset, the input image comprising one or more objects (Nguyen, page 14, right col, section B. ACQUIRED OPTICAL IMAGES: “Images are collected directly using cameras mounted on the UAVs. The UAVs are flown along power lines and circled around power masts to take pictures of the masts from different angles. For each power mast, around 20 images at 6048x4032 resolution are collected”); Nguyen does not explicitly disclose generating, by one or more deep neural networks, embeddings for one or more classification labels of the one or more objects in the input images, each embedding corresponding to a classification label and comprising a mapping between the classification label and a subset of feature vectors; Karan is in the same field of art of object detection using neural network. Further, Karan teaches generating, by one or more deep neural networks (Karan, ¶ [0031-0032]: “ some of the known labeled bounding boxes from the training datasets and the proposed bounding boxes can then be extracted using, for example, a convolutional neural network (CNN)”), embeddings for one or more classification labels of the one or more objects in the input images (Karan, ¶ [0031-0032 and 0052-0055]: “0054: the semantic space generator module 130 created a feature vector representation for the extracted features of each of the object bounding boxes and created a word vector representation for each of the respective object class labels associated with each of the object bounding boxes … 0055: The semantic space generator module 130 then embedded the created feature vectors and the word vectors into the semantic embedding space” Karan teaches create word vector, a representation for a class label of object detected in input image, then embed the vector into a semantic embedding space), each embedding corresponding to a classification label (Karan, ¶ [0031-0032 and 0052-0055]: “0054: a word vector representation for each of the respective object class labels with each of the object bounding boxes” ) and comprising a mapping between the classification label and a subset of feature vectors (Karan, ¶ [0032, 0034 and 0055]: “0032: In the semantic embedding space, embedded vectors that are related are closer together in the geometric embedding space than unrelated vectors; 0034: An object class, ŷi, for extracted features of a proposed bounding box, PNG media_image1.png 38 21 media_image1.png Greyscale i, can be predicted by finding a nearest embedded object class label based on a determined similarity score determined between the extracted features of the proposed bounding box and embedded object class labels corresponding to features of bounding boxes used to train the semantic embedding space”, see equation (3); Karan teaches embedded feature vector and embedded class vector are associated together based on their distance to each other in the semantic space, one feature vector will map to the nearest class vector and vice versa); Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Nguyen by incorporating method to generate embedded feature vector and class vector that is taught by Karan, to make a system to classify unseen objects; thus, one of ordinary skilled in the art would be motivated to combine the references since among its several aspects, the present invention recognizes there is a need to improve accuracy of object detection (Karan, ¶ [0025]: “classes of unseen objects can be determined without the need for training any of the unseen object classes. In addition, in accordance with present principles, background classes are defined to enable more accurate detection of objects and classes for unseen object classes”). The combination of Nguyen and Karan then teaches determining, by a plurality of defect classifiers (Nguyen, page 16, section C: “Finally, the detected top caps, poles, and cross arms are cropped from the input images and passed through their corresponding classifiers to identify faults.”, see annotated FIG. 5 below.), PNG media_image2.png 346 1245 media_image2.png Greyscale a likelihood of an object from the one or more objects in the input image containing a type of defect (Nguyen, page 16, right col, Algorithm 1: “each classifier CLF in C takes an image as input and outputs a label (cls_label) and a confidence score cls_conf” Nguyen teaches multiple defect classifiers that will output a type of defect for the object in input image, included a confidence score of the classification) (Karan, ¶ [0034]: “The trained semantic embedding space can be used to compute a similarity measure between a projected bounding box feature, ψi, of the proposed bounding box and an object class embedding (i.e., embedded features of an object bounding box), wj, for an object class label, yi.” Karan teaches computing similarity measure between a feature vector and a class vector), wherein each defect classifier from the plurality of defect classifiers is trained to determine a type of defect (Nguyen, page 15, left col, 2nd and 3rd paragraph: “The second dataset (DS2_Tc), which is used for training missing top cap detectors ... The third dataset (DS3_Po), which is used for training cracks in poles and woodpecker damage on poles detectors … The final dataset (DS4_Cr), which is used for training cracks on cross arms and rot damage on cross arms detectors,…” Nguyen teaches training dataset for different type of defects) based on the embeddings for the one or more classification labels (Karan,¶ [027-0031]: “For training the semantic embedding space, labeled bounding boxes of known/seen object classes as well as corresponding class labels of the bounding boxes are implemented. That is, in some embodiments, features of objects in provided/proposed labeled bounding boxes are extracted and embedded into the semantic embedding space along with a respective class label representative of the features of the bounding box”, ¶ [0065]: “Once the semantic embedding space has been created and trained in accordance with the present principles, zero-shot object detection can be performed” Karan teaches training a semantic embedding space from training dataset, the embedding space is then used to perform object detection); and generating an output image comprising a plurality of bounding boxes for the one or more objects in the input image, and an annotation corresponding a respective object from the one or more objects in the input image. (Nguyen, page 16, right col: “Output: A list of detected and classified components O (each item in O contains a label, bounding box coordinates, and a confidence score)”, see modified FIG. 7 below. Nguyen teaches outputting an image comprises bounding boxes for detected objects, each bounding box includes a label and a confidence score of the detection) PNG media_image3.png 1048 1160 media_image3.png Greyscale Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. CLAIM 2 Regarding claim 2, the combination of Nguyen and Karan teaches the method of Claim 1. In addition, the combination of Nguyen and Karan teaches generating, by one or more deep neural networks (Nguyen, page 17, right col, section A: “The component classifiers are built by fine-tuning the ResNet50_cvgj [28] model, …”), the plurality of bounding boxes for the one or more objects in the input image (Nguyen, page 16, Algorithm 1: “Component detector CD(I) that outputs labels (c_labels), bounding box coordinates (c_coords), and confidence scores (c_confs) of the detected components”), wherein each bounding box in the plurality of bounding boxes corresponds to an object from the one or more objects in the input image for the utility asset. (Nguyen, page 16, right col: “Output: A list of detected and classified components O (each item in O contains a label, bounding box coordinates, and a confidence score)”, see modified FIG. 7 above. Nguyen teaches outputting bounding boxes for detected sub components of the utility pole, such as top cap, cross arm, insulator, ...) CLAIM 3 Regarding claim 3, the combination of Nguyen and Karan teaches the method of Claim 1. In addition, the combination of Nguyen and Karan teaches generating, by one or more deep neural networks (Nguyen, page 17, right col, section A: “The component classifiers are built by fine-tuning the ResNet50_cvgj [28] model, …”), asset label data for the one or more objects in the input image (Nguyen, page 16, Algorithm 1: “Component detector CD(I) that outputs labels (c_labels), bounding box coordinates (c_coords), and confidence scores (c_confs) of the detected components”), wherein the asset label data comprises the one or more classification labels corresponding to the one or more objects in the input image, each classification label representing a type of utility asset. (Nguyen, page 16, right col: “Output: A list of detected and classified components O (each item in O contains a label, bounding box coordinates, and a confidence score)”, see modified FIG. 7 above. Nguyen teaches outputting bounding boxes for detected sub components of the utility pole, such as top cap, cross arm, insulator, …; each bounding box contains a label indicates the component) CLAIM 4 Regarding claim 4, the combination of Nguyen and Karan teaches the method of Claim 1. In addition, the combination of Nguyen and Karan teaches generating, by one or more deep neural networks, asset feature data for one or more objects in the input image (Nguyen, page 13, right col, section 2) FASTER R-CNN: “First, a base network (e.g., ResNet [3]) is utilized to extract features from images” Nguyen teaches extracting image features), wherein the asset feature data comprises a plurality of feature vectors, the feature vectors representing features of the one or more objects in the input image(Karan, ¶ [0031-0032]: “Features, such as deep features, of at least some of the known labeled bounding boxes from the training datasets and the proposed bounding boxes can then be extracted using, for example, a convolutional neural network (CNN). The extracted features for each bounding box, can be denoted as ϕ(bi) ∈ RD1. A respective feature vector can then be created for the extracted bounding box features…”, ¶ [0054-0055]: “the semantic space generator module 130 created a feature vector representation for the extracted features of each of the object bounding boxes” Karan teaches extracting features of from bounding boxes of detected objects, and creating a corresponding feature vector), and wherein the subset of feature vectors comprises at least one of the plurality of feature vectors. (Karan, ¶ [0032, 0034 and 0055]: “0032: In the semantic embedding space, embedded vectors that are related are closer together in the geometric embedding space than unrelated vectors; 0034: An object class, ŷi, for extracted features of a proposed bounding box, PNG media_image1.png 38 21 media_image1.png Greyscale i, can be predicted by finding a nearest embedded object class label based on a determined similarity score determined between the extracted features of the proposed bounding box and embedded object class labels corresponding to features of bounding boxes used to train the semantic embedding space”, see equation (3); Karan teaches one of feature vector (from the plurality of bounding boxes feature vectors ) is associated with the nearest class vector) CLAIM 5 Regarding claim 5, the combination of Nguyen and Karan teaches the method of Claim 1. In addition, the combination of Nguyen and Karan teaches the corresponding bounding box for an object indicates a position of the respective object in the output image. (Nguyen, page 16, Algorithm 1: “Mast detector MD(I) that outputs bounding box coordinates (m_coords) and confidence scores (m_confs) of the detected masts, Component detector CD(I) that outputs labels (c_labels), bounding box coordinates (c_coords), and confidence scores (c_confs) of the detected components”, see modified FIG. 7 above. Nguyen teaches locating the position of objects using bounding boxes that encloses the image region of detected object) CLAIM 7 PNG media_image4.png 809 1491 media_image4.png Greyscale Regarding claim 7, the combination of Nguyen and Karan teaches the method of Claim 1. In addition, the combination of Nguyen and Karan teaches the annotation comprises a classification label and a likelihood associated with the classification label. (Nguyen, page 16, Algorithm 1: “Mast detector MD(I) that outputs bounding box coordinates (m_coords) and confidence scores (m_confs) of the detected masts, Component detector CD(I) that outputs labels (c_labels), bounding box coordinates (c_coords), and confidence scores (c_confs) of the detected components”, see modified FIG. 7 below), the classification label indicating an asset type (see modified FIG. 7 below, first part of label) and defect status for the respective object (see modified FIG. 7 below, second part of label), and the likelihood associated with the classification label represents a probability of the respective object in the input image matching the asset type and the defect status. (see modified FIG. 7 below, last part of label) CLAIM 15 In regards to Claim 15, Nguyen teaches a system for joint asset and defect detection (Nguyen, Abstract: “vision-based power line inspection: (i) the lack of training data; (ii) class imbalance; and (iii) the detection of small components and faults.”), the system comprising: one or more computers and one or more storage devices (Nguyen, page 14, right col, section B: “The images are uploaded to the Microsoft Azure cloud” The Examiner notes Microsoft Azure cloud is a cloud platform provides both computing and storage system) storing instructions (Nguyen, page 16, right col, Algorithm 1) that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising: receiving input data comprising an input image of a utility asset, the input image comprising one or more objects (Nguyen, page 14, right col, section B. ACQUIRED OPTICAL IMAGES: “Images are collected directly using cameras mounted on the UAVs. The UAVs are flown along power lines and circled around power masts to take pictures of the masts from different angles. For each power mast, around 20 images at 6048x4032 resolution are collected”); Nguyen does not explicitly disclose generating, by one or more deep neural networks, embeddings for one or more classification labels of the one or more objects in the input images, each embedding corresponding to a classification label and comprising a mapping between the classification label and a subset of feature vectors; Karan is in the same field of art of object detection using neural network. Further, Karan teaches generating, by one or more deep neural networks (Karan, ¶ [0031-0032]: “ some of the known labeled bounding boxes from the training datasets and the proposed bounding boxes can then be extracted using, for example, a convolutional neural network (CNN)”), embeddings for one or more classification labels of the one or more objects in the input images (Karan, ¶ [0031-0032 and 0052-0055]: “0054: the semantic space generator module 130 created a feature vector representation for the extracted features of each of the object bounding boxes and created a word vector representation for each of the respective object class labels associated with each of the object bounding boxes … 0055: The semantic space generator module 130 then embedded the created feature vectors and the word vectors into the semantic embedding space” Karan teaches create word vector, a representation for a class label of object detected in input image, then embed the vector into a semantic embedding space), each embedding corresponding to a classification label (Karan, ¶ [0031-0032 and 0052-0055]: “0054: a word vector representation for each of the respective object class labels with each of the object bounding boxes” ) and comprising a mapping between the classification label and a subset of feature vectors (Karan, ¶ [0032, 0034 and 0055]: “0032: In the semantic embedding space, embedded vectors that are related are closer together in the geometric embedding space than unrelated vectors; 0034: An object class, ŷi, for extracted features of a proposed bounding box, PNG media_image1.png 38 21 media_image1.png Greyscale i, can be predicted by finding a nearest embedded object class label based on a determined similarity score determined between the extracted features of the proposed bounding box and embedded object class labels corresponding to features of bounding boxes used to train the semantic embedding space”, see equation (3); Karan teaches embedded feature vector and embedded class vector are associated together based on their distance to each other in the semantic space, one feature vector will map to the nearest class vector and vice versa); Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Nguyen by incorporating method to generate embedded feature vector and class vector that is taught by Karan, to make a system to classify unseen objects; thus, one of ordinary skilled in the art would be motivated to combine the references since among its several aspects, the present invention recognizes there is a need to improve accuracy of object detection (Karan, ¶ [0025]: “classes of unseen objects can be determined without the need for training any of the unseen object classes. In addition, in accordance with present principles, background classes are defined to enable more accurate detection of objects and classes for unseen object classes”). The combination of Nguyen and Karan then teaches determining, by a plurality of defect classifiers (Nguyen, page 16, section C: “Finally, the detected top caps, poles, and cross arms are cropped from the input images and passed through their corresponding classifiers to identify faults.”, see annotated FIG. 5 below.), PNG media_image2.png 346 1245 media_image2.png Greyscale a likelihood of an object from the one or more objects in the input image containing a type of defect (Nguyen, page 16, right col, Algorithm 1: “each classifier CLF in C takes an image as input and outputs a label (cls_label) and a confidence score cls_conf” Nguyen teaches multiple defect classifiers that will output a type of defect for the object in input image, included a confidence score of the classification) (Karan, ¶ [0034]: “The trained semantic embedding space can be used to compute a similarity measure between a projected bounding box feature, ψi, of the proposed bounding box and an object class embedding (i.e., embedded features of an object bounding box), wj, for an object class label, yi.” Karan teaches computing similarity measure between a feature vector and a class vector), wherein each defect classifier from the plurality of defect classifiers is trained to determine a type of defect (Nguyen, page 15, left col, 2nd and 3rd paragraph: “The second dataset (DS2_Tc), which is used for training missing top cap detectors ... The third dataset (DS3_Po), which is used for training cracks in poles and woodpecker damage on poles detectors … The final dataset (DS4_Cr), which is used for training cracks on cross arms and rot damage on cross arms detectors,…” Nguyen teaches training dataset for different type of defects) based on the embeddings for the one or more classification labels (Karan,¶ [027-0031]: “For training the semantic embedding space, labeled bounding boxes of known/seen object classes as well as corresponding class labels of the bounding boxes are implemented. That is, in some embodiments, features of objects in provided/proposed labeled bounding boxes are extracted and embedded into the semantic embedding space along with a respective class label representative of the features of the bounding box”, ¶ [0065]: “Once the semantic embedding space has been created and trained in accordance with the present principles, zero-shot object detection can be performed” Karan teaches training a semantic embedding space from training dataset, the embedding space is then used to perform object detection); and PNG media_image3.png 1048 1160 media_image3.png Greyscale generating an output image comprising a plurality of bounding boxes for the one or more objects in the input image, and an annotation corresponding a respective object from the one or more objects in the input image. (Nguyen, page 16, right col: “Output: A list of detected and classified components O (each item in O contains a label, bounding box coordinates, and a confidence score)”, see modified FIG. 7 below. Nguyen teaches outputting an image comprises bounding boxes for detected objects, each bounding box includes a label and a confidence score of the detection) Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. CLAIM 17 Regarding claim 17, the combination of Nguyen and Karan teaches the system of Claim 15. In addition, the combination of Nguyen and Karan teaches generating, by one or more deep neural networks (Nguyen, page 17, right col, section A: “The component classifiers are built by fine-tuning the ResNet50_cvgj [28] model, …”), the plurality of bounding boxes for the one or more objects in the input image (Nguyen, page 16, Algorithm 1: “Component detector CD(I) that outputs labels (c_labels), bounding box coordinates (c_coords), and confidence scores (c_confs) of the detected components”), wherein each bounding box in the plurality of bounding boxes corresponds to an object from the one or more objects in the input image for the utility asset. (Nguyen, page 16, right col: “Output: A list of detected and classified components O (each item in O contains a label, bounding box coordinates, and a confidence score)”, see modified FIG. 7 above. Nguyen teaches outputting bounding boxes for detected sub components of the utility pole, such as top cap, cross arm, insulator, ...) CLAIM 18 Regarding claim 18, the combination of Nguyen and Karan teaches the system of Claim 15. In addition, the combination of Nguyen and Karan teaches generating, by one or more deep neural networks (Nguyen, page 17, right col, section A: “The component classifiers are built by fine-tuning the ResNet50_cvgj [28] model, …”), asset label data for the one or more objects in the input image (Nguyen, page 16, Algorithm 1: “Component detector CD(I) that outputs labels (c_labels), bounding box coordinates (c_coords), and confidence scores (c_confs) of the detected components”), wherein the asset label data comprises the one or more classification labels corresponding to the one or more objects in the input image, each classification label representing a type of utility asset. (Nguyen, page 16, right col: “Output: A list of detected and classified components O (each item in O contains a label, bounding box coordinates, and a confidence score)”, see modified FIG. 7 above. Nguyen teaches outputting bounding boxes for detected sub components of the utility pole, such as top cap, cross arm, insulator, …; each bounding box contains a label indicates the component) CLAIM 19 Regarding claim 19, the combination of Nguyen and Karan teaches the system of Claim 15. In addition, the combination of Nguyen and Karan teaches generating, by one or more deep neural networks, asset feature data for one or more objects in the input image (Nguyen, page 13, right col, section 2) FASTER R-CNN: “First, a base network (e.g., ResNet [3]) is utilized to extract features from images” Nguyen teaches extracting image features), wherein the asset feature data comprises a plurality of feature vectors, the feature vectors representing features of the one or more objects in the input image(Karan, ¶ [0031-0032]: “Features, such as deep features, of at least some of the known labeled bounding boxes from the training datasets and the proposed bounding boxes can then be extracted using, for example, a convolutional neural network (CNN). The extracted features for each bounding box, can be denoted as ϕ(bi) ∈ RD1. A respective feature vector can then be created for the extracted bounding box features…”, ¶ [0054-0055]: “the semantic space generator module 130 created a feature vector representation for the extracted features of each of the object bounding boxes” Karan teaches extracting features of from bounding boxes of detected objects, and creating a corresponding feature vector), and wherein the subset of feature vectors comprises at least one of the plurality of feature vectors. (Karan, ¶ [0032, 0034 and 0055]: “0032: In the semantic embedding space, embedded vectors that are related are closer together in the geometric embedding space than unrelated vectors; 0034: An object class, ŷi, for extracted features of a proposed bounding box, PNG media_image1.png 38 21 media_image1.png Greyscale i, can be predicted by finding a nearest embedded object class label based on a determined similarity score determined between the extracted features of the proposed bounding box and embedded object class labels corresponding to features of bounding boxes used to train the semantic embedding space”, see equation (3); Karan teaches one of feature vector (from the plurality of bounding boxes feature vectors ) is associated with the nearest class vector) CLAIM 20 Regarding claim 20, the combination of Nguyen and Karan teaches the system of Claim 15. In addition, the combination of Nguyen and Karan teaches the annotation comprises a classification label and a likelihood associated with the classification label. (Nguyen, page 16, Algorithm 1: “Mast detector MD(I) that outputs bounding box coordinates (m_coords) and confidence scores (m_confs) of the detected masts, Component detector CD(I) that outputs labels (c_labels), bounding box coordinates (c_coords), and confidence scores (c_confs) of the detected components”, see modified FIG. 7 below), the classification label indicating an asset type (see modified FIG. 7 below, first part of label) and defect status for the respective object (see modified FIG. 7 below, second part of label), and the likelihood associated with the classification label represents a probability of the respective object in the input image matching PNG media_image4.png 809 1491 media_image4.png Greyscale the asset type and the defect status. (see modified FIG. 7 below, last part of label) Allowable Subject Matter Claims 6 and 16 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claims 8-14 are allowed. Pertinent Arts The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Minderer et al. (US-20230360365-A1) which is directed to an object detection system that can find objects in an image and label them using user-provided “query embeddings.” A query embedding can come from text, such as “giraffe” or “bird,” or from a query image showing an example object. The system first turns the input image into a set of object embeddings, which act like candidate object regions. It then predicts where each object is located by producing localization data, such as a bounding box. Next, it compares each object embedding against the query embeddings to estimate which category best matches each region. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to NHUT HUY (JEREMY) PHAM whose telephone number is (703)756-5797. The examiner can normally be reached Mo - Fr. 8:30am - 6pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, O'Neal Mistry can be reached on (313)446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /NHUT HUY PHAM/Examiner, Art Unit 2674 /ONEAL R MISTRY/Supervisory Patent Examiner, Art Unit 2674
Read full office action

Prosecution Timeline

Nov 01, 2024
Application Filed
Jul 28, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705688
APPARATUS AND METHOD WITH IMAGE PROCESSING FOR SPARSE DUAL-PIXEL IMAGE DATA
3y 9m to grant Granted Aug 11, 2026
Patent 12705913
METHOD AND ELECTRONIC DEVICE FOR RECOGNIZING TEXT IN IMAGE
3y 8m to grant Granted Aug 11, 2026
Patent 12700087
SYSTEMS AND METHODS FOR IMAGE GENERATION
3y 3m to grant Granted Aug 04, 2026
Patent 12700082
TRAINING A MACHINE LEARNING PROCESS FOR USE IN EVALUATING A SUBSTRATE
2y 6m to grant Granted Aug 04, 2026
Patent 12688672
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND PROGRAM
3y 1m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
81%
Grant Probability
99%
With Interview (+24.7%)
2y 10m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 72 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month