CTNF 18/794,164 CTNF 100542 Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Claim Rejections - 35 USC § 101 07-04-01 AIA 07-04 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim(s) 1-14 is/are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter as follows. Regarding claim(s) 1, 13, 14, the claims are directed to an abstract idea, namely mathematical operations and information processing. The claims are not integrated into a practical application and the claims lack an inventive concept. Furthermore, claim(s) 2-12 are also directed to an abstract idea, specifically, mathematical operations, image processing, data transmission, and object detection. Claim Rejections - 35 USC § 102 07-07-aia AIA 07-07 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – 07-08-aia AIA (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. 07-12-aia AIA (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. 07-15 AIA Claim s 1-6, 9-11, 13-14 are rejected under 35 U.S.C. 102( a)(1) and 35 U.S.C. 102(a)(2 ) as being anticipated by Turkelson et al. (US 20200193206 A1), hereinafter Turkelson . Regarding claim 1, Turkelson teaches A method for classifying input measurement data with respect to a given task using a given classifier, the method comprising the following steps: (Para. 6 see "Some aspects include a process including: obtaining, with a computer system, an image depicting an object within a context, wherein: the image is captured by a mobile computing device, the object is a member of an ontology of objects including a plurality of objects, and the context is a member of an ontology of contexts including a plurality of contexts; determining, with the computer system, with a trained context classification model, the context depicted by the image; determining, with the computer system, with a trained object recognition model, a first object identifier of the object based on the image and the context; and causing, with the computer system, the first object identifier of the object to be stored in memory." Para. 8 see "Some aspects include a tangible, non-transitory, machine-readable medium storing instructions that when executed by a data processing apparatus cause the data processing apparatus to perform operations including each of the above-mentioned processes." Para. 9 see "Some aspects include a system, including: one or more processors; and memory storing instructions that when executed by the processors cause the processors to effectuate operations of each the above-mentioned processes."). identifying, based on the given task, a relevant subset of the input measurement data that is of a higher relevancy with respect to the given task than the rest of the input measurement data; (Para. 26 see "Some embodiments augment computer-vision object detection by enriching a feature set by which objects are detected with a classification of a context in which the objects appear in an image. Examples include models that upweight kitchen utensils in response to classifying an image as depicting a scene in a kitchen as the image context, or upweight home improvement equipment in response to classifying an image as depicting a scene in a garage as the image context." Para. 33 see "Based on the identified vertical, different attributes (e.g., scores for dimensions) may be added to a feature vector (e.g., increasing its dimensionality) for an object recognition model, or different extant attributes of the feature vector may be weighted based on the vertical (e.g., by scaling the size of various scalars)." Para. 95 see "The identified context may function to (1) restrict a search to be narrowed to objects only related to the identified context, or (2) apply a weight to the search to weigh object related to the identified context greater than objects not related to the identified context."). determining, based on the input measurement data and the identified relevant subset, an enhanced input for the given classifier, such that, in the enhanced input, a portion of the input measurement data that corresponds to the identified relevant subset has a higher weight than other content of the input measurement data not corresponding to this identified subset; (Para. 26 see "Examples include models that upweight kitchen utensils in response to classifying an image as depicting a scene in a kitchen as the image context, or upweight home improvement equipment in response to classifying an image as depicting a scene in a garage as the image context." Para. 33 see "different extant attributes of the feature vector may be weighted based on the vertical (e.g., by scaling the size of various scalars)." Para. 95 see "objects related to a beach scene (e.g., beach balls, umbrellas, kites, sunscreen, coolers, towels, surfboards, etc.) may have their weights increased in the object recognition model, whereas objects unrelated to a beach scene (e.g., winter coats, snowboards, etc.) may not have their weights increased, or even may have their weights decreased."). providing the enhanced input to the given classifier, and obtaining an output from the classifier; (Para. 27 see "This scene classification vector may be input to the object recognition model as an enriched feature set along with the corresponding image itself for which objects are to be detected" Para. 58 see "This context classification vector, or a portion of that vector associated with the scene (e.g., a scene classification vector), may be input to an object recognition model as an enriched feature set along with the corresponding image itself for which objects are to be detected."). and determining a final classification result from the output. (Para. 6 see "determining, with the computer system, with a trained object recognition model, a first object identifier of the object based on the image and the context; and causing, with the computer system, the first object identifier of the object to be stored in memory." Para. 61 see "The output of the object recognition model may be one or more object identifiers indicating objects recognized as being present within a given image."). Regarding claim 2, Turkelson teaches The method of claim 1 . wherein the determining of the enhanced input includes cropping, from the input measurement data, a portion including the identified relevant subset. (Para. 63 see "For example, if an image is determined to include a first object at a first location within the image, the image may be cropped about a region of interest (ROI) centered about the first location, the region of interest may have its resolution, clarity, or prominence increased, or portions of the image not included within the region of interest may be compressed or otherwise have their resolution downscaled." Para. 107 see "In some embodiments, the native application may crop a portion of the image including first object 702 and input location 708 , and the cropped portion of the image may be input to the visual search system." Para. 115 see "the native application on mobile computing device 104 may crop the image to encompass only a portion of the image including the object of interest (e.g., first object 702 for distance D 1 being less than distance D 2 )."). Regarding claim 3, Turkelson teaches The method of claim 2 . wherein a size of the cropped portion is scaled over a size of the identified relevant subset by a predetermined scaling factor. (Para. 75 see "In some embodiments, a scaling factor may be applied to input location to obtain the coordinates" Para. 110 see "In some embodiments, the mapping from physical coordinates of driving lines 812 and sensing lines 810 may be 1:1 (e.g., each coordinate along each axes relates to a corresponding pixel along that axes), or a scaling factor may be applied."). Regarding claim 4, Turkelson teaches The method of claim 1 . wherein the input measurement data includes images, and the given task includes classifying types of objects shown in the images from a given set of types. (Para. 6 see "Some aspects include a process including: obtaining, with a computer system, an image depicting an object within a context, wherein: the image is captured by a mobile computing device, the object is a member of an ontology of objects including a plurality of objects, and the context is a member of an ontology of contexts including a plurality of contexts; determining, with the computer system, with a trained context classification model, the context depicted by the image; determining, with the computer system, with a trained object recognition model, a first object identifier of the object based on the image and the context; and causing, with the computer system, the first object identifier of the object to be stored in memory."). Regarding claim 5, Turkelson teaches The method of claim 4 . wherein the identifying of the relevant subset includes detecting, by a given object detector, bounding boxes surrounding instances of objects of types from the given set of types. (Para. 40 see "Some embodiments may output such scores for each of a plurality of objects in an object ontology (e.g., in an object detection vector) and, in some cases, bounding polygons (with vertices expressed in pixel coordinates) of each object. For example, a feature vector may be generated from an input image, where dimensions correspond to features (like edges, blobs, corners, colors, and the like) in the input image." Para. 96 see "first bounding box 404 may be placed around first object 404 and a second bounding box 408 may be placed around second object 410 . A confidence level may be computed that the object detected in each of bounding boxes 406 and 408 is a particular object from an object ontology, based at least in part on context 402 of image 400 ."). Regarding claim 6, Turkelson teaches The method of claim 1 . wherein: the given classifier is a classifier that has been trained on a generic set of classes; (Para. 28 see "In some embodiments, the object recognition model is trained to recognize (e.g., classify and locate) objects in an ontology of objects, only a small (e.g., less than 0.1%) subset of which may appear in any given image in some cases." Para. 68 see "In some embodiments, model subsystem 116 may train an object recognition model based on a training set including a plurality of images depicting different objects, where each image is labeled with an object identifier of the object from an object ontology depicted by the image."). and the identifying of the relevant subset includes extracting, from the input measurement data, information that is relevant with respect to a given subset of the generic set of classes. (Para. 28 see "In some embodiments, the object recognition model is trained to recognize (e.g., classify and locate) objects in an ontology of objects, only a small (e.g., less than 0.1%) subset of which may appear in any given image in some cases." Para. 33 see "Based on the identified vertical, different attributes (e.g., scores for dimensions) may be added to a feature vector (e.g., increasing its dimensionality) for an object recognition model, or different extant attributes of the feature vector may be weighted based on the vertical (e.g., by scaling the size of various scalars)."). Regarding claim 9, Turkelson teaches The method of claim 1 . wherein: the input measurement data is obtained using at least one sensor carried on board a vehicle, (Para. 91 see "For example, a mobile robot, autonomous vehicle, drone, mobile manipulator, assistive robots, and the like, may ingest video or images in real-time, determine a context of the image (e.g., a scene), determine objects within the image based on the determined context and the image, and then return the determined object and initially determined context to update, if necessary, the context." Para. 127 see "I/ O device interface 1030 may provide an interface for connection of one or more I/ O devices 1060 to computer system 1000 . ... I/ O devices 1060 may include, for example, ... pointing devices (e.g., a computer mouse or trackball), keyboards, keypads, touchpads, scanning devices, voice recognition devices, gesture recognition devices, printers, audio speakers, microphones, cameras, or the like. I/ O devices 1060 may be connected to computer system 1000 through a wired or wireless connection. I/ O devices 1060 may be connected to computer system 1000 from a remote location. I/ O devices 1060 located on remote computer system, for example, may be connected to computer system 1000 via a network and network interface 1040 ."). and the relevant subset and/or the enhanced input, but not the input measurement data, is transmitted for further processing over a vehicle bus network of the vehicle, and/or over a public land mobile network. (Para. 45 see "tap point information (or coordinates of other forms of user input) may be used to enhance or selectively process an image prior to being provided to a server. ... the enhancement or other form of processing may be performed additionally or alternatively by server-side operations of a search system. This may balance the tradeoff between reducing the processing time associated with server side image processing and latency issues associated with transmitting high-quality images to the server." Para. 107 see "the cropped portion of the image may be input to the visual search system. In some embodiments, the native application may apply a bounding box to first object 702 and may enhance a portion of the image within the bounding box, where the enhanced portion may be input to the visual search system alone, with the rest of the image, or with the rest of the image and a weight applied to the portion to indicate prominence of the portion. In some embodiments, the remaining portions of the image not including first object 702 may be down-scaled in resolution or otherwise compressed to reduce a file size of the image for the visual search. For example, if the visual search functionality resides, at least in part, on a remote server system, the reduced file size image may be transmitted faster to the remote server system and may also facilitate a faster search." Para. 128 see "Network interface 1040 may include a network adapter that provides for connection of computer system 1000 to a network. Network interface may 1040 may facilitate data exchange between computer system 1000 and other devices connected to the network. Network interface 1040 may support wired or wireless communication. The network may include an electronic communication network, such as the Internet, a local area network (LAN), a wide area network (WAN), a cellular communications network, or the like."). Regarding claim 10, Turkelson teaches The method of claim 1 . wherein: multiple enhanced inputs are supplied to the given classifier; and the determining of the final classification result includes aggregating outputs obtained from the classifier for the multiple enhanced inputs. (Para. 49 see "In some embodiments, one or more additional images may be obtained in the background, either spatially or temporally, and these images may subsequently be provided to the server as part of the same image processing job as that of the initially provided (compressed) image. By doing so, different objects, backgrounds, contexts, and visualization aspects (e.g., lighting, angle, etc.) may be analyzed in parallel processing with the initially sent image." Para. 91 see "In some embodiments, the object recognition model and context classification model may form a loop for dynamically analyzing captured video or images in real-time, and making adjustments based on the continuously evolving analysis."). Regarding claim 11, Turkelson teaches The method of claim 10 . wherein, in the aggregating, the outputs are weighted according to how much of the enhanced input belongs to the relevant subset of the input measurement data. (Para. 95 see "The identified context may function to (1) restrict a search to be narrowed to objects only related to the identified context, or (2) apply a weight to the search to weigh object related to the identified context greater than objects not related to the identified context."). Claim 13 is rejected under the same analysis as claim 1 above. Claim 14 is rejected under the same analysis as claim 1 above . Claim Rejections - 35 USC § 103 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-21-aia AIA Claim s 7-8 are rejected under 35 U.S.C. 103 as being unpatentable over Turkelson et al. (US 20200193206 A1), hereinafter Turkelson, in view of Radford et al.: "Learning Transferable Visual Models From Natural Language Supervision", arxiv.org cornell university, published 26 Feb 2021, [retrieved on 5-28-2026]. Retrieved from the internet <https://arxiv.org/abs/2103.00020>, hereinafter Radford . Regarding claim 7, Turkelson teaches The method of claim 1 . Turkelson does not teach wherein the given classifier is a classifier that is configured to: accept measurement data and a text prompt as inputs; and determine a classification score with respect to a class corresponding to the text prompt by rating a similarity between the measurement data and the text prompt. However, Radford teaches wherein the given classifier is a classifier that is configured to: accept measurement data and a text prompt as inputs; (Pg. 2, Fig. 1 desc. See "CLIP jointly trains an image encoder and at extencoder to predict the correct pairings of a batch of (image, text) training examples. At test time the learned text encoder synthesizes a zero-shot linear classifier by embedding the names or descriptions of the target dataset’s classes." Pg. 6, Sect. 3.1.2. see "To perform zero-shot classification, we reuse this capability. For each dataset, we use the names of all the classes in the dataset as the set of potential text pairings and predict the most probable (image, text) pair according to CLIP."). and determine a classification score with respect to a class corresponding to the text prompt by rating a similarity between the measurement data and the text prompt. (Pg. 2, Col. 1, Para. 1 see "by scoring target classes based on their dictionary of learned visual n-grams and predicting the one with the highest score." Pg. 6, Sect. 3.1.2. see "CLIP is pre-trained to predict if an image and a text snippet are paired together in its dataset. To perform zero-shot classification, we reuse this capability. For each dataset, we use the names of all the classes in the dataset as the set of potential text pairings and predict the most probable (image, text) pair according to CLIP. In a bit more detail, we first compute the feature embedding of the image and the feature embedding of the set of possible texts by their respective encoders. The cosine similarity of these embeddings is then calculated, scaled by a temperature parameter τ, and normalized into a probability distribution via a softmax."). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Turkelson to incorporate the teachings of Radford to accept measurement data and a text prompt as inputs and determine a classification score based on similarity. Doing so would predictably enable more flexible classification by allowing the system to check for different types of objects using a text prompt. Regarding claim 8, Turkelson in view of Radford teaches The method of claim 7 . Turkelson does not teach wherein the classifier is configured to: map the measurement data to a data representation in a latent space Z using a data encoder; map the text prompt to a text representation in the same latent space Z using a text encoder; and rate the similarity between the measurement data and the text prompt according to a distance between the data representation and the text representation in the latent space Z. However, Radford teaches wherein the classifier is configured to: map the measurement data to a data representation in a latent space Z using a data encoder; (Pg. 4, Col. 2, Para. 1 see "CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder" Pg. 6, Sect. 3.1.2. see "the image encoder is the computer vision backbone which computes a feature representation for the image"). map the text prompt to a text representation in the same latent space Z using a text encoder; (Pg. 2, Fig. 1 desc. See "CLIP jointly trains an image encoder and at extencoder to predict the correct pairings of a batch of (image, text) training examples. At test time the learned text encoder synthesizes a zero-shot linear classifier by embedding the names or descriptions of the target dataset’s classes." Pg. 4, Col. 2, Para. 1 see "jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings."). and rate the similarity between the measurement data and the text prompt according to a distance between the data representation and the text representation in the latent space Z. (Pg. 4, Col. 2, Para. 1 see "CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder" Pg. 6, Sect. 3.1.2. see "CLIP is pre-trained to predict if an image and a text snippet are paired together in its dataset. To perform zero-shot classification, we reuse this capability. For each dataset, we use the names of all the classes in the dataset as the set of potential text pairings and predict the most probable (image, text) pair according to CLIP. In a bit more detail, we first compute the feature embedding of the image and the feature embedding of the set of possible texts by their respective encoders. The cosine similarity of these embeddings is then calculated, scaled by a temperature parameter τ, and normalized into a probability distribution via a softmax." Fig. 3 see "logits = np.dot(I_e, T_e.T) * np.exp(t)" Examiner note: The dot product is used to determine the distance in latent space.). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Turkelson and Radford to incorporate the teachings of Radford to configure the classifier to map measurement data and text to the same latent space. Doing so would predictably enable more accurate and efficient matching by comparing the image and text description in the same language that the model understands . 07-21-aia AIA Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Turkelson et al. (US 20200193206 A1), hereinafter Turkelson, in view of Yamamoto et al. (US 20230215151 A1), hereinafter Yamamoto . Regarding claim 12, Turkelson teaches The method of claim 1 . wherein: the input measurement data is obtained from at least one sensor; (Para. 91 see "For example, a mobile robot, autonomous vehicle, drone, mobile manipulator, assistive robots, and the like, may ingest video or images in real-time, determine a context of the image (e.g., a scene), determine objects within the image based on the determined context and the image, and then return the determined object and initially determined context to update, if necessary, the context." Para. 127 see "I/ O device interface 1030 may provide an interface for connection of one or more I/ O devices 1060 to computer system 1000 . ... I/ O devices 1060 may include, for example, ... pointing devices (e.g., a computer mouse or trackball), keyboards, keypads, touchpads, scanning devices, voice recognition devices, gesture recognition devices, printers, audio speakers, microphones, cameras, or the like. I/ O devices 1060 may be connected to computer system 1000 through a wired or wireless connection. I/ O devices 1060 may be connected to computer system 1000 from a remote location. I/ O devices 1060 located on remote computer system, for example, may be connected to computer system 1000 via a network and network interface 1040 ."). Turkelson does not teach from the final classification result, an actuation signal is obtained; and a vehicle and/or a driving assistance system and/or a robot and/or a quality inspection system and/or a surveillance system and/or a medical imaging system, is actuated with the actuation signal. However, Yamamoto teaches from the final classification result, an actuation signal is obtained; and a vehicle and/or a driving assistance system and/or a robot and/or a quality inspection system and/or a surveillance system and/or a medical imaging system, is actuated with the actuation signal. (Para. 101 see "The vehicle control unit 32 controls each unit of the vehicle 1 . The vehicle control unit 32 includes the steering control unit 81 , the brake control unit 82 , the drive control unit 83 , a body system control unit 84 , a light control unit 85 , and a horn control unit 86 ." Para. 126 see "The object recognition filter 112 performs the object recognition processing such as semantic segmentation on the preprocessed image PC preprocessed by the preprocessing filter 111 , recognizes an object in units of pixels, and outputs an image including the recognition result as an object recognition result image PL." Para. 128 see "The operation control unit 63 recognizes an object of a subject in the image on the basis of the object recognition result in units of pixels in the object recognition result image PL, and controls the operation of the vehicle 1 on the basis of the recognition result."). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Turkelson to incorporate the teachings of Yamamoto to obtain an actuation signal from the final classification result and actuate a vehicle. Doing so would predictably enable safe automated operation by using classification outputs to determine where a vehicle may safely move and where it would be unsafe to move . Conclusion 07-96 AIA The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Dimond et al. (US 12046031 B2) discloses image target classification systems and methods to receive data associated with a scene and optimize a probability of detecting objects in an image . Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEXANDER J VAUGHN whose telephone number is (571) 272-5253. The examiner can normally be reached M-F 8:30-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ANDREW MOYER can be reached on (571) 272-9523. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ALEXANDER JOSEPH VAUGHN/Examiner, Art Unit 2675 /EDWARD PARK/Primary Examiner, Art Unit 2675 Application/Control Number: 18/794,164 Page 2 Art Unit: 2675 Application/Control Number: 18/794,164 Page 3 Art Unit: 2675 Application/Control Number: 18/794,164 Page 4 Art Unit: 2675 Application/Control Number: 18/794,164 Page 5 Art Unit: 2675 Application/Control Number: 18/794,164 Page 6 Art Unit: 2675 Application/Control Number: 18/794,164 Page 7 Art Unit: 2675 Application/Control Number: 18/794,164 Page 8 Art Unit: 2675 Application/Control Number: 18/794,164 Page 9 Art Unit: 2675 Application/Control Number: 18/794,164 Page 10 Art Unit: 2675 Application/Control Number: 18/794,164 Page 11 Art Unit: 2675 Application/Control Number: 18/794,164 Page 12 Art Unit: 2675 Application/Control Number: 18/794,164 Page 13 Art Unit: 2675 Application/Control Number: 18/794,164 Page 14 Art Unit: 2675 Application/Control Number: 18/794,164 Page 15 Art Unit: 2675 Application/Control Number: 18/794,164 Page 16 Art Unit: 2675 Application/Control Number: 18/794,164 Page 17 Art Unit: 2675 Application/Control Number: 18/794,164 Page 18 Art Unit: 2675