Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Applicant’s Application filed on 7/25/2025 has been reviewed.
Claims 1-20 have been examined.
Notice of Pre-AIA or AIA Status
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Specification
Applicant is reminded of the proper language and format for an abstract of the disclosure.
The abstract should be in narrative form and generally limited to a single paragraph on a separate sheet within the range of 50 to 150 words in length. The abstract should describe the disclosure sufficiently to assist readers in deciding whether there is a need for consulting the full patent text for details.
The language should be clear and concise and should not repeat information given in the title. It should avoid using phrases which can be implied, such as, “The disclosure concerns,” “The disclosure defined by this invention,” “The disclosure describes,” etc. In addition, the form and legal phraseology often used in patent claims, such as “means” and “said,” should be avoided.
The abstract of the disclosure is objected to because of legal phraseology “embodiments”, and implied phrase “present disclosure relate to”. A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b).
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
(Step 1) The claim(s) 1-20 recite(s) recite(s) a processor, system and method, and are directed toward statutory subject matter.
(Step 2A1-does the claim recite an abstract idea, law of nature, or natural phenomenon?)
The enumerated groupings of abstract ideas are defined as:
1) Mathematical concepts – mathematical relationships, mathematical formulas or equations, mathematical calculations (see MPEP § 2106.04(a)(2), subsection I);
2) Certain methods of organizing human activity – fundamental economic principles or practices (including hedging, insurance, mitigating risk); commercial or legal interactions (including agreements in the form of contracts; legal obligations; advertising, marketing or sales activities or behaviors; business relations); managing personal behavior or relationships or interactions between people (including social activities, teaching, and following rules or instructions) (see MPEP § 2106.04(a)(2), subsection II); and
3) Mental processes – concepts performed in the human mind (including an observation, evaluation, judgment, opinion) (see MPEP § 2106.04(a)(2), subsection III).
The limitation of claim 1 (similarly in claims 10 and 19) “obtain a digital asset; obtain one or more prompts requesting one or more metadata attribute values to be associated with the digital asset; based at least on a model processing a representation of the one or more prompts and a representation of the digital asset, generate a response to the one or more prompts, the response including the one or more metadata attribute values of the digital asset; and store the response using an index configured to facilitate retrieval of the digital asset as a search result candidate”, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind/manual process but for the recitation of generic computer components. That is, other than reciting “by a processor,” nothing in the claim element precludes the step from practically being performed in the mind/manually performance. For example, but for the “by a processor” language, “determining” in the context of this claim encompasses the user mental/manually perform the process.
The claims do recite a mental process when they contain limitations that can practically be performed in the human mind, including for example, observations, evaluations, judgments, and opinions. Examples of claims that recite mental processes include:
a claim to “collecting information, analyzing it, and displaying certain results of the collection and analysis,” where the data analysis steps are recited at a high level of generality such that they could practically be performed in the human mind, Electric Power Group v. Alstom, S.A., 830 F.3d 1350, 1353-54, 119 USPQ2d 1739, 1741-42 (Fed. Cir. 2016);
claims to “comparing BRCA sequences and determining the existence of alterations,” where the claims cover any way of comparing BRCA sequences such that the comparison steps can practically be performed in the human mind, University of Utah Research Foundation v. Ambry Genetics, 774 F.3d 755, 763, 113 USPQ2d 1241, 1246 (Fed. Cir. 2014);
a claim to collecting and comparing known information (claim 1), which are steps that can be practically performed in the human mind, Classen Immunotherapies, Inc. v. Biogen IDEC, 659 F.3d 1057, 1067, 100 USPQ2d 1492, 1500 (Fed. Cir. 2011); and
Further, if a claim recites a limitation that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper, the limitation falls within the mental processes grouping, and the claim recites an abstract idea. In this case, for claim 1, except for using generic elements such as processor, all other element can be performed by human mind as a mental process and/or performed manually using pencil and paper (The use of a physical aid (e.g., pencil and paper or a slide rule) to help perform a mental step (e.g., a mathematical calculation) does not negate the mental nature of the limitation, but simply accounts for variations in memory capacity from one person to another. For instance, in CyberSource, the court determined that the step of "constructing a map of credit card numbers" was a limitation that was able to be performed "by writing down a list of credit card transactions made from a particular IP address." In making this determination, the court looked to the specification, which explained that the claimed map was nothing more than a listing of several (e.g., four) credit card transactions. The court concluded that this step was able to be performed mentally with a pen and paper, and therefore, it qualified as a mental process. 654 F.3d at 1372-73, 99 USPQ2d at 1695. See also Flook, 437 U.S. at 586, 198 USPQ at 196 (claimed "computations can be made by pencil and paper calculations"); University of Florida Research Foundation, Inc. v. General Electric Co., 916 F.3d 1363, 1367, 129 USPQ2d 1409, 1411-12 (Fed. Cir. 2019) (relying on specification’s description of the claimed analysis and manipulation of data as being performed mentally "‘using pen and paper methodologies, such as flowsheets and patient charts’"); Symantec, 838 F.3d at 1318, 120 USPQ2d at 1360 (although claimed as computer-implemented, steps of screening messages can be "performed by a human, mentally or with pen and paper").) (MPEP 2106.04(a)(2).)
Thus, if a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind and/or manually performed, but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
(Step 2A2-Practical Application?)This judicial exception is not integrated into a practical application.
The courts have also identified limitations that did not integrate a judicial exception into a practical application:
• Merely reciting the words “apply it” (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f);
• Adding insignificant extra-solution activity to the judicial exception, as discussed in MPEP § 2106.05(g); and
• Generally linking the use of a judicial exception to a particular technological environment or field of use, as discussed in MPEP § 2106.05(h).
In particular, the claim only recites one additional element – using a processor to perform “…obtain…generate…store…”. The processor in performing the steps is recited at a high-level of generality (i.e., as a generic processor performing a generic computer function of the steps) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
(Step 2B- does the claim recite additional elements that amount to significantly more than the judicial exception?)
Limitations that the courts have found not to be enough to qualify as “significantly more” when recited in a claim with a judicial exception include:
i. Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, e.g., a limitation indicating that a particular function such as creating and maintaining electronic records is performed by a computer, as discussed in Alice Corp., 573 U.S. at 225-26, 110 USPQ2d at 1984 (see MPEP § 2106.05(f));
ii. Simply appending well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception, e.g., a claim to an abstract idea requiring no more than a generic computer to perform generic computer functions that are well-understood, routine and conventional activities previously known to the industry, as discussed in Alice Corp., 573 U.S. at 225, 110 USPQ2d at 1984 (see MPEP § 2106.05(d));
iii. Adding insignificant extra-solution activity to the judicial exception, e.g., mere data gathering in conjunction with a law of nature or abstract idea such as a step of obtaining information about credit card transactions so that the information can be analyzed by an abstract mental process, as discussed in CyberSource v. Retail Decisions, Inc., 654 F.3d 1366, 1375, 99 USPQ2d 1690, 1694 (Fed. Cir. 2011) (see MPEP § 2106.05(g)); or
iv. Generally linking the use of the judicial exception to a particular technological environment or field of use, e.g., a claim describing how the abstract idea of hedging could be used in the commodities and energy markets, as discussed in Bilski v. Kappos, 561 U.S. 593, 595, 95 USPQ2d 1001, 1010 (2010) or a claim limiting the use of a mathematical formula to the petrochemical and oil-refining fields, as discussed in Parker v. Flook, 437 U.S. 584, 588-90, 198 USPQ 193, 197-98 (1978) (MPEP § 2106.05(h)).
The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using a processor to perform the steps amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claims are not patent eligible.
“As explained by the Supreme Court, the addition of insignificant extra-solution activity does not amount to an inventive concept, particularly when the activity is well-understood or conventional. Parker v. Flook, 437 U.S. 584, 588-89, 198 USPQ 193, 196 (1978). In Flook, the Court reasoned that “[t]he notion that post-solution activity, no matter how conventional or obvious in itself, can transform an unpatentable principle into a patentable process exalts form over substance. A competent draftsman could attach some form of post-solution activity to almost any mathematical formula”. 437 U.S. at 590; 198 USPQ at 197; Id. (holding that step of adjusting an alarm limit variable to a figure computed according to a mathematical formula was “post-solution activity”). “
As to claims 2-9, 11-18 and 20, the claim further recites additional steps related to obtaining data, and processing data. The additional limitation further detailing with data processing and add insignificant extra-solution activity. Refining the abstract idea and/or add insignificant extra-solution activity does not make an abstract idea beyond the abstract idea itself. The claim recites mental process and/or manual process including steps of "obtain…generate…", “…render…”, “…convert…”, “…store…”, “…update…”, wherein the steps are recited at a high level of generality such that they could practically be performed in the human mind, or merely insignificant post-solution activity that outputs the result of performing the abstract idea, which is insufficient to integrate the abstract idea into a practical application (MPEP 2106.05(g).) (Electric Power Group v. Alstom, S.A., 830 F.3d 1350, 1353-54, 119 USPQ2d 1739, 1741-42 (Fed. Cir. 2016)). Thus, the claim does not mount to significantly more than the abstract idea.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1, 3-10 and 12-20 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by U.S. Patent Application Publication No. 20250217407 to Xin-Jing Wang (hereinafter “Wang”).
As to claim 1, Wang teaches one or more processors comprising one or more processing units to (computer implemented method in a system comprising processor and non-transitory computer readable storage medium, par. 0004-0026, 0078-0105, Fig. 5):
obtain a digital asset (Fig. 1, par. 0035, 0073, obtain images, i.e. “[0035] FIG. 1 is a diagram of an example multi-modal search-based object detection system 100 for identifying and providing images of utility assets that match an input query. The multi-modal search-based object detection system 100 includes the multi-modal search system 106, which is configured to obtain images 103 from image databases 101-1-101-N (collectively referred to as “image databases 101”)…. The multi-modal search system 106 can utilize scripting tools (e.g., by a datastore extractor 144) to extract the data from the datastores and update search tokens in the search index”);
obtain one or more prompts requesting one or more metadata attribute values to be associated with the digital asset (par. 0036-0038, obtain and map image feature (metadata) to images and image token, i.e. “[0036] The multi-modal search system 106 includes an offline substage 108 to perform offline processing of the images 103 and generate a token-based search index 118. The search index 118 includes search tokens 120-1-120-N (collectively referred to as “search tokens 120”), each search token 120 including an image identifier that corresponds to a particular image from images 103. A search token in the search tokens 120 is an embedded representation of features from the images 103. For example, a search token can include an identifier for an image, labels of objects represented in the image, and pixel positions in the image for a bounding box that corresponds to an object of the objects represented in the image. The search token can include any number of objects and bounding boxes, and provides a mapping of features (e.g., text features, image features) to a unique image. In some implementations, the search token is a lower dimensional representation (e.g., an embedding and/or encoding) of the features compared to the features of the original input image…. [0038] The vision transformer encoder 114 can perform a number of embedding and/or encoding processes to map image feature data into a search image token, which can be stored as part of the search image tokens 124-1-124-N (collectively “search image tokens 124”). The vision transformer encoder 114 can be configured to determine a subset of images from the images 103 that share similar visual features (e.g., object class, defect type, or both object class and defect type) based on the bounding boxes of the subset of images.”);
based at least on a model processing a representation of the one or more prompts and a representation of the digital asset, generate a response to the one or more prompts, the response including the one or more metadata attribute values of the digital asset (par. 0036-0038, generate a response by mapping image feature (metadata) to images and image token, i.e. “[0036] The multi-modal search system 106 includes an offline substage 108 to perform offline processing of the images 103 and generate a token-based search index 118. The search index 118 includes search tokens 120-1-120-N (collectively referred to as “search tokens 120”), each search token 120 including an image identifier that corresponds to a particular image from images 103. A search token in the search tokens 120 is an embedded representation of features from the images 103. For example, a search token can include an identifier for an image, labels of objects represented in the image, and pixel positions in the image for a bounding box that corresponds to an object of the objects represented in the image. The search token can include any number of objects and bounding boxes, and provides a mapping of features (e.g., text features, image features) to a unique image. In some implementations, the search token is a lower dimensional representation (e.g., an embedding and/or encoding) of the features compared to the features of the original input image…. [0038] The vision transformer encoder 114 can perform a number of embedding and/or encoding processes to map image feature data into a search image token, which can be stored as part of the search image tokens 124-1-124-N (collectively “search image tokens 124”). The vision transformer encoder 114 can be configured to determine a subset of images from the images 103 that share similar visual features (e.g., object class, defect type, or both object class and defect type) based on the bounding boxes of the subset of images.”)); and
store the response using an index configured to facilitate retrieval of the digital asset as a search result candidate (Fig. 3, par. 0057-0064, storing image token in search index, i.e. “[0064] Similar to the text token, the image token can be stored in a search index and includes a corresponding identifier for each image in the second subset of images. In some implementations, the multi-modal search system 106 includes a machine learning network (e.g., machine learning network 112) configured to generate image tokens (e.g., search image tokens 124). For example, the machine learning network can include a vision transformer encoder configured to generate image embeddings from a clustering of the shared visual features. The machine learning network can include a visual transformer encoder (e.g., visual transformer encoder 114) trained to generate embeddings of feature data (e.g., representing features of utility asset) based on the shared visual features. Embeddings of features represented by the shared visual features can be projected in a lower dimensional representation, e.g., relative to the feature data from the images in the second subset of images.”).
As to claim 3, Wang teaches the one or more processors of claim 1, wherein the one or more processing units are further to: render a plurality of images of the digital asset, each image of the plurality of images representing a unique viewing angle of the digital asset, and wherein a plurality of representations of the plurality of images are used by the model as input to generate the response (par. 0053-0054, plurality of representations of the digital asset, i.e. “Although each candidate image is depicted in FIG. 2 as having a corresponding bounding box (e.g., bounding box 214-1 for candidate image 212-1, bounding box 214-2 for candidate image 212-2, and bounding box 214-3 for candidate image 212-3), any number of bounding boxes can be rendered for a candidate image. The user interface 204 can be updated by providing the query interface results 148 from the multi-modal search system 106, as described in reference to FIG. 1 above.”).
As to claim 4, Wang teaches the one or more processors of claim 1, wherein the one or more processing units are further to: receive a user query that references the one or more metadata attribute values of the digital asset; based at least on the user query, obtain, using the index, the response; and based at least on the generating of the response, executing the user query by retrieving the response and the digital asset and causing presentation of the digital asset as a search result for the user query (par. 0053-0056, receiving user query reference metadata attribute “<rusty transformer>”and obtaining the result for presentation using index, i.e. “[0054] The user interface 204 shows resulting candidate images that match for one or both of the text input 104-1 and image input 104-2. Each of the candidate images from candidate images 212 can be rendered at a position of the user interface 204 that corresponds to the geographical position of the utility asset captured in the candidate image, relative to the geographical map 210 displayed by the user interface 204. The user interface 204 also includes a window 216 for additional information for the candidate image 212-1, displaying additional data from datastores such as information related to the asset. The information related to the asset can include any connected feeders (e.g., from a feeder map or electric grid utility data), an asset identifier utilized by a utility company, electrical power characteristics such as voltage, current, and any applicable electrical ratings. In some implementations, the information related to the asset depicted through window 216 can include inspection data (e.g., inspection dates and results), any known defects from the electrical utility database or inspection record, a priority for the utility asset, as well as an indicator for the associated utility company for the asset.”).
As to claim 5, Wang teaches the one or more processors of claim 4, wherein the one or more processing units are further to: convert at least one of the digital asset to a first embedding, the first embedding being a vector representation of a word or phrase that semantically represents the digital asset (par. 0042, generate embeddings of feature vectors, such as tokens, i.e. “0041] In some implementations, the machine learning network 112 can perform a variety of training techniques to improve embedding and encoding of utility asset features from images into tokens. …In some implementations, machine learning network 112 includes one or more fully or partially connected layers. Each of the layers can include one or more parameter values indicating an output of the layers. The layers of the model can generate embeddings of feature vectors from input images, including text annotations and bounding boxes.”); store the first embedding using the index (par. 0004-0006, 0008, store tokens in index); and based at least on accessing the first embedding using the index and determining a distance between the first embedding and a second embedding representing a user query, execute the user query by retrieving the digital asset and causing presentation of the digital asset as a search result for the user query (par. 0018, 0074-0075, the similarity score indicates a likelihood of a respective image matching the textual token and the image token, that is interpreted as distance between a first embedding and a second embedding, i.e. “[0018] In some implementations, the method further includes filtering, based on the textual token and the image token for the input, the images to obtain a filtered subset of images. The filtered subset of images can exclude images that do not include at least one token from the textual token and the image token corresponding to the input that match, from the search index, the textual token and the image token. The method further includes determining, based on the textual token and the image token, a similarity score for each image in the filtered subset of images, the similarity score indicates a likelihood of a respective image matching the textual token and the image token. The method further includes identifying the candidate images from the filtered subset of images, the candidate images each having a respective similarity score that exceeds a threshold value. In some implementations, the method further includes ranking the candidate images based on the similarity score of the respective candidate image.”).
As to claim 6, Wang teaches the one or more processors of claim 1, wherein the one or more processing units are further to: subsequent to executing a user query, receive a second prompt requesting a second metadata attribute value to be associated with the digital asset (Fig. 4, par. 0067-0077, input plurality of attributes for query and requesting for digital asset, i.e. “[0068] The multi-modal search system 106 provides, for display on a user device (e.g., client device 102, described in reference to FIG. 1 above), a user interface (e.g., user interface 204, described in reference to FIG. 2 above) configured to receive input representing a search query for one or more images depicting an electric grid asset (410). The user interface permits the input to include at least one of textual data or image data requesting the search query. The user interface can include a search bar (e.g., search bar 206) for a user of a client and/or user device to enter a text input. Examples of textual data (e.g., a text input) can include a keyword, a number of keywords, and/or a semantic phrase of words. The user interface can include an input mechanism (e.g., a button to enter a directory) to upload and submit an image query. In some implementations, the user interface includes a mechanism for entering a query bounding box (e.g., entering coordinates for the boundary box, drawing a bounding box over a portion of the input image) that encloses pixels of the input image to submit as an image query.”); generate, based at least on the model processing a second representation of the second prompt and the representation of the digital asset, a second response to the second prompt; and update the index by storing the second response using the index (Fig. 3, par. 0057-0066, process of update index by storing image token in the search index, i.e. “[0065] In some implementations, the visual transformer encoder is configured to generate the embedding based on a clustering of the shared visual features of utility assets in the second subset of images. ... In some implementations, the text transformer encoder can be a natural language processing model trained to generate textual tokens based on natural language inputs.”).
As to claim 7, Wang teaches the one or more processors of claim 1, wherein the response comprises a structured data format with a plurality of metadata attributes of the digital asset that are mapped to a corresponding metadata attribute value (par. 0006, 0053-0055, map results, i.e. “[0006] The multi-modal search system permits for efficient and dynamic display of query results on a client device. The multi-modal search system receives input in the form of a search query, which may be a form of text and/or image data. Examples of input can include keywords, semantic search, images (with or without annotations), and bounding boxes for the images. The multi-modal search system utilizes an encoding of utility asset features (textual features, image features) from the search input to generate the search tokens. From the search index of images from an image database, the multi-modal search system identifies candidate images by comparing image tokens and textual tokens in the search index to the search tokens. Thus, the multi-modal search system allows for an identification of candidate images that share textual and/or visual features to the search input and provides the candidate images on the user interface of the client device. The user interface can display a geographical map of an electric grid and overlay a candidate image from the query results onto a position of the geographical map.”).
As to claim 8, Wang teaches the one or more processors of claim 1, wherein the one or more prompts include at least one of a first natural language command or question issued by a user or a second natural language command or question issued by a language model agent (Fig. 1, 2, par. 0040-0043, 0061-0065, natural language inputs such as user input features of images “<rusty transformer>”).
As to claim 9, Wang teaches the one or more processors of claim 1, wherein the one or more processors is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system for generating synthetic data; a system for generating synthetic data using one or more large language models (LLMs); a system for generating synthetic data using one or more vision language models (VLMs); a system for generating synthetic data using one or more multi-modal language models; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (par. 0079-0080, a system implemented at least partially in a data center, i.e. “In addition, multiple computing devices may be connected, with each device providing portions of the operations (e.g., as a server bank, a group of blade servers, or a multi-processor system). In some implementations, the processor 502 is a single threaded processor. In some implementations, the processor 502 is a multi-threaded processor. In some implementations, the processor 502 is a quantum computer.” ).
As to claim 10, Wang teaches a data center system comprising a plurality of computing nodes, wherein two or more computing nodes of the plurality of computing nodes comprise one or more graphics processing units (GPUs) to (par. 0078-0083, a standard server 520, or multiple times in a group of such servers):
obtain one or more user-defined fields representing one or more metadata attributes associated with a digital asset (par. 0036-0038, generate a response by mapping image feature (metadata) to images and image token, i.e. “[0036] The multi-modal search system 106 includes an offline substage 108 to perform offline processing of the images 103 and generate a token-based search index 118. The search index 118 includes search tokens 120-1-120-N (collectively referred to as “search tokens 120”), each search token 120 including an image identifier that corresponds to a particular image from images 103. A search token in the search tokens 120 is an embedded representation of features from the images 103. For example, a search token can include an identifier for an image, labels of objects represented in the image, and pixel positions in the image for a bounding box that corresponds to an object of the objects represented in the image. The search token can include any number of objects and bounding boxes, and provides a mapping of features (e.g., text features, image features) to a unique image. In some implementations, the search token is a lower dimensional representation (e.g., an embedding and/or encoding) of the features compared to the features of the original input image…. [0038] The vision transformer encoder 114 can perform a number of embedding and/or encoding processes to map image feature data into a search image token, which can be stored as part of the search image tokens 124-1-124-N (collectively “search image tokens 124”). The vision transformer encoder 114 can be configured to determine a subset of images from the images 103 that share similar visual features (e.g., object class, defect type, or both object class and defect type) based on the bounding boxes of the subset of images.”));
receive one or more prompts requesting, in natural language, one or more metadata attribute values of the one or more metadata attributes to be associated with the digital asset (Fig. 1, 2, par. 0040-0043, 0061-0065, natural language inputs such as user input features of images “<rusty transformer>”);
provide at least one of: a representation of the one or more user-defined fields, a representation of the one or more prompts, or a representation of the digital asset as input into a model, wherein the model generates a response to the one or more prompts, the response including the one or more metadata attribute values of the digital asset (par. 0042, generate embeddings of feature vectors, such as tokens, i.e. “0041] In some implementations, the machine learning network 112 can perform a variety of training techniques to improve embedding and encoding of utility asset features from images into tokens. …In some implementations, machine learning network 112 includes one or more fully or partially connected layers. Each of the layers can include one or more parameter values indicating an output of the layers. The layers of the model can generate embeddings of feature vectors from input images, including text annotations and bounding boxes.”); and
store the response using an index, the index being configured to facilitate retrieval of data for a query associated with the digital asset (Fig. 3, par. 0057-0064, storing image token in search index, i.e. “[0064] Similar to the text token, the image token can be stored in a search index and includes a corresponding identifier for each image in the second subset of images. In some implementations, the multi-modal search system 106 includes a machine learning network (e.g., machine learning network 112) configured to generate image tokens (e.g., search image tokens 124). For example, the machine learning network can include a vision transformer encoder configured to generate image embeddings from a clustering of the shared visual features. The machine learning network can include a visual transformer encoder (e.g., visual transformer encoder 114) trained to generate embeddings of feature data (e.g., representing features of utility asset) based on the shared visual features. Embeddings of features represented by the shared visual features can be projected in a lower dimensional representation, e.g., relative to the feature data from the images in the second subset of images.”).
As to claim 12, Wang teaches the data center system of claim 10, wherein the one or more GPUs are further to: render a plurality of images of the digital asset, each image, of the plurality of images, representing a unique viewing angle of the digital asset, and wherein a plurality of representations of the plurality of images are used by the model as input to generate the response (par. 0053-0054, plurality of representations of the digital asset, i.e. “Although each candidate image is depicted in FIG. 2 as having a corresponding bounding box (e.g., bounding box 214-1 for candidate image 212-1, bounding box 214-2 for candidate image 212-2, and bounding box 214-3 for candidate image 212-3), any number of bounding boxes can be rendered for a candidate image. The user interface 204 can be updated by providing the query interface results 148 from the multi-modal search system 106, as described in reference to FIG. 1 above.”).
As to claim 13, Wang teaches the data center system of claim 10, wherein the one or more GPUs are further to: receive a user query that references the one or more metadata attribute values of the digital asset; based at least on the user query, obtain, using the index, the response; and based at least on the model generating the response, execute the user query by retrieving the response and the digital asset and cause presentation of the digital asset as a search result for the user query (par. 0053-0056, receiving user query reference metadata attribute “<rusty transformer>”and obtaining the result for presentation using index, i.e. “[0054] The user interface 204 shows resulting candidate images that match for one or both of the text input 104-1 and image input 104-2. Each of the candidate images from candidate images 212 can be rendered at a position of the user interface 204 that corresponds to the geographical position of the utility asset captured in the candidate image, relative to the geographical map 210 displayed by the user interface 204. The user interface 204 also includes a window 216 for additional information for the candidate image 212-1, displaying additional data from datastores such as information related to the asset. The information related to the asset can include any connected feeders (e.g., from a feeder map or electric grid utility data), an asset identifier utilized by a utility company, electrical power characteristics such as voltage, current, and any applicable electrical ratings. In some implementations, the information related to the asset depicted through window 216 can include inspection data (e.g., inspection dates and results), any known defects from the electrical utility database or inspection record, a priority for the utility asset, as well as an indicator for the associated utility company for the asset.”).
As to claim 14, Wang teaches the data center system of claim 13, wherein the one or more GPUs are further to: convert at least one of the digital asset or the response to a first embedding, the first embedding being a vector representation of a word or phrase that captures meaning in relation to other words or phrases(par. 0042, generate embeddings of feature vectors, such as tokens, i.e. “0041] In some implementations, the machine learning network 112 can perform a variety of training techniques to improve embedding and encoding of utility asset features from images into tokens. …In some implementations, machine learning network 112 includes one or more fully or partially connected layers. Each of the layers can include one or more parameter values indicating an output of the layers. The layers of the model can generate embeddings of feature vectors from input images, including text annotations and bounding boxes.”); store the first embedding using the index (par. 0004-0006, 0008, store tokens in index); and based at least on accessing the first embedding using the index and determining a distance between the first embedding and a second embedding representing the user query, executing the user query by retrieving the digital asset and causing presentation of the digital asset as a search result for the user query (par. 0018, 0074-0075, the similarity score indicates a likelihood of a respective image matching the textual token and the image token, that is interpreted as distance between a first embedding and a second embedding, i.e. “[0018] In some implementations, the method further includes filtering, based on the textual token and the image token for the input, the images to obtain a filtered subset of images. The filtered subset of images can exclude images that do not include at least one token from the textual token and the image token corresponding to the input that match, from the search index, the textual token and the image token. The method further includes determining, based on the textual token and the image token, a similarity score for each image in the filtered subset of images, the similarity score indicates a likelihood of a respective image matching the textual token and the image token. The method further includes identifying the candidate images from the filtered subset of images, the candidate images each having a respective similarity score that exceeds a threshold value. In some implementations, the method further includes ranking the candidate images based on the similarity score of the respective candidate image.”).
As to claim 15, Wang teaches the data center system of claim 10, wherein the one or more GPUs are further to: subsequent to executing a user query, receive a second prompt requesting a second metadata attribute value to be associated with the digital asset(Fig. 4, par. 0067-0077, input plurality of attributes for query and requesting for digital asset, i.e. “[0068] The multi-modal search system 106 provides, for display on a user device (e.g., client device 102, described in reference to FIG. 1 above), a user interface (e.g., user interface 204, described in reference to FIG. 2 above) configured to receive input representing a search query for one or more images depicting an electric grid asset (410). The user interface permits the input to include at least one of textual data or image data requesting the search query. The user interface can include a search bar (e.g., search bar 206) for a user of a client and/or user device to enter a text input. Examples of textual data (e.g., a text input) can include a keyword, a number of keywords, and/or a semantic phrase of words. The user interface can include an input mechanism (e.g., a button to enter a directory) to upload and submit an image query. In some implementations, the user interface includes a mechanism for entering a query bounding box (e.g., entering coordinates for the boundary box, drawing a bounding box over a portion of the input image) that encloses pixels of the input image to submit as an image query.”); generate, based at least on the model processing a second representation of the second prompt and the representation of the digital asset, a second response to the second prompt; and update the index by storing the second response using the index (Fig. 3, par. 0057-0066, process of update index by storing image token in the search index, i.e. “[0065] In some implementations, the visual transformer encoder is configured to generate the embedding based on a clustering of the shared visual features of utility assets in the second subset of images. ... In some implementations, the text transformer encoder can be a natural language processing model trained to generate textual tokens based on natural language inputs.”).
As to claim 16, Wang teaches the data center system of claim 10, wherein the response comprises a structured data format with a plurality of metadata attributes of the digital asset that are mapped to a corresponding metadata attribute value (par. 0006, 0053-0055, map results, i.e. “[0006] The multi-modal search system permits for efficient and dynamic display of query results on a client device. The multi-modal search system receives input in the form of a search query, which may be a form of text and/or image data. Examples of input can include keywords, semantic search, images (with or without annotations), and bounding boxes for the images. The multi-modal search system utilizes an encoding of utility asset features (textual features, image features) from the search input to generate the search tokens. From the search index of images from an image database, the multi-modal search system identifies candidate images by comparing image tokens and textual tokens in the search index to the search tokens. Thus, the multi-modal search system allows for an identification of candidate images that share textual and/or visual features to the search input and provides the candidate images on the user interface of the client device. The user interface can display a geographical map of an electric grid and overlay a candidate image from the query results onto a position of the geographical map.”).
As to claim 17, Wang teaches the data center system of claim 10, wherein the one or more prompts include at least one of a first natural language command or question issued by a user or a second natural language command or question issued by a language model agent (Fig. 1, 2, par. 0040-0043, 0061-0065, natural language inputs such as user input features of images “<rusty transformer>”).
As to claim 18, Wang teaches the data center system of claim 10, wherein the system is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system for generating synthetic data; a system for generating synthetic data using one or more large language models (LLMs); a system for generating synthetic data using one or more vision language models (VLMs); a system for generating synthetic data using one or more multi-modal language models; or a system incorporating one or more virtual machines (VMs) (par. 0079-0080, a system implemented at least partially in a data center, i.e. “In addition, multiple computing devices may be connected, with each device providing portions of the operations (e.g., as a server bank, a group of blade servers, or a multi-processor system). In some implementations, the processor 502 is a single threaded processor. In some implementations, the processor 502 is a multi-threaded processor. In some implementations, the processor 502 is a quantum computer.”).
As to claim 19, Wang teaches a method comprising (computer implemented method in a system comprising processor and non-transitory computer readable storage medium, par. 0004-0026, 0078-0105, Fig. 5):
obtaining one or more user-specified metadata attributes of a digital asset (par. 0036-0038, generate a response by mapping image feature (metadata) to images and image token, i.e. “[0036] The multi-modal search system 106 includes an offline substage 108 to perform offline processing of the images 103 and generate a token-based search index 118. The search index 118 includes search tokens 120-1-120-N (collectively referred to as “search tokens 120”), each search token 120 including an image identifier that corresponds to a particular image from images 103. A search token in the search tokens 120 is an embedded representation of features from the images 103. For example, a search token can include an identifier for an image, labels of objects represented in the image, and pixel positions in the image for a bounding box that corresponds to an object of the objects represented in the image. The search token can include any number of objects and bounding boxes, and provides a mapping of features (e.g., text features, image features) to a unique image. In some implementations, the search token is a lower dimensional representation (e.g., an embedding and/or encoding) of the features compared to the features of the original input image…. [0038] The vision transformer encoder 114 can perform a number of embedding and/or encoding processes to map image feature data into a search image token, which can be stored as part of the search image tokens 124-1-124-N (collectively “search image tokens 124”). The vision transformer encoder 114 can be configured to determine a subset of images from the images 103 that share similar visual features (e.g., object class, defect type, or both object class and defect type) based on the bounding boxes of the subset of images.”));
processing, by a multi-modal model, a representation of the digital asset and a representation of the one or more user-specified metadata attributes to generate a response comprising one or more metadata attribute values corresponding to the one or more user-specified metadata attributes (par. 0036-0038, generate a response by mapping image feature (metadata) to images and image token, i.e. “[0036] The multi-modal search system 106 includes an offline substage 108 to perform offline processing of the images 103 and generate a token-based search index 118. The search index 118 includes search tokens 120-1-120-N (collectively referred to as “search tokens 120”), each search token 120 including an image identifier that corresponds to a particular image from images 103. A search token in the search tokens 120 is an embedded representation of features from the images 103. For example, a search token can include an identifier for an image, labels of objects represented in the image, and pixel positions in the image for a bounding box that corresponds to an object of the objects represented in the image. The search token can include any number of objects and bounding boxes, and provides a mapping of features (e.g., text features, image features) to a unique image. In some implementations, the search token is a lower dimensional representation (e.g., an embedding and/or encoding) of the features compared to the features of the original input image…. [0038] The vision transformer encoder 114 can perform a number of embedding and/or encoding processes to map image feature data into a search image token, which can be stored as part of the search image tokens 124-1-124-N (collectively “search image tokens 124”). The vision transformer encoder 114 can be configured to determine a subset of images from the images 103 that share similar visual features (e.g., object class, defect type, or both object class and defect type) based on the bounding boxes of the subset of images.”));
storing the response in a structured data format(Fig. 3, par. 0057-0064, storing image token in search index, i.e. “[0064] Similar to the text token, the image token can be stored in a search index and includes a corresponding identifier for each image in the second subset of images. In some implementations, the multi-modal search system 106 includes a machine learning network (e.g., machine learning network 112) configured to generate image tokens (e.g., search image tokens 124). For example, the machine learning network can include a vision transformer encoder configured to generate image embeddings from a clustering of the shared visual features. The machine learning network can include a visual transformer encoder (e.g., visual transformer encoder 114) trained to generate embeddings of feature data (e.g., representing features of utility asset) based on the shared visual features. Embeddings of features represented by the shared visual features can be projected in a lower dimensional representation, e.g., relative to the feature data from the images in the second subset of images.”); and
indexing the response in the structured data forma to facilitate retrieval of the digital asset as a search result candidate (Fig. 3, par. 0057-0064, storing image token in search index, i.e. “[0064] Similar to the text token, the image token can be stored in a search index and includes a corresponding identifier for each image in the second subset of images. In some implementations, the multi-modal search system 106 includes a machine learning network (e.g., machine learning network 112) configured to generate image tokens (e.g., search image tokens 124). For example, the machine learning network can include a vision transformer encoder configured to generate image embeddings from a clustering of the shared visual features. The machine learning network can include a visual transformer encoder (e.g., visual transformer encoder 114) trained to generate embeddings of feature data (e.g., representing features of utility asset) based on the shared visual features. Embeddings of features represented by the shared visual features can be projected in a lower dimensional representation, e.g., relative to the feature data from the images in the second subset of images.”).
As to claim 20, Wang teaches the method of claim 19, wherein the method is performed by at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation;a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system for generating synthetic data; a system for generating synthetic data using one or more large language models (LLMs); a system for generating synthetic data using one or more vision language models (VLMs); a system for generating synthetic data using one or more multi-modal language models; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (par. 0079-0080, a system implemented at least partially in a data center, i.e. “In addition, multiple computing devices may be connected, with each device providing portions of the operations (e.g., as a server bank, a group of blade servers, or a multi-processor system). In some implementations, the processor 502 is a single threaded processor. In some implementations, the processor 502 is a multi-threaded processor. In some implementations, the processor 502 is a quantum computer.”).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 2 and 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang, and further in view of U.S. Patent Application Publication No. 20170185603 to Dentel et al. (hereinafter “Dentel”).
As to claim 2, Wang teaches the one or more processors of claim 1. Wang does not explicitly teach wherein the one or more processing units are further to: obtain a configuration file that includes: a user-defined field representing a metadata attribute associated with the one or more metadata attribute values of the digital asset; and the one or more prompts that include natural language characters input by a user, and wherein the model processes the configuration file to generate the response that includes the one or more metadata attribute values as claimed.
Dentel teaches wherein the one or more processing units are further to: obtain a configuration file that includes: a user-defined field representing a metadata attribute associated with the one or more metadata attribute values of the digital asset; and the one or more prompts that include natural language characters input by a user, and wherein the model processes the configuration file to generate the response that includes the one or more metadata attribute values (Fig. 7, par. 0006, 0062, 0065, configuration file, i.e. “The search query block 410 may comprise a representation of a set of source-objects “Manager” 411 that the first user previously inputted, a connector 412 that is included to make the form of the composed query command similar to a sentence written in natural language grammar, and display of a query constraint “all of” 413 that the first user previously inputted. Together, elements 411-413 indicate that the first user has inputted a query constraint corresponding to an operator configuration file, which may retrieve an intersection between a set of all manager objects and a to-be-entered set of objects.”)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Wang with the teaching of Dentel because they are in the same field of endeavor. One of ordinary skill in the art at the time of the invention would have been motivated to do so because the teaching of Dentel would allow Wang to facilitate “…the query-composition platform ensures that every component of a query command composed by the user is valid with respect to objects in the database system and with respect to other components of the query command. This guarantees that the completed query command is valid and will return a non-empty set of objects as search results. The query-composition platform, operable as an added layer of abstraction on top of the database system, may reduce the difficulty in efficiently searching the database system using query commands, particularly in at least two situations: when there is no comprehensive representation or index of the relationships among data objects (e.g., a social graph), such that the path of retrieving particular objects is often unclear and when the user's knowledge about the database system is limited.…” (Dentel, par. 0002-0007)
As to claim 11, Wang teaches the data center system of claim 10. Wang does not explicitly teach wherein the one or more GPUs are further to: obtain a configuration file that includes: the one or more user-defined fields; and the one or more prompts, wherein the model processes the configuration file to generate the response that includes the one or more metadata attribute values as claimed.
Dentel teaches wherein the one or more GPUs are further to: obtain a configuration file that includes: the one or more user-defined fields; and the one or more prompts, wherein the model processes the configuration file to generate the response that includes the one or more metadata attribute values (Fig. 7, par. 0006, 0062, 0065, configuration file, i.e. “The search query block 410 may comprise a representation of a set of source-objects “Manager” 411 that the first user previously inputted, a connector 412 that is included to make the form of the composed query command similar to a sentence written in natural language grammar, and display of a query constraint “all of” 413 that the first user previously inputted. Together, elements 411-413 indicate that the first user has inputted a query constraint corresponding to an operator configuration file, which may retrieve an intersection between a set of all manager objects and a to-be-entered set of objects.”)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Wang with the teaching of Dentel because they are in the same field of endeavor. One of ordinary skill in the art at the time of the invention would have been motivated to do so because the teaching of Dentel would allow Wang to facilitate “…the query-composition platform ensures that every component of a query command composed by the user is valid with respect to objects in the database system and with respect to other components of the query command. This guarantees that the completed query command is valid and will return a non-empty set of objects as search results. The query-composition platform, operable as an added layer of abstraction on top of the database system, may reduce the difficulty in efficiently searching the database system using query commands, particularly in at least two situations: when there is no comprehensive representation or index of the relationships among data objects (e.g., a social graph), such that the path of retrieving particular objects is often unclear and when the user's knowledge about the database system is limited.…” (Dentel, par. 0002-0007)
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANHTAI V TRAN whose telephone number is (571)270-5129. The examiner can normally be reached on Monday through Thursday from 8:00 AM to 4:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Charles Rones can be reached on (571)272-4085. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ANHTAI V TRAN/Primary Examiner, Art Unit 2168