Prosecution Insights
Last updated: August 17, 2026
Application No. 18/956,107

TRAINING MULTI-MODAL FOUNDATION MODEL

Non-Final OA §103
Filed
Nov 22, 2024
Priority
Aug 23, 2024 — CN 202411171380.5
Examiner
KIM, JONATHAN C
Art Unit
Tech Center
Assignee
Baidu Online Network Technology (Beijing) Co., Ltd.
OA Round
1 (Non-Final)
74%
Grant Probability
Favorable
1-2
OA Rounds
8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
269 granted / 366 resolved
+13.5% vs TC avg
Strong +39% interview lift
Without
With
+38.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
22 currently pending
Career history
387
Total Applications
across all art units

Statute-Specific Performance

§101
20.2%
-19.8% vs TC avg
§103
50.2%
+10.2% vs TC avg
§102
11.9%
-28.1% vs TC avg
§112
10.2%
-29.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 366 resolved cases

Office Action

§103
DETAILED ACTION This Office Action is in response to the correspondence filed by the applicant on 11/22/2024. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Receipt is acknowledged of certified copies of papers submitted under 35 U.S.C. 119(a)-(d), which papers have been placed of record in the file. Allowable Subject Matter Claims 4-8 and 11-15 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-3, 9-10, and 16-20 are rejected under 35 U.S.C. 103 as being unpatentable over CHONG (US 2024/0403605 A1), and in further view of YAN (Yan Y, Wen H, Zhong S, Chen W, Chen H, Wen Q, Zimmermann R, Liang Y. Urbanclip: Learning text-enhanced urban region profiling with contrastive language-image pretraining from the web. InProceedings of the ACM Web Conference 2024 2024 May 13 (pp. 4006-4017).). REGARDING CLAIM 1, CHONG discloses a method, comprising: obtaining first [urban] data of a first sample [urban region], wherein the first [urban] data comprises a plurality of first data segments that respectively correspond to a plurality of data modalities (CHONG Fig. 2 Image Training Data 212 and Sensor Training Data 214; Par 54 – “Other architectures and model-modalities combinations are possible, such as a transformer teacher for text and next word prediction, a GBDT teacher for image data, time-series data, tabular data, etc., and all such configurations (including any number of underlying modalities and associated teachers/student encoders) are within the contemplated scope of this disclosure.”); inputting the first urban data into a multi-modal foundation model (CHONG Fig. 2 Multimodal neural network 206; Par 55 – “In some embodiments, for each modality (e.g., image data, sensor data, etc.), the features learned by each student encoder (e.g., the image student encoder 208 and the sensor student encoder 210) are enforced to be similar to the teacher models via internal feature knowledge distillation according to the following procedure.”; Par 56 -- “For example, the CNN teacher 202 can be trained over image training data 212 and the GBDT teacher 204 can be trained over the sensor training data 214.”; Par 69 – “In some embodiments, the projection layer 216 is configured to project a feature vector (e.g., an output) obtained from the image student encoder 208 to a same dimension as the feature vector obtained from the CNN teacher 202.”; Par 73 – “In some embodiments, the concatenation of outputs from the plurality of student encoders is used to train the fusion neural network 218 of the multimodal neural network 206. In some embodiments, the fusion of modality-specific features from the respective encoders enables the fusion network 218, which is part of the multimodal network 206, to fully leverage the benefit of joint multimodal data in its prediction.”) to obtain respective predicted vector representations output by the multi-modal foundation model for the plurality of first data segments (CHONG Par 90 – “At block 808, data is received from the plurality of modalities. For example, the data can include image data and/or sensor data, although the data involved is not meant to be particularly limited. At block 810, a prediction is generated from an output layer of the trained fusion neural network.”); obtaining a plurality of general-purpose foundation models that are pre-trained (CHONG Fig. 2 CNN Teacher 202 and GBDT Teacher 204; Fig. 8 – “Train a plurality of unimodal teacher models 802 [Wingdings font/0xE0] For each of the unimodal teacher models, train a respective student encoder of a plurality of student encoders usings a knowledge distillation … 804”; In other words, the teacher models are pre-trained to train the student models.), wherein each general-purpose foundation model of the plurality of general-purpose foundation models corresponds to at least one data modality of the plurality of data modalities (CHONG Par 54 – “In some embodiments, the multimodal neural network training module 150 includes a plurality of unimodal teacher models (here, a CNN teacher 202 and a GBDT teacher 204) coupled to a multimodal neural network 206. The use of a CNN teacher and a GBDT teacher is for ease of discussion and illustration only. It should be understood that the particular teacher model architectures and respective modalities chosen need not be particularly limited. Other architectures and model-modalities combinations are possible, such as a transformer teacher for text and next word prediction, a GBDT teacher for image data, time-series data, tabular data, etc., and all such configurations (including any number of underlying modalities and associated teachers/student encoders) are within the contemplated scope of this disclosure.”); for each general-purpose foundation model of the plurality of general-purpose foundation models (CHONG Fig. 2 CNN Teacher 202 and GBDT Teacher 204): generating a vector representation label of a first data segment of a corresponding data modality by using the general-purpose foundation model (CHONG Fig. 2 Soft Labels 220; Par 74 – “In some embodiments, knowledge distillation includes a soft-label knowledge distillation from all teachers (e.g., the CNN teacher 202 and GBDT teacher 204) for all modalities (e.g., for image and sensor modalities). In some embodiments, average soft-labels 220 (softened probabilities) of the various teacher models are leveraged to further train the multimodal neural network 206.”); and determining a knowledge distillation loss of the general-purpose foundation model based on the vector representation label and a predicted vector representation of the first data segment (CHONG Pars 65-66 – “In some embodiments, a modified objective function is used to enforce similarity in layerwise features. For an image modality, for example, the modified objective function for the joint training of the multimodal neural network 206 is defined according to equation (1): … where the term L(y, σ(ŷ)) trains the multimodal network 206 to predict the groundtruth label y, the second term Ldistilling(Ti img,ϕj img) denotes the internal image feature distillation from layer i in the teacher Timg (e.g., CNN teacher 202) to layer j in the student encoder ϕimg, Note that, in equation (1), L(.,.) is the cross-entropy loss and Ldistilling (.,.) is one or more of squared L2 regression, cosine loss, and KL divergence for layers involving probability distributions (e.g., attention layers). δi,jΣ{0, 1} determines the pair(s) of teacher-student layers that take part in the knowledge distillation.”; Par 75 – “In some embodiments, knowledge distillation is enforced on an output layer 222 of the multimodal neural network 206. In some embodiments, knowledge distillation is denoted by a learning operation from the aggregated probability outputs of the ensemble of teachers via a loss term according to the equation (4): L(PTemp=τ , σ(y^;Temp=τ)) where the term σ(.; Temp=τ) is the softmax temperature, PTemp=τ is the average softened probabilities from the teacher encoders, and L(.,.) is the cross-entropy loss.”; In other words, the loss function is based on the soft-labels, PTemp, of the encoders and the predcited vector, y^.); determining an overall loss of the multi-modal foundation model based on at least respective knowledge distillation losses of the plurality of general-purpose foundation models (CHONG Par 31 – “In some embodiments, knowledge distillation is also enforced on the output (predictions) of the shared multimodal neural network. For example, average soft-labels (softened probabilities) from each of the teacher models can be used as additional guidance to train the multimodal network by enforcing the output layer of the multimodal network to approximate the aggregated probability outputs (aggregated soft-labels) of the ensemble of teachers via the loss term of the final objective function of the shared multimodal neural network.”; Par 76 – “The use of soft-label knowledge distillation in this manner further modifies the objective function as described with respect to equation (5): PNG media_image1.png 262 1000 media_image1.png Greyscale ”); and adjusting parameters of the multi-modal foundation model based on the overall loss (CHONG Par 60 – “Optimization: The goal of optimization is to adjust a model's parameters to minimize the loss function. This is typically done using an optimization algorithm such as gradient descent, and can depend on the specific architecture for a given task. For example, CNNs typically use stochastic gradient descent (SGD) with momentum, while RNNs often use adaptive moment estimation”; Also note that equation (5) shows finding the multi-modal parameters (e.g., θimg, θsensor) that minimize the loss function.). CHONG does not explicitly teach the [square-bracketed] limitation. In other words, CHONG discloses analyzing multimodal data and training a multimodal model based on pre-trained unimodal models by adjusting parameters based on an overall loss function, but does not explicitly teach the multimodal data are [urban] data. YAN discloses the [urban] data. YAN discloses a method/system for training a multimodal urban data (satellite images and textual description of a region) using pre-train unimodal models (Figure 2) to predict a target urban information (e.g., urban indicators including population, carbon emission, GDP ranking etc.) comprising: obtaining first [urban] data of a first sample [urban region], wherein the first [urban] data comprises a plurality of first data segments that respectively correspond to a plurality of data modalities (YAN pg. 4007 2nd Col – “The description text 𝑇𝑔 for an urban region 𝑔 contains several individual sentences. Such text can be generated manually or using image captioning tools. E.g., by leveraging the well-trained LLM’s profound understanding of general-purpose knowledge [30, 51, 89, 102], we can generate the summary text of a given region, especially including its spatial context (e.g., POIs) that significantly reflects its land function [19].”; pg. 4009 1st Col Textual Representation Learning – “For an urban region 𝑔, a highquality text summary 𝑇𝑔 is generated from LLMs through text generation and refinement. Similar to the prior visual representation, it is desirable to encode this summary into a latent textual representation. Normally, BERT-style [41] models with encoder-only architecture can be generalized to capture latent textual representations. … where M-MSA means masked multi-head self attention operation, and 𝒆𝑇𝐸 is the token representation of added location information. We also add a learnable [CLS] token to obtain the global information.”; Pg. 4010 1st Col – “4.1.1 Datasets. The datasets used in this paper include satellite imagery, textual description, and three urban indicators for four representative cities in China: Beijing, Shanghai, Guangzhou, and Shenzhen. The satellite images obtained from Baidu Map API have a fixed size of 256×256 with a spatial resolution of around 13 meters per pixel, which leads to an area of approximately 1 𝑘𝑚2. … There exists a one-to-many relationship between images and associated texts. We filter out low-quality descriptions and then adopt a random selection to choose one high-quality summary text that matches each satellite image.”; Pg. 4010 Table 1 – “Coverage; Location Description”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of CHONG to include multimodal urban data, as taught by YAN. One of ordinary skill would have been motivated to include multimodal urban data, in order to utilize a more powerful data analysis technology for urban area discovery so that city planning/development can improved. REGARDING CLAIM 2, CHONG in view of YAN discloses the method according to claim 1, wherein the plurality of first data segments comprise point-of-interest data corresponding to a text modality, and the plurality of general-purpose foundation models comprise a language foundation model corresponding to the text modality, and wherein the generating the vector representation label of the first data segment (CHONG Par 54 – “It should be understood that the particular teacher model architectures and respective modalities chosen need not be particularly limited. Other architectures and model-modalities combinations are possible, such as a transformer teacher for text and next word prediction, a GBDT teacher for image data, time-series data, tabular data, etc., and all such configurations (including any number of underlying modalities and associated teachers/student encoders) are within the contemplated scope of this disclosure.”; Par 77 – “During inference, the trained model (e.g., the multimodal neural network 206) takes in new input data, such as an image and a text description, and generates an output (also referred to as a “prediction 224”). The prediction 224 can include, for example, a predicted label and/or a response to an input text.”; YAN pg. 4007 2nd Col – “The description text 𝑇𝑔 for an urban region 𝑔 contains several individual sentences. Such text can be generated manually or using image captioning tools. E.g., by leveraging the well-trained LLM’s profound understanding of general-purpose knowledge [30, 51, 89, 102], we can generate the summary text of a given region, especially including its spatial context (e.g., POIs) that significantly reflects its land function [19].”; pg. 4009 1st Col Textual Representation Learning – “For an urban region 𝑔, a highquality text summary 𝑇𝑔 is generated from LLMs through text generation and refinement. Similar to the prior visual representation, it is desirable to encode this summary into a latent textual representation. Normally, BERT-style [41] models with encoder-only architecture can be generalized to capture latent textual representations. … where M-MSA means masked multi-head self attention operation, and 𝒆𝑇𝐸 is the token representation of added location information. We also add a learnable [CLS] token to obtain the global information.”; Pg. 4010 1st Col – “4.1.1 Datasets. The datasets used in this paper include satellite imagery, textual description, and three urban indicators for four representative cities in China: Beijing, Shanghai, Guangzhou, and Shenzhen. The satellite images obtained from Baidu Map API have a fixed size of 256×256 with a spatial resolution of around 13 meters per pixel, which leads to an area of approximately 1 𝑘𝑚2. … There exists a one-to-many relationship between images and associated texts. We filter out low-quality descriptions and then adopt a random selection to choose one high-quality summary text that matches each satellite image.”; Pg. 4010 Table 1 – “Coverage; Location Description”) comprises: inputting at least part of the point-of-interest data into the language foundation model (YAN Pg. 4008 Figure 2: Overall framework of the proposed UrbanCLIP – “Text Refinement [Wingdings font/0xE0] Tokenize [Wingdings font/0xE0] Liner Projection [Wingdings font/0xE0] Unimodal Text Encoder [Wingdings font/0xE0]”) to obtain a description text generated by the language foundation model for the first sample urban region (YAN pg. 4009 1st Col Textual Representation Learning – “For an urban region 𝑔, a highquality text summary 𝑇𝑔 is generated from LLMs through text generation and refinement. Similar to the prior visual representation, it is desirable to encode this summary into a latent textual representation. Normally, BERT-style [41] models with encoder-only architecture can be generalized to capture latent textual representations. … where M-MSA means masked multi-head self attention operation, and 𝒆𝑇𝐸 is the token representation of added location information. We also add a learnable [CLS] token to obtain the global information.”); and encoding the description text to obtain a vector representation label of the point-of-interest data (CHONG Fig. 2 – “CNN Teacher [Wingdings font/0xE0] Soft Labels 220”; Par 31 – “In some embodiments, knowledge distillation is also enforced on the output (predictions) of the shared multimodal neural network. For example, average soft-labels (softened probabilities) from each of the teacher models can be used as additional guidance to train the multimodal network by enforcing the output layer of the multimodal network to approximate the aggregated probability outputs (aggregated soft-labels) of the ensemble of teachers via the loss term of the final objective function of the shared multimodal neural network.”; YAN pg. 4007 2nd Col – “The description text 𝑇𝑔 for an urban region 𝑔 contains several individual sentences. Such text can be generated manually or using image captioning tools. E.g., by leveraging the well-trained LLM’s profound understanding of general-purpose knowledge [30, 51, 89, 102], we can generate the summary text of a given region, especially including its spatial context (e.g., POIs) that significantly reflects its land function [19].”; pg. 4009 1st Col Textual Representation Learning – “For an urban region 𝑔, a highquality text summary 𝑇𝑔 is generated from LLMs through text generation and refinement. Similar to the prior visual representation, it is desirable to encode this summary into a latent textual representation. Normally, BERT-style [41] models with encoder-only architecture can be generalized to capture latent textual representations. … where M-MSA means masked multi-head self attention operation, and 𝒆𝑇𝐸 is the token representation of added location information. We also add a learnable [CLS] token to obtain the global information.”). In other words, both CHONG and YAN teaches a text modality. YAN further teaches the text modality data include urban data representing textual description of a region, as explained in the rejection of claim 1. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of CHONG to include multimodal urban data, as taught by YAN. One of ordinary skill would have been motivated to include multimodal urban data, in order to utilize a more powerful data analysis technology for urban area discovery so that city planning/development can improved. REGARDING CLAIM 3, CHONG in view of YAN discloses the method according to claim 1, wherein the plurality of first data segments comprise image data corresponding to an image modality, and the plurality of general-purpose foundation models comprise a visual foundation model corresponding to the image modality, and wherein the generating the vector representation label of the first data segment (CHONG Fig. 2 – “Image training Data 212 [Wingdings font/0xE0] CNN Teacher 202 [Wingdings font/0xE0] Soft Labels 220”; Par 54 – “It should be understood that the particular teacher model architectures and respective modalities chosen need not be particularly limited. Other architectures and model-modalities combinations are possible, such as a transformer teacher for text and next word prediction, a GBDT teacher for image data, time-series data, tabular data, etc., and all such configurations (including any number of underlying modalities and associated teachers/student encoders) are within the contemplated scope of this disclosure.”; Par 77 – “During inference, the trained model (e.g., the multimodal neural network 206) takes in new input data, such as an image and a text description, and generates an output (also referred to as a “prediction 224”). The prediction 224 can include, for example, a predicted label and/or a response to an input text.”) comprises: inputting the image data into the visual foundation model to obtain a vector representation label output by the visual foundation model for the image data (CHONG Fig. 2 Soft Labels 220; Par 74 – “In some embodiments, knowledge distillation includes a soft-label knowledge distillation from all teachers (e.g., the CNN teacher 202 and GBDT teacher 204) for all modalities (e.g., for image and sensor modalities). In some embodiments, average soft-labels 220 (softened probabilities) of the various teacher models are leveraged to further train the multimodal neural network 206.”). REGARDING CLAIM 9, CHONG in view of YAN discloses the method according to claim 1. YAN further discloses wherein the first urban data comprises point-of-interest data of a text modality (YAN Pg. 4008 Figure 2: Overall framework of the proposed UrbanCLIP – “Text Refinement [Wingdings font/0xE0] Tokenize [Wingdings font/0xE0] Liner Projection [Wingdings font/0xE0] Unimodal Text Encoder [Wingdings font/0xE0]”) and image data of an image modality (YAN Pg. 4008 Figure 2: Overall framework of the proposed UrbanCLIP – “Patchify [Wingdings font/0xE0] Liner Projection [Wingdings font/0xE0] Unimodal Image Encoder [Wingdings font/0xE0]”), and the multi-modal foundation model comprises a point-of-interest encoder, an image encoder, and a multi-modal transformer (YAN Pg. 4008 Figure 2: Overall framework of the proposed UrbanCLIP – “Multimodal Interaction Module”; Pg. 4009 2nd Col Modality Interaction Task – “Unlike the previous studies where cross-modality interaction is shallow (e.g., via dot product-based similarity) [52, 57], UrbanCLIP emphasizes the deep inter-modal interaction learning through layers for a contextualized feature sequence. Motivated by [42, 97], Transformer-based decoder architecture is then leveraged to fuse unimodal visual and textual representations together as multimodal representations.”), and wherein the inputting the first urban data into the multi-modal foundation model to obtain respective predicted vector representations for the plurality of first data segments (YAN pg. 4008 2nd col – “Phase 2: In the urban indicator prediction phase, we utilize a frozen unimodal image encoder for downstream tasks, by simply fine-tuning outermost multi-layer perceptrons (MLPs) with a few trainable parameters. Furthermore, we offer two optional choices, which are a flexible infusion of other spatial modalities and prompt-guided urban indicator prediction.”; Pg. 4009 2nd Col – “Prediction Stage. Through optimizing the loss function in Eq. 5, we can obtain the final text-enhanced visual representations 𝒆𝑔 from the frozen image encoder. Given any satellite image 𝐼𝑔, we can use a simple MLP to predict urban indicators as Y𝑔 = MLP (𝒆𝑔).”) comprises: inputting the point-of-interest data into the point-of-interest encoder to obtain a point-of-interest embedding sequence of the point-of-interest data (YAN Pg. 4008 Figure 2: Overall framework of the proposed UrbanCLIP – “Text Refinement [Wingdings font/0xE0] Tokenize [Wingdings font/0xE0] Liner Projection [Wingdings font/0xE0] Unimodal Text Encoder [Wingdings font/0xE0]”), wherein the point-of-interest embedding sequence comprises an embedding of each point-of-interest data unit in the point-of-interest data (YAN pg. 4007 2nd Col – “The description text 𝑇𝑔 for an urban region 𝑔 contains several individual sentences. Such text can be generated manually or using image captioning tools. E.g., by leveraging the well-trained LLM’s profound understanding of general-purpose knowledge [30, 51, 89, 102], we can generate the summary text of a given region, especially including its spatial context (e.g., POIs) that significantly reflects its land function [19].”; pg. 4009 1st Col Textual Representation Learning – “For an urban region 𝑔, a highquality text summary 𝑇𝑔 is generated from LLMs through text generation and refinement. Similar to the prior visual representation, it is desirable to encode this summary into a latent textual representation. Normally, BERT-style [41] models with encoder-only architecture can be generalized to capture latent textual representations. … where M-MSA means masked multi-head self attention operation, and 𝒆𝑇𝐸 is the token representation of added location information. We also add a learnable [CLS] token to obtain the global information.”); inputting the image data into the image encoder to obtain an image embedding sequence of the image data (YAN Pg. 4008 Figure 2: Overall framework of the proposed UrbanCLIP – “Patchify [Wingdings font/0xE0] Liner Projection [Wingdings font/0xE0] Unimodal Image Encoder [Wingdings font/0xE0]”), wherein the image embedding sequence comprises an embedding of each image data unit in the image data (YAN Pg. 4008 2nd Col Visual Representation Learning – “For an urban region 𝑔 with its satellite imagery 𝐼𝑔, we first split it into a sequence of patches 𝐼𝑝 (the default patch size is 16×16), which are then linearly embedded into a dense vector: 𝒆𝐼𝑝 = W𝑝𝐼⊤𝑝 + 𝑏𝑝 , where W𝑝 and 𝑏𝑝 are learnable parameters. The learnable positional embeddings E are further added to provide information about the relative position of each patch: 𝒆𝐼𝐸 = 𝒆𝐼𝑝+ E.”); and inputting the point-of-interest embedding sequence and the image embedding sequence into the multi-modal transformer, so that the multi-modal transformer fuses and updates the point-of-interest embedding sequence and the image embedding sequence (YAN Pg. 4008 Figure 2: Overall framework of the proposed UrbanCLIP – “Multimodal Interaction Module”; Pg. 4009 2nd Col Modality Interaction Task – “Unlike the previous studies where cross-modality interaction is shallow (e.g., via dot product-based similarity) [52, 57], UrbanCLIP emphasizes the deep inter-modal interaction learning through layers for a contextualized feature sequence. Motivated by [42, 97], Transformer-based decoder architecture is then leveraged to fuse unimodal visual and textual representations together as multimodal representations. Specifically, UrbanCLIP employs multimodal decoder layers to effectively learn joint image-text representations, by leveraging unimodal textual encoder outputs and employing cross-attention mechanisms towards image encoder outputs. The key difference between multimodal cross-attention and unimodal MSA is that cross-attention uses visual modality as a query and text modality as key and value. Besides, to generate a text description for comprehensive urban region profiling, we introduce a language modeling loss LLM that enables the model to predict the next tokenized texts autoregressively with fine granularity. Hence, the multimodal decoder can learn to maximize the conditional likelihood of the paired text T: …”), to obtain a predicted vector representation of the point-of-interest data and a predicted vector representation of the image data (YAN Pg. 4013 Figure 6 – “Case study of most similar satellite imagery matching between Beijing and other three cities through text-enhanced visual representation by UrbaCLIP”; Pg. 4013 1st Col – “In particular, for a given satellite image from a source city, we compute the cosine similarity of visual representations among all others from different target cities. We assess whether there are commonalities in terms of urban indicators and description texts generated by UrbanCLIP. We further investigate the capability of generated text to guide associated image representations to focus on similar spatial information.”). In other words, both CHONG and YAN teaches iteratively adjusting the parameters of the models with a plurality of multimodal data in the dataset. YAN further teaches the multimodal data are urban data representing urban regions visually and textually as explained in the rejection of claim 1. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of CHONG to include multimodal urban data, as taught by YAN. One of ordinary skill would have been motivated to include multimodal urban data, in order to utilize a more powerful data analysis technology for urban area discovery so that city planning/development can improved. REGARDING CLAIM 10, CHONG in view of YAN discloses the method according to claim 9. YAN further discloses wherein the point-of-interest data comprises respective names, geographic positions, (YAN Pg. 4010 1st Col – “4.1.1 Datasets. The datasets used in this paper include satellite imagery, textual description, and three urban indicators for four representative cities in China: Beijing, Shanghai, Guangzhou, and Shenzhen. The satellite images obtained from Baidu Map API have a fixed size of 256×256 with a spatial resolution of around 13 meters per pixel, which leads to an area of approximately 1 𝑘𝑚2. … There exists a one-to-many relationship between images and associated texts. We filter out low-quality descriptions and then adopt a random selection to choose one high-quality summary text that matches each satellite image.”; Pg. 4010 Table 1 – “Coverage; Location Description”; ) and categories of a plurality of points of interest located in the first sample urban region (YAN Pg. 4010 1st col – “For more modalities, an example of a positive sample could be the combination of a satellite image, a text description, the majority of POI categories as parks, and the road network of a given area. ii) better interaction with existing modalities.”), and wherein the inputting the point-of-interest data into the point-of-interest encoder to obtain the point-of-interest embedding sequence of the point-of-interest data (YAN Pg. 4008 Figure 2: Overall framework of the proposed UrbanCLIP – “Text Refinement [Wingdings font/0xE0] Tokenize [Wingdings font/0xE0] Liner Projection [Wingdings font/0xE0] Unimodal Text Encoder”) comprises: inputting the point-of-interest data into the point-of-interest encoder, so that the point-of-interest encoder performs the following (YAN Pg. 4008 2nd Col – “Phase 1: We first generate a detailed location description via LLaMA-Adapter V2 (an image-to-text foundation model) for the satellite imagery crawled from Baidu Map, thus forming a set of high-quality image-text pairs. The image and text are then fed into two unimodal encoders separately. Lastly, a multimodal interaction module is designed to align the representation of the two modalities in the latent space, with an elaborately designed cross-attention mechanism and contrastive learning objective.”) operations: representing the respective names of the plurality of points of interest as a point-of-interest name sequence, wherein each token in the point-of-interest name sequence is one point-of-interest data unit (YAN pg. 4007 2nd Col – “The description text 𝑇𝑔 for an urban region 𝑔 contains several individual sentences. Such text can be generated manually or using image captioning tools. E.g., by leveraging the well-trained LLM’s profound understanding of general-purpose knowledge [30, 51, 89, 102], we can generate the summary text of a given region, especially including its spatial context (e.g., POIs) that significantly reflects its land function [19].”); and for each token in the point-of-interest name sequence, generating an embedding of the token based on at least one of: the token, a geographic position or a category of a point of interest corresponding to the token (YAN pg. 4009 1st Col Textual Representation Learning – “For an urban region 𝑔, a highquality text summary 𝑇𝑔 is generated from LLMs through text generation and refinement. Similar to the prior visual representation, it is desirable to encode this summary into a latent textual representation. Normally, BERT-style [41] models with encoder-only architecture can be generalized to capture latent textual representations. However, such traditional bidirectional attention may encounter low-rank issues [17], potentially weakening the model’s expressive capacity and yielding limited generative capabilities. Hence, we choose a decoder-only architecture for the text encoding module. The primary distinction is that the textual representation is acquired via causally masked multi-head self-attention: … where M-MSA means masked multi-head self attention operation, and 𝒆𝑇𝐸 is the token representation of added location information. We also add a learnable [CLS] token to obtain the global information.”; pg. 4010 1st Col – “For more modalities, an example of a positive sample could be the combination of a satellite image, a text description, the majority of POI categories as parks, and the road network of a given area. ii) better interaction with existing modalities. An intuitive way is adopting cross-attention mechanisms in UrbanCLIP.”). In other words, both CHONG and YAN teaches iteratively adjusting the parameters of the models with a plurality of multimodal data in the dataset. YAN further teaches the multimodal data are urban data representing urban regions visually and textually as explained in the rejection of claim 1. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of CHONG to include multimodal urban data, as taught by YAN. One of ordinary skill would have been motivated to include multimodal urban data, in order to utilize a more powerful data analysis technology for urban area discovery so that city planning/development can improved. REGARDING CLAIM 16, CHONG in view of YAN discloses the method according to claim 1, further comprising: obtaining second urban data of a second sample urban region, an information label of the second sample urban region (CHONG Fig. 2 Image Training Data 212 and Sensor Training Data 214; Par 54 – “Other architectures and model-modalities combinations are possible, such as a transformer teacher for text and next word prediction, a GBDT teacher for image data, time-series data, tabular data, etc., and all such configurations (including any number of underlying modalities and associated teachers/student encoders) are within the contemplated scope of this disclosure.”; YAN pg. 4007 2nd Col – “The description text 𝑇𝑔 for an urban region 𝑔 contains several individual sentences. Such text can be generated manually or using image captioning tools. E.g., by leveraging the well-trained LLM’s profound understanding of general-purpose knowledge [30, 51, 89, 102], we can generate the summary text of a given region, especially including its spatial context (e.g., POIs) that significantly reflects its land function [19].”; pg. 4009 1st Col Textual Representation Learning – “For an urban region 𝑔, a highquality text summary 𝑇𝑔 is generated from LLMs through text generation and refinement. Similar to the prior visual representation, it is desirable to encode this summary into a latent textual representation. Normally, BERT-style [41] models with encoder-only architecture can be generalized to capture latent textual representations. … where M-MSA means masked multi-head self attention operation, and 𝒆𝑇𝐸 is the token representation of added location information. We also add a learnable [CLS] token to obtain the global information.”), and a trained first multi-modal foundation model (CHONG – “CNN Teacher 202”; YAN Figure 2 – “Unimodal Image Encoder”), wherein the second urban data comprises a plurality of second data segments that respectively correspond to a plurality of data modalities (CHONG Fig. 2 Image Training Data 212 and Sensor Training Data 214; Par 54 – “Other architectures and model-modalities combinations are possible, such as a transformer teacher for text and next word prediction, a GBDT teacher for image data, time-series data, tabular data, etc., and all such configurations (including any number of underlying modalities and associated teachers/student encoders) are within the contemplated scope of this disclosure.”; YAN Pg. 4008 2nd Col – “Phase 1: We first generate a detailed location description via LLaMA-Adapter V2 (an image-to-text foundation model) for the satellite imagery crawled from Baidu Map, thus forming a set of high-quality image-text pairs. The image and text are then fed into two unimodal encoders separately. Lastly, a multimodal interaction module is designed to align the representation of the two modalities in the latent space, with an elaborately designed cross-attention mechanism and contrastive learning objective.”); inputting the second urban data into the first multi-modal foundation model (CHONG Fig. 2 Multimodal neural network 206; Par 55 – “In some embodiments, for each modality (e.g., image data, sensor data, etc.), the features learned by each student encoder (e.g., the image student encoder 208 and the sensor student encoder 210) are enforced to be similar to the teacher models via internal feature knowledge distillation according to the following procedure.”; Par 56 -- “For example, the CNN teacher 202 can be trained over image training data 212 and the GBDT teacher 204 can be trained over the sensor training data 214.”; Par 69 – “In some embodiments, the projection layer 216 is configured to project a feature vector (e.g., an output) obtained from the image student encoder 208 to a same dimension as the feature vector obtained from the CNN teacher 202.”; Par 73 – “In some embodiments, the concatenation of outputs from the plurality of student encoders is used to train the fusion neural network 218 of the multimodal neural network 206. In some embodiments, the fusion of modality-specific features from the respective encoders enables the fusion network 218, which is part of the multimodal network 206, to fully leverage the benefit of joint multimodal data in its prediction.”; YAN Pg. 4008 Figure 2: Overall framework of the proposed UrbanCLIP – “Text Refinement [Wingdings font/0xE0] Tokenize [Wingdings font/0xE0] Liner Projection [Wingdings font/0xE0] Unimodal Text Encoder [Wingdings font/0xE0]”) to obtain respective segment vector representations output by the first multi-modal foundation model for the plurality of second data segments (CHONG Par 90 – “At block 808, data is received from the plurality of modalities. For example, the data can include image data and/or sensor data, although the data involved is not meant to be particularly limited. At block 810, a prediction is generated from an output layer of the trained fusion neural network.”; YAN pg. 4007 2nd Col – “The description text 𝑇𝑔 for an urban region 𝑔 contains several individual sentences. Such text can be generated manually or using image captioning tools. E.g., by leveraging the well-trained LLM’s profound understanding of general-purpose knowledge [30, 51, 89, 102], we can generate the summary text of a given region, especially including its spatial context (e.g., POIs) that significantly reflects its land function [19].”; pg. 4009 1st Col Textual Representation Learning – “For an urban region 𝑔, a highquality text summary 𝑇𝑔 is generated from LLMs through text generation and refinement. Similar to the prior visual representation, it is desirable to encode this summary into a latent textual representation. Normally, BERT-style [41] models with encoder-only architecture can be generalized to capture latent textual representations. … where M-MSA means masked multi-head self attention operation, and 𝒆𝑇𝐸 is the token representation of added location information. We also add a learnable [CLS] token to obtain the global information.”; YAN pg. 4007 2nd Col – “The description text 𝑇𝑔 for an urban region 𝑔 contains several individual sentences. Such text can be generated manually or using image captioning tools. E.g., by leveraging the well-trained LLM’s profound understanding of general-purpose knowledge [30, 51, 89, 102], we can generate the summary text of a given region, especially including its spatial context (e.g., POIs) that significantly reflects its land function [19].”; pg. 4009 1st Col Textual Representation Learning – “For an urban region 𝑔, a highquality text summary 𝑇𝑔 is generated from LLMs through text generation and refinement. Similar to the prior visual representation, it is desirable to encode this summary into a latent textual representation. Normally, BERT-style [41] models with encoder-only architecture can be generalized to capture latent textual representations. … where M-MSA means masked multi-head self attention operation, and 𝒆𝑇𝐸 is the token representation of added location information. We also add a learnable [CLS] token to obtain the global information.” determining a region vector representation of the second sample urban region based on the respective segment vector representations of the plurality of second data segments (CHONG Fig. 2 Multimodal neural network 206; Par 55 – “In some embodiments, for each modality (e.g., image data, sensor data, etc.), the features learned by each student encoder (e.g., the image student encoder 208 and the sensor student encoder 210) are enforced to be similar to the teacher models via internal feature knowledge distillation according to the following procedure.”; Par 56 -- “For example, the CNN teacher 202 can be trained over image training data 212 and the GBDT teacher 204 can be trained over the sensor training data 214.”; Par 69 – “In some embodiments, the projection layer 216 is configured to project a feature vector (e.g., an output) obtained from the image student encoder 208 to a same dimension as the feature vector obtained from the CNN teacher 202.”; Par 73 – “In some embodiments, the concatenation of outputs from the plurality of student encoders is used to train the fusion neural network 218 of the multimodal neural network 206. In some embodiments, the fusion of modality-specific features from the respective encoders enables the fusion network 218, which is part of the multimodal network 206, to fully leverage the benefit of joint multimodal data in its prediction.”; YAN Pg. 4008 Figure 2: Overall framework of the proposed UrbanCLIP – “Multimodal Interaction Module”; Pg. 4009 2nd Col Modality Interaction Task – “Unlike the previous studies where cross-modality interaction is shallow (e.g., via dot product-based similarity) [52, 57], UrbanCLIP emphasizes the deep inter-modal interaction learning through layers for a contextualized feature sequence. Motivated by [42, 97], Transformer-based decoder architecture is then leveraged to fuse unimodal visual and textual representations together as multimodal representations.”); inputting the region vector representation of the second sample urban region (CHONG Fig. 2; Par 73 – “In some embodiments, the concatenation of outputs from the plurality of student encoders is used to train the fusion neural network 218 of the multimodal neural network 206. In some embodiments, the fusion of modality-specific features from the respective encoders enables the fusion network 218, which is part of the multimodal network 206, to fully leverage the benefit of joint multimodal data in its prediction.”) into a prediction network to obtain predicted information output by the prediction network for the second sample urban region (CHONG Par 90 – “At block 808, data is received from the plurality of modalities. For example, the data can include image data and/or sensor data, although the data involved is not meant to be particularly limited. At block 810, a prediction is generated from an output layer of the trained fusion neural network.”; YAN Pg. 4008 Figure 2: Overall framework of the proposed UrbanCLIP – “Multimodal Interaction Module”; Pg. 4009 2nd Col Modality Interaction Task – “Unlike the previous studies where cross-modality interaction is shallow (e.g., via dot product-based similarity) [52, 57], UrbanCLIP emphasizes the deep inter-modal interaction learning through layers for a contextualized feature sequence. Motivated by [42, 97], Transformer-based decoder architecture is then leveraged to fuse unimodal visual and textual representations together as multimodal representations. Specifically, UrbanCLIP employs multimodal decoder layers to effectively learn joint image-text representations, by leveraging unimodal textual encoder outputs and employing cross-attention mechanisms towards image encoder outputs. The key difference between multimodal cross-attention and unimodal MSA is that cross-attention uses visual modality as a query and text modality as key and value. Besides, to generate a text description for comprehensive urban region profiling, we introduce a language modeling loss LLM that enables the model to predict the next tokenized texts autoregressively with fine granularity. Hence, the multimodal decoder can learn to maximize the conditional likelihood of the paired text T: …”); and adjusting at least parameters of the prediction network based on the predicted information and the information label (CHONG Par 60 – “Optimization: The goal of optimization is to adjust a model's parameters to minimize the loss function. This is typically done using an optimization algorithm such as gradient descent, and can depend on the specific architecture for a given task. For example, CNNs typically use stochastic gradient descent (SGD) with momentum, while RNNs often use adaptive moment estimation”; Also note that equation (5) shows finding the multi-modal parameters (e.g., θimg, θsensor) that minimize the loss function.; YAN Pg. 4010 2nd Col – “To assess the prediction performance, we adopt three commonly used evaluation metrics: coefficient of determination (𝑅2), rooted mean squared error (RMSE), and mean absolute error (MAE) [34, 90]. Higher 𝑅2, and lower RMSE, MAE means better performance. As for the default implementation of UrbanCLIP, Vision Transformer (ViT) [18] and the first half of transformer decoder are applied to convert the satellite image and location description into their unimodal representations, respectively; and the rest of transformer decoder can be used for multimodal interaction to generate image-text representations. The parameter initialization follows the setting from [13, 33]. Adam optimizer is chosen to minimize the training loss during parameter learning.”). In other words, both CHONG and YAN teaches iteratively adjusting the parameters of the models with a plurality of multimodal data in the dataset. YAN further teaches the multimodal data are urban data representing urban regions visually and textually as explained in the rejection of claim 1. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of CHONG to include multimodal urban data, as taught by YAN. One of ordinary skill would have been motivated to include multimodal urban data, in order to utilize a more powerful data analysis technology for urban area discovery so that city planning/development can improved. REGARDING CLAIM 17, CHONG in view of YAN discloses the method according to claim 1, further comprising: obtaining urban data of a target urban region to be identified, wherein the urban data comprises a plurality of data segments that respectively correspond to the plurality of data modalities (CHONG Fig. 2 Image Training Data 212 and Sensor Training Data 214; Par 54 – “Other architectures and model-modalities combinations are possible, such as a transformer teacher for text and next word prediction, a GBDT teacher for image data, time-series data, tabular data, etc., and all such configurations (including any number of underlying modalities and associated teachers/student encoders) are within the contemplated scope of this disclosure.”; YAN Pg. 4008 Figure 2: Overall framework of the proposed UrbanCLIP – “Text Refinement [Wingdings font/0xE0] Tokenize [Wingdings font/0xE0] Liner Projection [Wingdings font/0xE0] Unimodal Text Encoder [Wingdings font/0xE0]” and “Patchify [Wingdings font/0xE0] Liner Projection [Wingdings font/0xE0] Unimodal Image Encoder [Wingdings font/0xE0]”); inputting the urban data into a trained first multi-modal foundation model to obtain respective segment vector representations output by the multi-modal foundation model for the plurality of data segments (CHONG Fig. 2 Multimodal neural network 206; Par 55 – “In some embodiments, for each modality (e.g., image data, sensor data, etc.), the features learned by each student encoder (e.g., the image student encoder 208 and the sensor student encoder 210) are enforced to be similar to the teacher models via internal feature knowledge distillation according to the following procedure.”; Par 56 -- “For example, the CNN teacher 202 can be trained over image training data 212 and the GBDT teacher 204 can be trained over the sensor training data 214.”; Par 69 – “In some embodiments, the projection layer 216 is configured to project a feature vector (e.g., an output) obtained from the image student encoder 208 to a same dimension as the feature vector obtained from the CNN teacher 202.”; Par 73 – “In some embodiments, the concatenation of outputs from the plurality of student encoders is used to train the fusion neural network 218 of the multimodal neural network 206. In some embodiments, the fusion of modality-specific features from the respective encoders enables the fusion network 218, which is part of the multimodal network 206, to fully leverage the benefit of joint multimodal data in its prediction.”; YAN Pg. 4008 Figure 2: Overall framework of the proposed UrbanCLIP – “Multimodal Interaction Module”; Pg. 4009 2nd Col Modality Interaction Task – “Unlike the previous studies where cross-modality interaction is shallow (e.g., via dot product-based similarity) [52, 57], UrbanCLIP emphasizes the deep inter-modal interaction learning through layers for a contextualized feature sequence. Motivated by [42, 97], Transformer-based decoder architecture is then leveraged to fuse unimodal visual and textual representations together as multimodal representations. Specifically, UrbanCLIP employs multimodal decoder layers to effectively learn joint image-text representations, by leveraging unimodal textual encoder outputs and employing cross-attention mechanisms towards image encoder outputs. The key difference between multimodal cross-attention and unimodal MSA is that cross-attention uses visual modality as a query and text modality as key and value. Besides, to generate a text description for comprehensive urban region profiling, we introduce a language modeling loss LLM that enables the model to predict the next tokenized texts autoregressively with fine granularity. Hence, the multimodal decoder can learn to maximize the conditional likelihood of the paired text T: …”); determining a region vector representation of the target urban region based on the respective segment vector representations of the plurality of data segments (CHONG Fig. 2; Par 73 – “In some embodiments, the concatenation of outputs from the plurality of student encoders is used to train the fusion neural network 218 of the multimodal neural network 206. In some embodiments, the fusion of modality-specific features from the respective encoders enables the fusion network 218, which is part of the multimodal network 206, to fully leverage the benefit of joint multimodal data in its prediction.”; YAN Pg. 4008 Figure 2: Overall framework of the proposed UrbanCLIP – “Multimodal Interaction Module”; Pg. 4009 2nd Col Modality Interaction Task – “Unlike the previous studies where cross-modality interaction is shallow (e.g., via dot product-based similarity) [52, 57], UrbanCLIP emphasizes the deep inter-modal interaction learning through layers for a contextualized feature sequence. Motivated by [42, 97], Transformer-based decoder architecture is then leveraged to fuse unimodal visual and textual representations together as multimodal representations. Specifically, UrbanCLIP employs multimodal decoder layers to effectively learn joint image-text representations, by leveraging unimodal textual encoder outputs and employing cross-attention mechanisms towards image encoder outputs. The key difference between multimodal cross-attention and unimodal MSA is that cross-attention uses visual modality as a query and text modality as key and value. Besides, to generate a text description for comprehensive urban region profiling, we introduce a language modeling loss LLM that enables the model to predict the next tokenized texts autoregressively with fine granularity. Hence, the multimodal decoder can learn to maximize the conditional likelihood of the paired text T: …”; Pg. 4013 Figure 6 – “Case study of most similar satellite imagery matching between Beijing and other three cities through text-enhanced visual representation by UrbaCLIP”; Pg. 4013 1st Col – “In particular, for a given satellite image from a source city, we compute the cosine similarity of visual representations among all others from different target cities. We assess whether there are commonalities in terms of urban indicators and description texts generated by UrbanCLIP. We further investigate the capability of generated text to guide associated image representations to focus on similar spatial information.”); and identifying target information of the target urban region based on the region vector representation by using a prediction network (CHONG Par 90 – “At block 808, data is received from the plurality of modalities. For example, the data can include image data and/or sensor data, although the data involved is not meant to be particularly limited. At block 810, a prediction is generated from an output layer of the trained fusion neural network.”; YAN Pg. 4013 Figure 6 – “Case study of most similar satellite imagery matching between Beijing and other three cities through text-enhanced visual representation by UrbaCLIP”; Pg. 4013 1st Col – “In particular, for a given satellite image from a source city, we compute the cosine similarity of visual representations among all others from different target cities. We assess whether there are commonalities in terms of urban indicators and description texts generated by UrbanCLIP. We further investigate the capability of generated text to guide associated image representations to focus on similar spatial information.”). In other words, both CHONG and YAN teaches iteratively adjusting the parameters of the models with a plurality of multimodal data in the dataset. YAN further teaches the multimodal data are urban data representing urban regions visually and textually as explained in the rejection of claim 1. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of CHONG to include multimodal urban data, as taught by YAN. One of ordinary skill would have been motivated to include multimodal urban data, in order to utilize a more powerful data analysis technology for urban area discovery so that city planning/development can improved. REGARDING CLAIM 18, CHONG in view of YAN discloses the method according to claim 17. YAN further discloses wherein the target information comprises any of the following information: a type, a population, a traffic flow, or an economic index (YAN Pg. 4011 1st Col – “UrbanCLIP achieves promising results across all three urban indicators, with carbon emission being the best, followed by population, and GDP ranking last. The average 𝑅2 improvement percentages for carbon emission, population and GDP prediction are 12.07%, 5.83% and 0.52%, respectively”; pg. 4008 2nd col – “Phase 2: In the urban indicator prediction phase, we utilize a frozen unimodal image encoder for downstream tasks, by simply fine-tuning outermost multi-layer perceptrons (MLPs) with a few trainable parameters. Furthermore, we offer two optional choices, which are a flexible infusion of other spatial modalities and prompt-guided urban indicator prediction.”; Pg. 4009 2nd Col – “Prediction Stage. Through optimizing the loss function in Eq. 5, we can obtain the final text-enhanced visual representations 𝒆𝑔 from the frozen image encoder. Given any satellite image 𝐼𝑔, we can use a simple MLP to predict urban indicators as Y𝑔 = MLP (𝒆𝑔).”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of CHONG to include an urban information as a target information, as taught by YAN. One of ordinary skill would have been motivated to include an urban information as a target information, in order to utilize a more powerful data analysis technology for urban area discovery so that city planning/development can improved. REGARDING CLAIM 19, CHONG in view of YAN discloses an electronic device, comprising: a processor; and a memory communicatively connected to the processor; wherein the memory stores instructions executable by the processor, and the instructions, when executed by the processor, cause the processor to perform operations comprising: performing the steps of claim 1; thus, it is rejected under the same rationale. REGARDING CLAIM 20, CHONG in view of YAN discloses a non-transitory computer readable storage medium storing computer instructions, wherein the computer instructions are configured to enable a computer to perform operations comprising: performing the steps of claim 1; thus, it is rejected under the same rationale. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JONATHAN C KIM whose telephone number is (571)272-3327. The examiner can normally be reached Monday to Friday 8:00 AM thru 4:00 PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew C Flanders can be reached at 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JONATHAN C KIM/Primary Examiner, Art Unit 2655
Read full office action

Prosecution Timeline

Nov 22, 2024
Application Filed
Jul 24, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12700403
METHOD, APPARATUS, ELECTRONIC DEVICE AND STORAGE MEDIUM FOR TEXT CONTENT MATCHING
2y 5m to grant Granted Aug 04, 2026
Patent 12688850
ELECTRONIC DEVICE AND CONTROL METHOD THEREFOR
2y 5m to grant Granted Jul 21, 2026
Patent 12670903
Short-Lived Repeat Voice Commands
3y 7m to grant Granted Jun 30, 2026
Patent 12670906
VEHICULAR RECORDING CONTROL DEVICE AND RECORDING METHOD
2y 4m to grant Granted Jun 30, 2026
Patent 12640148
Network Microphone Device With Command Keyword Conditioning
3y 6m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
74%
Grant Probability
99%
With Interview (+38.9%)
2y 5m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 366 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month