Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed 5/12/2026 have been fully considered but they are not persuasive.
Regarding applicant arguments for 35. U.S.C. 101 on page 9 and 14 the applicant argued “Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The rejection is moot with respect to canceled claims 4-8 and 13-17. To the extent the Examiner believes the rejection applies to the amended claims, the Applicant traverses this rejection as follows …,
Amended claim I does not merely recite generic data collection, generic analysis, and generic output. Instead, amended independent claim I defines that a camera in the electronic device acquires image data, and a text collector in the electronic device collects text data; the image data and the text data are aligned to obtain image and text modal data; the image modal data is divided into image blocks, the text modal data is divided into text symbols, which may reduce data processing volume of a single semantic encoding process, effectively improve the efficiency of semantic encoding processing, enhance the semantic encoding effect of image/text modal data, and improve the accuracy of semantic representation…,
Taken individually and as an ordered combination, these limitations do not amount to well-understood, routine, and conventional activity. The claimed combination provides a specific technical mechanism for improving computer-based multimodal token recognition, particularly by addressing the poor generality, poor generalization, and inaccurate cross-modal token recognition associated with conventional approaches. Therefore, amended claim 1 recites additional elements that amount to significantly more than any alleged judicial exception…,
Accordingly, claims 10 and 19 are patent eligible for at least the same reasons as amended claim 1. The dependent claims are patent eligible at least by virtue of their dependence. For at least these reasons discussed above, Applicant requests withdrawal of the§ 101 rejection.” The applicant argues that the amended claim limitations solve a technical problem as supported by the amended independent claims. However the amended limitations have not been examined rendering the argument moot and not persuasive. See the update 101 rejection.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-3, 9-12, 18-20 rejected under 35 U.S.C. 101 because the claimed invention is directed to abstract idea without significantly more. The claim(s) recite(s) significantly more. The subject matter eligibility test for products and process is describe below for claim 1 in view of dependent claims.
Regarding claim 1:
Step 1: Is the claim to a process machine manufacture or composition of matter?
Yes – Claim 1 recites a method and that falls under the statutory categories.
Step 2A Prong 1: Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes – The claim recites the following:
“determining the plurality of text tokens as a second token;” - The limitations recites a mental process of determining the plurality of image block tokens as the first token (see MPEP 2106.04(a)(2)III).
“determining the plurality of text tokens as a second token;” The limitations recites a mental process of determining the plurality of text tokens as the second tokens (see MPEP 2106.04(a)(2)III).
“determining a grounded token matching the cluster description information from a grounded dictionary as an initial grounded token,” - The limitations recites a mental process of determining a grounded token that matches the cluster description information from a grounded dictionary as the initial grounded token. (see MPEP 2106.04(a)(2)III).
“determining the similarity information as cluster description information between the first token and the second token, wherein the target image block token and the target text token are in a same data category obtained by clustering;” - The limitations recites a mental process of determining similarity information as cluster description information between the first token and the second token having similar data category. (see MPEP 2106.04(a)(2)III).
“obtaining an associated token [by a pre-trained grounded token fusion transformer] based on fusing and encoding the first token, the second token, and the initial grounded token, so that the first token and the second token are aligned in a space of the initial grounded token” - The limitations recites a mental process of obtaining an associated token that has a similarity that satisfies a preset condition. (see MPEP 2106.04(a)(2)III).
“obtaining a fused token by fusing the first token and the second token based on the associated token,” - The limitations recites a mental process of obtaining a fused token. (see MPEP 2106.04(a)(2)III).
“and determining the token information as a target shared token between the first modal data and the second modal data.” - The limitations recites a mental process of determining the token information as a target shared token. (see MPEP 2106.04(a)(2)III).
Step 2 Prong 2: Does the claim recite additional elements that integrate the judicial exception into a particular application? No –
The claim includes the additional element(s):
“A method for cross-modal token recognition a token based on artificial intelligence recognizing a token, performed by an electronic device, comprising: acquiring image data by a camera in the electronic device, and collecting text data by a text collector in the electronic device;”
The additional elements fall under Insignificant Extra-Solution Activity as mere data gathering by obtaining data form a camera. See MPEP 2106.5(g).
“obtaining first modal data and second modal data by aligning the image data and the text data, wherein the first modal is an image modal and the second modal is a text modal;”
The additional elements fall under Insignificant Extra-Solution Activity as mere data gathering by obtaining data for the modals. See MPEP 2106.5(g).
“dividing the first modal data into a plurality of image blocks, wherein the plurality of image blocks comprise pieces of image pixel information, inputting a first sequence with the pieces of image pixel information to a visual transformer in the electronic device, encoding the pieces of image pixel information in the first sequence by a multi-layer attention mechanism of the visual transformer to obtain a plurality of image block tokens, and determining the plurality of image block tokens as a first token;”
The additional elements fall under “apply it” as using a generic computer to divided the first modal data into block and using a transformer to produce a tokens. See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)).
“dividing the second modal data into a plurality of text symbols, inputting a second sequence with the plurality of text symbols into a text transformer in the electronic device, encoding the plurality of text symbols in the second sequence by a multi-layer attention mechanism of the text transformer to obtain a plurality of text tokens, and determining the plurality of text tokens as a second token;”
The additional elements fall under “apply it” as using a generic computer to divide the second modal data into a plurality of text symbols and using a transformer to encode them to produce text tokens. See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)).
“obtaining similarity information between a target image block token and a target text token,”
The additional elements fall under Insignificant Extra-Solution Activity as mere data gathering by obtaining similarity data for the modal. See MPEP 2106.5(g).
“wherein the grounded dictionary comprises pieces of cluster description information and grounded tokens matching the pieces of cluster description information;”
The additional elements fall under Insignificant Extra-Solution Activity as mere data gathering by obtaining similarity data for the modal. See MPEP 2106.5(g).
“[obtaining an associated token] by a pre-trained grounded token fusion transformer [based on fusing and encoding the first token, the second token, and the initial grounded token, so that the first token and the second token are aligned in a space of the initial grounded token]”
The additional elements fall under “apply it” as using a generic computer to obtain an associated token by implementing a pre-train grounded token fusion transformer. See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)).
“obtaining token information output by a pre-trained token decoder based on the fused token,”
The additional elements fall under Insignificant Extra-Solution Activity as mere data gathering by obtaining token information. See MPEP 2106.5(g).
Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception?
No - The claim does not include additional elements that are sufficient to amount to a significantly more than the judicial exemption. As an order whole, the claim is directed determining an association between modal tokens. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements of dividing, encoding, and parsing fall under using generic computer to apply an exemption and mere data gathering. The method does not improve on the function of a computer, transforms an article into another article, nor is it applied by a particular machine, making the claim not patent eligible.
Regarding claim 2:
Step 2A Prong 1:
“determining ” – The limitations recites a mental process of determining the target token between models (see MPEP 2106.04(a)(2)III).
Step 2A Prong 2, Step 2B: The additional element(s):
“wherein determining
obtaining a first target token by processing the first token based on the associated token;
obtaining a second target token by processing the second token based on the associated token;”
The additional elements fall under Insignificant Extra-Solution Activity as mere data gathering by obtaining data for the models based on association token. See MPEP 2106.5(g). The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application.
Regarding claim 3:
Step 2A Prong 1:
“wherein obtaining the first target token comprises: aligning the associated token and the first token, and determining the aligned first token as the first target token; and obtaining the second target token comprises: aligning the associated token and the second token, and determining the aligned second token as the second target token” – The limitations recites a mental process of aligning and determining the alignment of the tokens (see MPEP 2106.04(a)(2)III).
Step 2A Prong 2, Step 2B: The additional element(s):
No additional elements. The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application
Regarding claim 9:
Step 2A Prong 1:
“determining fusion weight information based on the similarity information; and”
The limitations recites a mental process of determining a weight to base on the similarity (see MPEP 2106.04(a)(2)III).
Step 2A Prong 2, Step 2B: The additional element(s):
“The method of claim 8, wherein obtaining the associated token by fusing and encoding the first token, the second token and the initial grounded token, comprises: obtaining the associated token by fusing and encoding the first token, the second token and the initial grounded token based on the fusion weight information.”
The additional elements fall under “apply it” as using a generic computer to obtain the associated token by fusing and encoding. See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)).
The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application.
Claims 10-12 and 18 recite a system and are analogous to the method of claims 1-3 and 9. Therefore, the rejections of claim 1-3, and 9 above applies to claims 10-12 and 18.
Claims 19-20 recite a computer readable medium product and are analogous to the method of claims 1-2. Therefore, the rejections of claim 1-2 above applies to claims 19-20.
Allowable Subject Matter
Claim 1-3, 9, 10-12, 18-20 would be allowable if rewritten or amended to overcome the rejection(s) under 35 U.S.C. 101, set forth in this Office action.
Regarding claim 1 and analogous 10 and 19, the limitation A method for cross-modal token recognition a token based on artificial intelligence recognizing a token, performed by an electronic device, comprising:
acquiring image data by a camera in the electronic device, and collecting text data by a text collector in the electronic device;
Wang et al. (US20200311542A1) (“Wang”) teaches a camera device acquiring an image and processing the image by text collector (Wang Para 0033 line 1-6, FIG. 1 shows an overview of a computing environment for training and applying an encoder component 102 that addresses at least the above technical challenges. The computing environment includes a training environment 104 that includes a training component 106 for producing a machine-trained model 108
para 0060, FIG. 6 shows a third application 602 of the encoder component 102. In this example, a digital camera 604 takes a digital photograph of a product [acquiring image data by a camera in the electronic device], here a book cover 606. An optical character recognition component 608 then converts the resultant image into input text. The encoder component 102 next transforms the input text into the embedding vector 124 in the same manner described above [, and collecting text data by a text collector in the electronic device].);
Regarding the limitations of claim 1:
obtaining first modal data and second modal data by aligning the image data and the text data, wherein the first modal is an image modal and the second modal is a text modal;
dividing the first modal data into a plurality of image blocks, wherein the plurality of image blocks comprise pieces of image pixel information, inputting a first sequence with the pieces of image pixel information to a visual transformer in the electronic device, encoding the pieces of image pixel information in the first sequence by a multi-layer attention mechanism of the visual transformer to obtain a plurality of image block tokens, and determining the plurality of image block tokens as a first token;
dividing the second modal data into a plurality of text symbols, inputting a second sequence with the plurality of text symbols into a text transformer in the electronic device, encoding the plurality of text symbols in the second sequence by a multi-layer attention mechanism of the text transformer to obtain a plurality of text tokens, and determining the plurality of text tokens as a second token;
obtaining similarity information between a target image block token and a target text token,
Luo, Huaishao, et al. "Clip4clip: An empirical study of clip for end to end video clip retrieval." arXiv preprint arXiv:2104.08860 (2021) (“Luo”) teaches a method of dividing video images into patches and processing the image with a Video Encode. Further Luo teaches diving text into tokens and processing them using Text Encoder (Transformer) and determining similarity information (Luo page 4, Fig 1,
PNG
media_image1.png
564
1072
media_image1.png
Greyscale
[wherein the first modal is an image modal and the second modal is a text modal;]
Figure 1: The framework of CLIP4Clip, which comprises three components, including two single-modal encoders and a similarity calculator. The model takes a video-text pair as input. For the input video, we first sample the input video into ordinal frames (images). Next, these image frames are reshaped into a sequence of flattened 2D patches. These patches are mapped to the 1D sequence of embeddings with a linear patch embedding layer and input to the image encoder for representation as in ViT (Dosovitskiy et al., 2021). Finally, the similarity calculator predicts the similarity score between the text representation and representation sequence of these frames [obtaining first modal data and second modal data by aligning the image data and the text data,]. We investigate three types of similarity calculators in this work, including parameter-free, sequential, and tight types. ⊗ means cosine similarity. We initial the two single-modal encoders with CLIP (ViT-B/32) (Radford et al., 2021)
Page 3,
PNG
media_image2.png
335
583
media_image2.png
Greyscale
Page 3, 3.1 Video Encoder
To get the video representation, we first extract the frames from the video clip and then encode them via a video encoder to obtain a sequence of features. In this paper, we adopt the ViT-B/32 (Dosovitskiy et al., 2021) with 12 layers and the patch size 32 as our video encoder. Concretely, we use the pretrained CLIP (ViT-B/32) (Radford et al., 2021) as our backbone and mainly consider transferring the image representation to video representation. The pre-trained CLIP (ViT-B/32) is effective for the video-text retrieval task in this paper [inputting a first sequence with the pieces of image pixel information to a visual transformer in the electronic device, encoding the pieces of image pixel information in the first sequence by a multi-layer attention mechanism of the visual transformer to obtain a plurality of image block tokens, and determining the plurality of image block tokens as a first token].
The ViT (Dosovitskiy et al., 2021) first extracts non-overlapping image patches, then performs a linear projection to project them into 1D tokens, and exploits the transformer architecture to model the interaction between each patch of the input image to get the final representation. Following the ViT and CLIP, we use the output from the [class] token as the image representation. For the input frame sequence of video
v
i
=
{
v
i
1
,
v
i
2
,
v
i
|
v
i
|
}
, the generated representation can denote as
Z
i
=
{
z
i
1
,
z
i
2
,
z
i
|
v
i
|
}
[dividing the first modal data into a plurality of image blocks, wherein the plurality of image blocks comprise pieces of image pixel information].
Figure 1.
PNG
media_image3.png
316
456
media_image3.png
Greyscale
[encoding the plurality of text symbols in the second sequence by a multi-layer attention mechanism of the text transformer to obtain a plurality of text tokens, and determining the plurality of text tokens as a second token;]
3.2 Text Encoder
We directly apply the text encoder from the CLIP to generate the caption representation. The text encoder is a Transformer (Vaswani et al., 2017) with the architecture modifications described in (Radford et al., 2019). It is a 12-layer 512-wide model with 8 attention heads. Following CLIP, the activations from the highest layer of the transformer at the [EOS] token are treated as the feature representation of the caption. For the caption
t
j
∈
T
, we denote the representation as
w
j
. [dividing the second modal data into a plurality of text symbols, inputting a second sequence with the plurality of text symbols into a text transformer in the electronic device,]
page 3,
PNG
media_image4.png
211
779
media_image4.png
Greyscale
[obtaining similarity information]
para 4, 3.3 Similarity Calculator, Parameter-free type line 1-9, According to the CLIP, the frames representation Zi and the caption representation wj have been layer normalized and linearly projected into a multi-modal embedding space through the large-scale pretraining on the image-text pairs. The natural idea is to employ a parameter-free type to calculate similarity directly with the image/frame from the video perspective [between a target image block token and a target text token].),
Regarding the limitation:
“and determining the similarity information as cluster description information between the first token and the second token, wherein the target image block token and the target text token are in a same data category obtained by clustering;-”
Messina, Nicola, et al. "Towards efficient cross-modal visual textual retrieval using transformer-encoder deep features." 2021 International Conference on Content-Based Multimedia Indexing (CBMI). IEEE, 2021 (“Messina”) teaches clustering visual and textual concepts from a training set to create a codebook- (Messina page 3 , Creating the codebook: The codebook can be produced by collecting a large amount of visual and textual concepts from the training set and then performing clustering as in the standard Bag of Visual Words model. For this reason, we produce a large set of mixed visual and textual concepts:
PNG
media_image5.png
73
205
media_image5.png
Greyscale
We downsample C so that |C|
-
~
100k concepts.
At this point, kmeans is used to produce p clusters. The p centroids represent our codebook of concepts. Given that the word and the visual word spaces correspond, it is also possible to create a common codebook by using only the textual words from all the sentences. If we follow this methodology, we can consider the top p most common words appearing in all the Sj of the training set that are also present in the English dictionary and which are not stop-words.
Page 4 Fig. 2,
PNG
media_image6.png
330
592
media_image6.png
Greyscale
);
Regarding the limitations:
determining a grounded token matching the cluster description information from a grounded dictionary as an initial grounded token, wherein the grounded dictionary comprises pieces of cluster description information and grounded tokens matching the pieces of cluster description information;
and “obtaining an associated token by a pre-trained grounded token fusion transformer based on fusing and encoding the first token, the second token, and the initial grounded token, so that the first token and the second token are aligned in a space of the initial grounded token, wherein the associated token is a cross-modal shared token among the first token and the second token having a similarity greater than a threshold”
Zhicheng Huang, Zhaoyang Zeng, Yupan Huang, Bei Liu, Dongmei Fu, Jianlong Fu, Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning (2021) (“Huang”) teaches an end to end framework that that learns from text and images to represent areas of an image to text by associating and using visual dictionary (Huang
PNG
media_image7.png
639
959
media_image7.png
Greyscale
Page 3. Approach, 3. Approach para 1 line 1-11
The overall architecture of our proposed vision-language pre-training framework SOHO is shown in Figure 2. SOHO is an end-to-end framework, which consists of a trainable CNN-based visual encoder, a visual dictionary (VD) embedding module, and a multi-layer Transformer. The visual encoder takes an image as input and produces the visual features. VD embedding module is designed to aggregate diverse visual semantic information into visual tokens with a proposed visual dictionary. The Transformer is adopted to fuse features from visual and language modalities, and produce task-specific output.
Page 8, 4.4. Visualization of Visual Dictionary
To share insights on what the proposed Visual Dictionary (VD) learned, we visualize some representative VD indices in Figure 3. As introduce in Sec 3.2, a VD index is correlated with many visual features, where each visual feature corresponds to an image patch. We randomly sample some indices from VD and visualize their corresponding image patches. As shown in Figure 3, the VD groups meaningful and consistent image patches into different indices, which reflects an abstraction of visual semantics. The visualization shows the strong capability of the learned VD. More cases can be found in supplementary material).
However Wang, Luo and Messina, Huang fail to teach, in combination or alone, “wherein the associated token is a cross-modal shared token among the first token and the second token having a similarity greater than a threshold” and “obtaining a fused token by fusing the first token and the second token based on the associated token, obtaining token information output by a pre-trained token decoder based on the fused token, and determining the token information as a target shared token between the first modal data and the second modal data” as recited in claim 1, in combination with the remaining features and elements of the claimed invention.
Independent claims 10 and 19 would be allowable for the same reasons cited in claim 1. The remaining claims would be allowable because they depend on one of allowable independent claims 1, 10 and 19.
Pertinent Prior Art
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Dosovitskiy, Alexey, et al. "An image is worth 16x16 words: Transformers for image recognition at scale." arXiv preprint arXiv:2010.11929 (2020) – teaches an Vision Transformer (ViT) that takes in image blocks for image recognition.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALFREDO CAMPOS whose telephone number is (571)272-4504. The examiner can normally be reached 7:00 - 4:00 pm M - F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J. Huntley can be reached at (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALFREDO CAMPOS/Examiner, Art Unit 2129
/MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129