DETAILED ACTION
This office action is in response to amendments filed on 09/22/2025.
Claims 1, 4, 10-11, 15-16, and 21 have been amended. Claims 2-3, 14, and 23 have been canceled. Claims 24-27 have been added. Claims 1, 4-13, 15-22, and 24-27 are pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after January 27, 2022, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Rejections under 35 USC § 112(b):
In light of applicant’s amendments to the claims (pg. 2-6), the previous rejections under 35 USC § 112(b) have been withdrawn. However, applicant’s amendments include claim limitations which are interpreted under 35 U.S.C. 112(f), and which lack corresponding structure disclosed in the written description. Therefore, new rejections under 35 USC § 112(a) and 112(b) have been introduced.
Prior Art Rejections:
Applicant's arguments regarding the prior art rejections have been fully considered but they are not persuasive.
Applicant argues (pg. 9-10) in regard to claim 1 that Chordia’s model architecture includes a BiLSTM between the text-based transformer and co-attention block, while the claimed model architecture feeds the output of the text-based transformer directly to the fusion processor’s cross-modal attention module. Examiner respectfully disagrees with the assertion that the claim precludes the placement of a BiLSTM between the text-based transformer and the fusion processor, and notes that Chordia’s model architecture including a text-based transformer, BiLSTM, and co-attention block falls within the broadest reasonable interpretation of the claimed “text-based uni-model transformer processor for identifying features of one or more items based on said text data and generating a text transformer processor output” and “fusion processor including a cross-modal attention module for combining the text transformer output and the image transformer output to form a multi-modal representation,” as in Chordia, the output of the text-based transformer, after being processed by the BiLSTM, is combined with the image representation using the co-attention block.
Applicant argues (pg. 10) in regard to claim 1 that Chordia discloses using a Resnet model to extract image embeddings rather than the claimed image-based transformer. Examiner respectfully notes that, as can be seen in the rejection below, the image-based transformer is taught by Dosovitskiy, and one of ordinary skill in the art would have been motivated to replace Chordia’s Resnet model with Dosovitskiy’s image-based transformer for the explicitly stated benefit of “attain[ing] excellent results compared to state-of-the-art convolutional networks while requiring substantially fewer computational resources to train” (Dosovitskiy, pg. 1, Abstract).
Applicant argues (pg. 10) in regard to claim 1 that while Dosovitskiy teaches an image-based transformer, it does not teach the claimed text-based transformer or fusion processor including a cross-modal attention module. Examiner respectfully notes that Dosovitskiy is not relied upon to teach these features, as these features are taught by Chordia.
Applicant argues (pg. 10-11) in regard to claim 1 that replacing Chordia’s Resnet model with Dosovitskiy’s image-based transformer would change the principle of operation of Chordia’s implementation, and that Chordia teaches away from making such a modification. Examiner respectfully disagrees. Examiner notes that substitution of a Resnet model for a Vision Transformer (ViT) amounts to a simple substitution of known alternatives for image embedding – both models perform the same role within the image classification task, and thus the principle of operation is not changed. Examiner additionally notes that “the prior art’s mere disclosure of more than one alternative does not constitute a teaching away from any of these alternatives because such disclosure does not criticize, discredit, or otherwise discourage the solution claimed….” MPEP 2141.02 (VI); In re Fulton, 391 F.3d 1195, 1201 (Fed. Cir. 2004).
Applicant argues (pg. 11) in regard to claims 3 and 7 that one skilled in the art would not have looked to Taniguchi and Gao to remedy the deficiencies of Chordia and Dosovitskiy because Taniguchi classifies a binary result, rather than categorizing items into many categories, and Gao is only concerned with products in the fashion category. Examiner notes that for obviousness, “when a work is available in one field of endeavor, design incentives and other market forces can prompt variations of it, either in the same field or a different one. If a person of ordinary skill can implement a predictable variation, § 103 likely bars its patentability. For the same reason, if a technique has been used to improve one device, and a person of ordinary skill in the art would recognize that it would improve similar devices in the same way, using the technique is obvious unless its actual application is beyond his or her skill.” KSR International Co. v. Teleflex Inc., 550 U.S. 398, at 417 (2007).
Applicant argues (pg. 13) regarding claim 10 that one of ordinary skill in the art would not have considered R. Li when searching for a solution for item categorization because R. Li requires category embeddings as an input for its retrieval model. Examiner respectfully notes that while R. Li may not disclose a transformer-based model for outputting item categorization, it is not relied upon to do so, and its e-commerce item retrieval model is clearly relevant to the claimed e-commerce product retrieval portal.
Applicant argues (pg. 14) in regard to claim 10 that one skilled in the art would not have looked to Zhuge to remedy the deficiencies of R. Li and Tan because Zhuge is only concerned with products in the fashion category. Examiner notes that for obviousness, “when a work is available in one field of endeavor, design incentives and other market forces can prompt variations of it, either in the same field or a different one. If a person of ordinary skill can implement a predictable variation, § 103 likely bars its patentability. For the same reason, if a technique has been used to improve one device, and a person of ordinary skill in the art would recognize that it would improve similar devices in the same way, using the technique is obvious unless its actual application is beyond his or her skill.” KSR International Co. v. Teleflex Inc., 550 U.S. 398, at 417 (2007).
Applicant argues (pg. 14) in regard to claim 10 that MY Li teaches that using a combination of a Seq2Seq model and transformer model produces slightly better results than using a transformer model alone, and thus MY Li teaches away from the claimed transformer model. Examiner notes that “the prior art’s mere disclosure of more than one alternative does not constitute a teaching away from any of these alternatives because such disclosure does not criticize, discredit, or otherwise discourage the solution claimed….” MPEP 2141.02 (VI); In re Fulton, 391 F.3d 1195, 1201 (Fed. Cir. 2004).
Applicant argues (pg. 14-15) in regard to claim 10 that the only way one of ordinary skill in the art would have combined four or more references is using impermissible hindsight bias. Examiner respectfully notes that reliance on a large number of references in a rejection does not, without more, weigh against the obviousness of the claimed invention. See In re Gorman, 933 F.2d 982, 18 USPQ2d 1885 (Fed. Cir. 1991). Further, the reasons that would have motivated one of ordinary skill in the art to combine the references are clearly explained in the rejection below, and do not rely on applicant’s disclosure.
Applicant’s arguments (pg. 16-17) regarding claim 15 mirror those regarding claim 1, and are found to be unpersuasive for the same reasons explained above.
Applicant argues (pg. 18-19) in regard to claim 21 that while Yuan teaches optimizing matching items to known categories, it does not disclose doing so by processing data with a transformer processor. Examiner respectfully notes that Yuan is not relied upon to teach this limitation. As can be seen in the rejection below, processing data with a transformer processor is taught by Tan, and thus the combination of Tan and Yuan teaches the claimed processing data with a transformer processor to characterize values that optimize matching items to known categories.
The prior art rejections have been updated to include the amended limitations and to clarify the reasoning given for the limitations that were not amended.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier.
Such claim limitation(s) is/are:
“a cross-modal attention module for combining…” in claim 1.
“a cross-modal attention module for combining…” in claim 10.
“a cross-modal attention module, wherein the cross modal attention module receives…” in claim 21.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112(a)
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 1, 4-13, 21-22, and 24-26 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
Specifically, the claim limitations identified above invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, while the specification mentions the implementation of the claimed cross-modal attention module (e.g. 0036), it is silent as to the structure of this module. According to MPEP 2181(II)(B), "When a claim containing a computer-implemented 35 U.S.C. 112(f) claim limitation is found to be indefinite under 35 U.S.C. 112(b) for failure to disclose sufficient corresponding structure (e.g., the computer and the algorithm) in the specification that performs the entire claimed function, it will also lack written description under 35 U.S.C. 112(a)."
Claim Rejections - 35 USC § 112(b)
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1, 4-13, 21-22, and 24-26 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The claim limitations identified above invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. Specifically, while the specification mentions the implementation of the claimed cross-modal attention module (e.g. 0036), it is silent as to the structure of this module. Therefore, independent claims 1, 10, and 21 and their dependent claims 4-9, 11-13, 22, and 24-26 are indefinite and are rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph.
For examination purposes, the recited ‘module’ will be interpreted as a computer-implemented functional block of code for performing the associated steps.
Applicant may:
(a) Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph;
(b) Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
(c) Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)).
If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either:
(a) Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
(b) Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The following are the references relied upon in the rejections below:
Chordia et al. “Large Scale Multimodal Classification Using an Ensemble of Transformer Models and Co-Attention” (2020)
Dosovitskiy et al. “AN IMAGE IS WORTH 16X16 WORDS: TRANSFORMERS FOR IMAGE RECOGNITION AT SCALE” (2020)
Gao et al. “FashionBERT: Text and Image Matching with Adaptive Loss for Cross-modal Retrieval” (2020)
R. Li et al. “From Semantic Retrieval to Pairwise Ranking: Applying Deep Learning in E-commerce Search” (2019)
Tan et al. “LXMERT: Learning Cross-Modality Encoder Representations from Transformers” (2019)
Zhuge et al. “Kaleido-BERT: Vision-Language Pre-training on Fashion Domain” (2021)
MY Li et al. “Don’t Classify, Translate: Multi-Level E-Commerce Product Categorization Via Machine Translation” (2018)
Tagliabue et al. “How to Grow a (Product) Tree Personalized Category Suggestions for eCommerce Type-Ahead” (2020)
Ahmadvand et al. “JointMap: Joint Query Intent Understanding For Modeling Intent Hierarchies in E-commerce Search” (2020)
Yuan et al. “eProduct: A Million-Scale Visual Search Benchmark to Address Product Recognition Challenges” (2021)
Bi et al “A Multimodal Late Fusion Model for E-Commerce Product” (2020)
Chen et al. “UNITER: UNiversal Image-TExt Representation Learning” (2020)
Zhu et al. “Multimodal Joint Attribute Prediction and Value Extraction for E-commerce Product” (2020)
Claims 1, 4, 5, 6, 8, and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Chordia in view of Dosovitskiy.
Claim 1:
Regarding claim 1, Chordia discloses: A system directed to item categorization comprising:
Chordia, pg. 1, Column 1, Section 1, Paragraph 2 “…In this paper we focus on the first task, namely the classification of large-scale multimodal (text and image) product data into product type codes in the catalog of Rakuten France.”
digital data storage with data input and output capable of outputting a set of item data that includes image data and text data associated individually with at least one item;
Chordia, pg. 3, Column 1, Section 4.1 “The data for the task contains 99K products, of which close to 84K items were in the training dataset. Each product in the listing is associated with an image, a French title and an optional description…”
a program controlled digital processor connected to and in communication with said digital data storage wherein said programmed processor is capable of item categorization of said stored data, said programmed processor including:
Chordia, pg. 1, Column 2, Section 3.1, Paragraph 1 “For the baseline, we use the deep features extracted from pretrained image and text models, and explore several techniques to fuse the features from these two modalities. We then learn a small network with two fully connected layers followed by a softmax classifier to predict the product type.”
Chordia, pg. 3, Column 2, Section 4.3, Paragraph 1 “All our models were implemented in Pytorch and trained on multiple NVIDIA GPUs…”
Discloses a program controlled digital processor (GPUs running Pytorch) capable of item categorization of the stored data (Rakuten France catalog).
a text-based uni-model transformer processor for identifying features of one or more items based on said text data and generating a text transformer processor output;
Chordia, pg. 2, Column 1, Section 3.2 “We encode text using a BERT-based transformer model...”
Chordia pg. 1, Column 2, Section 3.1, Paragraph 1 “In detail, we encode text for every product through CamemBERT which outputs embeddings for every token in the text across a 12 layered hidden state…”
[an image-based uni-model transformer processor for identifying features of one or more items based on said stored image data and generating an image transformer processor output; and]
a fusion processor including a cross-modal attention module for combining the text transformer output and the image transformer output to form a multi-modal representation and a multi-layer perception head to receive the multi-modal representation and generate an item classification prediction.
Chordia, pg. 1, Column 2, Section 3.1, Paragraph 1 “For the baseline, we use the deep features extracted from pretrained image and text models, and explore several techniques to fuse the features from these two modalities. We then learn a small network with two fully connected layers followed by a softmax classifier to predict the product type.”
Chordia, pg. 1, Section 1 “We combine language and visual representations using a modified version of the co-attention architecture proposed in Lu et al.”
Chordia does not appear to explicitly teach an image-based uni-model transformer processor for identifying features of one or more items based on said stored image data and generating an image transformer processor output;
However, Dosovitskiy teaches an image-based uni-model transformer processor for identifying features of one or more items based on said stored image data and generating an image transformer processor output;
Dosovitskiy, pg. 3, Figure 1 “Figure 1: Model overview. We split an image into fixed-size patches, linearly embed each of them, add position embeddings, and feed the resulting sequence of vectors to a standard Transformer encoder.”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the ResNet image backbone network taught by Chordia with the Vision Transformer (ViT) taught by Dosovitskiy to teach an image-based transformer processor because ViT is a well-known alternative to ResNets that provides equal or better image classification accuracy (See Dosovitskiy, pg. 1, Abstract “…We show that this reliance on CNNs is not necessary and a pure transformer applied directly to sequences of image patches can perform very well on image classification tasks…Vision Transformer (ViT) attains excellent results compared to state-of-the-art convolutional networks while requiring substantially fewer computational resources to train.”).
Claim 4:
Regarding claim 4, Chordia discloses: The system of claim 1 wherein tokens are concatenated to create the multi-modal representation to be received by the multi-layer perception head.
Chordia, pg. 2, Column 1, Paragraph 2 “To combine the text and images features, we implement the following functions (/): • Concatenation - Concatenate the image and text embedding features into a single dimension vector.”
Chordia, pg. 1, Column 2, Section 3.1, Paragraph 1 “For the baseline, we use the deep features extracted from pretrained image and text models, and explore several techniques to fuse the features from these two modalities. We then learn a small network with two fully connected layers followed by a softmax classifier to predict the product type.”
Claim 5:
Regarding claim 5, Chordia discloses: The system of claim 1 wherein said text-based transformer applies a BERT model.
Chordia, pg. 1, Column 2, Paragraph 2 “…Many variations of BERT have produced advances in language processing tasks. FlauBERT [10] and CamemBERT [12] are example of variations of BERT trained on large French corpora.”
Chordia, pg. 1, Column 2, Section 3.1, Paragraph 2 “In detail, we encode text for every product through CamemBERT which outputs embeddings for every token in the text across a 12 layered hidden state.”
Claim 6:
Regarding claim 6, Chordia and Dosovitskiy discloses: The system of claim 1 wherein said image-based transformer applies a ViT model.
Dosovitskiy, pg. 1, Abstract “…When pre-trained on large amounts of data and transferred to multiple mid-sized or small image recognition benchmarks (ImageNet, CIFAR-100, VTAB, etc.), Vision Transformer (ViT) attains excellent results compared to state-of-the-art convolutional networks while requiring substantially fewer computational resources to train.”
In combination, Chordia and Dosovitskiy disclose the system of claim 1 wherein said image-based transformer applies a ViT model.
Claim 8:
Regarding claim 8, Dosovitskiy discloses: The system of claim 6 wherein said image-based model is fine tuned for ViT L-16.
Dosovitskiy, pg. 5, Training & Fine-tuning. “…For ImageNet results in Table 2, we fine-tuned at higher resolution: 512 for ViT-L/16 and 518 for ViT-H/14, and also used Polyak & Juditsky (1992) averaging with a factor of 0.9999 (Ramachandran et al., 2019; Wang et al., 2020b).”
Claim 9:
Regarding claim 9, Chordia discloses: The system of claim 1 wherein one or more models are pre-trained on a pre-set dataset and implemented by a GPU for training.
Chordia, pg. 3, Column 2, Section 4.3, Paragraph 1 “All our models were implemented in Pytorch and trained on multiple NVIDIA GPUs. For the proposed approach, we used a ResNet152 [7] or ResNext-101 [20] pre-trained on ImageNet as the backbone network...”
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Chordia and Dosovitskiy in view of Gao.
Claim 7:
Regarding claim 7, the combination of Chordia/Dosovitskiy discloses: The system of claim 1 wherein said text-based model is fine tuned [for product title data].
Chordia, pg. 3, Column 1, Section 4.2, Paragraph 1 “To prepare the data for modeling, a first step was text preprocessing, which involved cleaning and processing the textual information…As a second step, we appended the product description to the title to create a single text corpus that serves as input to our deep learning model.”
Chordia, pg. “In Figure 3 we plotted the distribution of text length across products. The sequence length describing the product varies significantly with maximum length close to 2000 words. This was relevant to consider, as i) experimenting with various sequence length within the transformer models can help fine tune the model,”
Discloses the text-based model is fine-tuned on data (a combination of product description and title data).
The combination of Chordia/Dosovitskiy does not appear to explicitly teach wherein said text-based model is fine tuned specifically for product title data alone.
However, Gao teaches wherein said text-based model is fine tuned specifically for product title data alone.
Gao, pg. 1, Abstract “…we propose FashionBERT, which leverages patches as image features. With the pre-trained BERT model as the backbone network, FashionBERT learns high level representations of texts and images…”
Gao, pg. 7, Column 2, Paragraph 1 “<Text, Image> pairs are collected from fashion products in our Alibaba.com8 website, where the titles of products act as the text information…The fine-tune dataset of cross-modal retrieval is extracted from logs of the search engine. From search logs, the queries and their clicked products first compose of the click dataset. In consequence, <Query, Title, Image> triples are chosen as the fine-tune dataset, where “Title” and “Image” are from the
same clicked products…”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to fine-tune the BERT backbone taught by Chordia/Dosovitskiy with the product title data taught by Gao, because it is routine to fine tune on domain-specific data and then merge their embeddings (leads to higher precision and recall) and using concise product titles only, instead of a concatenation of product titles and descriptions, would yield a consistent and noise resistant model that is more robust to missing/varied description information and faster to train.
Claims 10 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over R. Li in view of Tan, Zhuge, and MY Li.
Claim 10:
Regarding claim 10, R. Li discloses: A data processing system for implementing an e-commerce portal offering goods online for purchase, said system comprising:
R. Li, pg. 1, Abstract “We introduce deep learning models to the two most important stages in product search at JD.com, one of the largest e-commerce platforms in the world…”
a search engine for receiving inquiries from users seeking information regarding products for online purchase;
R. Li, pg. 1, Abstract “We introduce deep learning models to the two most important stages in product search at JD.com, one of the largest e-commerce platforms in the world. Specifically, we outline the design of a deep learning system that retrieves semantically relevant items to a query within milliseconds, and a pairwise deep re-ranking system, which learns subtle user preferences...”
a storage connected to said portal for storing retrieval data regarding one or more products responsive to said user search request;
R. Li, pg. 1, Column 1, Section 1, Paragraph 1 “…Candidate Retrieval, which uses inverted indexes to efficiently retrieve candidates based on term matching...”
Discloses the storage and look up of retrieval data.
R. Li does not appear to explicitly teach a transformer processor for categorization of products based on image and text data associated with said products, wherein said transformer processor implements categorization using: a text based uni-model transformer processor to generate a text transformer processor output; an image based uni-model transformer processor to generate an image transformer processor output; and
a fusion processor including a cross-modal attention module for combining the text transformer output and the image transformer output to form a multi-modal representation and a multi-layer perception head to receive the multi-modal representation and output a single recommendation for a given product regarding its categorization;
However, Tan teaches a transformer processor [for categorization of products] based on image and text data [associated with said products], wherein said transformer processor implements categorization using: a text based uni-model transformer processor to generate a text transformer processor output; an image based uni-model transformer processor to generate an image transformer processor output; and
Tan, pg. 1, Abstract “…In LXMERT, we build a large-scale Transformer model that consists of three encoders: an object relationship encoder, a language encoder, and a cross-modality encoder....”
Tan, pg. 3, Column 2, Single-Modality Encoders “After the embedding layers, we first apply two transformer encoders (Vaswani et al., 2017), i.e., a language encoder and an object-relationship encoder, and each of them only focuses on a single modality (i.e., language or vision)…”
Discloses a processor that implements categorization using a text based transformer processor and an image based transformer processor with categorization clues generated by each.
a fusion processor including a cross-modal attention module for combining the text transformer output and the image transformer output to form a multi-modal representation and a multi-layer perception head to receive the multi-modal representation and [output a single recommendation for a given product regarding its categorization];
Tan, pg. 3-4, Section 2.2, Cross-Modality Encoder “The cross-attention sub-layer is used to exchange the information and align the entities between the two modalities in order to learn joint cross modality representations… Lastly, the
k
-th layer output
{
h
i
k
}
and
{
v
j
k
}
are produced by feed-forward sub-layers (‘FF’) on top of
{
h
^
i
k
}
and
{
v
^
j
k
}
.”
Discloses a cross-modal attention module for combining text and image representations, followed by a multi-layer perceptron.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the transformer processor taught by Tan with the data processing system implementing an e-commerce portal because the dual-transformer processor would architecture would provide a boost in product/item categorization for e-commerce and for more accurate classification (using information from both text and images).
The combination of R. Li/Tan does not appear to explicitly teach a transformer processor for categorization of products which output[s] a single recommendation for a given product regarding its categorization.
However, Zhuge teaches a transformer processor for categorization of products.
Zhuge, pg. 1, Column 1-2, Abstract “ Kaleido-BERT is conceptually simple and easy to extend to the existing BERT framework, it attains state-of-the-art results by large margins on four downstream tasks, including text retrieval (R@1: 4.03% absolute improvement), image retrieval (R@1: 7.13% abs imv.), category recognition (ACC: 3.28% abs imv.), and fashion captioning (Bleu4: 1.2 abs imv.)…”
output a single recommendation for a given product regarding its categorization
Zhuge, pg. 7, Section 4.2.3 “We consider a classification task that judges the category and sub-category of a product, such as {HOODIES, SWEATERS}, {TROUSERS, PANTS}. We directly use a FC layer after [CLS] for these tasks.”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to provide the disclosure Zhuge to explicitly teach a transformer processor configured to categorize products.
The combination of R. Li/Tan/Zhuge does not appear to explicitly teach a taxonomy data set comprising a categorization for products determined by said transformer processor, used for configuring a response to said search request to reflect product categorization.
However, MY Li teaches a taxonomy data set comprising a categorization for products determined by said transformer processor, used for configuring a response to said search request to reflect product categorization.
MY Li, pg. 1, Abstract “…we translate a product’s natural language description into a sequence of tokens representing a root-to-leaf path in a product taxonomy…”
MY Li, pg. 1, Abstract “…In addition, we demonstrate that our machine translation models can propose meaningful new paths between previously unconnected nodes in a taxonomy tree, thereby transforming the taxonomy into a directed acyclic graph (DAG)…”
So, the root-to-leaf category path is generated for each product, then they are merged into the existing taxonomy directed acrylic graph, thereby storing every product’s classification in the digital taxonomy dataset.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the taxonomy write-back taught by MY Li with the e-commerce search portal system taught by R. Li/Tan because it would provide the search engine an up-to-date category directed acyclic graph that can handle products efficiently/accurately (See MY Li., pg. 1, Abstract “…We discuss how the resultant taxonomy directed acyclic graph promotes user-friendly navigation, and how it is more adaptable to new products.”).
Claim 12:
Regarding claim 12, Zhuge discloses: The system of claim 10 wherein said transformer processors are trained to facilitate model accuracy in proper categorization of selected products.
Zhuge, pg. 6, Column 2, Section 4 “We evaluate our Kaleido-BERT on four VL tasks by transferring the pre-trained model to each target task and fine-tuning through end-to-end training.”
Zhuge, pg. 7, Column 1, Section 4.2, Subsection 3. Category/SubCategory Recognition (CR & SUB) “The category is a vital attribute for describing a product, and is especially useful in many real-life applications. We consider a classification task that judges the category and subcategory of a product, such as {HOODIES, SWEATERS}, {TROUSERS, PANTS}. We directly use a FC layer after [CLS] for these tasks.”
The model is trained on product category labels to maximize categorization accuracy.
Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over R. Li, Tan, Zhuge, and MY Li in view of Dosovitskiy.
Claim 11:
Regarding claim 11, Tan discloses: The system of claim 10 wherein said text based uni-model transformer processor implements a BERT text transformer model
Tan, pg. 8, Column 1, Footnote 10 “Since our language encoder is same as BERTBASE, except the number of layers (i.e., LXMERT has 9 layers and BERT has 12 layers), we load the top 9 BERT-layer parameters into the LXMERT language encoder.”
The combination of R. Li/Tan/Zhuge/MY Li does not appear to explicitly teach that said image based uni-model transformer processor implements a ViT image transformer model.
However, Dosovitskiy teaches said image based uni-model transformer processor implements a [ViT] image transformer model.
Dosovitskiy, pg. 1, Abstract “…When pre-trained on large amounts of data and transferred to multiple mid-sized or small image recognition benchmarks (ImageNet, CIFAR-100, VTAB, etc.), Vision Transformer (ViT) attains excellent results compared to state-of-the-art convolutional networks while requiring substantially fewer computational resources to train…”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the language-and-vision architecture/pipeline taught by R. Li/Tan/Zhuge/MY Li with the vision transformer taught by Dosovitskiy to explicitly teach ViT model because it is known to outperform state-of-the-art CNNs (which Tan’s LXMERT is based upon) (See Dosovitskiy, pg. 1, Abstract “…We show that this reliance on CNNs is not necessary and a pure transformer applied directly to sequences of image patches can perform very well on image classification tasks…Vision Transformer (ViT) attains excellent results compared to state-of-the-art convolutional networks while requiring substantially fewer computational resources to train.”).
Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over R. Li, Tan, Zhuge, and MY Li in view of Tagliabue.
Claim 13:
Regarding claim 13, the combination of R. Li/Tan/Zhuge/MY Li does not appear to explicitly disclose: The system of claim 10 wherein a grouping of products within a single category predicted by the transformer processor is provided in response to a user search request.
However, Tagliabue discloses wherein a grouping of products within a single category predicted by the transformer processor is provided in response to a user search request.
Tagliabue, pg. 1, Column 2, Paragraph 1 “…narrowing down candidate products explicitly by matching the selected categories, shops are able to present less noisy result pages and increase the perceived relevance of their search engine…In this work we present SessionPath, a scalable and personalized model to solve facet prediction for type-ahead suggestions: given a shopping session and candidate queries in the suggestion dropdown menu, the model is asked to predict the best category facet to help users narrow down search intent...”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine taxonomy write-back pipeline with the system taught by Li/Tan/Zhuge/MY Li because would enhance search results by providing the most relevant set of products based on the prediction (See Tagliabue, pg. 1, Column 2, Paragraph 1 “…narrowing down candidate products explicitly by matching the selected categories, shops are able to present less noisy result pages and increase the perceived relevance of their search engine…”).
Claims 15 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Chordia in view of Tan, and MY Li.
Claim 15:
Regarding claim 15, Chordia discloses: A data processing method to classify a large diverse data set of individual items many of which are associated with image and text data corresponding to product type and class, the method comprising:
Chordia, pg. 1, Column 1, Section 1, Paragraph 2 “…In this paper we focus on the first task, namely the classification of large-scale multimodal (text and image) product data into product type codes in the catalog of Rakuten France.”
inputting text data for a product into a first uni-model transformer processor to ascertain clues regarding what class the product fits;
Chordia, pg. 1, Column 2, Section 3.1, Paragraph 2 “In detail, we encode text for every product through CamemBERT which outputs embeddings for every token in the text across a 12 layered hidden state…”
Discloses inputting text data for a product into a first transformer (CamemBERT).
[inputting image data for said product into a second uni-model transformer processor to ascertain clues regarding what class the product fits;]
aggregating clues from said first [and second] uni-model transformer processor[s] into a final prediction regarding a class for that product, wherein said aggregating step includes a cross attention fusion process to form a multi-modal representation and a multi-layer perception head to receive the multi-modal representation and generate the final prediction; and
Chordia, pg. 2, Column 2, Paragraph 2 “Next, the image and text representations are passed to the co-attention block. We adopt the architecture and method of Lu et al. [11] with one modification. In their work, a hierarchy of attention elements is created between spatial maps and text at 3 levels: words, phrases and sentences. In contrast, our method jointly reasons between words and images…”
Chordia, pg. 3, Column 1, Paragraph 2 “The image and text vectors are concatenated and passed to a neural network with a softmax activation to obtain a vector of class probability scores….”
[outputting said final prediction in association with said product into a taxonomy data set stored for digital access.]
Chordia does not appear to explicitly teach inputting image data for said product into a second uni-model transformer processor to ascertain clues regarding what class the product fits / a second transformer processor (for image data)
However, Tan teaches inputting image data for said product into a second uni-model transformer processor to ascertain clues regarding what class the product fits / a second transformer processor (for image data)
Tan, pg. 2, Column 2, Section 2.1, Object-level Image Embeddings “Instead of using the feature map output by a convolutional neural network, we follow Anderson et al. (2018) in taking the features of detected objects as the embeddings of images. Specifically, the object detector detects m objects {o1, . . . , om} from the image (denoted by bounding boxes on the image in Fig. 1). Each object oj is represented by its position feature (i.e., bounding box coordinates) pj and its 2048-dimensional region-of-interest (RoI) feature fj…”
Tan, pg. 3, Column 2, Single-Modality Encoders “After the embedding layers, we first apply two transformer encoders (Vaswani et al., 2017), i.e., a language encoder and an object-relationship encoder, and each of them only focuses on a single modality (i.e., language or vision)…”
Discloses inputting image data into a second transformer (object-relationship encoder.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the separate modality specific encoders taught by Tan into Chordia’s pipeline to explicitly teach a second transformer/two separate transformers which would make the system capable of making classifications on both text and image data (providing the rich information of images, which would lead to better classification).
The combination of Chordia/Tan does not explicitly teach outputting said final prediction in association with said product into a taxonomy data set stored for digital access.
However, MY Li teaches outputting said final prediction in association with said product into a taxonomy data set stored for digital access.
MY Li, pg. 1, Abstract “E-commerce platforms categorize their products into a multi-level taxonomy tree with thousands of leaf categories...In this article, we propose a new paradigm based on machine translation. In our approach, we translate a product’s natural language description into a sequence of tokens representing a root-to-leaf path in a product taxonomy…”
MY Li, pg. 1, Abstract “…In addition, we demonstrate that our machine translation models can propose meaningful new paths between previously unconnected nodes in a taxonomy tree, thereby transforming the taxonomy into a directed acyclic graph…”
So, the root-to-leaf category path is generated for each product, then they are merged into the existing taxonomy directed acrylic graph, thereby storing every product’s classification in the digital taxonomy dataset.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to generate the final class predictions generated by the pipeline taught by Chordia/Tan and insert it as a root-leaf-path in a stored taxonomy as taught by MY Li because it would provide platforms with category trees that would be useful for downstream tasks (See MY Li., pg. 1, Abstract “…We discuss how the resultant taxonomy directed acyclic graph promotes user-friendly navigation, and how it is more adaptable to new products.”).
Claim 18:
Regarding claim 18, Chordia discloses: The method of claim 15 wherein the transformer processors are trained against a data set of products having known classifications.
Chordia, pg. 3, Column 1, Section 4.1 “The data for the task contains 99K products, of which close to 84K items were in the training dataset. Each product in the listing is associated with an image, a French title and an optional description…”
Chordia, pg. 3, Column 1, Section 4.1, Paragraph 2 “The goal of the task is to predict the product type code (PRC). These codes are numbers associated with a general product name.”
The product type code is the known classification.
Claims 16 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Chordia, Tan, and MY Li in view of Dosovitskiy.
Claim 16:
Regarding claim 16, Chordia and Tan disclose: The method of claim 15 wherein the first uni-model transformer processor uses BERT processing [and the second uni-model transformer processor uses ViT processing.]
Chordia, pg. 1, Column 2, Section 3.1, Paragraph 2 “In detail, we encode text for every product through CamemBERT which outputs embeddings for every token in the text across a 12 layered hidden state…”
Discloses BERT processing.
Tan, pg. 3, Column 2, Single-Modality Encoders “After the embedding layers, we first apply two transformer encoders (Vaswani et al., 2017), i.e., a language encoder and an object-relationship encoder, and each of them only focuses on a single modality (i.e., language or vision)…”
Discloses the image transformer processor.
The combination of Chordia/Tan/ MY Li does not explicitly teach that the second uni-model transformer processor uses ViT processing.
However, Dosovitskiy teaches the second uni-model transformer processor uses ViT processing.
Dosovitskiy, pg. 1, Abstract “…When pre-trained on large amounts of data and transferred to multiple mid-sized or small image recognition benchmarks (ImageNet, CIFAR-100, VTAB, etc.), Vision Transformer (ViT) attains excellent results compared to state-of-the-art convolutional networks while requiring substantially fewer computational resources to train…”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the ResNet image backbone network taught by Chordia with the Vision Transformer (ViT) taught by Dosovitskiy to teach an image-based transformer processor because ViT is a well-known alternative to ResNets that provides equal or better image classification accuracy.
Claim 17:
Regarding claim 17, Chordia discloses: The method of claim 16 wherein the transformer processors are encoded for operation on GPU based processors.
Chordia, pg. 3, Column 2, Section 4.3, Paragraph 1 “All our models were implemented in Pytorch and trained on multiple NVIDIA GPUs…”
Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Chordia, Tan, and MY Li in view of Ahmadvand.
Claim 19:
Regarding claim 19, the combination of Chordia/Tan/ MY Li does not appear to explicitly disclose: The method of claim 18 wherein the taxonomy set is used to facilitate responses to user queries made online to an ecommerce portal.
However, Ahmadvand teaches wherein the taxonomy set is used to facilitate responses to user queries made online to an ecommerce portal.
Ahmadvand, pg. 1, Abstract “…In this paper, we introduce Joint Query Intent Understanding (JointMap), a deep learning model to simultaneously learn two different high-level user intent tasks: 1) identifying a query’s commercial vs. non-commercial intent, and 2) associating a set of relevant product categories in taxonomy to a product query...”
Discloses a taxonomy being used to facilitate responses to queries.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to feed the taxonomy set taught by Chordia/Tan/MY Li to deep learning model (JointMap) taught by Ahmadvand because feeding up-to-date taxonomies to its query-understanding pipeline will directly boost its downstream search functions (See Ahmadvand, pg. 1, Abstract “An accurate understanding of a user’s query intent can help improve the performance of downstream tasks such as query scoping and ranking…”).
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Chordia, Tan, MY Li, and Ahmadvand in view of Yuan.
Claim 20:
Regarding claim 20, the combination of Chordia/Tan/ MY Li/Ahmadvand does not appear to explicitly disclose: The method of claim 19 wherein the number of products having text and image data that are processed exceeds one million which are classified in the taxonomy set into at least four categories.
However, Yuan teaches wherein the number of products having text and image data that are processed exceeds one million which are classified in the taxonomy set into at least four categories.
Yuan, pg. 1, Abstract “…We present eProduct as a training set and an evaluation set, where the training set contains 1.3M+ listing images with titles and hierarchical category labels, for model development, and the evaluation set includes 10,000 query and 1.1 million index images for visual search evaluation…”
Products are assigned across multiple levels of classification which necessarily involves at least four distinct categories for 1.3M+ pairings.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to employ the dataset taught by Yuan alongside the pipeline taught by Chordia/Tan/MY Li/Ahmadvand to specifically teach the number of products exceeding one million, which the pipeline would be capable of processing.
Claim 21 is rejected under 35 U.S.C. 103 as being unpatentable over Tan in view of Yuan and Bi.
Claim 21:
Regarding claim 21, Tan discloses: A computer implemented method of training a computerized classification system, comprising:
Tan, pg. 1, Abstract “…Next, to endow our model with the capability of connecting vision and language semantics, we pre-train the model with large amounts of image-and-sentence pairs, via five diverse representative pre-training tasks: masked language modeling, masked object prediction (feature regression and label classification), cross-modality matching, and image question answering…”
a. a first computer memory for storing a pre-determined set of training data comprising text data [associated with items within a known category];
b. a second computer memory for storing a pre-determined set of training data comprising image data [associated with items within the known category];
Tan, pg. 3, Column 2, Single-Modality Encoders “After the embedding layers, we first apply two transformer encoders (Vaswani et al., 2017), i.e., a language encoder and an object-relationship encoder, and each of them only focuses on a single modality (i.e., language or vision)…”
Tan, pg. 1, Abstract “…we pre-train the model with large amounts of image-and-sentence pairs, via five diverse representative pre-training tasks: masked language modeling, masked object prediction…”
The language encoder serves as a first computer memory and object-relationship encoder the second (both are preloaded via large-scale pre-training datasets) because they each have their own distinct parameters, implicitly acting as two separate memories.
c. processing said text data in a uni-modal text based transformer processor to characterize values [that optimize matching items to known categories];
Tan, pg. 1, Abstract “…In LXMERT, we build a large-scale Transformer model that consists of three encoders: an object relationship encoder, a language encoder, and a cross-modality encoder…”
The language encoder processes text data.
d. processing said image data in an image based transformer model to characterize values within the model [that optimize matching items to known categories]; and
Tan, pg. 1, Abstract “…In LXMERT, we build a large-scale Transformer model that consists of three encoders: an object relationship encoder, a language encoder, and a cross-modality encoder...”
The object relationship encoder processes image data.
e. [storing said characterized model values for use against data that has not been classified; and]
f. generating a classification prediction using a multi-layer perception head including a cross-modal attention module, wherein the cross modal attention module receives the values characterized by the text based uni-model transformer process and the image based uni-model transformer to form a multi-modal representation, wherein the multi-layer perception head receives the multi-modal representation and generates the classification prediction.
Tan, pg. 3-4, Section 2.2, Cross-Modality Encoder “The cross-attention sub-layer is used to exchange the information and align the entities between the two modalities in order to learn joint cross modality representations… Lastly, the
k
-th layer output
{
h
i
k
}
and
{
v
j
k
}
are produced by feed-forward sub-layers (‘FF’) on top of
{
h
^
i
k
}
and
{
v
^
j
k
}
.”
Discloses a cross-modal attention module for combining text and image representations, followed by a multi-layer perceptron for generating predictions.
Tan does not appear to explicitly teach data associated with items within a known category / optimizing matching items to known categories
However, Yuan teaches data associated with items within a known category / optimizing matching items to known categories
Yuan, pg. 1, Abstract “…We present eProduct as a training set and an evaluation set, where the training set contains 1.3M+ listing images with titles and hierarchical category labels, for model development, and the evaluation set includes 10,000 query and 1.1 million index images for visual search evaluation…”
Text and image data is tagged with a known category and is used to train (optimize) the model to match items to their categories.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to feed the product text/image data taught by Yuan to the transformer model taught by Tan because Tan’s transformer was built to be fine-tuned on aligned text and image data for classification.
The combination of Tan/Yuan does not appear to explicitly teach storing said characterized model values for use against data that has not been classified.
However, Bi teaches storing said characterized model values for use against data that has not been classified.
Bi, pg. 3, Column 2, Section 4.1 “…We trained the policy
network for 40 epochs and kept the checkpoint that has the highest
macro-f1 score on the validation set…”
By keeping the checkpoint they are saving the learned model parameters so they can be loaded later.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to store the trained model parameters taught by the pipeline of Tan/Yuan for later inference by simply checkpointing the best-performing weights during training as taught by Bi, because it is a known practice in machine learning that ensures fault tolerance.
Claim 22 is rejected under 35 U.S.C. 103 as being unpatentable over Tan, Yuan, and Bi in view of Chen.
Claim 22:
Regarding claim 22, the combination of Tan/Yuan/Bi does not appear to explicitly disclose: The method of claim 21 wherein said item classification system processes text and image data with transformer processors and an early fusion processor.
However, Chen teaches wherein said item classification system processes text and image data with transformer processors and an early fusion processor.
Chen, pg. 4, Section 3.1 “The model architecture of UNITER is illustrated in Figure 1. Given a pair of image and sentence, UNITER takes the visual regions of the image and textual tokens of the sentence as inputs. We design an Image Embedder and a Text Embedder to extract their respective embeddings. These embeddings are then fed into a multi-layer Transformer to learn a cross-modality contextualized embedding across visual regions and textual tokens…”
So, text and image features are simultaneously processed together up front in a single transformer (early fusion).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to integrate the early fusion taught by Chen with the pipeline taught by Tan/Yuan/Bi because Chen discloses clear performance advantages of processing image and text using early fusion, where the model can learn tightly coupled representations rather than merging unimodal features late (See Chen, pg. 1, Abstract “…Extensive experiments show that UNITER achieves new state of the art across six V+L tasks (over nine datasets)…”).
Claim 24 is rejected under 35 U.S.C. 103 as being unpatentable over Chordia in view of Dosovitskiy, and further in view of Tan.
Claim 24:
Regarding claim 24, the combination of Chordia/ Dosovitskiy does not appear to explicitly disclose: The system of claim 1, wherein the multi-modal representation is computed by: pairing key-values from the text transformer processor output with a query from the image transformer processor output, or pairing key-values from the image transformer processor output with a query from the text transformer processor output.
However, Tan teaches wherein the multi-modal representation is computed by: pairing key-values from the text transformer processor output with a query from the image transformer processor output, or pairing key-values from the image transformer processor output with a query from the text transformer processor output.
Tan, pg. 3-4, section 2.2, Cross-Modality Encoder “Inside the k-th layer, the bi-directional cross-attention sub-layer (‘Cross’) is first applied, which contains two uni-directional cross-attention sub-layers: one from language to vision and one from vision to language. The query and context vectors are the outputs of the (k-1)-th layer (i.e., language features
{
h
i
k
-
1
}
and vision features {
v
j
k
-
1
}):
h
^
i
k
=
C
r
o
s
s
A
t
t
L
→
R
h
i
k
-
1
,
v
1
k
-
1
,
…
,
v
m
k
-
1
v
^
j
k
=
C
r
o
s
s
A
t
t
R
→
L
v
j
k
-
1
,
h
1
k
-
1
,
…
,
h
n
k
-
1
The cross-attention sub-layer is used to exchange the information and align the entities between the two modalities in order to learn joint cross modality representations.”
Discloses computing cross modality representations via cross attention between language queries and vision key-values, and between vision queries and language key-values.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to fine-tune the BERT backbone taught by Chordia/Dosovitskiy with the cross-attention mechanism taught by Tan in order to “exchange the information and align the entities between the two modalities in order to learn joint cross modality representations” (Tan, pg. 4, section 2.2).
Claim 25 is rejected under 35 U.S.C. 103 as being unpatentable over Chordia in view of Dosovitskiy, and further in view of Zhu.
Claim 25:
Regarding claim 25, the combination of Chordia/ Dosovitskiy does not appear to explicitly disclose: The system of claim 1, wherein the cross-modal attention module comprises a visual gate to filter out visual noise.
However, Zhu teaches wherein the cross-modal attention module comprises a visual gate to filter out visual noise.
Zhu, pg. 3, section 2.4 “we need to avoid introducing noises resulted from when the image fails to represent some semantic meaning of words, such as abstract concepts. To achieve this, we design a global visual gate to filter out visual noise for any words that are irrelevant based on the visual signals.”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to fine-tune the BERT backbone taught by Chordia/Dosovitskiy with the visual noise gating taught by Zhu in order to “avoid introducing noises resulted from when the image fails to represent some semantic meaning of words, such as abstract concepts” (Zhu, pg. 3, section 2.4).
Claim 26 is rejected under 35 U.S.C. 103 as being unpatentable over R. Li in view of Tan, Zhuge, and MY Li, and further in view of Zhu.
Claim 26:
Regarding claim 26, the combination of R. Li/Tan/Zhuge/MY Li does not appear to explicitly disclose: The system of claim 10, wherein the cross-modal attention module comprises a visual gate to filter out visual noise.
However, Zhu teaches wherein the cross-modal attention module comprises a visual gate to filter out visual noise.
Zhu, pg. 3, section 2.4 “we need to avoid introducing noises resulted from when the image fails to represent some semantic meaning of words, such as abstract concepts. To achieve this, we design a global visual gate to filter out visual noise for any words that are irrelevant based on the visual signals.”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the language-and-vision architecture/pipeline taught by R. Li/Tan/Zhuge/MY Li with the visual noise gating taught by Zhu in order to “avoid introducing noises resulted from when the image fails to represent some semantic meaning of words, such as abstract concepts” (Zhu, pg. 3, section 2.4).
Claim 27 is rejected under 35 U.S.C. 103 as being unpatentable over Chordia in view of Tan, and MY Li., and further in view of Zhu
Claim 27:
Regarding claim 27, the combination of Chordia/Tan/MY Li does not appear to explicitly disclose: The method of claim 15, wherein said aggregating step further comprises filtering out visual noise with a visual gate.
However, Zhu teaches wherein said aggregating step further comprises filtering out visual noise with a visual gate.
Zhu, pg. 3, section 2.4 “we need to avoid introducing noises resulted from when the image fails to represent some semantic meaning of words, such as abstract concepts. To achieve this, we design a global visual gate to filter out visual noise for any words that are irrelevant based on the visual signals.”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to fine-tune the BERT backbone taught by Chordia/Tan/MY Li with the visual noise gating taught by Zhu in order to “avoid introducing noises resulted from when the image fails to represent some semantic meaning of words, such as abstract concepts” (Zhu, pg. 3, section 2.4).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BENJAMIN M ROHD whose telephone number is (571)272-6445. The examiner can normally be reached Mon-Thurs 8:00-6:00 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/B.M.R./Examiner, Art Unit 2147
/ERIC NILSSON/Primary Examiner, Art Unit 2151